Overview
LegionEdge Cloud is live for all accounts today. It's one panel for the compute side of AI work: train and fine-tune models, manage compute, and spin up inference endpoints — on our capacity or capacity you bring.
The panel runs on the same infrastructure our research runs on. The fleet that validates our open datasets and trains our models is the fleet your workloads schedule against.
Train, serve, observe
The catalog carries the open models plus the models we train ourselves, each with context lengths, quantizations, and per-token pricing in one place. Point a fine-tuning job at a dataset and a base model; promote the result to an inference endpoint when the evals clear.
Observability is built into the request path, not bolted on. Every request is inspectable — tokens, latency, tool calls — with log streams and traces at the endpoint level. If an endpoint misbehaves at 2 a.m., the panel can already show you why.
Our compute or yours
Run on our fleet — including the dedicated NVIDIA H200 capacity behind our research — with per-hour metering and hard budget caps you set yourself. Nothing scales past a limit you didn't write down.
Or bring your own: attach AWS or Google Cloud capacity and the panel schedules onto it with the same interface, the same observability, and your existing cloud bill. The translation layer comes from Foltrac, so one definition runs anywhere.
Availability
LegionEdge Cloud is open to every account today — no waitlist. Existing API usage appears in the panel automatically.
Training and fine-tuning start with the open-model families we publish distillation recipes for; more base models and more bring-your-own-compute providers follow through the year.