Tokens, as a service.
An AI foundry at hyperscale, on our own GPUs: open models served and trained, closed models bridged where required.
# any OpenAI-compatible SDK works
curl https://api.mosaicfoundry.eryce.com/v1/chat/completions \
-H "Authorization: Bearer $MOSAIC_KEY" \
-d '{
"model": "llama-4-maverick",
"messages": [{"role": "user",
"content": "Hello, Mosaic."}]
}'Serve, train and route. One foundry.
Most inference platforms stop at serving. Mosaic Foundry takes models in, tunes them on your data, and sends them out as tokens at whatever scale your product demands. And it keeps growing: through 2026 we are scaling our GPU fleet to a capacity of more than twenty trillion tokens a day, a quarter of a billion every second.
Inference
Production endpoints for open-source models, autoscaled on our GPU fleet. Pay per token, or reserve dedicated capacity with guaranteed throughput.
Training and fine-tuning
Fine-tune open models on your data, on dedicated clusters. The result deploys straight to an inference endpoint, in the same foundry.
The closed-model bridge
Where a task needs a frontier closed model, Mosaic Foundry routes to it through the same API and the same bill. One key, every model.
Open-source families in the foundry, and the closed models behind the bridge
Bridged where required: GPT, Claude and Gemini through the same API and the same invoice.
From your SDK to the first token in minutes.
Point
Swap the base URL in your existing OpenAI-style SDK and drop in a Mosaic Foundry key. Your first tokens flow in minutes.
Produce
Pick your models: open-source on our GPUs, closed through the bridge. Serverless by default, dedicated when you need guarantees.
Scale
Autoscaling handles the traffic, observability shows every token, and one invoice covers the whole foundry.
Routing and pricing, done in the open.
The two things this industry keeps complicated are the two things Mosaic Foundry keeps simple: which model runs your request, and what it costs.
One request in, the right model out.
Name a model, or describe the job. The Mosaic Foundry router sends every request to the cheapest model that clears your quality and latency bar, and fails over automatically when a provider blinks.
- Route by cost, latency or quality, per request
- Automatic failover and multi-provider redundancy
- Prompt caching: repeated context billed at a fraction of the price
- Batch lane for asynchronous workloads at a deep discount
- Structured outputs and tool calling preserved across models
Pricing you can read in one screen.
Every model has one published price per million tokens, in and out. No sales call to see a number, no hidden markup on bridged models, no egress surprises at the end of the month.
- One public price card, per million tokens, for every model
- Cached and batch tokens discounted automatically
- Spend caps, budgets and alerts per project and per key
- Bridged closed models at list price plus one disclosed fee
- Cost observability down to the single request
Flexible on the way in, industrial on the way out.
Drop-in API
OpenAI-compatible. Point your existing SDK at a new base URL and keep shipping; no rewrite, no new client.
Serverless or dedicated
Start per-token with zero commitment. Move hot workloads onto dedicated GPUs with reserved throughput when it pays off.
Autoscaling
Endpoints scale with your traffic and back down when it fades. Built for agentic workloads: tool calls, long contexts and reasoning models.
Observability
Tokens, latency, errors and cost per model and per key. Logs and usage exports your finance team will actually accept.
European residency
Prompts and outputs processed in the EU and Switzerland, zero retention by default. Built for GDPR and the EU AI Act.
One bill, every model
Open models on our metal and closed models through the bridge, on a single invoice with per-project breakdowns.
Mosaic Foundry runs on the same Eryce infrastructure described on our infrastructure page: our GPU fleet, our monitoring, our on-call. The foundry is not resold capacity; it is the machine we already operate.
Going live in Q4 2026.
Register now. Early-access partners onboard with our engineers, get direct input on the roadmap and are first through the door at launch.