Self-hosted frontier inference

Your hardware. Full precision. Nothing leaves the building.

We put open frontier models on the GPUs you already own, tune the serving stack, then layer on memory, routing and caching. The token bill stops chasing your growth and flattens into a floor you control.

Hosted API
$23,130
Deep Inference
$5,245
per month at 7B tokens / month · 4.4× cheaper, and the gap widens as you scale
1×H100
runs a 120B open model
100%
of weights stay on your servers
0
bytes leave your VPC
01 Economics

The bill you’re replacing.

Drag the volume. Every number is from a sourced cost model, not a blended average. A hosted API bills every token forever; your own cluster is a floor that barely moves until you are serving millions of requests a day.

Tokens / month 7B
avg RAG context tokens
400-token prompt plus your RAG · 600-token output · 2.5× peak · warm floor and ops included · hosted at list price. Latencies are directional estimates, costs follow the sourced model.
Hosted API
GPT-5.6 Terra
$23,117
per month
673 ms first token · 5.7s full reply
benchmark 72.7
Deep Inference
GPT-OSS-120B
$5,245
per month
573 ms first token · 11.5s full reply
benchmark 49.2
You save$17,872/ mo · 4.4× cheaper
GPUs deployed
1 × H100
holds p50 latency at 2.5× peak · billed as 1.0 GPU-eq per minute
Benchmark · overall
72.7 vs 49.2
the hosted model scores higher, route only the hard calls to it
The drop · 7B tokens / month
Hosted API · GPT-5.6 Terra + RAG$23,130
Deep Inference · GPT-OSS-120B on your GPUs$5,245
The levers · yours to toggle, we tune each one
Memory trim · cuts RAG 65% Model routing · 60% to a cheaper model Prefix caching · 40% cacheable Shorter outputs · 25% shorter Semantic cache Context compression Code offload
02 The wedge

Four ways to serve a frontier model.

Same model, four places to run it. The difference is who owns the hardware, who pays per token, and who gets paged at 3am. Only one column should ever be you, and it depends entirely on whether the GPUs are already yours.

Hosted APIRented GPU cloudDIY on your racksDeep Inferenceyour servers
who owns the weightsthe vendoryou, on their metalyouyou
cost at 7B tokens / moabout $23k / morental plus markupabout $5k, you build itabout $5k, we build it
how cost scaleswith every token, foreverwith usageflat, a floorflat, a floor
precisionundisclosedyours to setshrunk to fitfull BF16
time to productionan afternoondays3 to 6 months2 weeks
who runs it at 3amtheir SREyouyouus, on the term
data leaves your networkyesmaybenono
the right call whenunder ~2B tokens / moyou need capacity nowyou have a spare infra teamthe GPUs are already yours
Where we lose

Under roughly 2 billion tokens a month, or if you do not already own the GPUs, a hosted API is cheaper than us, and we will say so on the first call. We only win when the hardware is already depreciating on your balance sheet. That is the whole thesis, and it is why the calculator above will happily tell you to stay on the API.

03 Serving

Same GPU. 1.8× the tokens.

A naive deploy leaves most of the silicon idle: bad parallelism, no continuous batching, prefill and decode fighting for the same devices. We do not sell you more hardware. We fix the serving stack on the hardware you have.

Naive deploy · vLLM defaults1,850 units/s
Deep Inference tuned · same 1×H1003,330 units/s

Method · anchored on the one verified benchmark, a 70B at FP8 on 1×H100 moving 1,850 to 2,400 units/s on vLLM, plus 8 to 16% from TensorRT-LLM. The managed stack reaches about 1.8× effective through continuous batching, paged KV cache and a warm floor. Presets for other models scale from this anchor.

1.8×
throughput on the same GPUs
$0.75
per 1M tokens, GPT-OSS-120B self-hosted
0
bytes of weights or prompts leaving your VPC
04 Ownership

You own the cluster. We own the hard part.

Standing up a 16×H200 model across two racks is a distributed-systems problem, not a pip install. That problem is what we do, pre-built, tuned, and running entirely on your servers. When the term ends you keep the stack, the runbooks and the dashboards. Renew because it is working, not because you are trapped.

01

Assess

We benchmark your model on your hardware and set a concrete throughput and latency target to quote against.

week 0 · free
02

Deploy

The distributed stack goes onto your nodes: runtime, parallelism, autoscaling, observability. Endpoints go live.

weeks 1 to 2
03

Optimize

Memory, routing, caching and kernel work until we hit the number in the contract.

weeks 2 to 6
04

Hand over

Runbooks, Grafana dashboards, IaC for the control plane and a recorded walkthrough. Then on-call for the term.

ongoing
05 Engagement

Pick the depth.

Fixed-term engagements, quoted against your model, cluster size and support window. Your GPU and cloud costs never touch our invoice, they stay on your account, where they already are.

Deploy
2-week engagement
Most engagements
Optimize
4 to 6 weeks
Operate
3, 6 or 12 months
Distributed stack on your nodes
Multi-node parallelism configured
Autoscaling, routing, observability
Memory, prefix caching, shorter outputs
Model routing across the leaderboard
Committed units/s & cost in contract
Handover docs & runbooks
Engineer on-call2 weeks90 dayswhole term
Re-tuning as models change
Runtime & security upgrades, SLAs
Get a quote Get a quote Get a quote
Start with a benchmark

Send us your model. We’ll send back a number.

A no-charge benchmark on your own hardware. Ten days from the first call to a committed units/s figure and a monthly cost you can hold us to. No GPUs to rent from us, ever.

Book the benchmark →