2026-06-02

How in-metro inference beats Bedrock on price

Hyperscaler inference pricing bundles two costs that have nothing to do with serving a token: a data-transfer tax and the margin on a closed, rented model. Volt removes both, and the lower price is what's left.

Pods run reserved, multi-vendor GPUs in Tier III metros. Committing capacity for 36 months lands an NVIDIA B200 at $2.36/GPU/hr — ~44% below CoreWeave reserved, ~78% below CoreWeave on-demand. That reserved base is the foundation of the unit economics; utilization climbs from 55% toward 88% as a pod fills, and pod-level EBITDA turns positive in year two.

Open-weights models carry no per-token licensing markup. With Llama, Mistral, Gemma, and Phi, the cost is compute, not rent — so the savings reach the price sheet instead of a model vendor.

And because the architecture is zero-egress, there are no transfer bills to surprise you at month end. Llama 70B runs at $0.95/M tokens standard and $1.45/M on the sovereign tier — with data that never leaves the metro.