The short answer
- An owned eight-GPU B200 server costs about $16,600 a month over three years. It breaks even against a premium closed model at under one billion tokens a month, but against cheap hosted open models only above twenty billion.
- Staff, power contracts and idle capacity push the break-even point further out.
- For most organisations the case for self-hosting rests on control over data. Run the numbers for your own volume below.
How we calculate
API cost is monthly tokens times a blended price per million tokens. We blend input and output prices at a ratio of three to one, which is typical for document and search workloads.
Self-hosting cost is the monthly cost of the server plus the people who run it. For an owned server we use a vendor estimate of $22.70 per server-hour for an eight-GPU B200 system at 70 percent utilisation over three years,1 which is about $16,600 per month. Rental figures use list prices at Verda, a Finnish provider.2
Break-even volume is the monthly self-hosting cost divided by the API price per token: the volume at which both cost the same.
Break-even by API
The cheaper the API, the further away the break-even point
Against a premium closed model, an owned server pays off below one billion tokens a month. Against a cheap hosted open model, it takes more than twenty billion.
Break-even volume ranges from under 1 to over 22 billion tokens a month.
Monthly volume at which an owned B200 server matches the API bill, billion tokens
1 Owned 8-GPU B200 server at about $16,600 per month (3 years, 70% utilisation), divided by blended API price per million tokens (3 input : 1 output). Excludes staff costs.
Source: Lindstead calculation; server cost from Mercatus (vendor estimate, July 2026); API list prices, 27 September 2026
Lindstead
The comparison that matters most is quality-matched. If an open model on your own server does the job of a premium closed model, the break-even point is low. If a cheap hosted open model through an API would do the same job, it is very high. Independent research from Carnegie Mellon reaches a similar conclusion: payback in months for small models, around five years for the largest.3
Calculator
Enter your own volume and prices. All amounts in the same currency.
Result
What the numbers leave out
- Quality. A cheaper model that needs more human review is not cheaper.
- Utilisation. A server that idles at night still costs money; APIs do not.
- Price trends. API prices for a given capability have fallen steeply every year, while GPU prices rose in 2026.4
- Control. The value of keeping data out of third-party hands does not appear in a cost model, and is often the deciding factor.
Sources
- Mercatus, B200 server price and cost per hour, 16 July 2026 (vendor analysis). mercatus-ai.com/blog/b200-server-price
- Verda GPU pricing, accessed 27 September 2026. verda.com/pricing
- Carnegie Mellon University, cost-benefit analysis of on-premise LLM deployment, arXiv 2509.18101. arxiv.org/abs/2509.18101
- Epoch AI, LLM inference price trends, 12 March 2025. epoch.ai/data-insights/llm-inference-price-trends
- API list prices: Anthropic, OpenAI, DeepSeek and Mistral pricing pages, accessed 27 September 2026.