LLM inference cost is what an organisation pays to generate answers from a language model. Through an API it is the number of tokens times the price per token. On own GPUs it is server or rental cost plus people, power and idle capacity, divided by the tokens actually served. Volume and the choice of API decide which is cheaper.

The short answer

  1. API list prices differ by a factor of more than twenty: one billion tokens a month costs about €660 through Mistral Large 3 and about €17,540 through GPT-6 Astra.
  2. An owned eight-GPU B200 server costs about €14,560 a month whether it is busy or idle. It matches the bill of Claude Sonnet 5 at about 4.2 billion tokens a month, before staff costs.
  3. The comparison only holds if the open model on the own server does the same job as the API it replaces. That quality match, not the hardware price, moves the result most.

This guide explains the components of the bill and how to read them. The break-even calculator applies them to an organisation's own volume, and the guide to self-hosted LLMs covers the wider decision.

What is LLM inference cost?

Training a model is the one-off work of teaching it from data. Inference is every later use: each question, summary or classification the model produces. Few organisations train large models themselves; almost all of them pay for inference, every month, for as long as the application runs. That makes inference the recurring cost line that finance teams see on the invoice.

Inference is measured in tokens. A token is a fragment of text, often a word or part of a word. A request consists of input tokens (the question, instructions and any documents sent along) and output tokens (the answer the model writes). Document search and summarisation send far more input than they receive as output; drafting and code generation are closer to balanced.

Which components make up the bill?

The two options put the same costs in different places. An API bundles hardware, power and operations into a price per token. Own infrastructure turns them into fixed monthly costs that the organisation carries whether the GPUs are busy or not.

Component API Own GPUs What moves it
Tokens Price per million tokens Included in fixed cost Volume, prompt length
Hardware Included in price Purchase or rental GPU type and count
Throughput Provider's concern Tokens per GPU-hour Model, precision, batching
Idle capacity Not charged Paid for Load pattern, peaks
Power and hosting Included in price Owned servers only Energy contract, rack
People Integration only Operations and on-call Team size, skills

Sources: Lindstead cost framework; pricing structure from the API providers' pricing pages; throughput factors from the vLLM documentation; accessed 28 September 2026.

How is API inference priced?

APIs charge per million tokens, with separate prices for input and output. Output tokens cost more because the model generates them one at a time, while input is processed in parallel. Converted at the ECB reference rate of 25 September 2026, Anthropic's list price for Claude Sonnet 5 is €1.75 per million input tokens and €8.77 per million output tokens; OpenAI's GPT-6 Astra costs €8.77 and €43.85.12 Hosted open-weight models are far cheaper: DeepSeek V4-Pro costs €1.16 and €3.47 during peak hours and half that outside them, and Mistral Large 3 costs €0.44 and €1.32.34 Most providers also discount input that is repeated and cached, which matters for applications that send the same instructions with every request.

To compare providers with a single number, Lindstead blends the two prices at three input tokens for every output token, a ratio typical of document and search workloads. Claude Sonnet 5 then costs €3.51 per million tokens, GPT-6 Astra €17.54, DeepSeek V4-Pro €1.74 at peak and Mistral Large 3 €0.66. An organisation with a different mix should use its own ratio.

How fast are API prices falling?

Quickly. Epoch AI found that the price of reaching a given level of model performance fell by between 9 and 900 times per year, depending on the task and the performance level chosen, and cautioned that the fastest declines may not continue.5 A cost model that locks in today's API prices for three years therefore flatters own infrastructure. GPU rental moved the other way in 2026: Nebius raises its on-demand prices on 1 October 2026, and the median H200 rental rose about a quarter in a year, as the current GPU rental prices in Europe show.10

Exhibit 1

An API bill scales with volume; a server does not

At one billion tokens a month, API bills range from under €1,000 to over €17,000. A server costs the same at any volume, as long as it can carry the load.

At one billion tokens a month, only a premium closed API costs more than an owned eight-GPU server.

Monthly cost of one billion tokens through an API, and fixed monthly cost of own GPUs, € per month

Closed model API Hosted open-weight model API Own GPUs, fixed monthly cost
0 5,000 10,000 15,000 20,000 17,539 3,508 1,736 658 14,558 1,255 GPT-6 Astra Claude Sonnet 5 DeepSeek V4-Pro API³ Mistral Large 3 API Owned 8× B200¹ Rented 1× RTX PRO 6000²

1 Eight-GPU B200 server at €19.91 per server-hour (3 years, 70% utilisation, vendor estimate). Excludes staff. 2 On-demand list price at Verda (FI), €1.72 per GPU-hour for 730 hours. Suits models up to about 120 billion parameters. 3 Peak-hour price; DeepSeek charges half outside peak hours. All API prices blended at 3 input : 1 output tokens.

Source: OpenAI, Anthropic, DeepSeek and Mistral pricing pages; Mercatus (vendor analysis, 16 July 2026); verda.com; accessed 28 September 2026; converted at the ECB reference rate of 25 September 2026

What does it cost to run an LLM on your own GPUs?

Self-hosting turns a variable price per token into four blocks of cost. Only the first is easy to find on a price list.

Hardware: buy or rent

An owned server is paid once and written off over its life. A vendor analysis by Mercatus puts an eight-GPU B200 server at about €19.91 per server-hour at 70 percent utilisation over three years, or roughly €14,558 a month. The figure covers the purchase net of resale value, GPU power, colocation and maintenance, but lists no staff costs.6 It is a vendor estimate and should be read as such. Renting avoids the purchase: Verda, a Finnish provider, lists an RTX PRO 6000 with 96 GB at €1.72 per GPU-hour, about €1,255 for a full month, and an H200 at €4.07.7 Which GPU a model needs depends on its size; the guide to hardware needed per model size sets this out.

Throughput: tokens per GPU-hour

The hardware price only becomes a cost per token once it is divided by the number of tokens the server produces. That number depends on the model, the numeric precision, the length of prompts and answers, and how many requests the inference server processes together. Serving software such as vLLM raises throughput by batching requests and by managing the memory each conversation needs.8 No published figure replaces a measurement on the actual model and workload.

Utilisation: idle capacity is paid for

An API charges only for tokens used. A server that is sized for the morning peak and idles at night costs the same either way. Utilisation is therefore the lever that most often separates a good self-hosting case from a poor one, and the reason the Mercatus figure states its 70 percent assumption.

People, power and facilities

Owned hardware needs rack space, power, cooling and network capacity, and someone on call when it fails. Rented GPUs shift the facility costs to the provider but still need engineers for the inference server, model updates, access control and monitoring. These costs depend on the organisation and are not in any list price; they belong in the model as explicit assumptions.

How do you compare an API bill with your own infrastructure?

The comparison rests on one formula: the monthly cost of self-hosting divided by the API price per token gives the break-even volume, the number of tokens a month at which both options cost the same. Above that volume the own server is cheaper, provided it can carry the load. A worked example with the published assumptions:

  1. Volume. An organisation processes 2 billion tokens a month in a document assistant, three input tokens for each output token.
  2. API bill. At Claude Sonnet 5's blended €3.51 per million tokens, the bill is 2,000 × €3.51 = about €7,020 a month.
  3. Own server. The owned eight-GPU B200 server costs about €14,560 a month, before staff.
  4. Break-even. €14,560 divided by €3.51 per million tokens is about 4.2 billion tokens a month. At 2 billion, the API is cheaper.
  5. Different API. Against GPT-6 Astra at €17.54, the same volume costs about €35,080 a month and break-even falls to about 0.8 billion tokens. The server wins, if an open model on it does the job of GPT-6 Astra.

Staff costs push both break-even points further out. The break-even calculator runs the same steps with an organisation's own volume, prices and staff costs. For a decision, Lindstead prepares a three-year cost model with every assumption shown, covering API use, rented GPUs and owned servers.

Why is the choice of API the biggest variable?

Exhibit 1 shows a spread of more than twenty times between the cheapest and the most expensive API. Hardware prices vary far less. The break-even volume therefore depends mostly on which API the own server is measured against, and that choice has to be quality-matched: the question is whether the open model on the own hardware does the job of the API model it replaces.

If an open model can replace a premium closed model, the break-even volume is low. If a cheap hosted open-weight API would do the same job, it is very high: a rented RTX PRO 6000 at about €1,250 a month only matches the Mistral Large 3 API at about 1.9 billion tokens a month. A cost-benefit study with authors from Carnegie Mellon University reaches a similar conclusion: on-premise deployment breaks even within three months for small models, in 6 to 24 months for medium models and often only after more than two years for large models. It finds self-hosting viable mainly from about 50 million tokens a month or under strict data residency rules.9

What does a cost model leave out?

  • Quality and review. A cheaper model whose answers need more human checking is not cheaper.
  • Price trends. API prices for a given capability keep falling, while GPU prices rose in 2026.
  • Growth. Volume that doubles may need a second server; an API scales without a purchase decision.
  • Control over data. Keeping prompts and documents inside the organisation's own environment has no price in the model, and is often the deciding factor.

That last point weighs heaviest in financial services, healthcare and legal and professional services, where client and patient data are subject to strict confidentiality. The comparison of on-premise and private cloud AI covers where such workloads can run, and the Model Index lists which open models are available and under which licence.

Lindstead builds cost models that compare API use, rented GPUs and owned servers for an organisation's own workloads and volumes, and deploys the first use case on the chosen infrastructure. Schedule a meeting to review AI infrastructure costs.

Frequently asked questions

  • There is no single figure. Through an API the monthly cost is the number of tokens times the price per token, which ranges widely between providers. On own GPUs it is the server or rental cost plus people, power and idle capacity, largely independent of volume. The break-even calculator on lindstead.com shows both for a given volume.

  • Only at high and steady volume, and mainly when the open model on the own server replaces an expensive closed model. Against low-priced hosted open-weight APIs, an owned eight-GPU server needs tens of billions of tokens a month to break even. Many organisations self-host for control over data rather than for cost.

  • A token is a fragment of text, often a word or part of a word, that the model reads or writes. APIs charge per million tokens, with a separate and higher price for output tokens than for input tokens, and often a lower price for input that is repeated and cached.

  • It depends on the model, the numeric precision, the number of requests processed together and the length of prompts and answers. Inference servers such as vLLM raise throughput by batching requests. The figure has to be measured for the actual workload before capacity is bought.

Sources

  1. Anthropic, Claude API pricing, accessed 28 September 2026. platform.claude.com/docs/en/about-claude/pricing
  2. OpenAI, API pricing, accessed 28 September 2026. developers.openai.com/api/docs/pricing
  3. DeepSeek, models and pricing, accessed 28 September 2026. api-docs.deepseek.com/quick_start/pricing
  4. Mistral AI, API pricing, accessed 28 September 2026. mistral.ai/pricing/api
  5. Epoch AI, LLM inference prices have fallen rapidly but unequally across tasks, 12 March 2025, accessed 28 September 2026. epoch.ai/data-insights/llm-inference-price-trends
  6. Mercatus, B200 server price and cost per hour, 16 July 2026 (vendor analysis), accessed 28 September 2026. mercatus-ai.com/blog/b200-server-price
  7. Verda, GPU pricing, accessed 28 September 2026. verda.com/pricing
  8. Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention, SOSP 2023, accessed 28 September 2026. arxiv.org/abs/2309.06180; vLLM documentation, docs.vllm.ai
  9. Pan, Wang et al., A Cost-Benefit Analysis of On-Premise Large Language Model Deployment: Breaking Even with Commercial LLM Services, arXiv 2509.18101, v3 11 November 2025, accessed 28 September 2026. arxiv.org/abs/2509.18101
  10. Nebius, pricing including changes from 1 October 2026, and getdeploying.com, H200 price history, accessed 28 September 2026. nebius.com/prices