The main hardware requirement for an open-weight LLM is GPU memory: the model's weights plus working memory for every conversation must fit. Models up to about 30 billion parameters run on one 48 to 96 GB GPU. Around 120 billion needs one 80 GB GPU or more. The largest open models need a full eight-GPU server.
The short answer
- At 16-bit precision, weights take about 2 GB of GPU memory per billion parameters; 8-bit halves that, 4-bit quarters it.
- gpt-oss-120b runs on a single 80 GB GPU, according to its model card. Llama 4 Maverick and Mistral Large 3 need a full eight-GPU server.
- A European cloud such as Verda rents every GPU class by the hour, from about €1.35 for an L40S to about €7.41 for a B300, so sizing can be tested before hardware is bought.
This guide covers the hardware layer. The self-hosted LLM guide covers the full decision, and the guide to what drives inference cost turns the hardware into a cost per token.
Why is GPU memory the main constraint?
A language model generates each word by reading all of its active weights from GPU memory. The weights therefore have to sit in the memory of the GPU, or of several GPUs working together. Computing speed decides how fast answers come; memory decides whether the model runs at all.
The rule of thumb is simple arithmetic. Hugging Face puts the memory for the weights at roughly 2 GB per billion parameters in 16-bit precision (bfloat16 or float16).1 Stored at 8 bits a parameter takes half that, and at 4 bits a quarter. A 31-billion-parameter model such as Gemma 4 31B needs about 62 GB at 16-bit, and Google states that its unquantised weights fit on a single 80 GB H100.2
Mixture-of-experts models (MoE) change the speed, not the memory. Mistral Small 4 has 119 billion parameters, of which only 6 billion are active for each token.3 That makes it fast, but all 119 billion still have to be loaded. Memory is sized on total parameters; speed follows active parameters.
Working memory for context grows with users
On top of the weights, the GPU keeps a working memory for each conversation, the key-value cache (KV cache). NVIDIA describes its size as growing with the batch size, the sequence length, the number of layers and the hidden size of the model; for Llama 2 7B, one conversation of 4,096 tokens already takes about 2 GB.4 Long documents and many simultaneous users multiply that. The vLLM team found that serving systems before PagedAttention wasted 60 to 80 percent of this memory, and that its approach cut the waste to under 4 percent.5 Sizing therefore starts from expected context length and concurrency, not from the model alone.
What is quantisation and what does it change?
Quantisation stores weights at lower precision, for example 8 or 4 bits instead of 16. It cuts memory; the effect on speed depends on the hardware and the inference software. Hugging Face notes that it trades memory efficiency against accuracy, and that 4-bit models can in practice give different results than 8-bit or 16-bit versions.1 Several developers now publish models in reduced precision themselves: gpt-oss-120b ships with its expert weights in the 4-bit MXFP4 format.6 Any quantised model should be tested on the organisation's own tasks before use.
Precision decides the memory bill
The same model needs a quarter of the memory for its weights at 4-bit that it needs at 16-bit.
A 119-billion-parameter model needs about 238 GB for its weights at 16-bit, and about 60 GB at 4-bit.
GPU memory needed for model weights, by precision, GB
1 Weights only. Working memory for context and concurrent users (KV cache) comes on top. Real 4-bit files are somewhat larger, because some layers stay at higher precision.
Source: Parameter counts from the model cards of IBM, utter-project, Google and Mistral AI; 2 bytes per parameter at 16-bit per Hugging Face; 8-bit and 4-bit scaled by bit width; accessed 28 September 2026
How much GPU memory do popular open models need?
The table groups the models on the Model Index by size. Hardware in the first four rows is what the developer states in its model card or announcement; the last row follows the Model Index and is indicative.
| Size class | Example models | Weights at 8-bit | Minimum hardware |
|---|---|---|---|
| Up to 31B, dense | Gemma 4, Granite 4.2, EuroLLM-22B | up to 31 GB | 1× H100 80 GB for Gemma 4 31B |
| About 110B to 120B, MoE | gpt-oss-120b, Llama 4 Scout | 109 to 117 GB | 1× 80 GB GPU at 4-bit |
| About 400B, MoE | Llama 4 Maverick | about 400 GB | One 8× H100 server, FP8 |
| About 675B, MoE | Mistral Large 3 | about 675 GB | One 8× H200 or B200 server, FP8 |
| 1.6T to 2.8T, MoE | DeepSeek V4 Pro, Kimi K3, Qwen3.8 | 1.6 to 2.8 TB | Full 8× B200 or B300 server, or more |
Runs on one GPU Needs a multi-GPU server
Sources: model cards and announcements of Google, IBM, utter-project, OpenAI, Mistral AI, Meta, DeepSeek, Moonshot AI and Alibaba; weights at 1 byte per parameter; last row from the Lindstead Model Index (indicative); accessed 28 September 2026.
Three points stand out. gpt-oss-120b fits on one 80 GB GPU because it ships in 4-bit precision.6 For Mistral Small 4, Mistral states a minimum of "4x NVIDIA HGX H100, 2x NVIDIA HGX H200, or 1x NVIDIA DGX B200", well above the 238 GB its weights take at 16-bit.3 And Meta states that Llama 4 Maverick fits on a single H100 host in FP8, while the smaller Llama 4 Scout fits on one H100 with 4-bit quantisation.7 Mistral Large 3 fits a single node of B200 or H200 GPUs in FP8, or of H100 GPUs in the 4-bit NVFP4 format.8
Which GPU should you choose for a self-hosted LLM?
Six NVIDIA data-centre GPUs cover almost every enterprise deployment. They differ in memory and in how they connect to each other, which matters as soon as a model spans more than one card.
| GPU | Memory | Link between GPUs | Rental at Verda (FI) | Typical use |
|---|---|---|---|---|
| L40S | 48 GB | PCIe Gen4 | €1.35 per hour | Up to about 30B at 8-bit |
| RTX PRO 6000 | 96 GB | PCIe Gen5 | €1.72 per hour | Up to about 120B at 4-bit |
| H100 SXM | 80 GB | NVLink 900 GB/s | €3.12 per hour | gpt-oss-120b; 8× for Maverick |
| H200 SXM | 141 GB | NVLink 900 GB/s | €4.07 per hour | 8× for Mistral Large 3 |
| B200 | 180 GB | NVLink 1.8 TB/s | €6.01 per hour | 8× for the largest models |
| B300 | 288 GB | NVLink 1.8 TB/s | €7.41 per hour | 8× for the largest models |
PCIe card SXM module on an NVLink board
Sources: NVIDIA product pages and DGX user guides (memory, interconnect); verda.com/pricing, on-demand price per GPU, converted at the ECB reference rate of 25 September 2026; accessed 28 September 2026.
The European cloud GPU price index compares these rates across European and US providers, including H100 capacity at Scaleway and OVHcloud and the Nebius price increase of 1 October 2026.
What does a self-hosted LLM server look like?
Apart from the GPUs, an inference server needs CPUs and system memory to feed them, fast local storage for model files that run to hundreds of gigabytes, and network capacity towards the applications that call it. The GPUs dominate the price and the power draw.
What is the difference between PCIe and SXM GPUs?
PCIe cards such as the L40S and RTX PRO 6000 plug into a standard server slot. They talk to each other over the PCIe bus: PCIe Gen5 offers 128 GB/s in both directions combined.9 SXM modules such as the H100, H200 and B200 sit on an NVIDIA HGX board with NVLink, at 900 GB/s per GPU for H100 and H200 and 1.8 TB/s for B200.910 When a model is split across eight GPUs, the GPUs exchange data for every token; the faster link keeps them busy. For models that fit on one card, PCIe is sufficient and simpler to host.
What does a server rack need?
Power and cooling set the limit in most existing data centres. NVIDIA specifies a maximum of about 10.2 kW for a DGX H100 system and about 14.3 kW for a DGX B200 with eight B200 GPUs and 1,440 GB of GPU memory.1011 A single RTX PRO 6000 Server Edition draws up to 600 W and an L40S up to 350 W.9 Many server rooms built for ordinary IT cannot supply or cool 14 kW for one machine, which is one reason organisations rent eight-GPU capacity or place it in a colocation facility. The comparison of on-premise and private cloud AI weighs these options.
Should you start with one GPU or a full server?
Most organisations start small. A single rented GPU in a European cloud is enough to prove a use case with a model of up to about 120 billion parameters, and it produces the one number that published sizing tables cannot: measured throughput on the organisation's own prompts and documents. Owned hardware can then be sized from measured load rather than estimates.
Whether buying ever pays off depends on volume and on the API it replaces; the break-even calculator shows the volume at which an owned server matches an API bill. Lindstead provides infrastructure sizing for each workload, comparing rented GPUs, owned servers and API use.
How do you set up the server software?
The software stack has four layers: the operating system with GPU drivers; an inference server such as vLLM, which loads the model and batches requests;5 a gateway that handles authentication, access rights and logging; and monitoring for load, latency and errors. The self-hosted LLM guide describes the full architecture.
Hardware choices weigh heaviest where data may not leave the organisation, as in healthcare, the public sector and industrial companies protecting designs and process knowledge. Lindstead sizes GPU capacity for an organisation's own models and workloads and deploys the first use case on the chosen infrastructure. Schedule a meeting to discuss hardware plans.
Frequently asked questions
-
The one with enough memory for the chosen model at the chosen precision, plus room for context. For models up to about 30 billion parameters at 8-bit, a 48 GB L40S or a 96 GB RTX PRO 6000 is enough for the weights. Models of around 120 billion parameters need about 60 GB at 4-bit, so an 80 or 96 GB GPU. The largest open models need a full eight-GPU H200, B200 or B300 server.
-
Yes, if it is quantised. At 16-bit the weights of a 70-billion-parameter model take about 140 GB, which does not leave room for context even on a 141 GB H200. At 8-bit they take about 70 GB, which fits a 96 GB RTX PRO 6000 with some headroom. At 4-bit, about 35 GB, it fits a 48 GB L40S.
-
About 2 GB per billion parameters for the weights at 16-bit precision, about 1 GB at 8-bit and about 0.5 GB at 4-bit. Working memory for context and concurrent users comes on top and grows with the length of the conversations and the number of users.
-
No. Models up to about 30 billion parameters run on PCIe cards such as the L40S or RTX PRO 6000 in a standard server. H100, H200 and B200 GPUs on NVLink boards become necessary when a model has to be split across several GPUs, as with the 400-billion-parameter and larger open models.
Sources
- Hugging Face, Optimizing LLMs for speed and memory (Transformers documentation), accessed 28 September 2026. huggingface.co/docs/transformers/main/en/llm_tutorial_optimization
- Google, Gemma 4 model card and announcement, accessed 28 September 2026. ai.google.dev/gemma/docs/core/model_card_4
- Mistral AI, Mistral Small 4, 16 March 2026, accessed 28 September 2026. mistral.ai/news/mistral-small-4
- NVIDIA Technical Blog, Mastering LLM Techniques: Inference Optimization, 17 November 2023, accessed 28 September 2026. developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization
- vLLM, Easy, fast and cheap LLM serving with PagedAttention, 20 June 2023, accessed 28 September 2026. vllm.ai/blog/2023-06-20-vllm
- OpenAI, gpt-oss-120b model card, accessed 28 September 2026. huggingface.co/openai/gpt-oss-120b
- Meta, Llama 4 Maverick and Scout model cards, accessed 28 September 2026. huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct
- Mistral AI, Mistral Large 3 model card, accessed 28 September 2026. huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512
- NVIDIA, L40S, RTX PRO 6000 Blackwell Server Edition, H100 and H200 product pages, accessed 28 September 2026. nvidia.com/en-us/data-center/h200
- NVIDIA, DGX B200 and HGX product pages, and DGX B300 user guide, accessed 28 September 2026. nvidia.com/en-us/data-center/dgx-b200
- NVIDIA, DGX H100 datasheet, accessed 28 September 2026. nvidia.com, DGX H100 datasheet
- Other model cards: IBM Granite 4.2, utter-project EuroLLM-22B, DeepSeek V4 Pro, Moonshot AI Kimi K3 and Qwen3.8-2.4T-A95B on Hugging Face; Verda GPU pricing, verda.com/pricing; all accessed 28 September 2026.