A self-hosted LLM is a large language model that runs on infrastructure the organisation controls: its own servers, rented racks, or dedicated GPU capacity at a European cloud provider. Prompts, documents and outputs stay out of third-party hands. For European enterprises the case rests mainly on control and regulation; cost favours self-hosting only at high, steady volumes.
Summary
- Control is the business case; cost is a constraint. An owned eight-GPU B200 server costs about €14,560 a month. It matches a premium closed-model API below one billion tokens a month, but cheap hosted open models only above twenty billion.1
- Choosing a model is a licence and origin decision. The strongest open models come from Chinese labs, and some leading licences restrict large or EU-based companies. Permissive European and US options run on one or a few GPUs.
- Start in a European GPU cloud. A 117-billion-parameter model such as gpt-oss-120b fits on one 80 GB GPU,21 and EU-owned providers rent an H100 for roughly half the hyperscaler price. Prove the use case before buying hardware.
What is a self-hosted LLM?
A self-hosted LLM is an open-weight language model whose weights the organisation downloads and runs with its own inference software, on hardware it owns or rents exclusively. Requests never pass through the model maker. The organisation decides where the model runs, who may use it, what is logged and when the model is replaced.
Three things are often confused with it:
- A closed-model API. Services such as the OpenAI or Anthropic APIs run the model for you. The provider processes every prompt, even when contracts limit what it may do with the data.
- A private endpoint at a hyperscaler. Dedicated capacity in an EU region of AWS, Azure or Google Cloud gives isolation, but the operator remains a US company. That matters for the jurisdiction question below.
- A local LLM on a laptop. Tools that run small models on a workstation are useful for experiments. They are not a production service with access control, logging and support.
Is a private LLM the same as a self-hosted LLM?
Not necessarily. "Private LLM" is a marketing term that covers everything from a dedicated hyperscaler endpoint to a model on the organisation's own servers. When a vendor offers a private LLM, ask three questions: who operates the hardware, under which jurisdiction does that operator fall, and who can read prompts, outputs and logs? Only when the answer to all three is "the organisation itself, or a provider under EU jurisdiction only" is the deployment self-hosted in the sense this guide uses.
Why do European enterprises host their own LLM?
Five reasons recur in regulated organisations. They are listed in the order in which they usually carry weight in a board discussion.
- Control over sensitive data. Client files, patient records, privileged legal documents and trade secrets stay inside systems the organisation governs. For many use cases this is the only reason that matters.
- Jurisdiction. Under the US CLOUD Act, US providers must disclose data in their possession, custody or control "regardless of whether" it is stored inside or outside the United States.10 Microsoft France told the French Senate under oath that it could not guarantee such a request would never be granted.11
- Sector rules. DORA requires financial entities to manage ICT third-party risk and plan exits; in November 2025 the European supervisors designated 19 critical ICT providers, including AWS, Google Cloud and Microsoft.12 Health data is a special category under GDPR Article 9.17 NIS2 brings supply-chain duties to essential and important entities.
- Independence from provider changes. A hosted model can be changed, repriced or retired by its provider. A self-hosted model stays exactly as tested until the organisation decides to upgrade.
- Cost at volume. At high and steady volumes, owned or rented GPUs can be cheaper than paying per token. The section on cost below shows why this is rarely the main argument.
Demand is broad. In an Accenture survey of 1,928 organisations, 62 percent of European organisations said they were actively seeking sovereign AI solutions (vendor survey, November 2025).9 HSBC announced in December 2025 that it would use "self-hosted AI models that operate on HSBC's internal technology systems" as part of its partnership with Mistral AI.8 Sector specifics are covered on the pages for financial services, healthcare, legal and professional services and the public sector.
Where can you host an LLM: on-premise, private cloud or European GPU cloud?
There are four realistic places to run a self-hosted model. They differ in who owns the hardware, how much capital is needed up front, how much the organisation operates itself and which jurisdiction applies.
Four hosting options, from most to least control over the hardware
| Option | Hardware owner | Upfront outlay | You operate | Examples |
|---|---|---|---|---|
| On-premise | Your organisation | High | Everything, incl. facility | Own data centre |
| Colocation | Your organisation | High | Servers and software | Rented racks in an EU facility |
| EU-owned GPU cloud | EU-owned provider | None | Model and software | Verda, Nebius, Scaleway |
| US hyperscaler, EU region | US-owned provider | None | Model and software | AWS, Azure, Google Cloud |
EU jurisdiction only US jurisdiction can also apply
Sources: 18 U.S.C. 2713; provider websites; Lindstead classification. Accessed 28 September 2026.
On-premise gives the most control and the highest burden: power, cooling, spare parts and a team that can replace a failed GPU. Colocation removes the building but not the hardware. An EU-owned GPU cloud rents dedicated GPUs by the hour or month; the organisation runs the model and software, the provider runs the machines, and only EU law applies to the operator. A US hyperscaler region in the EU offers the broadest service catalogue, but the CLOUD Act question remains because jurisdiction follows ownership, not the location of the data centre.
EU-owned GPU clouds cost about half as much
European providers rent an NVIDIA H100 for roughly half the on-demand price of AWS and Azure, and about a third of Google Cloud.
EU providers rent an H100 at roughly half the hyperscaler price.
On-demand price of one NVIDIA H100, € per GPU-hour
1 Nebius price from 1 October 2026; until 30 September 14 percent lower. Scaleway (FR) lists from €2.73 per hour.
Source: Provider price pages and getdeploying.com, accessed 27 September 2026; OVHcloud via aggregator; converted at the ECB reference rate of 25 September 2026
The detailed trade-off between owning hardware and renting dedicated capacity is covered in on-premise or private cloud AI for regulated industries. Current prices per provider are tracked monthly in the European cloud GPU price index.
Which models can you self-host?
Any model whose weights are published can be self-hosted, but capability, size and licence terms differ widely. The table below is a shortlist from the Lindstead Model Index, grouped by the origin of the maker. Hardware is indicative for production inference.
Leading open-weight models by origin, September 2026
| Model | Size | Licence | Indicative hardware |
|---|---|---|---|
| Europe | |||
| Mistral Small 4 Mistral AI (FR) | 119B MoE, 6B active | Apache 2.0 | 4× H100 or 2× H200 |
| Mistral Large 3 Mistral AI (FR) | 675B MoE, 41B active | Apache 2.0 | 8× H200 (FP8) |
| Mistral Medium 3.5 Mistral AI (FR) | 128B dense | Modified MIT | 2× H200 (FP8) |
| EuroLLM-22B EU research consortium | 22B dense | Apache 2.0 | 1 GPU |
| United States | |||
| gpt-oss-120b OpenAI | 117B MoE | Apache 2.0 | 1× 80 GB GPU |
| Gemma 4 Google | Up to 31B | Apache 2.0 | 1 GPU |
| Llama 4 Maverick Meta | 400B MoE, 17B active | Llama 4 Community | 8× H100 |
| Granite 4.2 IBM | 3B to 30B | Apache 2.0 | 1 GPU |
| China | |||
| DeepSeek V4 Pro DeepSeek | 1.6T MoE, 49B active | MIT | 8× H200 or 8× B300 |
| Kimi K3 Moonshot AI | 2.8T MoE, 104B active | Kimi K3 License | 8× B200 minimum |
| Qwen3.8 (flagship) Alibaba | 2.4T MoE, 95B active | Qwen3.8-Max License | Multi-node |
Permissive licence Custom licence, review terms Restrictions for large or EU companies
Sources: model cards and licence texts on Hugging Face and the makers' websites, as listed in the Lindstead Model Index; accessed 28 September 2026. MoE: mixture of experts; "active" is the share of parameters used per token.
What do open-weight licences allow a large company to do?
"Open" does not mean unrestricted. The Mistral Medium 3.5 licence states that a company may not exercise any rights under it if its global consolidated monthly revenue "exceeds $20 million" in the preceding month; such companies need a commercial licence.18 Meta's Llama 4 policy withholds the multimodal rights from companies based in the EU.19 Several Chinese flagship models, including Kimi K3 and the largest Qwen3.8, now use custom licences whose terms differ for internal use and resale. Permissive Apache 2.0 or MIT licences remain available for, among others, Mistral Small 4 and Large 3, gpt-oss-120b, Gemma 4, Granite 4.2, EuroLLM-22B and DeepSeek V4 Pro. Have legal counsel read the licence of every model before it reaches production; Lindstead's model selection work includes this licence and origin review.
Can a European company use Chinese open models?
Legally, in most cases yes, when the licence allows it and the weights run on the organisation's own infrastructure: no data then flows to the maker. The open questions are behaviour and security. The US standards body NIST found that agents built on DeepSeek's most secure model were on average 12 times more likely than US frontier models to follow malicious instructions, and that DeepSeek models echoed four times as many inaccurate and misleading Chinese Communist Party narratives as US reference models.20 Running the weights locally removes the data transfer; it does not remove what was trained into the model. Many organisations therefore set a model-origin policy at board level and choose a European or US model one capability tier lower for sensitive work.
What hardware do you need to run an LLM yourself?
The first constraint is GPU memory. The model weights must fit in memory, with room left for the KV cache, which grows with context length and the number of simultaneous users. A useful rule: each parameter needs 2 bytes at 16-bit precision, 1 byte at 8-bit and half a byte at 4-bit. Quantisation to 8 or 4 bits shrinks the model and usually costs some quality, which should be measured on the organisation's own tasks.
Precision decides whether a model fits on one GPU
At 16-bit, a 117-billion-parameter model needs about 234 GB for its weights alone. At 4-bit it needs about 59 GB and fits on a single 80 GB H100.
At 4-bit, the weights of models up to about 130 billion parameters fit on one 80 GB GPU.
GPU memory needed for the model weights, GB
1 gpt-oss-120b is published with 4-bit (MXFP4) expert weights and, according to OpenAI, fits on a single 80 GB GPU. Parameter counts from the model cards on Hugging Face.
Source: Lindstead calculation: parameters times bytes per parameter, weights only, excluding KV cache and runtime overhead; model cards on Hugging Face, accessed 28 September 2026
For reference, an NVIDIA L40S has 48 GB, an H100 80 GB, an RTX PRO 6000 Blackwell Server Edition 96 GB and an H200 141 GB.2426 An eight-GPU DGX B200 system has 1,440 GB in total, or 180 GB per GPU.25 In practice, three tiers cover most enterprise needs:
- Up to about 30 billion parameters (Gemma 4, Granite 4.2, EuroLLM-22B): one L40S or RTX PRO 6000. Hetzner rents a dedicated RTX PRO 6000 server for €1,199 a month plus a €599 setup fee.7
- Around 120 billion parameters (gpt-oss-120b, Mistral Small 4): one 80 to 96 GB GPU for gpt-oss, two H200 for Mistral Small 4 at full precision.
- Frontier-scale open models (DeepSeek V4 Pro, Kimi K3, Mistral Large 3): a full eight-GPU H200 or B200 node.
Memory is only half of the sizing. Throughput, the number of tokens per second the server delivers to all users together, depends on the GPU, the inference server, the batch size and the length of prompts and answers. It should be measured with a realistic load test before hardware is bought. The full sizing method is in what hardware you need to run Llama, Mistral or Qwen.
What does a self-hosted LLM architecture look like?
A production deployment is more than a model on a GPU. The reference architecture below has seven layers. Requests flow from the application through the gateway to the inference server; retrieval adds company documents to the prompt; every step is logged and monitored.
Reference architecture for a self-hosted LLM, top to bottom
| Layer | Role | Common options |
|---|---|---|
| Applications | Chat, search, drafting and document workflows | Internal apps, chat front end |
| Gateway | Sign-in, access rights, quotas, routing, audit log | API gateway with SSO |
| Inference server | Loads the model, batches and streams requests | vLLM, SGLang, llama.cpp |
| Model store | Verified, versioned weights and licences | Safetensors files, registry |
| Retrieval | Finds company documents for each question | Vector index, search engine |
| Observability | Metrics, traces, quality and safety evaluation | Prometheus, OpenTelemetry |
| Hardware | GPUs, network, storage, power | H100, H200, RTX PRO 6000 |
Sources: vLLM, SGLang, llama.cpp and Safetensors documentation; OWASP Top 10 for LLM Applications 2025. Accessed 28 September 2026. Examples are illustrative, not a ranking.
Two design choices matter most. First, the gateway is the control point: it enforces single sign-on, maps users to the documents they may see, sets quotas and writes the audit log. Applications should never call the inference server directly. Second, retrieval must respect access rights. A retrieval layer that indexes all company documents without permissions will show confidential content to anyone who asks the right question.
Which inference server should you use?
The inference server loads the model into GPU memory, batches requests from many users and streams answers back. The common open-source options are:
- vLLM, which provides an HTTP server compatible with the OpenAI Completions, Chat Completions and Responses APIs, so existing applications can switch with little change.27
- SGLang, an open-source serving framework for language and multimodal models.29
- llama.cpp, for smaller models and CPU or mixed hardware, often used for edge and workstation deployments.30
Hugging Face's Text Generation Inference, once a common choice, is in maintenance mode; Hugging Face itself now recommends vLLM, SGLang and local engines such as llama.cpp.28 The right choice depends on the model architecture, the hardware and the operating team. Test two candidates with the real workload rather than relying on published benchmarks.
What does it cost to self-host an LLM, and when does it break even?
Self-hosting turns a variable cost per token into a fixed cost per server. Whether that pays off depends on volume, utilisation and, above all, the API it replaces. Lindstead's break-even analysis uses a vendor estimate of €19.91 per server-hour for an owned eight-GPU B200 system at 70 percent utilisation over three years, about €14,560 a month before staff.1
The cheaper the API, the further away the break-even point
Against a premium closed model, an owned server pays off below one billion tokens a month. Against a cheap hosted open model, it takes more than twenty billion.
Break-even volume ranges from under 1 to over 22 billion tokens a month.
Monthly volume at which an owned B200 server matches the API bill, billion tokens
1 Owned 8-GPU B200 server at about €14,560 per month (3 years, 70% utilisation), divided by blended API price per million tokens (3 input : 1 output). Excludes staff costs. Dollar prices converted at the ECB reference rate of 25 September 2026.
Source: Lindstead calculation; server cost from Mercatus (vendor estimate, July 2026); API list prices, 27 September 2026
Renting instead of buying lowers the entry point. At list prices, Verda rents an H100 for €3.09 an hour and Nebius for €3.95 from 1 October 2026, against €6.03 at AWS and €9.71 at Google Cloud.456 Dollar prices are converted at the ECB reference rate of 25 September 2026.
Four factors move the break-even point, usually further out:
- Staff. Someone must patch, monitor, upgrade and be on call. That cost is not in the server price.
- Utilisation. A server that idles at night and at weekends still costs money; an API does not.
- Price trends. API prices for a given level of capability have fallen steeply every year,3 while GPU rental prices rose in 2026.
- Quality. A cheaper model that needs more human review is not cheaper. Compare at equal quality on the organisation's own tasks.
Independent academic research reaches the same conclusion: owning hardware pays back within about three months for small models, within 6 to 24 months for medium models and often only after more than two years for the largest, and is mainly viable at high, steady volumes or under strict data-residency requirements.2 The cost logic is explained in full in LLM inference cost explained, and Lindstead's infrastructure and cost work models it for a specific organisation.
How do you secure and evaluate a self-hosted LLM?
Self-hosting removes third-party access to prompts and outputs. It does not make the model safe; it makes the organisation responsible for the risks a provider would otherwise manage. The OWASP Top 10 for LLM Applications 2025 is the most widely used checklist. It puts prompt injection first, followed by sensitive information disclosure, supply chain risk and data and model poisoning.32
- Prompt injection. The UK National Cyber Security Centre warns that prompt injection may never be fully mitigated.33 Limit what the model can do: least-privilege access to tools and documents, human approval for actions, and output checks before anything is executed.
- Supply chain. Download weights only from the maker's official repository, verify checksums and prefer the Safetensors format, which stores tensors without the code-execution risk of Python pickle files.31
- Poisoning. Research by Anthropic, the UK AI Security Institute and the Alan Turing Institute showed that around 250 poisoned documents can plant a backdoor in models of very different sizes.34 Control what goes into fine-tuning data and the retrieval index.
- Sensitive information. Prompts and logs contain the same confidential data as the source documents. Encrypt them, set retention periods and restrict who can read them.
How do you evaluate quality in production?
Build an evaluation set before go-live: a representative set of real questions from the use case with reference answers, reviewed by domain experts. Run it against every candidate model, again before every model or prompt change, and on a sample of live traffic. Log the model version, prompt template, retrieved documents and latency with every request, so that any answer can be traced back. Track refusal rate, answer accuracy on the evaluation set and user feedback over time. New open models appear monthly; a fixed evaluation set is what makes an upgrade decision defensible.
What does regulation require when you self-host?
Self-hosting changes who is responsible, not whether rules apply. The main frameworks for a European organisation are the following.
- AI Act roles. An organisation that uses an AI system under its own authority is a deployer. It becomes a provider of a high-risk system if it puts its name on the system or substantially modifies it (Article 25).15 Fine-tuning an open model makes the organisation a provider of a general-purpose model only in significant cases; the Commission's indicative criterion is training compute of more than a third of the original model's.16
- AI Act dates. The Digital Omnibus on AI entered into force on 27 July 2026. It postponed high-risk obligations to 2 December 2027 for stand-alone systems and to 2 August 2028 for AI in regulated products. Transparency duties under Article 50 apply from 2 August 2026. The AI literacy duty in Article 4 now requires measures that support literacy rather than a guaranteed level.14
- GDPR. Self-hosting avoids a transfer to a model provider, but the organisation still needs a lawful basis, purpose limitation and security of processing, and usually a data protection impact assessment under Article 35 for large-scale processing of sensitive data. A GPU cloud provider that hosts the data is a processor under Article 28.17
- DORA. Financial entities must include the GPU cloud or hardware supplier in their ICT third-party register and exit planning.12
- NIS2. In the Netherlands the Cyberbeveiligingswet has applied since 15 August 2026 and brings duties of care, including supply-chain security, to more than 8,000 organisations.13
The board-level view of how these rules interact is in the briefing Where should your AI run?
How do you deploy a self-hosted LLM, step by step?
The following sequence takes one use case from idea to production. Each step has a clear output, so that security, compliance and the business can sign off along the way.
- Pick one use case with sensitive data and a measurable outcome. Document review, internal search or drafting on confidential files are typical first candidates. Write down what "good" means in numbers.
- Classify the data. Determine which data categories the use case touches, which rules apply (GDPR Article 9, professional secrecy, DORA, NIS2) and therefore which hosting options are acceptable.
- Build the evaluation set. Collect real questions and reference answers with domain experts before choosing a model. This set decides every later choice.
- Shortlist models and check licences. Test two or three models from the Model Index against the evaluation set, and have each licence and the model origin reviewed.
- Choose the hosting option and size the hardware. Start with rented capacity at an EU-owned provider unless data rules require on-premise. Size memory from the model and throughput from a load test.
- Build the platform. Inference server, gateway with single sign-on and audit log, retrieval that respects access rights, monitoring and encrypted log storage.
- Secure and test. Run the OWASP checklist, red-team for prompt injection and data leakage, and rerun the evaluation set on the final configuration.
- Go live with a limited group, then hand over. Monitor quality and cost, document the runbook, train the operating team and set a schedule for model reviews and upgrades.
Lindstead supports European organisations in deploying a first production use case on infrastructure you control, from the use-case choice to the handover of a documented, tested system.
Frequently asked questions
-
A self-hosted LLM is a large language model that runs on infrastructure the organisation controls: its own servers, rented racks in a colocation facility, or dedicated GPU capacity at a cloud provider. The organisation downloads the model weights, runs an inference server and decides who can send prompts and read outputs. No third-party model provider sees the data.
-
Not always. Vendors use private LLM for anything from a dedicated endpoint at a US hyperscaler to a model on a server in the basement. A self-hosted LLM is the narrower case: the organisation runs the model itself. The useful questions are who operates the hardware, which jurisdiction the operator falls under, and who can access prompts and logs.
-
Only at high, steady volumes, and it depends on the API you compare against. On Lindstead's break-even analysis, an owned eight-GPU B200 server at about €14,560 a month matches a premium closed-model API below one billion tokens a month, but a cheap hosted open-model API only above twenty billion. Staff, power and idle capacity push the break-even point further out.
-
There is no single best model. The choice depends on the task, the languages, the hardware budget and the licence. Mistral Small 4, gpt-oss-120b and Gemma 4 are permissively licensed and run on one or a few GPUs. The strongest open models are Chinese, and several leading licences restrict large or EU-based companies. The Lindstead Model Index lists current options with licence and hardware.
-
Yes, but self-hosting does not make processing compliant by itself. The organisation still needs a lawful basis, purpose limitation, access control, retention rules for prompts and logs, and usually a data protection impact assessment for large-scale or sensitive processing. What self-hosting removes is the transfer of personal data to an external model provider.
-
Models up to about 30 billion parameters run on one 48 to 96 GB GPU. Around 120 billion parameters, gpt-oss-120b runs on one 80 to 96 GB GPU, while Mistral Small 4 at full precision needs two H200 GPUs. Frontier-scale open models such as DeepSeek V4 Pro need a full eight-GPU H200 or B200 node.
-
Serving a model on rented GPUs is the quick part. Most of the time goes into what surrounds the model: data classification, access control, retrieval over company documents, evaluation and sign-off by security and compliance. Buying hardware adds procurement and delivery time, which is why many organisations start in a European GPU cloud.
Sources
- Mercatus, B200 server price and cost per hour, 16 July 2026 (vendor analysis). www.mercatus-ai.com/blog/b200-server-price
- Pan et al. (Carnegie Mellon University and independent researchers), cost-benefit analysis of on-premise LLM deployment, arXiv 2509.18101. arxiv.org/abs/2509.18101
- Epoch AI, LLM inference price trends, 12 March 2025. epoch.ai/data-insights/llm-inference-price-trends
- Verda GPU pricing. verda.com/pricing
- Nebius pricing, including changes from 1 October 2026. nebius.com/prices
- getdeploying.com, NVIDIA H100 rental price comparison (AWS, Azure, Google Cloud). getdeploying.com/gpus/nvidia-h100
- Hetzner, GEX131 dedicated GPU server announcement. www.hetzner.com/pressroom/new-gex131/
- HSBC, HSBC and Mistral AI join forces to accelerate AI adoption across global bank, 1 December 2025. www.hsbc.com/news-and-views/news/media-releases/2025/hsbc-and-mistral-ai-join-forces-to-accelerate-ai-adoption-across-global-bank
- Accenture, Europe seeking greater AI sovereignty, November 2025 (vendor survey). newsroom.accenture.com/news/2025/europe-seeking-greater-ai-sovereignty-accenture-report-finds
- 18 U.S.C. 2713, Required preservation and disclosure of communications and records (CLOUD Act). www.law.cornell.edu/uscode/text/18/2713
- The Register, Microsoft exec admits it cannot guarantee data sovereignty, 25 July 2025. www.theregister.com/off-prem/2025/07/25/microsoft-exec-admits-it-cannot-guarantee-data-sovereignty/458553
- EIOPA, ESAs designate critical ICT third-party providers under DORA, 18 November 2025. www.eiopa.europa.eu/european-supervisory-authorities-designate-critical-ict-third-party-providers-under-digital-2025-11-18_en
- Rijksoverheid, Cyberbeveiligingswet and Wet weerbaarheid kritieke entiteiten in force from 15 August 2026, 7 July 2026. www.rijksoverheid.nl/actueel/nieuws/2026/07/07/cyberbeveiligingswet-en-wet-weerbaarheid-kritieke-entiteiten-vanaf-15-augustus-2026-van-kracht
- Lewis Silkin, The Digital Omnibus on AI enters into force today, 27 July 2026. www.lewissilkin.com/insights/2026/07/27/the-digital-omnibus-on-ai-enters-into-force-today-102nedo
- Regulation (EU) 2024/1689 (AI Act), EUR-Lex. eur-lex.europa.eu/eli/reg/2024/1689/oj
- European Commission, Guidelines on the scope of the obligations for general-purpose AI models, C(2025) 5045, 18 July 2025. ai-act-service-desk.ec.europa.eu/sites/default/files/2025-07/guidelines_on_the_scope_of_the_obligations_for_generalpurpose_ai_models_established_by_regulation_1cx2atxgq79us4n3x8jfgyy1qlm_118340-3.pdf
- Regulation (EU) 2016/679 (GDPR), EUR-Lex. eur-lex.europa.eu/eli/reg/2016/679/oj
- Mistral AI, Mistral-Medium-3.5-128B model card and licence, Hugging Face. huggingface.co/mistralai/Mistral-Medium-3.5-128B
- Meta, Llama 4 Acceptable Use Policy. www.llama.com/llama4/use-policy
- NIST CAISI, Evaluation of DeepSeek AI models finds shortcomings and risks, 30 September 2025. www.nist.gov/news-events/news/2025/09/caisi-evaluation-deepseek-ai-models-finds-shortcomings-and-risks
- OpenAI, gpt-oss-120b model card, Hugging Face. huggingface.co/openai/gpt-oss-120b
- EuroLLM-22B-Instruct-2512 model card, Hugging Face. huggingface.co/utter-project/EuroLLM-22B-Instruct-2512
- Google, Gemma 4 31B model card, Hugging Face. huggingface.co/google/gemma-4-31B-it
- NVIDIA, H100 and H200 specifications (80 GB and 141 GB). www.nvidia.com/en-us/data-center/h200/
- NVIDIA, DGX B200 specifications (1,440 GB across eight GPUs). www.nvidia.com/en-us/data-center/dgx-b200/
- NVIDIA, RTX PRO 6000 Blackwell Server Edition (96 GB) and L40S (48 GB) specifications. www.nvidia.com/en-us/data-center/rtx-pro-6000-blackwell-server-edition/
- vLLM documentation, online serving and OpenAI-compatible server. docs.vllm.ai/en/latest/serving/online_serving/
- Hugging Face, Text Generation Inference documentation (maintenance mode notice). huggingface.co/docs/text-generation-inference/index
- SGLang project repository. github.com/sgl-project/sglang
- llama.cpp project repository. github.com/ggml-org/llama.cpp
- Hugging Face, Safetensors documentation. huggingface.co/docs/safetensors/index
- OWASP, Top 10 for LLM Applications 2025. genai.owasp.org/llm-top-10/
- UK NCSC, on prompt injection, December 2025. www.ncsc.gov.uk/news/mistaking-ai-vulnerability-could-lead-to-large-scale-breaches
- Anthropic, UK AI Security Institute and Alan Turing Institute, A small number of samples can poison LLMs of any size, October 2025. www.anthropic.com/research/small-samples-poison
Method: desk research of public primary sources, all accessed 28 September 2026. Figures are published list prices and specifications, not Lindstead measurements. Vendor-sponsored data is labelled as such. Corrections: contact@lindstead.com.