Self-Hosting an LLM: What It Costs and When It Makes Sense
Key Takeaways
- →Self-hosting is a fixed cost: the GPU bills you whether it serves ten requests a minute or sits mostly idle, while a hosted API charges per token.
- →Self-hosting wins on sustained predictable volume, regulated data, latency control, and deep fine-tuning, not on price alone.
- →Use Ollama for local development and vLLM for production, since PagedAttention fits more concurrent requests on the same GPU.
- →Llama is no longer the default open-weight pick, and both Hugging Face Open LLM Leaderboards have been retired.
- →Read the license on the exact checkpoint: MIT, Apache 2.0 and custom community licenses are not interchangeable.

On this page⌄
Self-hosting an LLM means running an open-weight model on hardware you control (your own servers or rented GPUs) instead of paying a provider per API call. It is usually not the cheap option. A GPU costs the same whether it answers ten requests a minute or sits mostly idle, while a hosted API only charges for tokens you actually use. So the real question is not "is self-hosting cheaper," it is "can I keep the hardware busy enough to make owning it worthwhile."
This guide walks through when self-hosting wins, when a hosted API is the smarter call, how to pick a serving stack, and which open-weight models you can legally run in production today.
Self-hosted vs API: the break-even math
Renting a GPU has a real, current price. So does sending the same workload to a hosted API. You can compare them directly if you run the arithmetic yourself, on your own volume, instead of trusting someone else's fixed threshold.
Here is what four current GPU types cost to rent on demand today, from two providers we checked directly:
| GPU | Provider | On-demand price |
|---|---|---|
| NVIDIA H100 SXM | Lambda (8x instance) | $3.99/hr per GPU |
| NVIDIA H100 SXM | RunPod (Secure Cloud) | $3.49/hr |
| NVIDIA A100 SXM 80GB | Lambda (8x instance) | $2.79/hr per GPU |
| NVIDIA A100 SXM | RunPod (Secure Cloud) | $1.59/hr |
| NVIDIA L40S | RunPod (Secure Cloud) | $1.09/hr |
| NVIDIA RTX A6000 | Lambda (4x instance) | $1.09/hr per GPU |
| NVIDIA RTX A6000 | RunPod (Secure Cloud) | $0.53/hr |
(Sources: Lambda GPU Cloud pricing, RunPod GPU Cloud pricing, both checked 2026-09-26.)
Illustrative example: take the RunPod H100 SXM rate, $3.49/hr, and run it around the clock for a month, and 3.49 x 24 x 30 works out to roughly $2,513/month for one GPU, before any of the operational costs the next section covers. Compare that figure against a hosted API's price per output token (our AI agent cost guide sources these from OpenAI's own pricing page) and you get the monthly token volume you would need to send through that API to spend the same amount:
| If the API alternative is... | Output price per 1M tokens | Monthly tokens to match one $2,513/month H100 |
|---|---|---|
| Cheapest tier (gpt-5.6-luna) | $1.20 | ~2.1 billion tokens |
| Mid tier (gpt-5.6-terra) | $12.00 | ~209 million tokens |
| Higher tier (gpt-5.6-sol) | $20.00 | ~126 million tokens |
| Top tier (gpt-6-astra) | $50.00 | ~50 million tokens |
(API prices: OpenAI API pricing, checked 2026-09-26 for our AI agent cost guide and reused here.)
Rerun this with your own GPU choice, your own provider's rate, and your own model's price per token, and you get a real number instead of an estimate. Notice how far the answer moves depending on which API tier you would otherwise pay for: against the cheapest tier, self-hosting rarely breaks even; against the most expensive tier, it can pay for itself at a fraction of the volume. That gap is the actual decision, not a single number either way. Three things push the real answer against self-hosting, so treat these volumes as a floor, not a target. The table counts output tokens only; your API bill also pays for input tokens, which lowers the volume at which the API costs as much as the GPU. One GPU can only generate so many tokens per hour, so the higher volumes may need more than one. And the open-weight model that fits on a single H100 is not a like-for-like swap for a top-tier hosted model, so compare against the API tier whose quality you actually need.
Why the "cheaper" assumption is usually backwards
Renting or buying a GPU is a fixed cost. You pay for every hour it is powered on, whether it is serving one request or a thousand. A hosted API is a variable cost: you pay per token, and nothing more.
If your traffic is low, spiky, or you are still finding product-market fit, a self-hosted GPU spends much of its life waiting for work. That idle time is money already spent producing nothing. A hosted API has no idle time to pay for, because you are only ever billed for what you send it.
There is no single volume number that flips this calculation for every team, because it depends on your traffic pattern, the model size, and how well your serving stack packs requests onto the hardware. What holds in every case is the shape of the trade-off: self-hosting rewards steady, high volume that keeps the GPU busy; a hosted API rewards low or unpredictable volume, because you never pay for capacity you are not using.
Four reasons to self-host even before the cost math
Cost is one axis, not the only one. These are the situations where running your own model is the right call regardless of price:
- Sustained, high, and predictable volume. If you can keep GPUs busy most of the time, owning the serving layer starts to beat renting it token by token.
- Regulated data that cannot leave your network. Healthcare, finance, legal, and government workloads often carry contractual or statutory limits on where data is processed. If the data cannot cross your boundary, a hosted API is off the table regardless of price. This is the same driver behind enterprise security for private LLMs.
- Latency control. Owning the stack means you control placement, batching, and network hops, so you can hold tail latency to a target instead of accepting a shared-tenancy provider's variance.
- Deep fine-tuning. If you need to train on proprietary data and serve custom weights, self-hosting gives you full control over the model artifact and its lifecycle.
What actually changes on compliance. The obligation behind reason 2 does not disappear if you host the model yourself, it just moves. In the US, HIPAA requires a signed, written agreement (a business associate agreement) with any vendor that creates, receives, maintains, or transmits protected health information on a covered entity's behalf (45 CFR 164.502(e)). Under the GDPR, any third party processing personal data on your behalf is a processor, and Article 28 requires a contract setting out the subject matter, duration, nature, and purpose of that processing before you can use them at all (GDPR Article 28). And the EU AI Act, which entered into force on 1 August 2024 and became generally applicable on 2 August 2026, sets obligations for the providers and deployers of AI systems, whichever servers the model happens to run on (European Commission, AI Act regulatory framework). Self-hosting can simplify where the data physically sits, but it does not remove the paperwork: you still need the same agreements and the same audit trail, just with your own team holding them instead of a vendor's.
If none of these apply and your volume is low or unpredictable, the honest answer is usually to use a hosted API and revisit the decision once you have real usage data.
A five-question decision checklist
Turn the four reasons above into a checklist you can run against your own workload before committing either way:
- Volume. Is your monthly token volume steady and high enough to keep a GPU busy most of the time, or is it low or spiky?
- Data sensitivity. Does a contract, a regulator, or your own policy require the data to stay inside your network rather than cross to a third party?
- Latency SLA. Do you need to control tail latency and request placement yourself, or is shared-tenancy variance acceptable?
- Fine-tuning need. Do you need to train on your own data and serve the resulting weights, rather than call someone else's model?
- MLOps capacity. Do you already have, or can you hire, someone who will own GPU patching, driver and CUDA upgrades, health monitoring, and being on call when a node fails?
Answer yes to two or more, especially volume paired with MLOps capacity, and self-hosting is worth pricing out properly. Answer mostly no, and a hosted API is the honest choice until your usage changes.
vLLM vs Ollama: different tools for different stages
A common mistake is picking one serving tool for the whole lifecycle. vLLM and Ollama solve different problems.
Ollama bundles llama.cpp, handles quantization automatically, and exposes an OpenAI-compatible API. It is built for a single user, local development, and Apple Silicon. It is not built to serve many concurrent users under production load.
vLLM is an open-source library for LLM inference and serving, built for production traffic. It exposes an OpenAI-compatible API server (its maintainers have also added Anthropic Messages API and gRPC support) and, per its own project documentation, supports more than 200 model architectures on Hugging Face. Its core design, PagedAttention, manages the GPU memory used for attention keys and values so more concurrent requests fit on the same hardware without running out of memory. That is the mechanism that lets vLLM push a GPU from mostly idle toward heavily loaded, which is the difference between self-hosting that saves money and self-hosting that burns it.
vLLM does demand more setup: NVIDIA hardware, CUDA, and multi-GPU configuration for larger models. Ollama runs almost anywhere with almost no configuration. A workable pattern is to evaluate candidate models in a tool like LM Studio, develop against Ollama for its zero-friction local loop, then deploy on vLLM for production serving. Each tool does the job it is built for, and you never ship a laptop-grade server into a workload that needs real concurrency.
(Source: vLLM project on GitHub.)
The hidden costs nobody puts in the spreadsheet
The GPU bill is the visible number, but it is rarely what sinks a self-hosting project. The costs that get missed are operational: someone has to keep the serving stack patched, monitor GPU health, handle driver and CUDA upgrades, and be on call when a node fails overnight. A hosted API absorbs all of that inside its price. When you self-host, it becomes your team's job, and engineering time is not free.
There is also the cost of capacity you provision but rarely use. Say your traffic peaks well above its average during business hours. You either size hardware for that peak and pay for the idle GPU the rest of the day, or you accept slower responses during spikes. Autoscaling helps, but GPUs are slower and more expensive to spin up than a stateless web server, so the elasticity a hosted API gives you for free is genuinely hard to reproduce on your own infrastructure.
None of this makes self-hosting the wrong call. It means the decision should be sized honestly, with an owner, a budget, and an uptime target, not treated as a side project running on borrowed attention.
Which open-weight models, and under which license
Read the license on the exact weights you plan to deploy, not the license of the model family. Two things changed the open-weight picture in 2026, and a lot of self-hosting advice still describes the version before either change.
Llama is no longer the default answer. Meta has shifted its frontier development toward the closed Muse Spark line, so a current Llama model is best planned for as a maintained option rather than one that will keep getting capability jumps. It still runs, and the license has not changed: the Llama 4 Community License Agreement permits commercial use and self-hosting, with one threshold. Per Meta's own license text, if your product or service had more than 700 million monthly active users in the month before a given Llama 4 release, you must request a separate license from Meta before you may use the model at all. Below that threshold, there is no extra step. (Source: Llama 4 Community License Agreement.)
Both Hugging Face Open LLM Leaderboards have been retired. Version 1 was archived in June 2024, and version 2 was retired on 2025-03-13, with the Hugging Face team pointing people toward specialized community leaderboards rather than a single replacement. If your shortlist process starts by opening the old leaderboard, it now starts on an archived page. (Sources: Hugging Face leaderboard v1 archive, Hugging Face leaderboard v2 retirement notice.)
On the capability question that used to justify paying for a closed API: Epoch AI's Capabilities Index found that, since January 2026, the most capable open-weight models have lagged frontier closed models by an average of four months, with an average gap of eight index points, similar to the gap between GPT-5 and GPT-5.5 (Epoch AI). Four months is real, and small enough that for most business tasks the deciding factors are license, deployment cost, and data residency rather than raw capability. (Source: Epoch AI, "The open-closed ECI gap".)
Here is the current field, with the license first because it is the part you cannot undo once you have built on it:
| Model | License | What to check before you deploy |
|---|---|---|
| DeepSeek V4-Pro / V4-Flash | MIT | Confirmed on the model's own Hugging Face repository. Commercial self-hosting with no usage restrictions. |
| GLM-5.2 | MIT | Zhipu AI's (Z.ai's) current flagship. Full weights on Hugging Face under MIT. |
| Kimi K3 | Custom Moonshot terms | Its predecessor, Kimi K2, shipped under a modified MIT license that only adds a branding requirement above 100 million monthly active users or $20 million in monthly revenue (K2 license). Moonshot's K3 terms are stricter than K2's; read the license file on the model's own repository before you rely on it. |
| Qwen family | Mixed, checkpoint by checkpoint | Some Qwen checkpoints ship under plain Apache 2.0. Others ship under a Qwen Community License that requires a separate commercial license once you cross the license's own usage thresholds. The brand name tells you nothing; the LICENSE file in that specific model's repository does. |
| Mistral Large 3 | Apache 2.0 | Mistral's own model card and release announcement say Apache 2.0. Some of Mistral's older pricing-page language still implies commercial deployments need a separate license, so read the LICENSE file attached to the exact model rather than the marketing copy. |
| Gemma | Custom Gemma Terms of Use, not Apache 2.0 | Google's own terms page makes this explicit: the code wrapper may say Apache 2.0, but the model weights carry a separate Gemma Terms of Use with a prohibited-use policy, a requirement to pass those restrictions on to anyone you distribute to, and a right for Google to restrict usage it judges to violate the policy. Read it before you assume Gemma behaves like a normal open-source license. |
| Meta Muse Glimmer | Apache 2.0 | A 30B open-weight model built for local, on-device agents (InfoQ). Meta shipped the full, unmodified Apache 2.0 license text rather than a custom community license. |
Three things follow from that table:
"Open-weight" does not mean one license. MIT, Apache 2.0, and a range of custom community licenses with their own thresholds and restrictions all sit under the same "open-weight" label. They are not interchangeable, legally.
Check the checkpoint, not the family. Qwen is the clearest case where the brand and the specific model disagree: some Qwen checkpoints are Apache 2.0 and others are not. The same caution applies to Mistral and Gemma, where marketing language and the actual LICENSE file in a given repository can say different things.
Google's Gemma is the one to read most carefully. Despite being widely described as open source, its terms include a prohibited-use policy that legally binds anyone you redistribute the model to, and Google reserves the right to restrict use it judges non-compliant. That is a materially different deal from MIT or Apache 2.0, even though all three get called "open-weight."
If you are still comparing options at the capability level rather than the license level, our guide on how to choose an LLM for your business walks through the trade-offs against hosted APIs as well.
Putting it together
Estimate your real monthly token volume and how steadily you can keep a GPU loaded. If that volume is low or unpredictable, use a hosted API and stop there. Then check the four self-hosting drivers: if regulated data, latency control, or deep fine-tuning apply, self-hosting may be right even at lower volume, but you should expect to pay for the privilege. If you do self-host, serve on vLLM rather than a development tool, and drive utilization up through batching. Finally, pick a model your hardware can actually run, then read the license on the exact weights, not the family name.
Most of the pain in self-hosting comes from provisioning hardware before anyone measured the real load. The open-weight models are capable enough for most business tasks, vLLM is built for production concurrency, and several of the strongest current models ship under genuinely permissive licenses. What decides the outcome is whether the workload keeps the GPUs busy. Self-hosting is also frequently paired with retrieval, so if that is your direction, our walkthrough on how to build a RAG system covers the serving layer these models sit behind.
Sources
- Kimi-K2-Instruct LICENSE (Modified MIT License) -- Moonshot AI, Hugging Face (primary source, checked 2026-09-14)
- Kimi-K3 LICENSE ("Kimi K3 License") -- Moonshot AI, Hugging Face (primary source, checked 2026-09-14)
- Llama 4 Community License Agreement -- Meta (primary source, checked 2026-09-14)
- Open LLM Leaderboard archive | Hugging Face (primary source, checked 2026-09-14)
- Open LLM Leaderboard v2 retirement announcement -- Hugging Face Spaces discussion #1135 (primary source, checked 2026-09-14)
- Open models lag state-of-the-art closed models by 4 months -- Epoch AI (primary source, checked 2026-09-14)
- Meta Open-Sources Muse Glimmer: a 30B Local Agentic Model Optimised for On-Device Execution -- InfoQ (secondary source, checked 2026-09-14)
- Lambda GPU Cloud pricing (primary source, checked 2026-09-26)
- RunPod GPU Cloud pricing (primary source, checked 2026-09-26)
- OpenAI API pricing (primary source, checked 2026-09-26, reused from our AI agent cost guide)
- 45 CFR 164.502(e) -- Cornell Law School Legal Information Institute (secondary source, regulatory-text mirror, checked 2026-09-26)
- Art. 28 GDPR -- gdpr-info.eu (secondary source, regulatory-text mirror, checked 2026-09-26)
- Regulatory framework proposal on artificial intelligence -- European Commission (primary source, checked 2026-09-26)
Frequently Asked Questions
Is self-hosting an LLM cheaper than using a hosted API?+
At what volume does self-hosting an LLM start to pay off?+
Should I use vLLM or Ollama for self-hosting?+
Which open-weight LLMs are safe to use commercially in 2026?+
Is Llama still the model to self-host?+
When does a hosted API make more sense than self-hosting?+
Has open-weight model quality caught up to commercial APIs?+
Which leaderboard should I use to compare open-weight models?+
What's the break-even point for self-hosting vs an API?+
Is self-hosting more secure than a hosted API?+
Weighing self-hosted models against hosted APIs? We scope the architecture and the cost math before you buy a single GPU.
Talk to our AI engineering teamAbout the Author

Rajat Gautam
AI Engineer and Consultant
My work goes far beyond recommending tools - I design AI systems that integrate directly into your workflows, eliminate inefficiencies, and deliver measurable business impact. Every solution I build is tailored, practical, and built with long-term scalability in mind.
Need help with this?
Related Topics
Related Articles



Ready to transform your business with AI? Let's talk strategy.
Book a Free Strategy Call