Private AI & Security

Self-Hosting an LLM: What It Costs and When It Makes Sense

Rajat Gautam••15 min read•Updated
Share

Key Takeaways

  • →Self-hosting is a fixed cost: the GPU bills you whether it serves ten requests a minute or sits mostly idle, while a hosted API charges per token.
  • →Self-hosting wins on sustained predictable volume, regulated data, latency control, and deep fine-tuning, not on price alone.
  • →Use Ollama for local development and vLLM for production, since PagedAttention fits more concurrent requests on the same GPU.
  • →Llama is no longer the default open-weight pick, and both Hugging Face Open LLM Leaderboards have been retired.
  • →Read the license on the exact checkpoint: MIT, Apache 2.0 and custom community licenses are not interchangeable.
Self-Hosting an LLM: What It Costs and When It Makes Sense

Self-hosting an LLM means running an open-weight model on hardware you control (your own servers or rented GPUs) instead of paying a provider per API call. It is usually not the cheap option. A GPU costs the same whether it answers ten requests a minute or sits mostly idle, while a hosted API only charges for tokens you actually use. So the real question is not "is self-hosting cheaper," it is "can I keep the hardware busy enough to make owning it worthwhile."

This guide walks through when self-hosting wins, when a hosted API is the smarter call, how to pick a serving stack, and which open-weight models you can legally run in production today.

Self-hosted vs API: the break-even math

Renting a GPU has a real, current price. So does sending the same workload to a hosted API. You can compare them directly if you run the arithmetic yourself, on your own volume, instead of trusting someone else's fixed threshold.

Here is what four current GPU types cost to rent on demand today, from two providers we checked directly:

GPUProviderOn-demand price
NVIDIA H100 SXMLambda (8x instance)$3.99/hr per GPU
NVIDIA H100 SXMRunPod (Secure Cloud)$3.49/hr
NVIDIA A100 SXM 80GBLambda (8x instance)$2.79/hr per GPU
NVIDIA A100 SXMRunPod (Secure Cloud)$1.59/hr
NVIDIA L40SRunPod (Secure Cloud)$1.09/hr
NVIDIA RTX A6000Lambda (4x instance)$1.09/hr per GPU
NVIDIA RTX A6000RunPod (Secure Cloud)$0.53/hr

(Sources: Lambda GPU Cloud pricing, RunPod GPU Cloud pricing, both checked 2026-09-26.)

Illustrative example: take the RunPod H100 SXM rate, $3.49/hr, and run it around the clock for a month, and 3.49 x 24 x 30 works out to roughly $2,513/month for one GPU, before any of the operational costs the next section covers. Compare that figure against a hosted API's price per output token (our AI agent cost guide sources these from OpenAI's own pricing page) and you get the monthly token volume you would need to send through that API to spend the same amount:

If the API alternative is...Output price per 1M tokensMonthly tokens to match one $2,513/month H100
Cheapest tier (gpt-5.6-luna)$1.20~2.1 billion tokens
Mid tier (gpt-5.6-terra)$12.00~209 million tokens
Higher tier (gpt-5.6-sol)$20.00~126 million tokens
Top tier (gpt-6-astra)$50.00~50 million tokens

(API prices: OpenAI API pricing, checked 2026-09-26 for our AI agent cost guide and reused here.)

Rerun this with your own GPU choice, your own provider's rate, and your own model's price per token, and you get a real number instead of an estimate. Notice how far the answer moves depending on which API tier you would otherwise pay for: against the cheapest tier, self-hosting rarely breaks even; against the most expensive tier, it can pay for itself at a fraction of the volume. That gap is the actual decision, not a single number either way. Three things push the real answer against self-hosting, so treat these volumes as a floor, not a target. The table counts output tokens only; your API bill also pays for input tokens, which lowers the volume at which the API costs as much as the GPU. One GPU can only generate so many tokens per hour, so the higher volumes may need more than one. And the open-weight model that fits on a single H100 is not a like-for-like swap for a top-tier hosted model, so compare against the API tier whose quality you actually need.

Why the "cheaper" assumption is usually backwards

Renting or buying a GPU is a fixed cost. You pay for every hour it is powered on, whether it is serving one request or a thousand. A hosted API is a variable cost: you pay per token, and nothing more.

If your traffic is low, spiky, or you are still finding product-market fit, a self-hosted GPU spends much of its life waiting for work. That idle time is money already spent producing nothing. A hosted API has no idle time to pay for, because you are only ever billed for what you send it.

There is no single volume number that flips this calculation for every team, because it depends on your traffic pattern, the model size, and how well your serving stack packs requests onto the hardware. What holds in every case is the shape of the trade-off: self-hosting rewards steady, high volume that keeps the GPU busy; a hosted API rewards low or unpredictable volume, because you never pay for capacity you are not using.

Four reasons to self-host even before the cost math

Cost is one axis, not the only one. These are the situations where running your own model is the right call regardless of price:

  1. Sustained, high, and predictable volume. If you can keep GPUs busy most of the time, owning the serving layer starts to beat renting it token by token.
  2. Regulated data that cannot leave your network. Healthcare, finance, legal, and government workloads often carry contractual or statutory limits on where data is processed. If the data cannot cross your boundary, a hosted API is off the table regardless of price. This is the same driver behind enterprise security for private LLMs.
  3. Latency control. Owning the stack means you control placement, batching, and network hops, so you can hold tail latency to a target instead of accepting a shared-tenancy provider's variance.
  4. Deep fine-tuning. If you need to train on proprietary data and serve custom weights, self-hosting gives you full control over the model artifact and its lifecycle.

What actually changes on compliance. The obligation behind reason 2 does not disappear if you host the model yourself, it just moves. In the US, HIPAA requires a signed, written agreement (a business associate agreement) with any vendor that creates, receives, maintains, or transmits protected health information on a covered entity's behalf (45 CFR 164.502(e)). Under the GDPR, any third party processing personal data on your behalf is a processor, and Article 28 requires a contract setting out the subject matter, duration, nature, and purpose of that processing before you can use them at all (GDPR Article 28). And the EU AI Act, which entered into force on 1 August 2024 and became generally applicable on 2 August 2026, sets obligations for the providers and deployers of AI systems, whichever servers the model happens to run on (European Commission, AI Act regulatory framework). Self-hosting can simplify where the data physically sits, but it does not remove the paperwork: you still need the same agreements and the same audit trail, just with your own team holding them instead of a vendor's.

If none of these apply and your volume is low or unpredictable, the honest answer is usually to use a hosted API and revisit the decision once you have real usage data.

A five-question decision checklist

Turn the four reasons above into a checklist you can run against your own workload before committing either way:

  1. Volume. Is your monthly token volume steady and high enough to keep a GPU busy most of the time, or is it low or spiky?
  2. Data sensitivity. Does a contract, a regulator, or your own policy require the data to stay inside your network rather than cross to a third party?
  3. Latency SLA. Do you need to control tail latency and request placement yourself, or is shared-tenancy variance acceptable?
  4. Fine-tuning need. Do you need to train on your own data and serve the resulting weights, rather than call someone else's model?
  5. MLOps capacity. Do you already have, or can you hire, someone who will own GPU patching, driver and CUDA upgrades, health monitoring, and being on call when a node fails?

Answer yes to two or more, especially volume paired with MLOps capacity, and self-hosting is worth pricing out properly. Answer mostly no, and a hosted API is the honest choice until your usage changes.

vLLM vs Ollama: different tools for different stages

A common mistake is picking one serving tool for the whole lifecycle. vLLM and Ollama solve different problems.

Ollama bundles llama.cpp, handles quantization automatically, and exposes an OpenAI-compatible API. It is built for a single user, local development, and Apple Silicon. It is not built to serve many concurrent users under production load.

vLLM is an open-source library for LLM inference and serving, built for production traffic. It exposes an OpenAI-compatible API server (its maintainers have also added Anthropic Messages API and gRPC support) and, per its own project documentation, supports more than 200 model architectures on Hugging Face. Its core design, PagedAttention, manages the GPU memory used for attention keys and values so more concurrent requests fit on the same hardware without running out of memory. That is the mechanism that lets vLLM push a GPU from mostly idle toward heavily loaded, which is the difference between self-hosting that saves money and self-hosting that burns it.

vLLM does demand more setup: NVIDIA hardware, CUDA, and multi-GPU configuration for larger models. Ollama runs almost anywhere with almost no configuration. A workable pattern is to evaluate candidate models in a tool like LM Studio, develop against Ollama for its zero-friction local loop, then deploy on vLLM for production serving. Each tool does the job it is built for, and you never ship a laptop-grade server into a workload that needs real concurrency.

(Source: vLLM project on GitHub.)

The hidden costs nobody puts in the spreadsheet

The GPU bill is the visible number, but it is rarely what sinks a self-hosting project. The costs that get missed are operational: someone has to keep the serving stack patched, monitor GPU health, handle driver and CUDA upgrades, and be on call when a node fails overnight. A hosted API absorbs all of that inside its price. When you self-host, it becomes your team's job, and engineering time is not free.

There is also the cost of capacity you provision but rarely use. Say your traffic peaks well above its average during business hours. You either size hardware for that peak and pay for the idle GPU the rest of the day, or you accept slower responses during spikes. Autoscaling helps, but GPUs are slower and more expensive to spin up than a stateless web server, so the elasticity a hosted API gives you for free is genuinely hard to reproduce on your own infrastructure.

None of this makes self-hosting the wrong call. It means the decision should be sized honestly, with an owner, a budget, and an uptime target, not treated as a side project running on borrowed attention.

Which open-weight models, and under which license

Read the license on the exact weights you plan to deploy, not the license of the model family. Two things changed the open-weight picture in 2026, and a lot of self-hosting advice still describes the version before either change.

Llama is no longer the default answer. Meta has shifted its frontier development toward the closed Muse Spark line, so a current Llama model is best planned for as a maintained option rather than one that will keep getting capability jumps. It still runs, and the license has not changed: the Llama 4 Community License Agreement permits commercial use and self-hosting, with one threshold. Per Meta's own license text, if your product or service had more than 700 million monthly active users in the month before a given Llama 4 release, you must request a separate license from Meta before you may use the model at all. Below that threshold, there is no extra step. (Source: Llama 4 Community License Agreement.)

Both Hugging Face Open LLM Leaderboards have been retired. Version 1 was archived in June 2024, and version 2 was retired on 2025-03-13, with the Hugging Face team pointing people toward specialized community leaderboards rather than a single replacement. If your shortlist process starts by opening the old leaderboard, it now starts on an archived page. (Sources: Hugging Face leaderboard v1 archive, Hugging Face leaderboard v2 retirement notice.)

On the capability question that used to justify paying for a closed API: Epoch AI's Capabilities Index found that, since January 2026, the most capable open-weight models have lagged frontier closed models by an average of four months, with an average gap of eight index points, similar to the gap between GPT-5 and GPT-5.5 (Epoch AI). Four months is real, and small enough that for most business tasks the deciding factors are license, deployment cost, and data residency rather than raw capability. (Source: Epoch AI, "The open-closed ECI gap".)

Here is the current field, with the license first because it is the part you cannot undo once you have built on it:

ModelLicenseWhat to check before you deploy
DeepSeek V4-Pro / V4-FlashMITConfirmed on the model's own Hugging Face repository. Commercial self-hosting with no usage restrictions.
GLM-5.2MITZhipu AI's (Z.ai's) current flagship. Full weights on Hugging Face under MIT.
Kimi K3Custom Moonshot termsIts predecessor, Kimi K2, shipped under a modified MIT license that only adds a branding requirement above 100 million monthly active users or $20 million in monthly revenue (K2 license). Moonshot's K3 terms are stricter than K2's; read the license file on the model's own repository before you rely on it.
Qwen familyMixed, checkpoint by checkpointSome Qwen checkpoints ship under plain Apache 2.0. Others ship under a Qwen Community License that requires a separate commercial license once you cross the license's own usage thresholds. The brand name tells you nothing; the LICENSE file in that specific model's repository does.
Mistral Large 3Apache 2.0Mistral's own model card and release announcement say Apache 2.0. Some of Mistral's older pricing-page language still implies commercial deployments need a separate license, so read the LICENSE file attached to the exact model rather than the marketing copy.
GemmaCustom Gemma Terms of Use, not Apache 2.0Google's own terms page makes this explicit: the code wrapper may say Apache 2.0, but the model weights carry a separate Gemma Terms of Use with a prohibited-use policy, a requirement to pass those restrictions on to anyone you distribute to, and a right for Google to restrict usage it judges to violate the policy. Read it before you assume Gemma behaves like a normal open-source license.
Meta Muse GlimmerApache 2.0A 30B open-weight model built for local, on-device agents (InfoQ). Meta shipped the full, unmodified Apache 2.0 license text rather than a custom community license.

Three things follow from that table:

"Open-weight" does not mean one license. MIT, Apache 2.0, and a range of custom community licenses with their own thresholds and restrictions all sit under the same "open-weight" label. They are not interchangeable, legally.

Check the checkpoint, not the family. Qwen is the clearest case where the brand and the specific model disagree: some Qwen checkpoints are Apache 2.0 and others are not. The same caution applies to Mistral and Gemma, where marketing language and the actual LICENSE file in a given repository can say different things.

Google's Gemma is the one to read most carefully. Despite being widely described as open source, its terms include a prohibited-use policy that legally binds anyone you redistribute the model to, and Google reserves the right to restrict use it judges non-compliant. That is a materially different deal from MIT or Apache 2.0, even though all three get called "open-weight."

If you are still comparing options at the capability level rather than the license level, our guide on how to choose an LLM for your business walks through the trade-offs against hosted APIs as well.

Putting it together

Estimate your real monthly token volume and how steadily you can keep a GPU loaded. If that volume is low or unpredictable, use a hosted API and stop there. Then check the four self-hosting drivers: if regulated data, latency control, or deep fine-tuning apply, self-hosting may be right even at lower volume, but you should expect to pay for the privilege. If you do self-host, serve on vLLM rather than a development tool, and drive utilization up through batching. Finally, pick a model your hardware can actually run, then read the license on the exact weights, not the family name.

Most of the pain in self-hosting comes from provisioning hardware before anyone measured the real load. The open-weight models are capable enough for most business tasks, vLLM is built for production concurrency, and several of the strongest current models ship under genuinely permissive licenses. What decides the outcome is whether the workload keeps the GPUs busy. Self-hosting is also frequently paired with retrieval, so if that is your direction, our walkthrough on how to build a RAG system covers the serving layer these models sit behind.

Sources

Frequently Asked Questions

Is self-hosting an LLM cheaper than using a hosted API?+
Usually not. A GPU is a fixed cost you pay for every hour it is powered on, while a hosted API bills only for the tokens you use, so an idle self-hosted GPU spends money producing nothing. Self-hosting can still be the right call, but for reasons other than price.
At what volume does self-hosting an LLM start to pay off?+
There is no single volume number that flips the calculation for every team, because it depends on your traffic pattern, the model size, and how well your serving stack packs requests onto the hardware. The shape of the trade-off holds: self-hosting rewards steady, high volume that keeps the GPU busy, and a hosted API rewards low or unpredictable volume.
Should I use vLLM or Ollama for self-hosting?+
They solve different problems. Ollama bundles llama.cpp, handles quantization, and suits a single user, local development, and Apple Silicon, but it is not built to serve many concurrent users under production load. vLLM is built for production traffic, supports more than 200 model architectures, and uses PagedAttention to fit more concurrent requests on the same hardware. A workable pattern is to develop against Ollama and deploy on vLLM.
Which open-weight LLMs are safe to use commercially in 2026?+
It depends on the exact checkpoint, not the family. DeepSeek V4-Pro and V4-Flash and GLM-5.2 ship under MIT, Mistral Large 3 and Meta Muse Glimmer under Apache 2.0, while Kimi K3, the Qwen family, and Gemma carry custom terms with their own thresholds. Read the LICENSE file in the specific model's repository before you build on it.
Is Llama still the model to self-host?+
It still runs, but it is no longer the default answer. Meta has shifted its frontier development toward the closed Muse Spark line, so a current Llama model is best planned for as a maintained option rather than one that keeps getting capability jumps. The Llama 4 Community License Agreement permits commercial use and self-hosting, with a separate license required only above 700 million monthly active users.
When does a hosted API make more sense than self-hosting?+
When your volume is low or unpredictable, or you are still finding product-market fit. A self-hosted GPU spends much of its life waiting for work, and that idle time is money already spent, while a hosted API never bills you for capacity you are not using. Use a hosted API in that case and revisit once you have real usage data.
Has open-weight model quality caught up to commercial APIs?+
It has narrowed to a small gap. Epoch AI's Capabilities Index found that since January 2026 the most capable open-weight models have lagged frontier closed models by an average of four months, with an average gap of eight index points. Four months is real, but small enough that for most business tasks license, deployment cost, and data residency matter more than raw capability.
Which leaderboard should I use to compare open-weight models?+
Not the old Hugging Face Open LLM Leaderboards, since both have been retired. Version 1 was archived in June 2024 and version 2 was retired on 2025-03-13. Hugging Face points people toward specialized community leaderboards rather than a single replacement.
What's the break-even point for self-hosting vs an API?+
There is no single number, because it depends on your GPU's rental price and the API's price per token, and both change over time. The formula is: divide your GPU's fixed monthly cost by the API's price per million tokens, and that gives you the token volume, in millions, where the two options cost the same. The break-even section above works this out for one real GPU rate against four real API prices; rerun it with your own GPU choice and your own model's price to get your own number.
Is self-hosting more secure than a hosted API?+
Neither is inherently safer. What changes is who is responsible for the safety layer and where the data physically sits. Under the GDPR, a hosted API provider processing your data is a processor and needs a contract under Article 28 regardless of where its servers are. Under HIPAA, any vendor touching protected health information needs a signed business associate agreement whether it is a hosted API or a GPU host you rented yourself. Self-hosting moves the operational security work onto your own team instead of the provider's; it does not exempt you from the underlying legal requirements.

Weighing self-hosted models against hosted APIs? We scope the architecture and the cost math before you buy a single GPU.

Talk to our AI engineering team

About the Author

Rajat Gautam

Rajat Gautam

AI Engineer and Consultant

My work goes far beyond recommending tools - I design AI systems that integrate directly into your workflows, eliminate inefficiencies, and deliver measurable business impact. Every solution I build is tailored, practical, and built with long-term scalability in mind.

Need help with this?

Related Topics

self-hosting
open-weight LLMs
vLLM
LLM cost
private AI
GPU utilization

Related Articles

Ready to transform your business with AI? Let's talk strategy.

Book a Free Strategy Call