What open-weight models actually cost, and when running one yourself pays
Six of the thirteen models in our catalogue ship their weights. Their prices span twenty-five times, three of their licences are conditional or bespoke, and only one of them is small enough to run on a single machine. This prices the hardware each one needs against what its own API charges, and finds the case for self-hosting narrower than either side of the argument usually admits: it pays for a small model displacing an expensive one, and almost nowhere else.
Watch
What Open-Weight Models Actually Cost
~13 min spoken. Keeps playing while you work in another tab.
There is a video genre about this. Search for running an open model instead of paying for a coding assistant and you will find hundreds of thousands of them, most demonstrating the same thing: install Ollama, point your tooling at a local model, stop paying the subscription. The arithmetic in those videos is usually right. What almost none of them mention is that the cheapest way to stop paying frontier prices is not to run a model at all — it is to call a different one.
This article is about the gap between those two facts. Prices and scores come from the public catalogue our benchmark articles draw on — Artificial Analysis for prices and speed, Epoch AI and the benchmarks' own leaderboards for scores. Licences and deployment requirements come from each model's own card. Everything is as it stood on 21 September 2026.
"Open weights" is a licence, not a price
The six open-weight models in the catalogue are DeepSeek V4 Pro, Kimi K3, GLM-5.3-Flash, Qwen3.8-27B, MiniMax-M3 and Mistral Medium 3.5. Between the cheapest and the dearest there is a factor of twenty-five in the hosted price, which is a wider spread than between many open and closed models.
The numbers behind this chart
| Model | Access | Score | $/1M tokens |
|---|---|---|---|
| Claude Fable 5.1 | API | 91.4% | $20 |
| GPT-6 Astra | API | 89.9% | $20 |
| GPT-5.6 Sol | API | 89.5% | $8.00 |
| Claude Opus 5 | API | 89.1% | $10 |
| Gemini 3.8 Flash | API | 87.6% | $1.50 |
| Kimi K3 | Open weights | 85.0% | $6.00 |
| GLM-5.3-Flash | Open weights | 84.3% | $0.24 |
| Claude Sonnet 5 | API | 80.5% | $4.00 |
| DeepSeek V4 Pro 0813 | Open weights | 78.7% | $1.98 |
| MiniMax-M3 | Open weights | 65.2% | $0.53 |
| Mistral Medium 3.5 | Open weights | 50.6% | $3.00 |
| Claude Haiku 4.5 | API | 44.2% | $2.00 |
Two things in that picture are worth stopping on.
GLM-5.3-Flash is where the money is. At $0.24 per million tokens blended it scores 84.3 on Terminal-Bench 2.1 — against Claude Opus 5's 89.1 at $10. That is 95% of the score at a forty-second of the price. Z.ai gets there by activating 18 billion of its 320 billion parameters on each token, which is also why it is cheap to serve.
Kimi K3 costs more than Claude Sonnet 5. $6 blended against $4, for 85.0 on Terminal-Bench against Sonnet's 80.5. It is a better model on this task and it is not a cheaper one. Mistral Medium 3.5 makes the point less flatteringly still: at $3 blended it is the second-dearest open model here and the lowest-scoring of any model on the chart bar one, below DeepSeek V4 Pro at $1.98 and far below GLM at a twelfth of its price. Neither is an argument against those models on their merits. They are an argument against reading "open weights" as a synonym for "cheap".
There is a thirteenth model in the catalogue that is not in that chart, and its absence is the most useful thing on the page. Qwen3.8-27B has no price and no measured speed, because Artificial Analysis measures hosted endpoints and nobody is selling Qwen3.8-27B by the token in volume. The one model here you would genuinely self-host is the one the price-performance leaderboard cannot see. Hold that thought.
What a GPU has to beat
Renting the machine is the part these comparisons usually skip. Price it properly: RunPod's on-demand H100 80GB at $2.89 an hour, which is $2,111 a month running continuously. There are cheaper cards — the same price list has an RTX 4090 at $0.74 — but a consumer card running a heavily quantised model is a hobbyist rig, and the question here is what production costs.
The numbers behind this chart
| Model | Weights | $/1M blended | Tokens for $2,111 | Break-even |
|---|---|---|---|---|
| Claude Fable 5.1 | API only | $20 | 106M | 40 tok/s |
| GPT-6 Astra | API only | $20 | 106M | 40 tok/s |
| Claude Opus 5 | API only | $10 | 211M | 80 tok/s |
| GPT-5.6 Sol | API only | $8.00 | 264M | 100 tok/s |
| Kimi K3 | Open | $6.00 | 352M | 134 tok/s |
| Claude Sonnet 5 | API only | $4.00 | 528M | 201 tok/s |
| Mistral Medium 3.5 | Open | $3.00 | 704M | 268 tok/s |
| Claude Haiku 4.5 | API only | $2.00 | 1,056M | 401 tok/s |
| DeepSeek V4 Pro 0813 | Open | $1.98 | 1,066M | 405 tok/s |
| Gemini 3.8 Flash | API only | $1.50 | 1,408M | 535 tok/s |
| MiniMax-M3 | Open | $0.53 | 4,022M | 1,529 tok/s |
| GLM-5.3-Flash | Open | $0.24 | 8,909M | 3,387 tok/s |
Read the right-hand column. It is the throughput a self-hosted model would have to sustain, every second of the month, for one H100 to be worth what the tokens would have cost.
- Against Claude Fable 5.1 at $20 blended, break-even is 40 tokens a second.
- Against Claude Opus 5 at $10, break-even is 80 tokens a second. A 27B model on an H100 with batching clears that.
- Against GLM-5.3-Flash at $0.24, break-even is 3,387 tokens a second. Nothing you can put on one card approaches it.
So the finding is not that self-hosting is a myth. It is that self-hosting competes with expensive models, not with cheap ones — and the thing that makes it pay is never that the weights were free. It is that the model you would otherwise have rented was dear.
The videos telling people to run a local model rather than pay frontier prices are, on this arithmetic, correct. They are also solving a problem that a $0.24 API call solves more cheaply, with no GPU, no quantisation and no evenings lost to a serving stack.
And the break-even above is a floor rather than an estimate. It assumes the card runs every hour of the month; a machine used only in office hours needs roughly three times the throughput to justify itself. It ignores storage, egress, and the engineer-hours that keeping a model served will quietly consume.
One card, if you are lucky
Everything above assumed one H100. That assumption holds for exactly one of the six models.
The numbers behind this chart
| Model | Total parameters | Weights at 8-bit | H100 80GBs | Cost a month |
|---|---|---|---|---|
| Qwen3.8-27B | 27B | 27 GB | 1 | $2,111 |
| Mistral Medium 3.5 | 128B | 128 GB | 2 | $4,223 |
| GLM-5.3-Flash | 320B | 320 GB | 4 | $8,445 |
| MiniMax-M3 | 428B | 428 GB | 6 | $12,668 |
| Kimi K3 | 2800B | 2800 GB | 35 | $73,896 |
| DeepSeek V4 Pro | vLLM on a single 4xGB300 node, the card's own example | - | - | - |
A model has to keep all of its weights in memory to serve any of them, and that is true of a sparse model too: GLM-5.3-Flash activates 18 billion parameters per token, but all 320 billion have to be resident. At 8-bit — one byte per parameter, which is a floor that ignores the KV cache entirely — that is four H100s before a single request arrives.
Run the break-even again with the right number of cards and the case collapses:
| Model | Cards | Hardware a month | Break-even against its own API |
|---|---|---|---|
| Qwen3.8-27B | 1 | $2,111 | — (no first-party API) |
| Mistral Medium 3.5 | 2 | $4,223 | 535 tok/s |
| GLM-5.3-Flash | 4 | $8,445 | 13,549 tok/s |
| MiniMax-M3 | 6 | $12,668 | 9,175 tok/s |
| Kimi K3 | 35 | $73,896 | 4,683 tok/s |
Which produces the sharpest version of the finding in this article: self-hosting an open model to avoid that same model's API essentially never pays. The vendor serving it has better hardware utilisation than you will, and it is passing some of that on. GLM would need four cards and thirteen thousand tokens a second to beat a service that charges $0.24 a million.
DeepSeek V4 Pro does not appear in that table because its card publishes no parameter count. It publishes something more useful: its own serving example is a four-GB300 node. That is the vendor telling you what this costs.
So the option that survives is narrow and specific. It is not "run open models". It is run a small model on hardware you already need, to displace an expensive one — and Qwen3.8-27B, at 27 billion dense parameters on a single card, is the only model here that qualifies.
What you are actually buying
If the money argument only works in that one narrow case, why take open weights at all? Three reasons survive the arithmetic, and none of them is the sticker price.
Nobody can reprice you. This is the strongest one, and the catalogue dates it precisely. Gemini 3.8 Flash's $0.75 and $3.75 are introductory and double on 1 January 2027. GPT-5.6 Sol's $4 and $20 are promotional and published only as far as 21 November 2026. DeepSeek's rates double during weekday peak hours. MiniMax labels its 50% discount permanent, which is a promise and not a mechanism. A downloaded set of weights has no renewal date.
Nobody can retire it. A hosted model is withdrawn when its vendor decides; a local one runs until you stop it. For anything that has to behave identically in two years — a regulated workflow, a reproducible evaluation, a product whose output users have already accepted — that is worth more than a price break.
The data stays where you put it. The reason most often given and the one least affected by any of the numbers above. If the work cannot leave your network, the API price is not a competing option and the break-even calculation does not apply.
The licence is not the same for all six
"Open weights" is doing a lot of work as a phrase. Under it sit four different legal regimes.
The numbers behind this chart
| Model | Licence | What it asks of a commercial user |
|---|---|---|
| DeepSeek V4 Pro | MIT | Nothing. The card states the repository and the weights are licensed under the MIT License. |
| GLM-5.3-Flash | MIT | Nothing. MIT in the model card frontmatter. |
| Qwen3.8-27B | Apache 2.0 | Nothing, and an explicit patent grant on top. |
| Kimi K3 | Kimi K3 License | A custom licence covering both the code and the weights. Not MIT or Apache, so its terms have to be read rather than assumed. |
| MiniMax-M3 | MiniMax Community | A community licence rather than a standard open-source one, with its own commercial conditions. |
| Mistral Medium 3.5 | Modified MIT | Commercial and non-commercial use "with exceptions for companies with large revenue". The card publishes no threshold, so a large company cannot tell from it whether it qualifies. |
Three of the six ask nothing about who is deploying them. DeepSeek V4 Pro and GLM-5.3-Flash are MIT; Qwen3.8-27B is Apache 2.0 with an explicit patent grant, the least encumbered terms here.
The other three are conditional in ways worth knowing before a procurement conversation. Mistral Medium 3.5 is a Modified MIT licence permitting commercial use "with exceptions for companies with large revenue" — and the model card publishes no threshold, so a large company cannot tell from the card whether it qualifies. Kimi K3 ships under a bespoke Kimi K3 License, and MiniMax-M3 under a MiniMax Community licence. Both are readable, neither is standard, and neither can be assumed from the phrase "open weights" on a launch post.
Only one of the three turns explicitly on size, and it is the one that publishes no figure. The other two are bespoke for everybody: a startup and a bank read the same licence, and both have to read it.
What to do with this
If you are paying frontier prices for bulk work, the first move is not a GPU. Try GLM-5.3-Flash. It is 95% of Opus 5 on Terminal-Bench at a forty-second of the price, and the change is an API endpoint rather than a project.
If you are choosing an open model for the licence, read it. Three of the six are unconditional. The other three are not, and the only one with an explicit revenue exception publishes no figure for it.
If you are self-hosting to save money, do the sum in the second and third charts together. Take the blended price of the model you would otherwise call — not the frontier model, the cheapest one that does the job — then count the cards the model you want to host actually needs. One H100 against Claude Opus 5 is 80 tokens a second and worth a look. Four H100s against a $0.24 API is thirteen thousand, and is not.
If you are self-hosting for data residency, reproducibility or protection from a price change, none of the above applies. Those are the reasons open weights are worth having, and they were never about the sticker price.
One thing this article cannot tell you: what your hardware actually does. No model here has a measured self-hosted throughput in our catalogue, because throughput is a property of your machine and your serving stack rather than of the model. The break-even column is the number to measure against, and measuring it is the one step nobody can take for you.