Editorial illustration for What open-weight models actually cost, and when running one yourself pays
AI analysis / Latest briefings

What open-weight models actually cost, and when running one yourself pays

Six of the thirteen models in our catalogue ship their weights. Their prices span twenty-five times, three of their licences are conditional or bespoke, and only one of them is small enough to run on a single machine. This prices the hardware each one needs against what its own API charges, and finds the case for self-hosting narrower than either side of the argument usually admits: it pays for a small model displacing an expensive one, and almost nowhere else.

By TerraNet Technologies13 min read22 sources
Editorial illustration for What open-weight models actually cost, and when running one yourself pays
open weight models
open source llm cost
self hosting llm
llm cost optimization
glm 5.3 flash
open weights licence

Watch

What Open-Weight Models Actually Cost

Listen to this article

~13 min spoken. Keeps playing while you work in another tab.

There is a video genre about this. Search for running an open model instead of paying for a coding assistant and you will find hundreds of thousands of them, most demonstrating the same thing: install Ollama, point your tooling at a local model, stop paying the subscription. The arithmetic in those videos is usually right. What almost none of them mention is that the cheapest way to stop paying frontier prices is not to run a model at all — it is to call a different one.

This article is about the gap between those two facts. Prices and scores come from the public catalogue our benchmark articles draw on — Artificial Analysis for prices and speed, Epoch AI and the benchmarks' own leaderboards for scores. Licences and deployment requirements come from each model's own card. Everything is as it stood on 21 September 2026.

"Open weights" is a licence, not a price

The six open-weight models in the catalogue are DeepSeek V4 Pro, Kimi K3, GLM-5.3-Flash, Qwen3.8-27B, MiniMax-M3 and Mistral Medium 3.5. Between the cheapest and the dearest there is a factor of twenty-five in the hosted price, which is a wider spread than between many open and closed models.

Terminal-Bench 2.1 against price per million tokens12 models plotted by Terminal-Bench 2.1 score against price per million tokens.Terminal-Bench 2.1 against price per million tokensAPI accessOpen weights0%25%50%75%100%$0.30$1.00$3.00$10$30price per 1M tokens, log scaleClaude Fable 5.1 · 91.4% · $20/1MClaude Haiku 4.5 · 44.2% · $2.00/1MClaude Opus 5 · 89.1% · $10/1MClaude Sonnet 5 · 80.5% · $4.00/1MDeepSeek V4 Pro 0813 · 78.7% · $1.98/1MGemini 3.8 Flash · 87.6% · $1.50/1MGLM-5.3-Flash · 84.3% · $0.24/1MGPT-5.6 Sol · 89.5% · $8.00/1MGPT-6 Astra · 89.9% · $20/1MKimi K3 · 85.0% · $6.00/1MMiniMax-M3 · 65.2% · $0.53/1MMistral Medium 3.5 · 50.6% · $3.00/1MClaude Fable 5.1Claude Haiku 4.5Claude Opus 5Claude Sonnet 5DeepSeek V4 Pro 0813Gemini 3.8 FlashGLM-5.3-FlashGPT-5.6 SolGPT-6 AstraKimi K3MiniMax-M3Mistral Medium 3.5Source: Artificial Analysis, https://artificialanalysis.ai/ Licence Free API, attribution required. Retrieved2026-09-05.Price is the blended cost per million tokens at 3:1 input to output, from Artificial Analysis(https://artificialanalysis.ai/), attribution required.TerraNet Technologies · terranettechnologies.com

The numbers behind this chart

Model Access Score $/1M tokens
Claude Fable 5.1 API 91.4% $20
GPT-6 Astra API 89.9% $20
GPT-5.6 Sol API 89.5% $8.00
Claude Opus 5 API 89.1% $10
Gemini 3.8 Flash API 87.6% $1.50
Kimi K3 Open weights 85.0% $6.00
GLM-5.3-Flash Open weights 84.3% $0.24
Claude Sonnet 5 API 80.5% $4.00
DeepSeek V4 Pro 0813 Open weights 78.7% $1.98
MiniMax-M3 Open weights 65.2% $0.53
Mistral Medium 3.5 Open weights 50.6% $3.00
Claude Haiku 4.5 API 44.2% $2.00

Two things in that picture are worth stopping on.

GLM-5.3-Flash is where the money is. At $0.24 per million tokens blended it scores 84.3 on Terminal-Bench 2.1 — against Claude Opus 5's 89.1 at $10. That is 95% of the score at a forty-second of the price. Z.ai gets there by activating 18 billion of its 320 billion parameters on each token, which is also why it is cheap to serve.

Kimi K3 costs more than Claude Sonnet 5. $6 blended against $4, for 85.0 on Terminal-Bench against Sonnet's 80.5. It is a better model on this task and it is not a cheaper one. Mistral Medium 3.5 makes the point less flatteringly still: at $3 blended it is the second-dearest open model here and the lowest-scoring of any model on the chart bar one, below DeepSeek V4 Pro at $1.98 and far below GLM at a twelfth of its price. Neither is an argument against those models on their merits. They are an argument against reading "open weights" as a synonym for "cheap".

There is a thirteenth model in the catalogue that is not in that chart, and its absence is the most useful thing on the page. Qwen3.8-27B has no price and no measured speed, because Artificial Analysis measures hosted endpoints and nobody is selling Qwen3.8-27B by the token in volume. The one model here you would genuinely self-host is the one the price-performance leaderboard cannot see. Hold that thought.

What a GPU has to beat

Renting the machine is the part these comparisons usually skip. Price it properly: RunPod's on-demand H100 80GB at $2.89 an hour, which is $2,111 a month running continuously. There are cheaper cards — the same price list has an RTX 4090 at $0.74 — but a consumer card running a heavily quantised model is a hobbyist rig, and the question here is what production costs.

What $2,111 of rented GPU buys instead, as tokensFor each model, the tokens $2,111 a month would buy on its API instead of renting one H100 80GB, and the sustained throughput needed to break even: Claude Fable 5.1 106 million tokens, 40 tokens per second; GPT-6 Astra 106 million tokens, 40 tokens per second; Claude Opus 5 211 million tokens, 80 tokens per second; GPT-5.6 Sol 264 million tokens, 100 tokens per second; Kimi K3 352 million tokens, 134 tokens per second; Claude Sonnet 5 528 million tokens, 201 tokens per second; Mistral Medium 3.5 704 million tokens, 268 tokens per second; Claude Haiku 4.5 1,056 million tokens, 401 tokens per second; DeepSeek V4 Pro 0813 1,066 million tokens, 405 tokens per second; Gemini 3.8 Flash 1,408 million tokens, 535 tokens per second; MiniMax-M3 4,022 million tokens, 1529 tokens per second; GLM-5.3-Flash 8,909 million tokens, 3387 tokens per second.What $2,111 of rented GPU buys instead, as tokensOpen weightsAPI access only100M300M1B3B10BClaude Fable 5.140 tok/sGPT-6 Astra40 tok/sClaude Opus 580 tok/sGPT-5.6 Sol100 tok/sKimi K3134 tok/sClaude Sonnet 5201 tok/sMistral Medium 3.5268 tok/sClaude Haiku 4.5401 tok/sDeepSeek V4 Pro 0813405 tok/sGemini 3.8 Flash535 tok/sMiniMax-M31,529 tok/sGLM-5.3-Flash3,387 tok/stokens the same money buys on the API, log scalebreak-evenA H100 80GB on RunPod at $2.89 an hour is $2,111 a month running continuously. Each bar is what that same moneybuys as tokens on the model's own API, at its blended price per million tokens at 3:1 input to output.Break-even is the throughput a self-hosted copy would have to sustain, every second of the month, to be worththe machine. It is a floor and not a forecast: it ignores idle hours, storage, egress and the engineer-hoursthat running a model costs, all of which push the real figure higher.Prices from Artificial Analysis; GPU rate from the provider's own published on-demand pricing.TerraNet Technologies · terranettechnologies.com
Break-even assumes the GPU runs every hour of the month. A machine used only in office hours needs roughly three times the throughput to justify itself. No model here has a measured self-hosted throughput in the catalogue, because throughput is a property of your hardware rather than of the model. The break-even column is the number to measure against.

The numbers behind this chart

Model Weights $/1M blended Tokens for $2,111 Break-even
Claude Fable 5.1 API only $20 106M 40 tok/s
GPT-6 Astra API only $20 106M 40 tok/s
Claude Opus 5 API only $10 211M 80 tok/s
GPT-5.6 Sol API only $8.00 264M 100 tok/s
Kimi K3 Open $6.00 352M 134 tok/s
Claude Sonnet 5 API only $4.00 528M 201 tok/s
Mistral Medium 3.5 Open $3.00 704M 268 tok/s
Claude Haiku 4.5 API only $2.00 1,056M 401 tok/s
DeepSeek V4 Pro 0813 Open $1.98 1,066M 405 tok/s
Gemini 3.8 Flash API only $1.50 1,408M 535 tok/s
MiniMax-M3 Open $0.53 4,022M 1,529 tok/s
GLM-5.3-Flash Open $0.24 8,909M 3,387 tok/s

Read the right-hand column. It is the throughput a self-hosted model would have to sustain, every second of the month, for one H100 to be worth what the tokens would have cost.

  • Against Claude Fable 5.1 at $20 blended, break-even is 40 tokens a second.
  • Against Claude Opus 5 at $10, break-even is 80 tokens a second. A 27B model on an H100 with batching clears that.
  • Against GLM-5.3-Flash at $0.24, break-even is 3,387 tokens a second. Nothing you can put on one card approaches it.

So the finding is not that self-hosting is a myth. It is that self-hosting competes with expensive models, not with cheap ones — and the thing that makes it pay is never that the weights were free. It is that the model you would otherwise have rented was dear.

The videos telling people to run a local model rather than pay frontier prices are, on this arithmetic, correct. They are also solving a problem that a $0.24 API call solves more cheaply, with no GPU, no quantisation and no evenings lost to a serving stack.

And the break-even above is a floor rather than an estimate. It assumes the card runs every hour of the month; a machine used only in office hours needs roughly three times the throughput to justify itself. It ignores storage, egress, and the engineer-hours that keeping a model served will quietly consume.

One card, if you are lucky

Everything above assumed one H100. That assumption holds for exactly one of the six models.

How many H100 80GBs it takes to hold the weightsCards needed to hold each open-weight model at 8-bit, on 80GB H100 80GBs: Qwen3.8-27B 1, Mistral Medium 3.5 2, GLM-5.3-Flash 4, MiniMax-M3 6, Kimi K3 35, DeepSeek V4 Pro unstated.How many H100 80GBs it takes to hold the weightsFits on one cardNeeds a multi-GPU nodeQwen3.8-27B1 card · $2,111/moMistral Medium 3.52 cards · $4,223/moGLM-5.3-Flash4 cards · $8,445/moMiniMax-M36 cards · $12,668/moKimi K335 cards · $73,896/moDeepSeek V4 ProvLLM on a single 4xGB300 node, the card's own exampleH100 80GBs needed to hold the weights at 8-bit, log scaleWeights only, at 8-bit: one byte per parameter, divided by the 80GB on one H100 80GB. A sparse model must holdevery expert in memory to serve any of them, so the total parameter count is what counts here, not the activeone.This is a floor. The KV cache grows with context length and with how many requests are in flight, the runtimetakes its own share, and a production deployment wants headroom over all three. Real deployments need morecards than this, never fewer.Monthly cost is RunPod's published on-demand rate of $2.89 an hour, run continuously.TerraNet Technologies · terranettechnologies.com
Cards are rounded up from the weight memory alone. A model needing 1.6 cards needs two, and two cards that each hold half a model only work if the runtime can split it. Where a model card states its own serving configuration, that is used in place of the arithmetic.

The numbers behind this chart

Model Total parameters Weights at 8-bit H100 80GBs Cost a month
Qwen3.8-27B 27B 27 GB 1 $2,111
Mistral Medium 3.5 128B 128 GB 2 $4,223
GLM-5.3-Flash 320B 320 GB 4 $8,445
MiniMax-M3 428B 428 GB 6 $12,668
Kimi K3 2800B 2800 GB 35 $73,896
DeepSeek V4 Pro vLLM on a single 4xGB300 node, the card's own example - - -

A model has to keep all of its weights in memory to serve any of them, and that is true of a sparse model too: GLM-5.3-Flash activates 18 billion parameters per token, but all 320 billion have to be resident. At 8-bit — one byte per parameter, which is a floor that ignores the KV cache entirely — that is four H100s before a single request arrives.

Run the break-even again with the right number of cards and the case collapses:

Model Cards Hardware a month Break-even against its own API
Qwen3.8-27B 1 $2,111 — (no first-party API)
Mistral Medium 3.5 2 $4,223 535 tok/s
GLM-5.3-Flash 4 $8,445 13,549 tok/s
MiniMax-M3 6 $12,668 9,175 tok/s
Kimi K3 35 $73,896 4,683 tok/s

Which produces the sharpest version of the finding in this article: self-hosting an open model to avoid that same model's API essentially never pays. The vendor serving it has better hardware utilisation than you will, and it is passing some of that on. GLM would need four cards and thirteen thousand tokens a second to beat a service that charges $0.24 a million.

DeepSeek V4 Pro does not appear in that table because its card publishes no parameter count. It publishes something more useful: its own serving example is a four-GB300 node. That is the vendor telling you what this costs.

So the option that survives is narrow and specific. It is not "run open models". It is run a small model on hardware you already need, to displace an expensive one — and Qwen3.8-27B, at 27 billion dense parameters on a single card, is the only model here that qualifies.

What you are actually buying

If the money argument only works in that one narrow case, why take open weights at all? Three reasons survive the arithmetic, and none of them is the sticker price.

Nobody can reprice you. This is the strongest one, and the catalogue dates it precisely. Gemini 3.8 Flash's $0.75 and $3.75 are introductory and double on 1 January 2027. GPT-5.6 Sol's $4 and $20 are promotional and published only as far as 21 November 2026. DeepSeek's rates double during weekday peak hours. MiniMax labels its 50% discount permanent, which is a promise and not a mechanism. A downloaded set of weights has no renewal date.

Nobody can retire it. A hosted model is withdrawn when its vendor decides; a local one runs until you stop it. For anything that has to behave identically in two years — a regulated workflow, a reproducible evaluation, a product whose output users have already accepted — that is worth more than a price break.

The data stays where you put it. The reason most often given and the one least affected by any of the numbers above. If the work cannot leave your network, the API price is not a competing option and the break-even calculation does not apply.

The licence is not the same for all six

"Open weights" is doing a lot of work as a phrase. Under it sit four different legal regimes.

What "open weights" asks of you, licence by licenceSix open-weight models and their licences: DeepSeek V4 Pro under MIT, Nothing. The card states the repository and the weights are licensed under the MIT License.; GLM-5.3-Flash under MIT, Nothing. MIT in the model card frontmatter.; Qwen3.8-27B under Apache 2.0, Nothing, and an explicit patent grant on top.; Kimi K3 under Kimi K3 License, A custom licence covering both the code and the weights. Not MIT or Apache, so its terms have to be read rather than assumed.; MiniMax-M3 under MiniMax Community, A community licence rather than a standard open-source one, with its own commercial conditions.; Mistral Medium 3.5 under Modified MIT, Commercial and non-commercial use "with exceptions for companies with large revenue". The card publishes no threshold, so a large company cannot tell from it whether it qualifies..What "open weights" asks of you, licence by licenceNo condition on the deployerConditional or bespokeDeepSeek V4 ProMITNothing. The card states the repository and theweights are licensed under the MIT License.GLM-5.3-FlashMITNothing. MIT in the model card frontmatter.Qwen3.8-27BApache 2.0Nothing, and an explicit patent grant on top.Kimi K3Kimi K3 LicenseA custom licence covering both the code and theweights. Not MIT or Apache, so its terms have tobe read rather than assumed.MiniMax-M3MiniMax CommunityA community licence rather than a standardopen-source one, with its own commercialconditions.Mistral Medium 3.5Modified MITCommercial and non-commercial use "withexceptions for companies with large revenue". Thecard publishes no threshold, so a large companycannot tell from it whether it qualifies.Taken from the licence each model card names, and the condition each one states. Where a card gives no figurefor a threshold, none is invented here.Unconditional licences are listed first. The distinction that matters commercially is not how permissive alicence reads but whether anything in it turns on who is deploying the model.TerraNet Technologies · terranettechnologies.com
A licence summary is not legal advice, and a card's wording is not the licence file. Read the LICENSE in the repository before deploying any of these commercially.

The numbers behind this chart

Model Licence What it asks of a commercial user
DeepSeek V4 Pro MIT Nothing. The card states the repository and the weights are licensed under the MIT License.
GLM-5.3-Flash MIT Nothing. MIT in the model card frontmatter.
Qwen3.8-27B Apache 2.0 Nothing, and an explicit patent grant on top.
Kimi K3 Kimi K3 License A custom licence covering both the code and the weights. Not MIT or Apache, so its terms have to be read rather than assumed.
MiniMax-M3 MiniMax Community A community licence rather than a standard open-source one, with its own commercial conditions.
Mistral Medium 3.5 Modified MIT Commercial and non-commercial use "with exceptions for companies with large revenue". The card publishes no threshold, so a large company cannot tell from it whether it qualifies.

Three of the six ask nothing about who is deploying them. DeepSeek V4 Pro and GLM-5.3-Flash are MIT; Qwen3.8-27B is Apache 2.0 with an explicit patent grant, the least encumbered terms here.

The other three are conditional in ways worth knowing before a procurement conversation. Mistral Medium 3.5 is a Modified MIT licence permitting commercial use "with exceptions for companies with large revenue" — and the model card publishes no threshold, so a large company cannot tell from the card whether it qualifies. Kimi K3 ships under a bespoke Kimi K3 License, and MiniMax-M3 under a MiniMax Community licence. Both are readable, neither is standard, and neither can be assumed from the phrase "open weights" on a launch post.

Only one of the three turns explicitly on size, and it is the one that publishes no figure. The other two are bespoke for everybody: a startup and a bank read the same licence, and both have to read it.

What to do with this

If you are paying frontier prices for bulk work, the first move is not a GPU. Try GLM-5.3-Flash. It is 95% of Opus 5 on Terminal-Bench at a forty-second of the price, and the change is an API endpoint rather than a project.

If you are choosing an open model for the licence, read it. Three of the six are unconditional. The other three are not, and the only one with an explicit revenue exception publishes no figure for it.

If you are self-hosting to save money, do the sum in the second and third charts together. Take the blended price of the model you would otherwise call — not the frontier model, the cheapest one that does the job — then count the cards the model you want to host actually needs. One H100 against Claude Opus 5 is 80 tokens a second and worth a look. Four H100s against a $0.24 API is thirteen thousand, and is not.

If you are self-hosting for data residency, reproducibility or protection from a price change, none of the above applies. Those are the reasons open weights are worth having, and they were never about the sticker price.

One thing this article cannot tell you: what your hardware actually does. No model here has a measured self-hosted throughput in our catalogue, because throughput is a property of your machine and your serving stack rather than of the model. The break-even column is the number to measure against, and measuring it is the one step nobody can take for you.

AI Tools

    What open-weight models actually cost, and when running one yourself pays | TerraNet Technologies