All videos

Video briefing

What Open-Weight Models Actually Cost

GLM-5.3-Flash · Kimi K3 · DeepSeek V4 Pro · Qwen3.8-27B · MiniMax-M3 · Mistral Medium 3.5 · Claude Opus 5 · Claude Sonnet 5 · Claude Haiku 4.5 · Gemini 3.8 Flash · GPT-6 Astra 6:35

An analysis of open-weight LLM serving costs, hardware requirements, and API break-even points across twelve leading frontier models.

Transcript

What the video says

What open-weight models actually cost, and when running one yourself pays. From TerraNet Technologies. Quick thing first: if this is useful, like the video and subscribe. Online advice often insists that self-hosting open models beats paid coding subscriptions. Yet commentators overlook that an inexpensive endpoint can cut frontier expenses without deploying local hardware. Usually, calling a different endpoint achieves the goal directly.

Figure 1 plots twelve models on Terminal-Bench 2.1 against cost per million tokens, showing a stark trade-off between frontier capability and efficiency. At the ceiling, Claude Fable 5.1 and GPT-6 Astra touch ninety percent accuracy, but command twenty dollars. Just behind them, models like GPT-5.6 Sol cut that price by more than half with minimal performance loss. The real story, though, is GLM-5.3-Flash: delivering eighty-four percent accuracy for just twenty-four cents, capturing ninety-five percent of Opus performance at two percent of the cost.

Across the six open models, hosted rates vary twenty-five times over. GLM-5.3-Flash activates eighteen billion of its three hundred twenty billion parameters per token, allowing Z.ai to serve it for twenty-four cents per million. Meanwhile, Kimi K3 costs six dollars blended, outpricing Claude Sonnet 5 while scoring higher on terminal tasks. Mistral Medium 3.5 sits at three dollars, outpricing DeepSeek V4 Pro while landing lower on benchmarks. Crucially, Qwen3.8-27B appears nowhere on price charts because no vendor serves it in high volume. The one model suited to single-card deployment is absent from hosted listings.

Figure 2 calculates the break-even throughput to justify renting an H100 GPU for two thousand dollars a month. Against premium models like Claude Fable and GPT-6 Astra, self-hosting breaks even easily at just forty tokens per second. Mid-range options like Claude Opus and GPT-5.6 Sol push that target toward one hundred. However, against ultra-budget APIs, the economics collapse. Matching Gemini Flash or GLM requires saturating thousands of tokens continuously. Dedicated silicon pays off against top-tier APIs, but managed commodity models remain vastly cheaper for intermittent workloads.

Renting an H100 on RunPod totals two thousand dollars every month. A batched twenty-seven billion parameter model easily clears 80 tokens per second, making self-hosting viable against Claude Opus 5 at $10. However, beating GLM-5.3-Flash requires over 3387 tokens per second continuously throughout the month. If hardware runs purely during office hours, break-even throughput triples. Self-hosting does not undercut inexpensive hosted endpoints. It only breaks even when displacing high-priced proprietary providers. Before ordering silicon or spinning up cloud servers to lower costs, measure your real sustained traffic against these break-even thresholds.

Figure 3 maps the physical hardware costs of hosting model weights in memory at eight-bit precision. Compact architectures like Qwen-27B fit comfortably on a single two-thousand-dollar H100. Mid-tier setups, like GLM-5.3-Flash, scale reasonably across four cards at eight thousand dollars monthly. The real inflection hits at the frontier: massive models like Kimi K3 demand thirty-five cards and nearly seventy-four thousand dollars each month. Because full weights must remain resident regardless of sparse activation, inference hardware demands scale exponentially with total parameter footprint, not active compute.

Sparse architecture does not reduce GPU memory footprint. GLM-5.3-Flash activates 18 billion parameters per token, but all 320 billion parameters must sit in memory. That demands 4 H100 accelerators before answering a request, driving baseline hardware to eight thousand dollars every month. At that capacity, matching its own API requires sustaining thirteen thousand tokens every second. Commercial providers maintain utilisation rates a single enterprise cannot match. Avoid deploying local clusters merely to dodge an open model's native API; simply consume the hosted endpoint instead.

Figure 4 tracks licensing across the open-weight ecosystem, revealing a widening split between true open-source and conditional enterprise use. Models like DeepSeek V4 Pro and Qwen-27B offer friction-free commercial adoption under standard MIT and Apache 2.0 terms, complete with patent protections. However, vendors are increasingly fragmenting the landscape. Players like MiniMax, Kimi, and Mistral have moved toward bespoke community and modified licenses. This shifts the risk onto enterprise teams, turning deployment into a legal minefield governed by vague revenue gates and custom restrictions.

Self-hosting proves valuable when protecting operations rather than cutting token expenses. Cloud providers adjust rates over time, whereas downloaded weights carry no renewal date. Hosted models risk sudden deprecation or unexpected behavioral drift. Local deployments guarantee complete execution reproducibility, perpetual availability for compliance workflows, and absolute data boundary enforcement. When enterprise data cannot exit private network infrastructure, hosted API calculations become irrelevant. Treat self-hosting as an architectural investment in data sovereignty, regulatory adherence, and operational longevity rather than a simple price optimization tactic.

Do not purchase cluster hardware simply to replace expensive models. Start by routing bulk traffic to GLM-5.3-Flash. Evaluate self-hosting solely when displacing top-tier models like Claude Opus 5, and restrict internal deployments to governance, compliance, and guaranteed continuity.

Inspect our complete benchmarking catalogue, methodology, and licensing breakdown for all models on TerraNet Technologies.

Produced by TerraNet Technologies from the cited evidence behind the written article. Facts can change after the recorded date.

    What Open-Weight Models Actually Cost | TerraNet Technologies