ai-models
MiniMax-M3
A 428B open-weight MoE with adaptive thinking, native video, and 80.5% on SWE-bench Verified
Starts at
From $0.30/1M input tokens
Pricing tier: Usage-Based
Visit MiniMax-M3Independent software comparison
controllable reasoning and lighter serving vs. always-on thinking at massive scale
ai-models · low search interest
ai-models
A 428B open-weight MoE with adaptive thinking, native video, and 80.5% on SWE-bench Verified
Starts at
From $0.30/1M input tokens
Pricing tier: Usage-Based
Visit MiniMax-M3ai-models
Moonshot's 2.8T open-weight MoE with always-on thinking and native video understanding
Starts at
From $3/1M input tokens
Pricing tier: Usage-Based
Visit Kimi K3Expert analysis
MiniMax-M3 and Kimi K3 present developers with two fundamentally distinct scaling philosophies within open-weight multimodal models. Both architectures support one-million-token contexts, native video and visual understanding, and flexible self-hosting runtimes, yet they target conflicting operational requirements. MiniMax-M3 is built as an efficient 428B mixture-of-experts model activating only 23B parameters per token, pairing its footprint with switchable thinking modes and low base token pricing. In contrast, Moonshot's Kimi K3 deploys an immense 2.8T mixture-of-experts engine activating 104B parameters per token, mandating persistent reasoning on every query to achieve top-tier autonomous agent benchmarks. Engineering teams choosing between these systems must weigh the cost of always-on inference against the operational necessity of top-end terminal and web evaluation scores.
Feature matrix
Rows are grouped by capability, and each cell shows the wording from that vendor’s own documentation. “Not documented” means we found no cited source for that capability, which is not the same as the product lacking it.
| Capability | MiniMax-M3 | Kimi K3 |
|---|---|---|
| Starting price | From $0.30/1M input tokens | From $3/1M input tokens |
| Free plan | No | No |
| API available | Product API available | Product API available |
| Context window and output limit | 1M-token context, billed in two tiers at 512K | 1,048,576-token context |
| Thinking mode and effort control | Thinking enabled, adaptive or disabled | Thinking always on, at low, high or max effort |
| Tool use and agent support | Frontier-level long-horizon agentic work, in coding and cowork | 88.3 on Terminal-Bench 2.1, 91.2 on BrowseComp |
| Image and document input | Native multimodal: text, image and video | Text, image and video input through MoonViT-V2 |
| Speed and latency | Sparse attention for long context; a priority tier at 1.5x | Not documented |
| Caching, batch and speed pricing | Prompt-cache reads at a fifth of input | Cache hits at a tenth of a miss, with two TTL tiers |
| Open weights and licence terms | MiniMax Community licence, 428B parameters with 23B active | Kimi K3 License, 2.8T parameters with 104B active |
| Running it on your own hardware | Six runtimes including KTransformers and unsloth | vLLM, SGLang and TokenSpeed, with MXFP4 quantisation-aware training |
Model benchmarks
MiniMax-M3 runs on MiniMax-M3 and Kimi K3 on Kimi K3. These are the models’ scores, not the tools’: independent evaluations from Epoch AI, Artificial Analysis and Datacurve, each at the model’s best published effort setting, last read 2026-09-05. A dash means the model has not been scored on that benchmark yet.
| Benchmark | MiniMax-M3 | Kimi K3 |
|---|---|---|
| FrontierMath Tier 4 | – | 39.0% at max |
| FrontierMath Tiers 1–3 | – | 72.2% at max |
| ARC-AGI-2 | – | 60.4% † at max |
| Terminal-Bench 2.1 | 65.2% | 85.0% at max |
| DeepSWE | – | 68.5% at max |
| Humanity's Last Exam | 39.0% | 46.9% at max |
| Artificial Analysis Coding Index | 58.6% | 76.2% at max |
Sources: Artificial Analysis · Epoch AI · Datacurve.
† Relayed by the source from a vendor or external leaderboard rather than run by it.
Detailed comparison
The operational divide between these two architectures centers on how they allocate reasoning tokens. MiniMax-M3 introduces explicit control over its internal reasoning via a dedicated thinking parameter. Engineering teams can select between enabled, adaptive, or disabled modes. In disabled mode, the model suppresses internal reasoning tokens to minimize latency and maximize raw output throughput, which is vital for interactive user flows or simple bulk extractions. Its adaptive mode delegates the decision to the model, generating chain-of-thought only when problem difficulty warrants the expense. Kimi K3 eliminates this toggle entirely, enforcing always-on thinking. While users can adjust the reasoning effort across low, high, and max settings, the model always generates and bills reasoning content. For teams processing high query volumes where multi-step planning is unnecessary, Kimi K3 forces output-token latency and expense that cannot be bypassed.
On managed provider infrastructure, the cost profiles of these models diverge sharply. MiniMax-M3 prices standard input at $0.30 per million tokens and output at $1.20 per million tokens for prompts up to 512K tokens, doubling to $0.60 input and $2.40 output for prompts reaching across the full one-million-token boundary. Prompt-cache reads cost a fifth of the standard rate, sitting at $0.06 or $0.12 depending on the window tier. Kimi K3 charges $3.00 per million input tokens on a cache miss, with cache hits dropping to $0.30 per million tokens. However, Kimi K3 prices output tokens at $15.00 per million. Because Kimi K3 mandates thinking on every request, queries inevitably generate large volumes of reasoning tokens billed at this $15 output rate. Moonshot provides time-to-live caching tiers of 5 minutes and 1 hour with write fees of $3.00 and $6.00 per million tokens respectively. Consequently, teams with static system instructions can mitigate input costs on Kimi K3, but generation volume remains substantially more expensive than on MiniMax-M3.
While both vendors provide open weights, their hosting footprints sit in completely different infrastructure tiers. MiniMax-M3 contains 428B total parameters with only 23B active per token, supported across six documented runtimes including SGLang, vLLM, Transformers, KTransformers, unsloth, and ATOM. Its MiniMax Sparse Attention cuts compute requirements substantially, allowing teams to run inference on comparatively compact clusters. Kimi K3 scales to 2.8T total parameters with 104B active per token across 896 experts and 93 layers. Moonshot trained the model with quantisation-aware training to produce native MXFP4 weights with MXFP8 activations, avoiding the degradation of post-hoc conversion. Serving Kimi K3 locally via vLLM, SGLang, or TokenSpeed still demands extensive multi-node GPU clusters capable of holding a multi-trillion parameter architecture in memory. Teams must also inspect the legal constraints of each release, as both the MiniMax Community licence and the Kimi K3 License impose custom non-standard terms rather than standard permissive open-source terms like Apache 2.0 or MIT.
Evaluating these models for tool-use tasks reveals a trade-off between verifiable agentic benchmarks and generalist qualitative claims. Kimi K3 publishes industry-leading autonomous environment scores, including an 88.3 on Terminal-Bench 2.1, a 91.2 on BrowseComp, a 93.5 on GPQA Diamond, and a 67.5 on DeepSWE. These figures demonstrate verified proficiency in multi-step browser navigation and complex command-line workflows. MiniMax-M3 positions itself as excelling in coding and cowork, asserting frontier-level performance across long-horizon agentic benchmarks, but its model card publishes no specific benchmark metrics to verify these claims directly. Organizations prioritizing transparent, independently verified performance in command-line tools and autonomous browsing find documented validation in Kimi K3, whereas those deploying MiniMax-M3 must evaluate its agentic capabilities against internal tasks or third-party evaluations.
Best use case for MiniMax-M3
Teams requiring high-throughput multimodal processing and long contexts with the flexibility to turn off reasoning tokens to keep operational costs low.
Best use case for Kimi K3
Developers tackling complex terminal or browsing agent tasks who need top-tier reasoning capabilities and have the budget or multi-node infrastructure to sustain always-on thinking.
Decision framework
Select MiniMax-M3 if your architecture demands cost-efficient multimodal throughput, variable context windows, and granular latency management. It is particularly well suited for production systems processing high-volume text, images, or native video where the ability to turn thinking off keeps compute bills low. MiniMax-M3 is also the practical choice for engineering groups seeking to self-host an open-weight model on private hardware without maintaining massive multi-node clusters, thanks to its 23B active parameter MoE design and native 428B footprint. Choose Kimi K3 when autonomous problem-solving capabilities in terminal environments and web browsing override token budgets and hardware constraints. If your application specifically requires verified multi-step agency and can tolerate persistent reasoning token generation at Moonshot's $15 per million output rate, Kimi K3 delivers top published benchmark results. It is ideal for high-complexity workflows where the price of always-on thinking is justified by fewer task failures in programmatic environments.
Bottom line
MiniMax-M3 delivers an economically controllable 428B architecture with modular reasoning toggles and a $0.30 input token base rate, while Kimi K3 functions as a 2.8T reasoning engine engineered for maximum autonomous precision at a $15.00 output token rate. MiniMax-M3 is the sensible production engine for high-volume multimodal systems, long-context document scanning, and general software pipelines where compute thrift is vital. Its switchable thinking modes let developers eliminate reasoning overhead when answering routine queries, keeping operational margins intact. Kimi K3 is an uncompromising platform for multi-step agentic problem-solving. By mandating thinking across low, high, and max effort tiers, it trades per-request economy for validated problem-solving depth on benchmarks like Terminal-Bench 2.1 and BrowseComp. Deploy MiniMax-M3 for scalable, budget-sensitive multimodal throughput, and reserve Kimi K3 for complex autonomous workflows where agent success rates matter more than token consumption.
Sources and verification
The product facts have been checked against the sources below. The AI-assisted analysis was audited against these exact evidence records and approved by a human editor.
Last verified September 21, 2026
Last verified September 21, 2026
Editorial validation
Human-approvedApproved September 23, 2026 after an automated evidence audit using gemini-3.6-flash.
Read our comparison methodology and editorial policy, learn about TerraNet, or report a correction.
Common questions
No. Kimi K3 always operates with thinking enabled and returns reasoning tokens on every completion. While users can modulate the intensity of the reasoning using low, high, or max effort settings, the model does not support a mode that bypasses chain-of-thought generation.
MiniMax-M3 uses a split pricing model across its 1M-token context: requests up to 512K input tokens cost $0.30 per million, doubling to $0.60 per million for prompts between 512K and 1M tokens. Kimi K3 maintains a flat rate across its 1,048,576-token context, charging $3.00 per million input tokens on a cache miss and $0.30 per million on a cache hit.
MiniMax-M3 features 428B total parameters with only 23B active per token, making it viable on standard MoE inference setups across runtimes like SGLang, vLLM, and ATOM. Kimi K3 totals 2.8T parameters with 104B active, requiring substantial multi-node hardware even when leveraging its native MXFP4 quantisation-aware weights.
Neither model uses a standard open-source license like MIT or Apache 2.0. MiniMax-M3 is distributed under the custom MiniMax Community licence, while Kimi K3 is distributed under the proprietary Kimi K3 License. Both licenses require specific review of commercial use conditions before production deployment.
AI-assisted draft audited against the cited product evidence and approved by a human editor. Vendor pricing and capabilities can change after the recorded verification date.
Continue researching
Kimi K3 mandates always-on thinking across every interaction, generating reasoning tokens billed at the full fifteen-dollar per million output rate regardless of whether a query requires deep thought or straightforward retrieval. While its 2.8-trillion parameter mixture-of-experts architecture achieves strong agentic scores such as 88.3 on Terminal-Bench 2.1 and 91.2 on BrowseComp, deploying the weights on-premises requires significant multi-node infrastructure, even with 104 billion active parameters. Furthermore, its custom Kimi K3 License imposes proprietary terms rather than standard permissive open-source protections, leading engineering teams to explore alternatives when seeking lighter self-hosting footprints, granular reasoning controls, or different price-performance profiles.
Read guideMiniMax-M3 ties its hosted economics to a steep price break inside its 1M-token context window, doubling both input and output rates once a prompt crosses the 512,000-token threshold. For workloads operating deep within long contexts, standard input shifts from $0.30 to $0.60 per million tokens and output doubles from $1.20 to $2.40. Beyond the hosting ledger, deployment flexibility requires navigating the proprietary MiniMax Community licence rather than an unencumbered open-source standard like MIT or Apache. Teams assessing alternatives often do so to secure predictable linear pricing at extreme context depths, acquire unrestricted commercial licensing for self-hosted clusters, or tap into fully managed enterprise clouds without self-hosting responsibilities.
Read guideClaude Haiku 4.5 provides low-latency execution and high-volume cost efficiency at $1 per million input tokens, while Claude Sonnet 5 provides a 1M-token context window and autonomous tool planning at double the base token price. Teams optimizing for interactive user experiences, live customer support desks, and narrow margin footprints will find Haiku 4.5 the more practical fit. Conversely, projects requiring broad document synthesis, deep programmatic refactoring, and independent multi-turn agent loops will find Sonnet 5 essential despite its higher token counts and strict 400-error validation on sampling overrides.
Read guideClaude Opus 5 costs half as much as Claude Fable 5.1 on standard input and output tokens while delivering faster response times and flexible reasoning toggles, whereas Claude Fable 5.1 delivers Anthropic's deepest reasoning capabilities alongside slower comparative latency and strict programmatic constraints. For standard agentic engineering and general enterprise workloads, Opus 5 provides the more balanced operational foundation due to its $5 and $25 token rates, optional 2.5-times fast mode, and ability to disable thinking when latency matters. Claude Fable 5.1 belongs in pipelines where evaluations demonstrate that Opus 5 cannot resolve the underlying reasoning problem, provided the engineering stack can accommodate double the token expense, mandatory adaptive thinking, and an API contract that prohibits forced tool selection.
Read guideClaude Sonnet 5 gives you a fast, cost-efficient workhorse priced at $2 in and $10 out per million tokens for everyday production throughput, whereas Claude Opus 5 gives you a frontier reasoning engine priced at $5 in and $25 out per million tokens built for complex autonomy and cybersecurity tasks. While both tools deploy identical 1M-token context buffers and cloud availability across the Claude API, AWS Bedrock, Google Cloud, and Microsoft Foundry, they should not be treated as interchangeable endpoints. Teams operating customer-facing interfaces, high-frequency tool pipelines, and latency-sensitive features will find Sonnet 5 far easier to sustain financially and operationally. Conversely, engineering departments deploying agents for multi-file refactoring, vulnerability inspection, and high-effort reasoning should absorb the cost of Opus 5, reserving Sonnet for the surrounding orchestration layers.
Read guideGLM-5.3-Flash delivers an accessible, lightweight 18B active-parameter architecture under an unencumbered MIT license with rock-bottom operational costs, whereas Kimi K3 demands a massive multi-node 2.8T-parameter footprint, custom legal approval, and always-on reasoning tokens to unlock frontier-level agentic task completion. For high-volume production, multimodal file pipelines, and internal deployments on standard server setups, GLM-5.3-Flash provides the most practical and legally clear path forward. For demanding terminal control, automated web browsing, and multi-step programmatic problem solving where accuracy supersedes operational cost, Kimi K3 stands as the superior agentic reasoning tool.
Read guide