ai-models
Qwen3.8-27B
Apache 2.0 and 27B dense: the one frontier-class model in this catalogue a single machine can run
Starts at
Free plan available
Pricing tier: Free
Visit Qwen3.8-27BIndependent software comparison
Single-node open weights vs. massive MoE with hosted API
ai-models · low search interest
ai-models
Apache 2.0 and 27B dense: the one frontier-class model in this catalogue a single machine can run
Starts at
Free plan available
Pricing tier: Free
Visit Qwen3.8-27Bai-models
Moonshot's 2.8T open-weight MoE with always-on thinking and native video understanding
Starts at
From $3/1M input tokens
Pricing tier: Usage-Based
Visit Kimi K3Expert analysis
Deploying advanced multimodal intelligence forces engineering teams to choose between the operational autonomy of running an open-weight model on local hardware and the managed reasoning power of a massive cloud-hosted API. Qwen3.8-27B packages 27 billion dense parameters under an unencumbered Apache 2.0 license, giving organizations an independent vision-language system that can be self-hosted on an ordinary single GPU with zero recurring per-token overhead. Kimi K3 approaches multimodal reasoning through an immense 2.8-trillion-parameter mixture-of-experts model with 104 billion active parameters, accessible through Moonshot's hosted per-token API or downloadable for cluster-scale deployment under a custom license. This comparison analyzes the hardware footprint, reasoning controls, licensing structures, and benchmark metrics that determine which architecture fits your production roadmap.
Feature matrix
Rows are grouped by capability, and each cell shows the wording from that vendor’s own documentation. “Not documented” means we found no cited source for that capability, which is not the same as the product lacking it.
| Capability | Qwen3.8-27B | Kimi K3 |
|---|---|---|
| Starting price | Free plan available | From $3/1M input tokens |
| Free plan | Yes | No |
| API available | Limited API available | Product API available |
| Context window and output limit | 262,144 tokens natively, extensible to 1,000,000 with YaRN | 1,048,576-token context |
| Thinking mode and effort control | Thinking on by default, disable per request, three effort levels | Thinking always on, at low, high or max effort |
| Tool use and agent support | 84.3 on OSWorld-Verified, 73.0 on Terminal-Bench 2.1 | 88.3 on Terminal-Bench 2.1, 91.2 on BrowseComp |
| Image and document input | Native vision-language: images and video | Text, image and video input through MoonViT-V2 |
| Caching, batch and speed pricing | Not documented | Cache hits at a tenth of a miss, with two TTL tiers |
| Open weights and licence terms | Apache 2.0, 27B dense parameters | Kimi K3 License, 2.8T parameters with 104B active |
| Running it on your own hardware | Runs on one machine, with quantisations for the local runtimes | vLLM, SGLang and TokenSpeed, with MXFP4 quantisation-aware training |
Model benchmarks
Qwen3.8-27B runs on Qwen 3.8 27B and Kimi K3 on Kimi K3. These are the models’ scores, not the tools’: independent evaluations from Epoch AI, Artificial Analysis and Datacurve, each at the model’s best published effort setting, last read 2026-09-05. A dash means the model has not been scored on that benchmark yet.
No scores yet on any of these benchmarks: Qwen 3.8 27B.
| Benchmark | Qwen 3.8 27B | Kimi K3 |
|---|---|---|
| FrontierMath Tier 4 | – | 39.0% at max |
| FrontierMath Tiers 1–3 | – | 72.2% at max |
| ARC-AGI-2 | – | 60.4% † at max |
| Terminal-Bench 2.1 | – | 85.0% at max |
| DeepSWE | – | 68.5% at max |
| Humanity's Last Exam | – | 46.9% at max |
| Artificial Analysis Coding Index | – | 76.2% at max |
Sources: Artificial Analysis · Epoch AI · Datacurve.
† Relayed by the source from a vendor or external leaderboard rather than run by it.
Detailed comparison
The physical compute required to run each system establishes the primary operational divide. Qwen3.8-27B is engineered as a 27B dense model where every parameter activates on every token. Because of this compact dense architecture, it is the only model in its cohort small enough to run on an ordinary single-GPU or single-node machine. The model card documents self-hosting through SGLang, vLLM, TokenSpeed, Transformers, and Docker Model Runner, with published quantised variants enabling local runtimes like llama.cpp, Ollama, LM Studio, and Jan. This allows teams to expose an internal OpenAI-compatible endpoint without managing distributed clusters. In contrast, Kimi K3 is an architectural giant built on 2.8 trillion total parameters across 896 experts, activating 16 experts plus two shared experts per token to total 104 billion active parameters across 93 layers. Although Moonshot publishes weights produced via quantisation-aware training in MXFP4, serving 2.8 trillion parameters locally exceeds the capacity of single-node setups. For organizations without multi-node enterprise infrastructure, accessing Kimi K3 realistically requires integrating Moonshot's hosted cloud API.
Both architectures accommodate long-context ingestion and multi-step reasoning, but they offer very different degrees of runtime control. Qwen3.8-27B provides a native context window of 262,144 tokens, which can be extended to 1,000,000 tokens through a YaRN configuration change using a recommended factor of 4.0. Its thinking mode is enabled by default, but developers can disable thinking entirely on a per-request basis. Furthermore, Qwen3.8-27B allows engineers to tune reasoning depth across xhigh, medium, and low effort settings while preserving prior thinking context across conversational turns using the preserve_thinking parameter. Kimi K3 delivers a native 1,048,576-token context window without requiring secondary scaling configurations, but its reasoning mode cannot be deactivated. Kimi K3 always outputs reasoning_content across low, high, or max effort levels. For basic extraction or high-volume processing where extended reasoning adds unnecessary delay, Kimi K3 forces users to wait for and pay for reasoning generation on every call.
The legal terms and economic structures backing these models establish contrasting risk profiles. Qwen3.8-27B is distributed under the Apache 2.0 license with an explicit patent grant and no commercial revenue thresholds. The model weights are entirely free to download and run commercially, fixing the ongoing cost of ownership strictly to the physical hardware used to host it. There is no first-party per-token API documented for Qwen3.8-27B, meaning hosted consumption requires contracting with third-party providers. Conversely, Kimi K3 is published under the bespoke Kimi K3 License, which requires specific legal examination rather than relying on standard open-source assumptions. When consuming Kimi K3 via Moonshot's managed chat API, pricing is structured around usage at $3.00 per million input tokens on a cache miss, $0.30 per million on a cache hit, and $15.00 per million output tokens. Context caching introduces 5-minute and 1-hour time-to-live tiers priced at $3.00 and $6.00 per million write tokens. Because Kimi K3 cannot turn off thinking, all requests pay for reasoning generation at the $15.00 per million output rate.
Each model natively processes multimodal workloads spanning text, image, and video inputs. Kimi K3 handles visual data using a dedicated 401-million-parameter MoonViT-V2 vision encoder and establishes an impressive benchmark profile for autonomous tool use and research. Moonshot publishes Kimi K3 benchmark scores of 88.3 on Terminal-Bench 2.1, 91.2 on BrowseComp, 67.5 on DeepSWE, and 93.5 on GPQA Diamond, reflecting exceptional performance in coding and command-line automation. Qwen3.8-27B also demonstrates competitive agentic capabilities despite its smaller dense footprint, recording 84.3 on the OSWorld-Verified computer-use benchmark, 73.0 on Terminal-Bench 2.1, 61.7 on SWE-bench Pro, and 89.2 on GPQA Diamond. While Kimi K3 maintains a measurable lead on complex terminal operations and browse-driven tasks, Qwen3.8-27B delivers robust desktop navigation and multimodal tool use without introducing multi-node cluster costs or cloud API reliance.
Best use case for Qwen3.8-27B
Organizations requiring an Apache 2.0 vision-language model that can be served independently on a single machine without recurring API costs.
Best use case for Kimi K3
Teams needing high-capacity tool use and deep reasoning across million-token contexts who prefer a fully managed per-token API over multi-node self-hosting.
Decision framework
Select Qwen3.8-27B if you require an unencumbered Apache 2.0 license with patent coverage, predictable zero-token hardware costs, and the operational freedom to run vision and video workloads entirely on a single local machine. It is the practical choice for teams needing to turn off reasoning mode for high-throughput pipelines or adapt context up to one million tokens via YaRN without exposing data to external APIs. Select Kimi K3 if your application relies on cutting-edge agentic execution and autonomous tool use where leading benchmark metrics like 88.3 on Terminal-Bench 2.1 and 91.2 on BrowseComp justify the investment. Kimi K3 is built for teams with budgets for Moonshot's hosted per-token API who want an out-of-the-box 1,048,576-token context window and deep reasoning across every query, provided they can accommodate always-on thinking costs and the provisions of the custom Kimi K3 License.
Bottom line
Qwen3.8-27B provides a self-hostable 27B dense model that runs completely on a single machine under Apache 2.0, whereas Kimi K3 deploys an immense 2.8T mixture-of-experts engine governed by a custom license and consumed primarily through a paid per-token API. For organizations demanding complete operational control, fixed infrastructure spending, and the freedom to switch reasoning off to reduce latency, Qwen3.8-27B represents a uniquely accessible multimodal model. In contrast, for projects prioritizing top-tier autonomous tool use, a native million-token context, and deep chain-of-thought analysis, Kimi K3 delivers frontier-level performance, provided your team can accommodate mandatory reasoning tokens at the fifteen-dollar output rate and the legal parameters of Moonshot's bespoke agreement.
Sources and verification
The product facts have been checked against the sources below. The AI-assisted analysis was audited against these exact evidence records and approved by a human editor.
Last verified September 21, 2026
Last verified September 21, 2026
Editorial validation
Human-approvedApproved September 21, 2026 after an automated evidence audit using gemini-3.6-flash.
Read our comparison methodology and editorial policy, learn about TerraNet, or report a correction.
Common questions
No. Kimi K3 enforces thinking mode across all requests, always returning reasoning_content alongside standard responses. Although you can select between low, high, and max reasoning effort, the feature cannot be disabled, meaning every request incurs output charges at the $15.00 per million token rate.
Qwen3.8-27B features a native context length of 262,144 tokens. To reach 1,000,000 tokens, operators must apply a YaRN configuration change, with the official model card recommending a scaling factor of 4.0.
Both models document deployment support for vLLM, SGLang, TokenSpeed, and Docker Model Runner. However, Qwen3.8-27B also supports Transformers and published quantised variants for local desktop engines like llama.cpp, Ollama, LM Studio, and Jan.
Qwen3.8-27B is released under the standard Apache 2.0 license, which grants commercial rights and patent protections without revenue ceilings. Kimi K3 is distributed under the proprietary Kimi K3 License, requiring teams to review custom terms before deployment.
AI-assisted draft audited against the cited product evidence and approved by a human editor. Vendor pricing and capabilities can change after the recorded verification date.
Continue researching
Kimi K3 mandates always-on thinking across every interaction, generating reasoning tokens billed at the full fifteen-dollar per million output rate regardless of whether a query requires deep thought or straightforward retrieval. While its 2.8-trillion parameter mixture-of-experts architecture achieves strong agentic scores such as 88.3 on Terminal-Bench 2.1 and 91.2 on BrowseComp, deploying the weights on-premises requires significant multi-node infrastructure, even with 104 billion active parameters. Furthermore, its custom Kimi K3 License imposes proprietary terms rather than standard permissive open-source protections, leading engineering teams to explore alternatives when seeking lighter self-hosting footprints, granular reasoning controls, or different price-performance profiles.
Read guideQwen3.8-27B ships Apache 2.0 weights for a 27B dense vision-language model, but its operational profile requires engineering teams to manage their own hardware or negotiate hosting with third-party providers. Because the model card documents no first-party managed per-token API, organizations seeking a turnkey hosted endpoint with service commitments must look elsewhere. Deployments that require out-of-the-box one-million-token contexts face an additional operational step, as Qwen3.8-27B caps its native window at 262,144 tokens and requires a manual YaRN configuration adjustment with a 4.0 scaling factor to reach a million tokens. Furthermore, the model card omits a documented knowledge cutoff date, standard latency benchmarks, and formal retirement timelines. For engineering teams prioritizing native multi-modal context windows of a million tokens, fully managed API contracts with clear service horizons, or alternate architectural designs such as mixture-of-experts, evaluating alternative open-weight releases and hosted cloud models becomes an essential technical exercise.
Read guideClaude Haiku 4.5 provides low-latency execution and high-volume cost efficiency at $1 per million input tokens, while Claude Sonnet 5 provides a 1M-token context window and autonomous tool planning at double the base token price. Teams optimizing for interactive user experiences, live customer support desks, and narrow margin footprints will find Haiku 4.5 the more practical fit. Conversely, projects requiring broad document synthesis, deep programmatic refactoring, and independent multi-turn agent loops will find Sonnet 5 essential despite its higher token counts and strict 400-error validation on sampling overrides.
Read guideClaude Opus 5 costs half as much as Claude Fable 5.1 on standard input and output tokens while delivering faster response times and flexible reasoning toggles, whereas Claude Fable 5.1 delivers Anthropic's deepest reasoning capabilities alongside slower comparative latency and strict programmatic constraints. For standard agentic engineering and general enterprise workloads, Opus 5 provides the more balanced operational foundation due to its $5 and $25 token rates, optional 2.5-times fast mode, and ability to disable thinking when latency matters. Claude Fable 5.1 belongs in pipelines where evaluations demonstrate that Opus 5 cannot resolve the underlying reasoning problem, provided the engineering stack can accommodate double the token expense, mandatory adaptive thinking, and an API contract that prohibits forced tool selection.
Read guideClaude Sonnet 5 gives you a fast, cost-efficient workhorse priced at $2 in and $10 out per million tokens for everyday production throughput, whereas Claude Opus 5 gives you a frontier reasoning engine priced at $5 in and $25 out per million tokens built for complex autonomy and cybersecurity tasks. While both tools deploy identical 1M-token context buffers and cloud availability across the Claude API, AWS Bedrock, Google Cloud, and Microsoft Foundry, they should not be treated as interchangeable endpoints. Teams operating customer-facing interfaces, high-frequency tool pipelines, and latency-sensitive features will find Sonnet 5 far easier to sustain financially and operationally. Conversely, engineering departments deploying agents for multi-file refactoring, vulnerability inspection, and high-effort reasoning should absorb the cost of Opus 5, reserving Sonnet for the surrounding orchestration layers.
Read guideGLM-5.3-Flash delivers an accessible, lightweight 18B active-parameter architecture under an unencumbered MIT license with rock-bottom operational costs, whereas Kimi K3 demands a massive multi-node 2.8T-parameter footprint, custom legal approval, and always-on reasoning tokens to unlock frontier-level agentic task completion. For high-volume production, multimodal file pipelines, and internal deployments on standard server setups, GLM-5.3-Flash provides the most practical and legally clear path forward. For demanding terminal control, automated web browsing, and multi-step programmatic problem solving where accuracy supersedes operational cost, Kimi K3 stands as the superior agentic reasoning tool.
Read guide