GLM-5.3-Flash
Same categoryZ.ai's natively multimodal MoE: 320B parameters with 18B active, MIT-licensed, and the cheapest hosted rate here
GLM-5.3-Flash is a 320B parameter mixture-of-experts model activating 18B parameters per token, distributed under a permissive MIT license alongside a low-cost first-party hosted API.
Best for: Best for budget-conscious organizations that want both downloadable open weights under a standard open-source license and the cheapest hosted inference option in this category.
Consider: The developer documentation markets a 1M-token context window, but the model card notes evaluation at a 300,000-token maximum context length, and free cached-input storage is offered as a limited-time rate with no published expiration date.
From $0.15/1M input tokens · Product API available