GLM-5.3-Flash
Same categoryZ.ai's natively multimodal MoE: 320B parameters with 18B active, MIT-licensed, and the cheapest hosted rate here
GLM-5.3-Flash from Z.ai provides an MIT-licensed 320B parameter MoE architecture that activates only 18B parameters per token, paired with native video, image, text, and file ingestion. At an API baseline of $0.15 per million input tokens and $0.50 per million output tokens, it offers a remarkably affordable managed endpoint, while the model card documents six separate serving frameworks including vLLM, SGLang, and KTransformers alongside 115 quantisations for lightweight self-hosting.
Best for: Engineering teams that require fully open MIT weights, native visual and video ingestion, and very low hosted token rates.
Consider: Although marketed with a 1M-token context window, the model card notes evaluation at a 300,000-token maximum context length with a context management strategy, and its documentation does not publish a latency figure or fixed knowledge cutoff.
From $0.15/1M input tokens · Product API available