GLM 5.3 Flash

Z.ai · z-ai/glm-5.3-flash

← Back to leaderboard

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

GLM 5.3 Flash ranks #4 of 348 models in the default AI Benchmark Hub order (quality first, then cost). At $0.16 blended per million tokens, it is cheaper than 85% of models with published API pricing. Its 1,048,576-token context window is 4.0× the catalog median (262,144 tokens). A quality index of 89.7 places it at #4 of 41 models with an Arena Elo signal.

open weightsimagetexttext+image+video->textvideo

Context

Max context: 1048576
Max output: 131072

Pricing

Input / 1M: 0.07
Output / 1M: 0.25
Blend / 1M: 0.16

Quality

Quality index: 89.7

Provider

Provider: Z.ai
Moderated: no

Monthly cost at this blended rate

Estimated API spend if every token is billed at $0.16per million (the blend of this model's input and output prices). Real bills depend on the input/output mix.

Tokens / monthEst. cost (USD)
1,000,000$0.16
10,000,000$1.63
100,000,000$16.25

Related models

Same lab, similar price, or similar quality — useful next comparisons.

More from Z.ai

Closest API price

Closest quality index