GLM 5.3 FlashX

Z.ai · z-ai/glm-5.3-flashx

← Back to leaderboard

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...

GLM 5.3 FlashX ranks #157 of 359 models in the default AI Benchmark Hub order (quality first, then cost). At $0.81 blended per million tokens, it is cheaper than 56% of models with published API pricing. Its 1,048,576-token context window is 4.0× the catalog median (262,144 tokens).

imagetexttext+image+video->textvideo

Context

Max context: 1048576
Max output: 131072

Pricing

Input / 1M: 0.37
Output / 1M: 1.25
Blend / 1M: 0.81

Quality

Quality index:

Provider

Provider: Z.ai
Moderated: no

Monthly cost at this blended rate

Estimated API spend if every token is billed at $0.81per million (the blend of this model's input and output prices). Real bills depend on the input/output mix.

Tokens / monthEst. cost (USD)
1,000,000$0.81
10,000,000$8.10
100,000,000$81.00

Related models

Same lab, similar price, or similar quality — useful next comparisons.

More from Z.ai

Closest API price