GLM 5.3 FlashX
Z.ai · z-ai/glm-5.3-flashx
GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...
GLM 5.3 FlashX ranks #157 of 359 models in the default AI Benchmark Hub order (quality first, then cost). At $0.81 blended per million tokens, it is cheaper than 56% of models with published API pricing. Its 1,048,576-token context window is 4.0× the catalog median (262,144 tokens).
Context
Max output: 131072
Pricing
Output / 1M: 1.25
Blend / 1M: 0.81
Quality
Provider
Moderated: no
Monthly cost at this blended rate
Estimated API spend if every token is billed at $0.81per million (the blend of this model's input and output prices). Real bills depend on the input/output mix.
| Tokens / month | Est. cost (USD) |
|---|---|
| 1,000,000 | $0.81 |
| 10,000,000 | $8.10 |
| 100,000,000 | $81.00 |
Related models
Same lab, similar price, or similar quality — useful next comparisons.
More from Z.ai
- GLM 5.3 Flash$0.16 / 1M
- GLM 4.7 Flash$0.23 / 1M
- GLM 4.5 Air$0.49 / 1M
- GLM 4.6V$0.60 / 1M
- GLM 4.7$1.07 / 1M
Closest API price
- Qwen3.7 Plus$0.80 / 1M
- R1 Distill Llama 70B$0.80 / 1M
- Inkling Small$0.82 / 1M
- Perceptron Mk1$0.82 / 1M
- Qwen2.5 Coder 32B Instruct$0.83 / 1M