Qwen3 VL 8B Thinking is 2.0× cheaper per million tokens (blended). Gemini 3.7 Flash has the larger context window (1,048,576 vs 131,072 tokens, 8.0×). Qwen3 VL 8B Thinking is open-weights; the other is not.

Gemini 3.7 Flash vs Qwen3 VL 8B Thinking

Live catalog fields. Best-in-row is highlighted.

Gemini 3.7 Flash

Google · google__gemini-3.7-flash

Qwen3 VL 8B Thinkingopen

Qwen · qwen__qwen3-vl-8b-thinking

Identity

Fieldgoogle__gemini-3.7-flashqwen__qwen3-vl-8b-thinking
ProviderGoogleQwen
Sluggoogle__gemini-3.7-flashqwen__qwen3-vl-8b-thinking
Statuslivelive
Open weightsNoYes
License
HuggingFaceQwen/Qwen3-VL-8B-Thinking

Quality

Fieldgoogle__gemini-3.7-flashqwen__qwen3-vl-8b-thinking
Quality index (0–100)88.4

Cost

Fieldgoogle__gemini-3.7-flashqwen__qwen3-vl-8b-thinking
Input $ / 1M tokens$0.75$0.18
Cached input $ / 1M$0.07
Output $ / 1M tokens$3.75$2.10

Context

Fieldgoogle__gemini-3.7-flashqwen__qwen3-vl-8b-thinking
Max context tokens1,048,576131,072
Max output tokens65,53632,768

Modalities

Fieldgoogle__gemini-3.7-flashqwen__qwen3-vl-8b-thinking
Inputtext, image, video, file, audioimage, text
Outputtexttext