LLM arena — test models on the same prompt
An LLM arena is the honest UI for model choice: same prompt, multiple models, your eyes (or a blind vote). Specs tell you price and context; an arena tells you whether the answer is usable.
AI Benchmark Hub’s arena replays captured streams so you can watch several models on one template without burning a key on every page load. You still see time-to-first-token and tokens per second. For paid ids, bring your own key when you leave the replay path.
Popular frontier models by quality
Live OpenRouter pricing and Arena quality, cached about an hour. Not a static blog table.
| Model | Provider | Blend $ / 1M | Quality | Context |
|---|---|---|---|---|
| Claude Opus 5 | Anthropic | 15.00 | 91.4 | 1,000,000 |
| Claude Fable 5 | Anthropic | 30.00 | 89.4 | 1,000,000 |
| Qwen3.8 27B | Qwen | 1.71 | 87.2 | 1,000,000 |
| Gemini 3.7 Flash | 2.25 | 86.7 | 1,048,576 | |
| DeepSeek V4 Pro 0423 | DeepSeek | 1.43 | 86.4 | 1,024,000 |
Arena vs leaderboard
The leaderboard is for sorting 500 rows. The arena is for feeling 2–4 finalists. Use weights to get to three models, then stop clicking spreadsheets and read answers. Most teams skip this and regret it when the “#1” model writes the wrong tone.
Blind votes
If you can hide names, do. Brand bias is real: people vote GPT because it says GPT. Our arena supports blind-style comparison in the product UI; use it for stakeholder demos.
What we pre-compute
Replays exist so the arena is a product, not a surprise bill. They are snapshots. For a hiring-critical eval, re-run live. See the methodology page and README for how the roster is selected from OpenRouter.
Prompts worth using
One typical ticket, one adversarial edge case, one long-context paste if that’s your job. Three prompts beat a 40-item spreadsheet you never run. Save the winners as your golden set.
Elo on this site vs LMArena
LMArena’s public Elo feeds our quality index. Votes you cast here are for your decision, not automatically the global board. Read “what is Arena Elo” if you want the math in plain language.
FAQ
Is this the same as lmarena.ai?
- No. We reuse their public text Elo as a quality signal and run our own multi-model replay/arena UX with pricing and weights.
Why aren’t all models in the arena?
- Replays are generated for a derived roster (core labs, pinned ids). The leaderboard still lists the wider catalog.
Can I paste my own prompt?
- Yes, within the arena UI. Replays illustrate defaults; your prompt is the point.
Do ads run in the arena?
- No. We keep the benchmarking tool clean on purpose.
Related guides
Also try: LLM leaderboard, compare tool, live arena.