Best LLM for writing — prose, tone and editing
The best LLM for writing depends on whether you want a first draft, a ruthless editor, or a brand-voice rewriter. Crowd Elo correlates loosely with “sounds good in English chat,” which is closer to writing than it is to SQL — but your style guide still wins.
Start with quality-weighted rankings, then read outputs in the arena rather than trusting adjectives like ‘most literary.’ Check output token prices: writing jobs are output-heavy, so a model with cheap input and expensive output can surprise you.
Highest quality index right now
Live OpenRouter pricing and Arena quality, cached about an hour. Not a static blog table.
| Model | Provider | Blend $ / 1M | Quality | Context |
|---|---|---|---|---|
| Qwen3.8 Max (0902) | Qwen | 4.00 | 95.1 | 1,000,000 |
| Claude Opus 5 | Anthropic | 15.00 | 93.4 | 1,000,000 |
| Claude Fable 5 | Anthropic | 30.00 | 91.1 | 1,000,000 |
| GLM 5.3 Flash | Z.ai | 0.16 | 89.7 | 1,048,576 |
| Qwen3.8 27B | Qwen | 1.71 | 89.0 | 1,000,000 |
Draft vs edit
Drafting rewards fluency and idea coverage. Editing rewards following a markup scheme (“track changes,” “don’t invent facts”). Claude-class models are often shortlisted for edit-heavy workflows; cheaper models often suffice for brainstorming. Put both jobs in the arena: the winner may differ.
Brand voice
No public leaderboard measures your brand voice. Build a 10-sample eval: rewrite a paragraph to match a style sheet. Score with a human, not with another LLM if you can avoid it. Use this site to keep candidates cheap enough that you can afford the human loop.
Long-form and context
A novel chapter or a 40-page report needs context and a model that does not lose the plot in the middle. Compare max context and try a needle: ask for a fact from paragraph 12 of a pasted draft. If it fails, RAG chapters instead of buying a huge window you still can’t trust.
Multilingual writing
English-centric Elo will mis-rank models for other languages. Include Qwen and Gemini in the shortlist and test in the target language. See also the translation guide.
Cost control for content mills
If you generate thousands of product descriptions, output price dominates. Sort by cost, generate a sample, and only escalate to a flagship for hero pages. The monthly cost table on each model page at 10M and 100M tokens is the planning number.
FAQ
Is Claude the best writing model?
- It is a strong default for careful prose, not a monopoly. Compare current Claude, GPT, and Gemini rows on output price and arena samples.
What about open models for writing?
- Plenty are good enough for drafts. Filter open weights and compare quality index, then read the actual text.
How do I keep a consistent voice?
- A style sheet in the prompt plus a cheap model beats hopping flagship models every week.
Does quality index measure writing?
- Indirectly, via chat preference votes. Use it to shortlist, then read.
Related guides
Also try: LLM leaderboard, compare tool, live arena.