Best writing score
Qwen3-235B-A22B-Thinking-250788.3%
Qwen3-235B-A22B-Thinking-2507 is currently the best-ranked AI model for writing in WritingBench, scoring 88.3% across a dedicated generative-writing evaluation. Compare it for essays, books, and professional drafts; source checks and editing remain your responsibility.
Review boundary
Benchmark interpretation reviewed internally; not independently reviewed by a novelist, academic writing instructor, or practicing lawyer.
Scope and disclosure
This page ranks underlying language models for generative writing, not complete writing products, legal research systems, citation tools, or publishers. LLM Stats accepts no payment for ranking position. Read the methodology or report a correction.
WritingBench is the primary signal because it directly evaluates generative writing. The broader communication index is shown only as a cross-check for instruction retention and collaborative editing.
Best writing score
Qwen3-235B-A22B-Thinking-250788.3%
Value among leaders
Qwen3-235B-A22B-Instruct-2507$0.09 in · $0.55 out / 1M
Most context headroom
Qwen3-Next-80B-A3B-Instruct262K tokens
Use the rank to form a shortlist, then run your own blind read. A five-point benchmark lead may matter less than voice fit, correction time, privacy, or whether the product can work from your approved source set.
Communication cross-check
Claude Opus 4.6
Leads the broader communication index at 32.0. This is not substituted for the writing score.
Verification status: 15 of 15 displayed WritingBench records are not independently reproduced by LLM Stats. The export preserves source and verification fields.
A leaderboard cannot collapse craft, scholarship and professional duty into one number. Use the same evidence differently for each job.
Use AI to: Outline competing claims, test structure and revise for clarity.
Human check: Sources, quotations, reasoning and your institution’s AI policy.
Use AI to: Develop scenes, alternatives, summaries and revision passes.
Human check: Character arcs, factual continuity, rights, voice and full-manuscript context.
Use AI to: Organize supplied facts, compare drafts and simplify approved text.
Human check: Every authority, jurisdiction, confidentiality setting and professional duty.
See the legal AI guideUse AI to: Generate options, compress copy and apply a documented style system.
Human check: Brand voice, originality, factual claims, permissions and final editorial judgment.
Inspect the underlying benchmark matrix here. WritingBench is the direct writing measure; the remaining communication benchmarks add context for instruction following, multi-turn editing and professional interaction.
The useful metric is often minutes to trustworthy, publishable prose—not how impressive the first paragraph looks.
Give every model the same audience, purpose, source pack, voice constraints and word budget.
Preserve the first output, latency, token use, refusals and any tool or retrieval calls.
Hide model names. Have target readers score usefulness, structure, voice and factual support.
Track unsupported claims, citation errors, continuity breaks and minutes to publication-ready copy.
Blind score
Reader preference
Count
Unsupported claims
Minutes
Correction time
API + review
Total cost
Primary evidence. WritingBench contains 1,239 queries across six domains and 100 subdomains. Its query-dependent criteria evaluate requirements such as style, format and length instead of applying one generic writing rubric to every prompt.
Secondary evidence. The communication index aggregates broader instruction-following and multi-turn benchmarks. It helps assess sustained collaboration, but customer-service performance is not treated as proof of literary or academic quality.
Important limitation. The primary results are benchmark records rather than LLM Stats' own controlled book, essay or legal-writing trial. Model coverage and release recency vary; evaluator models can encode preferences; a score does not establish originality, factuality, confidentiality or professional fitness.
Sources. WritingBench paper, open evaluation code, LLM Stats benchmark records and provider pricing.
Refresh policy. Data refreshes hourly. Editorial conclusions change only when the evidence materially changes; the visible review date records the checked edition.
The first model in the live WritingBench ranking is the strongest default in the current dedicated writing evidence. The final choice should still be tested on your genre, source material, editing workflow and privacy requirements.
Start with the WritingBench leader, then test it on a fixed essay packet. Score thesis quality, argument structure, counterarguments, source fidelity and revision effort. Verify every citation and follow the applicable school or publisher policy.
Yes—for outlining, scene alternatives, continuity checks and revision passes. A long context window helps, but does not prove narrative judgment or voice. Keep a human-owned story bible and compare the model against complete chapters rather than isolated paragraphs.
AI can assist with approved drafting and revision workflows, but this model ranking is not a legal-product or confidentiality certification. Lawyers remain responsible for competence, client information, candor, supervision and checking the output under the rules that apply to them.
No. It ranks underlying models. Finished products can add retrieval, citations, document controls, collaboration, style memory, permissions and contractual privacy terms that the model score does not measure.
Use the same representative briefs, hide model names, preserve failures and measure reader preference plus correction time. Include a source-grounded task, a voice-matching task, a long revision and an adversarial fact-check.