Cold open
Show the consequence before explaining the setup.
Qwen3-235B-A22B-Thinking-2507 is the best-ranked starting point in WritingBench, not a YouTube-specific test. Compare the top models for strong hooks, natural spoken delivery, and less editing.
Updated September 5, 2026 · 15 WritingBench models · No paid placement
“The spreadsheet looked profitable. One hidden row changed the whole story.”
Hook: 12 spoken words before context
Review boundary
Editorial application reviewed internally; not reviewed by YouTube or a channel-specific retention analyst.
Scope and disclosure
The ranking uses WritingBench as the closest dedicated writing evidence. LLM Stats has not claimed that it is a YouTube retention benchmark. LLM Stats accepts no payment for ranking position. Read the methodology or report a correction.
WritingBench measures generative writing across domains. It is useful evidence for drafting quality, but it does not measure click-through rate, audience retention, factual accuracy, or whether a line works when spoken.
Alibaba Cloud / Qwen Team
WritingBench
88.3%
Evidence
Self-reported
Alibaba Cloud / Qwen Team
WritingBench
87.3%
Evidence
Self-reported
Alibaba Cloud / Qwen Team
WritingBench
86.7%
Evidence
Self-reported
Alibaba Cloud / Qwen Team
WritingBench
86.2%
Evidence
Self-reported
Alibaba Cloud / Qwen Team
WritingBench
85.5%
Evidence
Self-reported
Alibaba Cloud / Qwen Team
WritingBench
85.5%
Evidence
Self-reported
Alibaba Cloud / Qwen Team
WritingBench
85.2%
Evidence
Self-reported
The common failure is a polished essay wearing timestamps. Spoken writing needs shorter thought units, deliberate breath, concrete nouns, and open loops that are actually closed.
Show the consequence before explaining the setup.
Give the viewer a reason to trust the premise before asking for patience.
Move between claim, example, footage, tension, and interpretation.
Deliver the title promise; do not finish with a generic recap.
Give the top three models the same brief, sources, title, audience, target duration, and voice sample. Read the first 90 seconds aloud before looking at the rest.
Does the first sentence create a specific unanswered question?
Does it sound natural when spoken, not merely clean on the page?
Does every section earn the next one without padding?
Are facts, footage, and claims traceable to supplied sources?
Does the ending deliver the promise instead of summarizing it?
How many minutes are needed before the script sounds like the creator?