# Natural voice field-test protocol

Updated: 2026-09-04

This protocol accompanies the LLM Stats Best AI for Writers guide. It tests voice fidelity and editorial burden on your own writing; it does not replace the WritingBench model ranking.

## Setup

1. Select three finalist models from the live WritingBench ranking.
2. Create one 300-word brief representative of your real work.
3. Attach the same approved sources and two short, anonymized samples of your own writing.
4. Freeze the system prompt, user prompt, model version, temperature, tools, token limit, and word limit.
5. Generate one draft from each model and hide the model identities before review.

## Score each draft

- Voice fidelity: 1–5
- Specificity: 1–5
- Structural usefulness: 1–5
- Unsupported claims: count
- Rewritten sentences: count
- Time to usable draft: minutes
- Total API or seat cost: USD

## Decision rule

Reject any draft with consequential unsupported claims. Among the remaining drafts, prefer the model with the lowest correction time and rewrite burden, using voice fidelity as the tie-breaker. Preserve every prompt, output, score, and failure.

## Limitations

One brief cannot establish performance across every genre. Repeat the test for materially different work. Voice judgments are subjective, and exposure to a model's style can bias a reviewer. Do not include confidential material unless the selected product, contract, and organizational policy permit it.
