Legal knowledge
Questions covering doctrine, rules, terminology, and legal concepts across available evaluation sets.
Qwen3.7 Max currently leads the LLM Stats legal category, followed by Qwen3.7-Plus. This rank covers model evidence; evaluate professional software and advice separately.
Review boundary
Data methodology reviewed internally; not reviewed by a practicing attorney.
Scope and disclosure
This page ranks underlying model evidence. Professional software products require a separate workflow and controls review. LLM Stats accepts no payment for ranking position. Read the methodology or report a correction.
Evaluating the complete product?
Compare sources, workflow, integrations, security, and price in the professional buyer guide.
| Model | Best for | Top strength | Watch out | Cost · Context |
|---|---|---|---|---|
| Qwen3.7 Max Alibaba Cloud / Qwen Team | Alibaba's newest — strongest open-weight Asian frontier | Excellent multilingual coverage (50+ languages) | Western provider coverage lags | $1.25 / $3.75 1.0M ctx |
| Qwen3.7-Plus Alibaba Cloud / Qwen Team | Alibaba's newest — strongest open-weight Asian frontier | Excellent multilingual coverage (50+ languages) | Western provider coverage lags | — |
| Qwen3.6 Plus Alibaba Cloud / Qwen Team | Mature Qwen generation — strong all-rounder | Open weights, broad language support | 3.7 line now ahead on the hardest tasks | $0.50 / $3.00 1.0M ctx |
| Qwen3.5-397B-A17B Alibaba Cloud / Qwen Team | Earlier Qwen 3 — still capable, especially MoE variants | MoE architecture gives strong quality at low active-parameter cost | Newer versions lead it | — |
| Qwen3.5-122B-A10B Alibaba Cloud / Qwen Team | Earlier Qwen 3 — still capable, especially MoE variants | MoE architecture gives strong quality at low active-parameter cost | Newer versions lead it | — |
| Qwen3.5-27B Alibaba Cloud / Qwen Team | Earlier Qwen 3 — still capable, especially MoE variants | MoE architecture gives strong quality at low active-parameter cost | Newer versions lead it | $0.30 / $2.40 262K ctx |
Alibaba's newest — strongest open-weight Asian frontier
Alibaba's newest — strongest open-weight Asian frontier
Mature Qwen generation — strong all-rounder
Earlier Qwen 3 — still capable, especially MoE variants
Earlier Qwen 3 — still capable, especially MoE variants
Earlier Qwen 3 — still capable, especially MoE variants
Benchmarks are controlled evidence. They help narrow a model shortlist, but they do not reproduce a complete professional workflow.
Questions covering doctrine, rules, terminology, and legal concepts across available evaluation sets.
Issue spotting and applying a rule to facts—not merely retrieving a memorized definition.
Extracting obligations, comparing provisions, classifying clauses, and following long-document context.
Whether the evidence is broad enough to trust a rank and where sparse results require more caution.
Favor legal reasoning plus search and long-context strength.
A model score cannot establish citation validity or current-law coverage.
Look for instruction following, long context, extraction, and structured output.
Test against your playbook, governing law, and known issue set.
Writing quality matters only after authority, facts, and requested posture are supplied.
Review every citation, representation, cross-reference, and defined term.
Consistency, context handling, and cost can matter more than the top aggregate score.
Measure both recall and false positives on representative documents.
The legal category combines model results from available legal benchmarks rather than relying on a single exam score. The live table exposes the contributing evaluations and coverage, so readers can distinguish a broadly measured model from one represented by only a narrow test.
When a category index is available, LLM Stats uses an uncertainty-aware conservative score. When it is not, the editorial summary falls back to benchmark coverage and normalized performance. Scores compare models inside this category; they are not probabilities of correct legal advice.
We keep product and model evidence separate. A model may reason well on a legal benchmark while the product around it lacks current case law, a citator, matter permissions, retention controls, or a defensible professional workflow.
High-consequence use: The leaderboard is comparative research, not legal, accounting, tax, or investment advice. Validate the selected model inside the exact product, data, jurisdiction, and review process you plan to use.
01
Record model version, date, benchmark set, and category.
02
Open the visible benchmark table and note missing results.
03
Compare normalized results and evidence coverage.
04
Run a separate known-answer test in the deployed product.
The answer changes as new results arrive. The live ranking above names the current leader using the legal category's available model evidence. Treat it as a capability signal, then evaluate the professional product, legal sources, jurisdiction, confidentiality terms, and workflow around that model.
Depending on the evaluation, they measure legal knowledge, rule application, issue spotting, classification, contract understanding, and exam-style reasoning. They generally do not measure privilege handling, live case-law completeness, citation treatment, client communication, or professional responsibility.
No. Benchmark performance is evidence about bounded tasks. It does not supply licensure, ethical duties, accountability, current jurisdiction-specific research, negotiation judgment, or an attorney-client relationship.
Not automatically. Compare the top models, then test the finalists inside an approved product using representative documents and known answers. Source grounding, permissions, retention, integration, latency, cost, and correction time may matter more than a small ranking difference.