Internal API instruction following (hard)
Progress Over Time
Interactive timeline showing model performance evolution on Internal API instruction following (hard)
Internal API instruction following (hard) Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | OpenAI | — | — | — | ||
| 2 | OpenAI | — | — | — | ||
| 3 | OpenAI | — | — | — | ||
| 4 | OpenAI | — | 1.0M | $2.00 / $8.00 | ||
| 5 | OpenAI | — | 1.0M | $0.40 / $1.60 | ||
| 6 | OpenAI | — | 1.0M | $0.10 / $0.40 | ||
| 7 | OpenAI | — | 128K | $2.50 / $10.00 |
What is Internal API instruction following (hard)?
Internal API instruction following (hard) benchmark - specific documentation not found in official sources
Internal API instruction following (hard) is a text benchmark evaluating models on structured output and general tasks. LLM Stats tracks 7 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.6.
Compare leaders on the best AI for structured output and best AI for general leaderboards.
Current leaders
GPT-5 from OpenAI currently leads the Internal API instruction following (hard) leaderboard with a score of 0.640 across 7 evaluated AI models.
FAQ
Common questions about the Internal API instruction following (hard) benchmark and leaderboard.