Nexus
Progress Over Time
Interactive timeline showing model performance evolution on Nexus
Nexus Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | 405B | — | — | |||
| 2 | 70B | — | — | |||
| 3 | 8B | — | — | |||
| 4 | 3B | — | — |
What is Nexus?
NexusRaven benchmark for evaluating function calling capabilities of large language models in zero-shot scenarios across cybersecurity tools and API interactions
Nexus is a text benchmark evaluating models on general and tool calling tasks. LLM Stats tracks 4 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.6.
Compare leaders on the best AI for general and best AI for tool calling leaderboards.
Current leaders
Llama 3.1 405B Instruct from Meta currently leads the Nexus leaderboard with a score of 0.587 across 4 evaluated AI models.
FAQ
Common questions about the Nexus benchmark and leaderboard.