BFCL-V4
Progress Over Time
Interactive timeline showing model performance evolution on BFCL-V4
BFCL-V4 Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Shanghai AI Laboratory | 744B | — | — | ||
| 2 | Alibaba Cloud / Qwen Team | — | 1.0M | $1.25 / $3.75 | ||
| 3 | InclusionAI | 124B | 131K | $0.06 / $0.18 | ||
| 4 | Alibaba Cloud / Qwen Team | — | — | — | ||
| 4 | Alibaba Cloud / Qwen Team | 397B | 262K | $0.45 / $3.00 | ||
| 6 | Alibaba Cloud / Qwen Team | 122B | 262K | $0.29 / $2.40 | ||
| 7 | Alibaba Cloud / Qwen Team | 27B | 262K | $0.26 / $2.60 | ||
| 8 | Alibaba Cloud / Qwen Team | 1.0T | 256K | $1.20 / $6.00 | ||
| 9 | Alibaba Cloud / Qwen Team | 35B | 262K | $0.14 / $1.00 | ||
| 10 | Alibaba Cloud / Qwen Team | 9B | 262K | $0.10 / $0.15 | ||
| 11 | Amazon | — | — | — | ||
| 12 | 30B | 131K | $0.16 / $0.65 | |||
| 13 | Amazon | — | 1.0M | $0.30 / $2.50 | ||
| 14 | Amazon | — | — | — | ||
| 15 | ByteDance | — | 256K | $0.25 / $2.00 | ||
| 16 | Liquid AI | 3B | — | — | ||
| 17 | 3B | 131K | $0.03 / $0.12 | |||
| 18 | 8B | 131K | $0.06 / $0.25 | |||
| 19 | Alibaba Cloud / Qwen Team | 4B | — | — | ||
| 20 | Alibaba Cloud / Qwen Team | 2B | — | — | ||
| 21 | Liquid AI | 3B | — | — | ||
| 22 | Alibaba Cloud / Qwen Team | 800M | — | — |
What is BFCL-V4?
Berkeley Function Calling Leaderboard V4 (BFCL-V4) evaluates LLMs on their ability to accurately call functions and APIs, including simple, multiple, parallel, and nested function calls across diverse programming scenarios.
BFCL-V4 is a text benchmark evaluating models on agents and tool calling tasks. LLM Stats tracks 22 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.8.
Compare leaders on the best AI for agents and best AI for tool calling leaderboards.
Current leaders
Atria Dawn Preview from Shanghai AI Laboratory currently leads the BFCL-V4 leaderboard with a score of 0.770 across 22 evaluated AI models.
FAQ
Common questions about the BFCL-V4 benchmark and leaderboard.