The AI arena is free today

Open Superagent

BFCL-V4

Progress Over Time

Interactive timeline showing model performance evolution on BFCL-V4

State-of-the-art frontier
Open
Proprietary

BFCL-V4 Leaderboard

22 models
ContextCostLicense
1
Shanghai AI Laboratory
Shanghai AI Laboratory
744B——
2
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
—1.0M$1.25 / $3.75
3
InclusionAI
InclusionAI
124B131K$0.06 / $0.18
4
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
———
4
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
397B262K$0.45 / $3.00
6
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
122B262K$0.29 / $2.40
7
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
27B262K$0.26 / $2.60
8
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0T256K$1.20 / $6.00
9
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B262K$0.14 / $1.00
10
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
9B262K$0.10 / $0.15
11———
1230B131K$0.16 / $0.65
13—1.0M$0.30 / $2.50
14———
15
ByteDance
ByteDance
—256K$0.25 / $2.00
16
Liquid AI
Liquid AI
3B——
173B131K$0.03 / $0.12
188B131K$0.06 / $0.25
19
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
4B——
20
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2B——
21
Liquid AI
Liquid AI
3B——
22
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
800M——
Notice missing or incorrect data?
About this benchmark

What is BFCL-V4?

Berkeley Function Calling Leaderboard V4 (BFCL-V4) evaluates LLMs on their ability to accurately call functions and APIs, including simple, multiple, parallel, and nested function calls across diverse programming scenarios.

BFCL-V4 is a text benchmark evaluating models on agents and tool calling tasks. LLM Stats tracks 22 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.8.

Compare leaders on the best AI for agents and best AI for tool calling leaderboards.

Current leaders

Atria Dawn Preview from Shanghai AI Laboratory currently leads the BFCL-V4 leaderboard with a score of 0.770 across 22 evaluated AI models.

1Atria Dawn PreviewShanghai AI Laboratory77.0%
2Qwen3.7 MaxAlibaba Cloud / Qwen Team75.0%
3Ling 3.0 FlashInclusionAI73.0%

FAQ

Common questions about the BFCL-V4 benchmark and leaderboard.

What is the BFCL-V4 benchmark?

Berkeley Function Calling Leaderboard V4 (BFCL-V4) evaluates LLMs on their ability to accurately call functions and APIs, including simple, multiple, parallel, and nested function calls across diverse programming scenarios.

What is the BFCL-V4 leaderboard?

The BFCL-V4 leaderboard ranks 22 AI models based on their performance on this benchmark. Currently, Atria Dawn Preview by Shanghai AI Laboratory leads with a score of 0.770. The average score across all models is 0.602.

What is the highest BFCL-V4 score?

The highest BFCL-V4 score is 0.770, achieved by Atria Dawn Preview from Shanghai AI Laboratory.

How many models are evaluated on BFCL-V4?

22 models have been evaluated on the BFCL-V4 benchmark, with 0 verified results and 22 self-reported results.

What categories does BFCL-V4 cover?

BFCL-V4 is categorized under agents and tool calling. The benchmark evaluates text models.

What is the best open-source model on BFCL-V4?

Atria Dawn Preview by Shanghai AI Laboratory is the top-ranked open-source model on BFCL-V4, with a score of 0.770 (rank #1).

Which model offers the best value on BFCL-V4?

Among models scoring within 10% of the leader, Ling 3.0 Flash from InclusionAI is the cheapest, at $0.06 per million input tokens with a score of 0.730.

How recent are the BFCL-V4 leaderboard results?

The BFCL-V4 leaderboard was last updated in October 2026 and currently includes 22 evaluated models.