The AI arena is free today

Open Superagent

BFCL-V4

Progress Over Time

Interactive timeline showing model performance evolution on BFCL-V4

State-of-the-art frontier
Open
Proprietary

BFCL-V4 Leaderboard

21 models
ContextCostLicense
1
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$1.25 / $3.75
2
InclusionAI
InclusionAI
124B131K$0.06 / $0.18
3
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
397B262K$0.45 / $3.00
3
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
5
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
122B262K$0.29 / $2.40
6
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
27B262K$0.26 / $2.60
7
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0T256K$1.20 / $6.00
8
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B262K$0.14 / $1.00
9
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
9B262K$0.10 / $0.15
10
1130B131K$0.16 / $0.65
121.0M$0.30 / $2.50
13
14
ByteDance
ByteDance
256K$0.25 / $2.00
15
Liquid AI
Liquid AI
3B
163B131K$0.03 / $0.12
178B131K$0.06 / $0.25
18
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
4B
19
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2B
20
Liquid AI
Liquid AI
3B
21
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
800M
Notice missing or incorrect data?
About this benchmark

What is BFCL-V4?

Berkeley Function Calling Leaderboard V4 (BFCL-V4) evaluates LLMs on their ability to accurately call functions and APIs, including simple, multiple, parallel, and nested function calls across diverse programming scenarios.

BFCL-V4 is a text benchmark evaluating models on agents and tool calling tasks. LLM Stats tracks 21 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.8.

Compare leaders on the best AI for agents and best AI for tool calling leaderboards.

Current leaders

Qwen3.7 Max from Alibaba Cloud / Qwen Team currently leads the BFCL-V4 leaderboard with a score of 0.750 across 21 evaluated AI models.

1Qwen3.7 MaxAlibaba Cloud / Qwen Team75.0%
2Ling 3.0 FlashInclusionAI73.0%
3Qwen3.5-397B-A17BAlibaba Cloud / Qwen Team72.9%

FAQ

Common questions about the BFCL-V4 benchmark and leaderboard.

What is the BFCL-V4 benchmark?

Berkeley Function Calling Leaderboard V4 (BFCL-V4) evaluates LLMs on their ability to accurately call functions and APIs, including simple, multiple, parallel, and nested function calls across diverse programming scenarios.

What is the BFCL-V4 leaderboard?

The BFCL-V4 leaderboard ranks 21 AI models based on their performance on this benchmark. Currently, Qwen3.7 Max by Alibaba Cloud / Qwen Team leads with a score of 0.750. The average score across all models is 0.594.

What is the highest BFCL-V4 score?

The highest BFCL-V4 score is 0.750, achieved by Qwen3.7 Max from Alibaba Cloud / Qwen Team.

How many models are evaluated on BFCL-V4?

21 models have been evaluated on the BFCL-V4 benchmark, with 0 verified results and 21 self-reported results.

What categories does BFCL-V4 cover?

BFCL-V4 is categorized under agents and tool calling. The benchmark evaluates text models.

What is the best open-source model on BFCL-V4?

Qwen3.5-397B-A17B by Alibaba Cloud / Qwen Team is the top-ranked open-source model on BFCL-V4, with a score of 0.729 (rank #3).

Which model offers the best value on BFCL-V4?

Among models scoring within 10% of the leader, Ling 3.0 Flash from InclusionAI is the cheapest, at $0.06 per million input tokens with a score of 0.730.

How recent are the BFCL-V4 leaderboard results?

The BFCL-V4 leaderboard was last updated in September 2026 and currently includes 21 evaluated models.