Skip to content
ModelCap

Benchmark leaderboard

BFCL V4 function-calling leaderboard

The Berkeley Function Calling Leaderboard scores how accurately a model calls tools across simple, parallel, multi-turn and agentic function-calling tasks; the overall score is out of 100. This page lists every model ModelCap tracks with a published BFCL V4 result, its best evaluated configuration, the source's own rank, and the model's ModelCap Index position, API price and context window.

Snapshot as of 12 April 2026

Claude Opus 4.5 from Anthropic leads the BFCL V4 function-calling leaderboard among the 44 tracked models with 77.5 (Function calling configuration), per the source snapshot published 12 April 2026.

Tracked models with a result
44
Entries on the source board
83
Open-weight models listed
17

BFCL V4 standings (44 models)

Sorted by published score, one row per model. Source: UC Berkeley (Gorilla).

BFCL V4 function-calling leaderboard: every tracked model's published score, source rank, ModelCap Index position, price and context
#ModelScoreConfigurationSource rankModelCap IndexOutput $/1MContextWeightsPublished
1Claude Opus 4.5Anthropic77.5Function calling+1 more evaluated#1 / 83$25.00200KAPI only12 April 2026
2Claude Sonnet 4.5Anthropic73.2Function calling+1 more evaluated#2 / 83$15.001MAPI only12 April 2026
3GLM 4.6Z.ai72.4Function calling · Thinking#4 / 83$1.75205KOpen weights12 April 2026
4Claude Haiku 4.5Anthropic68.7Function calling+1 more evaluated#6 / 83#70$5.00200KAPI only12 April 2026
5o3OpenAI63.1Prompt+1 more evaluated#7 / 83#53$8.00200KAPI only12 April 2026
6DeepSeek V3.2 ExpDeepSeek56.7Prompt · Thinking+1 more evaluated#12 / 83$0.41164KOpen weights12 April 2026
7Gemini 2.5 FlashGoogle56.2Function calling+1 more evaluated#13 / 83$2.501MAPI only12 April 2026
8xLAM 2 32b fc rSalesforce54.7Function calling#16 / 83Restricted license12 April 2026
9GPT-4.1OpenAI54.0Function calling+1 more evaluated#17 / 83$8.001MAPI only12 April 2026
10o4 MiniOpenAI53.2Function calling+1 more evaluated#18 / 83#107$4.40200KAPI only12 April 2026
11Qwen3 235B A22B Instruct 2507Qwen52.2Prompt+1 more evaluated#20 / 83#60$0.35262KOpen weights12 April 2026
12Nanbeige4 3B Thinking 2511Nanbeige51.4Function calling#22 / 83Open weights12 April 2026
13GPT-4.1 MiniOpenAI50.5Function calling+1 more evaluated#23 / 83$1.601MAPI only12 April 2026
14Qwen3 32BQwen48.7Function calling+1 more evaluated#24 / 83#140$0.28131KOpen weights12 April 2026
15Command ACohere46.5Function calling#27 / 83#137$10.00256KGated access12 April 2026
16Qwen3 8BQwen42.6Function calling+1 more evaluated#30 / 83$0.455131KOpen weights12 April 2026
17Qwen3 30B A3B Instruct 2507Qwen41.4Function calling+1 more evaluated#32 / 83#108$0.193262KOpen weights12 April 2026
18xLAM 2 3b fc rSalesforce41.2Function calling#33 / 83Restricted license12 April 2026
19Qwen3 14BQwen41.0Function calling+1 more evaluated#34 / 83$0.24131KOpen weights12 April 2026
20Mistral Medium 3Mistral AI37.7Default+1 more evaluated#36 / 83$2.00131KAPI only12 April 2026
21Mistral Small 3.2 24BMistral AI37.2Function calling+1 more evaluated#38 / 83#134$0.25256KOpen weights12 April 2026
22Gemini 2.5 Flash LiteGoogle36.9Function calling+1 more evaluated#39 / 83$0.401MAPI only12 April 2026
23Qwen3 4B Instruct 2507Qwen35.7Function calling+1 more evaluated#40 / 83#154262KOpen weights12 April 2026
24GPT-4.1 NanoOpenAI33.1Function calling+1 more evaluated#42 / 83$0.401MAPI only12 April 2026
25Llama 3.3 70B InstructMeta31.9Function calling#46 / 83#174$0.32131KGated access12 April 2026
26xLAM 2 1b fc rSalesforce30.4Function calling#48 / 83Restricted license12 April 2026
27Gemma 3 12BGoogle30.4Prompt#49 / 83#161$0.15131KGated access12 April 2026
28Gemma 3 27BGoogle29.5Prompt#51 / 83#128$0.45131KGated access12 April 2026
29Phi 4Microsoft28.8Prompt#52 / 83#217$0.1416KOpen weights12 April 2026
30Qwen3 1.7BQwen28.4Function calling#53 / 83Open weights12 April 2026
31Llama 4 ScoutMeta28.1Function calling#54 / 83#172$0.301.3MGated access12 April 2026
32granite 3.1 8b instructIBM Granite27.1Function calling#59 / 83Open weights12 April 2026
33granite 3.2 8b instructIBM Granite26.9Function calling#62 / 83Open weights12 April 2026
34Llama 3.1 8B InstructMeta25.8Prompt#64 / 83#229$0.08131KGated access12 April 2026
35Nova Pro 1.0Amazon25.0Function calling#66 / 83#188$3.20300KAPI only12 April 2026
36Qwen3 0.6BQwen23.9Function calling+1 more evaluated#68 / 83Open weights12 April 2026
37granite 20b functioncallingIBM Granite23.2Function calling#69 / 83Open weights12 April 2026
38Nova Micro 1.0Amazon22.3Function calling#70 / 83#224$0.14128KAPI only12 April 2026
39Llama 3.2 3B InstructMeta22.0Function calling#73 / 83#252$0.33131KGated access12 April 2026
40Gemma 3 4BGoogle19.6Prompt#76 / 83#187$0.10131KGated access12 April 2026
41granite 4.0 350mIBM Granite19.0Function calling#77 / 83Open weights12 April 2026
42Llama 3.2 1B InstructMeta10.8Function calling#81 / 83#258$0.20160KGated access12 April 2026
43Llama 3 1 Nemotron Ultra 253B v1NVIDIA10.0Function calling#82 / 83#141Restricted license12 April 2026
44gemma 3 1b itGoogle7.2Prompt#83 / 83Gated access12 April 2026

How ModelCap uses BFCL V4

BFCL V4 ranks models by function-calling accuracy. ModelCap ingests the board as published, matches each entry to a catalogue model with a reviewed identity, and shows the source's score, interval and rank unchanged. Where a model has an admitted result, it feeds the ModelCap Index as agent evidence alongside the other public boards; the methodology documents the weighting and the identity rules.

BFCL V4 leaderboard: common questions

What is the BFCL V4 benchmark?

The Berkeley Function Calling Leaderboard scores how accurately a model calls tools across simple, parallel, multi-turn and agentic function-calling tasks; the overall score is out of 100. It is published by UC Berkeley (Gorilla).

Which AI model leads BFCL V4 right now?

Claude Opus 4.5 (Anthropic) holds the top BFCL V4 score among the models ModelCap tracks, at 77.5 as of the source snapshot published 12 April 2026.

How many models are ranked on the BFCL V4 leaderboard here?

44 tracked models have a published BFCL V4 result on ModelCap; the source board itself lists 83 entries. Each row shows the model's best evaluated configuration.

What is the best open-weight model on BFCL V4?

GLM 4.6 from Z.ai is the highest-scoring model with openly downloadable weights on this board, at 72.4.

Which model offers the best value on BFCL V4?

Among the ten highest-scoring priced models, Qwen3 235B A22B Instruct 2507 has the lowest listed output price at $0.35 per 1M tokens while scoring 52.2.

How is BFCL V4 scored?

The board ranks models by function-calling accuracy; higher scores are better. ModelCap shows the source's own score, interval and rank and never re-runs the evaluation.

Does BFCL V4 decide the ModelCap Index rank?

Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; BFCL V4 contributes as agent evidence where a model has an admitted result. The methodology page documents the weighting.

How recent are the BFCL V4 results?

The newest BFCL V4 publication ModelCap holds is dated 12 April 2026. The page re-renders every minute from the sealed dataset, so it reflects the latest refresh of the source board.

Which BFCL V4 model ranks highest on the ModelCap Index?

Claude Haiku 4.5 is the highest-scoring model on this board that currently holds a ModelCap Index position (#70).

Explore further