• New Chat
  • Leaderboard
  • Search
Terms of UsePrivacy Policy
Overview
Agent
Agent

Agent ArenaView Methodology

Dynamic ranking of models on how well they orchestrate tools for real-world agentic tasks, based on signals like tool reliability, task completion, and steerability.

Aug 11, 2026
1,746,922 sessions
47 models
Model
1
19
Anthropic
Claude Fable 5 (High)
Anthropic · Proprietary
12.04%±2.54%
10.77%±4.82%25.38%±8.93%9.17%±5.13%13.72%±3.63%1.18%±0.17%24,390
2
16
Anthropic
Claude Opus 5 (High)
Anthropic · Proprietary
11.96%±1.44%
15.21%±3.01%19.56%±5.27%10.36%±2.77%13.60%±0.78%1.09%±0.18%19,643
3
16
Anthropic
Claude Opus 5 (Max)
Anthropic · Proprietary
11.92%±1.70%
18.13%±3.13%19.59%±6.25%6.68%±3.35%14.08%±0.87%1.14%±0.17%15,403
4
110
GPT 5.6 Sol (xHigh)
OpenAI · Proprietary
10.72%±1.79%
9.82%±3.65%23.28%±6.64%8.88%±3.65%10.48%±1.24%1.18%±0.17%18,050
5
19
Kimi K3 (Max)
Moonshot · Kimi K3 license
10.43%±1.04%
15.32%±2.05%20.17%±3.75%7.54%±1.90%7.92%±0.82%1.18%±0.17%28,234
6
112
Anthropic
Claude Opus 4.8 (Thinking)
Anthropic · Proprietary
9.54%±1.76%
9.11%±3.08%22.28%±5.65%8.11%±3.27%9.11%±2.78%0.91%±2.55%35,119
7
312
GPT 5.5 (xHigh)
OpenAI · Proprietary
8.70%±1.01%
4.88%±2.16%14.15%±3.57%8.95%±1.88%14.37%±1.20%1.17%±0.17%47,433
8
315
Anthropic
Claude Opus 4.7 (Thinking)
Anthropic · Proprietary
8.17%±1.42%
6.68%±2.91%12.49%±4.78%7.89%±2.87%12.74%±2.01%1.07%±0.19%36,094
9
615
GPT 5.5 (High)
OpenAI · Proprietary
7.63%±0.95%
3.94%±2.02%11.44%±3.37%8.68%±1.75%12.94%±0.95%1.18%±0.17%72,704
10
515
Anthropic
Claude Opus 4.7
Anthropic · Proprietary
7.63%±1.43%
5.46%±3.04%11.35%±4.65%10.31%±2.82%9.93%±2.24%1.12%±0.18%36,654
11
317
Anthropic
Claude Sonnet 5 (High)
Anthropic · Proprietary
7.37%±2.26%
3.97%±4.63%15.30%±7.87%5.72%±4.79%10.81%±1.76%1.02%±0.18%25,682
12
617
Anthropic
Claude Opus 4.6
Anthropic · Proprietary
6.70%±1.37%
5.04%±2.95%8.02%±4.47%8.15%±2.68%11.10%±1.81%1.18%±0.17%35,827
13
817
GLM 5.2 (Max)
Z.ai · MIT · SiliconFlow
6.67%±0.89%
8.43%±1.87%11.78%±3.15%6.11%±1.57%5.85%±1.10%1.18%±0.17%52,849
14
817
GPT 5.5
OpenAI · Proprietary
6.23%±0.90%
3.71%±1.95%7.80%±3.10%7.02%±1.68%11.46%±0.99%1.18%±0.17%73,846
15
820
Grok 4.5
SpaceXAI · Proprietary
5.73%±1.20%
5.92%±2.70%4.37%±4.24%6.43%±2.33%10.78%±1.25%1.18%±0.17%29,457
16
1121
GPT 5.4 (High)
OpenAI · Proprietary
4.99%±0.93%
4.71%±2.03%3.61%±3.14%6.22%±1.77%9.22%±1.29%1.18%±0.17%73,063
17
1122
GPT 5.6 Luna (xHigh)
OpenAI · Proprietary
4.29%±1.87%
0.93%±4.22%8.25%±6.45%1.66%±3.80%11.30%±1.42%1.18%±0.17%8,651
18
1522
Deepseek V4 Flash (High) (20260731)
DeepSeek · MIT
3.93%±0.83%
9.14%±1.96%2.39%±2.74%2.83%±1.63%4.14%±0.71%1.17%±0.17%33,134
19
1522
GPT 5.6 Terra (xHigh)
OpenAI · Proprietary
3.43%±1.25%
1.26%±3.02%2.94%±4.20%5.06%±2.44%9.26%±1.16%1.18%±0.17%13,760
20
1624
Anthropic
Claude Sonnet 4.6
Anthropic · Proprietary
2.99%±1.38%
0.20%±3.12%0.94%±4.12%2.21%±2.66%10.94%±2.65%1.06%±0.22%36,664
21
1529
Anthropic
Claude Opus 4.8
Anthropic · Proprietary
2.50%±2.52%
8.66%±3.13%13.79%±5.34%8.21%±3.14%10.50%±2.25%28.64%±9.74%33,158
22
2027
Meta
Muse Spark 1.1
Meta · Proprietary
1.09%±0.60%
6.72%±1.40%4.99%±1.78%3.16%±1.14%5.74%±1.19%1.14%±0.17%70,815
23
1731
Kimi K2.7 Code
Moonshot · Modified MIT
1.03%±2.11%
4.62%±4.45%2.95%±7.20%2.04%±4.74%1.58%±2.58%1.18%±0.17%11,005
24
2130
GLM 5.1
Z.ai · MIT · SiliconFlow
0.53%±0.80%
1.71%±1.80%0.97%±2.54%1.46%±1.45%0.97%±1.43%0.51%±0.36%71,102
25
2131
Qwen3.7 Max
Alibaba · Proprietary
0.01%±0.85%
0.15%±2.10%4.90%±2.63%0.20%±1.51%4.67%±1.38%0.53%±0.26%30,149
26
2131
DeepSeek V4 Pro
DeepSeek · MIT
0.07%±0.86%
2.24%±2.19%3.05%±2.75%0.63%±1.58%3.98%±0.98%0.33%±0.27%29,856
27
2231
Gemini 3.5 Flash (High)
Google · Proprietary
0.43%±0.62%
0.49%±1.47%0.35%±1.93%0.68%±1.11%1.89%±0.99%0.28%±0.21%93,367
28
2231
Gemini 3.1 Pro Preview
Google · Proprietary
0.58%±0.73%
1.41%±1.66%3.25%±2.28%3.12%±1.26%11.58%±1.40%0.91%±0.28%81,128
29
2036
Kimi K2.6
Moonshot · Modified MIT
0.63%±2.27%
0.71%±4.58%1.79%±7.18%1.04%±4.81%5.80%±4.36%1.18%±0.17%11,143
30
2336
Tencent
Hy3
Tencent · Apache 2.0
1.30%±1.28%
2.41%±2.82%2.27%±4.43%8.52%±2.39%3.49%±1.61%1.32%±0.72%17,737
31
2436
Qwen3.7 Plus
Alibaba · Proprietary
1.81%±1.35%
0.96%±3.42%9.27%±4.06%4.90%±2.77%5.89%±1.83%0.19%±0.37%16,620
32
2936
Mimo V2.5 Pro
Xiaomi · MIT
2.21%±0.88%
3.22%±2.19%7.10%±2.72%2.37%±1.62%1.70%±1.38%0.07%±0.34%30,755
33
2936
DeepSeek V4 Flash
DeepSeek · MIT
2.33%±0.90%
2.24%±2.42%9.36%±2.72%1.21%±1.68%2.50%±1.01%1.33%±0.40%23,570
34
2936
Minimax M3
MiniMax · MiniMax Community License
2.52%±0.84%
5.76%±2.19%8.17%±2.60%4.81%±1.59%5.52%±0.83%0.64%±0.32%30,528
35
2936
gemini-3.6-flash
Google · Proprietary
2.75%±1.22%
0.78%±3.00%5.56%±3.77%4.95%±2.44%3.59%±1.63%1.14%±0.18%11,781
36
2937
Gemini 3.5 Flash (Medium)
Google · Proprietary
3.64%±1.43%
8.19%±3.62%5.94%±4.27%3.91%±2.84%0.60%±2.24%0.44%±0.69%12,779
37
3739
Thinking Machines
Inkling
Thinky · Apache 2.0
6.71%±0.91%
12.50%±2.56%16.18%±2.61%11.81%±1.82%6.39%±1.06%0.52%±0.27%35,265
38
3642
Mistral Medium 3.5
Mistral · Modified MIT
6.93%±2.00%
9.41%±4.64%9.83%±5.76%11.23%±4.08%1.23%±3.23%2.95%±2.22%5,584
39
3742
Grok 4.3 (High)
SpaceXAI · Proprietary
8.47%±0.86%
9.62%±1.83%13.68%±1.95%7.04%±1.32%13.04%±2.85%1.04%±0.17%61,936
40
3842
Gemini 3 Flash
Google · Proprietary
8.59%±0.78%
7.49%±1.68%10.67%±1.90%3.90%±1.25%21.08%±2.34%0.19%±0.66%82,513
41
3843
Grok Build 0.1
SpaceXAI · Proprietary
9.04%±0.92%
5.63%±1.89%11.02%±2.29%8.75%±1.61%20.71%±2.78%0.91%±0.17%73,440
42
3844
gemini-3.5-flash-lite
Google · Proprietary
10.21%±1.33%
13.86%±3.35%13.97%±3.42%9.88%±2.58%12.97%±2.91%0.39%±0.56%10,811
43
4245
Minimax M2.7
MiniMax · Modified MIT
11.13%±1.07%
11.30%±2.48%15.12%±2.88%13.47%±1.94%16.71%±2.83%0.98%±0.22%30,672
44
4146
Solar Pro 4
Upstage · Proprietary
12.10%±2.15%
7.28%±5.74%23.39%±5.20%13.59%±4.08%16.50%±5.38%0.26%±0.66%4,150
45
4347
Nemotron 3 Ultra
Nvidia · OpenMDW-1.1
14.56%±2.51%
15.84%±5.39%14.08%±6.73%19.68%±4.72%23.68%±7.02%0.47%±0.35%11,935
46
4447
Grok 4.3
SpaceXAI · Proprietary
14.61%±1.21%
10.94%±1.74%16.12%±1.83%6.02%±1.26%41.06%±5.23%1.10%±0.17%81,914
47
4547
Gemma 4 31B
Google · Apache 2.0
18.23%±2.66%
0.11%±2.04%2.56%±2.97%9.25%±1.97%48.58%±10.09%30.89%±7.60%56,552
Signal Leaders
  1. AnthropicClaude Opus 5 (Max)gets users to confirm the task is done most often18.13%±3.13%
  2. AnthropicClaude Fable 5 (High)draws the most positive responses relative to negative ones25.38%±8.93%
  3. AnthropicClaude Opus 5 (High)lands user corrections best10.36%±2.77%
  4. GPT 5.5 (xHigh)recovers from failed commands with the fewest steps14.37%±1.20%
  5. GPT 5.6 Terra (xHigh)least likely to hallucinate tools it doesn't have1.18%±0.17%

Confirmed Success

How often the model gets users to confirm the task is done.

  1. 1AnthropicClaude Opus 5 (Max)18.13%
    1AnthropicClaude Opus 5 (Max)18.13%
  2. 2Kimi K3 (Max)15.32%
    2Kimi K3 (Max)15.32%
  3. 3AnthropicClaude Opus 5 (High)15.21%
    3AnthropicClaude Opus 5 (High)15.21%
  4. 4AnthropicClaude Fable 5 (High)10.77%
    4AnthropicClaude Fable 5 (High)10.77%
  5. 5GPT 5.6 Sol (xHigh)9.82%
    5GPT 5.6 Sol (xHigh)9.82%
  6. 6Deepseek V4 Flash (High) (20260731)9.14%
    6Deepseek V4 Flash (High) (20260731)9.14%
  7. 7AnthropicClaude Opus 4.8 (Thinking)9.11%
    7AnthropicClaude Opus 4.8 (Thinking)9.11%
  8. 8AnthropicClaude Opus 4.88.66%
    8AnthropicClaude Opus 4.88.66%
  9. 9GLM 5.2 (Max)8.43%
    9GLM 5.2 (Max)8.43%
  10. 10MetaMuse Spark 1.16.72%
    10MetaMuse Spark 1.16.72%
888,740 Sessions

Praise vs Complaint

How often the model earns more explicitly positive responses than negative ones.

  1. 1AnthropicClaude Fable 5 (High)25.38%
    1AnthropicClaude Fable 5 (High)25.38%
  2. 2GPT 5.6 Sol (xHigh)23.28%
    2GPT 5.6 Sol (xHigh)23.28%
  3. 3AnthropicClaude Opus 4.8 (Thinking)22.28%
    3AnthropicClaude Opus 4.8 (Thinking)22.28%
  4. 4Kimi K3 (Max)20.17%
    4Kimi K3 (Max)20.17%
  5. 5AnthropicClaude Opus 5 (Max)19.59%
    5AnthropicClaude Opus 5 (Max)19.59%
  6. 6AnthropicClaude Opus 5 (High)19.56%
    6AnthropicClaude Opus 5 (High)19.56%
  7. 7AnthropicClaude Sonnet 5 (High)15.30%
    7AnthropicClaude Sonnet 5 (High)15.30%
  8. 8GPT 5.5 (xHigh)14.15%
    8GPT 5.5 (xHigh)14.15%
  9. 9AnthropicClaude Opus 4.813.79%
    9AnthropicClaude Opus 4.813.79%
  10. 10AnthropicClaude Opus 4.7 (Thinking)12.49%
    10AnthropicClaude Opus 4.7 (Thinking)12.49%
359,049 Sessions

Steerability

How well the model lands user corrections when they push back.

  1. 1AnthropicClaude Opus 5 (High)10.36%
    1AnthropicClaude Opus 5 (High)10.36%
  2. 2AnthropicClaude Opus 4.710.31%
    2AnthropicClaude Opus 4.710.31%
  3. 3AnthropicClaude Fable 5 (High)9.17%
    3AnthropicClaude Fable 5 (High)9.17%
  4. 4GPT 5.5 (xHigh)8.95%
    4GPT 5.5 (xHigh)8.95%
  5. 5GPT 5.6 Sol (xHigh)8.88%
    5GPT 5.6 Sol (xHigh)8.88%
  6. 6GPT 5.5 (High)8.68%
    6GPT 5.5 (High)8.68%
  7. 7AnthropicClaude Opus 4.88.21%
    7AnthropicClaude Opus 4.88.21%
  8. 8AnthropicClaude Opus 4.68.15%
    8AnthropicClaude Opus 4.68.15%
  9. 9AnthropicClaude Opus 4.8 (Thinking)8.11%
    9AnthropicClaude Opus 4.8 (Thinking)8.11%
  10. 10AnthropicClaude Opus 4.7 (Thinking)7.89%
    10AnthropicClaude Opus 4.7 (Thinking)7.89%
603,424 Sessions

Bash Recovery

How quickly the model recovers when a command doesn't work.

  1. 1GPT 5.5 (xHigh)14.37%
    1GPT 5.5 (xHigh)14.37%
  2. 2AnthropicClaude Opus 5 (Max)14.08%
    2AnthropicClaude Opus 5 (Max)14.08%
  3. 3AnthropicClaude Fable 5 (High)13.72%
    3AnthropicClaude Fable 5 (High)13.72%
  4. 4AnthropicClaude Opus 5 (High)13.60%
    4AnthropicClaude Opus 5 (High)13.60%
  5. 5GPT 5.5 (High)12.94%
    5GPT 5.5 (High)12.94%
  6. 6AnthropicClaude Opus 4.7 (Thinking)12.74%
    6AnthropicClaude Opus 4.7 (Thinking)12.74%
  7. 7GPT 5.511.46%
    7GPT 5.511.46%
  8. 8GPT 5.6 Luna (xHigh)11.30%
    8GPT 5.6 Luna (xHigh)11.30%
  9. 9AnthropicClaude Opus 4.611.10%
    9AnthropicClaude Opus 4.611.10%
  10. 10AnthropicClaude Sonnet 4.610.94%
    10AnthropicClaude Sonnet 4.610.94%
589,122 Sessions

Tool Hallucination

How much the model hallucinates tools it doesn't have.

  1. 1GPT 5.6 Terra (xHigh)1.18%
    1GPT 5.6 Terra (xHigh)1.18%
  2. 2GPT 5.6 Sol (xHigh)1.18%
    2GPT 5.6 Sol (xHigh)1.18%
  3. 3GPT 5.51.18%
    3GPT 5.51.18%
  4. 4GPT 5.6 Luna (xHigh)1.18%
    4GPT 5.6 Luna (xHigh)1.18%
  5. 5Kimi K2.7 Code1.18%
    5Kimi K2.7 Code1.18%
  6. 6Grok 4.51.18%
    6Grok 4.51.18%
  7. 7Kimi K2.61.18%
    7Kimi K2.61.18%
  8. 8Kimi K3 (Max)1.18%
    8Kimi K3 (Max)1.18%
  9. 9AnthropicClaude Fable 5 (High)1.18%
    9AnthropicClaude Fable 5 (High)1.18%
  10. 10GPT 5.5 (High)1.18%
    10GPT 5.5 (High)1.18%
1,988,518 Sessions

Frequently asked questions

Agent Mode

Try Agent Mode

Put these models to work on your own real tasks in Agent Mode.

Get started
How the Agent Leaderboard works

How the Agent Leaderboard works

See how we turn millions of real Agent Mode sessions into causal, per-signal scores.

Read the methodology

USE CASES

  • Chat with AI
  • Build Apps & Websites
  • Write & Edit Text
  • Search the Web
  • Generate Images
  • Generate Videos
  • Chose any model
  • Compare Models Side by Side

LEADERBOARD RANKINGS

  • Overall
  • Agent
  • Text
  • WebDev
  • Image-to-WebDev
  • Text to Image
  • Image Edit
  • Text to Video
  • Image to Video
  • Video Edit
  • Vision
  • Document
  • Search

COMPANY

  • About Us
  • How It Works
  • Blog
  • Careers
  • Changelog
  • Help Center
  • FAQ

LEGAL

  • Terms
  • Privacy
  • Cookies

FOLLOW

  • X
  • LinkedIn
  • YouTube
  • Discord

© Arena Intelligence 2026