Skip to main content
v2026.11,772 entries · CC-BY 4.0

Aggregated rankings

Best LLM Leaderboard — Top AI Models Ranked

Which model is actually best depends on the task. We pull live rankings from Artificial Analysis and Arena and break them out by intelligence, agentic/coding, and writing, so you can pick the model that fits the work rather than a single generic 'best AI' claim.

Last updated:

Editorial disclosure: Some links on this page are CASRAI referral links. If you sign up through one, CASRAI may earn a commission at no extra cost to you — this helps fund our nonprofit mission. We only recommend tools our editorial team has independently researched. Read our full disclosure policy →

Overall

Best overall / most intelligent

Artificial Analysis Intelligence Index — a composite reasoning benchmark across math, science, and general problem-solving.

#ModelVendorIntelligence Index
1Claude Opus 5 (Max)Anthropic63
1Claude Opus 5 (xHigh)Anthropic63
3Claude Fable 5 (with fallback)Anthropic62
4Claude Opus 5 (High)Anthropic61

Source: Artificial Analysis — Models leaderboard, retrieved 14 August 2026.

Agentic & coding

Best for agentic and coding workflows

Two independent measures: a composite benchmark index, and real-world agent-session outcomes.

Artificial Analysis Agentic Index

Measures tool use, planning, autonomy, and complex problem-solving via GDPval-AA v2 and 𝜏³-Banking.

#ModelVendorAgentic Index
1Claude Opus 5 (Adaptive Reasoning, Max)Anthropic59
1Grok 4.6 (High)xAI59
3Claude Opus 5 (Adaptive Reasoning, xHigh)Anthropic58

Source: Artificial Analysis — Agentic capability index, retrieved 14 August 2026.

Arena Agent leaderboard — real-world sessions

Net task-completion improvement across live coding-agent-style sessions (task confirmation, command recovery, tool-hallucination rate) — 1.79M+ sessions, 48 models scored.

#ModelVendorNet improvementSessions
1Claude Opus 5 (High)Anthropic+12.19%19,739 sessions
2Claude Fable 5 (High)Anthropic+12.01%24,417 sessions
3Claude Opus 5 (Max)Anthropic+11.95%15,515 sessions
4GPT 5.6 Sol (xHigh)OpenAI+10.86%18,091 sessions
5Kimi K3 (Max)Moonshot+10.60%28,369 sessions
6Claude Opus 4.8 (High)Anthropic+9.78%35,151 sessions
7GPT 5.5 (xHigh)OpenAI+8.90%47,586 sessions

Source: Arena — Agent leaderboard, retrieved 14 August 2026.

Writing

Best for writing

Arena's Text leaderboard — head-to-head human preference across math, coding, creative writing, and open-ended text.

#ModelVendorRating
1Claude Fable 5Anthropic1506 ±5
2Claude Opus 4.6 (High)Anthropic1505 ±4
3Claude Opus 4.7 (High)Anthropic1502 ±4
4Muse Spark 1.2 (xHigh)Meta1498 ±10
5Claude Opus 4.6Anthropic1497 ±3
6Claude Opus 4.7Anthropic1494 ±4
7Claude Opus 5 (High)Anthropic1493 ±5
8Qwen3.8 MaxAlibaba1491 ±8
9Gemini 3.7 Flash (High)Google1490 ±8
10Claude Opus 5 (Max)Anthropic1489 ±7

Source: Arena — Text leaderboard, retrieved 14 August 2026.

Speed & cost

Fastest and cheapest

For high-volume or latency-sensitive use, throughput and price per token often matter more than top-line intelligence score.

Fastest (output tokens/sec)

Celeris-1

1,592 t/s

Fastest, runner-up

Mercury 2

1,110 t/s

Lowest latency (time to first answer)

Command A+

0.40s

Cheapest (blended $/M tokens)

GPT-5.6 Luna (Low)

$0.01

Cheapest, runner-up

MiMo-V2.5

$0.01

Source: Artificial Analysis — Models leaderboard, retrieved 14 August 2026.

After you draft

Polish and verify the output

Whichever model you draft with, running the output through a dedicated paraphrasing pass and an AI-detection check before submission is standard practice for research writing — see our full AI disclosure guidance for what to disclose and when.

We don't link QuillBot's "AI Humanizer" product from CASRAI — a tool built to evade AI detection runs directly counter to the disclosure practices we recommend elsewhere on this site. See our commercial disclosure for how this and every other referral link on the page works.

Frequently asked questions

Common questions

What is the best LLM right now?
By raw intelligence benchmarks, Claude Opus 5 (Max/xHigh) and Claude Fable 5 currently lead Artificial Analysis's Intelligence Index at 62–63. "Best" depends heavily on the task, though — the model that tops a general reasoning benchmark isn't necessarily the best choice for coding, agentic tool-use, or long-form writing, which is why we break rankings out by category below rather than naming one universal winner.
What is the best LLM for agentic and coding tasks?
On Artificial Analysis's Agentic Index, Claude Opus 5 (Adaptive Reasoning, Max) and Grok 4.6 (High) are tied at the top (59), with Claude Opus 5 (xHigh) close behind (58). On real-world agent-session data from Arena's Agent leaderboard — which scores task completion, command recovery, and tool-hallucination rate across nearly 1.8M sessions — Claude Opus 5 (High) leads at +12.19% net improvement, followed by Claude Fable 5 (High) and Claude Opus 5 (Max).
Is Gemini 3.7 Flash good for agentic work?
It's a credible value pick rather than an outright leader. Gemini 3.7 Flash ranks #17 of 188 models on Artificial Analysis's Intelligence Index (score 56) but is the fastest model measured on that board at 340 output tokens/sec, and is priced at $0.75/$3.75 per million input/output tokens — well under the category median. Artificial Analysis specifically highlights its performance on agentic benchmarks (real-world work tasks, tool use, terminal operations). For high-volume agentic pipelines where throughput and cost matter as much as top-line intelligence score, it's a reasonable choice; for single-shot complex reasoning, the higher-ranked Claude and Grok models score higher.
What is the best LLM for writing?
Arena's Text leaderboard — which scores head-to-head human preference across math, coding, creative writing, and open-ended text — currently has Claude Fable 5 on top (1506), followed by Claude Opus 4.6 (High) and Claude Opus 4.7 (High). Whichever model drafts the text, we'd still recommend a dedicated paraphrasing and AI-detection pass before submission — see our tool picks below.
How often is this leaderboard updated?
We refresh this page against Artificial Analysis and Arena's live boards periodically; it was last checked 14 August 2026. Because model releases and re-evaluations happen continuously, treat the linked source boards as the live source of truth and this page as a dated snapshot.
Does CASRAI run its own LLM benchmarks?
No. Every ranking on this page is sourced and attributed to a third-party leaderboard (Artificial Analysis, Arena) rather than an in-house evaluation. We aggregate and contextualize published results for a research-and-procurement audience; we don't claim to have independently reproduced them.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.