Aggregated rankings
Best LLM Leaderboard — Top AI Models Ranked
Which model is actually best depends on the task. We pull live rankings from Artificial Analysis and Arena and break them out by intelligence, agentic/coding, and writing, so you can pick the model that fits the work rather than a single generic 'best AI' claim.
Last updated:
Overall
Best overall / most intelligent
Artificial Analysis Intelligence Index — a composite reasoning benchmark across math, science, and general problem-solving.
| # | Model | Vendor | Intelligence Index |
|---|---|---|---|
| 1 | Claude Opus 5 (Max) | Anthropic | 63 |
| 1 | Claude Opus 5 (xHigh) | Anthropic | 63 |
| 3 | Claude Fable 5 (with fallback) | Anthropic | 62 |
| 4 | Claude Opus 5 (High) | Anthropic | 61 |
Source: Artificial Analysis — Models leaderboard, retrieved 14 August 2026.
Agentic & coding
Best for agentic and coding workflows
Two independent measures: a composite benchmark index, and real-world agent-session outcomes.
Artificial Analysis Agentic Index
Measures tool use, planning, autonomy, and complex problem-solving via GDPval-AA v2 and 𝜏³-Banking.
| # | Model | Vendor | Agentic Index |
|---|---|---|---|
| 1 | Claude Opus 5 (Adaptive Reasoning, Max) | Anthropic | 59 |
| 1 | Grok 4.6 (High) | xAI | 59 |
| 3 | Claude Opus 5 (Adaptive Reasoning, xHigh) | Anthropic | 58 |
Source: Artificial Analysis — Agentic capability index, retrieved 14 August 2026.
Arena Agent leaderboard — real-world sessions
Net task-completion improvement across live coding-agent-style sessions (task confirmation, command recovery, tool-hallucination rate) — 1.79M+ sessions, 48 models scored.
| # | Model | Vendor | Net improvement | Sessions |
|---|---|---|---|---|
| 1 | Claude Opus 5 (High) | Anthropic | +12.19% | 19,739 sessions |
| 2 | Claude Fable 5 (High) | Anthropic | +12.01% | 24,417 sessions |
| 3 | Claude Opus 5 (Max) | Anthropic | +11.95% | 15,515 sessions |
| 4 | GPT 5.6 Sol (xHigh) | OpenAI | +10.86% | 18,091 sessions |
| 5 | Kimi K3 (Max) | Moonshot | +10.60% | 28,369 sessions |
| 6 | Claude Opus 4.8 (High) | Anthropic | +9.78% | 35,151 sessions |
| 7 | GPT 5.5 (xHigh) | OpenAI | +8.90% | 47,586 sessions |
Source: Arena — Agent leaderboard, retrieved 14 August 2026.
Writing
Best for writing
Arena's Text leaderboard — head-to-head human preference across math, coding, creative writing, and open-ended text.
| # | Model | Vendor | Rating |
|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 1506 ±5 |
| 2 | Claude Opus 4.6 (High) | Anthropic | 1505 ±4 |
| 3 | Claude Opus 4.7 (High) | Anthropic | 1502 ±4 |
| 4 | Muse Spark 1.2 (xHigh) | Meta | 1498 ±10 |
| 5 | Claude Opus 4.6 | Anthropic | 1497 ±3 |
| 6 | Claude Opus 4.7 | Anthropic | 1494 ±4 |
| 7 | Claude Opus 5 (High) | Anthropic | 1493 ±5 |
| 8 | Qwen3.8 Max | Alibaba | 1491 ±8 |
| 9 | Gemini 3.7 Flash (High) | 1490 ±8 | |
| 10 | Claude Opus 5 (Max) | Anthropic | 1489 ±7 |
Source: Arena — Text leaderboard, retrieved 14 August 2026.
Speed & cost
Fastest and cheapest
For high-volume or latency-sensitive use, throughput and price per token often matter more than top-line intelligence score.
Fastest (output tokens/sec)
Celeris-1
1,592 t/s
Fastest, runner-up
Mercury 2
1,110 t/s
Lowest latency (time to first answer)
Command A+
0.40s
Cheapest (blended $/M tokens)
GPT-5.6 Luna (Low)
$0.01
Cheapest, runner-up
MiMo-V2.5
$0.01
Source: Artificial Analysis — Models leaderboard, retrieved 14 August 2026.
After you draft
Polish and verify the output
Whichever model you draft with, running the output through a dedicated paraphrasing pass and an AI-detection check before submission is standard practice for research writing — see our full AI disclosure guidance for what to disclose and when.
We don't link QuillBot's "AI Humanizer" product from CASRAI — a tool built to evade AI detection runs directly counter to the disclosure practices we recommend elsewhere on this site. See our commercial disclosure for how this and every other referral link on the page works.
Frequently asked questions
Common questions
- What is the best LLM right now?
- By raw intelligence benchmarks, Claude Opus 5 (Max/xHigh) and Claude Fable 5 currently lead Artificial Analysis's Intelligence Index at 62–63. "Best" depends heavily on the task, though — the model that tops a general reasoning benchmark isn't necessarily the best choice for coding, agentic tool-use, or long-form writing, which is why we break rankings out by category below rather than naming one universal winner.
- What is the best LLM for agentic and coding tasks?
- On Artificial Analysis's Agentic Index, Claude Opus 5 (Adaptive Reasoning, Max) and Grok 4.6 (High) are tied at the top (59), with Claude Opus 5 (xHigh) close behind (58). On real-world agent-session data from Arena's Agent leaderboard — which scores task completion, command recovery, and tool-hallucination rate across nearly 1.8M sessions — Claude Opus 5 (High) leads at +12.19% net improvement, followed by Claude Fable 5 (High) and Claude Opus 5 (Max).
- Is Gemini 3.7 Flash good for agentic work?
- It's a credible value pick rather than an outright leader. Gemini 3.7 Flash ranks #17 of 188 models on Artificial Analysis's Intelligence Index (score 56) but is the fastest model measured on that board at 340 output tokens/sec, and is priced at $0.75/$3.75 per million input/output tokens — well under the category median. Artificial Analysis specifically highlights its performance on agentic benchmarks (real-world work tasks, tool use, terminal operations). For high-volume agentic pipelines where throughput and cost matter as much as top-line intelligence score, it's a reasonable choice; for single-shot complex reasoning, the higher-ranked Claude and Grok models score higher.
- What is the best LLM for writing?
- Arena's Text leaderboard — which scores head-to-head human preference across math, coding, creative writing, and open-ended text — currently has Claude Fable 5 on top (1506), followed by Claude Opus 4.6 (High) and Claude Opus 4.7 (High). Whichever model drafts the text, we'd still recommend a dedicated paraphrasing and AI-detection pass before submission — see our tool picks below.
- How often is this leaderboard updated?
- We refresh this page against Artificial Analysis and Arena's live boards periodically; it was last checked 14 August 2026. Because model releases and re-evaluations happen continuously, treat the linked source boards as the live source of truth and this page as a dated snapshot.
- Does CASRAI run its own LLM benchmarks?
- No. Every ranking on this page is sourced and attributed to a third-party leaderboard (Artificial Analysis, Arena) rather than an in-house evaluation. We aggregate and contextualize published results for a research-and-procurement audience; we don't claim to have independently reproduced them.








