Evidence hub

AI benchmarks, with the footnotes left in.

Model version, date, harness, source, and limitation—all visible. We use benchmarks as supporting evidence, never as a substitute for workflow fit.

Arena snapshot: 2026-09-02 · Evidence reviewed: 2026-09-07

Benchmark evidence policy

Benchmarks are supporting evidence, not our final ranking. We identify the exact model/version, source, snapshot date and harness. Product scores are not treated as model scores, and results from different harness versions are not compared directly.

Exact model/versionSnapshot dateHarness disclosedSource linkedStale after 30 days

Current snapshot

Arena Text: representative model listings

Arena Text is a public leaderboard built from anonymous, pairwise human preference battles. Rank is the model's position in that snapshot; Arena rating is the statistical preference score; ± value is the approximate 95% confidence interval around that rating; and preference votes are the recorded battle outcomes contributing to the snapshot—not votes for the product alone.

Interpretation
Grok grok-4.1#531459±3 (95% CI) 66,382Representative public xAI model listing; not necessarily the product's current default. [1]
DeepSeek deepseek-v4-pro#551458±4 (95% CI) 54,225Representative public DeepSeek model listing. [1]
Claude claude-sonnet-4-5-20250929#621455±3 (95% CI) 79,668Older public Claude listing retained for reproducible snapshot context; newer variants may rank differently. [1]
Gemini gemini-2.5-pro#751446±2 (95% CI) 122,612Specific Gemini model version, not the whole Gemini product. [1]
ChatGPT chatgpt-4o-latest-20250326#781443±3 (95% CI) 80,696Historical public ChatGPT model listing; do not interpret as the current default model's rank. [1]
Mistral Le Chat mistral-medium-3.5#1031427±7 (95% CI) 11,017Representative Mistral model listing with a wider confidence interval. [1]

Download current snapshot CSV

Read this correctly: Arena ratings summarize human preference in anonymous pairwise battles. A higher rank does not prove better factuality, lower cost, stronger privacy, or a better end-user product. Methodology [2] ↗

Snapshot change

What moved since 2026-08-31.

All six tracked Arena ratings were unchanged. Five ranks moved down by 1–3 places as the leaderboard changed; Claude stayed at #62. This is leaderboard drift, not evidence that the products became worse.

Product familyExact modelRankVote changeArena rating
Grokgrok-4.1#51 → #53 (down 2)+8unchanged
DeepSeekdeepseek-v4-pro#54 → #55 (down 1)+11unchanged
Claudeclaude-sonnet-4-5-20250929#62 → #62 (unchanged)-19unchanged
Geminigemini-2.5-pro#73 → #75 (down 2)+23unchanged
ChatGPTchatgpt-4o-latest-20250326#76 → #78 (down 2)+9unchanged
Mistral Le Chatmistral-medium-3.5#100 → #103 (down 3)+11unchanged

Buying decision: Do not add, cancel, or switch a subscription because a model moved a few leaderboard places while its rating stayed fixed. Re-run the same real task in the products you can actually buy.

Coding-agent evidence

Keep the full configuration attached.

A coding score belongs to the agent, exact model, reasoning effort, harness version, trial policy, and integrity checks—not to one product name.

Verified agent configuration

Cursor · Terminal-Bench 2.1

79.3% ± 1.5% · Cursor CLI + Grok 4.5 · high effort · 445 trials · run 2026-07-09

Integrity metadata: 9.0% reward-hack disqualifications.

Maintainer-verified configuration snapshot. Agent, model, effort and integrity metadata are inseparable from the score.

Open submission record [10] ↗

Source registry

What we trust—and what each source misses.

[1]

Arena Text Leaderboard

Measures: Human preference in anonymous model-vs-model battles across open-ended text tasks.

Use it for: General assistants and model-family context

Do not miss: A product such as ChatGPT or Claude can change its default model; Arena ranks specific model versions, not the whole product.

Open source ↗
[2]

Arena-Rank methodology

Measures: Bradley-Terry ratings with reweighting and confidence intervals from pairwise preference data.

Use it for: Explaining how Arena scores are produced

Do not miss: Human preference can reward style and verbosity; it is not an objective factuality test.

Open source ↗
[3]

SWE-bench Verified

Measures: Percentage of 500 human-validated real GitHub issues resolved by a coding model or agent.

Use it for: Coding models and coding-agent context

Do not miss: Scores depend heavily on the agent harness. mini-SWE-agent v1.x and v2.x results are not directly comparable.

Open source ↗
[4]

LiveBench

Measures: Objective, frequently refreshed tasks across reasoning, math, coding, language, instruction following and data analysis.

Use it for: Contamination-limited general-model context

Do not miss: Always state the release; questions and model configurations change between snapshots.

Open source ↗
[5]

Artificial Analysis methodology

Measures: Independent intelligence, latency, speed, price, speech, image and video model evaluations.

Use it for: Cost/speed context and non-text modality methodology

Do not miss: API model measurements do not equal the complete end-user product experience.

Open source ↗
[6]

Terminal-Bench 2.1

Measures: Accuracy of an agent + model + reasoning-effort configuration on 89 complex containerized terminal tasks.

Use it for: Coding agents and terminal-capable development tools

Do not miss: Never present this as a model-only score. Keep agent, exact model, effort, trial count and integrity metadata visible.

Open source ↗
[7]

Artificial Analysis Image methodology

Measures: Blind pairwise image preference with exact model/version and matched prompt conditions.

Use it for: Text-to-image and image-editing model context

Do not miss: Keep text-to-image and editing pools separate; old model versions are historical evidence only.

Open source ↗
[8]

Artificial Analysis Video methodology

Measures: Modality-separated video quality, latency and price comparisons.

Use it for: Text-to-video and image-to-video model context

Do not miss: Do not compare Elo across different video modalities or output settings.

Open source ↗
[9]

Artificial Analysis TTS methodology

Measures: Provider-voice and controlled-voice human-preference comparisons for text-to-speech.

Use it for: Voice-generation model context

Do not miss: Controlled-voice and provider-voice arenas answer different questions and must remain separate.

Open source ↗
[10]

Terminal-Bench: Cursor CLI + Grok 4.5 submission

Measures: Maintainer-verified Terminal-Bench 2.1 submission record for one exact Cursor CLI configuration.

Use it for: A configuration-specific Cursor benchmark snapshot

Do not miss: This is Cursor CLI + Grok 4.5 at high effort, not a universal score for Cursor or Grok.

Open source ↗

Where benchmarks belong

Match the evidence to the buying decision.

01

General assistants

Arena preference + LiveBench objective tasks + price/latency context. Never collapse them into one magic number.

02

Coding assistants

SWE-bench only when the exact model, agent harness and release are disclosed. Product UX still needs separate evaluation.

03

Image, video & audio

Use modality-specific quality and latency methods. Product outputs, rights, editing control, and consistency matter more than text benchmarks.

Keep/Cut Weekly