[1]
Arena Text Leaderboard
Measures: Human preference in anonymous model-vs-model battles across open-ended text tasks.
Use it for: General assistants and model-family context
Do not miss: A product such as ChatGPT or Claude can change its default model; Arena ranks specific model versions, not the whole product.
Open source ↗
[2]
Arena-Rank methodology
Measures: Bradley-Terry ratings with reweighting and confidence intervals from pairwise preference data.
Use it for: Explaining how Arena scores are produced
Do not miss: Human preference can reward style and verbosity; it is not an objective factuality test.
Open source ↗
[3]
SWE-bench Verified
Measures: Percentage of 500 human-validated real GitHub issues resolved by a coding model or agent.
Use it for: Coding models and coding-agent context
Do not miss: Scores depend heavily on the agent harness. mini-SWE-agent v1.x and v2.x results are not directly comparable.
Open source ↗
[4]
LiveBench
Measures: Objective, frequently refreshed tasks across reasoning, math, coding, language, instruction following and data analysis.
Use it for: Contamination-limited general-model context
Do not miss: Always state the release; questions and model configurations change between snapshots.
Open source ↗
[5]
Artificial Analysis methodology
Measures: Independent intelligence, latency, speed, price, speech, image and video model evaluations.
Use it for: Cost/speed context and non-text modality methodology
Do not miss: API model measurements do not equal the complete end-user product experience.
Open source ↗
[6]
Terminal-Bench 2.1
Measures: Accuracy of an agent + model + reasoning-effort configuration on 89 complex containerized terminal tasks.
Use it for: Coding agents and terminal-capable development tools
Do not miss: Never present this as a model-only score. Keep agent, exact model, effort, trial count and integrity metadata visible.
Open source ↗
[7]
Artificial Analysis Image methodology
Measures: Blind pairwise image preference with exact model/version and matched prompt conditions.
Use it for: Text-to-image and image-editing model context
Do not miss: Keep text-to-image and editing pools separate; old model versions are historical evidence only.
Open source ↗
[8]
Artificial Analysis Video methodology
Measures: Modality-separated video quality, latency and price comparisons.
Use it for: Text-to-video and image-to-video model context
Do not miss: Do not compare Elo across different video modalities or output settings.
Open source ↗
[9]
Artificial Analysis TTS methodology
Measures: Provider-voice and controlled-voice human-preference comparisons for text-to-speech.
Use it for: Voice-generation model context
Do not miss: Controlled-voice and provider-voice arenas answer different questions and must remain separate.
Open source ↗
[10]
Terminal-Bench: Cursor CLI + Grok 4.5 submission
Measures: Maintainer-verified Terminal-Bench 2.1 submission record for one exact Cursor CLI configuration.
Use it for: A configuration-specific Cursor benchmark snapshot
Do not miss: This is Cursor CLI + Grok 4.5 at high effort, not a universal score for Cursor or Grok.
Open source ↗