Live model ranking radar

ModelRank

A practical ranking page for general LLMs, image, video, coding, audio, OCR/document, and vision models. It prioritizes public APIs and raw CSVs, then uses gated sources as cross-checks.

Source LMArena Text Style Control Upstream updated: Aug 3, 2026 · Fetched: Aug 4, 2026
Original source

Ranking source

Rank Model Organization Score Access Evidence
1
claude-fable-5 Style-controlled Arena rating ↑
anthropic 1,509 Proprietary 17,799 votes Source
2
claude-opus-4-6-thinking Style-controlled Arena rating ↑
anthropic 1,505 Proprietary 67,203 votes Source
3
claude-opus-4-7-thinking Style-controlled Arena rating ↑
anthropic 1,502 Proprietary 54,781 votes Source
4
claude-opus-4-6 Style-controlled Arena rating ↑
anthropic 1,497 Proprietary 70,970 votes Source
5
qwen3.8-max Style-controlled Arena rating ↑
alibaba 1,496 Proprietary 3,327 votes Source
6
claude-opus-4-7 Style-controlled Arena rating ↑
anthropic 1,492 Proprietary 55,992 votes Source
7
claude-opus-5-high Style-controlled Arena rating ↑
anthropic 1,492 Proprietary 10,704 votes Source
8
claude-opus-5-max Style-controlled Arena rating ↑
anthropic 1,490 Proprietary 5,124 votes Source
9
muse-spark-1.1 Style-controlled Arena rating ↑
meta 1,490 Proprietary 12,086 votes Source
10
muse-spark Style-controlled Arena rating ↑
meta 1,488 Proprietary 13,486 votes Source
11
gemini-3-pro Style-controlled Arena rating ↑
google 1,486 Proprietary 41,241 votes Source
12
gemini-3.1-pro-preview Style-controlled Arena rating ↑
google 1,485 Proprietary 89,133 votes Source
13
claude-opus-4-8-thinking Style-controlled Arena rating ↑
anthropic 1,484 Proprietary 35,303 votes Source
14
gpt-5.6-sol-xhigh Style-controlled Arena rating ↑
openai 1,483 Proprietary 10,716 votes Source
15
gemini-3.6-flash Style-controlled Arena rating ↑
google 1,483 Proprietary 8,575 votes Source
16
gpt-5.5-high Style-controlled Arena rating ↑
openai 1,482 Proprietary 49,978 votes Source
17
gpt-5.4-high Style-controlled Arena rating ↑
openai 1,477 Proprietary 60,264 votes Source
18
gpt-5.5 Style-controlled Arena rating ↑
openai 1,476 Proprietary 51,172 votes Source
19
gpt-5.2-chat-latest-20260210 Style-controlled Arena rating ↑
openai 1,476 Proprietary 34,036 votes Source
20
gemini-3.5-flash-high Style-controlled Arena rating ↑
google 1,475 Proprietary 10,472 votes Source
21
qwen3.7-max-preview Style-controlled Arena rating ↑
alibaba 1,475 Proprietary 3,695 votes Source
22
claude-opus-4-8 Style-controlled Arena rating ↑
anthropic 1,475 Proprietary 35,877 votes Source
23
grok-4.20-beta1 Style-controlled Arena rating ↑
xai 1,474 Proprietary 26,576 votes Source
24
gemini-3.5-flash-medium Style-controlled Arena rating ↑
google 1,474 Proprietary 18,778 votes Source
25
gpt-5.5-instant Style-controlled Arena rating ↑
openai 1,473 Proprietary 25,680 votes Source

Data sources

Every source below is tagged by sync method and confidence so future Codex updates can decide what to automate or keep as a reference.

API sync High confidence

LMArena Text Style Control

Primary mixed open/closed LLM ranking. Style control reduces preference bias caused by response formatting and tone.

Metric
Style-controlled Arena rating, higher is better
Cadence
Sync every 30 minutes from the official latest split.
Open and closed models mixed Original source
Manual High confidence

Artificial Analysis Intelligence Index

Strong benchmark-based cross-check for intelligence, price, and speed. Its official API requires an API key and must remain server-side.

Metric
Intelligence Index, higher is better
Cadence
API-ready after a server-side Artificial Analysis key is configured.
Open and closed models mixed Original source
API sync High confidence

LMArena Text-to-Image

Best automatic source for mixed open/closed image-generation rankings.

Metric
Arena rating, higher is better
Cadence
Sync every 30 minutes; upstream latest split updates when LMArena republishes.
Open and closed models mixed Original source
API sync High confidence

LMArena Image Edit

Complements text-to-image with editing models such as GPT Image, Gemini image, Seedream, and others.

Metric
Arena rating, higher is better
Cadence
Sync every 30 minutes.
Open and closed models mixed Original source
API sync High confidence

LMArena Text-to-Video

Good primary source for text-to-video, including ByteDance, Kuaishou, xAI, Google, and other providers.

Metric
Arena rating, higher is better
Cadence
Sync every 30 minutes.
Open and closed models mixed Original source
API sync High confidence

LMArena Image-to-Video

Secondary video source for image-to-video workflows.

Metric
Arena rating, higher is better
Cadence
Sync every 30 minutes.
Open and closed models mixed Original source
API sync High confidence

SWE-bench Verified

Best structured source for agentic coding on real GitHub issues.

Metric
Resolved percentage, higher is better
Cadence
Sync hourly; source is a Hugging Face benchmark leaderboard API.
Open and closed models mixed Original source
API sync High confidence

LMArena WebDev

Complements SWE-bench with web-development preference rankings.

Metric
Arena rating, higher is better
Cadence
Sync every 30 minutes.
Open and closed models mixed Original source
API sync High confidence

Open ASR Leaderboard

Strong source for speech recognition. Audio generation leaderboards still need a secondary curated source.

Metric
Average WER, lower is better
Cadence
Sync daily or hourly; benchmark API is structured and reproducible.
Open and closed models mixed Original source
CSV sync Medium confidence

OCRBench v2 English

Useful for OCR and text-centric visual understanding; treat as medium-confidence until mirrored into D1 snapshots.

Metric
Average score, higher is better
Cadence
Sync daily from raw CSV; upstream cadence is less formal than HF benchmark API.
Open and closed models mixed Original source
CSV sync Medium confidence

OCRBench v2 Chinese

Chinese OCR and document-understanding companion table.

Metric
Average score, higher is better
Cadence
Sync daily from raw CSV.
Open and closed models mixed Original source
API sync High confidence

LMArena Document

Good live source for document-style multimodal tasks.

Metric
Arena rating, higher is better
Cadence
Sync every 30 minutes.
Open and closed models mixed Original source
API sync High confidence

LMArena Vision

Broad visual reasoning and multimodal model comparison.

Metric
Arena rating, higher is better
Cadence
Sync every 30 minutes.
Open and closed models mixed Original source
HTML watch Watch source

Artificial Analysis Video

High-quality public leaderboard, but parsing HTML is more fragile than using benchmark APIs.

Metric
Video Arena ELO, higher is better
Cadence
Use as a cross-check unless a stable public API is available.
Open and closed models mixed Original source
HTML watch Watch source

Aider Polyglot

Good benchmark for editing across C++, Go, Java, JavaScript, Python, and Rust, but no stable public JSON found.

Metric
Pass rate, higher is better
Cadence
Manual or HTML-parse fallback; useful as a code-editing cross-check.
Open and closed models mixed Original source