API 同步 高可信
LMArena Text Style Control
开闭源混合大模型主榜;Style Control 用于降低回答格式和语气带来的人类偏好偏差。
- 指标
- 风格控制后的 Arena rating,越高越好
- 频率
- 每 30 分钟从官方 latest split 同步。
覆盖综合大语言模型、图片、视频、Coding、音频、OCR/文档和视觉模型的实用榜单页。优先使用公开 API 与 raw CSV,需要 Key 的来源只作交叉参考。
| 排名 | 模型 | 机构 | 分数 | 开放状态 | 证据 |
|---|---|---|---|---|---|
| 1 | claude-fable-5 Style-controlled Arena rating ↑ | anthropic | 1,509 | Proprietary | 17,799 票 来源 |
| 2 | claude-opus-4-6-thinking Style-controlled Arena rating ↑ | anthropic | 1,505 | Proprietary | 67,203 票 来源 |
| 3 | claude-opus-4-7-thinking Style-controlled Arena rating ↑ | anthropic | 1,502 | Proprietary | 54,781 票 来源 |
| 4 | claude-opus-4-6 Style-controlled Arena rating ↑ | anthropic | 1,497 | Proprietary | 70,970 票 来源 |
| 5 | qwen3.8-max Style-controlled Arena rating ↑ | alibaba | 1,496 | Proprietary | 3,327 票 来源 |
| 6 | claude-opus-4-7 Style-controlled Arena rating ↑ | anthropic | 1,492 | Proprietary | 55,992 票 来源 |
| 7 | claude-opus-5-high Style-controlled Arena rating ↑ | anthropic | 1,492 | Proprietary | 10,704 票 来源 |
| 8 | claude-opus-5-max Style-controlled Arena rating ↑ | anthropic | 1,490 | Proprietary | 5,124 票 来源 |
| 9 | muse-spark-1.1 Style-controlled Arena rating ↑ | meta | 1,490 | Proprietary | 12,086 票 来源 |
| 10 | muse-spark Style-controlled Arena rating ↑ | meta | 1,488 | Proprietary | 13,486 票 来源 |
| 11 | gemini-3-pro Style-controlled Arena rating ↑ | 1,486 | Proprietary | 41,241 票 来源 | |
| 12 | gemini-3.1-pro-preview Style-controlled Arena rating ↑ | 1,485 | Proprietary | 89,133 票 来源 | |
| 13 | claude-opus-4-8-thinking Style-controlled Arena rating ↑ | anthropic | 1,484 | Proprietary | 35,303 票 来源 |
| 14 | gpt-5.6-sol-xhigh Style-controlled Arena rating ↑ | openai | 1,483 | Proprietary | 10,716 票 来源 |
| 15 | gemini-3.6-flash Style-controlled Arena rating ↑ | 1,483 | Proprietary | 8,575 票 来源 | |
| 16 | gpt-5.5-high Style-controlled Arena rating ↑ | openai | 1,482 | Proprietary | 49,978 票 来源 |
| 17 | gpt-5.4-high Style-controlled Arena rating ↑ | openai | 1,477 | Proprietary | 60,264 票 来源 |
| 18 | gpt-5.5 Style-controlled Arena rating ↑ | openai | 1,476 | Proprietary | 51,172 票 来源 |
| 19 | gpt-5.2-chat-latest-20260210 Style-controlled Arena rating ↑ | openai | 1,476 | Proprietary | 34,036 票 来源 |
| 20 | gemini-3.5-flash-high Style-controlled Arena rating ↑ | 1,475 | Proprietary | 10,472 票 来源 | |
| 21 | qwen3.7-max-preview Style-controlled Arena rating ↑ | alibaba | 1,475 | Proprietary | 3,695 票 来源 |
| 22 | claude-opus-4-8 Style-controlled Arena rating ↑ | anthropic | 1,475 | Proprietary | 35,877 票 来源 |
| 23 | grok-4.20-beta1 Style-controlled Arena rating ↑ | xai | 1,474 | Proprietary | 26,576 票 来源 |
| 24 | gemini-3.5-flash-medium Style-controlled Arena rating ↑ | 1,474 | Proprietary | 18,778 票 来源 | |
| 25 | gpt-5.5-instant Style-controlled Arena rating ↑ | openai | 1,473 | Proprietary | 25,680 票 来源 |
每个来源都标注同步方式和可信度,方便 Codex 后续决定哪些自动抓取、哪些只做参考。
开闭源混合大模型主榜;Style Control 用于降低回答格式和语气带来的人类偏好偏差。
适合交叉核对综合智能、价格与速度;官方 API 需要 Key,且必须只放在服务端。
适合作为开闭源混合图片生成排行的主同步源。
补充图片编辑模型,例如 GPT Image、Gemini 图像、Seedream 等。
文生视频主同步源,覆盖字节、快手、xAI、Google 等厂商。
图生视频工作流的补充同步源。
真实 GitHub issue 修复类 Agent Coding 的最佳结构化来源。
补充 SWE-bench,覆盖网页开发偏好排行。
语音识别强来源;音频生成榜单仍需要二级人工/解析来源补充。
适合 OCR 和文字密集视觉理解;建议同步到 D1 后做中等可信展示。
中文 OCR 与文档理解补充榜单。
文档类多模态任务的可同步来源。
通用视觉推理和多模态模型对比。
公开质量较高,但 HTML 解析比 benchmark API 更脆弱。
覆盖 C++、Go、Java、JavaScript、Python、Rust 的代码编辑榜单,但暂未发现稳定 JSON。