🏟️ Arena AI 排行榜
LMSYS 团队运营的真实人类盲投 ELO 榜 · 6 大维度 · 23,792,030 投票 · 自动 1 小时刷新
实时抓取缓存 1h
💬 Text Leaderboard · Top 30
通用文本对话 ELO 榜(盲投最多) · 384 模型(榜单共 384 名)·11,137,435 票
| Rank | Model | Provider | Score (95% CI) | Votes | $ / M | Context |
|---|---|---|---|---|---|---|
| 1 | claude-fable-5 Anthropic | Anthropic | 1509 ±6 | 17,799 | $10 / $50 | 1M |
| 2 | claude-opus-4-6-thinking Anthropic | Anthropic | 1505 ±4 | 67,203 | $5 / $25 | 1M |
| 3 | claude-opus-4-7-thinking Anthropic | Anthropic | 1502 ±4 | 54,781 | $5 / $25 | 1M |
| 4 | claude-opus-4-6 Anthropic | Anthropic | 1497 ±4 | 70,970 | $5 / $25 | 1M |
| 5 | claude-opus-4-7 Anthropic | Anthropic | 1492 ±4 | 55,992 | $5 / $25 | 1M |
| 6 | claude-opus-5-high Anthropic | Anthropic | 1492 ±6 | 10,704 | $5 / $25 | 1M |
| 7 | claude-opus-5-max Anthropic | Anthropic | 1490 ±9 | 5,124 | $5 / $25 | 1M |
| 8 | muse-spark-1.1 Meta | Meta | 1490 ±6 | 12,086 | $1.25 / $4.25 | N/A |
| 9 | muse-spark Meta | Meta | 1488 ±6 | 13,486 | N/A | N/A |
| 10 | gemini-3-pro | 1486 ±4 | 41,241 | $2 / $12 | 1M | |
| 11 | gemini-3.1-pro-preview | 1485 ±3 | 89,133 | $2 / $12 | 1M | |
| 12 | Moonshot | kimi-k3-max | 1485 ±10 | 3,556 | $3 / $15 | 1M |
| 13 | claude-opus-4-8-thinking Anthropic | Anthropic | 1484 ±5 | 35,303 | $5 / $25 | 1M |
| 14 | OpenAI | gpt-5.6-sol-xhigh | 1483 ±6 | 10,716 | N/A | N/A |
| 15 | gemini-3.6-flash | 1483 ±7 | 8,575 | $1.50 / $7.50 | 1M | |
| 16 | OpenAI | gpt-5.5-high | 1482 ±4 | 49,978 | $5 / $30 | 1.1M |
| 17 | OpenAI | gpt-5.4-high | 1477 ±4 | 60,264 | $2.50 / $15 | 1.1M |
| 18 | gemini-3.5-flash-high | 1476 ±7 | 10,006 | $1.50 / $9 | 1M | |
| 19 | OpenAI | gpt-5.5 | 1476 ±4 | 51,172 | $5 / $30 | 1.1M |
| 20 | OpenAI | gpt-5.2-chat-latest-20260210 | 1476 ±4 | 34,036 | $1.75 / $14 | 128K |
| 21 | Alibaba | qwen3.7-max-preview | 1475 ±10 | 3,695 | $1.48 / $4.42 | 1M |
| 22 | claude-opus-4-8 Anthropic | Anthropic | 1475 ±5 | 35,877 | $5 / $25 | 1M |
| 23 | SpaceXAI | grok-4.20-beta1 | 1474 ±5 | 26,576 | N/A | N/A |
| 24 | gemini-3.5-flash-medium | 1474 ±5 | 18,778 | $1.50 / $9 | 1M | |
| 25 | OpenAI | gpt-5.5-instant | 1473 ±5 | 25,680 | $5 / $30 | 1.1M |
| 26 | gemini-3-flash | 1473 ±4 | 30,635 | $0.50 / $3 | 1M | |
| 27 | claude-opus-4-5-20251101-thinking-32k Anthropic | Anthropic | 1473 ±4 | 36,973 | $5 / $25 | 200K |
| 28 | SpaceXAI | grok-4.20-beta-0309-reasoning | 1472 ±4 | 61,765 | $2 / $6 | 2M |
| 29 | claude-sonnet-4-6 Anthropic | Anthropic | 1472 ±4 | 61,253 | $3 / $15 | 1M |
| 30 | SpaceXAI | grok-4.20-multi-agent-beta-0309 | 1471 ±4 | 60,405 | $2 / $6 | 2M |
关于数据:Arena.ai 排行榜基于人类对模型两两盲投结果,由 LMSYS Chatbot Arena 团队运营,是 LLM 评测领域公信力最高的 ELO 排名之一。所有分数均含 95% 置信区间,投票数越大分数越稳定。本页 实时抓取 arena.ai/leaderboard/ 6 个端点(text / webdev / vision / agent / search / text-to-image),缓存 1 小时。