分享一个刚整理的榜单合集:6 个主流大模型评测榜,我把每家官方最新的头部数据都翻出来了,排个队给大家看个热闹。数据快照基本都在 2026 年 8 月中下旬,更新也算及时:LMArena 08-12 / AA 指数 08-06 / HLE 08-19 / LiveBench 06-25 / GPQA 08-14 / SWE 08-18。

先看六榜榜首

榜单榜首厂商分数
LMArena 综合Claude Fable 5AnthropicElo 1507
AA 智能指数Claude Opus 5Anthropic63
HLE 人类最后的考试Claude Fable 5Anthropic55.5%
LiveBenchClaude Fable 5 MaxAnthropic83.0
GPQA DiamondGemini 3.7 FlashGoogle94.8%
SWE-bench VerifiedClaude Opus 5Anthropic96%

六榜里 Anthropic 拿了五个第一,只剩 GPQA 被 Google 抢走——这轮的格局相当清晰。

LMArena:真人盲测偏好榜(节选 Top 20)

靠真实用户盲测投票跑出来的榜,最反映"人觉得谁好用"。Elo 归一化到 0-100:

#模型Elo
1claude-fable-5100
2claude-opus-4-6-high99.7
3claude-opus-4-7-high99.2
4muse-spark-1.2 (xHigh)98.6
5claude-opus-4-698.3
6claude-opus-4-797.7
6claude-opus-5-high97.7
8claude-opus-5-max97.2
9qwen3.8-max97.1
10muse-spark-1.196.8
10kimi-k3-max96.8
12muse-spark96.6
13gemini-3.1-pro-preview96.3
14gemini-3-pro96.2
15gemini-3.6-flash-high96
16claude-opus-4-8-high95.5
16gpt-5.5-high95.5
18gpt-5.6-sol-xhigh95.4
19gemini-3.5-flash-high94.7
20gpt-5.594.6

头几名咬得很紧,Elo 差距都在个位数。

AA 智能指数:独立机构的能力标尺(节选 Top 15)

Artificial Analysis 用九个难任务加权合成 0-100 智能指数,比单看刷题靠谱。全表 265 个模型、257 个有分:

#模型指数
1Claude Opus 5 (max)63
1Claude Opus 5 (xhigh)63
3Claude Fable 5 (with fallback)62
4Claude Opus 5 (high)61
4GPT-5.6 Sol (max)61
4Grok 4.6 (high)61
7Kimi K3 (max)60
7GLM-5.3 (max)60
9GPT-5.6 Sol (xhigh)59
9Claude Opus 5 (medium)59
11Qwen3.8 Max58
11Qwen3.8 2.4T A95B58
13GPT-5.6 Sol (high)57
13Muse Spark 1.2 (xhigh)57
13GPT-5.6 Terra (max)57

Opus 5 拿到 63,上 6 字头的独一档。

HLE:人类最后的考试(31 个有成绩模型全列)

2500 道连人都难答的专家题,模型普遍翻车,中位数只有个位数。满分 100%:

#模型得分
1Claude Fable 5 (with fallback)55.5%
2Claude Opus 5 (max)54.9%
3Claude Opus 5 (xhigh)54.4%
4Claude Opus 5 (high)52.8%
5Claude Opus 5 (medium)51.3%
6GPT-5.6 Sol (max)49.5%
7Gemini 3.7 Flash (high)47.9%
8Kimi K3 (max)46.9%
9Muse Spark 1.2 (xhigh)45.5%
10Qwen3.8 Max43%

榜首 55.5%,离 60% 都还差得远。这榜确实是给 AI 泼冷水专用。

LiveBench:客观题综合(44 模型全列,节选 Top 15)

Abacus.AI 出品,每题都有标准答案,满分 100:

#模型综合分
1Claude Fable 5 Max Effort83.0
2GPT-5.6 Sol Max Effort81
3GPT-5.5 Thinking xHigh Effort80.2
4Claude 5 Opus Thinking Max Effort80.1
5Smaug-Agentic79.5
6Kimi K379.2
7Gemini 3.7 Flash High78.8
8Qwen 3.8 Max78.5
9Grok 4.678
10GPT-5.4 Thinking xHigh Effort78
11Muse Spark 1.2 xHigh Effort78
12GPT-5.6 Terra Max Effort77.9
13DeepSeek V4 Pro 081377.4
14Gemini 3.1 Pro Preview High77
15Claude 4.7 Opus Thinking xHigh Effort76.5

GPQA Diamond:榜首换人了(节选 Top 15)

研究生级别的自然科学题。我用 Epoch AI 官方 CSV 的 263 模型版本,注意这榜榜首被 Google 抢走了:

#模型得分
1gemini-3.7-flash high94.8%
2gpt-5.4-pro-2026-03-05 xhigh94.6%
3gemini-3.1-pro-preview high94.4%
4gemini-3.6-flash high94.1%
5grok-4.6 high94%
5gpt-5.5-pre-release xhigh94%
7gpt-5.5-pro-pre-release xhigh93.9%
7claude-opus-5 max93.9%
9gpt-5.6-sol max93.5%
10grok-4.5 high93.4%

94.8%,这榜头部已经接近刷满。

SWE-bench Verified:真实仓库修 bug(节选 Top 20)

让 AI 在真实 GitHub 仓库里修 bug,500 道验证过的题:

#模型解决率
1Claude Opus 596%
2Claude Mythos 595.5%
3Claude Fable 595%
4Claude Opus 4.888.6%
5Claude Opus 4.7 (Adaptive)87.6%
6Claude Sonnet 585.2%
7GPT-5.3 Codex85%
8Ornith-1.0-397B82.4%
9Claude Opus 4.580.9%
10Claude Opus 4.680.8%
11DeepSeek V4 Pro 081380.6%
12MiniMax M380.5%
13Qwen3.7 Max80.4%
14Kimi K2.680.2%
14Inkling-Small80.2%
16GPT-5.280%
17Claude Sonnet 4.679.6%
18DeepSeek V4 Pro (High)79.4%
19DeepSeek V4 Flash 073179%
20Qwen3.6 Plus78.8%

榜首 96%,前几名挤在 95-96%,基本到顶了。

几句总结

  • Anthropic 五榜第一,只剩 GPQA 被 Google 抢走;OpenAI 的 GPT-5.6 系列紧咬第二梯队
  • 开源模型头部(Qwen3.8、DeepSeek V4、GLM-5.3、Kimi K3)在各榜都能进前 30,性价比依然能打
  • HLE 中位数极低,说明"考试式基准"对难推理题的区分度一直很强,短期还是别指望模型通考
  • 榜单数值随发布随时变,看的时候注意快照日期,别拿旧数开杠

我没法预测未来涨跌,也不替任何一家站台。数据都来自各家官方或独立复评机构的公开榜单。想深挖的可以自己点开官网表格看全量。

数据来源:[1] LMArena 官方 leaderboard(arena.ai),2026-08-12 快照;[2] Artificial Analysis Intelligent Index v4.1.1,2026-08-06;[3] HLE 为 artificialanalysis.ai 实时复评,2026-08-19 抓取;[4] LiveBench 官方 2026-06-25 版;[5] GPQA 用 Epoch AI 官方 CSV(epoch.ai/data)2026-08 快照;[6] SWE-bench Verified 为 BenchLM.ai 实测榜 2026-08-18。

A roundup of six major LLM leaderboards with each source's latest head-of-pack data. Snapshots fall in mid-to-late August 2026: LMArena 08-12 / AA Index 08-06 / HLE 08-19 / LiveBench 06-25 / GPQA 08-14 / SWE 08-18.

Podium summary

LeaderboardLeaderVendorScore
LMArenaClaude Fable 5AnthropicElo 1507
AA Intelligence IndexClaude Opus 5Anthropic63
HLEClaude Fable 5Anthropic55.5%
LiveBenchClaude Fable 5 MaxAnthropic83.0
GPQA DiamondGemini 3.7 FlashGoogle94.8%
SWE-bench VerifiedClaude Opus 5Anthropic96%

Anthropic takes five of six; Google steals GPQA.

Highlights per board

  • LMArena (blind human votes, Elo normalized 0–100): claude-fable-5 at 100, opus variants 97–99.7, then muse-spark-1.2, qwen3.8-max (97.1) and kimi-k3-max (96.8) leading the open models. Gaps in the top tier are single-digit.
  • AA Intelligence Index (265 models): Claude Opus 5 (max) is alone above 60 at 63; GPT-5.6 Sol, Grok 4.6, Kimi K3 and GLM-5.3 fill the 60–61 tier; Qwen3.8 Max leads open at 58.
  • HLE ("Humanity's Last Exam", 2,500 expert questions): median is brutally low; best is 55.5% (Claude Fable 5). The cold-water benchmark.
  • LiveBench (44 models, objective answers): Claude Fable 5 Max 83.0, GPT-5.6 Sol 81, Kimi K3 79.2, Qwen 3.8 Max 78.5.
  • GPQA Diamond (Epoch AI CSV): Google's Gemini 3.7 Flash takes #1 at 94.8% — the top of this board is nearly saturated.
  • SWE-bench Verified (real-repo bug fixing): Claude Opus 5 at 96%, top five within 95–96%; DeepSeek V4 Pro at 80.6% and Qwen3.7 Max at 80.4% lead the open pack.

Takeaways

  • Anthropic's five first places define this round; OpenAI's GPT-5.6 family sits close behind
  • Open-weight heads (Qwen3.8, DeepSeek V4, GLM-5.3, Kimi K3) crack the top 30 on every board
  • HLE's low median shows exam-style benchmarks still discriminate hard at the frontier
  • Scores change with every release — always check the snapshot date before arguing
Sources: LMArena official leaderboard (2026-08-12); Artificial Analysis Intelligent Index v4.1.1 (2026-08-06); HLE via artificialanalysis.ai (2026-08-19); LiveBench official (2026-06-25); GPQA via Epoch AI CSV (2026-08); SWE-bench Verified via BenchLM.ai (2026-08-18).
返回文章列表