AIUpdateWatch Capability Index and official scorecards
A transparent AIUpdateWatch-owned index comes first. Separate official benchmark tabs follow underneath so general reasoning, mathematics, coding, agents, long context and document understanding are never blended without disclosure.
Morning edition: Friday, July 24, 2026 · Data cutoff Jul 23, 2026, 8:30 PM (America/New_York)
AIUpdateWatch-owned benchmark layer
AIUpdateWatch Capability Index v1.0 — open-model evidence edition
Version 1.0
A transparent weighted view built from official benchmark feeds, not a competitor composite. This first edition focuses on open-weight models because the official aggregate currently provides the strongest multi-benchmark coverage for that group. Proprietary models remain visible in the official benchmark tabs below.
AIUpdateWatch Capability Index
Weighted score across covered categories. Minimum publication coverage: 75%.
ChinaInternational
Qwen3.5-397B-A17BChina72.9%
Step-3.5-FlashChina70.7%
Kimi K2.5China69.4%
GLM-5China68.6%
Qwen3.5-27BChina67.9%
Qwen3.5-35B-A3BChina66.8%
DeepSeek-V3.2China61.4%
NVIDIA Nemotron 3 Super 120B-A12BInternational59.7%
Item
Origin
Value
Qwen3.5-397B-A17B
China
72.9%
Step-3.5-Flash
China
70.7%
Kimi K2.5
China
69.4%
GLM-5
China
68.6%
Qwen3.5-27B
China
67.9%
Qwen3.5-35B-A3B
China
66.8%
DeepSeek-V3.2
China
61.4%
NVIDIA Nemotron 3 Super 120B-A12B
International
59.7%
Published methodology
Missing evidence is never silently scored as zero
A model receives an index score only when at least 75% of the category weight is covered. The score is renormalized across covered categories, and the coverage percentage stays visible beside every result.
Reasoning & knowledge
30%
Mathematics
20%
Software engineering
25%
Agents & tools
15%
Document understanding
10%
Methodology and source
Reasoning and knowledge: 30% — average of available GPQA, Humanity’s Last Exam and MMLU-Pro scores.
Mathematics: 20% — average of available AIME 2026 and HMMT 2026 scores.
Software engineering: 25% — average of available SWE-Pro and SWE-bench Verified scores.
Agents and tools: 15% — Terminal-Bench score.
Document understanding: 10% — olmOCR score when available.
A model receives an index score only when at least 75% of the total category weight is covered. Missing categories are not treated as zero. The score is renormalized across covered categories and coverage remains visible beside it.
Every model name opens its official provider page. Component values are the published raw-score averages defined in the methodology. “Not evaluated” means the official feed did not contain a comparable score for that category.
This is an editorial index, not a claim that one number determines the best model for every task. Raw official scorecards remain available immediately below it.
Official benchmark tabs
Inspect each methodology separately
The tabs do not mix incompatible benchmark versions. Each panel keeps its original scale, model naming and source attribution. Model rows stay on AIUpdateWatch; one compact source drawer appears below each scorecard.
General capability
Stanford HELM Capabilities — official snapshot
Same HELM release
Mean score across the current HELM Capabilities scenario set. Higher is better. Scores are shown as percentages for readability.
Rank
Origin
Provider
Model
Mean score
1
International
OpenAI
GPT-5 mini (2025-08-07)
81.9%
2
International
OpenAI
o4-mini (2025-04-16)
81.2%
3
International
OpenAI
o3 (2025-04-16)
81.1%
4
International
OpenAI
GPT-5 (2025-08-07)
80.7%
5
International
Google
Gemini 3 Pro (Preview)
79.9%
6
China
Qwen
Qwen3 235B A22B Instruct 2507 FP8
79.8%
7
International
xAI
Grok 4 (0709)
78.5%
8
International
Anthropic
Claude 4 Opus, extended thinking
78%
9
International
OpenAI
gpt-oss-120b
77%
10
China
Moonshot AI
Kimi K2 Instruct
76.8%
Source and methodology
HELM Capabilities provides prompt-level transparency and reproducible evaluations across a curated general-capability scenario set.
AIME 2026 and HMMT 2026 — selected official-feed rows
Two exams kept visible
Raw percentages from the OpenEvals official benchmark aggregation. The two exams remain separate columns; the displayed mean is only a convenience for this tab.
Rank
Origin
Provider
Model
AIME 2026
HMMT 2026
Mean
1
China
StepFun
Step-3.5-Flash
96.7%
86.4%
91.5%
2
China
Moonshot AI
Kimi K2.5
95.8%
87.1%
91.5%
3
China
Z.ai
GLM-5
95.8%
86.4%
91.1%
4
China
Qwen
Qwen3.5-397B-A17B
93.3%
87.9%
90.6%
5
China
DeepSeek
DeepSeek-V3.2
94.2%
84.1%
89.1%
6
China
Qwen
Qwen3.5-35B-A3B
93.3%
81.8%
87.6%
7
International
NVIDIA
NVIDIA Nemotron 3 Super 120B-A12B
90%
84.8%
87.4%
8
China
Qwen
Qwen3.5-27B
90.8%
81.1%
85.9%
Source and methodology
The two official-feed exam scores remain separate. AIUpdateWatch displays their arithmetic mean only as a reading aid.
Terminal-Bench 2.1 — verified international and Chinese agent–model results
Agent + model system
Terminal-Bench 2.1 evaluates 89 terminal tasks across software engineering, system administration, data processing, model training and security. Scores are not model-only ratings: the agent harness, effort setting, tool behavior and cost are part of each result. This selected table keeps one official entry per model so repeated harness submissions do not dominate the comparison.
Rank
Origin
Provider
Model
Agent
Effort
Accuracy
Cost
1
International
Anthropic
Fable 5
Claude Code Anthropic
xhigh
83.8% ± 1.2%
$552.67
2
International
OpenAI
GPT-5.5
Codex OpenAI
xhigh
83.1% ± 1.1%
$2,059.19
3
International
xAI
Grok 4.5
Cursor CLI Cursor
high
79.3% ± 1.5%
$134.09
4
International
Anthropic
Claude Opus 4.8
Claude Code Anthropic
high
78.9% ± 1.3%
$286.94
5
International
OpenAI
GPT-5.6 Terra
Codex OpenAI
max
78.4% ± 1.3%
$421.15
6
International
Meta
Muse Spark 1.1
mini-SWE-agent Princeton
xhigh
76.2% ± 1.2%
$198.05
7
International
Google
Gemini 3 Pro
Terminus 2 Terminal-Bench
high
73.9% ± 1.3%
$224.44
8
China
Z.ai
GLM-5.1
Claude Code Anthropic
max
58.7% ± 1.2%
$277.14
Source and methodology
At the July 24 cutoff, the official maintained leaderboard included GLM-5.1 as its Chinese-model entry. Kimi K3, GLM-5.2, Qwen3.7 Max, MiniMax-M3 and DeepSeek V4 Pro were not added here because no official Terminal-Bench 2.1 row for those exact models appeared on the maintained leaderboard at verification time.
This tab is document/OCR evaluation, not a universal vision score. General chat models should not be assigned an OCR result unless the official feed contains one.
Rank
Origin
Provider
Model
Score
1
International
Datalab
chandra-ocr-2
85.9%
2
China
RedNote HiLab
dots.mocr
83.9%
3
International
LightOn AI
LightOnOCR-2-1B
83.2%
4
International
Datalab
chandra
83.1%
5
International
Infiny AI
Infinity-Parser-7B
82.5%
6
International
AllenAI
olmOCR-2-7B FP8
82.4%
7
China
PaddlePaddle
PaddleOCR-VL
80%
8
China
Baidu
Qianfan-OCR
79.8%
9
China
DeepSeek
DeepSeek-OCR-2
76.3%
10
China
Z.ai
GLM-OCR
75.2%
Source and methodology
These are document/OCR specialist scores. They are not substituted for general image reasoning or chat quality.