AIUpdateWatch Daily Intelligence

AIUpdateWatch Capability Index and official scorecards

A transparent AIUpdateWatch-owned index comes first. Separate official benchmark tabs follow underneath so general reasoning, mathematics, coding, agents, long context and document understanding are never blended without disclosure.

Morning edition: Friday, July 24, 2026 · Data cutoff Jul 23, 2026, 8:30 PM (America/New_York)

AIUpdateWatch-owned benchmark layer

AIUpdateWatch Capability Index v1.0 — open-model evidence edition

Version 1.0

A transparent weighted view built from official benchmark feeds, not a competitor composite. This first edition focuses on open-weight models because the official aggregate currently provides the strongest multi-benchmark coverage for that group. Proprietary models remain visible in the official benchmark tabs below.

AIUpdateWatch Capability Index

Weighted score across covered categories. Minimum publication coverage: 75%.

ChinaInternational
ItemOriginValue
Qwen3.5-397B-A17BChina72.9%
Step-3.5-FlashChina70.7%
Kimi K2.5China69.4%
GLM-5China68.6%
Qwen3.5-27BChina67.9%
Qwen3.5-35B-A3BChina66.8%
DeepSeek-V3.2China61.4%
NVIDIA Nemotron 3 Super 120B-A12BInternational59.7%
Published methodology

Missing evidence is never silently scored as zero

A model receives an index score only when at least 75% of the category weight is covered. The score is renormalized across covered categories, and the coverage percentage stays visible beside every result.

Reasoning & knowledge
30%
Mathematics
20%
Software engineering
25%
Agents & tools
15%
Document understanding
10%
Methodology and source
  1. Reasoning and knowledge: 30% — average of available GPQA, Humanity’s Last Exam and MMLU-Pro scores.
  2. Mathematics: 20% — average of available AIME 2026 and HMMT 2026 scores.
  3. Software engineering: 25% — average of available SWE-Pro and SWE-bench Verified scores.
  4. Agents and tools: 15% — Terminal-Bench score.
  5. Document understanding: 10% — olmOCR score when available.
  6. A model receives an index score only when at least 75% of the total category weight is covered. Missing categories are not treated as zero. The score is renormalized across covered categories and coverage remains visible beside it.
Open official aggregated benchmark data (opens in a new tab)

Visible components and coverage

AIUpdateWatch Capability Index table

Verified 2026-07-24

Every model name opens its official provider page. Component values are the published raw-score averages defined in the methodology. “Not evaluated” means the official feed did not contain a comparable score for that category.

RankOriginProviderModelIndexCoverageReasoningMathSoftwareAgentsDocument
1ChinaQwenQwen3.5-397B-A17B — open official model page in a new tab72.9%90%68.3%90.6%76.4%52.5%Not evaluated
2ChinaStepFunStep-3.5-Flash — open official model page in a new tab70.7%90%63.7%91.5%74.4%51%Not evaluated
3ChinaMoonshot AIKimi K2.5 — open official model page in a new tab69.4%90%75%91.5%60.8%43.2%Not evaluated
4ChinaZ.aiGLM-5 — open official model page in a new tab68.6%90%58.3%91.1%72.8%52.4%Not evaluated
5ChinaQwenQwen3.5-27B — open official model page in a new tab67.9%90%65.3%85.9%72.4%41.6%Not evaluated
6ChinaQwenQwen3.5-35B-A3B — open official model page in a new tab66.8%90%64%87.6%69.2%40.5%Not evaluated
7ChinaDeepSeekDeepSeek-V3.2 — open official model page in a new tab61.4%90%69.4%89.1%42.8%39.6%Not evaluated
8InternationalNVIDIANVIDIA Nemotron 3 Super 120B-A12B — open official model page in a new tab59.7%90%60.4%87.4%53.7%31%Not evaluated

This is an editorial index, not a claim that one number determines the best model for every task. Raw official scorecards remain available immediately below it.

Official benchmark tabs

Inspect each methodology separately

The tabs do not mix incompatible benchmark versions. Each panel keeps its original scale, model naming and source attribution. Model rows stay on AIUpdateWatch; one compact source drawer appears below each scorecard.

General capability

Stanford HELM Capabilities — official snapshot

Same HELM release

Mean score across the current HELM Capabilities scenario set. Higher is better. Scores are shown as percentages for readability.

RankOriginProviderModelMean score
1InternationalOpenAIGPT-5 mini (2025-08-07)81.9%
2InternationalOpenAIo4-mini (2025-04-16)81.2%
3InternationalOpenAIo3 (2025-04-16)81.1%
4InternationalOpenAIGPT-5 (2025-08-07)80.7%
5InternationalGoogleGemini 3 Pro (Preview)79.9%
6ChinaQwenQwen3 235B A22B Instruct 2507 FP879.8%
7InternationalxAIGrok 4 (0709)78.5%
8InternationalAnthropicClaude 4 Opus, extended thinking78%
9InternationalOpenAIgpt-oss-120b77%
10ChinaMoonshot AIKimi K2 Instruct76.8%
Source and methodology

HELM Capabilities provides prompt-level transparency and reproducible evaluations across a curated general-capability scenario set.

Open Stanford HELM Capabilities (opens in a new tab)

Global model evidence coverage

This matrix shows which capability areas the daily registry monitors. It does not claim that every model has a verified score in every category.

OriginProviderModelGeneralCodingImage / multimodalAgents / toolsLong context
InternationalOpenAIGPT-5.6 SolTrackedTrackedTrackedTrackedNot evaluated
InternationalOpenAIGPT-5.6 TerraTrackedTrackedTrackedTrackedNot evaluated
InternationalOpenAIGPT-5.6 LunaTrackedNot evaluatedNot evaluatedNot evaluatedNot evaluated
InternationalAnthropicClaude Fable 5TrackedTrackedTrackedTrackedNot evaluated
InternationalxAIGrok 4.5TrackedTrackedNot evaluatedTrackedTracked
InternationalGoogleGemini 3.5 FlashTrackedTrackedTrackedTrackedNot evaluated
ChinaDeepSeekV4 FlashTrackedTrackedNot evaluatedNot evaluatedTracked
ChinaAlibaba CloudQwen3.7-MaxTrackedTrackedNot evaluatedTrackedNot evaluated
ChinaMoonshot AIKimi K3TrackedTrackedNot evaluatedTrackedTracked
ChinaZhipu AIGLM-5.2TrackedTrackedNot evaluatedTrackedTracked
ChinaBaiduERNIE 5.0TrackedNot evaluatedTrackedNot evaluatedTracked
ChinaByteDanceDoubao Seed 2.1TrackedTrackedTrackedTrackedNot evaluated
ChinaMiniMaxMiniMax-M3TrackedTrackedTrackedTrackedTracked
ChinaStepFunStep 3.7 FlashTrackedTrackedTrackedTrackedNot evaluated
ChinaTencentHunyuan A13BTrackedTrackedNot evaluatedTrackedTracked

Responsible benchmark rules

  1. 1

    Compare the exact model version, date, reasoning setting and tool access.

  2. 2

    Use coding benchmarks for coding, document benchmarks for OCR and long-context benchmarks for long documents.

  3. 3

    Publish the index weights, component scores, source date and missing-evidence rule.

  4. 4

    Keep raw official scorecards directly underneath every editorial index.

  5. 5

    Measure cost, latency, reliability and task success alongside benchmark scores.