LLM Landscape 2026: Intelligence Leaderboard and Model Guide
Leaderboard Methodology
Both tables use the AA (Artificial Analysis) Intelligence Index v4.1, which aggregates nine agentic-weighted evaluations — GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR — into a single normalized integer. AA now labels that same nine-eval suite v4.1.1; newly scored models (GLM-5.3-Flash, Qwen3.8-Flash-Next, Qwen3.8-27B) are taken from the current listing. Previously published integers in these tables are left unchanged unless AA re-ran a specific checkpoint — compare within v4.1/v4.1.1, not against pre-v4.1 numbers. Scores shown use each model's highest published effort tier (typically max / xhigh). Context windows use comma-separated token counts; missing public data is "—". Pricing is per million tokens (input / output).
Table 1 — Top 25 by vendor: exactly one representative per company — the provider's newest clearly superior flagship or best overall product — to surface geographic and corporate diversity. Table 2 — Top 25 by model: pure AA Index ranking; multiple entries from Anthropic, OpenAI, Google, and others are expected. Models with limited or suspended API access (Fable 5, Mythos 5) are included with availability notes because they materially affect the competitive picture.
Top 25 LLMs by Vendor — Company Diversity Leaderboard (August 2026)
| Rank | Model | Capability Index (AA Index) | Context Window (tokens) | Input Cost ($/M tokens) | Output Cost ($/M tokens) | Notes |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 Anthropic | 61 | 1,000,000 | $5.00 | $25.00 | New Jul 24 New AA leader; succeeds Opus 4.8 at the same price; comes close to Fable 5's intelligence at half the cost |
| 2 | Grok 4.6 xAI | 61 | 500,000 | $2.00 | $6.00 | New Aug 12 Ties Claude Opus 5; two Grok generations shipped in five weeks (4.5 → 4.6); built for long-running agents |
| 3 | GLM-5.3 Z.ai | 60 | 1,000,000 | $1.40 | $4.40 | New Aug 14 Open Weight Ties Kimi K3 for #1 open-weight, 1 pt behind Opus 5; same 753B base as GLM-5.2, gains from post-training alone; weights staged ~Aug 28. Flash sibling (AA 57, $0.15/$0.50) shipped Aug 26 |
| 4 | Kimi K3 Moonshot AI | 60 | 1,049,000 | $3.00 | $15.00 | Open Weight Jul 16 launch, weights opened Jul 26; 2.8T params, native multimodal — largest open-weight model to date |
| 5 | GPT-5.6 Sol OpenAI | 59 | 1,050,000 | $4.00 | $20.00 | GA since Jul 9 (June 26 preview ended); price cut -20%/-33% on Aug 21; most token-efficient frontier model (~15k tokens/task) |
| 6 | Qwen3.8 Max Alibaba | 58 | 1,000,000 | — | — | New Aug 3 Open Weight 2.4T params (95B active); open weights Aug 12 — first Qwen-Max-class model released open. Efficiency siblings: Flash-Next (AA 56) and dense 27B (AA 52) |
| 7 | Gemini 3.7 Flash Google | 56 | 1,000,000 | $0.75 | $3.75 | New Aug 13 Fastest reasoning model on the market (~340 t/s); Gemini 3.5 Pro still unreleased, 3+ months after its May announcement — Gemini 4 now in pretraining |
| 8 | DeepSeek V4 Pro-0813 DeepSeek AI | 53 | 1,000,000 | $2.19 | $8.76 | Open Weight Dated GA checkpoint; +9 pts vs the original V4 Pro; V4 Flash-0731 and an experimental Flash Vision sibling cover budget/multimodal |
| 9 | MiniMax-M3 MiniMax | 44 | 1,000,000 | — | — | Open Weight Strong SWE-bench Verified (~80.5%) |
| 10 | MiMo-V2.5-Pro Xiaomi | 42 | 1,000,000 | — | — | Successor to MiMo-V2-Pro; pricing not publicly disclosed |
| 11 | NVIDIA Nemotron 3 Super 120B NVIDIA | 25 | 1,000,000 | $0.30 | $0.75 | Open Weight Enterprise self-hosting value tier; Mamba-2/MoE architecture |
| 12 | Mistral Large 3 Mistral | 16 | 256,000 | $0.50 | $1.50 | Open Weight EU-based; Apache 2.0; v4.1 agentic re-weighting lowered composite vs v4.0 |
| 13 | Nova Premier Amazon | 13 | 1,000,000 | $2.50 | $12.50 | Hyperscaler representative; deep AWS integration |
| 14 | Llama 4 Scout Meta | 10 | 10,000,000 | — | — | Open Weight Context-window outlier; 10M tokens for corpus-scale self-hosting |
| 15 | Command A Cohere | 8 | 256,000 | $2.50 | $10.00 | Enterprise RAG and tool-use focus |
| 16 | Solar Pro 2 Upstage | 8 | — | — | — | 31B Korean frontier; strong regional/language performance; v4.1 composite lower than legacy v4.0 ranking |
| 17 | ERNIE 4.5 300B A47B Baidu | — | — | — | — | Best verifiable ERNIE-family public entry; AA v4.1 pending |
| 18 | Granite 4.0 H Small IBM | 5 | — | — | — | Open Weight Enterprise governance and open-deployment focus |
| 19 | Jamba 1.7 Large AI21 | 5 | — | — | — | Hybrid SSM/Transformer architecture for long-input efficiency |
| 20 | Yi-Lightning 01.AI | — | — | — | — | Vendor-diversity slot; public AA v4.1 score not yet published |
| 21 | Sonar Reasoning Pro Perplexity | — | 128,000 | — | — | Search-augmented reasoning API; AA v4.1 pending |
| 22 | Reka Flash 3 Reka | — | — | — | — | Multimodal agentic model; AA v4.1 pending |
| 23 | Hunyuan-A13B-Instruct Tencent | — | — | — | — | Chinese hyperscaler representative; AA v4.1 pending |
| 24 | Stable LM 2 12B Stability AI | — | — | — | — | Open Weight Community/open-deployment slot; AA v4.1 pending |
| 25 | Evo-Ukiyoe Sakana AI | — | — | — | — | Evolutionary-model research lab; specialized rather than general-purpose frontier |
Top 25 Models by AA Index v4.1 — Pure Capability Leaderboard (August 2026)
This table ranks the twenty-five highest-scoring models on the AA Index v4.1 regardless of vendor — expect multiple Anthropic, OpenAI, and Google entries. Effort tiers are max / xhigh unless noted.
| Rank | Model | Capability Index (AA v4.1) | Context Window (tokens) | Input Cost ($/M tokens) | Output Cost ($/M tokens) | Notes |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 Anthropic | 61 | 1,000,000 | $5.00 | $25.00 | New Jul 24 New AA leader; succeeds Opus 4.8 two months after it shipped, at the same price |
| 2 | Grok 4.6 (high) xAI | 61 | 500,000 | $2.00 | $6.00 | New Aug 12 Ties Opus 5; ~4x cheaper output than Opus 5; long-running agent and visual-work focus |
| 3 | Claude Fable 5 Anthropic | 60 | 1,000,000 | $10.00 | $50.00 | Limited access Mythos-class with safety classifiers; suspended June 12, staged return expected; Opus 5 fallback on blocked queries |
| 4 | GLM-5.3 Z.ai | 60 | 1,000,000 | $1.40 | $4.40 | New Aug 14 Open Weight Ties Kimi K3 for open-weight #1; same 753B/40B-active base as GLM-5.2; weights staged ~Aug 28. Flash sibling (AA 57) is the cost play, not a replacement |
| 5 | Kimi K3 Moonshot AI | 60 | 1,049,000 | $3.00 | $15.00 | Open Weight Jul 16 launch, weights opened Jul 26; 2.8T params; native multimodal; largest open-weight model to date |
| 6 | GPT-5.6 Sol (max) OpenAI | 59 | 1,050,000 | $4.00 | $20.00 | GA since Jul 9; most token-efficient frontier model (~15k tokens/task, ~$1.04/task) |
| 7 | Qwen3.8 Max Alibaba | 58 | 1,000,000 | — | — | New Aug 3 Open Weight 2.4T params (95B active); open weights Aug 12; leads on PaperBench and OSWorld-Verified |
| 8 | GLM-5.3-Flash Z.ai | 57 | 1,000,000 | $0.15 | $0.50 | New Aug 26 Open Weight Native multimodal; 320B/18B-active; MIT; stealth-tested as Ox Alpha on OpenRouter; ~10× cheaper than GLM-5.3 |
| 9 | Claude Opus 4.8 Anthropic | 56 | 1,000,000 | $5.00 | $25.00 | Superseded by Opus 5; SWE-bench Pro 69.2%; still pinned in many production stacks |
| 10 | Gemini 3.7 Flash (high) Google | 56 | 1,000,000 | $0.75 | $3.75 | New Aug 13 Fastest reasoning model on the market (~340 t/s); intro pricing |
| 11 | Qwen3.8-Flash-Next Alibaba | 56 | 262,000 | — | — | New Aug 26 Open Weight Qwen4 architecture preview; 125B/6B-active + 51B n-gram embeddings; 256K native (1M via YaRN) |
| 12 | GPT-5.6 Terra (max) OpenAI | 55 | 1,050,000 | — | — | Matches GPT-5.5 at roughly half the cost; GA since Jul 9, price cut Jul 30 |
| 13 | GPT-5.5 OpenAI | 55 | 1,050,000 | $5.00 | $30.00 | Prior OpenAI flagship; superseded by the GPT-5.6 family |
| 14 | Claude Opus 4.7 Anthropic | 54 | 1,000,000 | $5.00 | $25.00 | Two Opus generations back; still strong for pinned production workflows |
| 15 | Grok 4.5 xAI | 54 | 500,000 | $2.00 | $6.00 | Jul 8 release, superseded by Grok 4.6 in just 5 weeks; #1 on agentic tool use (τ³-Banking) at launch |
| 16 | Claude Sonnet 5 Anthropic | 53 | 1,000,000 | $2.00 | $10.00 | Default Free/Pro model since Jul 1; Terminal-Bench 2.1 80.4%; intro pricing through Aug 31 |
| 17 | DeepSeek V4 Pro-0813 DeepSeek AI | 53 | 1,000,000 | $2.19 | $8.76 | Open Weight Dated GA checkpoint, +9 pts vs the original V4 Pro |
| 18 | DeepSeek V4 Flash-0731 DeepSeek AI | 52 | 1,000,000 | — | — | Open Weight MIT-licensed; 284B total / 13B active MoE; re-post-trained checkpoint, same architecture as the April preview |
| 19 | Qwen3.8-27B (xhigh) Alibaba | 52 | 262,000 | $0.43 | $2.55 | New Aug 14 Open Weight Dense 27B; Apache 2.0; native multimodal; laptop-class, ties GPT-5.6 Luna |
| 20 | GLM-5.2 (max) Z.ai | 51 | 1,000,000 | $1.40 | $4.40 | Open Weight Superseded by GLM-5.3; June 13 GA; SWE-bench Pro 62.1%; MIT license |
| 21 | GPT-5.6 Luna (max) OpenAI | 51 | 1,050,000 | — | — | Matches/exceeds Gemini 3.5 Flash and GLM-5.2 at lower cost; access expanded to Free/Go users Aug 6 |
| 22 | Gemini 3.5 Flash (high) Google | 50 | 1,000,000 | $1.50 | $9.00 | Superseded by Gemini 3.7 Flash; still GA and widely deployed |
| 23 | Claude Sonnet 4.6 (max) Anthropic | 47 | 1,000,000 | $3.00 | $15.00 | Superseded by Sonnet 5; still pinned in production agent stacks |
| 24 | Gemini 3.1 Pro Preview Google | 46 | 1,000,000 | $2.00 | $12.00 | Remains Google's Pro-tier flagship while Gemini 3.5 Pro stays unreleased; multimodal strength |
| 25 | MiniMax-M3 MiniMax | 44 | 1,000,000 | — | — | Open Weight Strong multimodal and agentic scores; SWE-bench Verified ~80.5% |
Ox Alpha (also styled 0x Alpha) is not a separate model — Z.ai stealth-tested GLM-5.3-Flash anonymously on OpenRouter and OpenCode from ~Aug 20 before the Aug 26 unmasking. Claude Mythos 5 shares Fable 5's underlying weights and AA score (~60) but is restricted to Project Glasswing partners (cyber/biology safeguards lifted). DeepSeek V4 Flash Vision Exp (Aug 21) is an experimental, API-only vision sibling built on the V4 Flash-0731 backbone — no AA composite score yet, so it stays off the ranked tables; it splits roughly evenly against Claude Opus 4.8 on dedicated visual benchmarks (ALE, ZeroBench) and has no public weights. MiMo-V2.5-Pro, Qwen3.5 397B, NVIDIA Nemotron 3 Super, Grok 4.3/4.20, GPT-5.4, o3/o4-mini, gpt-oss-120B, Kimi K2.6/K2.7-Code, and the Gemma 4 family fell out of this cycle's top 25 amid the new arrivals but remain relevant lower-cost options — see the model selector for full listings.
Key Takeaways
Key Performance Metrics
Task-Specific Leaders
| Model | Benchmark Leadership |
|---|---|
| Claude Opus 5 | New AA leader (61) · available today · same price as Opus 4.8 |
| Grok 4.6 | Ties Opus 5 (61) · $2/$6 · long-running agents |
| Claude Fable 5 | SWE-bench Pro 80.3% · limited access · Mythos-class |
| GLM-5.3 / Kimi K3 | Tied open-weight leaders (60) · 1 pt behind Opus 5 |
| GPT-5.6 Sol | GA Jul 9 · most token-efficient frontier model · AA 59 |
| Qwen3.8 Max | Open weights Aug 12 · leads PaperBench, OSWorld-Verified |
| GLM-5.3-Flash | AA 57 · $0.15/$0.50 · native multimodal · former Ox Alpha |
Context Window Champions
| Model | Tokens |
|---|---|
| Llama 4 Scout | 10,000,000 |
| Opus 5 · GLM-5.3 · GLM-5.3-Flash · Kimi K3 · Qwen3.8 Max | 1,000,000–1,049,000 |
| GPT-5.6 Sol · DeepSeek V4 Pro-0813 · Nemotron | 1,000,000–1,050,000 |
| Qwen3.8-Flash-Next · Qwen3.8-27B | 262,000 (1M YaRN) |
| Grok 4.6 | 500,000 |
Cost Efficiency
| Tier | Models | Output $/M |
|---|---|---|
| Best Value | GLM-5.3-Flash · Nemotron · Gemini 3.7 Flash | ~$0.50–$3.75 |
| Mid-Range | Grok 4.6 · Sonnet 5 · Qwen3.8-27B | $2.55–$10.00 |
| Flagship | GPT-5.6 Sol · Gemini 3.1 Pro | $12.00–$20.00 |
| Premium | Claude Opus 5 · GPT-5.5 · Fable 5 | $25.00–$50.00 |
Specialized Performance Highlights
Speed & Latency
Open-Weight Excellence
| Model | Key Strength |
|---|---|
| GLM-5.3 / Kimi K3 | AA 60 tied #1 open-weight · 1 pt behind Opus 5 |
| GLM-5.3-Flash | AA 57 · $0.15/$0.50 · 320B/18B · MIT · native multimodal |
| Qwen3.8 Max | AA 58 · 2.4T params · first open Qwen-Max model |
| Qwen3.8-Flash-Next | AA 56 · Qwen4 preview · 125B/6B active |
| Qwen3.8-27B | AA 52 · dense 27B · Apache 2.0 · laptop-class |
| DeepSeek V4 Pro-0813 | AA 53 · dated GA checkpoint · best $/task open weight |
| Llama 4 Scout | 10M-token context · corpus-scale tasks |
Model Selection Guide
Industry Impact & Future Trends (2026)
The 2026 LLM landscape is defined by the fastest release cadence yet across both closed and open models, and a closing open/closed intelligence gap:
Coding & Agents
Regulatory & Access
Open-Weight & Local AI
Conclusion
The August 2026 update covers the busiest stretch this landscape has tracked. Anthropic shipped Claude Opus 5 (AA 61, Jul 24) as the new leader among generally-available models — just two months after Opus 4.8. OpenAI took GPT-5.6 Sol to general availability (Jul 9) and cut its price twice since. xAI shipped and superseded an entire model generation (Grok 4.5 → Grok 4.6) in five weeks, with 4.6 tying Opus 5's AA score outright. Google shipped Gemini 3.7 Flash while Gemini 3.5 Pro remains unreleased three months after its announcement. Late August then unmasked the OpenRouter stealth model Ox Alpha as GLM-5.3-Flash (AA 57, $0.15/$0.50) and added Alibaba's Qwen3.8-Flash-Next (AA 56) and dense Qwen3.8-27B (AA 52). Open weights had their best stretch yet: Kimi K3 and GLM-5.3 both reached AA 60, a single point off the new closed-model leader, while Qwen3.8 Max became Alibaba's first open Max-class model and DeepSeek shipped two more dated V4 checkpoints plus an experimental vision variant.
Strategic Takeaway (2026)
Looking ahead: Gemini 3.5 Pro (or a Gemini 4/3.7 Pro successor) and GLM-5.3's pending open-weight release (~Aug 28) are still the two most consequential unresolved threads — GLM-5.3-Flash weights are already public. Watch for whether Kimi K3 or GLM-5.3 pulls ahead once both settle, whether Qwen3.8-Flash-Next's Qwen4 architecture graduates to a production Max-class successor, and whether DeepSeek's Flash Vision Exp graduates from experimental to a full open-weight release. Expect another refresh within a few weeks given this cycle's pace.