robotmunki.commodel-selector

LLM Model Selector

Top Models — Default View

Top 8 by capability score — apply levers to filter by task, cost, context, or vendor

Aug 2026 data
1

Claude Opus 5

Anthropic
AA 61

New Jul 24 — new AA leader; succeeds Opus 4.8 at the same $5/$25 price; near Fable 5 intelligence at half the cost.

$25/M out1M ctx
AA 61NewAA v4.1 leaderAgents
2

Grok 4.6

xAI
AA 61

New Aug 12 — ties Claude Opus 5 for the top AA score at a quarter of the cost; built for long-running agents.

$6/M out500K ctx
AA 61NewTies leaderAgents
3

Claude Fable 5

Anthropic
AA 60

Mythos-class with safety classifiers, 1 pt below Opus 5; access suspended June 12, staged return; Opus 5 fallback on blocked queries.

$50/M out1M ctx
AA 60Mythos-classLimited access
4

Claude Mythos 5

Anthropic
AA 60

Same weights as Fable 5 without cyber/bio classifiers — Project Glasswing partners only.

$50/M out1M ctx
AA 60GlasswingRestricted
5

Kimi K3

Moonshot AI
AA 60

Jul 16 launch, weights opened Jul 26 — 2.8T params, native multimodal; ties GLM-5.3 for open-weight #1, 1 pt behind Opus 5.

$15/M out1.05M ctxopen
AA 60Open weightTies GLM-5.32.8T
6

GLM-5.3

Z.ai
AA 60

New Aug 14 — ties Kimi K3 for open-weight #1, 1 pt behind Opus 5; same 753B base as GLM-5.2, gains from post-training alone; weights staged ~Aug 28. Flash sibling (AA 57) is the cost play.

$4.4/M out1M ctxopen
AA 60Open weightTies Kimi K3New
7

GPT-5.6 Sol

OpenAI
AA 59

GA since Jul 9 (June 26 preview ended); price cut -20%/-33% on Aug 21; most token-efficient frontier model.

$20/M out1.05M ctx
AA 59GANext-genToken-efficient
8

Qwen3.8 Max

Alibaba
AA 58

New Aug 3 (open weights Aug 12) — 2.4T params (95B active); first Qwen-Max-class model shipped open; leads PaperBench, OSWorld-Verified. Flash-Next (AA 56) and 27B (AA 52) are the efficiency siblings.

N/A1M ctxopen
AA 58NewOpen weight2.4T

Lever Reference

1
Capability

AA composite index (MMLU-Pro, SWE-bench, GPQA, ARC-AGI, AIME, context)

2
Task Fit

Peak AA score does not equal best task fit (e.g. coding)

3
Context Window

Standard ≤262K · Large 200K–1M · Massive 1M–10M

4
Cost Per Token

Budget < $1/M · Mid $1–10/M · Premium $10+/M output

5
Deployment

API-only vs open-weight self-host vs both

6
Vendor Origin

Geographic / residency-style grouping for vendor choice