Rubric v2026-08-14
How we score
Scores have two layers: raw benchmark numbers, and the RouteVault Scale — a published table that maps those numbers to plain-English grades. Every grade is traceable to the table below; nothing is hand-adjusted.
Technical layer
Default value weights (Explorer sliders can change these locally):
- Reasoning 32% — Artificial Analysis intelligence index
- Coding 25% — AA coding index
- Tool use 18% — AA agentic index
- General 15% — same intelligence series, tracked separately for future benches
- Context 7% — derived from advertised context window
- Speed 3% — unused until we store a latency series
Value score = weighted bench average ÷ (input+output $/M) × 10. Models missing a benchmark are scored without it rather than counted as zero.
RouteVault Scale
- Project Planner — agentic ≥ 45 and intelligence ≥ 50
- Researcher — intelligence ≥ 52, or ≥ 45 with ≥ 400k context
- Senior Developer — coding ≥ 65 (or ≥ 50 if not cheap)
- Junior Developer — coding ≥ 35
- Budget Everyday — combined list price under $2.50 / M, or leftover generalists
- Specialist — thin bench coverage or narrow models
What is live vs accumulated
Price history is weekly OpenRouter list prices from the public OpenRouterList ledger, from first appearance up to six months. We only show prices from weeks a model actually existed. Capability indexes are last-known AA snapshots held across that price series.