Rubric v2026-08-14

How we score

Two layers. Technical numbers stay numbers. The RouteVault Scale is a published mapping, not a vibe.

Technical layer

Default value weights (Explorer sliders can change these locally):

  • Reasoning 32% — Artificial Analysis intelligence index
  • Coding 25% — AA coding index
  • Tool use 18% — AA agentic index
  • General 15% — same intelligence series, tracked separately for future benches
  • Context 7% — derived from advertised context window
  • Speed 3% — unused until we store a latency series

Value score = weighted bench average ÷ (input+output $/M) × 10. Missing benches are omitted, not filled with zeros pretending to be scores.

RouteVault Scale

  • Project Planner — agentic ≥ 45 and intelligence ≥ 50
  • Researcher — intelligence ≥ 52, or ≥ 45 with ≥ 400k context
  • Senior Developer — coding ≥ 65 (or ≥ 50 if not cheap)
  • Junior Developer — coding ≥ 35
  • Budget Everyday — combined list price under $2.50 / M, or leftover generalists
  • Specialist — thin bench coverage or narrow models

What is live vs accumulated

Price history is weekly OpenRouter list prices from the public OpenRouterList ledger, from first appearance up to six months. We do not invent prices from before a model existed, and we do not reconstruct a random walk. Capability indexes are last-known AA snapshots held across that price series.

Current labels