The current state of AI models.
A live catalog of 347 models from OpenRouter, release timeline from Epoch AI, and geography of origin. Pricing shown in your currency of choice — real-time USD from OpenRouter, CAD converted at the Bank of Canada noon rate.
Release cadence, month by month
166 notable language models tracked by Epoch AI over the last 24 months. Chart = total per month; notable releases listed below with dates.
The line is fairly flat — 6-10 notable language models per month over the last two years, with occasional spikes when a lab clusters releases (Anthropic and DeepSeek both did this in mid-2026). The steady pace matters more than the peaks: the frontier isn't sprinting, it's compounding.
Geography of the AI model landscape
Country of origin for the 347 models in the OpenRouter catalog. Editorial signal, not a scientific measure — some labs release many variants of one architecture.
The US still supplies the plurality of models, but China is now solidly #2 at over a quarter of the catalog — three years ago that was under 5%. The share is driven by open-weights releases from DeepSeek, Alibaba (Qwen), Tencent, and MiniMax. That's the story worth tracking: not who leads on any given benchmark, but who's shipping models people can actually deploy.
347 of 347 models
Filter to what you're trying to compare. CAD pricing per million tokens. Sorted by input price, cheapest first.
Popular models: monthly cost side-by-side
Preset workload: 1M input + 500K output tokens per month. Values compute as (input price × 1M) + (output price × 0.5M), then converted to CAD at 1.4169. Popular list is manually curated; every price is real-time from OpenRouter.
Same workload, huge spread — Fable 5 is roughly 115× the cost of Gemini Flash Lite for the exact same tokens. For most bulk workloads — content generation, summarization, categorization — the budget tier does the job, and reserving frontier models like Opus and Fable for the 5% of work that needs it saves 50-100× on run cost. Volume matters more than model choice for the routine layer.
How they compare on capability
Artificial Analysis composite indices (0-100), pulled live from each model's OpenRouter listing. Intelligence = general reasoning & knowledge · Coding = programming benchmarks · Agentic = multi-step tool use. Some models don't publish scores.
Fable 5 leads all three benchmarks — but so does its price tag ($49.59for our preset workload). The interesting cluster is the second tier: Opus, Grok, Sonnet all land within 3 points of each other on intelligence, but Sonnet costs a third of Opus and less than half of Grok. That's the tradeoff to actually reason about — a rounding-error capability gap for meaningful cost savings.
Capability vs. cost
Every popular model plotted by Artificial Analysis intelligence score (right = smarter) against monthly cost for the preset workload (bottom = cheaper, log scale). The value frontier lives in the bottom-right corner.
Look at the bottom-right of the chart: DeepSeek V4 Pro and Sonnet 5anchor the value frontier — high intelligence scores paired with costs a fraction of the frontier tier. Fable and Opus live in the top-right premium quadrant where you're paying a real premium for the last few points of capability. GPT-5 Mini and Gemini 2.5 Pro cluster at similar cost with lower intelligence — the middle of the chart is crowded with reasonable-but-not-differentiated options. The story: for volume work, pick something on the value frontier; reserve the premium tier for the specific tasks that actually need frontier reasoning.
David’s picks
Curated as of July 2026. What I’d actually reach for, and why.
My default for the most complex work. When a problem has real depth or needs sustained, careful reasoning, this is what I reach for.
Where I start. I lean on it for early planning and the first passes of development, before the shape of the work is settled.
The workhorse for execution and everyday tasks. When something doesn't need a ton of context or extensive thinking, Sonnet handles it fast.
Not a work tool. Claude does the real work; for my non-work, day-to-day questions I mostly just talk to Grok's voice AI.