← All issues

The AI Enablement Brief · Jul 16, 2026

The 100x Spread

Token costs are climbing faster than productivity. The fix isn't spending less on AI. It's knowing which model actually deserves your money.

I hear a lot of chatter around token costs increasing while productivity increases at a much slower pace. It’s leaving most companies wondering if investing in AI tools and solutions is even worth it.

I understand the doubt. The invoice is real, and it arrives every month with a number on it. The productivity gain is real too, but it shows up slower, spread across a dozen teams, and it never fits neatly on a line item.

But I think most companies are asking the wrong question. The question isn’t whether AI is worth the spend. It’s whether you’re paying premium prices for work that was never premium to begin with.

Most of the time, you are.

The Spread Nobody Looks At

I put together a data hub on my site that tracks every model available through OpenRouter: what they cost, where they come from, and how they score on intelligence and capability. It’s sitting at 347 models as of last week.

Here’s what jumped out the first time I sorted it by price.

Gemini 2.5 Flash Lite runs about $0.14 per million input tokens. Claude Fable 5 runs about $14.17. That’s a hundred times the price for the same million tokens going in. On the way out, the gap widens: roughly $0.57 against $70.85. A hundred and twenty-five times.

Same task. Same prompt. Two invoices that don’t belong in the same conversation.

Now here’s the part that matters. The intelligence scores don’t spread anything like that. Fable 5 sits around 98. Sonnet 5 sits around 92, at roughly a fifth of the price. DeepSeek V4 Pro lands near 88 for about $0.62 per million input tokens.

Cost scales like a rocket. Capability scales like a staircase.

That gap between the two curves is where your token budget is quietly going.

Right Model, Right Task

I’ve said before that model choice barely matters, and I still think that’s true for most teams picking a tool to standardize on. The difference between the frontier models is small enough that integration and context matter more than the badge on the box.

But that was about picking one model to live in. This is different. When you’re running agents, at volume, on repeatable tasks, you’re not making one choice. You’re making that choice ten thousand times a month, and the invoice compounds every time.

At that scale, the reality is that models vary widely in cost and intelligence, and choosing the right model for the right task is becoming a real skill worth spending time learning.

Sonnet 5 is cost effective and scores high on intelligence. That’s fine for most day-to-day work: drafting, summarizing, classification, the reporting pass, the QA sweep. The stuff that fills most of the day.

Fable 5 should be saved for long-running, complex tasks that require extensive thinking. Genuine reasoning over a large surface. Work where being wrong is expensive.

Running Fable 5 on a task Sonnet 5 handles at 92 isn’t ambition. It’s a rounding error you pay for ten thousand times.

The Four Zones

When I built the hub, the models sorted themselves into four zones, and I ended up mapping them that way.

Value frontier is high capability, low cost. This is where the interesting work is happening right now, and it’s fuller than it was six months ago.

Premium is high on both. Real, and worth it, for a narrow band of work.

Bulk floor is low on both. Cheap, capable enough for volume, and honestly fine for a lot of what gets automated.

Avoid is the one that should make you look twice: low capability, high cost. These models exist for reasons that have nothing to do with your workload.

Most teams I talk to have never mapped their spend against those zones. They picked a model in 2025, wired it into everything, and never revisited it while the frontier moved underneath them.

The Local Turn

My prediction is that we’ll start seeing closed, local models gain popularity as the large LLMs keep releasing expensive models that aren’t going to be sustainable for most brands to maintain.

The tracker backs this up more than I expected. Of the 347 models, 47% are open-weights. That’s 163 models you can run yourself, on your own hardware, with no meter running.

That’s not a fringe corner anymore. That’s half the field.

And the geography is worth noticing too: 44% of the models come from the US, 26% from China. The pressure on price isn’t coming from one place, and it isn’t letting up.

What To Do With This

Four concrete starting points, in the order I’d do them.

Find out where the tokens actually go. Most teams can’t answer this. Pull your usage by task, not by month. You’ll almost always find one or two jobs eating the majority of the budget, and they’re rarely the ones that need the intelligence.

Route by task, not by default. Pick your top three workloads and assign a model to each deliberately. The reporting pass doesn’t need what the strategy work needs.

Test the drop before you argue about it. Take one workload, run it on a cheaper model for a week, and compare the output side by side. This is a twenty-minute test that people spend months theorizing about instead.

Watch the open-weights column. You don’t have to move your agents local this quarter. But if you haven’t looked at what’s runnable in-house lately, the answer changed while you weren’t looking.

Where This Lands

The companies wondering whether AI is worth the investment are usually the ones who never made a choice. They defaulted to the biggest model on the menu, wired it into every task, and read the invoice as a verdict on the technology.

It isn’t a verdict on the technology. It’s a bill for not choosing.

The skill isn’t picking the best model. It’s picking the right one.

The full tracker is live at davidzagury.ca/data/ai-models, updated with live pricing.

Enjoyed this?
Get every issue in your inbox — weekly, free.
David Zagury
David's Digital Twin
Online
David Zagury
Hi — I'm David's AI twin. I've read all his writing and know his professional background well. Ask me anything about his work in media or AI.
Powered by Claude · AI can make mistakes