top of page
Search

The Week the Frontier Got Cheaper

  • mahdinaser
  • Jul 26
  • 3 min read

Anthropic, Google, and a Chinese open-weight lab all made frontier-grade AI cheaper in the same seven days — while the money behind it got a lot bigger.

If you blinked this week, you probably missed a model launch. The back half of July has become the most crowded stretch the AI industry has ever seen, with new systems arriving every few days. But this week had a theme worth pausing on, and it wasn't raw capability. It was price. Three of the most powerful models on the planet just got dramatically cheaper to run, and the capital funding them got dramatically larger. Here is what actually happened, and why it matters.

Anthropic's Opus 5 chases the frontier at half the price

On July 24, Anthropic released Claude Opus 5, and the pitch was as much about economics as intelligence. Anthropic says Opus 5 comes close to the frontier intelligence of its top-end Fable 5 model “at half the price,” landing at $5 per million input tokens and $25 per million output tokens. On the company's own Frontier-Bench v0.1 evaluation, it claims Opus 5 sets a new state of the art and more than doubles the score of Opus 4.8, the previous workhorse, at a lower cost per task.

The interesting part isn't the benchmark; it's the strategy. Opus 5 shipped simultaneously on Anthropic's own platform and on AWS, Google Cloud, and Microsoft Foundry. Frontier-tier coding and reasoning are quietly becoming something you rent by the token, and the labs are now competing on the meter as hard as they compete on the model.

Google floods the cheap tier — and stalls at the top

Google DeepMind reinforced the same idea from the opposite end. On July 21 it shipped three new models — Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber — all built for fast, cheap, high-volume work. Gemini 3.6 Flash reportedly trims token usage by up to 17% versus its predecessor, which is its own kind of price cut.

What Google didn't ship is more telling. There was still no Gemini 3.5 Pro. The flagship has slipped past several targets after, by multiple accounts, failing to clear Google's own internal quality bar. Product lead Logan Kilpatrick said the team is testing 3.5 Pro with partners and hopes to “land soon,” while also revealing that it has begun its most ambitious pre-training run yet for Gemini 4. Winning the cheap-and-fast tier while stalling at the top is a strange, telling place to be.

The largest open-weight model yet comes from Beijing

The most striking release came from China. Moonshot AI's Kimi K3 is a 2.8-trillion-parameter open-weight model — the largest openly available model yet — built as a sparse mixture-of-experts that activates only about 16 of its 896 experts on any given token, keeping inference costs down despite its enormous size. It pairs that with a one-million-token context window.

The benchmarks drew attention: Kimi K3 topped the Frontend Code Arena with 1,679 points, edging out Claude Fable 5 and OpenAI's GPT-5.6 Sol on that particular test. Take any single leaderboard with a grain of salt, but the direction is unmistakable. A frontier-class model you can download and run yourself, for free, arriving the same week two U.S. labs cut their hosted prices, changes the math for everyone. Moonshot says the full weights drop on July 27.

The money is moving faster than the models

None of this is cheap to build, which is why the funding news matters as much as the model news. Fireworks AI, the Nvidia-backed inference startup, raised a $1.5 billion Series D at a $17.5 billion valuation — on the back of crossing $1 billion in annualized revenue and serving more than 40 trillion tokens a day.

The bet is specific: as frontier intelligence commoditizes, the value shifts to whoever can serve customized models fastest and cheapest. That's the through-line of the whole week. And it all landed against an already frantic backdrop — OpenAI's GPT-5.6 family and its ChatGPT Work agent shipped on July 9, and xAI's Grok 4.5 the day before. The cadence is relentless, and it is not slowing down.

The takeaway

This wasn't a week about one breakthrough model. It was a week about deflation. When Anthropic, Google, and a Chinese open-weight lab all make frontier-grade AI cheaper within days of each other — and investors pour billions into serving it — the edge stops being “who has the smartest model” and becomes “who can deliver it for less.” That shift is quieter than a benchmark record, but it will matter far more.

Comments


bottom of page