top of page
Search

The Week the AI Race Moved Down the Stack

  • mahdinaser
  • Jun 28
  • 3 min read

Google lost its stars, OpenAI taped out its first chip, and Mistral quietly turned document reading into an enterprise weapon — the week the AI race moved down the stack.

Some weeks bring a shinier model that nudges a benchmark a few points higher. This past week brought something more structural. The action moved away from chatbots and one layer down the stack — to the people who build frontier models, the silicon that serves them, and the unglamorous tooling that feeds them data. None of it makes for a viral demo, and all of it matters more. Three stories from the last seven days tell that better than any leaderboard.

Google’s stars walk out the door

The loudest story was a talent drain at Google DeepMind. Noam Shazeer — a co-author of the 2017 paper “Attention Is All You Need” that introduced the transformer, and a co-lead of Gemini — announced on June 18 that he was leaving for OpenAI. Two days later, John Jumper, who shared the 2024 Nobel Prize in Chemistry for AlphaFold, said he was joining Anthropic. By June 24, two more Gemini contributors, Jonas Adler and Alexander Pritzel, had followed Jumper to Anthropic.

Markets noticed. Alphabet shares fell as much as 7% on June 22 — their worst day in months — erasing an estimated $250 billion in value, according to CNBC and Bloomberg. The reaction looks outsized for four hires until you remember how concentrated this expertise is. Money buys compute, but it doesn’t easily buy back the handful of people who know how to build a frontier model. With both OpenAI and Anthropic eyeing public offerings, the equity math now tilts toward the challengers, and Google is suddenly defending its bench rather than raiding everyone else’s.

OpenAI starts making its own silicon

On June 24, OpenAI and Broadcom unveiled Jalapeño, OpenAI’s first custom AI chip. Unlike Nvidia’s general-purpose GPUs, Jalapeño is an inference-focused ASIC — purpose-built to run finished models cheaply and efficiently rather than to train them. The companies say it went from first design to tape-out in roughly nine months, with initial deployments targeted for the end of 2026.

The strategic message is louder than the chip. For years OpenAI rented essentially all of its compute. Owning the inference silicon — the part of the stack that runs every single user query — is how you control cost and capacity at scale. Paired with Broadcom’s manufacturing, it is the clearest sign yet that OpenAI wants to “build the full stack” behind its products rather than rent it, in the company’s own words.

Mistral makes the boring stuff strategic

While headlines chased chips and talent, Mistral shipped something deceptively mundane on June 23: OCR 4, a document-understanding model. It reads 170 languages, returns not just text but paragraph-level bounding boxes and per-word confidence scores, and — crucially — runs self-hosted in a single container, so regulated businesses never have to send sensitive documents to someone else’s cloud. Pricing starts around $4 per 1,000 pages.

That last detail is the whole point. The hardest part of enterprise AI usually isn’t reasoning — it’s getting clean, trustworthy data out of the messy PDFs, scans, and forms that real businesses run on. Bad inputs quietly poison every fancy model downstream. By making high-quality OCR cheap, structured, and private, Mistral is repositioning itself less as a chatbot maker and more as enterprise plumbing — and plumbing, unglamorous as it is, is where a lot of the durable money lives.

The throughline

Notice what none of these stories were about: a flashy new chatbot. The week’s real contest played out one layer down — over the researchers who build the models, the chips that serve them, and the pipelines that feed them data. The frontier labs increasingly look less like model shops and more like vertically integrated platforms competing on talent, silicon, and data all at once. That is a more expensive game, and a much harder one to fake with a single slick demo.

The takeaway

The model is no longer the whole story. This week the AI race turned on people, chips, and plumbing — the parts of the stack that never make a viral screenshot but quietly decide who can ship intelligence at scale. Watch the stack, not just the leaderboard.
 
 
 

Comments


bottom of page