Here's how fast this actually moves. In the six weeks before I finished writing this: Anthropic shipped Mythos 5 and Fable 5 on June 9. OpenAI answered on July 9 with the whole GPT‑5.6 family, Luna, Terra, and Sol. Google quietly shelved Gemini 3.5 Pro, shipped three other Gemini models instead, and started teasing Gemini 4. Anthropic came back again on July 24 with Opus 5. LMArena's own leaderboard had to go down and rebuild itself in the middle of all of it[1]. Four labs, six weeks, one leaderboard that couldn't keep up with its own subject. By the time you read this sentence, at least one of those names is already old news. That's not a disclaimer I'm making about this post. That's the actual subject of it. So I went and asked the people building these systems what they think the race is even for. Their answers contradicted each other on almost everything. Except for one thing, and it's the part that never makes it into the keynote.
Everybody admits it in the interview
Ilya Sutskever spent a decade making the case that scale wins. So it means something when, in November 2025, he says this about the systems his own life's work helped build:
"These models somehow just generalize dramatically worse than people. It's super obvious. That seems like a very fundamental thing."Ilya Sutskever, Dwarkesh Podcast, Nov 25 2025[2]
Two months later, Demis Hassabis said basically the same thing, in different words, while running the lab that ships Gemini. His worry: the runaway commercial success of chatbots might be pulling the field's attention away from the real problem, not toward it.
"Text alone would not get you to the endgame faster."Demis Hassabis, via Semafor, Jan 2026[3]
Richard Sutton wrote "The Bitter Lesson," the essay every scale-is-all-you-need argument quotes like scripture. In the same breath as everyone above, he said the part scaling maximalists don't love to hear:
"Large language models are about mimicking people, doing what people say you should do. They're not about figuring out what to do."Richard Sutton, Dwarkesh Podcast, Sept 26 2025[4]
And Karpathy said it better than any of them, in seven words:
"Today's frontier LLM research is not about building animals. It is about summoning ghosts."Andrej Karpathy, "Animals vs Ghosts," Oct 1 2025[5]
Four people. Four labs. Four completely different sets of incentives. And the same diagnosis, arrived at independently: predicting the next token well is not the same thing as reliably doing something in the world. That gap, between sounding right and being right enough to bet on, is the whole story of this post.
Even the person with the least reason to admit it, admits it
Here's the data point that actually got to me. Not a quote, a drift. Sam Altman runs the best-funded, most model-obsessed lab on the planet, the one with the least reason on Earth to say the model itself isn't the whole answer. Watch what happened to his own explanation of what's missing, over eighteen months[6]:
I want to be a little cynical here, because it's earned. Everyone's stated bottleneck lines up suspiciously well with what their own lab already happens to be good at. Amodei runs a safety-first lab that can't outspend the frontier on compute, and he names interpretability as the crisis. Altman is sitting on more compute and distribution than anyone alive, and he names infrastructure, his own backyard. Sutskever can't outspend OpenAI, Google, or Anthropic on hardware, and he says ideas beat scale, which happens to be exactly what someone without the hardware needs to be true. None of that makes any of them wrong. It just means you shouldn't read these as neutral technical opinions. They're bets, and everyone above already put their chips down before they told you what they believe.
Patch the current thing, or replace it
Once you separate the diagnosis from the treatment, the field splits into three camps, and none of them are bluffing. Real money, real published research, real companies built around each bet.
This is our read of where each person has put their own research program. Treat the spacing as illustrative, not a measured scale.
The "replace" camp isn't rhetoric. It's funded, and it has proofs.
Yann LeCun quit Meta in November 2025 and raised a billion dollars for AMI Labs on an argument he'd been making in public for three years[7]. It's not just a slogan. There's real math under it. Picture generating text as a walk through a tree, where only some branches lead somewhere correct. If a model has some probability of stepping off that tree at any given token, the odds of the whole answer staying correct over many tokens fall off exponentially:
Every wrong step drags the next one with it. That's his basis for calling autoregressive LLMs "doomed"[8]. Let's actually run the numbers, because the formula is more convincing than the soundbite:
Sutton's bet is a different shape entirely. He's not arguing about architecture, he's arguing about what the model learns from in the first place. His claim: training on human-written text is this decade's version of hand-coded chess heuristics. It works great, right up until the moment something learns from its own experience instead of copying ours. He didn't just write an essay about it. He started a company, Oak Lab, to go build it[10].
The extend camp's own favorite example got taken apart by its own logic
Here's a story I love, because it's the field roasting itself in public. In February 2024, Berkeley's AI research group argued the future belongs to compound systems, lots of models and tools stitched together, not one giant model. Their star exhibit was Medprompt, a clever ensembling pipeline that beat raw GPT‑4 by nine points on medical exams[11]. Nine months later, people from that same research circle published the paper that quietly took their own exhibit apart:
"Even without prompting techniques, o1-preview largely outperforms the GPT‑4 series with Medprompt… [in-context steering] may no longer be an effective approach for reasoning-native models."Nori et al., "From Medprompt to o1," Nov 2024[11]
A model trained with RL on verifiable rewards ate, in under a year, everything a clever hand-built pipeline had to sweat to fake. Watching the bitter lesson happen in real time, inside one model generation, is the best evidence the extend camp has going for it. It doesn't prove extending is enough forever. It proves the floor keeps moving under anyone betting against the base model.
Reliability doesn't arrive as a byproduct
Fine, so everyone agrees the missing piece is reliability, not raw intelligence. Great. How far has that piece actually moved? This is where it stops being comfortable. The 2025 to 2026 literature here is a lot more specific, and a lot more uncomfortable, than "AI agents keep getting better every month" lets on.
Benchmarks overstate what actually ships
METR tracks how long a task an AI agent can handle on its own before it starts falling apart. Credit where it's due: they turned around and fact-checked their own number, which is rarer in this field than it should be. They took Claude Sonnet 4.5's measured "time horizon" (the task length it completes correctly half the time) and regraded it against a harder, more honest bar. Would a real software maintainer actually merge the pull request it produced[12]?
an automated test-passing check
real maintainer merge decisions
Karpathy has a name for why a smarter model doesn't just fix this on its own: "the march of nines." Each extra decimal point of reliability costs roughly a full project's worth of work, same as the last one. It isn't a discount you earn from the model getting smarter. Pair that with what he calls "jagged intelligence": the same model that just cleared a PhD-level problem will turn around and insist 9.11 is bigger than 9.9[5]. I've watched this happen. It's funny until it's in production.
Voting across models doesn't average the errors away. It can't.
There's a comforting story everyone tells themselves: sure, one model might be unreliable, but run three and take a vote and the errors should cancel out. A 2026 audit tested that story directly, across 30 models, and it falls apart in a way you can actually measure[13]:
of 3 models could beat the best one alone
voting actually captured that gain
There is a real fix. It's just not "run three models and see what two of them say." The Weaver method combines 33 weak, individually unreliable verifiers using actual weak-supervision statistics, needing almost no ground truth to calibrate against, and it took one open model from 72.2% to 93.4% on math and reasoning benchmarks. It beat a noticeably stronger single model doing that[15]. It works. It just needs real statistics behind it, not a show of hands.
More agents help, and hurt, depending on the shape of the task
In December 2025, someone finally tested this properly. 180 configurations, four benchmarks, one direct question: does adding more agents to a task actually help[16]? The answer is yes, and also no, depending entirely on whether the task breaks apart cleanly.
decomposable, parallelizable
sequential, constraint-satisfaction
Same lever. Opposite sign, depending which task you pull it on. The same study found what they call a "baseline paradox": once a single agent already clears about 45% accuracy alone, adding more agents actively makes it worse, with error amplification measured as high as 17.2× in the bad regime. Whether coordinating multiple agents helps is now something you can calculate, not something you can assume. That single fact should change how a lot of "multi-agent" products get built, ours included.
Timelines compressed. Certainty didn't.
Researcher surveys have yanked their own predicted arrival dates forward, hard, over the last few years. Worth showing exactly how much, and exactly how shaky the ground under that trend actually is.
AI Impacts survey of published AI researchers
Here's the finding that should worry the field more than any timeline number. The AAAI polled 475 researchers plus a 24-person expert panel on what's actually blocking progress right now. Compute didn't win. Data didn't win.
75% of researchers agreed that a lack of rigor in evaluating AI systems is holding back AI research progress. Only 8% disagreed. Not FLOPs. Not tokens. The field looked at itself and said the bottleneck is evaluation[19].
Does any of this actually show up in the numbers?
A research consensus is one thing. Money actually moving toward it is a different, more honest test. This is the part where the two lined up more than I expected when I started pulling numbers.
Enterprise generative-AI spend more than tripled in a single year. That's a disclosed figure from a 495-company survey, not somebody's extrapolation off a napkin:
3.2× YoY. Menlo Ventures, "State of Generative AI in the Enterprise," Dec 9 2025 (n=495 US enterprise AI decision-makers)[20].
Two analyst firms sized the AI orchestration category completely independently, different methods, different base years, and landed on almost the same growth rate. That agreement matters more than either dollar figure on its own:
Now shrink that down to just the coding-agent slice. Claude and OpenAI together hold 63% of code-generation spend, out of $8.4B in total enterprise LLM API spend[20]. Code-gen itself is worth $1.9B of that pool, which works out to a 23% share. Multiply that share against the orchestration market above and you get:
And here's the part the market already settled on its own, without waiting for anyone's opinion. Routing a prompt to whichever model is cheapest for the job used to be something you'd build. Now it's something you just get. AWS shipped it into Bedrock in April 2025 and claimed a 60% cost saving over running everything through one model[22]. Microsoft had an equivalent in Copilot by November. Even OpenAI shipped GPT‑5 with its own internal router, then walked part of it back in December 2025 after free-tier users complained they'd lost the ability to just pick a model themselves. Sit with that reversal for a second. The complaint wasn't "routing is bad." It was "I didn't ask you to hide it from me." People wanted the choice made in front of them, not behind their back.
The checklist, not the vibe
Line all of this up and you don't get a vibe. You get a checklist you could actually go verify yourself:
The binding problem is reliability and generalization, not raw model IQ, and that diagnosis comes from the people with the least reason to hand it to a critic for free.
Reliability doesn't fall out of a bigger model. Verification, self-correction, and voting all have specific, currently unsolved failure modes, and the correlated-error problem gets worse, not better, as models improve.
Whether coordinating multiple models helps or actively hurts is now a measurable property of the task, not a matter of taste.
Cost-based routing between models is already a commodity feature inside three hyperscalers. Building a company around that specific claim, alone, is picking a fight you can't win on price.
The field's own practitioners rank evaluation rigor, not compute, not data, as the most binding constraint on progress. That's an unglamorous, under-invested answer, and it's also the one thing a governance and confidence layer is actually built to address.
We're building Hydra, a local-first control plane that routes AI coding work across models instead of betting on just one. I didn't write any of the above to justify that. I wrote it blind, on purpose, and only went back to check our own roadmap against it afterward, which is a more uncomfortable exercise than it sounds when you actually do it. The honest result was mixed. Some of it held up better than I expected. Some of it needed fixing on the spot.
Take the voting tax from three sections back: naive majority vote across models captures only 10 to 19% of the improvement that's mathematically sitting on the table, because stronger models make correlated mistakes, not independent ones. That's not a reason to give up on combining models. It's a reason to stop voting and start accumulating evidence properly. Here's the actual math behind Hydra's confidence-routing layer (see First Principles for the full derivation), not a metaphor for it:
That's the part of the roadmap this research made us more confident about, not less. The part it made us less confident about is any race-and-vote swarm mode that doesn't check for that diversity in the first place. Those need to become confidence-weighted, not majority-counted, or they're just an expensive way to make the same mistake three times and call it consensus.
That's the whole point of doing it this way around. If the research had come back saying the real value is in routing by cost, I'd have had to say so, and go rework the plan. It didn't. It said the value is in confidence, verification, and governance, which happens to be where we'd already started walking. Good news checked twice is still good news. It's just better news when you were willing to hear the bad kind too.
What this rests on
- Model releases, Jun to Jul 2026: Anthropic, "Claude Fable 5 and Claude Mythos 5," Jun 9 2026, and "Introducing Claude Opus 5," Jul 24 2026; TechCrunch, "OpenAI launches its new family of models with GPT-5.6," Jul 9 2026; TechCrunch, "Google releases three new Gemini models, but no 3.5 Pro," Jul 21 2026; LMArena's live leaderboard for the July restore and rebaseline (exact cause not independently confirmed here). anthropic.com/news/claude-opus-5
- Sutskever, I.: "We're moving from the age of scaling to the age of research." Dwarkesh Podcast, Nov 25 2025. dwarkesh.com/p/ilya-sutskever-2
- Hassabis, D.: via Semafor, Jan 2026; own essay "A Framework for Frontier AI and the Dawning of a New Age," Jul 14 2026. demishassabis.substack.com
- Sutton, R.: Dwarkesh Podcast, Sept 26 2025; "Era of Experience" (Silver & Sutton, DeepMind, 2025). dwarkesh.com/p/richard-sutton
- Karpathy, A.: "Animals vs Ghosts," Oct 1 2025; YC AI Startup School talk ("the march of nines"), Jun 19 2025. karpathy.bearblog.dev/animals-vs-ghosts
- Altman, S.: Stratechery interviews, Mar 20 2025 & Apr 28 2026; OpenAI Podcast Ep.1, Jun 18 2025. stratechery.com
- AMI Labs: $1.03B seed, Mar 10 2026; LeCun's Meta departure, Nov 11 2025. techcrunch.com/2026/03/09
- LeCun, Y.: "Auto-regressive LLMs are doomed," Paris AI Action Summit, Feb 2025; original (1−e)ⁿ argument, X, Mar 2023. LeJEPA, arXiv:2511.08544
- Miranda, B.: technical review of LeCun's error-compounding argument, Stanford CS, May 26 2026. cs.stanford.edu/people/brando9
- Sutton, R.: OaK architecture keynote, RLC 2025, Aug 30 2025; Oak Lab launch coverage. mlq.ai/news
- Zaharia, Khattab, Chen et al.: "The Shift from Models to Compound AI Systems," BAIR blog, Feb 18 2024; Nori et al., "From Medprompt to o1," arXiv:2411.03590, Nov 6 2024. arxiv.org/abs/2411.03590
- METR: "Measuring AI Ability to Complete Long Software Tasks," arXiv:2503.14499, Mar 2025; "Many SWE-bench-Passing PRs Would Not Be Merged into Main," Mar 10 2026. metr.org/blog
- Kim, S.: "Are Diversity Metrics Measuring Diversity?," arXiv:2607.20768, Jul 2026. arxiv.org/html/2607.20768v1
- Zhou, Xu, Zhou, Singh, Gui, Joty: "Variation in Verification," arXiv:2509.17995, 2025. arxiv.org/pdf/2509.17995
- Saad-Falcon, J. et al.: "Shrinking the Generation-Verification Gap with Weak Verifiers" (Weaver), arXiv:2506.18203, 2025. arxiv.org/html/2506.18203
- Kim, Liu et al. (MIT / Google DeepMind): "Towards a Science of Scaling Agent Systems," arXiv:2512.08296, Dec 9 2025. arxiv.org/abs/2512.08296
- AI Impacts: "Thousands of AI Authors on the Future of AI," arXiv:2401.02843, Jan 2024 (rev. Oct 2025), n=2,778. arxiv.org/abs/2401.02843
- Metaculus: "Date Weakly General AI is Publicly Known" community question, live tracker. metaculus.com/questions/3479
- AAAI: 2025 Presidential Panel on the Future of AI Research, report, Mar 7 2025 (24-member panel + 475-respondent survey). aaai.org (PDF)
- Menlo Ventures: "2025: The State of Generative AI in the Enterprise," Dec 9 2025 (n=495); "2025 Mid-Year LLM Market Update," Jul 31 2025. menlovc.com
- Grand View Research & MarketsandMarkets: AI Orchestration Platform Market reports, 2025. grandviewresearch.com
- AWS: "Amazon Bedrock Intelligent Prompt Routing," general availability announcement, Apr 2025. aws.amazon.com/whats-new
- Wald, A. & Wolfowitz, J.: sequential probability ratio test and its optimality proof, Annals of Mathematical Statistics, 1945 & 1948. Classical statistics, cited here for the log-likelihood accumulation mechanism, not a new result.