HYDRA blog All posts
field notes · frontier AI research

The diagnosis is unanimous.
The cure isn't.

I spent the better part of three weeks doing something dumber than it sounds: reading the actual papers instead of the takes about the papers. What came out the other side changed my mind about a few things I was pretty confident of. Every number below is sourced. The ones we ran the arithmetic on ourselves, we say so.

17 min, no filler· 23 sources, all linked· Aug 9, 2026
⊢ primary source ▣ named survey/report ◈ our own arithmetic ○ reported, unverified

Here's how fast this actually moves. In the six weeks before I finished writing this: Anthropic shipped Mythos 5 and Fable 5 on June 9. OpenAI answered on July 9 with the whole GPT‑5.6 family, Luna, Terra, and Sol. Google quietly shelved Gemini 3.5 Pro, shipped three other Gemini models instead, and started teasing Gemini 4. Anthropic came back again on July 24 with Opus 5. LMArena's own leaderboard had to go down and rebuild itself in the middle of all of it[1]. Four labs, six weeks, one leaderboard that couldn't keep up with its own subject. By the time you read this sentence, at least one of those names is already old news. That's not a disclaimer I'm making about this post. That's the actual subject of it. So I went and asked the people building these systems what they think the race is even for. Their answers contradicted each other on almost everything. Except for one thing, and it's the part that never makes it into the keynote.

01 · the thing nobody puts in the keynote

Everybody admits it in the interview

Ilya Sutskever spent a decade making the case that scale wins. So it means something when, in November 2025, he says this about the systems his own life's work helped build:

"These models somehow just generalize dramatically worse than people. It's super obvious. That seems like a very fundamental thing."Ilya Sutskever, Dwarkesh Podcast, Nov 25 2025[2]

Two months later, Demis Hassabis said basically the same thing, in different words, while running the lab that ships Gemini. His worry: the runaway commercial success of chatbots might be pulling the field's attention away from the real problem, not toward it.

"Text alone would not get you to the endgame faster."Demis Hassabis, via Semafor, Jan 2026[3]

Richard Sutton wrote "The Bitter Lesson," the essay every scale-is-all-you-need argument quotes like scripture. In the same breath as everyone above, he said the part scaling maximalists don't love to hear:

"Large language models are about mimicking people, doing what people say you should do. They're not about figuring out what to do."Richard Sutton, Dwarkesh Podcast, Sept 26 2025[4]

And Karpathy said it better than any of them, in seven words:

"Today's frontier LLM research is not about building animals. It is about summoning ghosts."Andrej Karpathy, "Animals vs Ghosts," Oct 1 2025[5]

Four people. Four labs. Four completely different sets of incentives. And the same diagnosis, arrived at independently: predicting the next token well is not the same thing as reliably doing something in the world. That gap, between sounding right and being right enough to bet on, is the whole story of this post.

Even the person with the least reason to admit it, admits it

Here's the data point that actually got to me. Not a quote, a drift. Sam Altman runs the best-funded, most model-obsessed lab on the planet, the one with the least reason on Earth to say the model itself isn't the whole answer. Watch what happened to his own explanation of what's missing, over eighteen months[6]:

Sam Altman, in his own words
Mar 2025"I think the models just aren't smart enough yet."
Jun 2025"The moment it gets to a problem it can't solve, it falls apart."
Apr 2026"I no longer think of the harness and the model as entirely separable things… we don't even have a primitive [for an agent's login vs. a human's]."
Capability gap becomes reliability gap becomes infrastructure gap. Same person, three answers, and each one walks further away from "just make the model bigger." That's not somebody losing an argument. That's somebody paying close attention and updating in public.

I want to be a little cynical here, because it's earned. Everyone's stated bottleneck lines up suspiciously well with what their own lab already happens to be good at. Amodei runs a safety-first lab that can't outspend the frontier on compute, and he names interpretability as the crisis. Altman is sitting on more compute and distribution than anyone alive, and he names infrastructure, his own backyard. Sutskever can't outspend OpenAI, Google, or Anthropic on hardware, and he says ideas beat scale, which happens to be exactly what someone without the hardware needs to be true. None of that makes any of them wrong. It just means you shouldn't read these as neutral technical opinions. They're bets, and everyone above already put their chips down before they told you what they believe.

02 · where they violently disagree

Patch the current thing, or replace it

Once you separate the diagnosis from the treatment, the field splits into three camps, and none of them are bluffing. Real money, real published research, real companies built around each bet.

Altman · Amodeiscale + reasoning scaffolds
HassabisGemini scales, Genie adds a leg
Sutskever"age of research," undisclosed
Suttonexperience, not imitation
LeCunJEPA, non-generative
extend the current architecturereplace it outright

This is our read of where each person has put their own research program. Treat the spacing as illustrative, not a measured scale.

The "replace" camp isn't rhetoric. It's funded, and it has proofs.

Yann LeCun quit Meta in November 2025 and raised a billion dollars for AMI Labs on an argument he'd been making in public for three years[7]. It's not just a slogan. There's real math under it. Picture generating text as a walk through a tree, where only some branches lead somewhere correct. If a model has some probability of stepping off that tree at any given token, the odds of the whole answer staying correct over many tokens fall off exponentially:

P(correct) = (1 − e)n

Every wrong step drags the next one with it. That's his basis for calling autoregressive LLMs "doomed"[8]. Let's actually run the numbers, because the formula is more convincing than the soundbite:

LeCun's exponential-drift argument, worked out
assume e = 1%per-token chance of stepping off the correct path
at n = 100 tokens(0.99)¹⁰⁰ ≈ 36.6%
at n = 500 tokens(0.99)⁵⁰⁰ ≈ 0.7%
That's our own arithmetic on LeCun's formula. He never publishes a specific value for e, so we picked one to make the shape of the curve visible. It's not a settled argument, either: a Stanford review found the math checks out, but the assumptions holding it up (constant error rate, independent mistakes, no way to recover) get shakier every time someone ships a model that can catch and fix its own error mid-answer[9]. Real critique, not a dismissal.

Sutton's bet is a different shape entirely. He's not arguing about architecture, he's arguing about what the model learns from in the first place. His claim: training on human-written text is this decade's version of hand-coded chess heuristics. It works great, right up until the moment something learns from its own experience instead of copying ours. He didn't just write an essay about it. He started a company, Oak Lab, to go build it[10].

The extend camp's own favorite example got taken apart by its own logic

Here's a story I love, because it's the field roasting itself in public. In February 2024, Berkeley's AI research group argued the future belongs to compound systems, lots of models and tools stitched together, not one giant model. Their star exhibit was Medprompt, a clever ensembling pipeline that beat raw GPT‑4 by nine points on medical exams[11]. Nine months later, people from that same research circle published the paper that quietly took their own exhibit apart:

"Even without prompting techniques, o1-preview largely outperforms the GPT‑4 series with Medprompt… [in-context steering] may no longer be an effective approach for reasoning-native models."Nori et al., "From Medprompt to o1," Nov 2024[11]

A model trained with RL on verifiable rewards ate, in under a year, everything a clever hand-built pipeline had to sweat to fake. Watching the bitter lesson happen in real time, inside one model generation, is the best evidence the extend camp has going for it. It doesn't prove extending is enough forever. It proves the floor keeps moving under anyone betting against the base model.

03 · what the data says about the actual gap

Reliability doesn't arrive as a byproduct

Fine, so everyone agrees the missing piece is reliability, not raw intelligence. Great. How far has that piece actually moved? This is where it stops being comfortable. The 2025 to 2026 literature here is a lot more specific, and a lot more uncomfortable, than "AI agents keep getting better every month" lets on.

Benchmarks overstate what actually ships

METR tracks how long a task an AI agent can handle on its own before it starts falling apart. Credit where it's due: they turned around and fact-checked their own number, which is rarer in this field than it should be. They took Claude Sonnet 4.5's measured "time horizon" (the task length it completes correctly half the time) and regraded it against a harder, more honest bar. Would a real software maintainer actually merge the pull request it produced[12]?

~50 min
time horizon, graded against
an automated test-passing check
→
~8 min
the same model, regraded against
real maintainer merge decisions
the overstatement, worked out
50 min ÷ 8 min≈ 6.25×
METR's own framing"roughly a sevenfold overstatement"
Call it six to seven times, take your pick. And it's getting worse, not better: METR measured the gap widening roughly 9.6 percentage points a year[12]. The benchmark isn't catching up to reality. It's drifting further from it.

Karpathy has a name for why a smarter model doesn't just fix this on its own: "the march of nines." Each extra decimal point of reliability costs roughly a full project's worth of work, same as the last one. It isn't a discount you earn from the model getting smarter. Pair that with what he calls "jagged intelligence": the same model that just cleared a PhD-level problem will turn around and insist 9.11 is bigger than 9.9[5]. I've watched this happen. It's funny until it's in production.

Voting across models doesn't average the errors away. It can't.

There's a comforting story everyone tells themselves: sure, one model might be unreliable, but run three and take a vote and the errors should cancel out. A 2026 audit tested that story directly, across 30 models, and it falls apart in a way you can actually measure[13]:

~100%
of trials where some combination
of 3 models could beat the best one alone
vs.
10–19%
of trials where simple majority
voting actually captured that gain
the "voting tax," worked out
100% − 10–19%≈ 81–90%
That's our framing of their reported numbers, not a line lifted straight from the paper. Here's the mechanism underneath it: stronger models make more similar mistakes, not more independent ones. That's the exact opposite of what a vote needs to be true, and it gets worse, not better, as the models you're voting across get smarter[14].

There is a real fix. It's just not "run three models and see what two of them say." The Weaver method combines 33 weak, individually unreliable verifiers using actual weak-supervision statistics, needing almost no ground truth to calibrate against, and it took one open model from 72.2% to 93.4% on math and reasoning benchmarks. It beat a noticeably stronger single model doing that[15]. It works. It just needs real statistics behind it, not a show of hands.

More agents help, and hurt, depending on the shape of the task

In December 2025, someone finally tested this properly. 180 configurations, four benchmarks, one direct question: does adding more agents to a task actually help[16]? The answer is yes, and also no, depending entirely on whether the task breaks apart cleanly.

Finance-Agent
decomposable, parallelizable
+80.9%
PlanCraft
sequential, constraint-satisfaction
−70%

Same lever. Opposite sign, depending which task you pull it on. The same study found what they call a "baseline paradox": once a single agent already clears about 45% accuracy alone, adding more agents actively makes it worse, with error amplification measured as high as 17.2× in the bad regime. Whether coordinating multiple agents helps is now something you can calculate, not something you can assume. That single fact should change how a lot of "multi-agent" products get built, ours included.

04 · how fast, and how sure

Timelines compressed. Certainty didn't.

Researcher surveys have yanked their own predicted arrival dates forward, hard, over the last few years. Worth showing exactly how much, and exactly how shaky the ground under that trend actually is.

AI Impacts survey of published AI researchers

2022 survey 2060 2023 survey 2047
Median year for a 50% chance of "High-Level Machine Intelligence." n=2,778 published authors[17].

Metaculus community forecast

early 2022 ~2042 mid-2026 Oct 2028
Median date for "weakly general AI," ~1,800 forecasters[18].
the compression rate, worked out
predicted date moved2042 → 2028 = 14 years pulled in
calendar time elapsedearly 2022 → mid-2026 ≈ 4.5 years
14 ÷ 4.5≈ 3.1 years pulled in, per year that passed
That's our arithmetic across exactly two snapshots, so treat it as a feel for the pace, not a trend line you'd bet a company on. There are early, low-confidence signs in 2026 that the compression is reversing a little, including inside Anthropic's own internal policy posture[18].

Here's the finding that should worry the field more than any timeline number. The AAAI polled 475 researchers plus a 24-person expert panel on what's actually blocking progress right now. Compute didn't win. Data didn't win.

75% of researchers agreed that a lack of rigor in evaluating AI systems is holding back AI research progress. Only 8% disagreed. Not FLOPs. Not tokens. The field looked at itself and said the bottleneck is evaluation[19].

05 · follow the money

Does any of this actually show up in the numbers?

A research consensus is one thing. Money actually moving toward it is a different, more honest test. This is the part where the two lined up more than I expected when I started pulling numbers.

Enterprise generative-AI spend more than tripled in a single year. That's a disclosed figure from a 495-company survey, not somebody's extrapolation off a napkin:

$11.5B
2024
$37.0B
2025

3.2× YoY. Menlo Ventures, "State of Generative AI in the Enterprise," Dec 9 2025 (n=495 US enterprise AI decision-makers)[20].

Two analyst firms sized the AI orchestration category completely independently, different methods, different base years, and landed on almost the same growth rate. That agreement matters more than either dollar figure on its own:

the orchestration-market anchor
Grand View Research$9.76B (2024) → $58.9B (2033), 22.4% CAGR
MarketsandMarkets$11.0B (2025) → $30.2B (2030), 22.3% CAGR
Different base years, different endpoints, nearly identical growth rate. The agreement is the signal here, not either headline number by itself[21].

Now shrink that down to just the coding-agent slice. Claude and OpenAI together hold 63% of code-generation spend, out of $8.4B in total enterprise LLM API spend[20]. Code-gen itself is worth $1.9B of that pool, which works out to a 23% share. Multiply that share against the orchestration market above and you get:

a derived slice, shown as math, not asserted
$9–11B × 23%≈ $2.1–2.5B → "$2–3B" today
That's a derivation, not a published figure. Nobody sizes this exact slice, so we're showing you the multiplication instead of just handing you a number and asking you to trust it.

And here's the part the market already settled on its own, without waiting for anyone's opinion. Routing a prompt to whichever model is cheapest for the job used to be something you'd build. Now it's something you just get. AWS shipped it into Bedrock in April 2025 and claimed a 60% cost saving over running everything through one model[22]. Microsoft had an equivalent in Copilot by November. Even OpenAI shipped GPT‑5 with its own internal router, then walked part of it back in December 2025 after free-tier users complained they'd lost the ability to just pick a model themselves. Sit with that reversal for a second. The complaint wasn't "routing is bad." It was "I didn't ask you to hide it from me." People wanted the choice made in front of them, not behind their back.

06 · so where does that leave anyone building on top of this

The checklist, not the vibe

Line all of this up and you don't get a vibe. You get a checklist you could actually go verify yourself:

01

The binding problem is reliability and generalization, not raw model IQ, and that diagnosis comes from the people with the least reason to hand it to a critic for free.

02

Reliability doesn't fall out of a bigger model. Verification, self-correction, and voting all have specific, currently unsolved failure modes, and the correlated-error problem gets worse, not better, as models improve.

03

Whether coordinating multiple models helps or actively hurts is now a measurable property of the task, not a matter of taste.

04

Cost-based routing between models is already a commodity feature inside three hyperscalers. Building a company around that specific claim, alone, is picking a fight you can't win on price.

05

The field's own practitioners rank evaluation rigor, not compute, not data, as the most binding constraint on progress. That's an unglamorous, under-invested answer, and it's also the one thing a governance and confidence layer is actually built to address.

We're building Hydra, a local-first control plane that routes AI coding work across models instead of betting on just one. I didn't write any of the above to justify that. I wrote it blind, on purpose, and only went back to check our own roadmap against it afterward, which is a more uncomfortable exercise than it sounds when you actually do it. The honest result was mixed. Some of it held up better than I expected. Some of it needed fixing on the spot.

Take the voting tax from three sections back: naive majority vote across models captures only 10 to 19% of the improvement that's mathematically sitting on the table, because stronger models make correlated mistakes, not independent ones. That's not a reason to give up on combining models. It's a reason to stop voting and start accumulating evidence properly. Here's the actual math behind Hydra's confidence-routing layer (see First Principles for the full derivation), not a metaphor for it:

why confidence beats a vote, worked out
3 calibrated headseach independently 75% accurate
log-likelihood ratio per voteln(0.75 / 0.25) ≈ 1.10
threshold for 95% confidenceln(0.95 / 0.05) ≈ 2.94
votes needed2.94 ÷ 1.10 ≈ 2.7 → 3
This is Wald's sequential probability ratio test, decades-old statistics, not something we invented[23]. Accumulate log-likelihood instead of counting hands, and three genuinely diverse, calibrated heads cross a 95% confidence threshold on their own, no ten-model swarm required. The word doing all the work in that sentence is "diverse." Run three copies of the same model family and you're right back in the correlated-error problem from earlier. Run one Anthropic model, one open-weight local model, and one from somewhere else entirely, and the math actually holds.

That's the part of the roadmap this research made us more confident about, not less. The part it made us less confident about is any race-and-vote swarm mode that doesn't check for that diversity in the first place. Those need to become confidence-weighted, not majority-counted, or they're just an expensive way to make the same mistake three times and call it consensus.

That's the whole point of doing it this way around. If the research had come back saying the real value is in routing by cost, I'd have had to say so, and go rework the plan. It didn't. It said the value is in confidence, verification, and governance, which happens to be where we'd already started walking. Good news checked twice is still good news. It's just better news when you were willing to hear the bad kind too.

references

What this rests on

  1. Model releases, Jun to Jul 2026: Anthropic, "Claude Fable 5 and Claude Mythos 5," Jun 9 2026, and "Introducing Claude Opus 5," Jul 24 2026; TechCrunch, "OpenAI launches its new family of models with GPT-5.6," Jul 9 2026; TechCrunch, "Google releases three new Gemini models, but no 3.5 Pro," Jul 21 2026; LMArena's live leaderboard for the July restore and rebaseline (exact cause not independently confirmed here). anthropic.com/news/claude-opus-5
  2. Sutskever, I.: "We're moving from the age of scaling to the age of research." Dwarkesh Podcast, Nov 25 2025. dwarkesh.com/p/ilya-sutskever-2
  3. Hassabis, D.: via Semafor, Jan 2026; own essay "A Framework for Frontier AI and the Dawning of a New Age," Jul 14 2026. demishassabis.substack.com
  4. Sutton, R.: Dwarkesh Podcast, Sept 26 2025; "Era of Experience" (Silver & Sutton, DeepMind, 2025). dwarkesh.com/p/richard-sutton
  5. Karpathy, A.: "Animals vs Ghosts," Oct 1 2025; YC AI Startup School talk ("the march of nines"), Jun 19 2025. karpathy.bearblog.dev/animals-vs-ghosts
  6. Altman, S.: Stratechery interviews, Mar 20 2025 & Apr 28 2026; OpenAI Podcast Ep.1, Jun 18 2025. stratechery.com
  7. AMI Labs: $1.03B seed, Mar 10 2026; LeCun's Meta departure, Nov 11 2025. techcrunch.com/2026/03/09
  8. LeCun, Y.: "Auto-regressive LLMs are doomed," Paris AI Action Summit, Feb 2025; original (1−e)ⁿ argument, X, Mar 2023. LeJEPA, arXiv:2511.08544
  9. Miranda, B.: technical review of LeCun's error-compounding argument, Stanford CS, May 26 2026. cs.stanford.edu/people/brando9
  10. Sutton, R.: OaK architecture keynote, RLC 2025, Aug 30 2025; Oak Lab launch coverage. mlq.ai/news
  11. Zaharia, Khattab, Chen et al.: "The Shift from Models to Compound AI Systems," BAIR blog, Feb 18 2024; Nori et al., "From Medprompt to o1," arXiv:2411.03590, Nov 6 2024. arxiv.org/abs/2411.03590
  12. METR: "Measuring AI Ability to Complete Long Software Tasks," arXiv:2503.14499, Mar 2025; "Many SWE-bench-Passing PRs Would Not Be Merged into Main," Mar 10 2026. metr.org/blog
  13. Kim, S.: "Are Diversity Metrics Measuring Diversity?," arXiv:2607.20768, Jul 2026. arxiv.org/html/2607.20768v1
  14. Zhou, Xu, Zhou, Singh, Gui, Joty: "Variation in Verification," arXiv:2509.17995, 2025. arxiv.org/pdf/2509.17995
  15. Saad-Falcon, J. et al.: "Shrinking the Generation-Verification Gap with Weak Verifiers" (Weaver), arXiv:2506.18203, 2025. arxiv.org/html/2506.18203
  16. Kim, Liu et al. (MIT / Google DeepMind): "Towards a Science of Scaling Agent Systems," arXiv:2512.08296, Dec 9 2025. arxiv.org/abs/2512.08296
  17. AI Impacts: "Thousands of AI Authors on the Future of AI," arXiv:2401.02843, Jan 2024 (rev. Oct 2025), n=2,778. arxiv.org/abs/2401.02843
  18. Metaculus: "Date Weakly General AI is Publicly Known" community question, live tracker. metaculus.com/questions/3479
  19. AAAI: 2025 Presidential Panel on the Future of AI Research, report, Mar 7 2025 (24-member panel + 475-respondent survey). aaai.org (PDF)
  20. Menlo Ventures: "2025: The State of Generative AI in the Enterprise," Dec 9 2025 (n=495); "2025 Mid-Year LLM Market Update," Jul 31 2025. menlovc.com
  21. Grand View Research & MarketsandMarkets: AI Orchestration Platform Market reports, 2025. grandviewresearch.com
  22. AWS: "Amazon Bedrock Intelligent Prompt Routing," general availability announcement, Apr 2025. aws.amazon.com/whats-new
  23. Wald, A. & Wolfowitz, J.: sequential probability ratio test and its optimality proof, Annals of Mathematical Statistics, 1945 & 1948. Classical statistics, cited here for the log-likelihood accumulation mechanism, not a new result.
Live-verify before external use: a few items above rest on secondary paraphrase where primary pages blocked automated fetches. Treat direct quotes from essays/arXiv/official blogs as high-confidence, podcast-transcript quotes as verified-but-secondary, and anything marked "reported" as needing one more independent check before it leaves this document.