Whereâs the âintelligence explosionâ?
One of my fundamental beliefs about the world is that Ramez Naam ought to blog more. Ramez is one of the worldâs greatest futurists â he predicted the solar and battery revolutions long before these were widely understood. If you were reading Ramez in 2011, you were able to understand the future of both energy technology and climate change, long before other people did. His earlier book More than Human is still a great guide to the kind of biological enhancements that AI might make possible. Ramez is also an excellent science fiction author, having written a trilogy of novels in which nanotechnological telepathy is distributed as a party drug (Iâm not sure if he actually expects that to happen, but itâs a very cool idea).
Unfortunately, although he does have a Substack (which you should absolutely follow), Ramez does not blog regularly. However, after having a lengthy private debate with him about Recursive Self-Improvement, I was able to prevail upon him to write up his thoughts for my blog.
To say that RSI is a big deal in the AI world would be a colossal understatement. Among AI researchers, entrepreneurs, and AI safety people, thereâs a widespread belief that as AI gets better at improving itself, there will be a âfast takeoffâ or âFOOMâ, in which AIâs capabilities âtake offâ and create a technological Singularity. This event is a staple of science fiction, including works by my favorite sci-fi author, Vernor Vinge.
A lot of people in the industry believe that this moment is now close at hand, and are racing toward that prize:
But Ramez â normally among the most wide-eyed of techno-optimists â is highly skeptical that weâll see anything like the âFOOMâ of Vernor Vinge novels. In this lengthy, well-researched post, he explains his skepticism.
Personally, Iâm agnostic. Ramezâs case necessarily rests on a lot of assumptions; although itâs cogently laid out, I think the real answer is that weâll just have to wait and see whether the Singularity arrives. But even more fundamentally, I donât know how much this debate matters in the practical sense â even without the kind of Singularity depicted in sci-fi novels, AI capabilities are improving so rapidly that theyâre already superhuman in many respects, and soon will probably be strongly superhuman in most or all dimensions. The AI of 2040 is going to look godlike, whether or not it explodes into an actual god in 2027.
Still, itâs a very interesting argument, and Ramezâs thoughts on the future of technology are always worth listening to.
1. AI is Helping Improve Itself
AI is already helping improve itself. The question is whether even fully autonomous recursive self-improvement (RSI) would cause a runaway intelligence explosion.
The theory is that each generation of AI could build a better successor, faster than the last generation did. That could lead to a âfast takeoff,â with capabilities surging to artificial superintelligence (ASI) in a year, months, or even days.
Hereâs my take: Given our best current data, the AI self-improvement loop would need to be roughly 5â10Ă stronger to sustain itself, let alone run away. Iâll explain this math in section 8. I expect incredibly rapid AI progress by the standards of nearly any other technology. But the evidence we have doesnât suggest a sudden explosion to incomprehensible superintelligence anytime soon.
I could be wrong. Forecasters have repeatedly underestimated AI progress! I could well be next. One thing thatâs clear is that we need better data. For now, letâs work with what we can measure, and stay open to breakthroughs that could change the picture.
How Strong Is the Feedback Loop?
Figure 1. How strong is the self-improvement loop? Model.
Contents
Hereâs the case, with links to each part:
Key charts: The feedback loop ¡ Measured vs. forecast progress ¡ Diminishing returns
2. What Does RSI Mean?
People use ârecursive self-improvementâ to mean everything from AI boosting the productivity of human researchers to AI bootstrapping itself to incomprehensible intelligence. Hereâs my taxonomy: productivity gains (Type 1), increasing autonomy while still facing diminishing returns (Types 2â4), and a runaway loop to superintelligence if we can ever find accelerating returns (Type 5).
Figure 2. Five types of AI self-improvement.
Weâve made real progress on Types 1 and 2: AI helps both researchers and engineers inside of AI companies, and powerful models can train and improve smaller ones. We havenât yet seen clear evidence for Type 3 (though Alibaba just made some strong claims) and certainly not for Type 4. I do expect autonomous self-improvement to arrive at some point. Iâm skeptical that it leads to Type 5 - runaway super-intelligence - without a major conceptual breakthrough.
There are plenty of other definitions of RSI, which can be a bit confusing. Wecoâs four levels of RSI are close to mine. For a broader tour of all the things people mean when they say âRSIâ, read Tom Cunninghamâs comprehensive guide.
We Already Have Narrow Superintelligence
I do expect narrow superintelligence in highly verifiable domains. Think chess, Go, formal math, parts of computer science and coding. Highly verifiable domains are largely formal and structured types of work where machines can generate unlimited training data, with perfect or near-perfect verification of correct vs incorrect, and do so entirely in software without waiting on the physical world or humans. Thatâs an ideal setting for AI learning.
Figure 3. What makes a domain highly verifiable?
In fact, we already have narrow superintelligence in game plang. Weâre seeing it happen now in the most formal parts of math, in particular in proofs and in finding counter-examples that disprove major conjectures. For example, OpenAI recently reported an AI-generated proof resolving the NavierâStokes existence and smoothness problem. Parts of software development are also extremely verifiable, while others are a bit less crisp (such as understanding what humans want).
That isnât the same as broad superintelligence. Even our most powerful models need far more training data than humans, struggle to learn reliably from ongoing experience, and fail in surprising ways on tasks people find straightforward. Superhuman math doesnât automatically mean superhuman judgment everywhere else.
3. Real AI Research is Harder than Benchmarks or Forecasts
Benchmarks and forecasts suggest that AI models should reliably succeed at coding tasks that take humans hours, without human help. The real world is messier. OpenAIâs internal data shows much shorter stretches of autonomous work on research tasks.
In its Research Acceleration / RSI report, OpenAI showed how often its models completed tasks with and without human help, grouped by how long a human would need to do the work.
Figure 4. OpenAIâs internal research tasks. Source.
Even on tasks that would take a human less than 15 minutes, OpenAIâs models succeeded without human intervention only 86% of the time. The estimated task length at 80% success was roughly 15 minutes over the first seven months of the year. Julyâs results were similar to the whole period average.
Fully autonomous RSI would require an AI to string together a great many research tasks reliably, stretching out over complex tasks that humans need weeks or months to accomplish. OpenAIâs data suggests that we arenât close.
Anthropic also released a graph showing how Claude accelerates AI research. It shows that internal AI models collaborate on or even lead more than 90% of R&D tasks. Thatâs objectively impressive. At the same time, the graph reports zero cases of AI autonomously completing AI R&D tasks.
Figure 5. Claudeâs role in internal AI R&D. Source.
These are incredible tools. But they still need skilled people to set direction and get them back on track.
The Gap Between Benchmarks and Reality
For years, METR has been publishing a chart showing what length of coding task (measured in human hours to complete) best-in-class AI models can achieve. Itâs been called the most important graph in AI. METRâs Mythos Preview evaluation estimated that the model could succeed at 80% of coding tasks that took humans three hours.
Figure 6. METRâs 80% task horizons. Source.
From ECI Scores to METR Task Horizons
Epochâs own rule of thumb is that every five additional points of ECI (their overall benchmark of AI capability) correspond to roughly a doubling of METRâs task horizon. Using that formula, weâd expect GPT 5.6 Sol and GPT 6 Astra to be 80% successful at completing tasks of around 4 hours and 11 hours of human length, respectively.
Another estimate (a forecast) of AI task length comes from the AI 2027 scenario, which estimated that by July 2026, frontier AIs would be 80% successful accomplishing tasks of around 11 hours. Fairly similar.
The AI 2027 Tracker charts all of these.
Figure 7. The AI 2027 Tracker. Source.
Inside OpenAI, though, the July research-task horizon at 80% success was roughly 15 minutes.
Hereâs the gap:
Measured Progress vs. AI 2027 and ECI-extrapolated METR
Figure 8. Forecasts, benchmarks, and real AI research. Tracker ¡ OpenAI.
A four-hour benchmark horizon is about 16 times longer than OpenAIâs research horizon. AI 2027âs 11-hour forecast is about 44 times longer. Of course, the tasks being performed by researchers at OpenAI arenât the same as those in the METR benchmark. So we should expect some discrepancy. This, however, goes well beyond that.
Actual AI research at OpenAI is an order of magnitude or more harder than metrics, benchmarks, or forecasts suggest. That should make us wary of relying too much on benchmarks, or of saying that future scenarios like AI 2027 are âon track.â The authors of the related AI 2040 project still describe AI 2027 as roughly the future they expect, and say reality is tracking closer to it than even they expected. Thatâs not what we see from within OpenAI. This isnât an apples-to-apples comparison, but the difference is remarkable. AI 2027 appears to be substantially over-optimistic in this regard.
In January of this year, Nathan Witkin made a case that the METR graph was exaggerating progress. The real world data suggests that at least some of his critiques were correct. The gap between benchmarks, forecasts, and data gleaned from actual use of AI should influence our expectations about the future.
4. The Sharp Diminishing Returns to Impressive AI Numbers
OpenAIâs report also shows impressive increases in AI token usage, in compute spend per researcher, and in lines of code written. But these arenât results. Theyâre intermediate measures. How much progress do they actually drive?
Researchers used 124x more tokens per person. Engineers shipped roughly 7x as many lines of code per person. Researchers ran 1.6x as many experiments per researcher vs OpenAIâs 2025 whole year average.
Figure 9. Token use inside OpenAI. Source.
Figure 10. Experiment pace inside OpenAI. Source.
From More Tokens to More Experiments
Figure 11. From tokens to code to experiments. Source.
More tokens and code donât tell us much on their own. The 1.6Ă experiment pace is closer to useful research output. Even that doesnât mean AI is improving 1.6Ă faster.
An enormous increase in AI output has accompanied a much smaller increase in experiments run.
This isnât a controlled experiment. We donât know what would happen if researchers switched back to an older model. But it gives us a useful view of AI-assisted research inside a frontier lab.
Itâs not just OpenAI. Anthropic reports that their engineers are now producing 8x as many lines of code per person as they did in 2024 - somewhat similar to OpenAI. Anthropic also sees significant diminishing returns between productivity and AI progress. Hereâs a direct quote from its Mythos Preview system card:
âProductivity uplift does not translate one-for-one to capabilities progress. We surveyed technical staff on the productivity uplift they experience from Claude Mythos Preview relative to zero AI assistance. The distribution is wide and the geometric mean is on the order of 4Ă. [âŚ] We estimate that reaching 2Ă on overall progress via this channel would require uplift roughly an order of magnitude larger than what we observe.â- Anthropic, Claude Mythos Preview System Card; emphasis mine
Translation: To double the pace of AI progress, Anthropic estimates that AI would need to increase the productivity of their employees by roughly a factor of 40 relative to no AI assistance.
Figure 12. Anthropicâs productivity-to-progress estimate. Source.
This is an estimate, not a measurement of progress. Even the 4Ă productivity figure comes from an opt-in survey of 130 Anthropic staff. I put more weight on OpenAIâs logged experiments, though the two sources measure different things.
We donât yet know how much those extra experiments are accelerating AI improvement, if at all. In general, there are also steeply diminishing returns of more experiments in most branches of science. That means that a 60% increase in experiment pace could be on the order of a 10% boost to AI improvement pace. (A power law exponent of 0.2, for those who want to do the math.) Thatâs speculation for now. Weâll learn more as the labs publish results.
Test Time Compute Also Has Diminishing Returns
What about giving the same AI model more time to think?
That scales badly also. In OpenAIâs recently publicized results on unsolved math problems, success rises roughly with the log of compute over the range shown. It shows logarithmic diminishing returns. In plain English, each additional doubling of compute for a model buys roughly the same gain in success rate, while costing twice as much.
Figure 13. Test-time compute and math performance. Source.
What About Agent Swarms?
What if we throw more agents at it instead? A common RSI / ASI idea is that once we have AIs at a certain capability level, we can just spawn more copies and put them to work.
Adding agents can get tasks done faster and sometimes reach a higher capability level. But on the three benchmarks in Toby Ordâs analysis, expanding a swarm buys less improvement per token than letting one agent think longer.
His rough rule of thumb is a square root. If one agent can accomplish a task in 10 hours, then 100 agents could accomplish it in one hour. The speedup is 10, the square root of the number of agents (100). But to get this speedup, you increase the total cost in tokens or run time compute by the same factor. So going from one to 100 agents can get a task done in one tenth the time. But itâll be ten times as expensive.
Parallel agents can save time, at a much higher compute cost.
Another challenge is that agents often think alike. In a study comparing LLMs with 467 people, the first ten AI responses offered collective creativity comparable to about eight to ten people. After that, roughly two extra AI responses added as much as one extra human response. A separate study across model families also found less diversity in AI responses. That doesnât mean every agent has the same idea. But a hundred copies may offer less variety than a hundred different researchers.
None of this makes swarms useless-or safe. Lisan al-Gaib makes a strong case for parallel agent swarms as a potent cyber-weapon in âAccidental Scaling.â I donât share all of his assessment of what swarms have accomplished. In math, for example, I think he gives far too much credit to the swarm and not enough to the better internal model that OpenAI used.
OpenAI says the model behind its NavierâStokes result was developed through âlarge-scale reinforcement learning on top of a previously pretrained model.â Formal math is a highly verifiable domain, which makes it a particularly good fit for that approach: Machines can generate nearly limitless amounts of training data, and verify that solutions are correct or incorrect, all in software. My guess is that this modelâs full results will show an especially large improvement in math.
OpenAIâs Noam Brown made the central point explicitly: he wouldnât give multi-agent methods even 10% of the credit for the NavierâStokes result.
I do think Lisan makes good points about cybersecurity. If youâre searching for a security vulnerability at a target site and can divide the search among agents, speed may justify a huge token bill. Swarms can be dangerous even when theyâre inefficient.
Iâm less convinced that this scales to research breakthroughs. Inventing something like the transformer probably takes more than searching a space someone has already defined.
5. Better Models Matter More Than More Copies
Building a better model can bring gains that extra thinking time or more copies of the old model canât. Look at the gap between Astra and OpenAIâs internal model on the same math problems.
Figure 14. Better models versus more thinking time. Source.
Thatâs the strongest version of the RSI argument: a more capable AI could do research that todayâs model canât do, however many copies we run.
But building that better model also runs into diminishing returns. More training data, more training compute, larger models, and more reinforcement-learning (RL) compute all show diminishing returns in published scaling studies. Making dense models larger usually raises the compute needed for each output token, too. None of these routes gives us a free pass around the problem.
Figure 15. Diminishing returns to scaling. Chinchilla ¡ ScaleRL ¡ OpenAI.
Those scaling results give us reason to expect diminishing returns when AI helps build the next model, too.
6. Weâre Not Seeing Runaway Acceleration
AI capabilities are rising quickly. But the public data doesnât show a sustained acceleration. To the extent that AI tools are boosting productivity, they may be being offset by the problems growing harder. Or we may simply be early. Either way, the trend isnât showing a fast takeoff.
Figure 16. Frontier ECI gains since January 2024. Source.
The public ECI frontier-the best score among models released by each date-has gained about 16 points a year on a trend fitted from January 2024 through September 2026. Thatâs blisteringly fast progress, but this period doesnât show a runaway surge.
Hereâs the same frontier in absolute ECI points, through July 2026, to put it in perspective.
Figure 17. The absolute frontier ECI score. Source.
The public frontier also canât tell us everything happening inside the labs. Anthropic gives us a closer look in the Opus 5.5 system card, using its own version of the index, AECI.
Figure 18. Anthropicâs fitted capability trend. Source.
Eli Lifland, a co-author of AI 2027 and AI 2040, saw the apparent trend break as a warning that we were heading toward an intelligence explosion:
âAnthropic is probably right here [that they hadnât reached dangerous levels of AI self-improvement], but alarm bells should be going off! Our processes are not ready to handle an intelligence explosion and we appear to be going full-steam ahead toward one.â
- Eli Lifland, On Mythosâs AI R&D Capabilities
What looked like acceleration now appears more consistent with a one-time jump. The level went up. The rate hasnât kept climbing.
Keeping Up the Pace Takes Exponentially More Resources
Achieving those gains has required an enormous increase in the inputs to AI. For example, consider computing power. Epochâs estimates of AI chip capacity, measured in NVIDIA H100 equivalents, show roughly 127-fold growth in just over three years (including projections at the end of this period).
Figure 19. AI chip capacity and frontier ECI. Source: Epoch AI.
This is total AI chip capacity, including inference. Still, the increase is striking: vastly more computing capacity has accompanied much steadier gains in measured capability.
The broader picture looks similar. Here are six inputs alongside capability gains, going back to February 2023.
Figure 20. Six inputs alongside frontier ECI. Epoch chip data ¡ SemiAnalysis workload shares.
Everywhere we look, AI has diminishing returns. It gets more expensive in treasure and talent to make each step forward. More of every input has been required to maintain steady gains in AI capabilities.
Weâve been able to scale these inputs because, until recently, the cost was within the scope of what hyperscalers could pay from their profits. That is no longer the case. From this point forward, future AI investment will increasingly depend on AI revenues going up. And the scale of the numbers - 3% of US GDP is now going into AI infrastructure - suggests that eventually the growth rate will decline. If investment growth does slow, to anything less than its current blistering exponential pace, capability progress could slow too. Even if investment growth continues (which I expect for the foreseeable future) a slowdown from its current exponential growth rate to a more modest one (which I also expect) could lead to a slower pace of progress. Better AI research tools may be needed to offset that.
Better AI May Be Needed Just to Maintain the Pace
The day when we need better AI tools just to continue the pace of AI progress may already have arrived. Not because investment is slowing, but because the problem of improving AI itself gets harder at each step.
Hereâs Anthropic in the Mythos 5.1 system card:
âwe believe that internal usage of recent AI models has been a key factor in maintaining the current rate of progress, but we do not yet see clear signs of dramatic acceleration beyond that rate.â- Anthropic, Claude Fable 5.1 & Claude Mythos 5.1 System Card, section 2.3 â emphasis theirs.
The key word is maintaining-and Anthropic italicized that word in its own system card. Increasingly capable AI may be essential just to keep the pace of improvement where it is.
Gains on Other Benchmarks Donât All Carry Through to Research
Opus 5.5 improves substantially on several coding and computer use benchmarks. But on CoBench, Anthropicâs benchmark built from historical AI R&D problems, it gains just 2.6 percentage points over Opus 5, within the reported error bars.
Figure 21. Opus 5.5 benchmark gains. Source.
Why the smaller gain here? Maybe AI research is simply harder than other tasks. Bear in mind that CoBench isnât testing the ability to produce significant discoveries. Itâs much more limited in scope. It asks models to investigate historical AI R&D problems using code, logs, and documents. Thatâs useful research debugging and productivity work, but it doesnât directly test whether a model can invent a new architecture or make a conceptual breakthrough.
The evidence on open-ended research suggests another obstacle: coming up with useful ideas that havenât already been tried.
7. Why Does Progress Get Harder?
Better Ideas Get Harder to Find
Why do useful new ideas often get harder to find?
Tom Cunningham and Manish Shetty have a useful apple-picking metaphor. An AI can pick the low-hanging fruit quickly, while humans can still reach ideas the AI canât.
Once those apples are picked, another copy of the same agent finding them again doesnât help. A stronger model can reach higher. To add my own flourish, the apples may also get sparser and farther apart as you climb. The RSI question is whether each harvest gives us enough to build a better apple-picker.
Figure 22. The apple-picking model of AI R&D. Source.
This pattern shows up across R&D. Bloom and colleagues document fields where research effort grows while research productivity falls. A famous example is Eroomâs Law: in the historical drug-development data, the inflation-adjusted R&D cost per new approved drug roughly doubled every nine years.
Figure 23. Eroomâs Law in drug development. Source.
Pharma has other complications, including regulation, difficult clinical trials, and rising expectations for safety. Existing treatments can also raise the bar for a useful new drug. But some of this difficulty may also be that the low-hanging fruit has been picked.
Lessons from Software R&D
Stockfish, the chess engine, gives us a more direct look at software research. We have records of experiments aimed at improving it and the gains that followed. This gives us a real-world dataset to look at the gains of experimentation in software. As a result, several RSI models draw on this data. That said, not all the improvements came from these experiments. Several important ideas also came from outside the project, so we shouldnât give its experiments all the credit.
Epochâs analysis of software R&D estimates returns to research effort at about 0.83 for Stockfish, a bit slower than linear. These are diminishing returns, but gentle ones. These returns, however, are improvements in computational efficiency. And more compute does not turn directly into more AI capability. As we saw earlier, AI capability also has steep diminishing returns from adding more computational power. So we shouldnât read that 0.83 as the return from experimentation to AI capability itself. AI capability grows much more slowly than compute, as weâve seen already.
Andrej Karpathyâs autoresearch demonstration gets closer to the process we want to understand. A âteacherâ AI agent changes a smaller âstudentâ AI modelâs training code, runs it, checks the result, and tries again. The teacher agent itself doesnât improve, but it is able to improve the âlearnerâ. This is my Type 2: A stronger AI improves a weaker one.
One public run, posted by an agent operating on Karpathyâs behalf, reported 89 experiments over roughly 7.5 hours. About 92% of that sessionâs gain arrived by run 44. Gains came quickly, then slowed. The setup was deliberately small, with a five-minute training budget per experiment. But the agent could change the architecture, optimizer, and training settings; it wasnât limited to a handful of knobs.
Figure 24. Gains in one autoresearch run. Source.
A later public run got further, so the first run hadnât hit a hard ceiling. This is a useful early example of autonomous research, and yet another place where we see the diminishing returns endemic in AI research. That said, this was a very early experiment. I expect future systems to do much better. This particular AI improvement loop will likely grow stronger.
From More Activity to Better Ideas
This is where the distinction matters. More tokens can buy more code, and more code can help us run more experiments. But experiments only improve AI if they uncover something useful.
Figure 25. From AI activity to useful improvements.
AI Still Struggles With Big Research Ideas
The bigger question is whether AI can come up with ambitious new research ideas or conceptual breakthroughs.
Anthropicâs description of Opus 5.5 is blunt:
âAs with previous models, it is weaker on open-ended research: internal users report that it mostly tests incremental ideas and prefers less ambitious hypotheses, and in our human-run biology exercise, it deferred to the published literature and struggled to develop novel ideas (Section 2.2.2).â- Anthropic, Claude Opus 5.5 System Card, section 2.3.3; emphasis mine
METRâs assessment in the same card identifies what may still be missing:
âThis is highly uncertain, but we expect that full automation of AI R&D will require large improvements in foresight, prediction, creating oneâs own feedback loops, and generally other skills that might typically be referred to as researcher âjudgementâ or âtasteâ.â- METR, quoted in the Claude Opus 5.5 System Card, section 2.3.6
In these examples, humans still supply much of the direction and judgment.
Future models will probably get better at this. But in the worldâs stockpile of potential training data, we have many more examples of incremental work than of breakthroughs. I wonder whether that makes novelty harder to learn. Thatâs speculation, but worth watching.
This is also tough to address by simply running more copies of the AI. A huge number of parallel agents can help with the incremental improvements or searching over a large set of parameters, but for breakthrough ideas they may run into the homogeneity problem: More parallel agents still think alike.
8. The Self-Improvement Loop Doesnât Look Strong Enough
How far are we from the self-improvement loop being strong enough to sustain itself, or to propel itself into runaway super-intelligence? Can we quantify this?
We can make a rough estimate. Better AI helps with research; useful research produces better AI. For the loop to sustain itself, each round must produce enough gains to propel the system through the next loop, even as improvements get harder to discover.
Figure 26. The AI self-improvement loop. Model.
In a recent paper, The Economics of Recursive Self-Improvement, Tom Cunningham and colleagues modeled this from the standpoint of how much more productivity every point of additional ECI produces from an AI. They ask first and foremost what that number would need to be to create a self-sustaining feedback loop. And secondly, they try to determine what that productivity-per-ECI-point number is today.
First, they find a self-sustaining RSI threshold of roughly 15% more research productivity per extra ECI point. In their model, thatâs about where better AI would generate enough progress to sustain the loop.
The picture below shows the idea. At the threshold, each cycle of gains powers the next. Above the threshold, the feedback loop accelerates. Below the threshold, the feedback loop is too weak, and the rate of improvement it brings drops on each cycle. This model isolates the software loop; outside investment can still drive rapid progress.
Figure 27. Three illustrative feedback paths. Source.
Updating this slightly with data from the Stockfish experiments puts the threshold a little higher, at roughly 19% per ECI point. I wouldnât put much weight on that precise difference. Both estimates are uncertain. But they give us a way to think about the strength of the feedback loop and a rough band at which self-sustaining or runaway RSI may begin.
How Fast Are Gains Coming Now?
The second thing Cunningham and team do is make a rough estimate that the current AI productivity gain is about 9% per ECI point. Thatâs below their self-sustaining threshold.
I like the model. OpenAIâs newer data, however, suggests the loop may be quite a bit weaker.
Cunninghamâs estimate of 9% productivity gain per ECI point is based on Anthropicâs survey of 130 staff, who reported roughly 4Ă the productivity theyâd have without AI. Cunningham and colleagues compare that with a 16-point capability gain since early Claude Code.
That comparison assumes the earlier tools added little or no productivity, so âno AIâ is a reasonable starting point. The authors say this explicitly. Iâm not sure the assumption holds for the same researchers doing the same work, but thatâs a smaller issue.
The authors themselves know that this is a rough calculation, and warn that the 4Ă survey estimate is probably too high.
OpenAIâs newer data gives us a firmer way to check the number: Actual logged experiments over time, rather than human estimates of their own productivity with and without AI. I put more weight on this for three reasons:
Direct and broad measurement. Instead of relying on surveys, OpenAI actually tracked and measured experiments run on their infrastructure. That means they didnât rely on researchers estimating their own productivity, which can be far off.
Full sample, not opt-in. Similarly, OpenAIâs data catches every active experimenter, while Anthropicâs only reflects the 130 employees who took the time to answer the survey â and who therefore may not be a representative set.
Enormously more data. We donât know how many experiments are in the 32 weeks of OpenAI data, but itâs likely at least tens of thousands of individual examples and possibly hundreds of thousands.
Any way you slice it, the new OpenAI data, released after Cunninghamâs paper was drafted, is a larger, more comprehensive, more representative, and almost certainly more accurate dataset than Anthropicâs internal opt-in survey of employees.
Now letâs use OpenAIâs experiment data to calibrate the productivity gain per ECI point. We know that in August, OpenAI researchers ran ~1.6Ă as many experiments per person per month as the 2025 average. If we pair that with roughly 16 points of frontier ECI improvement, it works backward to about 3% productivity gain per point of ECI. By contrast, 9% compounded over 16 points would mean roughly 4Ă productivity.
Figure 28. Comparing productivity estimates. OpenAI methods.
Hereâs OpenAIâs published weekly series alongside that hypothetical path of 9% more productivity per additional ECI point. The blue line ends at ~1.6Ă. The red line shows what 9% per point would imply if 16 ECI points were spread across this period. That doesnât match what we see from OpenAIâs data. I want to be clear here that all data sets are noisy. We donât know exactly what model researchers were using on what days, or whether the new experiments were also higher quality than old experiments. We need more experiments and more data to further calibrate these numbers. Working with what we do have, what we see is a quite low boost to productivity from each additional ECI point.
Figure 29. Experiment pace versus a hypothetical path. Source.
Even that 3% could give better models too much credit. OpenAI also used far more tokens and had more compute for experiments. Those could account for some of the increase in experiment pace. So the range is probably a bit lower.
I use 2â3% productivity gain per ECI point as a working assumption, allowing for some help from those other inputs. This is still a rough estimate, albeit one thatâs based on the best real-world data we have.
Figure 30. Productivity estimates and the takeoff threshold. Source.
With those assumptions, 2â3% per ECI point against a 15â19% threshold leaves a roughly five- to tenfold gap. Thatâs a big gap, though its size depends on how well experiment counts capture useful research and whether the assumed capability change is right.
Figure 31. Diminishing returns around the loop. Source.
AI is helping build better AI. Under this estimate, though, each turn of the loop adds less than the last. The feedback would have to become much stronger to sustain itself.
9. What Could Accelerate This?
This software loop sits alongside faster chips, bigger data centers, more training data, and greater investment. Those can keep driving rapid progress even if the loop canât sustain itself.
The loop itself could strengthen too. Better training data, memory, and research judgment could all help.
A breakthrough on the scale of the Transformer architecture in 2017 could change the picture much more. That would be a good reason to revisit these estimates.
Better researchers might also run fewer experiments and learn more from each one. A handful of better ideas can matter more than a mountain of routine runs.
Still, diminishing returns in machine learning arenât new. Cortes and colleagues were fitting machine learning scaling curves in 1993: More examples reduced error, following a power law with diminishing returns. These diminishing returns and harsh scaling laws are as old as machine learning. They didnât appear for the first time with transformers or LLMs or deep learning. That doesnât prove todayâs relationships will last forever. But until we see evidence that weâve found a new approach that scales without these inhibitors, we should plan for diminishing returns as likely to be with us for some time.
Software, Hardware, and Economic Feedback
That said, the world is more than just software. Tom Davidson, Basil Halperin, Thomas Houlden, and Anton Korinek model software progress, hardware progress, and economic feedback together. Better AI helps design better chips; better chips support better AI; economic growth finances more investment in both. Several feedback loops can combine to overcome diminishing returns even when one loop alone canât. I think itâs fantastic that someone has attempted a model that integrates all these different avenues of improving AI through software, hardware, and economics.
But I have questions about the software loop itself. In their central calibration, fully automating software research puts that loop roughly at the threshold for explosive growth, even without help from better hardware or broader economic growth. Recall that Cunninghamâs model puts the self-sustaining threshold at roughly 15% more research productivity per additional ECI point, while our estimate using OpenAIâs experimental data puts todayâs gains at only 2-3%. These models use different measures, so we canât equate their numbers directly. But the contrast matters: their fully automated software loop reaches the threshold, while our best estimate from current data puts todayâs loop far below it.
Having AI do all the research doesnât eliminate the diminishing returns inherent to improving AI, or the broader problem of useful ideas getting harder to find. This is the distinction between Type 4 and Type 5 in the taxonomy above. An AI might autonomously design, train, and test its successor, and still need exponentially more resources to make each additional step forward. Closing the loop doesnât tell us whether itâs strong enough to sustain itself.
The authors do account for diminishing returns. The concern is whether their calibration overestimates how much useful AI research each round of software improvement produces. Diminishing returns appear to be fundamental to machine learning. We see them in training, in test-time compute, and in the search for better algorithms. Full autonomy could remove human bottlenecks without removing any of those constraints.
Weâve already seen this within autonomous research. In the Karpathy autoresearch example above, most of the gains arrived early, and more experiments bought progressively less improvement. That was a small experiment with a fixed teacher model, not a test of fully autonomous RSI. It doesnât settle the question. But it illustrates why removing the human from an experiment loop doesnât, by itself, remove diminishing returns.
I do expect the feedback loop to get stronger over time. Better AI should become better at research. But based on our best current data, reaching self-sustaining feedback requires a loop roughly five to ten times stronger than todayâs. Treating fully automated software research as already at that threshold is a substantial leap, before we add the benefits of hardware improvements or economic growth. I could be wrong, but Iâd like to see evidence that autonomy brings enough additional useful discoveries to close that gap.
On hardware, I have some further reservations. The model doesnât explicitly include the years it can take to turn a chip design into deployed hardware. The authors discuss physical bottlenecks, and Iâd like to see manufacturing and construction delays built into the predictions.
I also wonder how much past chip progress came from better ideas, and how much depended on ever more expensive factories and equipment. If we give researchers too much credit for gains that also needed those investments, we could overestimate what faster AI research alone would produce.
Even with those reservations, this is the most compelling paper and model Iâve seen for combining feedback loops in software, hardware, and economics to understand how fast they could push AI forward. Iâm not convinced it establishes that a fast AI takeoff is possible under realistic conditions. More data could help us calibrate that judgment. But it gives us a useful framework for understanding what could happen beyond the software layer alone.
This is an important paper that helps us model AI as part of a broader economy that might have larger feedback loops around it. I appreciate it, and Iâm glad they wrote it.
10. We Need More Data
These estimates rest on less data than Iâd like. I might be putting too much weight on a few observations and reaching a comforting conclusion I want to believe. We need better measurements, shared often enough to catch changes as they happen.
When OpenAI released its research data, Cheryl Wu welcomed the disclosure and pointed out how much was still missing. More tokens and experiments are useful things to know about. We also need to see how they turn into better algorithms and more capable AI.
Figure 32. Cheryl Wu on OpenAIâs research data. Source.
Now Wu, Arjun Ramani, and Basil Halperin, with their colleagues at the Elasticity Institute, have written a concrete proposal: How to Measure RSI. It lists eight things the labs could share to help answer these questions. Check it out.
Figure 33. Eight proposals for measuring RSI. Source.
Iâd especially like to see how much useful research each new model adds, holding resources roughly constant, and how that research translates into better AI. Thatâs how weâll learn whether the loop is getting stronger.
What the Future Holds
AI is already helping build better AI. Itâs improving at a stupendous pace, and I expect that to continue. We already have narrow superintelligence in chess and Go. I expect increasingly superhuman performance in parts of formal math, coding, and cybersecurity, and any other verifiable domain where machines can generate training data and verify success at machine speed. Those are powerful capabilities. That doesnât mean weâre close to super-intelligence for less verifiable, messier, open-ended work - or to a general ASI.
Iâm skeptical of a fast takeoff to super-intelligence, but evidence matters more than hunches. Letâs collect the data we need to get a clearer picture of whatâs happening. Including evidence that could change our minds. If better AI starts producing enough useful research to make the next round easier, I want to know. If the gains keep shrinking, I want to know that too.
This is a great post! I wanted to build on the research being harder than the benchmarks and forecasts.
People with management experience know how hard it can be to specify and direct what needs to get done, and how this is qualitatively different than just executing on a task. Specification itself is hard and doesn't always have just one solution.
Benchmarks necessarily do some of the specification work. For people who haven't seen it yet, go look at the OpenAI GDPval prompts. This is meant to model performance on economically valuable, real-world tasks. The prompt is very clearly doing the specification work that wouldn't be done so neatly "in the wild".
Verifiable domains like math are then where AI excels because the verifier *is* the spec.
Ashby's Law of Requisite Variety says that only variety can absorb variety: for a system to successfully control or regulate another (or its environment), it needs to be able to tackle a greater variety of states/responses than what its controlling.
You can think of specification as this regulation done in advance: it anticipates the ways a task can go wrong and constrains the work against them.
Human accumulate variance through experience, and the collaborative process within organizations are also part of the variety that can absorb variety. This is why METR says that full automation will require more "foresight, prediction, creating oneâs own feedback loops" but I think it also depends not just on AI feedback loops, but AI-human feedback loops. That's part of the variety process, and why I'm skeptical of reaching Type 4 (humans add no research value)
I don't think super intelligence will come soon, either. Reading this, it makes me think that intelligence has an innate cost. The human brain weighs 1300 grams, a humming bird's weighs 0.13 games. Humans are clearly far more intelligent than hummingbirds, but 10000x? The math in the article also shows there is exponential resources required to sustain linear capability growth.
Part of intelligence is being able to draw conclusions across different pieces of seemingly unrelated information. However, this is a network effect, which is exponential. There is no magic sauce which can join together infinitely large information spaces. As we learn more about intelligence, I believe we will discover a power law, similar to the inverse square law for energy dissipation. Every level of intelligence requires a certain amount of cross connectivity, and capability cannot grow beyond that. Exponential takeoff is impossible.
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content â general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached â you'll always get the same 5 for this article.