Industry Insights

When progress starts improving itself

Inside the workshop where AI is learning to improve AI

By Merlin Fitz-Hall

An AI agent changes some training code, runs an experiment and checks the result. If the change helps, it stays. If it makes things worse, out it goes. Then the agent tries again.

Repeat for two days.

In Andrej Karpathy’s March account, roughly 700 changes produced about 20 improvements. A great many ideas went into the bin. Research, it turns out, remains research even when the researcher has no use for a kettle.

Karpathy subsequently tested whether the improvements carried over to a larger training run. Human judgement was still part of the process. Nevertheless, a small research workshop had acquired a computational night shift, working on the machinery used to build AI.

Elon Musk’s response on 9 March was characteristically restrained: “We are in the Singularity.” Farzad’s 8 October video brought that claim back into view.

I think the workshop is more interesting than the proclamation. Something important is happening inside it: intelligence is becoming an input into producing more intelligence. The tools are beginning to improve the tools.

A workshop with a feedback loop

Recursive self-improvement, or RSI, is the idea that an improvement feeds back into the system doing the improving. A research agent alters its working methods, tests the result and retains a better version. That version becomes the starting point for another round.

At the ambitious end, an AI system could develop a more capable successor, which becomes better at developing the next one. If each generation accelerates the process, familiar expectations about technological progress could unravel. That possibility sits behind much of the singularity discussion.

There are several levels between a useful feedback loop and an accelerating succession of superhuman researchers. The evidence already reaches further along that path than many people realise. It also stops well short of proving the whole journey.

Karpathy’s autoresearch project provides a deliberately bounded example: training code, a fixed experiment budget, a measurable objective. Its value is in making the experiment-and-test cycle cheap enough to repeat relentlessly.

Google DeepMind’s AlphaEvolve connects Gemini’s code proposals to automated evaluation. It found a 23% speed-up in a computational routine used during Gemini training, saving about 1% of overall training time. The system helped improve the machinery behind its own underlying technology. It did not independently produce a succession of new Geminis.

Then there is Weco’s September AIDE² study. Over eight autonomous days, its system accepted seven successive improvements to a research agent’s code, including how it searched and managed memory. Gains transferred to held-out tasks under fixed evaluation budgets. The foundation models remained fixed; the software organising their work improved.

That is bounded recursive improvement, demonstrated. A separate test of whether the improved agent had become a better improver was inconclusive. The distinction matters: producing a better research tool and reliably accelerating the production of better research tools are different achievements.

Inside OpenAI, agents now undertake well-defined research assignments representing days of human work, under supervision. People choose the priorities, judge promising results and decide what to scale or deploy. Computational work is feeding into the development of future systems through a research organisation that still depends on people.

The workshop is changing at several benches simultaneously. We can see that much without declaring the building has achieved escape velocity.

My bet on the next twelve months

By October 2027, I expect AI to be doing a substantially greater share of the work that improves AI, including research planning within bounded tasks, experimentation and refinement of the agents themselves. I expect the practical gains to be large enough that treating today’s capabilities as a stable basis for many long projects will become an expensive mistake.

This is my forecast, conditional on these gains continuing to transfer into useful work. Systems with broadly human-level usefulness across substantial areas of knowledge work belong in serious planning scenarios for that period. Their scope and reliability, and how quickly organisations can use them, remain uncertain.

The AI 2027 scenario explores how automated research might compound. Its dates help tell that story; they are assumptions to test against events. Anthropic’s own account says that fully autonomous successor development has not been reached and RSI is not inevitable.

My confidence is stronger about the direction than the timetable. The decisive evidence will be improved research systems repeatedly producing better successors, with gains that hold up outside their original tests. More experiments only help if more valuable knowledge comes out.

The invention still has to leave the workshop

Suppose tomorrow’s AI produces a brilliant result. The proof checks out.

What happens next?

A discovery needs significance: does it solve an important problem? It needs an application, engineering and a route into actual use. A beautiful new bridge design must become a bridge that people can safely cross, in a particular place, using materials that someone can supply.

Each stage asks different questions. A valid calculation cannot tell us by itself whether we should build the bridge, whether the foundations suit the ground or how traffic will work while the old one is closed.

These stages can become more automated too. Danaher’s planned autonomous laboratory, announced this month for operation at scale in early 2027, aims to connect AI design with robotic manufacture, testing and feedback. Its performance remains to be demonstrated, but the intended loop reaches into the physical world.

The path from discovery to delivery contains plenty of friction. Some of it will resist acceleration; some will become the next thing we learn to improve. Treating every present bottleneck as permanent would be a remarkably adventurous form of conservatism.

Build for the tools that may arrive during the build

For business leaders, the immediate challenge is painfully practical. A capability that looks bespoke and expensive at the start of a programme could become widely available before it goes live. The project plan may arrive in production as a historical document.

That is a hypothesis to test against particular investments. It does not establish that buyers have collectively stopped buying, or that every long programme is doomed.

It does argue for useful results sooner, shorter commitments where possible, and deliberate opportunities to change components. At each stage, ask what has become easier since the last decision. A project that cannot absorb improving tools may spend heavily recreating something the market has just made ordinary.

I also expect the visible machinery to recede. Most people will want a trusted assistant to organise a result. Specialist agents and services can work underneath it. Managing a committee of bots should rarely become the user’s new occupation.

This could move a substantial share of software’s value away from the interface people click. Business context, dependable data, integrations and controls still have work to do behind the assistant. Some suppliers will earn a more valuable role there. Others will find that the part customers paid for has become easy to substitute.

Neither an established brand nor a large installed base guarantees protection. Equally, a disappearing interface need not mean the disappearance of useful infrastructure.

The test is whether the complete workflow improves. DORA’s research with Google engineers describes how verification and integration can consume gains from faster code production. A working demo is an invitation to investigate what remains between the demonstration and dependable operation.

Leave room on the bench

I find the prospect of this expanding workshop enormously hopeful. More people could gain the capacity to test an idea, pursue a scientific question or build something that previously required resources they could never assemble.

Human and computational intelligence can widen one another’s possibilities. Keeping that partnership beneficial requires deliberate choices about control, safety and who receives the gains. Faster progress makes those choices more consequential.

There may never be a universally agreed morning when someone opens the workshop door and announces that the singularity has arrived.

We can still put the tools to work. And leave room on the bench for what we discover together.


Merlin Fitz-Hall is a computational intelligence consultant working with the team at Zenpath.

Source notes

  • Musk’s original remark is dated 9 March 2026. Farzad’s 8 October video prompted this article; the technical examples link to their primary sources.
  • Karpathy’s post reports the experimental run; the autoresearch repository explains its fixed-budget method. Karpathy performed the later transfer check.
  • AlphaEvolve’s training benefit concerns a computational routine and total training time, rather than overall research productivity.
  • AIDE² demonstrates improvement of an agent’s software scaffold with fixed foundation models. Whether its improved agent was a better improver remained unresolved by the reported test.
  • OpenAI and Anthropic describe their own internal use. These are primary accounts from organisations with a direct interest in the technology, and they retain human supervision.
  • AI 2027 is a scenario. Danaher’s announcement describes planned capabilities and timing. Neither is treated as an achieved outcome.
  • DORA’s March 2026 article analyses 1,110 open-ended survey responses from Google software engineers, collected in Q3 2025. Its findings concern that population, rather than the entire software market.
Industry Insights