I spent three decades building software that did precisely what it was told, no more and no less. SMAC ran the queries I designed. FedMine, which I built alone over four years before an acquisition let me retire, tracked federal spending exactly as I'd specified, down to the last appropriations line. Software, in my working life, was a very obedient kind of clay. It never occurred to me that it might one day pick up its own hands and start shaping itself.
That assumption quietly expired this year, and almost nobody noticed it happen.
A Machine Wrote the Science
On March 26, 2026, Nature published a paper describing a system its makers call, without much modesty, “The AI Scientist.”[1] It came from Sakana AI, a Tokyo lab, working with researchers at the University of British Columbia, the Vector Institute, and Oxford. The system does something that would have sounded like science fiction at any other point in the history of science: given a broad research direction, it generates its own novel hypotheses, searches and reads the relevant literature, designs and runs the experiments, writes the results up as a full paper complete with figures, and then puts the work through review.
Not metaphorical review. Actual peer review. An earlier version of the system had already produced a manuscript that went, unedited, into the double-blind review process of a workshop at ICLR, one of the top machine-learning conferences, with the organizers' permission. It scored 6.33 out of 10, above the average human acceptance threshold and higher than 55 percent of the human-authored papers in the same pool.[2] Under a protocol agreed in advance, the researchers withdrew it before publication. But the point had been made. A machine had written science good enough to pass the humans whose job was to judge it.
The score a fully AI-generated paper received from human peer reviewers at an ICLR 2025 workshop. Hypothesis, code, experiments, figures and prose were all produced by the machine.
The Reviewer in the Machine
The part of the paper I keep returning to isn't the writing, though. It's the reviewing. Sakana's team also built an Automated Reviewer, an AI trained to play Area Chair, the senior role that weighs conflicting referee opinions and renders a verdict, and benchmarked it against thousands of real human decisions from OpenReview's conference archives. It reached 69 percent balanced accuracy, comparable to human reviewers. On one measure, F1 score, its agreement with the actual outcomes exceeded the level of agreement human reviewers managed with each other in NeurIPS's well-known 2021 consistency experiment.[1]
We have grown used to AI generating things faster than we can evaluate them. This is the other half of that problem quietly resolving itself: AI evaluating things, at human competence, so that the generation doesn't have to wait for us at all. The bottleneck that has throttled science since the invention of the journal, that a finding is only as fast as the humans available to judge it, has started to dissolve.
Every station that once needed a human
Generate novel hypotheses from a broad research direction; search and read the literature.
Write the code, design and run the experiments, collect the results.
Draft the full paper, figures and all. One passed human peer review.
An Automated Reviewer plays Area Chair, matching human accuracy. Then the loop begins again.
Forty Years Ahead of the Hardware
Then, in the same month I'm writing this, Sakana made a second announcement that matters more than the first, because it explains why the first one wasn't an accident. The company formally launched a Recursive Self-Improvement Lab, and it named as Chief Scientific Advisor a German computer scientist named Jürgen Schmidhuber, who has spent nearly forty years waiting for this moment.[3][4]
In 1987, as a diploma student, Schmidhuber described a program that could rewrite its own learning algorithm,[5] an idea with no way to be tested at any meaningful scale, because the compute needed to try it wouldn't exist for another three decades. He kept building on it anyway: work on meta-learning, on world models, and, in the early 2000s, on something he called the Gödel Machine, a theoretical system that could formally prove an improvement to itself was actually an improvement before making it.[6] For most of his career, this was pure theory, decades ahead of the hardware that could run it, the kind of research a person does on faith. He is now advising a lab that builds the thing he described in his diploma thesis, nearly forty years before anyone could switch it on.
For most of his career, this was pure theory: the kind of research a person does on faith.
Small Compute, Compounding Returns
Sakana's own path there has a shape worth noticing. In 2024, a project called LLM-Squared used language models to invent better ways of training language models, and produced DiscoPOP, a state-of-the-art preference-optimization algorithm discovered and written entirely by an LLM. In 2025, the Darwin Gödel Machine, named pointedly after Schmidhuber's old thought experiment, maintained an evolving population of AI agents that rewrote their own code and more than doubled their own baseline score on SWE-bench, a standard software-engineering benchmark, from 20 to 50 percent.[7] A related system called ShinkaEvolve reached a state-of-the-art solution to a classic optimization problem using only about 150 samples, and went on to discover a new load-balancing method for mixture-of-experts models,[8] a degree of efficiency that has nothing to do with the brute-force scaling most of the industry has bet on.
The strategic argument underneath all of it, which Sakana states plainly, is that recursive self-improvement may not require the two or three hyperscale compute clusters currently racing each other in the United States and China.[3] It might run, instead, on the more modest but well-designed compute a country like Japan actually has. That echoes a question I raised about China's Kimi K3 and what it really runs on: how much of the frontier is raw scale, and how much is ingenuity under constraint? If Sakana is right, the capability to build self-improving intelligence stops being a resource monopoly and becomes something closer to an engineering discipline anyone disciplined enough can practice.
2024 — an AI invents a better way to train AI. 2025 — AI agents rewrite their own code and double their benchmark score. 2026 — an AI Scientist passes peer review, an AI reviewer matches human judgment, and a lab is founded to close the loop on purpose.
The Notebook Writes Itself
I don't think most people have registered what it means that the loop has closed: that an AI Scientist can now, in a bounded but real sense, improve the AI that builds the next AI Scientist. It is one thing to say a machine can answer our questions. It is another thing entirely to say a machine can now formulate its own questions, design its own experiments, judge its own conclusions, and hand the improved version of itself the pen to do it again, faster, on the next iteration. That is not a tool. That is the beginning of a process with a compounding rate of return, the possibility I first sketched in The Mehan Dispatch, “When the Machine Becomes Its Own Teacher”, and compounding processes are the only kind of process in the history of the universe that have ever produced something genuinely new fast enough to notice within a human lifetime: stars, life, us. Our institutions, as I argued in The Clock Speed Problem, are not built to keep pace with one.
Schmidhuber waited nearly forty years for the hardware to catch up to an idea he had as a young man doing arithmetic on paper about programs that could rewrite themselves. History rarely announces the moment it turns; it just quietly stops needing the people who used to be indispensable to it, and keeps going. The scientific method spent four centuries, from Galileo's telescope onward, building the machinery of the peer-reviewed journal. It may take considerably less time to get from here to whatever comes after us being the ones who read the journal at all. The future doesn't wait for us to notice it's already begun. It just keeps a very good lab notebook, and lately, it's started writing the notebook itself.
Sources & Further Reading
Primary sources are Sakana AI's own announcements and papers; independent coverage is noted where it adds dates or context.
- Sakana AI, “The AI Scientist: Towards Fully Automated AI Research, Now Published in Nature” (March 26, 2026), including the Automated Reviewer's 69% balanced accuracy and F1 comparison with the NeurIPS 2021 consistency experiment. Paper: “Towards end-to-end automation of AI research,” Nature. ↩
- Sakana AI, “The AI Scientist Generates its First Peer-Reviewed Scientific Publication”: the ICLR 2025 workshop experiment, scores of 6, 7 and 6, organizer permission, and pre-agreed withdrawal. ↩
- Sakana AI, “Introducing Sakana AI's Recursive Self-Improvement (RSI) Lab” (September 2026), including the argument that recursive self-improvement is achievable on modest, sample-efficient compute. ↩
- The Decoder, “Sakana AI hires Jürgen Schmidhuber” (September 24, 2026), on his role as Chief Scientific Advisor to the RSI Lab. ↩
- Jürgen Schmidhuber, Evolutionary Principles in Self-Referential Learning, diploma thesis, TU Munich, 1987. ↩
- Jürgen Schmidhuber, “Gödel Machines: Self-Referential Universal Problem Solvers Making Provably Optimal Self-Improvements” (2003). ↩
- Sakana AI, LLM-Squared and DiscoPOP (2024); Zhang et al., “Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents” (2025), SWE-bench 20.0% → 50.0%. ↩
- Sakana AI, “ShinkaEvolve” (2025): sample-efficient program evolution, including the ~150-sample result and a new mixture-of-experts load-balancing loss. ↩