🔬 Research / /via sakana.ai / updated 17h ago

Sakana AI’s ‘AI Scientist’ Project Lands Nature Publication on Automated Research

Sakana AI’s fully automated “AI Scientist” research agent has been formally described in a new open-access paper in Nature. The work, developed with collaborators at UBC, the Vector Institute, and the University of Oxford, details a system that can generate ideas, run machine learning experiments, and write complete papers that pass human peer review. It matters because it shows AI systems beginning to automate core parts of scientific research, hinting at a future where AI-generated science scales far beyond human capacity while raising new questions about rigor and oversight.

#SakanaAI#UniversityofBritishColumbia#UBC#VectorInstitute#UniversityofOxford#NeurIPS#ICLR
~/ Research/ Sakana AI’s ‘AI Scientist’ Project Lands Nature...

Sakana AI has reached a new milestone in automated research, with its "AI Scientist" project now described in a paper published in Nature. The work lays out a vision of an AI agent, powered by foundation models, that can execute the entire machine learning research lifecycle from end to end. The Nature publication is open access and consolidates both the initial system and its improved successor, AI Scientist-v2, along with new insights into how AI-generated science can be scaled.

The AI Scientist first drew attention when Sakana AI released a preprint showing that an agent could be given a starting code template and then autonomously generate novel research ideas, design and run experiments, and write a full paper. Building on this, the team introduced AI Scientist-v2, which produced what they describe as the first fully AI-generated paper to pass a rigorous human peer-review process. That paper went through blind review at the ICLR 2025 "I Can’t Believe It’s Not Better" workshop, where one manuscript achieved an average score that exceeded the typical human acceptance threshold and outperformed a substantial share of human-authored submissions before being voluntarily withdrawn.

The Nature paper credits a close collaboration between Sakana AI, the University of British Columbia (UBC), the Vector Institute, and the University of Oxford. Together, the authors provide a comprehensive description of the AI Scientist’s architecture and workflow, from initial idea generation through to automated reviewing. The publication also situates the project within broader advances in foundation models, emphasizing how improvements in the underlying models have enabled the system to tackle more open-ended research directions in AI.

Under the hood, the AI Scientist is designed to take a broad research direction in machine learning and then autonomously handle tasks that would traditionally be split across a human research team. It generates novel ideas, searches for and reads relevant literature, and then designs, programs, and conducts computational experiments using a parallelized agentic tree search. Once results are in, the system writes the entire paper in LaTeX, even receiving feedback on its figures from a foundation model with vision capabilities, and then passes the draft to an automated reviewing module.

A key innovation described in the Nature article is the "Automated Reviewer," built to evaluate AI-generated science at scale without exhausting human reviewers. This reviewer is prompted to act like an Area Chair, ensembling several independent reviews into a final decision based on official NeurIPS guidelines. When benchmarked against thousands of human decisions from the OpenReview dataset and compared with results from a prominent NeurIPS consistency experiment, the Automated Reviewer closely matches human performance and in some measures exceeds the agreement levels seen among human reviewers themselves.

The team reports that the Automated Reviewer also aligns well with human judgments on AI papers from top conferences such as ICLR, including papers that appeared after the model’s own training cutoff. Crucially, by using this reviewer to score papers generated by different foundation models, they observed a clear scaling law: as the capabilities of the underlying models improve, the quality of the generated papers increases. This pattern suggests that as compute becomes cheaper and model capabilities continue to rise, future versions of the AI Scientist could become substantially more capable, potentially pushing beyond current human performance benchmarks in research workflows.

Why this matters

The publication of the AI Scientist in Nature signals that AI-driven automation is moving from speculative demos into the center of mainstream scientific discourse. By demonstrating that an AI system can not only generate ideas and run experiments but also produce papers that pass blind human peer review, Sakana AI and its collaborators are testing where the boundary of human-only science actually lies. The work raises practical questions for conferences, journals, and research institutions about how to evaluate AI-generated work, how to ensure methodological rigor, and how to integrate automated systems into existing peer-review and authorship norms.

Despite the breakthrough, the authors are explicit about current limitations. The AI Scientist can still produce naive or underdeveloped ideas, struggle with deep methodological rigor, and make obvious mistakes such as hallucinated citations or duplicated figures in appendices. For now, it remains confined to computational experiments and does not yet operate in physical laboratories or other domains where experimental constraints and safety considerations are more complex. The Nature paper frames these shortcomings as early-stage constraints, noting a broader machine learning trend in which new capabilities, once they begin to work at all, tend to become superhuman as scale and improved core models rapidly enhance performance.

Looking ahead, Sakana AI expects that the playbook documented in the Nature publication will be adapted to domains beyond machine learning, with the potential to catalyze truly open-ended scientific discovery. If the observed scaling laws hold, future versions of the AI Scientist could generate and test far more hypotheses than human teams can feasibly manage, especially in fields where computational experimentation is viable. For the wider ecosystem of AI research, peer review, and scientific publishing, this project offers an early glimpse of a world where automated agents are not just tools but active participants in the creation and evaluation of new knowledge.

source sakana.ai →
share
𝕏 FB
← cd ../news