🔬 Research / /via arxiv.org / updated -116m ago

DeepScientist Claims AI Can Outpace Human Research on Three AI Tasks

DeepScientist used more than 20,000 GPU hours to generate about 5,000 scientific ideas and experimentally validate roughly 1,100 of them. The Westlake University system surpassed human-designed state-of-the-art methods across three frontier AI tasks, according to its arXiv paper. Its goal-oriented, memory-driven search offers a model for AI research systems focused on measurable scientific progress rather than open-ended idea generation.

#WestlakeUniversity#DeepScientist#arXiv#ResearAI
~/ Research/ DeepScientist Claims AI Can Outpace Human Resea...

A research team from Westlake University has introduced DeepScientist, an autonomous AI system designed to pursue scientific discoveries over month-long timelines. In an arXiv paper, the team says the system generated about 5,000 unique scientific ideas and experimentally validated approximately 1,100 of them after consuming more than 20,000 GPU hours.

The system reportedly surpassed human-designed state-of-the-art methods on three frontier AI tasks: agent failure attribution, LLM inference acceleration, and AI text detection. The paper gives improvement figures of 183.7%, 1.9%, and 7.9%, respectively, although the supplied text does not specify which figure corresponds to each task.

DeepScientist is built around a goal-driven formulation of scientific discovery as a Bayesian Optimization problem. Rather than asking an AI system to produce research ideas without a tightly defined objective, the researchers direct it toward finding a novel method that maximizes a target performance metric.

Its workflow is organized around a repeated “hypothesize, verify, and analyze” process. A cumulative Findings Memory records earlier research knowledge and experimental results, allowing the system to balance exploration of unfamiliar approaches with exploitation of directions that have already shown promise.

The approach uses hierarchical evaluation to avoid giving every idea the same level of experimental scrutiny. Promising findings are selectively promoted to higher-fidelity validation, while large-scale parallel exploration supplies a broader pool of hypotheses for the system to examine.

The researchers began from human state-of-the-art methods associated with the three selected tasks and asked DeepScientist to conduct continuous research. The paper says the system operated for about a month on 16 H800 GPUs, with its progress on AI text detection described as comparable over two weeks to several years of human research.

Why this matters

DeepScientist targets a central weakness of earlier AI Scientist systems: the ability to generate plausible research proposals without reliably producing work that addresses important human-defined problems. By tying autonomous discovery to explicit benchmarks and iterative validation, the system moves the focus from idea volume to measurable improvement over established methods.

The results also point to a potentially different role for AI in research. Instead of functioning only as an assistant for literature review or experiment design, a system with persistent memory and automated testing could search through many candidate methods and concentrate resources on the strongest ones. The scale of the reported experiment, however, also shows that this style of discovery depends on substantial computational experimentation.

DeepScientist presents its findings as large-scale evidence that an autonomous AI system can progressively surpass human-designed state of the art on scientific tasks. The team identifies the work as a step toward fully automated discovery, while the system’s open-source project and code are intended to support further investigation of how such research loops perform across other problems.

source arxiv.org →
share
𝕏 FB
← cd ../news