The AI scientist is about to leave the computer
An AI agent receives a tray of protein samples. It instructs a liquid handler to transfer them, tells a robotic arm where to move the plate and asks a reader to measure the result. It notices that one liquid behaves differently from another, adjusts the flow rate and tries again.
This is no longer entirely speculative. On 27 August, Anthropic announced the Model Hardware Standard, or MHS: a common interface intended to let AI agents communicate with programmable scientific instruments. Early demonstrations suggest that agents can already coordinate laboratory equipment and optimise parts of an experiment in a closed loop.
The bigger ambition is to connect two capabilities that have largely developed separately. AI systems can reason over scientific papers, datasets and possible experiments. Automated laboratories can perform repetitive physical tasks. Connect the two and an agent could, in principle, propose an idea, test it, interpret the outcome and decide what to do next.
That would turn AI from a scientific adviser into something closer to a laboratory colleague. The commercial question is whether companies can make AI-driven scientific discovery reliable enough to produce results that are not merely fast, but useful, reproducible and valuable.
Anthropic connects AI agents to laboratory equipment
Most laboratory equipment was not designed with AI in mind. Microscopes, plate readers, liquid handlers and robotic arms often use different control systems and proprietary software. Connecting several machines into one automated workflow can require weeks or months of bespoke engineering.
Anthropic’s Model Hardware Standard provides a shared software layer through which an AI agent can discover, monitor and control compatible devices. Anthropic says MHS can reduce some integration work from weeks or months to hours or minutes. The standard is model-agnostic and works with programmable equipment; Anthropic plans to make it open source after further safety work.
In one proof of concept, Genentech used Claude and MHS to coordinate a liquid handler, robotic arm and plate reader for a standard protein-concentration assay. Claude conducted trial transfers, analysed the readings and adjusted the liquid-handling speed. It arrived at different settings for water and a viscous protein solution, which Genentech’s automation specialists considered reasonable.
At Carnegie Mellon University, researchers reported running serial-dilution experiments roughly three times faster after an agent was connected to equipment spread across three computers with incompatible interfaces.
Quantum computing company QuEra, meanwhile, reported that an agent maintained a laser lock for 99.3% of a test period, compared with 58% for its previous scripted system. These are partner-reported demonstrations rather than general benchmarks, but they show what a common interface between AI agents and physical hardware might enable.
MHS is not itself an AI scientist. It is connective tissue: a way for an agent to acquire eyes, hands and instrument readings inside a physical laboratory.
Laboratory automation is not scientific autonomy
Laboratories have used robots for decades. Traditional laboratory automation executes a predefined sequence: transfer this liquid, heat that sample and record the measurement.
Scientific autonomy requires something more difficult. The system must interpret an ambiguous result, decide whether an experiment has failed and choose what should happen next. That demands both reasoning and an understanding of the physical world, precisely where current AI models remain unreliable.
Genentech’s experiment exposed the distinction. When bubbles formed in a viscous solution, Claude initially responded by retrying the operation with different parameters. That agitated the liquid and created still more bubbles. Researchers had to explain that the error came from foam and guide the system towards gentler handling. Once told, Claude retained the lesson for the remainder of the run.
It is an instructive failure. An AI model may understand laboratory instructions without possessing the tacit knowledge of someone who has spent years watching real liquids misbehave.
FutureHouse, a non-profit building AI systems for science, sets a demanding definition of an AI scientist: a system capable of independently generating hypotheses, designing and running experiments, interpreting results and revising its ideas.
By FutureHouse’s own assessment, no existing system, including its own, currently fully meets that standard.
A-Lab closes the experimental loop
One of the clearest demonstrations of closed-loop experimentation came from A-Lab at Lawrence Berkeley National Laboratory.
The self-driving laboratory combined computational predictions, information extracted from scientific literature, machine learning, robotic synthesis and X-ray diffraction. It selected experimental recipes, made materials, analysed the results and used that information to choose subsequent experiments.
Over 17 days, A-Lab performed 353 experiments and produced 36 of 57 targeted inorganic compounds, a success rate of 63%, with minimal human intervention.
Early reports frequently quoted 41 successful materials from 58 targets. A January 2026 correction to the Nature paper revised the peer-reviewed record to the lower totals.
That is an important achievement, but its boundaries matter. Researchers defined the field of possible targets, constructed the laboratory and determined the system’s objectives. A-Lab demonstrated autonomy inside a carefully engineered scientific environment, not an artificial researcher capable of wandering into any laboratory and inventing its own research programme.
Robin shows how humans and AI may divide the work
FutureHouse’s Robin system offers another model for AI-driven scientific discovery. Robin connects specialised agents for literature searches, hypothesis generation and data analysis.
In a study published in Nature, it identified ripasudil (an existing drug) as a potential candidate for dry age-related macular degeneration.
Robin proposed experiments, interpreted the resulting data and updated its hypotheses. Human researchers, however, carried out the wet-lab work. The paper therefore describes the process as “semi-autonomous” rather than fully autonomous.
Its cell-based results were promising, but in-vivo studies would still be required before drawing conclusions about ripasudil’s therapeutic value.
FutureHouse says the project moved from its initial question to a scientific paper in roughly two and a half months. The important point is not that human scientists disappeared. They worked at a different level: performing and supervising physical experiments while AI agents searched more literature and analysed more possible explanations than a small team could manage alone.
The commercial race to build AI science factories
A commercial technology stack is now forming around autonomous scientific discovery.
Anthropic is working on the interface between AI agents and physical machines.
Lila Sciences is building integrated “AI Science Factories” that combine scientific AI models with automated laboratories, initially targeting areas including therapeutics, materials, chemicals and energy. The company emerged from stealth in 2025 and has raised $550 million, including a $350 million Series A.
Lila has reported testing more than 950,000 messenger RNA sequences and using the resulting data to develop unusually stable mRNA. It also says a three-person team used the platform to produce an experimental in-vivo CAR-T therapy that outperformed a benchmark in non-human primates within six months.
Periodic Labs is also building physical laboratories in which AI systems can run experiments and learn from the results, particularly in materials science. Its backers include Andreessen Horowitz, NVentures, Jeff Bezos, Eric Schmidt and Google DeepMind chief scientist Jeff Dean. Public evidence of independently reproduced discoveries remains limited, reflecting how early the company is.
UK-based CuspAI is taking a materials-focused route, using generative AI models to search for compounds with desirable properties. Its collaborations include Hyundai Motor Group and a wider AI Materials Foundry involving industrial and research partners.
As with AI-generated drug molecules, however, proposing a new material is only the beginning. Researchers must still synthesise it and demonstrate that it works outside the model.
Faster scientific discovery still needs slower safeguards
The danger is that autonomous laboratories could manufacture confidence as quickly as they manufacture samples.
An AI discovery system needs a traceable record of every decision, parameter change, instrument reading and discarded result. Negative outcomes must be preserved rather than quietly filtered out. Promising findings must be reproduced, preferably on different equipment and by independent researchers.
Recent scientific reviews of self-driving laboratories identify scalability, generalisability and complete experimental provenance as central requirements for the field.
Medicine presents an additional constraint. AI can accelerate the search for a drug candidate, but it cannot remove the need for toxicity studies, clinical trials and long-term observation inside the human body. Faster discovery does not automatically mean faster regulatory approval — and certainly does not guarantee a successful medicine.
The AI scientist, then, is not about to replace the human scientist. It is beginning to leave the computer under supervision, entering laboratories built around tightly specified tasks, instrument safeguards and human-defined goals.
If systems such as MHS can make laboratory equipment easier for agents to control, the design–build–test–learn cycle could run far more quickly and, in some settings, through the night. The decisive test will be whether those cycles produce knowledge that survives replication.
A machine that can perform a thousand experiments while its human colleagues sleep would be extraordinarily useful. A machine that can fail a thousand times overnight without anyone understanding why would merely accelerate confusion.
The winners in autonomous science will not be the companies whose laboratories run fastest. They will be the ones whose results can still be trusted when the robots stop.
Liked this article? You can support our independent journalism via our page on Buy Me a Coffee. It helps keep MoveTheNeedle.news focused on depth, not clicks.