Scientific discovery is driven by scientists generating hypotheses for complex problems that undergo rigorous experimental validation. To augment this process, we introduce Co-Scientist, a multi-agent artificial intelligence (AI) system built on Gemini for structured scientific thinking and hypothesis generation. Co-Scientist aims to help scientists discover new original knowledge. Conditioned on their research objectives and previous scientific evidence, it formulates demonstrably novel research hypotheses for experimental verification. The system’s design involves agents continuously generating, critiquing and refining hypotheses accelerated by scaling test-time compute. Key contributions include (1) a multi-agent architecture with an asynchronous task execution framework for flexible compute scaling, and (2) a tournament evolution process for self-improving hypotheses generation. Automated evaluations show continued benefits of test-time compute scaling, improving hypothesis quality over time. Although this is a general-purpose system, we focus the validation in three biomedical applications: drug repurposing; novel-target discovery1; and explaining mechanisms of antimicrobial resistance2. Specifically, Co-Scientist helped to identify new drug-repurposing candidates and synergistic combination therapies for acute myeloid leukaemia that were validated through in vitro experiments. These real-world validations demonstrate the potential of Co-Scientist to accelerate scientific discovery and usher in an era of AI-empowered scientists
Extract authors, key findings, references, and an executive summary using AI.
Scientific discovery requires generating and experimentally validating complex hypotheses, a process increasingly challenged by the sheer volume of literature and required multidisciplinary depth. To address this, the authors introduce Co-Scientist, a multi-agent artificial intelligence system built on Gemini that functions as a structured scientific thinking engine. Operating through an asynchronous task execution framework and specialized worker agents—including Generation, Reflection, Ranking, Evolution, Proximity, and Meta-review agents—Co-Scientist scales test-time compute to iteratively generate, critique, and refine research hypotheses and proposals. Automated evaluations across diverse scientific research goals demonstrate that scaling test-time compute leads to continuous improvements in hypothesis quality, as measured by progressive increases in Elo ratings. Furthermore, on expert-curated biomedical benchmarks, Co-Scientist outperformed other state-of-the-art language and reasoning models, and blinded human expert evaluations confirmed its superior ratings in both novelty and impact. The system's practical utility was validated through end-to-end wet-laboratory experiments across three distinct biomedical domains. In oncology, Co-Scientist identified single-agent drug repurposing candidates (such as binimetinib and the novel IRE1alpha inhibitor KIRA6) and synergistic multi-drug combinations that inhibited acute myeloid leukaemia cell viability. In hepatology, it discovered novel epigenetic targets and anti-fibrotic compounds (including FDA-approved vorinostat) using human hepatic organoids. In microbiology, it independently recapitulated an unpublished mechanism of bacterial mobile genetic element transfer. While subject to limitations such as reliance on open-access literature and the potential for model hallucinations, Co-Scientist represents a significant step toward AI-assisted scientific augmentation. By combining rigorous automated reasoning, expert-in-the-loop collaboration, and wet-laboratory validation, the system demonstrates the potential to accelerate discovery and help researchers resolve grand challenges in biomedicine and science.
Scientific discovery is driven by scientists generating hypotheses for complex problems that undergo rigorous experimental validation. To augment this process, we introduce Co-Scientist, a multi-agent artificial intelligence (AI) system built on Gemini for structured scientific thinking and hypothesis generation. Co-Scientist aims to help scientists discover new original knowledge. Conditioned on their research objectives and previous scientific evidence, it formulates demonstrably novel research hypotheses for experimental verification. The system’s design involves agents continuously generating, critiquing and refining hypotheses accelerated by scaling test-time compute. Key contributions include (1) a multi-agent architecture with an asynchronous task execution framework for flexible compute scaling, and (2) a tournament evolution process for self-improving hypotheses generation. Automated evaluations show continued benefits of test-time compute scaling, improving hypothesis quality over time. Although this is a general-purpose system, we focus the validation in three biomedical applications: drug repurposing; novel-target discovery; and explaining mechanisms of antimicrobial resistance. Specifically, Co-Scientist helped to identify new drug-repurposing candidates and synergistic combination therapies for acute myeloid leukaemia that were validated through in vitro experiments. These real-world validations demonstrate the potential of Co-Scientist to accelerate scientific discovery and usher in an era of AI-empowered scientists.
1.Co-Scientist is a multi-agent artificial intelligence system built on Gemini designed to act as a structured scientific thinking engine for automated hypothesis generation and literature-based reasoning.
2.The system utilizes an asynchronous task execution framework and specialized agents including Generation, Reflection, Ranking, Evolution, Proximity, and Meta-review agents.
3.Automated evaluations across diverse scientific research goals demonstrate that scaling test-time compute continuously improves hypothesis quality over time as measured by Elo ratings.
The discussion highlights Co-Scientist as a structured scientific thinking engine built on a multi-agent Gemini architecture that successfully mirrors the scientific method through a 'generate, debate, evolve' paradigm. Practical applications were demonstrated across acute myeloid leukaemia drug repurposing, liver fibrosis target discovery, and microbial gene transfer mechanisms. Limitations include reliance on open-access literature (omitting paywalled or negative results), potential propagation of erroneous findings, and inherent model limitations like hallucinations. Future directions involve enhancing literature search and factual verification capabilities, integrating direct reasoning over public databases and multimodal data, incorporating reinforcement learning from human and experimental feedback, expanding evaluations across broader scientific disciplines, and eventually integrating with laboratory automation platforms for closed-loop autonomous discovery.