MIA: Shreya Johri, Evaluating AI agents in biological discovery; primer by Maha Shady
Broad Institute
0:00 / 0:00
MIA: Shreya Johri, Evaluating AI agents in biological discovery; primer by Maha Shady
497 просмотров · 2 месяца назад
Broad Institute
35,3 тыс. подписчиков
497 просмотров · 2 месяца назад
Models, Inference, and Algorithms | April 15, 2026
Broad Institute of MIT and Harvard
Seminar: Evaluating the autonomous and copilot limitations of AI agents for biological discovery
Shreya Johri
Graduate Student, Harvard University
Abstract:
Recent advances in large language models (LLMs) have improved their ability to execute structured analytical workflows, including standard bioinformatic pipelines. However, computational biology rarely consists of deterministic pipeline execution alone. Biological datasets are heterogeneous and noisy, and meaningful discovery often requires open-ended hypothesis generation and iterative reasoning over multimodal evidence. The extent to which emerging agentic AI systems can support this mode of scientific discovery remains poorly characterized. Here, we systematically evaluate the capabilities and limitations of agentic AI for biological discovery using multimodal oncology datasets spanning 15 cancer types. We benchmark 10 analysis tasks designed to vary in biological reasoning complexity, including replication of canonical workflows, tumor-program characterization, tumor-microenvironment analysis, and immune-cell discovery tasks. We also benchmark autonomous and human-copilot agent configurations. Our results delineate the current boundaries of agentic AI in computational biology and provide a framework for evaluating AI systems designed to support scientific discovery.
Primer: AI agents in biomedical research
Maha Shady
Graduate Student, Harvard University
Abstract:
LLM-based agents are increasingly used in biomedical research pipelines, for literature synthesis, data analysis, hypothesis generation, and clinical decision support. This talk provides an overview of how these systems work and where they break. We cover the core architectural components of single and multi-agent systems, as well as current evaluation benchmarks and failure mechanisms specific to biomedical applications. The talk may serve as a practical guide for developing and evaluating these systems, with consideration of failure modes most consequential in biomedical settings.
About MIA:
The Models, Inference & Algorithms (MIA) Initiative at the Broad Institute supports learning and collaboration across the interface of biology and medicine with mathematics, statistics, machine learning, and computer science. Our weekly meetings are open and pedagogical, emphasizing lucid exposition of computational ideas over rapid-fire communication of results.
MIA is hosted by the Eric and Wendy Schmidt Center at the Broad Institute.
Relevant Links:
MIA Website: https://www.broadinstitute.org/mia
MIA YouTube Playlist: https://broad.io/MIAPlaylist
Copyright Broad Institute, 2026. All rights reserved.