Health & Bioscience

Research in health and biomedical sciences has a unique potential to improve peoples’ lives, and includes work ranging from basic science that aims to understand biology, to diagnosing individuals’ diseases, to epidemiological studies of whole populations. We recognize that our strengths in machine learning, large-scale computing, and human-computer interaction can help accelerate the progress of research in this space. By collaborating with world-class institutions and researchers and engaging in both early-stage research and late-stage work, we hope to help people live healthier, longer, and more productive lives.

Recent Publications

Wearable Foundation Models Should Go Beyond Static Encoders
Tong Xia
Dong Ma
Ting Dang
Cecilia Mascolo
Hyungjun Yoon
Sung-Ju Lee
Yu Yvonne Wu
Jing Han
Evelyn Zhang
Qiang Yang
arXiv (2026) (to appear)
Preview abstract Wearable foundation models (WFMs), trained on large volumes of data collected by affordable, always-on devices, have demonstrated strong performance on short-term, well-defined health monitoring tasks, including activity recognition, fitness tracking, and cardiovascular signal assessment. However, most existing WFMs primarily map short temporal windows to predefined labels via static encoders, emphasizing retrospective prediction rather than reasoning over evolving personal history, context, and future risk trajectories. As a result, they are poorly suited for modeling chronic, progressive, or episodic health conditions that unfold over weeks, months or years. Hence, we argue that WFMs must move beyond static encoders and be explicitly designed for longitudinal, anticipatory health reasoning. We identify three foundational shifts required to enable this transition: (1) Structurally rich data, which goes beyond isolated datasets or outcome-conditioned collection to integrated multimodal, long-term personal trajectories, and contextual metadata, ideally supported by open and interoperable data ecosystems; (2) Longitudinal-aware multimodal modeling, which prioritizes long-context inference, temporal abstraction, and personalization over cross-sectional or population-level prediction; and (3) Agentic inference systems, which move beyond static prediction to support planning, decision-making, and clinically grounded intervention under uncertainty. Together, these shifts reframe wearable health monitoring from retrospective signal interpretation toward continuous, anticipatory, and human-aligned health support. View details
Toward a test of medical AI superintelligence
Ethan Goh
David Wu
Chase Walton
Liam McCoy
Anastasia Perez
Laura Wegner
Fateme Nateghi Haredasht
Luyang Luo
Kathleen Lacar
Thomas Buckley
Austin Schoeffler
Peter Brodeur
Kameron C. Black
John Havlik
John Rumsfeld
Daniel Lopez-martinez
Paxton Maeder-York
Karan Singhal
David Gunning
Bon Ku
Haider Warraich
Shantanu Nundy
Vishnu Ravi
Arnold Milstein
Jason Hom
Kevin Schulman
Pranav Rajpurkar
Arjun Manrai
Robert Wachter, MD
Eric Topol
Eric horvitz
Adam Rodman
Jonathan Chen
Nature Medicine (2026)
Preview abstract Researchers urgently need a rigorous, task-based framework to define and measure medical AI ‘superintelligence’, because existing benchmarks are misleading and insufficient. View details
Towards expert-level medical AI for real-time video consultations
Mahvish Nagda
Jihyeon Lee
Matthew Thompson
CJ Park
Tim Strother
Roma Ruparel
Teya Bergamaschi
Suhana Bedi
Meet Shah
Pavel Dubov
Toshiyuki Fukuzawa
Sam Schmidgall
Craig Schiff
Joseph Xu
Aliya Rysbek
Yana Lunts
Jan Freyberg
Rebecca Hemenway
David Racz
Carey Radebaugh
Joelle Barral
Kavi Goel
Kat Chou
James Manyika
Gregory Wayne
Yun Liu
Ethan Goh
Christina Chen
Ryutaro Tanno
arXiv (2026)
Preview abstract Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated feasibility, but not reached clinician-level performance. Here, we provide the first demonstration of expert-level AI in real-time clinical video consultations using AMIE (Articulate Medical Intelligence Explorer) in a video configuration. AMIE (Video) is a Gemini-based multi-agent system integrating low-latency dialogue, clinical reasoning, and real-time audio-visual perception. To guide development, we established a taxonomy and automated evaluations for clinical audio-visual cues in telehealth settings. In a randomized Objective Structured Clinical Examination (OSCE) study with 30 primary care physicians (PCPs), 15 patient actors and 100 clinical scenarios, we compared AMIE (Video), its text-only counterpart AMIE (Text), and PCPs consulting via video. Clinical evaluators rated AMIE (Video) on par or better than PCPs in history-taking, diagnosis, management, and physical observation and examination. Patient actors preferred AMIE's approach to assessing and explaining conditions, while PCPs were preferred for rapport and partnership building. In modality ablation, patient actors preferred AMIE (Video)'s interface over text chat for communicative effectiveness, convenience, and feeling understood. Limitations remain in fine anatomical precision, subtle affective nuances, and high-frequency movements. While further research is needed before real-world translation, these results mark an important milestone toward AI systems capable of augmenting care across the sensory complexity of clinical practice. View details
Nine changes needed to deliver a radical transformation in biodiversity measurement
Neil Burgess
Andy Purvis
Scott J. Goetz
William Sutherland
Anil Madhavapeddy,
Tanya Birch
PNAS Perspective (2026)
Preview abstract Biodiversity is declining in many parts of the world. The measurement and monitoring of biological diversity are fundamental to the assessment of the causes and consequences of environmental changes, identification of key areas for the protection of biodiversity or ecosystem services, determining the effectiveness of actions, and the creation of decision-support tools critical to the maintenance of a sustainable planet. The measurement of biodiversity is rapidly changing due to advances in citizen science, image recognition, acoustic monitoring, environmental DNA, genomics, remote sensing and artificial intelligence. In this perspective, we outline the exciting opportunities that these developments offer, but also consider the challenges, especially the potential poisoning of data by AI, lack of standardisation across methods, coverage gaps in data, concerns over losing databases, and undervaluing of on-the-ground expertise and data-generation. Our key recommendations are (1) ensure new technologies are calibrated with existing data; (2) use emerging technologies to fill data gaps; (3) create living databases of trusted information to increase reliability of data and reduce the risk of poisoning by false - or AI hallucinated - information; (4) ensure data generation is valued; (5) ensure the respect and incorporation of Indigenous Knowledge; (6) increase in-country capacity in the tropics; and (7) increase the resilience of global datasets to technical and societal change. Radical new collaborations are needed between computer scientists, engineers, molecular biologists, data scientists, field ecologists, citizen scientists, Indigenous peoples, and local communities to create the rigorous, resilient, accessible biodiversity information systems required to underpin policies and practices that ensure the maintenance and restoration of ecological systems. View details
SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment
Joe Breda
Fadi Yousif
Beszel Hawkins
Marinela Cotoi
Miao Liu
Ray Luo
Sam Schmidgall
Girish Narayanswamy
Samuel Solomon
Max Xu
Longfei Shangguan
Bhavna Daryani
Buddy Herkenham
Cara Tan
Mark Malhotra
Shwetak Patel
Zach Wasson
Dimitrios Antos
Bob Lou
Matthew Thompson
Jonathan Richina
Anupam Pathak
Nichole Young-Lin
Jake Sunshine
Daniel McDuff
Arxiv preprint, 2605.040 (2026) (to appear)
Preview abstract Language models excel at diagnostic assessments on curated medical case-studies and vignettes, performing on par with, or better than, clinical professionals. However, existing studies focus on complex scenarios with rich context making it difficult to draw conclusions about how these systems perform for patients reporting symptoms in everyday life. We deployed SymptomAI, a set of conversational AI agents for end-to-end patient interviewing and differential diagnosis (DDx), via the Fitbit app in a study that randomized participants (N=13,917) to interact with five AI agents. This corpus captures diverse communication and a realistic distribution of illnesses from a real world population. A subset of 1,228 participants reported a clinician-provided diagnosis, and 517 of these were further evaluated by a panel of clinicians during over 250 hours of annotation. SymptomAI DDx were significantly more accurate (OR = 2.56, p < 0.001) than those from independent clinicians given the same dialogue in a blinded randomized comparison. Moreover, agentic strategies which conduct a dedicated symptom interview that elicit additional symptom information before providing a diagnosis, perform substantially better than baseline, user-guided conversations (p < 0.001). An auxiliary analysis on 1,509 conversations from a general US population panel validated that these results generalize beyond wearable device users. We used SymptomAI diagnoses as labels for all 13,917 participants to analyze over 500,000 days of wearable metrics across nearly 400 unique conditions. We identified strong associations between acute infections and physiological shifts (e.g., OR > 7 for influenza). While limited by self-reported ground truth, these results demonstrate the benefits of a dedicated and complete symptom interview compared to a user-guided symptom discussion, which is the default of most consumer LLMs. View details
Large-scale, interpretable gene regulatory network inference through biologically informed matrix factorization
Soel Micheletti
Viola Fanfani
Julia Vogt
John Quackenbush
Jonas Fischer
Alexander Marx
Panagiotis Mandros
bioRxiv (2026)
Preview abstract Gene regulatory networks (GRNs) provide a mechanistic framework for understand- ing how transcription factors coordinate gene expression to establish cellular identity and phenotype. Methods that integrate gene expression with motif-derived regulatory priors and other sources of biological information have substantially advanced gene regulatory network inference by reconstructing condition-specific regulatory architecture. These approaches estimate the evidence supporting regulatory interactions and have proven remarkably successful in a wide range of biological applications. A complementary view of regulatory networks, however, seeks to estimate the effect of those interactions on gene expression itself, providing a framework in which regulatory edges can be interpreted as activating or inhibitory influences on transcription. We developed Giraffe, a biologically informed matrix factorization framework that jointly estimates transcription factor activities and gene regulatory networks by integrating gene expression, motif-based regulatory priors, and transcription factor protein-protein interactions. Giraffe estimates signed partial regulatory effects whose magnitude and sign can be interpreted as the strength and direction of transcriptional regulation. Building directly on the biological framework established by methods such as PANDA, Giraffe provides a complementary representation of gene regulatory networks that emphasizes mechanistic interpretation while remaining scalable, flexible, and computationally efficient. Across synthetic benchmarks, six human tissues, yeast transcription factor perturbation experiments, and liver hepatocellular carcinoma, Giraffe accurately recon- structs regulatory interactions while distinguishing activating from inhibitory regulation with high accuracy. The inferred networks recover known features of tissue-specific regulation, correctly classify regulatory effects in transcription factor perturbation experiments, and identify biologically coherent changes in regulatory programs associated with liver cancer. Together, these results demonstrate that estimating the direction of transcriptional regulation provides a complementary perspective on gene regulatory networks that facilitates biological interpretation and hypothesis generation. View details
×