Since its launch in 2021, the Cambridge Centre for AI in Medicine (CCAIM) has worked to advance the use of artificial intelligence and machine learning across healthcare, with the ambition of transforming prevention, diagnosis, treatment and clinical development.
Clinical trials remain one of the greatest bottlenecks in modern medicine. They are often slow, expensive and operationally complex, making it difficult to identify early which patients are most likely to benefit, which may be harmed, and which trial designs are most likely to succeed. Addressing these challenges has been a central part of CCAIM’s mission from the very beginning and remains one of the Centre’s most important research priorities.
CCAIM’s vision has always been that AI should not simply be applied to existing clinical-trial workflows. Instead, advances in machine learning provide an opportunity to fundamentally rethink clinical development itself. By enabling more adaptive, personalised and data-driven approaches, AI has the potential to improve how trials are designed, how decisions are made and how new knowledge is generated, ultimately supporting researchers, clinicians, developers and healthcare systems in delivering better outcomes.
Over the past several years and building on research that predates the Centre itself, CCAIM and its affiliated researchers have advanced this vision through a number of interconnected areas, including causal AI, digital twins, synthetic data, adaptive clinical trials, AI for pharmacology, treatment-effect estimation and more recently, agentic AI systems for clinical development. Together, these advances form the scientific foundations of CCAIM’s approach to AI-enabled clinical trials.
This page brings together publications, articles and resources that reflect this research agenda and its contribution to CCAIM’s broader mission of advancing AI-enabled medicine and healthcare.
Causal AI
At the centre of this vision is causal AI. Clinical trials are not only about predicting outcomes, but they are also about making better decisions. What would happen if we changed the dose, the eligibility criteria, the comparator, the endpoint, the site mix, or the treatment strategy? Which patients are likely to benefit, and under what assumptions? What evidence is needed before a decision can be trusted?
The van der Schaar Lab has pioneered methods for causal effect inference and individualized treatment-effect estimation that help answer these questions. Our work enables researchers to better understand how treatments affect different patients, predict the consequences of interventions and generate evidence that can support more informed clinical and regulatory decisions.
- Causal Effect Inference (Research Pillar)
- Nonparametric estimation of heterogeneous treatment effects: From theory to learning algorithms (AISTATS 2021)
- Using Machine Learning to Individualize Treatment Effect Estimation: Challenges and Opportunities (Clinical Pharmacology & Therapeutics)
- Causal machine learning for predicting treatment outcomes (Nature Medicine)
- Bayesian Inference of Individualized Treatment Effects using Multi-task Gaussian Processes (NeurIPS 2017)
- From Real-World Patient Data to Individualized Treatment Effects Using Machine Learning: Current and Future Methods to Address Underlying Challenges (Clinical Pharmacology & Therapeutics, 2020)
- Gradient-Based Causal Tree Ensembles: A Backbone Architecture for Heterogeneous Treatment Effects (ICML 2026)
- Identifiable Nonlinear Differentiable Causal Discovery via Independence and Adaptive Group Sparsity (ICML 2026)
- Overlap-weighted orthogonal meta-learner for treatment effect estimation over time (ICLR 2026)
- Treatment Effect Estimation for Optimal Decision-Making (NeurIPS 2025)
- AutoCATE: End-to-End, Automated Treatment Effect Estimation (ICML 2025)
- Active Feature Acquisition for Personalised Treatment Assignment (AISTATS 2025)
- Quantifying Aleatoric Uncertainty of the Treatment Effect: A Novel Orthogonal Learner (NeurIPS 2024)
- Causal Deep Learning
- ODE Discovery for Longitudinal Heterogeneous Treatment Effects Inference (ICLR 2024)
Digital Twins
This causal foundation naturally leads to digital twins. In our vision, digital twins are not generic simulators or decorative AI tools. They are dynamic, adaptive, causally grounded models of patients, diseases, populations, trials and healthcare systems. They allow us to ask: what could happen before we run the trial? Which assumptions are fragile? Which patients should be enrolled? How might disease trajectories evolve? How should we adapt if new evidence emerges?
The lab’s work on digital twins provides a route to stress-test trials before launch, support in-silico evidence generation, and support more informed clinical development decisions. Together, these technologies offer a path towards clinical trials that can continuously learn, adapt and improve over time.
- Understanding Digital Twins (Sep 2025)
- Nature Biotech Q&A (Oct 2025) “Applications of digital twins in medicine”
- Inspiration Exchange 40
- Revolutionising Healthcare 38
- Decision-Targeted Digital Twins (DT²) (ICML 2026)
- Continuously Updating Digital Twins using Large Language Models (ICML 2025)
- HDTwin – Automatically Learning Hybrid Digital Twins of Dynamical Systems (NeurIPS 2024)
- SyncTwin: Treatment Effect Estimation with Longitudinal Outcomes (NeurIPS 2021)
- HDTwinGen – Automatically Learning Hybrid Digital Twins of Dynamical Systems (NeurIPS 2024)
Synthetic Data
A third major pillar is synthetic data. Clinical development is constrained by limited access to high-quality, representative, privacy-preserving datasets. Over many years, the van der Schaar Lab has developed methods and tools for generating, evaluating and applying synthetic healthcare data to help address these challenges.
Synthetic data can support privacy-preserving data access, data augmentation, fairness, benchmarking and machine learning model development, while enabling researchers to work with realistic healthcare data in a safe and scalable way. This work includes open-source software such as SynthCity and more recent work on making synthetic data generation accessible to clinicians and healthcare researchers.
- Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes (ICML2024)
- Synthetic data in biomedicine via generative artificial intelligence (Nature Reviews Bioengineering, 2024)
- How Faithful is your Synthetic Data? Sample-level Metrics for Evaluating and Auditing Generative Models (ICML 2022)
- Time-series Generative Adversarial Networks (NeurIPS 2019)
- PATE-GAN: Generating Synthetic Data with Differential Privacy Guarantees (ICLR 2019)
- Synthcity: facilitating innovative use cases of synthetic data in different data modalities (NeurIPS 2023)
- Synthetic data for privacy-preserving clinical risk prediction (Scientific Reports, 2024)
- SynthCraft: An AI partner for synthetic data generation to support data access and augmentation in healthcare(PLOS Digital Health, 2026)
Adaptive Clinical Trials
The lab has also been at the forefront of adaptive clinical trials. Traditional trials often rely on decisions and assumptions made at the outset, with limited ability to respond to new information as it emerges. Adaptive trials offer a different paradigm: learning during the trial while preserving scientific validity, transparency and trust.
Over the years, our research has developed machine learning approaches that can support better patient allocation, cohort enrichment, treatment adaptation, and decision-making under uncertainty. The goal is not uncontrolled flexibility, but the ability to adapt when the evidence supports it, helping clinical trials become more efficient, informative and responsive.
- Revolutionizing Clinical Trials: A Manifesto for AI-Driven Transformation (2025)
- Adaptive Identification of Populations with Treatment Benefit in Clinical Trials: Machine Learning Challenges and Solutions (ICML 2023)
- Adaptive Experiment Design with Synthetic Controls (AISTATS 2024)
- Towards Regulatory-Confirmed Adaptive Clinical Trials: Machine Learning Opportunities and Solutions (AISTATS 2025)
- Machine learning for clinical trials in the era of COVID-19 (Stat Biopharm Res, 2020)
AI for Pharmacology
Another key area of research is AI for pharmacology and pharmacometrics. Clinical development depends on understanding dose, response, toxicity, pharmacokinetics, pharmacodynamics, disease progression and patient heterogeneity. The van der Schaar Lab’s work combines mechanistic modelling with modern machine learning to improve prediction of drug response, enable precision dosing, learn dynamical systems, and connect pharmacological theory with clinical data. By connecting pharmacological theory with clinical data, this work aims to support not only more efficient clinical trials, but also the scientific and biological questions at the heart of drug development.
- Data-Driven Discovery of Dynamical Systems in Pharmacology using Large Language Models (NeurIPS 2024)
- From Real-World Patient Data to Individualized Treatment Effects Using Machine Learning: Current and Future Methods to Address Underlying Challenges (Clinical Pharmacology & Therapeutics, 2020)
- Synthetic Model Combination: A new machine-learning method for pharmacometric model ensembling (CPT Pharmacometrics Syst Pharmacol, 2023)
- The potential and pitfalls of artificial intelligence in clinical pharmacology (CPT Pharmacometrics Syst Pharmacol, 2023)
- Bridging the Worlds of Pharmacometrics and Machine Learning (Clin Pharmacokinet, 2023)
- Integrating Expert ODEs into Neural ODEs: Pharmacology and Disease Progression (NeurIPS 2021)
Agentic AI for clinical trials
Most recently, the lab has been developing agentic AI for clinical trials. Bringing a new treatment to patients requires decisions across many areas, including biology, statistics, operations, regulation, safety, recruitment, and real-world evidence. As the volume and complexity of information continue to grow, there is an increasing interest in how AI systems can help researchers make sense of that evidence and support better decision-making.
In our vision, AI agents can become active reasoning partners: monitoring evidence, identifying fragile assumptions, stress-testing designs, integrating data streams, detecting drift, and helping teams ask better questions before failures occur. In our vision, these agents do not replace clinical, statistical or regulatory expertise. Rather they augment it, helping to create a new layer of clinical-development intelligence.
- White Paper: Clinical Trials as Continuously Learning Systems.
The opportunity now is to bring these advances together. Clinical trials are not transformed by any single technology, but by combining complementary approaches that address different parts of the development process. Causal AI helps answer the right questions, digital twins enable simulation and scenario testing, synthetic data expands access to information while protecting privacy, adaptive methodologies support better decision-making during trials and pharmacological AI connects mechanisms and treatment response, and AI agents helps researchers navigate increasingly complex evidence, assumptions and decisions across the development lifecycle.
Together, these technologies offer a new foundation for transforming clinical trials from isolated, rigid studies into continuously learning evidence systems. Before a trial begins, we can stress-test its design. During the trial, we can monitor assumptions and adapt responsibly. After the trial, we can feed the evidence back into digital twins, causal models, synthetic-data engines and agentic systems that improve the next trial.
The van der Schaar Lab has spent the past decade developing the scientific foundations for this transformation. The challenge now is translation: building validated, auditable, regulator-aligned tools that can be used in real clinical-development settings. If successful, this could make trials faster, safer, more predictive, more efficient, and ultimately better able to bring the right treatments to the right patients sooner.
This page brings together a selection of the publications, articles, software tools and resources that contribute to this vision of AI-enabled clinical development.
