
🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
About this episode
From the show’s notesCan a model predict how a cell responds to a genetic perturbation it has never seen? Xaira Therapeutics' new virtual-cell model, X-Cell, is a 4.9-billion-parameter diffusion language model trained on X-Atlas/Pisces — the largest genome-wide CRISPRi Perturb-seq dataset ever built, spanning 25.6 million single cells across 16 biological contexts. Bo Wang (Chief AI Scientist) and Ci Chu (Chief Discovery Officer) explain why observational atlases can describe biology but can't predict what happens when you intervene, why they abandoned autoregression for a diffusion "editing" approach, and how a model trained on immortalized cell lines predicted perturbation responses in primary T cells from real donors. Plus the counterintuitive scaling result: X-Cell scales like an LLM on training loss, but generalization is bottlenecked by data diversity, not compute — a conversation about why, in AI for science, the hard part isn't the model, it's the data.
Bios
Read the show’s notes in full
Bo Wang is Chief AI Scientist at Xaira Therapeutics and a co-senior author of X-Cell. He is also an Associate Professor at the University of Toronto, a CIFAR AI Chair at the Vector Institute, and Chief AI Scientist at University Health Network, where his lab pioneered scGPT — one of the first single-cell foundation models — and BioReason. His work centers on foundation models that learn the underlying biology of the cell; X-Cell is initialized from scGPT's own encoder weights.
Ci Chu is Chief Discovery Officer at Xaira Therapeutics and a co-senior author of X-Cell. He leads the high-throughput biology behind Xaira's data-generation engine — including the industrialized Perturb-seq platform that produced X-Atlas/Pisces — and previously led functional genomics and high-content phenotyping work at insitro. His throughline is pairing large-scale, high-quality experimental data with the most capable models.
Timestamps
"This is the first time that someone can put together not just one, but seven genome-wide Perturb-seq campaigns together." (0:26) "We are nowhere near the same kind of massive, high-quality data [in cell modeling] that we have in protein design... it's a data limitation issue." (0:9:46) "Instead of typing, think of diffusion language models as editing: you iteratively generate a sentence from a very vague, rough start and then refine it." (0:43:11) "I tell my students: in the era of agentic AI, the role of the scientist is shifting from coding to debugging the AI's outputs." (1:23:07) "You can't have the cake and eat it—currently, to sequence a cell, you have to kill it. The real 'magic wand' for virtual cells would be temporal sequencing of live cells." (1:28:43)





