Why Jingtian Zhou is building high-resolution maps of the 3D genome
Each cell in our body inherits the same DNA, yet becomes one of hundreds of cell types. We make immune cells that hunt down viruses, brain cells that pattern our memories, and pacemaker cells that generate electricity to drive our heartbeat. To understand how cells can have the exact same genetic code, but do completely different things, Arc Science Fellow Jingtian Zhou has built a powerful 3D genome toolkit, developing both the wet-lab assays that measure how DNA physically folds inside the nuclei of single cells and the algorithms that make that data interpretable.
In recent years, he’s helped generate a single-cell atlas of 3D genome organization and DNA methylation across the human body in 86,689 nuclei from 16 tissues, resolving 35 major cell types and 206 subtypes, out in Science; developed a spatial multiomics technology called Spatial Hi-C-RNA that simultaneously maps genome-wide 3D chromatin architecture and gene expression directly in intact tissue at near-single-cell resolution, out in Cell; and served as the primary computational expert for a publication in Cancer Cell that revealed how chronic stress recruits bone marrow cells to accelerate brain tumor progression.
In this discussion, Zhou shares his reasons for studying spatial technologies, what it takes to faithfully scale single-cell datasets from hundreds to millions of cells, and how to bridge wet- and dry-lab work by grounding predictive models in real biological data.
How did you get your start in 3D genomics?
In my undergraduate studies, I worked on a project reconstructing 3D genome structures using data from Hi-C, a sequencing technology to measure genome architecture in 3D. Before that, my impression of 3D genome structure was based on the globular images from textbooks. I thought it was unstable, chaotic, and extremely stochastic, kind of an “instant noodle-like” state.
At the time, Hi-C was quite new; only a few labs were working on those types of experiments. When reading those works, I started to realize that the genome has a highly stable, organized architecture within its apparent disorder. Features like topologically associating domains (TADs), chromatin compartments, and chromatin loops are conserved within cell types and directly correlate with gene regulation.
An interesting finding from those early studies is that 3D genome structures are very conserved across cell types, and even across species. This is contradictory to what we see from other epigenome signatures, whose dynamics are associated with cell-type specific gene expression. But a few pieces of evidence suggested that very different cell types may still harbor specific chromatin interactions regulating cell-type-specific gene expression. We wanted to zoom past the global similarities and see if specific genomic pixels or loops showed cell-type specificity, which drove us to develop single-cell Hi-C.
You’ve been involved in several major studies of single-cell and spatial genomics, including Spatial Hi-C-RNA and a whole-body snm3C-seq atlas. What motivates you to build these resources?
When I entered the field, Hi-C was an extremely expensive and restrictive assay. The first time I formulated an algorithm to analyze single-cell 3D genomic data, we had data for less than ten cells. So, we decided to build the wet lab technology to generate our own data, which is how snm3C-seq was born.
Biologically, the core question we wanted to resolve was how to physically link distal regulatory elements, like enhancers, to the specific target genes they regulate. Methods like open chromatin or DNA methylation assays can identify regulatory elements, but they cannot show you which genes those elements are actually interacting with.
The question I keep coming back to is: how much does genome conformation vary between cell states? Answering this can help model the relationship between regulators and target genes. Both the human body atlas (snm3C-seq) and the spatial multi-omics paper (Spatial Hi-C-RNA) share a deep focus on 3D genome architecture across cell states. While snm3C-seq is a comprehensive reference atlas of single-cell 3D chromatin folding and methylation across the whole body, Spatial Hi-C-RNA is a spatially resolved multi-omics technology that lets us map these features within their intact tissue context.
Spatial-Hi-C-RNA measures chromatin folding and transcription in the same tissue sections. What does chromatin tell you that the RNA does not?
Chromatin is a good indicator of genome rearrangement–insertion, deletion, reversion, duplication, and so on–which is quite common in cancer cells. With chromatin data, we have higher power to detect those structural variants, which is how we found copy number changes at several oncogenes and tumor suppressors. Given that those variants accumulate but are almost never erased, we can use that information to track cell lineages and tumor subclones, which links genetics with epigenetics in the tumor.
And, in previous multiomic studies with different technologies, we and others have suggested that the time course of dynamics differs by modality, or more generally, during cell state switching. Our spatial Hi-C-RNA data isn’t yet high-resolution enough to model these differences, so we’re working on optimizing our protocol further to push our understanding forward.
What did you find when you expanded your single-cell datasets from the brain to the rest of the body?
In the brain, different layers of the epigenome are highly correlated. The highest gene expression perfectly corresponds to stronger chromatin loops, open chromatin, and active DNA methylation. Almost every mapping algorithm works beautifully in the brain because everything is tightly aligned.
But when we moved to non-brain organs, everything became much more subtle and localized. What surprised me most was discovering modality inconsistency, where different epigenetic layers disagreed on cell-type boundaries.
For example, in skeletal muscle, slow-twitch and fast-twitch myofibers are clearly classified by both DNA methylation and 3D chromatin conformation. But we found a population of intermediate cells with characteristics of both muscle stem cells and mature muscle fibers. They had lost their stem-cell chromatin loops but still retained the methylation signature of stem cells. We think that these discordant cells are caught in active cell-state transitions that occur at different molecular velocities. The 3D physical architecture of the genome remodels first, while the DNA methylation signatures lag behind.
A core focus of your research is transitioning from descriptive maps to predictive sequence-to-function models. How do your single-cell atlases inform how you design and train computational models?
With our large datasets, we mapped the 3D genome and epigenome across hundreds of different cell types, developmental stages, and tissues. Now, we want to build a predictive computational model to interrogate any genomic mutation and see its cell-type-specific functional impact. We realized it would be far more powerful to save this collective regulatory context inside a machine learning model that you can query, rather than just viewing static profiles in a web browser.
We’re currently developing a model called Bolero that does exactly this. By bridging single-cell atlas embeddings with a raw DNA sequence model, we can computationally predict variant effects, chromatin accessibility, and gene expression across hundreds of cell types in different contexts—across brain regions, developmental stages, and most interestingly, upon neural activation, without training on neural activity data. We can even predict cell states that haven't been profiled in the wet lab yet.
What are the most valuable causal experiments for validating model predictions based on correlative atlas data?
In our lab, we perform forward mapping from the DNA sequence to function for our sequence-to-function models. The most direct way to introduce causal validation into these models is by incorporating CRISPR screening data. We are currently working with colleagues to integrate functional genomics screening data into our model architectures.
When it comes to our clinical disease predictions, establishing a causal “gold standard” is much more challenging. We cannot easily run global causal screens for cell types associated with a complex disease. But we can perform targeted screens at the gene or genetic variant level and measure their direct causal impact across different cellular subtypes. This targeted, subtype-specific screening is where we can bridge correlative predictions with causal mechanisms.
You apply your computational expertise to basic biology and clinical questions. Where has your work taken you this year?
Recently, we published a paper on stress-induced glioma growth. Our collaborators had a specific biological hypothesis about how psychological stress accelerates tumor growth, and my role was to lead the single-cell RNA-seq computational efforts to identify the pro-tumorigenic cell types involved. We ended up finding that among immune cells, a certain type of stress-associated macrophages was mediating brain-bone marrow crosstalk and shielding the tumor so it could grow fast under chronic stress. More recently, with collaborators at UCLA and Salk Institute, we’ve started to expand our analyses of glioblastoma and prostate cancer to the 3D genome, mapping the chromatin regulators shaping the heterogeneous cell state transitions in those tumors.
While basic single-cell RNA-seq has become widely accessible and can be handled by standard coding pipelines, resolving the physical, 3D genome structure of primary tissues at single-cell resolution requires highly specialized computational tools. Because I can generate and implement these tools, I get to work on a wide range of biological questions. It keeps things interesting.
Your work sits at a unique intersection of coding and cell biology. How do you view the relationship between wet-lab experimentation and dry-lab computational prediction?
I fell in love with coding because of its consistency. In the wet lab, experiments can fail for reasons you can't easily see. But computationally, when you run the same code twice, it always gives you the exact same result. Computational predictions also let us test millions of hypotheses in a single run, so we get experimental results faster.
But computational predictions always need to land onto real biology to make an impact. We also need rigorously designed wet-lab data to validate these predictions, which is why my lab focuses on developing both the experimental technologies to generate high-quality data and the computational frameworks required to analyze it. We developed the first comprehensive, scalable algorithm to reconstruct and analyze single-cell 3D chromatin conformation, and we continue to scale that algorithm as our datasets grow, from just 8 single cells early in my PhD to millions of cells today.
Zhou, J., Wu, Y., Liu, H., Tian, W., Castanon, R.G., Bartlett, A., Zhang, Z., Yao, G., Shi, D., Clock, B., Marcotte, S., Nery, J.R., Liem, M., Claffey, N., Boggeman, L., Barragan, C., Arrojo e Drigo, R., Weimer, A.K., Shi, M., Cooper-Knock, J., Zhang, S., Snyder, M.P., Preissl, S., Ren, B., O'Connor, C., Chen, S., Luo, C., Dixon, J.R., & Ecker, J.R. (2026). Human body single-cell atlas of three-dimensional genome organization and DNA methylation. Science, 393(6809). https://doi.org/10.1126/science.adx0673
Guo, P., Cui, Y., He, J., Waldman, A.J., Zhu, J., Chen, Y., Huang, Z., Zhou, J., Phillips-Cremins, J.E., & Deng, Y. (2026). Integrative spatial profiling of 3D genome organization and gene expression in tissue. Cell. https://doi.org/10.1016/j.cell.2026.07.039
Yang, Z., Zhou, J., Fei, F., Zheng, B., Yang, S.X., Pearce, T.M., Zhang, P., Huang, L., Yue, J., Shen, Q., Du, Y., Liao, X., Cheng, S., Li, L., Liang, J., Liu, L., Wan, X., Taylor, M.D., Ecker, J.R., Zhou, S., Rich, J.N., & Zhao, L. (2026). Macrophage-mediated brain-bone marrow crosstalk promotes chronic stress-induced glioma growth. Cancer Cell, 44. https://doi.org/10.1016/j.ccell.2026.05.018