Generating high-quality Perturb-seq data for the 2026 Virtual Cell Challenge
At Arc, we believe virtual cell models can fill a critical gap: predicting the effects of genetic mutations, environmental changes, and small molecule treatments in biological systems we cannot readily test at-scale or cost-effectively in the lab. Realizing that vision requires an ambitious effort, generating causal, single-cell resolution data at a scale that does not yet exist, so models can learn to reliably predict any cell type's biological response to a perturbation. We're building large-scale perturbation datasets as a resource toward that goal, the first step to generate the quality and breadth of data virtual cell models need to make trustworthy predictions.
This year's Virtual Cell Challenge further pushes the field on context generalization, a task to accurately predict perturbation effects in unobserved biological systems. Last year's Challenge asked contestants to use models to predict the transcriptional effects of genetic perturbations in a single cellular context after seeing a few perturbation profiles. This year, we make the leap to accurate zero-shot predictions across six cell lines representing several axes of biological diversity (Figure 1).
Enabling and evaluating these predictions requires high-quality data for each cellular context: transcriptional profiles of unperturbed cells, which we provide to contestants as input, and accurate ground truth perturbation data for scoring model performance. Consistent data quality is particularly critical for cross-context generalization, where technical artifacts and noisy measurements can obscure true biological signal across contexts. Here, we offer a glimpse into Arc's process for generating high-quality perturbation data across distinct cellular contexts (Figure 2) - a massive effort that required months of close collaboration across Arc’s experimental and computational Technology Center teams.
Arc’s approach to data generation
Most biological data available today for ML training is observational. It captures how gene expression and phenotypes correlate within and across samples, but it does not reveal how any particular sample arrived at its observed state. Building AI models that predict how cells respond to a perturbation – genetic, environmental, or chemical – requires causal data that directly measures the consequence of a specific intervention on a cell’s phenotype, connecting cause and effect (Rood et al., 2024).
At Arc, we use the functional genomics technique called Perturb-seq to generate causal data. Perturb-seq pairs a targeted genetic perturbation, in this case CRISPR interference (CRISPRi)-mediated knockdown of a single gene, with the transcriptional profile that results. This directly links cause to effect (perturbation → transcriptional profile). Each cell in a Perturb-seq experiment receives one perturbation, and single-cell RNA sequencing (scRNA-seq) captures both the perturbation identity and the cell's transcriptional response. Because Perturb-seq captures both the perturbation and its transcriptomic readout in the same cell, it generates the causal data that observational datasets cannot: thousands of controlled genetic experiments run in parallel, where each perturbation’s effect is quantified across the many cells that receive it. We discuss Perturb-seq and CRISPRi in more detail in last year’s post describing the 2025 Challenge dataset (“Behind the Data of the Virtual Cell Challenge | Arc Institute,” 2025).
Scaling data generation across cellular contexts
For this year’s competition, we chose six cell lines that capture biological diversity across dimensions relevant to model generalization: tissue of origin, cell lineage, and disease status (Figure 3). These cell lines differ from the H1 embryonic stem cells from the 2025 Challenge, and this difference in identity means they may respond differently to perturbation.
To generate the 2026 Challenge dataset, we executed Perturb-seq consistently across cell lines to minimize technical artifacts that could impact transcriptional readouts. We first validated each cell line for high on-target gene knockdown, applying a stringent CRISPRi activity threshold before proceeding. We used the same sgRNA sequences (RNA sequence that directs CRISPRi to a specific target gene) in every cell line and delivered them within a narrow multiplicity of infection range (the ratio of sgRNA-containing delivery particles to cells in the experiment). We collected cells at the same time point after perturbation introduction and processed them using a consistent protocol, minimizing differences in transcriptional profiles due to time or handling. We performed scRNA-seq profiling and sequencing using the same technology platforms, and processed sequencing data using the same pipeline across all cell lines. Finally, we applied a suite of stringent quality control metrics at each step of the Perturb-seq workflow (Figure 4), ensuring that only data from high-performing experiments was included in this year’s Challenge.
The Perturb-seq experiments each included 100 non-targeting sgRNAs (NTCs), which serve as a negative control. NTCs define the baseline transcriptional state of each cell line. We performed deep profiling across ~50,000 control cells per experiment to produce a stable reference distribution against which we compare perturbed cells. NTCs matter because the CRISPRi machinery itself can shift transcription independent of any specific gene repression. NTCs capture this background signal so we can distinguish it from true perturbation effects. Because we pool these controls within the same experiment as the perturbations, they also help account for other sources of variation between samples. The performance of the NTCs defines data quality alongside the other metrics, confirming we capture true biological changes from our targeted gene silencing.
In each cell line, we profiled more than 10,000 perturbations in addition to the NTCs. Our experimental teams handled over 1 billion cells to ultimately generate deeply-sequenced transcriptional profiles from over 30 million cells across the six cell lines. With this scale, we captured a median of 500 cells per perturbation per cell line.
The perturbations and NTCs included in this year’s Challenge were curated from the larger set of 10,000 perturbations and 100 NTCs. To ensure we selected genes with high-quality data and consistent data depth, we chose perturbations that produced at least 80% median on-target gene expression reduction and subsampled to 400 cells per perturbation (Figure 5). We took great care during this data curation process to ensure the competition encompassed confident biological data, creating an ideal ground-truth to measure performance.
We use 10x Genomics Flex probe-based technology to enable high throughput scRNA-seq data generation while minimizing technical noise and batch effects (“Behind the Data of the Virtual Cell Challenge | Arc Institute,” 2025; Swinderman et al., 2026). Custom probe designs recapture the sgRNAs that define our perturbation, ensuring cells we consider “perturbed” show high detection levels of sgRNAs and reduce uncertainty in our perturbation assignments.
Prioritizing sequencing depth strengthens signal detection for lowly expressed genes and reduces data loss, both critical for capturing subtle perturbation effects. To reach our desired depth of median 20,000 unique RNA transcripts per cell (or UMIs), we generated >800 billion sequencing reads per cell line on the UG100 platform from Ultima Genomics. The scale of the sequencing data we produced for this year’s Challenge also prompted an evolution in how we store and process sequencing data. To keep pace, we developed cyto (Teyssier and Dobin, 2026), Arc’s bioinformatics pipeline for rapid 10x Flex processing.
Seven Arc Technology Center teams, Functional Genomics, Molecular Engineering, Multiomics Profiling, Next Generation Sequencing Platform, Bioinformatics, Infrastructure, and Machine Learning, contributed to generating and processing this data. This scale reflects the reality of producing high-quality perturbational data: it demands substantial laboratory infrastructure, sequencing capacity, specialized expertise, and cross-team coordination.
Toward building a general context virtual cell
Virtual cell data generation and curation will continue to evolve as models do. The scale of data necessary for a comprehensive virtual cell that can complete the context generalization task with high accuracy will be massive. These six cell lines are the tip of the iceberg, and Arc and others in the field will continue to build causal data at scale to fuel these models. While Arc has focused on genetic perturbations, the Allen Institute has generated causal data for their cytokine signaling studies in T cells (“10x Flex v2 Enables High Throughput IL-6 Signaling Inhibitor Analysis in T cells,” n.d.). Like Arc, the Allen Institute has opted to use 10x Flex profiling for their large-scale cell profiling efforts. Other notable causal data generation includes the Tahoe-100M chemical perturbation data (Zhang et al., 2025), and cytokine perturbation data in immune cells from the Fabian Theis and Georg Seelig groups (Oesinghaus et al., 2025), which used the Parse scRNA-seq and Ultima Genomics UG100 sequencing platforms.
These atlasing efforts will continue to grow, and our experimental decisions evolve as we receive the real-time feedback from our model development team. Next, we will refine how we iterate between model learning and data generation, bringing us closer to an optimized lab-in-the-loop.
Bibliography
10x Flex v2 Enables High Throughput IL-6 Signaling Inhibitor Analysis in T cells [WWW Document], n.d. https://doi.org/10.57785/vpde-ss91
Behind the Data of the Virtual Cell Challenge | Arc Institute [WWW Document], 2025. URL https://arcinstitute.org/news/behind-the-data-virtual-cell-challenge (accessed 8.19.26).
Oesinghaus, L., Becker, S., Vornholz, L., Papalexi, E., Pangallo, J., Monifar, A.A., Liu, J., La Fleur, A., Shulman, M., Marrujo, S., Hariadi, B., Curca, C., Suyama, A., Nigos, M., Sanderson, O., Nguyen, H., Tran, V.K., Sapre, A.A., Kaplan, O., Schroeder, S., Salvino, A., Gallareta-Olivares, G., Koehler, R., Geiss, G., Rosenberg, A.B., Roco, C.M., Seelig, G., Theis, F.J., 2025. A single-cell cytokine dictionary of human peripheral blood. bioRxiv 2025.12.12.693897. https://doi.org/10.64898/2025.12.12.693897
Replogle, J.M., Bonnar, J.L., Pogson, A.N., Liem, C.R., Maier, N.K., Ding, Y., Russell, B.J., Wang, X., Leng, K., Guna, A., Norman, T.M., Pak, R.A., Ramos, D.M., Ward, M.E., Gilbert, L.A., Kampmann, M., Weissman, J.S., Jost, M., 2022. Maximizing CRISPRi efficacy and accessibility with dual-sgRNA libraries and optimal effectors. eLife 11, e81856. https://doi.org/10.7554/eLife.81856
Rood, J.E., Hupalowska, A., Regev, A., 2024. Toward a foundation model of causal cell and tissue biology with a Perturbation Cell and Tissue Atlas. Cell 187, 4520–4545. https://doi.org/10.1016/j.cell.2024.07.035
Swinderman, J.T., Tung, P.-Y., Winters, A., Goudy, L., Wilson, C.M., Bonds, L.R., Teyssier, N., Agrawal, A., Dobin, A., Hua, T., Goodarzi, H., Feng, F.Y., Marson, A., Burke, D.P., Hsu, P.D., Roohani, Y.H., Konermann, S., Kosicki, M., Li, N., Gilbert, L.A., 2026. Scalable probe-based single-cell transcriptional profiling for virtual cell perturbation mapping and synthetic biology phenotyping. https://doi.org/10.64898/2026.02.04.703058
Teyssier, N., Dobin, A., 2026. cyto: ultra high-throughput processing of 10x-flex single cell sequencing. https://doi.org/10.64898/2026.01.21.700936
Zhang, J., Ubas, A.A., Borja, R. de, Svensson, V., Thomas, N., Thakar, N., Lai, I., Winters, A., Khan, U., Jones, M.G., Thompson, J.D., Tran, V., Pangallo, J., Papalexi, E., Sapre, A., Nguyen, H., Sanderson, O., Nigos, M., Kaplan, O., Schroeder, S., Hariadi, B., Marrujo, S., Salvino, C.C.A., Olivares, G.G., Koehler, R., Geiss, G., Rosenberg, A., Roco, C., Merico, D., Alidoust, N., Goodarzi, H., Yu, J., 2025. Tahoe-100M: A Giga-Scale Single-Cell Perturbation Atlas for Context-Dependent Gene Function and Cellular Modeling. https://doi.org/10.1101/2025.02.20.639398