An Empirical Comparison of Virtual Cell Models: Perturbation Prediction, Representation, and the Baseline Gap
Abstract
The “AI virtual cell” has become a rallying vision for computational biology: a learned model that represents cell state and simulates how cells respond to genetic and chemical perturbations. A rapidly growing zoo of candidate models—from purpose-built perturbation predictors (GEARS, CPA, chemCPA, scGen) to single-cell foundation-model embeddings (Geneformer, scGPT, scFoundation, UCE) to the recent integrated State model—is now marketed under this banner, but they are trained, evaluated and reported so heterogeneously that it is unclear which capabilities are real. We assemble a unified benchmark of eleven methods spanning five model families across six task families and nine public datasets, holding data splits, preprocessing and metrics fixed, and paying explicit attention to strong simple baselines and to how the choice of metric shapes conclusions. Three results stand out. First, on unseen single-gene perturbations no method dramatically beats a well-tuned additive/linear baseline: the best model (State) improves over a cell-mean baseline by 26% versus 19% for the linear baseline and 22% for GEARS, and several foundation-model predictors fall below the linear baseline—echoing recent critical benchmarks. Second, the apparent success of virtual cells is highly metric-dependent: an all-genes Pearson correlation, still widely reported, is maximized by a model that predicts no change and barely separates good from useless predictors; differential-expression and discrimination metrics tell a very different story. Third, the picture is not uniformly negative—deep models earn clear, reproducible gains on combinatorial (non-additive) perturbations, on cross-context transfer to unseen cell types, and above all on representation tasks (annotation and integration), where foundation embeddings decisively beat linear baselines. We connect this representation-strong/prediction-weak profile to interpretability evidence that these models encode organized, largely correlational cell-type structure rather than causal regulatory logic. We argue that the field is closer to a useful virtual microscope than a virtual simulator, and we release the benchmark to make that distinction measurable.
Related articles
Related articles are currently not available for this article.