PertResolve / measurement-resolution framework

Measurement resolution constrains fine-grained perturbation prediction.

A benchmark cannot support biological distinctions that its own measurements do not reproducibly resolve.

PertResolve separates measurement questions from model questions, quantifies which perturbation distinctions reproduce in the experiment, and interprets prediction performance only at that supported resolution.

470 coding-variant conditions 321,043 cells in the allele arm 31 perturbation configurations 14 public resources
Framework overviewFigure 1
PertResolve framework separating measurement resolution from model prediction
Detection and identification are measurement questions; prediction is a model question. PertResolve evaluates them separately before interpreting benchmark scores.
Core principle

Measure the distinction first. Interpret the prediction second.

01 / Framework

Four empirical questions should not be collapsed into one score.

PertResolve keeps measurement support and model performance separate. Detection asks whether a perturbation differs from reference; identification asks whether it differs from relevant competitors; split-half reproducibility asks whether the experiment recovers the same structure twice; model-ranking resolution asks which predictor gaps the benchmark can reliably order.

01 / Detection

Can the perturbation be distinguished from reference?

PertResolve measures whether a perturbation response separates reproducibly from the reference population under disjoint cell sampling.

Detection answers whether there is a measurable response. It does not imply that nearby perturbations can be identified.

02 / Study design

A fine-grained allele case and a broader perturbation panel answer complementary questions.

The allele-resolved arm tests whether experiments and models can distinguish variants of the same gene. The broader panel asks whether measurement-resolution constraints recur across genetic, chemical and cytokine perturbations, cellular contexts, donors and readout modalities.

470coding-variant conditions
321kcells in the allele-resolved arm
31perturbation configurations
14public data resources
Allele datasetVariantsContextAssay
TP5398A549Perturb-seq
KRAS92A549Perturb-seq
GATA1254Human HSPCsBase editing
JAK126HT-29scSNV-seq

The scientifically relevant competitor is often a sibling perturbation, not an unrelated condition.

A model may recover a strong gene-level programme while failing to identify the correct allele. PertResolve therefore evaluates within-gene discrimination explicitly and then tests the same measurement logic across a broader perturbation landscape.

TP53KRASGATA1JAK1
03 / Fine-grained evidence

The experiment itself can fail to resolve the allele distinction being scored.

A disjoint split-half measurement provides an empirical reference for within-gene discrimination. TP53 and KRAS remain near chance, GATA1 shows partial structure, whereas JAK1 supports substantially stronger allele-level recovery. The four genes therefore expose different measurement regimes before any model is compared.

TP53

98 variants
0.472split-half PDS · n=59 scored · 95% CI 0.441–0.505
Pairwise-resolvable pairs
0.0%
Native-depth unrankable*
100%
Top-1 split-half recovery
0.45%

KRAS

92 variants
0.481split-half PDS · n=51 scored · 95% CI 0.431–0.533
Pairwise-resolvable pairs
0.0%
Native-depth unrankable*
100%
Top-1 split-half recovery
1.05%

GATA1

254 variants
0.578split-half PDS · n=172 scored · 95% CI 0.552–0.604
Pairwise-resolvable pairs
3.8%
Native-depth unrankable*
97.6%
Top-1 split-half recovery
0.81%

JAK1

26 variants
0.807split-half PDS · n=26 scored · 95% CI 0.729–0.880
Pairwise-resolvable pairs
78.5%
Native-depth unrankable*
10.0%
Top-1 split-half recovery
32.6%

* Native-depth eligibility differs from the full variant count. The unrankable estimates use n=98 TP53, 92 KRAS, 254 GATA1 and 20 JAK1; the split-half PDS uses the scored counts printed above.

Measurement-limited comparison

Near-chance split-half recovery changes what a model score can mean.

For TP53 and KRAS, two disjoint measurements of the same allele do not reliably recover within-gene identity. A near-chance model score in this regime cannot be interpreted simply as evidence that the model failed to learn a distinction that the experiment itself supports.

Model-limited comparison

JAK1 separates measurement support from model capability.

The JAK1 split-half reference reaches 0.807, while the canonical gene-wise evaluated-predictor summary reports 0.448 for its highest PDS entry. Here the experimental response contains reproducible allele structure, so the remaining prediction gap is not explained by the same measurement limitation.

Candidate geometry

Detecting an effect is easier than identifying the right sibling perturbation.

Pairwise resolvability ranges from 0% for TP53 and KRAS to 78.5% for JAK1. Fine-grained interpretation therefore depends on local competitor geometry rather than perturbation-versus-control separation alone.

Prediction metrics

Model ordering changes with the question being scored.

Across five released perturbation-prediction models, the conventional response-correlation summary and the fine-grained perturbation-discrimination score do not induce the same ordering.

scGen has the largest Pearson-Δ among these five models (0.201) but PDS 0.474, whereas PerturbNet has the largest PDS (0.530) but Pearson-Δ 0.134. This is why PertResolve reports response prediction and perturbation identification as different questions rather than treating one metric as a substitute for the other.

Released modelPDSPDS 95% CIPearson-Δ
CellFlow0.4930.467–0.5200.156
Biolord0.4790.453–0.5050.139
scGen0.4740.449–0.5000.201largest Pearson-Δ in this table
PerturbNet0.5300.504–0.5550.134largest PDS in this table
scVIDR0.4860.461–0.5110.183
470 → 4variant conditions → gene-level representations

Representation resolution must match evaluation resolution.

Gene-keyed perturbation interfaces evaluated in the manuscript comparison can accept all 470 conditions but collapse variants of the same gene onto the same perturbation representation. A gene-constant reference gives PDS 0.478 (95% CI 0.450–0.504), illustrating why allele-level scoring is a category mismatch when the model input cannot express allele identity.

04 / Across-dataset generalization

Measurement resolution varies sharply across public perturbation resources — and within the same resource.

The broader panel asks whether perturbation responses are reproducibly detectable across biological and technical regimes. Detection-side resolution spans almost the full 0–1 range, shifts across cell contexts and donors, and can differ between RNA and protein readouts.

05 / What shapes resolution

More cells help estimation, but they do not by themselves create biological separation.

Cross-dataset depth is an incomplete proxy for resolution. Controlled analyses instead point to the local separation between a perturbation and its closest competitors, together with sampling depth and benchmark size, as the quantities that determine what can be reproducibly distinguished and ranked.

ρ = 0.961nearest-competitor signal vs split-half PDS

Local competitor geometry is highly informative.

In the controlled resolution sweep, the median nearest-neighbour signal tracks the split-half discrimination reference more closely than the global signal summary. Fine-grained resolution is therefore governed by the closest confusable alternatives, not by average separation alone.

depth ≠ resolutioncross-dataset comparison

The deepest experiment is not necessarily the most resolvable.

Parse 10M is the deepest point in the cell-depth comparison, yet donor 1 reaches only 0.256 detection fraction. GSE306429 A549 reaches 1.00 with far fewer cells per perturbation. Depth reduces sampling uncertainty but cannot manufacture an absent or weak response distinction.

no single gateindependent measurement and ranking axes

Resolution is not one scalar pass/fail property.

The controlled sweep shows monotonic improvement of the split-half discrimination reference with local signal, but model-order recovery is not itself monotonic. Whether a benchmark can rank predictors also depends on the score gap and number of evaluated conditions.

Prospective experiment design

A small disjoint pilot can flag which perturbations are likely to become rankable.

Using only 25 pilot cells per perturbation, the learned pilot score predicts the predefined depth-100 rankability label in three genetic screens with high AUROC. A training-free signal-to-noise score retains substantial discrimination without fitting the learned pilot model.

Pilot cells and evaluation cells are disjoint; rankability is defined from split-half signal exceeding the uncertainty width in the held-out depth-100 evaluation slice.

Dataset25-cell learned-effect AUROC25-cell training-free SNR AUROC
Replogle0.9600.934
Norman0.9160.851
Adamson0.8940.856
06 / Interpretation

The same prediction score has different meaning in different measurement regimes.

PertResolve does not replace prediction metrics with a composite resolution score. It supplies the measurement context needed to interpret them: whether the target distinction is experimentally supported, whether the model representation can express it, and whether the benchmark has enough resolution to order competing predictors.

I

Detectable is not identifiable.

A perturbation can differ clearly from control yet remain indistinguishable from its nearest biological competitors. Detection and identification therefore answer different measurement questions.

II

Measurement-limited is not model-limited.

TP53 and KRAS illustrate weak within-gene measurement recovery; JAK1 provides a contrasting regime where the experiment supports allele discrimination but current evaluated predictors still leave a large gap.

III

Benchmark ranking is its own problem.

Even when perturbation responses are measurable, a finite benchmark may not reliably order small model-score differences. Model-ranking resolution must therefore be assessed separately from perturbation detection or identification.

07 / Software

Apply the same diagnostic questions to your own representation.

The NumPy-first API accepts a cell-by-feature matrix or fixed representation and reports the measurement quantities separately rather than manufacturing a composite verdict.

From expression matrix to resolution report.

Use the bundled demo for a deterministic interface check, or provide your own perturbation labels and reference condition.

resolution_report.py
from pertresolve.resolution import resolution_report

report = resolution_report(
    X,
    labels,
    control="non-targeting",
    depth=50,
)

print(report.summary())

Python ≥ 3.10 · NumPy / pandas core · MIT License

PPertResolve
2026
Paper & resources

Measurement resolution constrains fine-grained perturbation prediction

Bo Li, Chengyang Zhang, Mengran Li, Bob Zhang, Lin Wang, Zhenchao Tang, Jun Liu, Chengliang Liu, Chen Wei, Yuhao Yi, Jiancheng Lv, and Yang Zhang

The public repository includes the software, benchmark metadata, canonical result tables, analysis entry points and manuscript resources used by this project page.