DrugDis / component-resolved evaluation

Disentangling general and context‑specific effects in drug‑response prediction.

Strong overall accuracy can be driven by general drug and sample tendencies even when context-specific effects are recovered poorly.

DrugDis separates measured and predicted responses into additive effects and drug–sample interactions on the same observed pairs, then asks which component a model recovers, at what magnitude and with what error.

3,141,680 drug–sample pairs 986 cell lines 54,180 compounds 11 resources
Framework overviewFigure 1
The total response is split into an additive component and a drug-sample interaction component; the benchmark dataset integrates eleven response resources; the additive component carries 75.8% of response variance on 1.76% of the degrees of freedom.
DrugDis separates broad drug and sample tendencies from drug–sample-specific deviations, then evaluates the two components independently.
Core principle

Ask which predictive capability improved—not only whether the aggregate score increased.

01 / Framework

One aggregate score mixes distinct predictive capabilities.

DrugDis uses one support-defined orthogonal decomposition for both measurements and predictions. It separates component recovery from interaction magnitude and prediction error, with repeated measurements providing an empirical reference where available.

01 / Orthogonal decomposition

Which part of the response is additive, and which is interaction?

On the observed drug–sample pairs, the total response splits exactly into an additive component, the best additive fit of drug and sample marginals, and a drug–sample interaction component orthogonal to it.

The interaction component is defined relative to the observed pairs.

02 / Key findings

Aggregate performance can hide what a model actually recovers.

Across 3.14 million drug–sample pairs, component-resolved evaluation changes how model performance, representations and generalization should be interpreted.

75.8%

General response structure dominates variance.

Additive drug and sample effects occupy only 1.76% of the degrees of freedom but carry 75.8% of response variance. The additive share remains 59–83% within each of the five largest resources.

0.86 → 0.32

High aggregate accuracy can coexist with weak interaction recovery.

Under held-out cell lines, the same model reaches 0.86 total-response correlation but only 0.32 interaction correlation.

77% ↔ 64%

Different holdouts expose different failure modes.

Interactions absorb 77% of squared error for held-out cell lines, whereas additive effects absorb 64% for held-out compounds. The two regimes test different predictive capabilities.

score ≠ utility

Better component recovery does not guarantee a better independent outcome.

Component supervision improves interaction direction in selected regimes, but component-based model selection does not reduce the prespecified selectivity error on independent GDSC2 measurements.

Representation benchmark

Apparent representation advantages depend on the evaluated component, decoder and distribution shift; they are not properties of an embedding alone.

03 / Benchmark

Large-scale evaluation across unseen cells, unseen compounds and a new biological system.

The cell-line benchmark integrates eleven DROMA response resources on one CCLE-derived transcriptomic input. Patient-derived organoids are evaluated separately in zero-shot transfer with no organoid fine-tuning.

3.14Munique drug–sample pairs
986cancer cell lines
54,180compounds
11response resources
EvaluationSamplesCompoundsPairs
Cell-line benchmark98654,1803,141,680
Primary organoid set100784,886
All organoid cohorts17314510,010

Designed to separate generalization from composition effects.

Models are tested with held-out cell lines, held-out compounds and zero-shot organoid transfer. Because NCI60 contributes 68.5% of source response rows, additive variance is also decomposed within each major resource and after excluding NCI60.

04 / Software

Apply the same component-resolved evaluation to your own predictions.

Given measured and predicted responses on the same drug–sample pairs, the API returns total, additive and interaction recovery, interaction amplitude and exact component-wise error attribution.

From a response table to a component profile.

Run the bundled synthetic demo for an interface check, or pass your own test pairs with measured and predicted responses. The repository also holds every script, split manifest and result table behind the manuscript.

evaluate.py
from drugdis.evaluate import component_profile

scores = component_profile(
    test_frame,          # SMILES, Sample_ID per row
    y="Sensitivity",
    yhat="prediction",
)

print(scores["rawPCC"], scores["sharedPCC"],
      scores["intPCC"], scores["R2_interaction"])

Python 3.11+ · NumPy / pandas / SciPy core · MIT License

DDrugDis
2026
05 / Paper & resources

DrugDis: Disentangling general and context‑specific effects in drug‑response prediction

Bo Li, Chengliang Liu, Yuzhong Peng, Bob Zhang, Qing Wang, Pinxian Zeng, Mengran Li, Shenghui Huang, Chuxia Deng, and Yang Zhang

The public repository includes the code, the prespecified split manifests, the canonical result tables and the manuscript resources used on this page. The processed inputs are hosted on Hugging Face.