PIE: cross-dataset generalization to Replogle

PIE predicts the transcriptional response to a perturbation in a cell context using biological knowledge embeddings, context control expression, and pooled response evidence from the training set.

Code and CLI documentation: ArcInstitute/PIE
Datasets and knowledge sources: PIE collection

For every measured gene, it returns:

Output Meaning
p_de Predicted probability of differential expression
lfc_pred Predicted log2 fold change
delta_p_pred Predicted change in mean expression on the preprocessed expression scale

This repository contains one model from the replogle_xdataset experiment. It trains on Tahoe, Jiang, ARC VCC 25, and Orion, and evaluates on Replogle without training or fine-tuning on Replogle response labels. The benchmark measures transfer across datasets for perturbations that are either seen or unseen in training.

Checkpoint and files

README.md
best_auprc.ckpt
config.yaml
data_stats.json
eval/
  test_seen/{metrics_best_auprc.csv,granular_best_auprc.csv}
  test_unseen/{metrics_best_auprc.csv,granular_best_auprc.csv}

best_auprc.ckpt is the Lightning checkpoint with the highest validation binary_auprc. Validation uses the four training datasets. The checkpoint embeds its model configuration, pinned asset references, and fitted data statistics; the adjacent YAML and JSON files make those settings inspectable.

Data and training

The saved run uses the following preprocessed datasets.

Dataset Role Pinned HF revision
Tahoe Training and validation 597d6ded9f9a148331388484ffb193489121c702
Jiang Training and validation 99242356dfcf58e049d867703c5743df8d1790bf
ARC VCC 25 Training and validation e7dda6065316959958b87583ddf0a15f6e69c37d
Orion Training and validation 0950c8aa5ce5b33dc9d6f0c5eade3ae445ca0ffb
Replogle-Nadig essential Evaluation; anchors gene-axis ordering 20c9faef76fc96fdc809871340bd93499413dfe7

Replogle remains in the data configuration for evaluation, with no training or validation pairs. Its response labels do not contribute to training or pooled response evidence. Its gene axis anchors the union's ordering; Replogle evaluation reports predictions for its 6,642 measured genes. Target-context control expression is an inference input, so zero-shot here means no target response-label training or fine-tuning.

Canonical splits live in PIE splits under replogle_xdataset/. The training split contains 64,260 context–perturbation pairs, and validation contains 957.

Knowledge sources are esm2, ncbi_text, string_space, depmap_gene_effect, context_text, perturbation_text, smiles, l1000_tas, prism_secondary, and jump_morphology, plus gene_text for gene queries. All come from PIE sources at revision cb1aaa4e7655605bdc70a9bd77bbd62016b8c7d7.

Evaluation

Both settings evaluate best_auprc.ckpt on HepG2, Jurkat, K562, and RPE1 in Replogle, with no Replogle response-label training or fine-tuning. Seen/unseen refers to whether the perturbation identifier occurs in the cross-dataset train.json.

Setting Target dataset Perturbations Test pairs
test_seen Replogle Seen in training 2,286
test_unseen Replogle Unseen in training 3,173

Values below are mean ± one sample standard deviation across the four Replogle cell lines (ddof=1), with equal weight per cell line.

Setting AUPRC Jaccard Direction match Spearman DE count Spearman LFC L1 discrimination
test_seen 0.3802 ± 0.0242 0.1523 ± 0.0176 0.7401 ± 0.0236 0.6231 ± 0.0542 0.5130 ± 0.0421 0.6312 ± 0.0414
test_unseen 0.2463 ± 0.0560 0.1196 ± 0.0010 0.6466 ± 0.0121 0.5016 ± 0.0114 0.2698 ± 0.0614 0.5309 ± 0.0140

Metric names in the CSVs are binary_auprc, sig_jaccard, direction_match, spearman_nsig, spearman_lfc, and discrimination_score_l1. Higher is better for all six.

Download and use

Install PIE with pip install arc-pie (or, from a checkout of the PIE code repository, uv sync --frozen for the pinned environment). Set the writable data, runs and cache directories in common.sh (copy common.sh.example from the code repository):

pip install arc-pie
hf auth login

hf download arcinstitute/PIE_replogle_xdataset \
  best_auprc.ckpt config.yaml data_stats.json \
  --local-dir models/replogle_xdataset

Evaluate both Replogle test settings:

SPLITS=hf://datasets/arcinstitute/PIE_splits@396ab9563175ee887750c9eed7ccaea6f5fdbf50
for setting in test_seen test_unseen; do
  pie eval experiment_name=replogle_xdataset \
    run_dir=models/replogle_xdataset ckpt=best_auprc \
    split_path=$SPLITS/replogle_xdataset/$setting.json row_set=$setting
done

To write predictions for the unseen-perturbation test pairs:

SPLITS=hf://datasets/arcinstitute/PIE_splits@396ab9563175ee887750c9eed7ccaea6f5fdbf50
pie infer experiment_name=replogle_xdataset \
  run_dir=models/replogle_xdataset ckpt=best_auprc \
  rows_kind=split rows_path=$SPLITS/replogle_xdataset/test_unseen.json \
  output_path=models/replogle_xdataset/infer/test_unseen.parquet

Prediction parquet files contain one row per pair, arrays of the three outputs, and the gene axis in file metadata. For a new context, prepare a controls_only=true dataset, provide a query JSON file, and use the preprocessed_dirs override as described in the PIE documentation. The original training assets remain necessary for response evidence. Runtime split files must match the training split hash embedded in the checkpoint.

License

The PIE model checkpoints and accompanying files are released under the Arc Research Institute PIE Model Non-Commercial License and are subject to the PIE Model Acceptable Use Policy. The PIE code is licensed separately under CC BY-NC-SA 4.0; see the code repository.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train arcinstitute/PIE_replogle_xdataset

Collection including arcinstitute/PIE_replogle_xdataset