MPAC

MPAC (Malinois with Parallel Aggregated Cross-validation) predicts cis-regulatory activity of 200 bp human sequences in K562, HepG2 and SK-N-SH, and the allelic skew caused by non-coding variants.

Identifying non-coding variant effects at scale via machine learning models of cis-regulatory reporter assays.

This repository holds the 110 published checkpoints, converted to safetensors from the Zenodo deposit with no retraining or modification.

Usage

from modeling_mpac import MPACEnsemble

# Loads the ten models that held chr7 out of training
ensemble = MPACEnsemble.from_pretrained("saarantras1/MPAC", chromosome=7, device="cuda")

preds = ensemble.predict(["ACGT" * 50], device="cuda")   # (n, 3): K562, HepG2, SKNSH

Use from_pretrained and predict rather than loading a checkpoint or calling the model directly: they select the ensemble that did not train on your query's chromosome, and they add the MPRA vector context and average over both strands. Skipping either step returns plausible-looking but wrong numbers instead of an error.

predict follows vcf_predict.py from the upstream code base, which generated the published predictions: the reverse strand is the reverse complement of the 200 bp insert placed back in the forward-orientation vector, matching the assay, rather than a reverse complement of the whole 600 bp construct.

MPAC covers autosomes only; from_pretrained raises on chrX, chrY and anything else with no held-out fold.

For allelic skew, pass matched reference and alternate contexts. Each variant is tiled into eighteen 200 bp windows at stride 10 and averaged, reproducing the scheme behind the published predictions.

out = ensemble.predict_skew(ref_contexts, alt_contexts, device="cuda")
out["skew"]   # (n, 3), alt minus ref

Command-line tools for MPAC can be found at: Reilly-Lab-Yale/coda_mpac.

Which provide greater control of prediction parameters including (but not limited to): window number step size strand reduction insert/full plasmid reverse complement window aggregation (average, max, min, etc.)

An example prediction for SNVs from a VCF can be found in the methods of the publication and below:

To generate MPAC predictions for SNVs from the command line, first define your variants in a VCF-like TSV (no VCF header required) with the following tab-separated columns:

Column Description
CHROM Chromosome
POS Position
ID Variant ID (can be blank or . if none)
REF Reference allele
ALT Alternate allele
QUAL Quality (can be . if null)
FILTER Filter status (can be . if null)
INFO Can be . if null; predictions populate this column in the output

SNVs

SNV predictions are handled with vcf_predict.py:

python vcf_predict.py \
    --artifact_path {10X $MODEL} \
    --vcf_file ${VCF} \
    --fasta_file ${FASTA} \
    --output ${OUTPUT} \
    --relative_start 9 \
    --relative_end 180 \
    --step_size 10 \
    --strand_reduction mean \
    --window_reduction mean \
    --feature_ids K562 HepG2 SKNSH
Argument Description
--artifact_path Chromosome holdout models to ensemble
--vcf_file Variants of interest
--fasta_file Reference genome
--output Output path
--relative_start First window variant position
--relative_end Final window variant position
--step_size Window step size
--strand_reduction Reduction method for fwd/rev strand predictions
--window_reduction Reduction method for window predictions
--feature_ids Labels for predictions in the output

Additonal arguments can be found in vcf_predict.py

Small indels

Small indels (≤ 10 bp recommended) are handled the same way with vcf_predict_indel.py, with modification to the plasmid sequence padding loop.

Haplotypes

Haplotype predictions are handled similarly with vcf_predict_haplotype.py, with modification to the windowing loop to ensure all windows contain the desired variants.

Output

Output is a VCF-like TSV with predictions occupying the INFO column, see below:

Example output

CHROM   POS         ID              REF  ALT  INFO
chr22   11121724    COSV106573183   A    G    K562__ref=0.3535322;HepG2__ref=0.29076257;...

Citation

@article{butts2025mpac,
  title   = {Identifying non-coding variant effects at scale via machine learning
             models of cis-regulatory reporter assays},
  author  = {Butts, John C. and Rong, Stephen and Gosai, Sager J. and
             Castro, Rodrigo I. and Noon, Mackenzie and Adeniran, Kehinde and
             Ghosh, Rohit and Sabeti, Pardis C. and Tewhey, Ryan and Reilly, Steven K.},
  journal = {bioRxiv},
  year    = {2025},
  doi     = {10.1101/2025.04.16.648420}
}

License

CC-BY-4.0, matching the Zenodo deposit. modeling_mpac.py derives from the MIT licensed model code in sjgosai/boda2 and retains that notice.

Downloads last month
124
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support