MPAC
MPAC (Malinois with Parallel Aggregated Cross-validation) predicts cis-regulatory activity of 200 bp human sequences in K562, HepG2 and SK-N-SH, and the allelic skew caused by non-coding variants.
This repository holds the 110 published checkpoints, converted to safetensors from the Zenodo deposit with no retraining or modification.
Usage
from modeling_mpac import MPACEnsemble
# Loads the ten models that held chr7 out of training
ensemble = MPACEnsemble.from_pretrained("saarantras1/MPAC", chromosome=7, device="cuda")
preds = ensemble.predict(["ACGT" * 50], device="cuda") # (n, 3): K562, HepG2, SKNSH
Use from_pretrained and predict rather than loading a checkpoint or calling the
model directly: they select the ensemble that did not train on your query's
chromosome, and they add the MPRA vector context and average over both strands.
Skipping either step returns plausible-looking but wrong numbers instead of an error.
predict follows vcf_predict.py from the upstream code base, which generated the
published predictions: the reverse strand is the reverse complement of the 200 bp
insert placed back in the forward-orientation vector, matching the assay, rather
than a reverse complement of the whole 600 bp construct.
MPAC covers autosomes only; from_pretrained raises on chrX, chrY and anything else
with no held-out fold.
For allelic skew, pass matched reference and alternate contexts. Each variant is tiled into eighteen 200 bp windows at stride 10 and averaged, reproducing the scheme behind the published predictions.
out = ensemble.predict_skew(ref_contexts, alt_contexts, device="cuda")
out["skew"] # (n, 3), alt minus ref
Command-line tools for MPAC can be found at: Reilly-Lab-Yale/coda_mpac.
Which provide greater control of prediction parameters including (but not limited to): window number step size strand reduction insert/full plasmid reverse complement window aggregation (average, max, min, etc.)
An example prediction for SNVs from a VCF can be found in the methods of the publication and below:
To generate MPAC predictions for SNVs from the command line, first define your variants in a VCF-like TSV (no VCF header required) with the following tab-separated columns:
| Column | Description |
|---|---|
CHROM |
Chromosome |
POS |
Position |
ID |
Variant ID (can be blank or . if none) |
REF |
Reference allele |
ALT |
Alternate allele |
QUAL |
Quality (can be . if null) |
FILTER |
Filter status (can be . if null) |
INFO |
Can be . if null; predictions populate this column in the output |
SNVs
SNV predictions are handled with vcf_predict.py:
python vcf_predict.py \
--artifact_path {10X $MODEL} \
--vcf_file ${VCF} \
--fasta_file ${FASTA} \
--output ${OUTPUT} \
--relative_start 9 \
--relative_end 180 \
--step_size 10 \
--strand_reduction mean \
--window_reduction mean \
--feature_ids K562 HepG2 SKNSH
| Argument | Description |
|---|---|
--artifact_path |
Chromosome holdout models to ensemble |
--vcf_file |
Variants of interest |
--fasta_file |
Reference genome |
--output |
Output path |
--relative_start |
First window variant position |
--relative_end |
Final window variant position |
--step_size |
Window step size |
--strand_reduction |
Reduction method for fwd/rev strand predictions |
--window_reduction |
Reduction method for window predictions |
--feature_ids |
Labels for predictions in the output |
Additonal arguments can be found in vcf_predict.py
Small indels
Small indels (≤ 10 bp recommended) are handled the same way with vcf_predict_indel.py, with modification to the plasmid sequence padding loop.
Haplotypes
Haplotype predictions are handled similarly with vcf_predict_haplotype.py, with modification to the windowing loop to ensure all windows contain the desired variants.
Output
Output is a VCF-like TSV with predictions occupying the INFO column, see below:
Example output
CHROM POS ID REF ALT INFO
chr22 11121724 COSV106573183 A G K562__ref=0.3535322;HepG2__ref=0.29076257;...
Citation
@article{butts2025mpac,
title = {Identifying non-coding variant effects at scale via machine learning
models of cis-regulatory reporter assays},
author = {Butts, John C. and Rong, Stephen and Gosai, Sager J. and
Castro, Rodrigo I. and Noon, Mackenzie and Adeniran, Kehinde and
Ghosh, Rohit and Sabeti, Pardis C. and Tewhey, Ryan and Reilly, Steven K.},
journal = {bioRxiv},
year = {2025},
doi = {10.1101/2025.04.16.648420}
}
License
CC-BY-4.0, matching the Zenodo deposit. modeling_mpac.py derives from the MIT
licensed model code in sjgosai/boda2 and retains
that notice.
- Downloads last month
- 124