--- license: cc-by-4.0 library_name: mpac tags: - biology - genomics - dna - mpra - cis-regulatory - variant-effect-prediction pipeline_tag: other --- # MPAC MPAC (Malinois with Parallel Aggregated Cross-validation) predicts cis-regulatory activity of 200 bp human sequences in K562, HepG2 and SK-N-SH, and the allelic skew caused by non-coding variants. [Identifying non-coding variant effects at scale via machine learning models of cis-regulatory reporter assays](https://doi.org/10.1101/2025.04.16.648420). This repository holds the 110 published checkpoints, converted to safetensors from the [Zenodo deposit](https://doi.org/10.5281/zenodo.15178434) with no retraining or modification. ## Usage ```python from modeling_mpac import MPACEnsemble # Loads the ten models that held chr7 out of training ensemble = MPACEnsemble.from_pretrained("saarantras1/MPAC", chromosome=7, device="cuda") preds = ensemble.predict(["ACGT" * 50], device="cuda") # (n, 3): K562, HepG2, SKNSH ``` Use `from_pretrained` and `predict` rather than loading a checkpoint or calling the model directly: they select the ensemble that did not train on your query's chromosome, and they add the MPRA vector context and average over both strands. Skipping either step returns plausible-looking but wrong numbers instead of an error. `predict` follows `vcf_predict.py` from the upstream code base, which generated the published predictions: the reverse strand is the reverse complement of the 200 bp insert placed back in the forward-orientation vector, matching the assay, rather than a reverse complement of the whole 600 bp construct. MPAC covers autosomes only; `from_pretrained` raises on chrX, chrY and anything else with no held-out fold. For allelic skew, pass matched reference and alternate contexts. Each variant is tiled into eighteen 200 bp windows at stride 10 and averaged, reproducing the scheme behind the published predictions. ```python out = ensemble.predict_skew(ref_contexts, alt_contexts, device="cuda") out["skew"] # (n, 3), alt minus ref ``` Command-line tools for MPAC can be found at: [Reilly-Lab-Yale/coda_mpac](https://github.com/Reilly-Lab-Yale/coda_mpac). Which provide greater control of prediction parameters including (but not limited to): window number step size strand reduction insert/full plasmid reverse complement window aggregation (average, max, min, etc.) An example prediction for SNVs from a VCF can be found in the methods of the publication and below: To generate MPAC predictions for SNVs from the command line, first define your variants in a VCF-like TSV (no VCF header required) with the following tab-separated columns: | Column | Description | | --- | --- | | `CHROM` | Chromosome | | `POS` | Position | | `ID` | Variant ID (can be blank or `.` if none) | | `REF` | Reference allele | | `ALT` | Alternate allele | | `QUAL` | Quality (can be `.` if null) | | `FILTER` | Filter status (can be `.` if null) | | `INFO` | Can be `.` if null; predictions populate this column in the output | ### SNVs SNV predictions are handled with `vcf_predict.py`: ```bash python vcf_predict.py \ --artifact_path {10X $MODEL} \ --vcf_file ${VCF} \ --fasta_file ${FASTA} \ --output ${OUTPUT} \ --relative_start 9 \ --relative_end 180 \ --step_size 10 \ --strand_reduction mean \ --window_reduction mean \ --feature_ids K562 HepG2 SKNSH ``` | Argument | Description | | --- | --- | | `--artifact_path` | Chromosome holdout models to ensemble | | `--vcf_file` | Variants of interest | | `--fasta_file` | Reference genome | | `--output` | Output path | | `--relative_start` | First window variant position | | `--relative_end` | Final window variant position | | `--step_size` | Window step size | | `--strand_reduction` | Reduction method for fwd/rev strand predictions | | `--window_reduction` | Reduction method for window predictions | | `--feature_ids` | Labels for predictions in the output | Additonal arguments can be found in vcf_predict.py ### Small indels Small indels (≤ 10 bp recommended) are handled the same way with `vcf_predict_indel.py`, with modification to the plasmid sequence padding loop. ### Haplotypes Haplotype predictions are handled similarly with `vcf_predict_haplotype.py`, with modification to the windowing loop to ensure all windows contain the desired variants. ### Output Output is a VCF-like TSV with predictions occupying the `INFO` column, see below: ### Example output ``` CHROM POS ID REF ALT INFO chr22 11121724 COSV106573183 A G K562__ref=0.3535322;HepG2__ref=0.29076257;... ``` ## Citation ```bibtex @article{butts2025mpac, title = {Identifying non-coding variant effects at scale via machine learning models of cis-regulatory reporter assays}, author = {Butts, John C. and Rong, Stephen and Gosai, Sager J. and Castro, Rodrigo I. and Noon, Mackenzie and Adeniran, Kehinde and Ghosh, Rohit and Sabeti, Pardis C. and Tewhey, Ryan and Reilly, Steven K.}, journal = {bioRxiv}, year = {2025}, doi = {10.1101/2025.04.16.648420} } ``` ## License CC-BY-4.0, matching the Zenodo deposit. `modeling_mpac.py` derives from the MIT licensed model code in [sjgosai/boda2](https://github.com/sjgosai/boda2) and retains that notice.