File size: 5,365 Bytes
73c88ac
 
 
 
 
 
 
 
 
 
 
 
 
2c3b493
73c88ac
2c3b493
 
 
73c88ac
2c3b493
 
73c88ac
2c3b493
 
 
73c88ac
 
 
 
 
 
2c3b493
 
73c88ac
2c3b493
73c88ac
 
2c3b493
 
 
 
73c88ac
365ec0b
 
 
 
 
2c3b493
 
73c88ac
8906ce3
09d118b
 
 
 
 
 
 
 
8906ce3
 
 
 
 
 
 
 
 
 
 
 
a364167
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
73c88ac
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2c3b493
73c88ac
2c3b493
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
---
license: cc-by-4.0
library_name: mpac
tags:
  - biology
  - genomics
  - dna
  - mpra
  - cis-regulatory
  - variant-effect-prediction
pipeline_tag: other
---

# MPAC

MPAC (Malinois with Parallel Aggregated Cross-validation) predicts cis-regulatory
activity of 200 bp human sequences in K562, HepG2 and SK-N-SH, and the allelic skew
caused by non-coding variants.

[Identifying non-coding variant effects at scale via machine learning models of
cis-regulatory reporter assays](https://doi.org/10.1101/2025.04.16.648420).

This repository holds the 110 published checkpoints, converted to safetensors from
the [Zenodo deposit](https://doi.org/10.5281/zenodo.15178434) with no retraining or
modification.

## Usage

```python
from modeling_mpac import MPACEnsemble

# Loads the ten models that held chr7 out of training
ensemble = MPACEnsemble.from_pretrained("saarantras1/MPAC", chromosome=7, device="cuda")

preds = ensemble.predict(["ACGT" * 50], device="cuda")   # (n, 3): K562, HepG2, SKNSH
```

Use `from_pretrained` and `predict` rather than loading a checkpoint or calling the
model directly: they select the ensemble that did not train on your query's
chromosome, and they add the MPRA vector context and average over both strands.
Skipping either step returns plausible-looking but wrong numbers instead of an error.

`predict` follows `vcf_predict.py` from the upstream code base, which generated the
published predictions: the reverse strand is the reverse complement of the 200 bp
insert placed back in the forward-orientation vector, matching the assay, rather
than a reverse complement of the whole 600 bp construct.

MPAC covers autosomes only; `from_pretrained` raises on chrX, chrY and anything else
with no held-out fold.

For allelic skew, pass matched reference and alternate contexts. Each variant is tiled into eighteen
200 bp windows at stride 10 and averaged, reproducing the scheme behind the
published predictions.

```python
out = ensemble.predict_skew(ref_contexts, alt_contexts, device="cuda")
out["skew"]   # (n, 3), alt minus ref
```

Command-line tools for MPAC can be found at:
[Reilly-Lab-Yale/coda_mpac](https://github.com/Reilly-Lab-Yale/coda_mpac).

Which provide greater control of prediction parameters including (but not limited to):
  window number
  step size
  strand reduction
  insert/full plasmid reverse complement
  window aggregation (average, max, min, etc.)
  
An example prediction for SNVs from a VCF can be found in the methods of the publication and below:

To generate MPAC predictions for SNVs from the command line, first define your variants in a VCF-like TSV (no VCF header required) with the following tab-separated columns:

| Column | Description |
| --- | --- |
| `CHROM` | Chromosome |
| `POS` | Position |
| `ID` | Variant ID (can be blank or `.` if none) |
| `REF` | Reference allele |
| `ALT` | Alternate allele |
| `QUAL` | Quality (can be `.` if null) |
| `FILTER` | Filter status (can be `.` if null) |
| `INFO` | Can be `.` if null; predictions populate this column in the output |

### SNVs

SNV predictions are handled with `vcf_predict.py`:

```bash
python vcf_predict.py \
    --artifact_path {10X $MODEL} \
    --vcf_file ${VCF} \
    --fasta_file ${FASTA} \
    --output ${OUTPUT} \
    --relative_start 9 \
    --relative_end 180 \
    --step_size 10 \
    --strand_reduction mean \
    --window_reduction mean \
    --feature_ids K562 HepG2 SKNSH
```

| Argument | Description |
| --- | --- |
| `--artifact_path` | Chromosome holdout models to ensemble |
| `--vcf_file` | Variants of interest |
| `--fasta_file` | Reference genome |
| `--output` | Output path |
| `--relative_start` | First window variant position |
| `--relative_end` | Final window variant position |
| `--step_size` | Window step size |
| `--strand_reduction` | Reduction method for fwd/rev strand predictions |
| `--window_reduction` | Reduction method for window predictions |
| `--feature_ids` | Labels for predictions in the output |

Additonal arguments can be found in vcf_predict.py

### Small indels

Small indels (≤ 10 bp recommended) are handled the same way with `vcf_predict_indel.py`, with modification to the plasmid sequence padding loop.

### Haplotypes

Haplotype predictions are handled similarly with `vcf_predict_haplotype.py`, with modification to the windowing loop to ensure all windows contain the desired variants.

### Output

Output is a VCF-like TSV with predictions occupying the `INFO` column, see below:

### Example output

```
CHROM   POS         ID              REF  ALT  INFO
chr22   11121724    COSV106573183   A    G    K562__ref=0.3535322;HepG2__ref=0.29076257;...
```
## Citation

```bibtex
@article{butts2025mpac,
  title   = {Identifying non-coding variant effects at scale via machine learning
             models of cis-regulatory reporter assays},
  author  = {Butts, John C. and Rong, Stephen and Gosai, Sager J. and
             Castro, Rodrigo I. and Noon, Mackenzie and Adeniran, Kehinde and
             Ghosh, Rohit and Sabeti, Pardis C. and Tewhey, Ryan and Reilly, Steven K.},
  journal = {bioRxiv},
  year    = {2025},
  doi     = {10.1101/2025.04.16.648420}
}
```

## License

CC-BY-4.0, matching the Zenodo deposit. `modeling_mpac.py` derives from the MIT
licensed model code in [sjgosai/boda2](https://github.com/sjgosai/boda2) and retains
that notice.