Update README.md
#1
by buttsjoc - opened
README.md
CHANGED
|
@@ -48,8 +48,7 @@ than a reverse complement of the whole 600 bp construct.
|
|
| 48 |
MPAC covers autosomes only; `from_pretrained` raises on chrX, chrY and anything else
|
| 49 |
with no held-out fold.
|
| 50 |
|
| 51 |
-
For allelic skew, pass matched reference and alternate contexts
|
| 52 |
-
upstream of the variant, the variant, 190 bp downstream. Each is tiled into eighteen
|
| 53 |
200 bp windows at stride 10 and averaged, reproducing the scheme behind the
|
| 54 |
published predictions.
|
| 55 |
|
|
@@ -58,9 +57,33 @@ out = ensemble.predict_skew(ref_contexts, alt_contexts, device="cuda")
|
|
| 58 |
out["skew"] # (n, 3), alt minus ref
|
| 59 |
```
|
| 60 |
|
| 61 |
-
Command-line
|
| 62 |
-
[
|
| 63 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 64 |
## Citation
|
| 65 |
|
| 66 |
```bibtex
|
|
|
|
| 48 |
MPAC covers autosomes only; `from_pretrained` raises on chrX, chrY and anything else
|
| 49 |
with no held-out fold.
|
| 50 |
|
| 51 |
+
For allelic skew, pass matched reference and alternate contexts. Each variant is tiled into eighteen
|
|
|
|
| 52 |
200 bp windows at stride 10 and averaged, reproducing the scheme behind the
|
| 53 |
published predictions.
|
| 54 |
|
|
|
|
| 57 |
out["skew"] # (n, 3), alt minus ref
|
| 58 |
```
|
| 59 |
|
| 60 |
+
Command-line tools for MPAC can be found at:
|
| 61 |
+
[Reilly-Lab-Yale/coda_mpac](https://github.com/Reilly-Lab-Yale/coda_mpac).
|
| 62 |
+
|
| 63 |
+
Which provide greater control of prediction parameters including (but not limited to):
|
| 64 |
+
window number
|
| 65 |
+
step size
|
| 66 |
+
strand reduction
|
| 67 |
+
insert/full plasmid reverse complement
|
| 68 |
+
window aggregation (average, max, min, etc.)
|
| 69 |
+
|
| 70 |
+
An example prediction for SNVs from a VCF can be found in the methods of the publication and below:
|
| 71 |
+
|
| 72 |
+
python vcf_predict.py --artifact_path {10X $MODEL} \ CHROMOSOME HOLDOUT MODELS TO ENSEMBLE
|
| 73 |
+
--vcf_file ${VCF} \ VARIANTS OF INTEREST
|
| 74 |
+
--fasta_file ${FASTA} \ REFERENCE GENOME
|
| 75 |
+
--output ${OUTPUT} \ OUTPUT PATH
|
| 76 |
+
--relative_start 9 \ START WINDOW
|
| 77 |
+
--relative_end 180 \ END WINDOW
|
| 78 |
+
--step_size 10 \ N WINDOWS
|
| 79 |
+
--strand_reduction mean \ REDUCTION METHOD OF FWD/REV STRAND PREDICTIONS
|
| 80 |
+
--window_reduction mean \ REDUCTION METHOD OF PREDICTION WINDOWS
|
| 81 |
+
|
| 82 |
+
And include specialized prediction methods for:
|
| 83 |
+
small indels (<= 10bp recommended) (vcf_predict_indel.py)
|
| 84 |
+
haplotypes (vcf_predict_haplotype.py)
|
| 85 |
+
|
| 86 |
+
Which can be run in the same way as vcf_predict.py but with modifications to handle sequence padding (vcf_predict_indel.py) or windowing (vcf_predict_haplotype.py)
|
| 87 |
## Citation
|
| 88 |
|
| 89 |
```bibtex
|