Update README.md
#2
by buttsjoc - opened
README.md
CHANGED
|
@@ -69,21 +69,70 @@ Which provide greater control of prediction parameters including (but not limite
|
|
| 69 |
|
| 70 |
An example prediction for SNVs from a VCF can be found in the methods of the publication and below:
|
| 71 |
|
| 72 |
-
|
| 73 |
-
|
| 74 |
-
|
| 75 |
-
--
|
| 76 |
-
|
| 77 |
-
|
| 78 |
-
|
| 79 |
-
|
| 80 |
-
|
| 81 |
-
|
| 82 |
-
|
| 83 |
-
|
| 84 |
-
|
| 85 |
-
|
| 86 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 87 |
## Citation
|
| 88 |
|
| 89 |
```bibtex
|
|
|
|
| 69 |
|
| 70 |
An example prediction for SNVs from a VCF can be found in the methods of the publication and below:
|
| 71 |
|
| 72 |
+
To generate MPAC predictions for SNVs from the command line, first define your variants in a VCF-like TSV (no VCF header required) with the following tab-separated columns:
|
| 73 |
+
|
| 74 |
+
| Column | Description |
|
| 75 |
+
| --- | --- |
|
| 76 |
+
| `CHROM` | Chromosome |
|
| 77 |
+
| `POS` | Position |
|
| 78 |
+
| `ID` | Variant ID (can be blank or `.` if none) |
|
| 79 |
+
| `REF` | Reference allele |
|
| 80 |
+
| `ALT` | Alternate allele |
|
| 81 |
+
| `QUAL` | Quality (can be `.` if null) |
|
| 82 |
+
| `FILTER` | Filter status (can be `.` if null) |
|
| 83 |
+
| `INFO` | Can be `.` if null; predictions populate this column in the output |
|
| 84 |
+
|
| 85 |
+
### SNVs
|
| 86 |
+
|
| 87 |
+
SNV predictions are handled with `vcf_predict.py`:
|
| 88 |
+
|
| 89 |
+
```bash
|
| 90 |
+
python vcf_predict.py \
|
| 91 |
+
--artifact_path {10X $MODEL} \
|
| 92 |
+
--vcf_file ${VCF} \
|
| 93 |
+
--fasta_file ${FASTA} \
|
| 94 |
+
--output ${OUTPUT} \
|
| 95 |
+
--relative_start 9 \
|
| 96 |
+
--relative_end 180 \
|
| 97 |
+
--step_size 10 \
|
| 98 |
+
--strand_reduction mean \
|
| 99 |
+
--window_reduction mean \
|
| 100 |
+
--feature_ids K562 HepG2 SKNSH
|
| 101 |
+
```
|
| 102 |
+
|
| 103 |
+
| Argument | Description |
|
| 104 |
+
| --- | --- |
|
| 105 |
+
| `--artifact_path` | Chromosome holdout models to ensemble |
|
| 106 |
+
| `--vcf_file` | Variants of interest |
|
| 107 |
+
| `--fasta_file` | Reference genome |
|
| 108 |
+
| `--output` | Output path |
|
| 109 |
+
| `--relative_start` | First window variant position |
|
| 110 |
+
| `--relative_end` | Final window variant position |
|
| 111 |
+
| `--step_size` | Window step size |
|
| 112 |
+
| `--strand_reduction` | Reduction method for fwd/rev strand predictions |
|
| 113 |
+
| `--window_reduction` | Reduction method for window predictions |
|
| 114 |
+
| `--feature_ids` | Labels for predictions in the output |
|
| 115 |
+
|
| 116 |
+
Additonal arguments can be found in vcf_predict.py
|
| 117 |
+
|
| 118 |
+
### Small indels
|
| 119 |
+
|
| 120 |
+
Small indels (≤ 10 bp recommended) are handled the same way with `vcf_predict_indel.py`, with modification to the plasmid sequence padding loop.
|
| 121 |
+
|
| 122 |
+
### Haplotypes
|
| 123 |
+
|
| 124 |
+
Haplotype predictions are handled similarly with `vcf_predict_haplotype.py`, with modification to the windowing loop to ensure all windows contain the desired variants.
|
| 125 |
+
|
| 126 |
+
### Output
|
| 127 |
+
|
| 128 |
+
Output is a VCF-like TSV with predictions occupying the `INFO` column, see below:
|
| 129 |
+
|
| 130 |
+
### Example output
|
| 131 |
+
|
| 132 |
+
```
|
| 133 |
+
CHROM POS ID REF ALT INFO
|
| 134 |
+
chr22 11121724 COSV106573183 A G K562__ref=0.3535322;HepG2__ref=0.29076257;...
|
| 135 |
+
```
|
| 136 |
## Citation
|
| 137 |
|
| 138 |
```bibtex
|