Files changed (1) hide show
  1. README.md +64 -15
README.md CHANGED
@@ -69,21 +69,70 @@ Which provide greater control of prediction parameters including (but not limite
69
 
70
  An example prediction for SNVs from a VCF can be found in the methods of the publication and below:
71
 
72
- python vcf_predict.py --artifact_path {10X $MODEL} \ CHROMOSOME HOLDOUT MODELS TO ENSEMBLE
73
- --vcf_file ${VCF} \ VARIANTS OF INTEREST
74
- --fasta_file ${FASTA} \ REFERENCE GENOME
75
- --output ${OUTPUT} \ OUTPUT PATH
76
- --relative_start 9 \ START WINDOW
77
- --relative_end 180 \ END WINDOW
78
- --step_size 10 \ N WINDOWS
79
- --strand_reduction mean \ REDUCTION METHOD OF FWD/REV STRAND PREDICTIONS
80
- --window_reduction mean \ REDUCTION METHOD OF PREDICTION WINDOWS
81
-
82
- And include specialized prediction methods for:
83
- small indels (<= 10bp recommended) (vcf_predict_indel.py)
84
- haplotypes (vcf_predict_haplotype.py)
85
-
86
- Which can be run in the same way as vcf_predict.py but with modifications to handle sequence padding (vcf_predict_indel.py) or windowing (vcf_predict_haplotype.py)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
87
  ## Citation
88
 
89
  ```bibtex
 
69
 
70
  An example prediction for SNVs from a VCF can be found in the methods of the publication and below:
71
 
72
+ To generate MPAC predictions for SNVs from the command line, first define your variants in a VCF-like TSV (no VCF header required) with the following tab-separated columns:
73
+
74
+ | Column | Description |
75
+ | --- | --- |
76
+ | `CHROM` | Chromosome |
77
+ | `POS` | Position |
78
+ | `ID` | Variant ID (can be blank or `.` if none) |
79
+ | `REF` | Reference allele |
80
+ | `ALT` | Alternate allele |
81
+ | `QUAL` | Quality (can be `.` if null) |
82
+ | `FILTER` | Filter status (can be `.` if null) |
83
+ | `INFO` | Can be `.` if null; predictions populate this column in the output |
84
+
85
+ ### SNVs
86
+
87
+ SNV predictions are handled with `vcf_predict.py`:
88
+
89
+ ```bash
90
+ python vcf_predict.py \
91
+ --artifact_path {10X $MODEL} \
92
+ --vcf_file ${VCF} \
93
+ --fasta_file ${FASTA} \
94
+ --output ${OUTPUT} \
95
+ --relative_start 9 \
96
+ --relative_end 180 \
97
+ --step_size 10 \
98
+ --strand_reduction mean \
99
+ --window_reduction mean \
100
+ --feature_ids K562 HepG2 SKNSH
101
+ ```
102
+
103
+ | Argument | Description |
104
+ | --- | --- |
105
+ | `--artifact_path` | Chromosome holdout models to ensemble |
106
+ | `--vcf_file` | Variants of interest |
107
+ | `--fasta_file` | Reference genome |
108
+ | `--output` | Output path |
109
+ | `--relative_start` | First window variant position |
110
+ | `--relative_end` | Final window variant position |
111
+ | `--step_size` | Window step size |
112
+ | `--strand_reduction` | Reduction method for fwd/rev strand predictions |
113
+ | `--window_reduction` | Reduction method for window predictions |
114
+ | `--feature_ids` | Labels for predictions in the output |
115
+
116
+ Additonal arguments can be found in vcf_predict.py
117
+
118
+ ### Small indels
119
+
120
+ Small indels (≤ 10 bp recommended) are handled the same way with `vcf_predict_indel.py`, with modification to the plasmid sequence padding loop.
121
+
122
+ ### Haplotypes
123
+
124
+ Haplotype predictions are handled similarly with `vcf_predict_haplotype.py`, with modification to the windowing loop to ensure all windows contain the desired variants.
125
+
126
+ ### Output
127
+
128
+ Output is a VCF-like TSV with predictions occupying the `INFO` column, see below:
129
+
130
+ ### Example output
131
+
132
+ ```
133
+ CHROM POS ID REF ALT INFO
134
+ chr22 11121724 COSV106573183 A G K562__ref=0.3535322;HepG2__ref=0.29076257;...
135
+ ```
136
  ## Citation
137
 
138
  ```bibtex