TIGR-Tas sgRNA activity prediction
Default pretrained TIGR-Tas model for CRISPRa sgRNA activity ranking. Model ID: Cuiting0906/TIGR-Tas_sgRNA_pred. Code and installation.
Prediction
Predict one sgRNA
After installing the tigr and ag environments, provide one guide, its hg38 location and strand, the gene strand, and pre-edit ATAC bigWigs:
python scripts/predict_guide.py \
--guide ACGTACGTACGTACGTAC \
--chrom chr1 --target-start 123456789 \
--gene-strand + --guide-strand - \
--atac cell_rep1.bw cell_rep2.bw \
--output outputs/predictions.csv
The script creates the internal manifest and features automatically. It downloads model.pt and returns one activity score. Higher scores indicate higher predicted activity. For offline prediction, download this file and pass --checkpoint /path/to/model.pt.
Batch prediction from CSV
For a CSV batch from one target cell or ATAC condition, use:
python scripts/predict_batch.py --input guides.csv --atac cell_rep1.bw cell_rep2.bw --output outputs/batch_predictions.csv
The required CSV columns are guide, chrom, target_start, gene_strand, and guide_strand; sample_id and gene are optional. Place the default hg38 FASTA and AlphaGenome all-folds model at the repository paths documented in the code README.
Inputs and architecture
The model has 4,991,137 trainable parameters. An 18-nt guide and Watson/Crick direction are encoded alongside three independent context branches:
- Coarse frozen AlphaGenome embeddings:
[N, 32, 3072], broader DNA regulatory context. - Fine frozen AlphaGenome embeddings:
[N, 256, 1536], local sequence context. - Pre-edit ATAC features:
[N, 256, 2], binned accessibility and validity.
Each branch uses independent guide-query cross-attention before late fusion. AlphaGenome transfers sequence-derived regulatory information learned from functional-genomics data; measured ATAC adds cell-specific context. AlphaGenome weights and feature caches are not contained in this checkpoint.
Training and limitations
The fixed recipe uses 10% positive replacement sampling, AdamW learning rate 1e-4, weight decay 0, global dropout 0.1, fine Transformer dropout 0.05, batch size 32 and seven epochs of binary cross-entropy training.
This model was selected during retrospective development and refit on all 2,102 labeled guides from 13 genes (81 positives). Architecture and training choices were informed by the available evaluation data. There is no independent held-out evaluation of this released checkpoint. Historical branch comparisons in the code repository used an earlier recipe and are not performance estimates for this model.
Scores are not calibrated experimental success probabilities. Validate on new loci and biological replicates before prospective use. The model is intended for CRISPRa research; CRISPRi and other systems have not been validated. Original experimental data and feature caches are not publicly deposited with this release.
File and terms
model.pt contains only the PyTorch state dictionary, architecture settings and fixed training metadata. The provided code loads it with weights_only=True and strict parameter matching.
The source code is MIT-licensed. These weights are a separate research artifact: use and redistribution require author permission and compliance with applicable source-data and AlphaGenome terms. No additional open-source license is granted to the weights here. Contact the author through the code repository.