Buckets:

cgeorgiaw's picture
|
download
raw
2.15 kB
# GenBank Annotation Output Schema
Published final annotations live under:
`annotations/<division>/<shard>.parquet`
Rows are one row per source contig.
| Column | Type | Purpose |
|---|---|---|
| `source_key` | string | Stable local key: `<assembly_accession>|<record_name>` |
| `division` | string | GenBank division/group used by the pipeline |
| `record_name` | string | FASTA/GenBank contig record ID |
| `assembly_accession` | string | GenBank assembly accession |
| `organism_name` | string | Assembly organism name from GenBank metadata |
| `taxid` | int64 | NCBI taxonomy ID |
| `genome_size` | int64 | Assembly genome size from GenBank metadata |
| `ftp_path` | string | NCBI FTP directory for the assembly; enough to recover the source FASTA path |
| `aligned_bp_length` | int64 | Number of bp represented by the probability arrays after 6-mer alignment |
| `annotation_run_id` | string | Pipeline run identifier used when unpacking/publishing annotations |
| `annotation_run_timestamp_utc` | string | UTC timestamp associated with the annotation run |
| `annotation_git_commit` | string | Git commit of this pipeline when annotations were unpacked/submitted |
| `manifest_id` | string | Manifest/run identifier for the source assembly selection |
| `sequence` | string | Source contig sequence prefix represented by the probability arrays. Optional; omitted for Hub-scale outputs by default. |
| `pred_prob_positive_strand_cds` | list<float16> | Per-bp P(CDS) on positive strand, aligned to source sequence prefix |
| `pred_prob_negative_strand_cds` | list<float16> | Per-bp P(CDS) on negative strand, aligned to source sequence prefix |
Intentionally omitted from Hub-scale final outputs:
- raw packed blocks and raw packed-block probability outputs.
- `sequence`, unless the run was unpacked without `--omit-sequence`.
Mapping back to source FASTA:
1. Use `ftp_path` to locate the assembly directory.
2. The FASTA file basename is the final path component of `ftp_path` plus `_genomic.fna.gz`.
3. Use `record_name` to find the contig in that FASTA.
4. Probability arrays cover the first `aligned_bp_length` bp of that contig.

Xet Storage Details

Size:
2.15 kB
·
Xet hash:
5603bf46c22e2bfc9a7efd20d34ff70039070fbee427e7318d2431eb429a4af2

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.