Buckets:
| # GenBank Annotation Output Schema | |
| Published final annotations live under: | |
| `annotations/<division>/<shard>.parquet` | |
| Rows are one row per source contig. | |
| | Column | Type | Purpose | | |
| |---|---|---| | |
| | `source_key` | string | Stable local key: `<assembly_accession>|<record_name>` | | |
| | `division` | string | GenBank division/group used by the pipeline | | |
| | `record_name` | string | FASTA/GenBank contig record ID | | |
| | `assembly_accession` | string | GenBank assembly accession | | |
| | `organism_name` | string | Assembly organism name from GenBank metadata | | |
| | `taxid` | int64 | NCBI taxonomy ID | | |
| | `genome_size` | int64 | Assembly genome size from GenBank metadata | | |
| | `ftp_path` | string | NCBI FTP directory for the assembly; enough to recover the source FASTA path | | |
| | `aligned_bp_length` | int64 | Number of bp represented by the probability arrays after 6-mer alignment | | |
| | `annotation_run_id` | string | Pipeline run identifier used when unpacking/publishing annotations | | |
| | `annotation_run_timestamp_utc` | string | UTC timestamp associated with the annotation run | | |
| | `annotation_git_commit` | string | Git commit of this pipeline when annotations were unpacked/submitted | | |
| | `manifest_id` | string | Manifest/run identifier for the source assembly selection | | |
| | `sequence` | string | Source contig sequence prefix represented by the probability arrays. Optional; omitted for Hub-scale outputs by default. | | |
| | `pred_prob_positive_strand_cds` | list<float16> | Per-bp P(CDS) on positive strand, aligned to source sequence prefix | | |
| | `pred_prob_negative_strand_cds` | list<float16> | Per-bp P(CDS) on negative strand, aligned to source sequence prefix | | |
| Intentionally omitted from Hub-scale final outputs: | |
| - raw packed blocks and raw packed-block probability outputs. | |
| - `sequence`, unless the run was unpacked without `--omit-sequence`. | |
| Mapping back to source FASTA: | |
| 1. Use `ftp_path` to locate the assembly directory. | |
| 2. The FASTA file basename is the final path component of `ftp_path` plus `_genomic.fna.gz`. | |
| 3. Use `record_name` to find the contig in that FASTA. | |
| 4. Probability arrays cover the first `aligned_bp_length` bp of that contig. | |
Xet Storage Details
- Size:
- 2.15 kB
- Xet hash:
- 5603bf46c22e2bfc9a7efd20d34ff70039070fbee427e7318d2431eb429a4af2
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.