Buckets:

cgeorgiaw's picture
|
download
raw
2.99 kB

GenBank Annotation Output Schema

Published final annotations live under:

Single-file shard output:

annotations/<division>/<shard>.parquet

Chunked shard output:

annotations/<division>/<shard>/<shard>.chunk-00000.parquet

Rows are one row per source contig segment. Most contigs have one segment; very large contigs are split to keep each Parquet row below writer limits.

Column Type Purpose
source_key string Stable local key: `
division string GenBank division/group used by the pipeline
record_name string FASTA/GenBank contig record ID
assembly_accession string GenBank assembly accession
organism_name string Assembly organism name from GenBank metadata
taxid int64 NCBI taxonomy ID
genome_size int64 Assembly genome size from GenBank metadata
ftp_path string NCBI FTP directory for the assembly; enough to recover the source FASTA path
aligned_bp_length int64 Total number of bp represented for the source contig after 6-mer alignment
segment_start_bp int64 0-based inclusive start coordinate of this row within the aligned source contig
segment_end_bp int64 0-based exclusive end coordinate of this row within the aligned source contig
segment_bp_length int64 Number of aligned bp represented by this row
segment_index int32 0-based segment index for this source contig
segment_count int32 Number of output rows needed to represent this source contig
annotation_run_id string Pipeline run identifier used when unpacking/publishing annotations
annotation_run_timestamp_utc string UTC timestamp associated with the annotation run
annotation_git_commit string Git commit of this pipeline when annotations were unpacked/submitted
manifest_id string Manifest/run identifier for the source assembly selection
sequence string Source contig sequence prefix represented by the probability arrays. Optional; omitted for Hub-scale outputs by default.
pred_prob_positive_strand_cds list Per-bp P(CDS) on positive strand for segment_start_bp:segment_end_bp
pred_prob_negative_strand_cds list Per-bp P(CDS) on negative strand for segment_start_bp:segment_end_bp

Intentionally omitted from Hub-scale final outputs:

  • raw packed blocks and raw packed-block probability outputs.
  • sequence, unless the run was unpacked without --omit-sequence.

Mapping back to source FASTA:

  1. Use ftp_path to locate the assembly directory.
  2. The FASTA file basename is the final path component of ftp_path plus _genomic.fna.gz.
  3. Use record_name to find the contig in that FASTA.
  4. Probability arrays cover segment_start_bp:segment_end_bp within the first aligned_bp_length bp of that contig. Concatenate rows ordered by segment_index to reconstruct full-contig arrays when segment_count > 1.

Xet Storage Details

Size:
2.99 kB
·
Xet hash:
5ee78ad70918ac7fc11c98241448291b33bba1a3b165deba74395911efa57578

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.