Buckets:

GenerTeam's picture
|
download
raw
3.05 kB

Carbon-A training data

This repository contains the Parquet training shards and the evaluation files used by Carbon-A.

Bucket layout

Carbon-A-training-data/
├── train/
│   ├── vertebrate_mammalian/
│   ├── vertebrate_other/
│   ├── invertebrate/
│   ├── plant/
│   ├── fungi/
│   └── protozoa/
└── eval/
    ├── seen/
    ├── unseen/
    └── nonstandard_code/

The bucket contains 4,561 training shards. These are already the sliding-window-augmented training examples: each Parquet row is a genomic window of 98,304 bp. The evaluation design contains 28 seen accessions, 14 unseen accessions, and one nonstandard_code accession. The domain breakdown is:

Domain Seen Unseen Nonstandard code Total
Mammals 5 4 0 9
Other vertebrates 3 5 0 8
Invertebrates 5 5 0 10
Plants 5 0 0 5
Fungi 5 0 0 5
Protozoa 5 0 1 6
Total 28 14 1 43

Training-data construction

The training corpus uses paired RefSeq GenBank Flat Files and FASTA records from RefSeq assemblies. A record is eligible when its genomic sequence is at least 98,304 bp. The six domains are mammals, other vertebrates, invertebrates, plants, fungi, and protozoa.

Each example contains a genomic window and two nucleotide-resolution binary target tracks, one per strand. CDS intervals from annotated isoforms are unioned on each strand; opposite-strand overlaps remain separate. Windows without CDS labels are retained, so the data include both coding and genomic background sequence. The model uses non-overlapping 6-mer tokens while the targets remain nucleotide aligned.

Domain-specific overlapping strides increase sampling of the rarer domains. A random 0--5 bp start offset changes both the crop boundary and the phase of the 6-mer grid.

Domain Species RefSeq accessions Coding genes (M) Source bases (B) Coding bases (%) Stride (bp) Windows (M) Augmented bases (B)
Mammals 236 242 4.89 974.99 0.86 98,304 9.96 978.82
Other vertebrates 443 444 9.06 1,205.34 1.37 98,304 12.32 1,211.47
Invertebrates 428 429 6.54 493.24 2.21 65,536 7.53 740.37
Plants 182 182 6.95 335.42 2.57 49,152 6.79 667.60
Fungi 644 644 6.43 19.99 45.17 4,096 4.39 431.61
Protozoa 114 114 1.13 4.03 42.41 1,024 3.28 322.30
Total 2,047 2,055 34.98 3,033.00 1.82 — 44.27 4,352.18

The raw, pre-augmentation source contains 3.033 Tbp. Sliding-window augmentation expands the represented sequence to 4.35 Tbp across 44.27 million windows, which are the examples stored under train/. The shorter fungi and protozoa strides compensate for their small share of the raw base count.

License

Apache 2.0.

Xet Storage Details

Size:
3.05 kB
·
Xet hash:
bf42179e08be86811c7cf40ed38a081d891199dc2d3986affae24603de3b263d

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.