Buckets:
Carbon-A training data
This repository contains the Parquet training shards and the evaluation files used by Carbon-A.
Bucket layout
Carbon-A-training-data/
├── train/
│ ├── vertebrate_mammalian/
│ ├── vertebrate_other/
│ ├── invertebrate/
│ ├── plant/
│ ├── fungi/
│ └── protozoa/
└── eval/
├── seen/
├── unseen/
└── nonstandard_code/
The bucket contains 4,561 training shards. These are already the
sliding-window-augmented training examples: each Parquet row is a genomic
window of 98,304 bp. The evaluation design contains 28
seen accessions, 14 unseen accessions, and one nonstandard_code
accession. The domain breakdown is:
| Domain | Seen | Unseen | Nonstandard code | Total |
|---|---|---|---|---|
| Mammals | 5 | 4 | 0 | 9 |
| Other vertebrates | 3 | 5 | 0 | 8 |
| Invertebrates | 5 | 5 | 0 | 10 |
| Plants | 5 | 0 | 0 | 5 |
| Fungi | 5 | 0 | 0 | 5 |
| Protozoa | 5 | 0 | 1 | 6 |
| Total | 28 | 14 | 1 | 43 |
Training-data construction
The training corpus uses paired RefSeq GenBank Flat Files and FASTA records from RefSeq assemblies. A record is eligible when its genomic sequence is at least 98,304 bp. The six domains are mammals, other vertebrates, invertebrates, plants, fungi, and protozoa.
Each example contains a genomic window and two nucleotide-resolution binary target tracks, one per strand. CDS intervals from annotated isoforms are unioned on each strand; opposite-strand overlaps remain separate. Windows without CDS labels are retained, so the data include both coding and genomic background sequence. The model uses non-overlapping 6-mer tokens while the targets remain nucleotide aligned.
Domain-specific overlapping strides increase sampling of the rarer domains. A random 0--5 bp start offset changes both the crop boundary and the phase of the 6-mer grid.
| Domain | Species | RefSeq accessions | Coding genes (M) | Source bases (B) | Coding bases (%) | Stride (bp) | Windows (M) | Augmented bases (B) |
|---|---|---|---|---|---|---|---|---|
| Mammals | 236 | 242 | 4.89 | 974.99 | 0.86 | 98,304 | 9.96 | 978.82 |
| Other vertebrates | 443 | 444 | 9.06 | 1,205.34 | 1.37 | 98,304 | 12.32 | 1,211.47 |
| Invertebrates | 428 | 429 | 6.54 | 493.24 | 2.21 | 65,536 | 7.53 | 740.37 |
| Plants | 182 | 182 | 6.95 | 335.42 | 2.57 | 49,152 | 6.79 | 667.60 |
| Fungi | 644 | 644 | 6.43 | 19.99 | 45.17 | 4,096 | 4.39 | 431.61 |
| Protozoa | 114 | 114 | 1.13 | 4.03 | 42.41 | 1,024 | 3.28 | 322.30 |
| Total | 2,047 | 2,055 | 34.98 | 3,033.00 | 1.82 | — | 44.27 | 4,352.18 |
The raw, pre-augmentation source contains 3.033 Tbp. Sliding-window
augmentation expands the represented sequence to 4.35 Tbp across 44.27 million
windows, which are the examples stored under train/. The shorter fungi and
protozoa strides compensate for their small share of the raw base count.
License
Apache 2.0.
- Total size
- 2.01 TB
- Files
- 4,881
- Last updated
- Oct 9
- Pre-warmed CDN
- US EU US EU