Buckets:

GenerTeam's picture
|
download
raw
3.05 kB
# Carbon-A training data
This repository contains the Parquet training shards and the evaluation
files used by Carbon-A.
## Bucket layout
```text
Carbon-A-training-data/
├── train/
│ ├── vertebrate_mammalian/
│ ├── vertebrate_other/
│ ├── invertebrate/
│ ├── plant/
│ ├── fungi/
│ └── protozoa/
└── eval/
├── seen/
├── unseen/
└── nonstandard_code/
```
The bucket contains 4,561 training shards. These are already the
sliding-window-augmented training examples: each Parquet row is a genomic
window of 98,304 bp. The evaluation design contains 28
`seen` accessions, 14 `unseen` accessions, and one `nonstandard_code`
accession. The domain breakdown is:
| Domain | Seen | Unseen | Nonstandard code | Total |
| --- | ---: | ---: | ---: | ---: |
| Mammals | 5 | 4 | 0 | 9 |
| Other vertebrates | 3 | 5 | 0 | 8 |
| Invertebrates | 5 | 5 | 0 | 10 |
| Plants | 5 | 0 | 0 | 5 |
| Fungi | 5 | 0 | 0 | 5 |
| Protozoa | 5 | 0 | 1 | 6 |
| **Total** | **28** | **14** | **1** | **43** |
## Training-data construction
The training corpus uses paired RefSeq GenBank Flat Files and FASTA records
from RefSeq assemblies. A record is eligible when its genomic
sequence is at least 98,304 bp. The six domains are mammals, other
vertebrates, invertebrates, plants, fungi, and protozoa.
Each example contains a genomic window and two nucleotide-resolution binary
target tracks, one per strand. CDS intervals from annotated isoforms are
unioned on each strand; opposite-strand overlaps remain separate. Windows
without CDS labels are retained, so the data include both coding and genomic
background sequence. The model uses non-overlapping 6-mer tokens while the
targets remain nucleotide aligned.
Domain-specific overlapping strides increase sampling of the rarer domains.
A random 0--5 bp start offset changes both the crop boundary and the phase of
the 6-mer grid.
| Domain | Species | RefSeq accessions | Coding genes (M) | Source bases (B) | Coding bases (%) | Stride (bp) | Windows (M) | Augmented bases (B) |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| Mammals | 236 | 242 | 4.89 | 974.99 | 0.86 | 98,304 | 9.96 | 978.82 |
| Other vertebrates | 443 | 444 | 9.06 | 1,205.34 | 1.37 | 98,304 | 12.32 | 1,211.47 |
| Invertebrates | 428 | 429 | 6.54 | 493.24 | 2.21 | 65,536 | 7.53 | 740.37 |
| Plants | 182 | 182 | 6.95 | 335.42 | 2.57 | 49,152 | 6.79 | 667.60 |
| Fungi | 644 | 644 | 6.43 | 19.99 | 45.17 | 4,096 | 4.39 | 431.61 |
| Protozoa | 114 | 114 | 1.13 | 4.03 | 42.41 | 1,024 | 3.28 | 322.30 |
| **Total** | **2,047** | **2,055** | **34.98** | **3,033.00** | **1.82** | — | **44.27** | **4,352.18** |
The raw, pre-augmentation source contains 3.033 Tbp. Sliding-window
augmentation expands the represented sequence to 4.35 Tbp across 44.27 million
windows, which are the examples stored under `train/`. The shorter fungi and
protozoa strides compensate for their small share of the raw base count.
## License
Apache 2.0.

Xet Storage Details

Size:
3.05 kB
·
Xet hash:
bf42179e08be86811c7cf40ed38a081d891199dc2d3986affae24603de3b263d

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.