lvwerra's picture
lvwerra HF Staff
Update README.md
c6b2b4a verified
|
Raw History Blame Contribute Delete
20.3 kB
---
title: Carbon-A Database Explorer
emoji: 🧬
colorFrom: green
colorTo: blue
sdk: gradio
sdk_version: 6.6.0
python_version: '3.10'
app_file: app.py
pinned: false
short_description: Find genome annotations with on-demand bucket reads
---
Search a dated snapshot of all published annotation files in [HuggingFaceBio/genbank-annotations](https://huggingface.co/buckets/HuggingFaceBio/genbank-annotations) by assembly accession, contig accession, or the exact `source_key`. A local SQLite index locates records; probability data stays in the bucket and loads on demand. View strand-specific CDS probabilities, inspect provenance, and download original segment rows as Parquet.
Future features and open design decisions are tracked in the [feature backlog](FEATURE_BACKLOG.md).
## GenBank taxonomy
Both tabs open with a short description of the dataset, read from `content/about.md`. Edit that file to change the wording; no code change is needed. Text after a `<!-- more -->` line, if present, goes into a closed "How the annotations were made" disclosure.
The landing page opens with an interactive **tree of eukaryotic life**, using NCBI's current GenBank assembly summary and taxonomy. It is drawn as a Sankey diagram running **down the page**: each group is a horizontal bar split into annotated (green) and not-annotated (gray) assemblies, and flows carry both parts from a group into its lineages on the level below. Expanding a group adds a level underneath rather than a column to the right, so a full lineage — eukaryotes down to humans is about thirty levels — is one ordinary page scroll instead of an endless sideways one. The chart itself only scrolls on screens too narrow to keep the labels legible. Shortcut buttons above the tree jump straight to a lineage: animals, mammals, humans, birds, ray-finned fish, insects, jellyfish, green plants and fungi. Each group shows a silhouette, its scientific name and its English name; counts and percentages live in the tooltip, which appears over a group or over the flux arriving at it. The flux feeding the current selection stays highlighted, so the trunk you have walked reads against the lineages you passed over. Each group is labelled with its scientific name and, where one exists, a plain-English name underneath, so the Latin backbone of the tree (Opisthokonta, Eumetazoa, Ecdysozoa) stays readable without prior knowledge. The two views sit inside one curved panel, switched by a full-width toggle at the top: the **Genome Atlas** holds the tree and its lineage controls, and **Database** holds accession search, segment visualization, and downloads over the full published-file snapshot. A **Show database results** button carries the current group across, handing the Database tab one annotated assembly from the selected lineage; it says so plainly when a group has none. RefSeq exploration can be added beneath the tree.
The viewer is restricted to **Eukaryota (NCBI taxid 2759) and its descendants**, including search, breadcrumbs, and parent navigation. The model predicts eukaryotic annotations; bacteria, archaea, and viruses are outside this display. The current eukaryotic snapshot contains **70,395 assemblies**, of which **17,561 (24.9%)** have published annotations. The underlying SQLite snapshot retains the full NCBI inventory for refresh/provenance, while the interface uses only the eukaryotic subtree and its denominators. Assemblies with unresolved taxonomy cannot be assigned to that subtree and are excluded.
The inventory accepts `GCA_` accessions with `version_status=latest`. Each assembly contributes once to its taxon and each ancestor. These counts describe genome assemblies, not all GenBank nucleotide records. An assembly counts as having annotations when it occurs in the publication inventory; this includes partial assemblies and does not imply full base coverage. Exact accession versions are matched to the NCBI snapshot. Published versions outside that snapshot are excluded from percentages.
Click a group to open its direct lineages as a new level below; every ancestor level stays visible, so the full path from Eukaryota is always on screen. Clicking a group in an earlier level replaces the levels after it. Every ancestor is already on screen, so going back is a click on a level above; the shortcuts jump elsewhere entirely. Nodes support Enter/Space keyboard navigation. The opening view expands the actual animal/fungal ancestor, Opisthokonta. Each level shows six direct lineages, an expandable remainder, and the selected lineage if it falls outside the top six. Bar and flow widths follow assembly counts within a level; every level is rescaled to the full width. “Other lineages” is an explicitly labeled display aggregate, not an invented clade; its counts sum the remaining siblings. Repeated expansion reaches all siblings without truncating them. Search reaches individual taxa by scientific name, taxon ID, or English name, so "sponges", "jellyfish" or "house mouse" all work; matches appear as a list with the closest first and already selected, and Enter opens it. Level spacing does not encode time. Expanding nudges the page down only when the new level has fallen past the viewport.
English names come from `data/common_names.json`, built by `refresh_common_names.py` from NCBI Taxonomy's `genbank common name` and `common name` entries (about 9,200 taxa) plus a curated table of roughly 190 short glosses for the large unranked clades NCBI leaves unnamed. Glosses are plain-language summaries of the living members of a clade, not formal synonyms; the scientific name stays the primary label and a gloss that merely repeats it is dropped at build time. Rebuild it with `python refresh_common_names.py` once `.cache/coverage/taxdump.tar.gz` is present; the build fails if a curated taxid no longer matches its scientific name in the snapshot.
Each group also carries a clade silhouette from [PhyloPic](https://www.phylopic.org), fetched by `refresh_icons.py` into `data/icons/` with the map and credits in `data/icons.json`. The fetcher resolves each taxon's NCBI ID to a PhyloPic node and prefers that node's representative image, falling back to a clade search only when the representative one is not free to use. Both searches exclude NonCommercial and ShareAlike images: NC is unsettled for a Space run by a company, and SA would reach into the repository as soon as an icon were adapted. What ships is CC0, Public Domain Mark and CC BY, stored byte-for-byte as PhyloPic serves it, with every CC BY artist credited in the "Silhouette credits" panel. Icons are fetched for groups with at least 100 assemblies; smaller groups borrow their nearest illustrated ancestor at render time. Gradio serves them as static files through `gr.set_static_paths`, so each is cached once by the browser instead of riding along in every re-render.
The graphic uses server-rendered SVG/HTML with scoped styles in `atlas.css` and delegated click/keyboard handlers in `atlas.js`; no external graphics library or image assets are required. Include these files when deploying.
The annotation workspace shares the atlas palette and typography through `style.py` and `app.css`: warm panels, serif section headings, green controls, and matching tables, downloads, and disclosures. CDS plots use a solid green positive strand and a dashed warm-brown negative strand; binary labels use green. The fixed light palette stays consistent with the landing graphic under either OS color preference. Include both style files when deploying; Gradio 6 receives the theme and stylesheet in `launch()`.
The **Show database results** jump needs a taxon-to-accession table, which `refresh_assemblies.py` writes into `data/taxonomy.sqlite` as an `assemblies` table: one row per annotated eukaryotic assembly, about 33,600 of them. It reads the NCBI assembly summary and the existing `data/coverage.json`, and deliberately leaves the counts alone — rebuilding those here would re-read a newer summary and shift every total in the interface away from the published snapshot. Only annotated assemblies are stored, so the jump never offers an accession with nothing behind it.
### Refresh the taxonomy snapshot
```bash
python refresh_taxonomy.py --download
```
This downloads roughly 2 GB of **metadata only** to `.cache/coverage/`, then atomically rebuilds `data/taxonomy.sqlite`. To rebuild from already downloaded metadata, omit `--download`. No sequence or probability data is downloaded. The builder uses the Python standard library. The SQLite file contains taxon relationships, names, counts, and source URLs/checksums/dates; it is independent of `data/catalog.sqlite`.
Upload the new `data/taxonomy.sqlite` to the Space and restart/redeploy to activate it. The app reads the bundled snapshot and does not download NCBI metadata at startup. Refreshes are manual; the displayed dates distinguish the taxonomy build from the annotation inventory. Accession navigation from taxonomy and base-level completeness remain in the backlog.
### Refresh annotation coverage
```bash
python refresh_coverage.py
python refresh_taxonomy.py
```
Use a Hugging Face token with read access to both `HuggingFaceBio/genbank-annotations` and `HuggingFaceBio/GENERanno-packed-inference-input`. The first command lists current published objects, identifies assemblies, and writes `data/coverage.json`. It reads Parquet footers and assembly-accession columns only; sequences and probabilities stay in the buckets. By default 12 workers process shards; change this with `--workers`.
For completed shards, the inventory joins `_SUCCESS.json` to assembly IDs in the corresponding packed `sources.parquet`. This relies on the pipeline's strict unpack contract: the publication script runs with `--strict`, checks part counts, and writes its completion marker after publishing annotations. It also assumes the input shard namespace still identifies the input used by that marker; the older markers do not contain input content hashes. This is assembly-presence evidence, not an audit of every output base. Shards without completion markers, including older single-file outputs, are inventoried directly from annotation metadata.
Assembly IDs are deduplicated across contigs, segments, and shards. The inventory records source/marker hashes and a fingerprint of published objects. Metadata results cached under `.cache/coverage/shards/` are reused only when those fingerprints match. A failed shard aborts publication of the new inventory; rerunning resumes using cached successes. Bucket listing is a dated observation, not an atomic bucket transaction.
The second command rebuilds taxonomy counts using that inventory and the cached NCBI files. Use `--download` to also refresh NCBI metadata. Upload both `data/coverage.json` (inventory and provenance) and `data/taxonomy.sqlite` (runtime counts) to the Space and redeploy. This coverage-only refresh does not expand accession lookup. For a synchronized full refresh, use the full-index workflow below; it derives atlas assembly presence directly from indexed annotation metadata.
## Run locally
```bash
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python app.py
```
Open http://localhost:7860. Remote reads require a Hugging Face login (`hf auth login`) or an `HF_TOKEN` with bucket read access. For the original offline sample, run `GENBANK_DATA_MODE=sample python app.py`; that mode requires no token.
## Deploy to a Space
The preview Space is [HuggingFaceBio/genbank-annotation-explorer](https://huggingface.co/spaces/HuggingFaceBio/genbank-annotation-explorer), with private visibility.
To deploy or update it, upload the app code and documentation, including `data/catalog_snapshot.json` and `data/taxonomy.sqlite`; exclude `.cache/` and local environments. Store the full SQLite accession index under a versioned `indexes/accession/` path in the annotation bucket. The small `data/catalog_snapshot.json` pointer pins its path, content hash, size, and inventory fingerprint. The Space downloads this index into its local cache on startup (about 16 GB); subsequent searches use local SQLite. Do not upload the large `data/catalog.sqlite` to the Space repository. Set `HF_TOKEN` as a Space secret with read access to the bucket. Include `data/sample.parquet` and `data/manifest.json` if offline sample mode is desired. Credentials stay on the server and must never be committed. The README metadata configures the runtime, following the [Gradio Spaces documentation](https://huggingface.co/docs/hub/spaces-sdks-gradio). Keep the Space private for the initial preview.
## Accession index snapshot
The current metadata index covers **600,175 published Parquet objects**, **115,891,939 segment rows**, and **33,722 assembly versions**. It represents **26,312,134,513,090 row-bases**, summing segment lengths across stored rows; duplicate or overlapping publications are not deduplicated in this base total. The SQLite file is **15.97 GB**. The inventory was captured on 2026-10-05; later publications require another refresh.
The atlas counts exact accession-version matches to the existing GenBank reference snapshot: **33,594/70,395 eukaryotic assemblies (47.72%)**. This is assembly presence, including partial assemblies, rather than base completeness. Atlas assembly IDs now come directly from the same full annotation metadata index used by accession search. See `data/index_refresh.json` for refresh provenance and counts.
The catalog stores object paths/hashes once per file and shared assembly/run metadata once per context. Contig and assembly accession indexes support exact and versionless searches; source keys are reconstructed when canonical and stored separately otherwise. Sequence and probability columns remain in the bucket. The app loads at most 200 result rows per search and 200 browse rows at startup.
## Retrieval and storage
- SQLite contains indexed accession fields, shared metadata, object paths and hashes, row groups, and row positions. Probability arrays and sequence are excluded.
- The Space opens the index read-only. Search uses indexed accession fields, and results are capped at 200 with the total match count shown. Narrow to a contig accession when necessary; pagination is not yet implemented.
- Remote reads use `HfFileSystem` and PyArrow, fetching the selected row group with all source columns so downloads preserve the original record.
- An in-memory LRU cache holds at most 512 MiB of Arrow table buffers. Fills are serialized to avoid duplicate concurrent downloads. Temporary read/plot buffers require additional memory; this is not a total process memory limit.
- Row groups with more than 512 MiB of uncompressed Parquet data are rejected before decoding. Larger source groups will need smaller chunks or another retrieval path.
- On a cache miss, the object content hash is checked before and after reading. Changed objects are rejected with an index-refresh message. Cached data remains tied to the indexed snapshot. Objects are not versioned by this app; replacement of a source requires rebuilding the index.
- Older whole-contig records derive display coordinates `[0, aligned_bp_length)`. Their downloaded Parquet rows retain the original schema.
- Authentication and retrieval errors are shown separately from accession misses. A failed segment selection clears the previous plot and download.
## Rebuild the full remote index
```bash
python list_annotation_objects.py --output .cache/full-index/objects.jsonl
python build_snapshot.py --inventory .cache/full-index/objects.jsonl --output data/catalog.sqlite --parts 8 --workers-per-part 32
python coverage_from_index.py --index data/catalog.sqlite --output data/coverage.json
python refresh_taxonomy.py
```
The builder reads scalar metadata and Parquet footers, coalescing adjacent byte ranges; it does not decode sequence or probability arrays. An object/content-hash ledger and checkpoint database make the build resumable. Existing finalized indexes are reused on refresh, with changed/deleted objects removed and new objects indexed. Failed objects prevent publication. Building the full collection can take substantial time and should run outside the Space. Upload the resulting index to a new versioned bucket path, then publish its pointer, coverage inventory, and taxonomy database together after completion.
For larger builds, split a single inventory into disjoint partitions, build each to its own SQLite file, then use `merge_full_index.py --inventory ... --parts part-0.sqlite part-1.sqlite ... --output data/catalog.sqlite`. The merge checks that every inventoried object and row is present before creating the final search indexes. `build_index.py` remains available for a small development subset; it does not produce the full snapshot.
## Lookup behavior
- Exact versioned accessions match only that version; unknown versions do not fall back.
- Unversioned accessions resolve matching versions in the index. Input ignores case and surrounding whitespace.
- Assembly lookups return indexed contigs/segments for that assembly, subject to the result limit. Indexed coverage may be incomplete.
- A miss means “not in this dated published snapshot,” not “absent from GenBank”; files published after the inventory time need a refresh.
- Split contigs stay as separate segments, preserving their source coordinates and segment counts.
## Data and coordinates
The viewer offers **Probabilities** and **Binary labels** modes, with an adjustable threshold defaulting to **0.5**. Binary mode shows one combined CDS/background track: `max(P_positive, P_negative) > threshold`. Either strand can make a base CDS; exact ties are background. This combines the per-strand binary decisions with OR. Validation uses per-strand `argmax` between background and CDS and scores those strand labels; the combined viewer track is not a separate reproduction of those validation metrics. Stored float16 probabilities and inference-window averaging also prevent exact reconstruction of original evaluation logits.
The stepped track preserves exact labels using their change positions when the combined track has at most 20,000 transitions. For denser regions, a clearly labeled overview shows 1 when any base in a bin exceeds the threshold; zoom in by changing the start/end fields for exact labels. Thresholding happens before binning, so it never thresholds an averaged probability. Mode and threshold changes update the current region without resetting its coordinates. Probability mode retains the two strand tracks. Downloads continue to contain original probabilities.
The bucket contains model predictions: positive- and negative-strand per-base P(CDS), metadata, and sometimes sequence. It does not provide curated gene names or gene models. Coordinates are 0-based, end-exclusive. The initial viewport shows the entire selected segment; update the start/end fields to inspect a smaller region. Large windows are averaged into at most 1,200 bins per strand. Downloads retain original probabilities and source columns without downsampling.
`data/manifest.json` records the bucket object, content hash, source row indices, selection limits, and sample checksum. The default sample is limited to 24 complete segments and 2 million bases from one fungi chunk. It is for interface development, not a representative biological sample. Full per-base columns are loaded only for the selected segment; the accession index holds metadata only.
## Rebuild the sample
Use a Hugging Face login or `HF_TOKEN` with bucket read access. Never commit tokens.
```bash
pip install -r requirements-data.txt
hf auth login
python prepare_sample.py
```
The default source object is about 31 MB; automatic downloads are capped at 100 MB. Override `--source`, `--max-rows`, or `--max-bp` to choose another bounded sample. If the source already exists locally:
```bash
python prepare_sample.py --local-file /path/to/source.parquet
```
Source rows are selected intact, without truncating their probabilities. This command rebuilds only the optional offline sample, independently of the remote SQLite index.
## Checks
```bash
python -m unittest discover -s tests -v
```