|
Download source/README.md from rustambekurokov/bgc-setnet: direct link, hf CLI and curl.
- Browser
- Download file 12.2 kB
-
https://huggingface.co/rustambekurokov/bgc-setnet/resolve/main/source/README.md
- Command line
-
hf download hf://rustambekurokov/bgc-setnet/source/README.md
-
curl -L -o README.md https://huggingface.co/rustambekurokov/bgc-setnet/resolve/main/source/README.md
12.2 kB
| # BGC retrieval rebuild | |
| This directory is an isolated, leakage-free replacement for the legacy experiment. | |
| It reads legacy inputs but never rewrites them. The primary development task is | |
| within-species retrieval of held-out MIBiG-reference groups in the | |
| *Streptomyces griseus* atlas. Curated within-species product labels are still | |
| required for a publication-grade biological claim. | |
| ## Scientific invariants | |
| 1. Model inputs are restricted to frozen ESM-2 gene embeddings and relative gene | |
| position. Alignment statistics and upstream product/activity predictions are | |
| forbidden. | |
| 2. Product/reference groups are split before pairs or batches are constructed. | |
| Test groups are excluded from pretraining, supervised training, tuning, and | |
| checkpoint selection. | |
| 3. Atlas-to-MIBiG assignments are silver training labels. Curated *S. griseus* | |
| BGC–product mappings are required for the primary biological result; broad | |
| multi-genus benchmarking is secondary stress evidence only. | |
| 4. Ensemble weights are selected on validation once. Tied baseline scores use | |
| expected metrics under uniform ordering within each tie group. | |
| 5. Every checkpoint and metric file records configuration, split, input, code, | |
| and environment fingerprints. | |
| ## Layout and workflow | |
| - `configs/main.yaml`: canonical one-GPU campaign configuration. | |
| - `src/bgc_retrieval`: data contracts, model, training, evaluation, and audit CLI. | |
| - `data/manifests`: tracked hashes, split assignments, and exclusion reports. | |
| - `artifacts/<run_id>`: immutable run outputs; never overwrite an existing run. | |
| From this directory, create a modern environment and install the package: | |
| ```bash | |
| python -m venv .venv | |
| .venv/bin/pip install -e '.[dev]' | |
| bgc-rebuild audit --config configs/main.yaml | |
| bgc-rebuild build-splits --config configs/main.yaml | |
| bgc-rebuild train --config configs/main.yaml --stage phase1 | |
| bgc-rebuild train --config configs/main.yaml --stage phase2 | |
| bgc-rebuild evaluate --config configs/main.yaml | |
| ``` | |
| The legacy CSV has MIBiG-reference assignments but no curated product identity | |
| table. `build-splits` therefore creates a **silver development split** by reference | |
| ID and refuses to call it product-level. The primary development endpoint is | |
| held-out, family-aggregated Recall@50 within the *S. griseus* atlas. The staged | |
| multi-genus benchmark measures out-of-distribution correlation between BGC | |
| embedding similarity and product-structure Tanimoto; it is not the paper's | |
| species-scoped primary endpoint. Exact-product retrieval remains secondary because | |
| most exact-product groups are singletons. | |
| Prepare and run the secondary stress endpoint with: | |
| ~~~bash | |
| python scripts/prepare_external_benchmark.py | |
| python scripts/embed_external_proteins.py --resume | |
| python scripts/evaluate_external.py \ | |
| --checkpoint artifacts/<training-run>/phase2_best.pt \ | |
| --run-id <external-evaluation-run> | |
| ~~~ | |
| `scripts/dgx_campaign.sh CAMPAIGN_TAG` runs the complete internal and external | |
| campaign on a CUDA host. It reuses the host's CUDA-enabled PyTorch installation, | |
| refuses to overwrite immutable run directories, and accepts | |
| `EXTERNAL_BOOTSTRAP_SAMPLES` for a shorter smoke campaign. The default is 1,000. | |
| ## 2026-08-11 scope correction and residual redesign report | |
| This section is the authoritative status for the rebuild and supersedes any earlier | |
| wording that treats the multi-genus benchmark as the primary paper endpoint. | |
| ### Authoritative biological scope | |
| The project is restricted to **within-species retrieval in _Streptomyces | |
| griseus_**: given known members of a held-out compound/reference family, rank | |
| other BGCs from the _S. griseus_ atlas. Generalization to other species or genera | |
| is not a primary claim. | |
| The legacy atlas contains 6,953 BGCs from 182 deduplicated genome accessions. Of | |
| those BGCs, 5,325 have silver MIBiG-reference assignments used by the rebuild. | |
| The frozen group-disjoint split contains: | |
| | Split | BGCs | reference groups | | |
| |---|---:|---:| | |
| | train | 4,211 | 98 | | |
| | validation | 676 | 21 | | |
| | test | 438 | 21 | | |
| Only 16 test groups contain more than the four required reference BGCs and | |
| therefore contribute to the reported retrieval summaries. The atlas itself has a | |
| `genome_id` column but no organism/species column. Its _S. griseus_ identity is | |
| currently a dataset-provenance assertion, not a machine-verifiable row-level | |
| contract. A taxonomy manifest for all 182 assemblies is required before | |
| submission. | |
| ### Why the previous analysis went wrong | |
| 1. **The scientific task drifted.** The rebuild promoted the released | |
| BGC-clustering benchmark to the primary endpoint even though the paper studies | |
| only _S. griseus_. That external sequence set contains 1,343 BGCs spanning | |
| hundreds of taxa: the unambiguous metadata contain 239 genera, and only 11 | |
| sequence records mention _S. griseus_. The resulting 1,270-BGC evaluation was | |
| principally an out-of-distribution, multi-genus test. It is useful as a stress | |
| test but does not answer the paper's within-species question. | |
| 2. **The labels do not constitute gold product identity.** The atlas | |
| `compound_family` values were produced from MIBiG alignment evidence. The | |
| rebuild correctly calls them silver `mibig_reference_id` groups, but the | |
| legacy manuscript sometimes described them as known compound families or | |
| exact products. The leakage-free _S. griseus_ result is therefore development | |
| evidence and is marked `publication_eligible: false`. | |
| 3. **The original encoder destroyed a strong baseline representation.** SetNet | |
| projected every 1,280-dimensional ESM gene vector into 256 dimensions, passed | |
| it through two induced-attention blocks, and pooled to another 256-dimensional | |
| vector. It had 2,272,512 trainable parameters and no raw-ESM residual path. | |
| This was excessive for only 98 training groups and allowed optimization to | |
| overwrite useful pretrained geometry. | |
| 4. **The contrastive target was mis-specified.** Supervised contrastive loss | |
| treated every different silver group as an equally strong negative. Related | |
| metabolites and noisy MIBiG assignments can therefore become false negatives. | |
| The loss optimized discrete silver-group separation, not the biological notion | |
| of alternative routes to the same or related chemistry. | |
| 5. **Checkpoint selection and the broad stress endpoint were misaligned.** | |
| Checkpoints were selected by silver validation Recall@50, while the promoted | |
| external endpoint was Spearman correlation with continuous product-structure | |
| Tanimoto similarity. Across the five original SetNet seeds, better validation | |
| recall did not reliably imply better external structural correlation. | |
| 6. **Legacy model selection overstated evidence.** Earlier work selected ensemble | |
| weights on the evaluated data, defined "hard" cases using weakness of the Pfam | |
| baseline itself, and treated many repeated retrieval draws as if they were | |
| independent observations. The rebuild fixed these issues with validation-only | |
| selection, expected tie-aware metrics, family aggregation, paired | |
| family-level tests, and Holm correction. | |
| 7. **The effective test size is small.** Although the atlas contains thousands of | |
| BGCs, uncertainty is governed by 16 eligible held-out groups, not by BGC-pair | |
| counts or 100 repeated draws per group. This produces wide intervals and makes | |
| small metric changes unconvincing. | |
| The prior multi-genus five-seed result remains recorded as a secondary | |
| out-of-distribution stress test: SetNet Spearman was | |
| `0.088884 ± 0.023216`, raw ESM was `0.114198`, and BiG-SCAPE was | |
| `0.261492`. Those numbers must not be used to accept or reject the | |
| _S. griseus_-scoped paper. | |
| ### Residual Gene Weighting v1 | |
| The executed redesign uses only frozen ESM embeddings and relative gene position: | |
| - a 397,057-parameter gene-importance gate, 5.7 times smaller than SetNet; | |
| - zero-initialized gate logits, making the starting representation exactly raw | |
| ESM mean pooling; | |
| - weighted pooling in the original 1,280-dimensional ESM space instead of a | |
| destructive 256-dimensional bottleneck; | |
| - soft Pfam/family relational distillation so related non-identical groups are | |
| not automatically hard negatives; | |
| - a frozen raw-ESM residual score with validation candidates including | |
| `alpha=0`; | |
| - a Pfam hybrid grid including pure Pfam; | |
| - checkpoint and mixture selection on validation only; | |
| - five deterministic seeds plus a five-checkpoint score ensemble. | |
| The residual construction guarantees: | |
| ```text | |
| cosine(residual_i, residual_j) | |
| = (1 - alpha) * cosine(raw_i, raw_j) | |
| + alpha * cosine(weighted_i, weighted_j) | |
| ``` | |
| Thus `alpha=0` reproduces raw ESM exactly. | |
| ### Executed DGX result | |
| All 26 tests passed before execution and again after artifact retrieval using | |
| the DGX environment's standard-library test runner. Training seeds 20260810–20260814 each | |
| early-stopped after 13 epochs. Every seed selected epoch 0 as its best checkpoint; | |
| best validation Recall@50 was approximately `0.4656–0.4658`, while later | |
| training reduced it to `0.4068–0.4346`. | |
| All five individual evaluations and the five-checkpoint ensemble independently | |
| selected: | |
| - learned residual weight: `alpha=0.0`; | |
| - hybrid learned-branch weight: `alpha=0.0`; | |
| - raw-ESM weight within the raw/Pfam hybrid: `beta=0.4`. | |
| Validation therefore rejected the trained gene-weighting branch. The fixed | |
| _S. griseus_ test results were: | |
| | Method | Recall@50 | MRR | MAP | nDCG@50 | | |
| |---|---:|---:|---:|---:| | |
| | raw ESM mean | 0.792383 | 0.253128 | 0.724006 | 0.805809 | | |
| | trained weighted ESM, five-seed mean | 0.792284 | 0.252970 | 0.723550 | 0.805465 | | |
| | Pfam Jaccard max | 0.877924 | 0.305931 | 0.847764 | 0.903320 | | |
| | validation-selected raw/Pfam hybrid | 0.878192 | 0.307223 | 0.849608 | 0.905321 | | |
| The trained weighted representation was slightly worse than raw ESM. The | |
| validation-selected residual correctly fell back to raw ESM. The raw/Pfam hybrid | |
| improved Recall@50 over Pfam by only `0.000268`; the paired family-level | |
| Holm-adjusted p-value was `0.498`. This is not a meaningful or statistically | |
| supported improvement. | |
| **Conclusion:** Residual Gene Weighting v1 successfully prevents degradation but | |
| does not rescue the learned-method claim. The current evidence supports Pfam | |
| Jaccard as the strongest _S. griseus_ method. Further model tuning on the same 16 | |
| test groups would be test-set overfitting. | |
| ### Reproducibility | |
| Canonical configuration and entry points: | |
| ```bash | |
| python scripts/train_residual.py \ | |
| --config configs/residual_griseus.yaml \ | |
| --run-id paper-v2-residual-seed-20260810 | |
| python scripts/evaluate_residual.py \ | |
| --config configs/residual_griseus.yaml \ | |
| --checkpoint artifacts/paper-v2-residual-seed-20260810/residual_best.pt \ | |
| --run-id paper-v2-residual-seed-20260810-evaluation | |
| python scripts/analyze_residual_campaign.py \ | |
| --seed 20260810 --seed 20260811 --seed 20260812 \ | |
| --seed 20260813 --seed 20260814 \ | |
| --ensemble-run paper-v2-residual-five-seed-ensemble-evaluation \ | |
| --output-dir artifacts/paper-v2-residual-analysis | |
| ``` | |
| Primary result files: | |
| - `artifacts/paper-v2-residual-analysis/aggregate_summary.csv` | |
| - `artifacts/paper-v2-residual-analysis/paired_comparisons.csv` | |
| - `artifacts/paper-v2-residual-analysis/analysis_metadata.json` | |
| - `artifacts/paper-v2-residual-five-seed-ensemble-evaluation/summary.csv` | |
| - `artifacts/paper-v2-residual-campaign.log` | |
| - `artifacts/paper-v2-residual-sha256.txt` (58 files; manifest SHA-256 | |
| `2cee5df67735490d886b94e642375bb771785bb7547ab02ed90ad000f4e511ee`) | |
| ### Required next scientific step | |
| Do not tune another architecture against the same test groups. First: | |
| 1. create a checked taxonomy manifest proving the 182 assemblies are | |
| _S. griseus_; | |
| 2. curate or obtain gold BGC–product labels within those genomes; | |
| 3. version a new untouched family-level holdout or use fully nested group | |
| cross-validation; | |
| 4. pre-register the primary metric and baseline comparisons; | |
| 5. only then train a chemistry-aligned residual model. | |
| ### Documentation rule | |
| After every future request that changes code, data, experiments, results, scope, | |
| or scientific interpretation, update this README in the same request with a dated | |
| summary and artifact paths. | |