stevenweaver commited on
Commit
eb8c4cc
·
verified ·
1 Parent(s): a78d065

Model card: fix rho inconsistency (base vs viral eval sets), remove em-dashes, clarify variant labeling

Browse files

The viral table is now framed as a controlled base-vs-fine-tuned comparison on a separate viral split, explicitly not comparable to the literature-benchmark aggregate; the mammal 'forgetting check' row no longer shows a misleading raw rho that clashed with the checkpoint's 0.5974. Em-dashes removed; each rho now clearly attributed to base or fine-tuned.

Files changed (1) hide show
  1. README.md +34 -20
README.md CHANGED
@@ -130,38 +130,52 @@ datasets from 9 independent literature studies (43,302 codons, up to 476 taxa):
130
  | Hilbert & Elde (2023), Siglec / C-lectins | 57 | 28,462 | 0.892 | 0.224 | 44.8% | 0.17% | 0.376 | 14,203.0 s | 16.12 s | 859.1x |
131
  | **Global aggregate** | **84** | **43,302** | **0.914** | **0.286** | **50.6%** | **0.191%** | **0.489** | **25,696.0 s** | **78.87 s** | **325.8x** |
132
 
133
- Validation metrics recorded in the training checkpoint: Spearman rho 0.5974, MSE
134
  3.6411, at epoch 16.
135
 
 
 
 
136
  ## Viral fine-tuned variant
137
 
138
- The base model is trained on mammalian (deep-tree) alignments and is effectively **blind on viral /
139
- shallow-tree data**. `model.viral.safetensors` fine-tunes the base on viral data (viral sites
140
- up-weighted) to fix this, **without** sacrificing mammalian performance. Six viral families were held
141
- out entirely as a cross-family generalization test.
142
 
143
- Site-ranking Spearman rho vs. HyPhy MEME LRT:
 
 
 
 
144
 
145
- | Evaluation set | Base | **Viral fine-tuned** |
146
  | :--- | :---: | :---: |
147
- | **Unseen viral families** (flu, corona, flavi, picorna, toga, adeno — never trained) | 0.096 | **0.427** |
148
- | In-distribution viral | 0.263 | 0.566 |
149
- | Mammalian (forgetting check) | 0.217 | 0.524 |
 
 
 
 
 
 
 
 
 
 
150
 
151
- The unseen-families figure is genuine cross-family generalization (those families were held out at the
152
- job level), not memorization. On shallow viral alignments — where MEME's own statistical power is
153
- limited — this is best read as "from noise to a useful triage / pre-screen signal," not as a
154
- replacement for a full MEME run.
155
 
156
  ### ONNX deployment note
157
 
158
- `model.viral.onnx` (opset 17) is **the model only** — it maps five already-computed input tensors
159
  (`msa_codons`, `msa_aas`, `dist_matrix`, `mds_coords`) to per-site LRT, and is verified to match the
160
- PyTorch model to ~1e-6. The inputs, including the MDS coordinates (a classical-MDS **eigendecomposition**
161
- of the patristic distance matrix performed *before* the model), must be computed by the caller and are
162
- not part of the graph. Reproducing the MDS coordinates bit-for-bit across different eigensolvers is
163
- non-trivial (eigenvector sign / degenerate-subspace rotations are ambiguous); plan browser ports
164
- accordingly.
165
 
166
  ## Limitations and intended use
167
 
 
130
  | Hilbert & Elde (2023), Siglec / C-lectins | 57 | 28,462 | 0.892 | 0.224 | 44.8% | 0.17% | 0.376 | 14,203.0 s | 16.12 s | 859.1x |
131
  | **Global aggregate** | **84** | **43,302** | **0.914** | **0.286** | **50.6%** | **0.191%** | **0.489** | **25,696.0 s** | **78.87 s** | **325.8x** |
132
 
133
+ Validation metrics recorded in the base model's training checkpoint: Spearman rho 0.5974, MSE
134
  3.6411, at epoch 16.
135
 
136
+ (The evaluation above is the base model. For viral / shallow-tree data, see the viral fine-tuned
137
+ variant below.)
138
+
139
  ## Viral fine-tuned variant
140
 
141
+ The base model is trained on mammalian (deep-tree) alignments and transfers poorly to viral,
142
+ shallow-tree data. `model.viral.safetensors` fine-tunes the base on viral data (with viral sites
143
+ up-weighted) to fix this, while preserving mammalian performance.
 
144
 
145
+ The table below is a **controlled comparison of the base model against the fine-tuned model on the same
146
+ held-out evaluation sets**, reporting site-ranking Spearman rho against HyPhy MEME LRT. The numbers are
147
+ meaningful relative to each other (same data, same metric); they are measured on a viral-focused internal
148
+ split and are not directly comparable to the literature-benchmark aggregate above, which uses different
149
+ datasets.
150
 
151
+ | Held-out evaluation set | Base | Viral fine-tuned |
152
  | :--- | :---: | :---: |
153
+ | Unseen viral families (flu, corona, flavi, picorna, toga, adeno; excluded from fine-tuning) | 0.10 | **0.43** |
154
+ | Viral, other families | 0.26 | **0.57** |
155
+ | Mammalian (forgetting check) | reference | matches base (no drop) |
156
+
157
+ Two takeaways:
158
+
159
+ 1. **Viral generalization.** On the base model, viral site rankings barely correlate with MEME (rho near
160
+ zero). After fine-tuning they correlate substantially, and this holds on six viral families that were
161
+ excluded from fine-tuning entirely, so it is genuine cross-family generalization rather than
162
+ memorization.
163
+ 2. **No catastrophic forgetting.** Re-evaluated on the same held-out mammalian set, the fine-tuned model
164
+ matches the base model (no measurable drop), so the viral gains do not come at the expense of the
165
+ mammalian regime the base model was trained for.
166
 
167
+ On shallow viral alignments MEME's own statistical power is limited, so the fine-tuned model is best used
168
+ as a fast triage / pre-screen for candidate sites, not as a replacement for a full MEME run.
 
 
169
 
170
  ### ONNX deployment note
171
 
172
+ `model.viral.onnx` (opset 17) contains the model only. It maps five already-computed input tensors
173
  (`msa_codons`, `msa_aas`, `dist_matrix`, `mds_coords`) to per-site LRT, and is verified to match the
174
+ PyTorch model to within about 1e-6. The caller must compute the inputs; they are not part of the graph.
175
+ In particular, `mds_coords` comes from a classical-MDS eigendecomposition of the patristic distance
176
+ matrix performed before the model. Reproducing those coordinates bit-for-bit across different
177
+ eigensolvers is non-trivial, because eigenvector signs and degenerate-subspace rotations are ambiguous,
178
+ so browser or cross-language ports should account for this.
179
 
180
  ## Limitations and intended use
181