Model card: fix rho inconsistency (base vs viral eval sets), remove em-dashes, clarify variant labeling
Browse filesThe viral table is now framed as a controlled base-vs-fine-tuned comparison on a separate viral split, explicitly not comparable to the literature-benchmark aggregate; the mammal 'forgetting check' row no longer shows a misleading raw rho that clashed with the checkpoint's 0.5974. Em-dashes removed; each rho now clearly attributed to base or fine-tuned.
README.md
CHANGED
|
@@ -130,38 +130,52 @@ datasets from 9 independent literature studies (43,302 codons, up to 476 taxa):
|
|
| 130 |
| Hilbert & Elde (2023), Siglec / C-lectins | 57 | 28,462 | 0.892 | 0.224 | 44.8% | 0.17% | 0.376 | 14,203.0 s | 16.12 s | 859.1x |
|
| 131 |
| **Global aggregate** | **84** | **43,302** | **0.914** | **0.286** | **50.6%** | **0.191%** | **0.489** | **25,696.0 s** | **78.87 s** | **325.8x** |
|
| 132 |
|
| 133 |
-
Validation metrics recorded in the training checkpoint: Spearman rho 0.5974, MSE
|
| 134 |
3.6411, at epoch 16.
|
| 135 |
|
|
|
|
|
|
|
|
|
|
| 136 |
## Viral fine-tuned variant
|
| 137 |
|
| 138 |
-
The base model is trained on mammalian (deep-tree) alignments and
|
| 139 |
-
shallow-tree data
|
| 140 |
-
up-weighted) to fix this,
|
| 141 |
-
out entirely as a cross-family generalization test.
|
| 142 |
|
| 143 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 144 |
|
| 145 |
-
|
|
| 146 |
| :--- | :---: | :---: |
|
| 147 |
-
|
|
| 148 |
-
|
|
| 149 |
-
| Mammalian (forgetting check) |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 150 |
|
| 151 |
-
|
| 152 |
-
|
| 153 |
-
limited — this is best read as "from noise to a useful triage / pre-screen signal," not as a
|
| 154 |
-
replacement for a full MEME run.
|
| 155 |
|
| 156 |
### ONNX deployment note
|
| 157 |
|
| 158 |
-
`model.viral.onnx` (opset 17)
|
| 159 |
(`msa_codons`, `msa_aas`, `dist_matrix`, `mds_coords`) to per-site LRT, and is verified to match the
|
| 160 |
-
PyTorch model to
|
| 161 |
-
|
| 162 |
-
|
| 163 |
-
non-trivial
|
| 164 |
-
|
| 165 |
|
| 166 |
## Limitations and intended use
|
| 167 |
|
|
|
|
| 130 |
| Hilbert & Elde (2023), Siglec / C-lectins | 57 | 28,462 | 0.892 | 0.224 | 44.8% | 0.17% | 0.376 | 14,203.0 s | 16.12 s | 859.1x |
|
| 131 |
| **Global aggregate** | **84** | **43,302** | **0.914** | **0.286** | **50.6%** | **0.191%** | **0.489** | **25,696.0 s** | **78.87 s** | **325.8x** |
|
| 132 |
|
| 133 |
+
Validation metrics recorded in the base model's training checkpoint: Spearman rho 0.5974, MSE
|
| 134 |
3.6411, at epoch 16.
|
| 135 |
|
| 136 |
+
(The evaluation above is the base model. For viral / shallow-tree data, see the viral fine-tuned
|
| 137 |
+
variant below.)
|
| 138 |
+
|
| 139 |
## Viral fine-tuned variant
|
| 140 |
|
| 141 |
+
The base model is trained on mammalian (deep-tree) alignments and transfers poorly to viral,
|
| 142 |
+
shallow-tree data. `model.viral.safetensors` fine-tunes the base on viral data (with viral sites
|
| 143 |
+
up-weighted) to fix this, while preserving mammalian performance.
|
|
|
|
| 144 |
|
| 145 |
+
The table below is a **controlled comparison of the base model against the fine-tuned model on the same
|
| 146 |
+
held-out evaluation sets**, reporting site-ranking Spearman rho against HyPhy MEME LRT. The numbers are
|
| 147 |
+
meaningful relative to each other (same data, same metric); they are measured on a viral-focused internal
|
| 148 |
+
split and are not directly comparable to the literature-benchmark aggregate above, which uses different
|
| 149 |
+
datasets.
|
| 150 |
|
| 151 |
+
| Held-out evaluation set | Base | Viral fine-tuned |
|
| 152 |
| :--- | :---: | :---: |
|
| 153 |
+
| Unseen viral families (flu, corona, flavi, picorna, toga, adeno; excluded from fine-tuning) | 0.10 | **0.43** |
|
| 154 |
+
| Viral, other families | 0.26 | **0.57** |
|
| 155 |
+
| Mammalian (forgetting check) | reference | matches base (no drop) |
|
| 156 |
+
|
| 157 |
+
Two takeaways:
|
| 158 |
+
|
| 159 |
+
1. **Viral generalization.** On the base model, viral site rankings barely correlate with MEME (rho near
|
| 160 |
+
zero). After fine-tuning they correlate substantially, and this holds on six viral families that were
|
| 161 |
+
excluded from fine-tuning entirely, so it is genuine cross-family generalization rather than
|
| 162 |
+
memorization.
|
| 163 |
+
2. **No catastrophic forgetting.** Re-evaluated on the same held-out mammalian set, the fine-tuned model
|
| 164 |
+
matches the base model (no measurable drop), so the viral gains do not come at the expense of the
|
| 165 |
+
mammalian regime the base model was trained for.
|
| 166 |
|
| 167 |
+
On shallow viral alignments MEME's own statistical power is limited, so the fine-tuned model is best used
|
| 168 |
+
as a fast triage / pre-screen for candidate sites, not as a replacement for a full MEME run.
|
|
|
|
|
|
|
| 169 |
|
| 170 |
### ONNX deployment note
|
| 171 |
|
| 172 |
+
`model.viral.onnx` (opset 17) contains the model only. It maps five already-computed input tensors
|
| 173 |
(`msa_codons`, `msa_aas`, `dist_matrix`, `mds_coords`) to per-site LRT, and is verified to match the
|
| 174 |
+
PyTorch model to within about 1e-6. The caller must compute the inputs; they are not part of the graph.
|
| 175 |
+
In particular, `mds_coords` comes from a classical-MDS eigendecomposition of the patristic distance
|
| 176 |
+
matrix performed before the model. Reproducing those coordinates bit-for-bit across different
|
| 177 |
+
eigensolvers is non-trivial, because eigenvector signs and degenerate-subspace rotations are ambiguous,
|
| 178 |
+
so browser or cross-language ports should account for this.
|
| 179 |
|
| 180 |
## Limitations and intended use
|
| 181 |
|