methscope / README.md
zhouwanding's picture
hg38_celltype: violation framework, 33 classes; .ubjx -> .clfx
2a5d398 verified
|
Raw
History Blame Contribute Delete
6.69 kB
---
license: agpl-3.0
library_name: methscope
tags:
- dna-methylation
- methylation
- wgbs
- single-cell
- imputation
- deconvolution
- cell-type-annotation
---
# MethScope — pretrained models
Pretrained models for [**methscope-cli**](https://github.com/zhou-lab/methscope-cli):
ultra-fast analysis of sparse DNA methylomes via MRMP (Most Recurrent Methylation
Pattern) encoding. Each model is a **single self-contained bundle** that already
carries its MRMP feature definition (and labels / signature), so a query `.cg`
runs directly — no separate annotation files.
## Fetch & run
```sh
# fetch one model (or --local-dir models to grab all)
hf download zhou-lab/methscope hg38_celltype.clfx --local-dir .
methscope classify query.cg hg38_celltype.clfx # cell type / label
methscope deconv mixture.cg hg38_65celltypes.refx # cell-type proportions (NNLS)
methscope upscale -o out.cg hg38_wg.updecx query.cg # genome-wide CpG upscaling
methscope inspect hg38_sex.clfx # framework, labels, features
```
The file extension names the model's **role** — `.clfx` classifier, `.refx`
deconvolution reference, `.updecx` upscale decoder. All three are the same
`MSBNDL1` container and are detected by magic, never by name. Classifiers used
to be `.ubjx`, which described a payload format (a UBJSON booster) rather than a
role and had already stopped being accurate for the non-booster frameworks;
`.ubjx` is still accepted everywhere and nothing needs renaming locally.
Small query `.cg` test fixtures (typed cells, a deconvolution mixture, an
upscale input + truth) live in the GitHub repo
[`zhou-lab/methscope_data`](https://github.com/zhou-lab/methscope_data) under
`test/`, alongside this catalog and the MRMP-construction reproducibility archive.
## Models
| file | task | framework | labels / notes |
|------|------|-----------|----------------|
| `hg38_wg.updecx` | whole-genome CpG upscaling | `UPDEC2` | all 29,401,795 hg38 CpGs; the primary upscaler (2.8 GB) |
| `hg38_10k1.updecx` | CpG upscaling (single block) | MLP decoder | block 10k1 (10,000 CpGs); small demo model |
| `hg38_celltype.clfx` | cell-type annotation | `violation` | 33 human cell types; **no fitted parameter** — see below |
| `mm10_celltype.ubjx` | cell-type annotation (brain) | xgboost | 41 mouse-brain cell types; rename to `mm10_celltype_brain.clfx` pending a Liu 2021 rebuild |
| `hg38_sex.clfx` | sex prediction | logistic | Female, Male (XCI `Xa_hi`/`Xa_lo` markers) |
| `hg38_65celltypes.refx` | cell-type deconvolution (NNLS) | refx | 65 types = 58 Zhou + 7 Loyfer organ/blood; split MRMP, 15,300 patterns |
`methscope inspect <model>` prints the exact framework, full label list, and (for
linear models) per-feature weights. Classifiers carry a required framework `kind`
mark (`xgboost` / `violation` / `threshold` / `logistic`); `classify` rejects an
unmarked bundle. Upscale decoders (`.updecx`) run via `upscale`, deconvolution
references (`.refx`, `kind=refx`) via `deconv`.
## The `violation` framework
`hg38_celltype.clfx` has **no trained parameters at all**. Every MRMP
already carries a binstring — the methylation state each reference class takes
at that pattern — so the reference *is* a per-class prediction. The model just
asks which class the query contradicts least:
```
E1 = patterns the class expects methylated; E0 = expects unmethylated
w_j = sqrt(n_cpg_j)
v1 = Σ_{j∈E1, observed} w_j·[β_j ≤ 0.5] / Σ_{j∈E1, observed} w_j
v0 = Σ_{j∈E0, observed} w_j·[β_j > 0.5] / Σ_{j∈E0, observed} w_j
predict argmin_class (v1 + v0), needing ≥20 observed patterns per side
```
`confidence` is the **margin** to the runner-up, not a probability. It is
calibrated (accuracy 0.757 / 0.820 / 0.899 at margin ≥ 0.02 / 0.05 / 0.10) and,
unlike a softmax, it collapses toward zero when *no* class fits — which is how
the model signals that the query's true type is missing from the reference.
Records too sparse to clear the minimum are reported `NA` rather than guessed.
Held out **across cohorts** (each class's reference and its test samples come
from different studies), 113 donor-resolved pseudobulks at ≥100k covered CpGs:
**accuracy 0.842, macro-F1 0.853** — against 0.873 / 0.868 for an xgboost model
on the identical MRMP. It trades ~3 points of normal-tissue accuracy for
behaviour that holds up on tumours.
**Single-cell tumours**, where it earns that trade. Bian is split by anatomical
site because that is the whole point — a colorectal cancer cell sitting in liver
must still read as colon:
| cohort | GI fraction | top call |
|---|---|---|
| Bian 2018 CRC, **liver metastases** | **100.0%** | `Enterocyte.(Colon)` |
| Bian 2018 CRC, **lymph metastases** | **100.0%** | `Enterocyte.(Colon)` |
| Bian 2018 CRC, primaries (L/R colon, rectum) | 97.1–100.0% | `Enterocyte.(Colon)` |
| Bian 2018 adjacent normal (within-patient control) | 100.0% | `Enterocyte.(Colon)` |
| Gaiti 2019 CLL | — | `B.Cell` |
| Chaligne 2021 glioma | — | `Oligodendrocyte` |
Nothing in the reference anticipates the metastatic case, and no bulk cohort can
separate origin from site the way these cells do. It also resists the
mitotic-age confound that dominates trained models on this feature space: the
"most-hypomethylated-class sink" sending globally demethylated tumours to
`Acinar.(Pancreas)` drops from 30.0% of TCGA calls to 2.3%, with no
residualisation and no tumour data anywhere in its construction.
Accuracy is coverage-dependent: 0.163 below 8k covered CpGs, 0.719 at 8–32k,
plateauing near 0.85 above ~130k. **Do not use below ~30k covered CpGs.**
Limitations: `Macrophage` (0.27) and `Monocyte` (0.74) bleed into
`Dendritic.Cell` — the macrophage reference here is BLUEPRINT's, which is
monocyte-derived and differentiated *in vitro*, a poor match for tissue
macrophages. `T.Cell.CD8` leaks to `CD4` at 27%. Every remaining error is a
within-lineage confusion; none cross lineages. The atlas is also incomplete, and
a missing origin surfaces as a **low margin** rather than an error — glioma has
the lowest median margin of any cohort tested, because neither OPC nor astrocyte
is a class. Margins are not comparable between models with different class
counts, so recalibrate any threshold per model. For a bulk tumour, deconvolve
rather than classify: a classifier assigns one label to a mixture.
Rebuild from any MRMP artifact, with no training data, in about 7 seconds:
```sh
methscope classify-train --framework violation -p 1000 -o out.clfx ref.mrmp
```
See the [methscope-cli documentation](https://zhou-lab.github.io/methscope-cli/)
for the full command reference and runnable examples.