|
Download README.md from int-brain-lab/ea-encoder-unit: direct link, hf CLI and curl.
- Browser
- Download file 4.8 kB
-
https://huggingface.co/int-brain-lab/ea-encoder-unit/resolve/main/README.md
- Command line
-
hf download hf://int-brain-lab/ea-encoder-unit/README.md
-
curl -L -o README.md https://huggingface.co/int-brain-lab/ea-encoder-unit/resolve/main/README.md
4.8 kB
| license: cc-by-4.0 | |
| library_name: pytorch | |
| tags: | |
| - neuroscience | |
| - electrophysiology | |
| - neuropixels | |
| - unit-level-encoder | |
| - international-brain-lab | |
| # Ephys Atlas unit-level interpolation model (2026_W41) | |
| Predicts **which kinds of neurons you expect to record at a brain position**, and their expected | |
| waveform phenotype. Trained by the | |
| [International Brain Laboratory](https://www.internationalbrainlab.com/) on spike-sorted units of | |
| the Ephys Atlas release `2026_W41` (56497 training units). | |
| The model has three stages: | |
| 1. a multimodal autoencoder embeds each unit's multi-channel **waveform**, 3D | |
| **autocorrelogram** and **population coupling** (waveform, acg, stpc) into a | |
| 60-d latent -- its phenotype; | |
| 2. a **25-component Gaussian mixture** over the standardised latents, whose | |
| components read as putative cell types. Component shapes are global; the **mixture weights | |
| depend on the molecular and cellular context** (PCA volumes of the molecular and cellular | |
| atlases) at the unit's position; | |
| 3. a **context-local exemplar readout** maps the model back to 10 interpretable waveform features: `depolarisation_slope`, `recovery_slope`, `repolarisation_slope`, `spatial_spread_um`, `tip_val`, `spike_width_secs`, `predepolarisation_width_secs`, `spike_amplitude`, `peak_to_trough_ratio_log`, `polarity`. At a position, each component is represented by its training members whose molecular context resembles the position's (the 4096 nearest in a 3-d context key, shrunk toward all of the component's members -- the more so the farther the context lies from the training data, and entirely where there is no context). So the context sets the mixture weights **and** selects which training exemplars represent each component; the latent density depends on it only through the weights. | |
| > **The input is a position.** `predict` takes `x, y, z` (IBL/Allen frame, metres) and returns the | |
| > expected phenotype there: the local component weights times each component's expected features | |
| > there. `mixture_weights` returns the local putative cell-type composition itself. | |
| ## Quickstart | |
| ```python | |
| import pandas as pd | |
| from ephysatlas import load_pretrained | |
| model = load_pretrained("int-brain-lab/ea-encoder-unit", revision="2026_W41") | |
| df = pd.read_parquet("example/positions_sample.parquet") # positions (x, y, z) | |
| out = model.predict(df) # one pred_<feature> column each | |
| weights = model.mixture_weights(df) # putative cell-type composition | |
| draws = model.sample(df, n_samples=16) # training units drawn from the predictive distribution | |
| model.selftest() # reproduces the shipped outputs | |
| ``` | |
| Units you recorded yourself can be embedded with `model.encode(waveform, acg, stpc)` (arrays shaped | |
| as in `config.json`), then assigned to a putative cell type with `model.assign(z)`. | |
| ## What ships here | |
| - `autoencoder.pt`, `shared_latent_scaler.joblib` -- the phenotype encoder and its latent scaler. | |
| - `global_gmm.joblib` -- the mixture's component means and covariances. | |
| - `context_transform.joblib`, `context_weight_model_bundle.pt` -- context -> mixture weights. | |
| - `knn_bank.npz` -- standardised latents, phenotype features, GMM components and context keys of the training units (the exemplars; part of the model, like the stored points of any nearest-neighbour model). | |
| - `component_feature_expectations.npz` -- each component's expected features over the whole | |
| brain. | |
| - `agea_vol_pca.npy`, `merfish_vol_pca.npy` -- the molecular and cellular context volumes, the same frozen | |
| volumes as the channel-level model `int-brain-lab/ea-encoder-channel` of this vintage. | |
| - `split.json`, `config.json`, `preprocessing/unit_stats.npz`, `results/` -- the probe split, the | |
| training configuration, the training-data statistics and held-out evaluation summaries. | |
| The recorded per-unit dataset (waveforms, ACGs, positions) is **not** here: it stays on IBL S3 and | |
| is only needed to retrain the model or to reproduce the atlas-wide analyses. | |
| > **First run downloads the Allen volume.** Sampling the context constructs an `AllenAtlas`, which | |
| > fetches the Allen CCF volume (public, no account, a few hundred MB) once per machine. | |
| ## Reproducibility | |
| **Pin the revision.** `revision="2026_W41"` is an immutable tag; omitting it resolves to `main`, | |
| which tracks the currently recommended model. `ephysatlas_model.json` records the training-time | |
| environment and random seed; `model.selftest()` verifies your install reproduces the shipped | |
| output (no IBL account needed). | |
| ## Citation | |
| Please cite the International Brain Laboratory Ephys Atlas. Model id `2026_W41_unit`, vintage | |
| `2026_W41`. | |