Ephys Atlas channel transformer (2026_W39, Cosmos)
Predicts the brain region of each Neuropixels recording channel from electrophysiological
features alone -- no histology required. Unlike the per-channel region classifier, this model is a
transformer that reads a whole probe at once: every channel attends to every other, using the
channel's physical depth as the positional signal, so neighbouring channels inform each other.
Trained by the International Brain Laboratory on the
Ephys Atlas feature release 2026_W39.
What you can and cannot do without IBL access. The model runs for anyone. Computing the input features from raw Neuropixels data needs
ibleatools. To try the model immediately, use the bundled sample underexample/
Quickstart
import pandas as pd
from ephysatlas import load_pretrained
model = load_pretrained("int-brain-lab/ea-decoder-channel-transformer", revision="2026_W39")
df = pd.read_parquet("example/features_sample.parquet") # or your own features
out = model.predict(df)
print(out[["predicted_acronym", "prediction_probability", "seed_agreement"]].head())
load_pretrained is the entry point for every ephysatlas model, whatever its family -- it reads
ephysatlas_model.json and returns the right wrapper. Use it rather than importing a concrete
class, so your code keeps working as the package evolves.
predict returns one row per input channel, indexed identically to the input:
predicted_acronym, its Allen predicted_atlas_id, the seed-averaged prediction_probability,
a seed_agreement column (fraction of the 5 seed models voting for the winner -- the
natural uncertainty signal), and a p_<acronym> column per class.
The prediction columns are namespaced so that df.join(out) works: the feature table already
carries histology-derived acronym / atlas_id columns, and predictions must not shadow them.
Which weights are used
This release is an ensemble of 5 models, one per random seed, all trained on the same
split. By default predict averages them. Pass estimator="global" to use the first seed alone
-- a fraction of the inference cost, but seed_agreement then comes back as NaN, since no other
seed was consulted. The two modes disagree on a small fraction of channels, so pick one per
analysis.
Note the ensemble is over seeds, not folds: each member saw the same training data from a different random initialisation. It measures the model's sensitivity to initialisation, not held-out generalisation.
Inputs
- 51 features, listed in
ephysatlas_model.jsonunderinputs.features. Every one must be present;predictraises and names anything missing. - A
axial_umcolumn, the channel's depth along the probe. This is the model's positional signal, not a feature -- without it there is no geometry andpredictraises. - Indexed by
(pid, channel), one row per recording channel. Rows are grouped bypidand each probe is read as a whole, so pass complete probes: scoring a handful of loose channels asks the model to attend over a probe that does not exist, and the answer will differ from the same channels scored in context. - Channels carrying any non-finite feature are dropped, so the returned frame may be shorter than the input. Join on the index rather than assuming row alignment.
- Must be the denoised aggregated features calculated from ibeatool -- . Run
model.selftest()to confirm your install reproduces the shipped output before trusting it.
Performance
Predicts 13 Cosmos regions. Splits are by insertion (pid), so no channel from a
test insertion appears in training. Per-run evaluation -- confusion matrices, per-seed prediction
tables and training curves -- is written beside each training run rather than quoted here, so
that this card cannot drift from the numbers it claims.
Limitations
- Trained on IBL Neuropixels 1.0 recordings in mouse. Transfer to NP2, other species or other rigs is untested.
- Channel depths are assumed uniform per channel index across probes (NP1 geometry). Probes with a different channel map are out of scope.
- Coverage follows IBL brain-wide-map targeting; rare regions are under-represented.
Cosmosis a coarse parcellation.- Because a prediction depends on the whole probe, truncated or heavily masked probes are scored in a context the model did not see in training.
Reproducibility
Pin the revision. revision="2026_W39" is an immutable tag. Omitting revision resolves
to main, which tracks whichever model is currently recommended and will change when a new
feature vintage is published -- fine for a first look, not for anything you publish or re-run.
ephysatlas_model.json records the architecture under config.model_config, the per-seed list
under artifacts.seeds, and the training seeds under training.seeds. Verify your install
reproduces the shipped output:
model.selftest()
Citation
Please cite the International Brain Laboratory Ephys Atlas. Model id ea-decoder-channel-transformer,
feature vintage 2026_W39.