Ephys Atlas channel transformer (2026_W39, Cosmos)

Predicts the brain region of each Neuropixels recording channel from electrophysiological features alone -- no histology required. Unlike the per-channel region classifier, this model is a transformer that reads a whole probe at once: every channel attends to every other, using the channel's physical depth as the positional signal, so neighbouring channels inform each other. Trained by the International Brain Laboratory on the Ephys Atlas feature release 2026_W39.

What you can and cannot do without IBL access. The model runs for anyone. Computing the input features from raw Neuropixels data needs ibleatools. To try the model immediately, use the bundled sample under example/

Quickstart

import pandas as pd
from ephysatlas import load_pretrained

model = load_pretrained("int-brain-lab/ea-decoder-channel-transformer", revision="2026_W39")
df = pd.read_parquet("example/features_sample.parquet")   # or your own features
out = model.predict(df)
print(out[["predicted_acronym", "prediction_probability", "seed_agreement"]].head())

load_pretrained is the entry point for every ephysatlas model, whatever its family -- it reads ephysatlas_model.json and returns the right wrapper. Use it rather than importing a concrete class, so your code keeps working as the package evolves.

predict returns one row per input channel, indexed identically to the input: predicted_acronym, its Allen predicted_atlas_id, the seed-averaged prediction_probability, a seed_agreement column (fraction of the 5 seed models voting for the winner -- the natural uncertainty signal), and a p_<acronym> column per class.

The prediction columns are namespaced so that df.join(out) works: the feature table already carries histology-derived acronym / atlas_id columns, and predictions must not shadow them.

Which weights are used

This release is an ensemble of 5 models, one per random seed, all trained on the same split. By default predict averages them. Pass estimator="global" to use the first seed alone -- a fraction of the inference cost, but seed_agreement then comes back as NaN, since no other seed was consulted. The two modes disagree on a small fraction of channels, so pick one per analysis.

Note the ensemble is over seeds, not folds: each member saw the same training data from a different random initialisation. It measures the model's sensitivity to initialisation, not held-out generalisation.

Inputs

  • 51 features, listed in ephysatlas_model.json under inputs.features. Every one must be present; predict raises and names anything missing.
  • A axial_um column, the channel's depth along the probe. This is the model's positional signal, not a feature -- without it there is no geometry and predict raises.
  • Indexed by (pid, channel), one row per recording channel. Rows are grouped by pid and each probe is read as a whole, so pass complete probes: scoring a handful of loose channels asks the model to attend over a probe that does not exist, and the answer will differ from the same channels scored in context.
  • Channels carrying any non-finite feature are dropped, so the returned frame may be shorter than the input. Join on the index rather than assuming row alignment.
  • Must be the denoised aggregated features calculated from ibeatool -- . Run model.selftest() to confirm your install reproduces the shipped output before trusting it.

Performance

Predicts 13 Cosmos regions. Splits are by insertion (pid), so no channel from a test insertion appears in training. Per-run evaluation -- confusion matrices, per-seed prediction tables and training curves -- is written beside each training run rather than quoted here, so that this card cannot drift from the numbers it claims.

Limitations

  • Trained on IBL Neuropixels 1.0 recordings in mouse. Transfer to NP2, other species or other rigs is untested.
  • Channel depths are assumed uniform per channel index across probes (NP1 geometry). Probes with a different channel map are out of scope.
  • Coverage follows IBL brain-wide-map targeting; rare regions are under-represented.
  • Cosmos is a coarse parcellation.
  • Because a prediction depends on the whole probe, truncated or heavily masked probes are scored in a context the model did not see in training.

Reproducibility

Pin the revision. revision="2026_W39" is an immutable tag. Omitting revision resolves to main, which tracks whichever model is currently recommended and will change when a new feature vintage is published -- fine for a first look, not for anything you publish or re-run.

ephysatlas_model.json records the architecture under config.model_config, the per-seed list under artifacts.seeds, and the training seeds under training.seeds. Verify your install reproduces the shipped output:

model.selftest()

Citation

Please cite the International Brain Laboratory Ephys Atlas. Model id ea-decoder-channel-transformer, feature vintage 2026_W39.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support