BiMU: Active Continual Learning with Metaplastic Binary Bayesian Neural Networks
ICML 2026 · Paper (arXiv:2605.30198) · OpenReview · Code · Project page
Kellian Cottart, Théo Ballet, Djohan Bonnet, Damien Querlioz (Université Paris-Saclay, CNRS, C2N)
BiMU (Binary Metaplasticity from Uncertainty) is a Bayesian continual-learning rule for binary neural networks: each one-bit weight keeps a probability of being +1 or −1, and its own uncertainty sets how far it may move, so the network keeps learning on long non-stationary streams without replay or task boundaries, and uses the same uncertainty to decide which samples are worth labelling.
Model details
| Method | Bayesian learning rule for binary neural networks, with uncertainty-driven active learning |
| Weights | Bernoulli over ±1, one latent parameter λ per weight (mean tanh λ, variance 1 − tanh² λ); a single bit per weight at inference |
| Key hyperparameter | Memory window N: how much evidence each weight keeps |
| Benchmarks | Permuted MNIST (1000 tasks), OpenLORIS-Object, Animals; STM32 microcontroller measurements |
| Framework | JAX, Equinox, Optax |
| Code | github.com/kellian-cottart/active-continual-learning-bayesianbinn, one script per table and figure |
Intended use
- Continual learning on streams that drift over time, one sample at a time, without a replay buffer or task boundaries.
- Always-on edge devices where each weight must fit in one bit at inference and backward passes are expensive.
- Active learning: the disagreement between K sampled binary networks decides whether an input is worth a label and an update.
- A baseline for Bayesian binary neural networks and binary continual learning.
How to use
This repository documents the method; the code, configurations and reproduction scripts are on GitHub.
git clone https://github.com/kellian-cottart/active-continual-learning-bayesianbinn
cd active-continual-learning-bayesianbinn
conda env create -f environment.yml && conda activate aclbbnn
# reproduce the 1000-task Permuted MNIST result
python main.py --config main-pmnist-1000tasks-100neurons/bimu --n_iterations 5 --ood fashion --gpu 0 --verbose
BiMU itself is an Optax gradient transformation, usable in any JAX training loop:
import optax
from optimizers.bimu import bimu
optimizer = bimu(lr=1.0, lr_max=1.0, N=1000) # N: the memory window
opt_state = optimizer.init(params) # params: the latent λ of each binary weight
updates, opt_state = optimizer.update(grads, opt_state, params)
params = optax.apply_updates(params, updates)
Results (from the paper)
- 1000 tasks of Permuted MNIST, one sample at a time: 90.30% mean accuracy on the last five tasks, against 41.12% for BayesBiNN, 29.35% for the straight-through estimator and 10.27% for Synaptic Metaplasticity, with strong out-of-distribution detection throughout.
- OpenLORIS-Object, online continual learning (12 tasks in one pass, online binary head on frozen VGG19 features): 73.61% mean accuracy with features compressed to 1,024 dimensions, against 72.01% for BayesBiNN, 62.82% for Synaptic Metaplasticity and 52.88% for STE; 89.19% at 8,192 and 90.62% at the full 25,088 dimensions, with an epistemic OOD-detection AUC of 1.00 at 1,024 and 8,192.
- OpenLORIS-Object, active continual learning on a long-tailed stream: 88.70% accuracy while labelling and updating on only 3.1% of the stream, a 32× reduction, against 87.76% when training on every sample; 90.91% with a 4.0% budget (25× fewer).
- On an STM32 microcontroller (NUCLEO-64, 216 MHz): learning only when unsure costs 45.0 ms of compute per data point instead of 559.4 ms (12.4× less), for 89.30% vs 90.98% accuracy.
Limitations
- The OpenLORIS-Object experiments train an online binary classifier on top of a frozen ImageNet-pretrained feature extractor, not a fully binary network end to end.
- The memory window N trades stability against plasticity: too small and the network forgets (27.09% mean accuracy at N = 100), too large and it saturates (83.70% at N = 100,000).
- Uncertainty estimates need K Monte Carlo forward passes per input.
Citation
@inproceedings{cottart2026active,
title = {Active Continual Learning with Metaplastic Binary Bayesian Neural Networks},
author = {Cottart, Kellian and Ballet, Th{\'e}o and Bonnet, Djohan and Querlioz, Damien},
booktitle = {Proceedings of the 43rd International Conference on Machine Learning (ICML)},
series = {Proceedings of Machine Learning Research},
volume = {306},
year = {2026},
eprint = {2605.30198},
archivePrefix = {arXiv},
url = {https://openreview.net/forum?id=SPZd0HVyiS}
}
BiMU is the binary counterpart of MESU: Bayesian continual learning and forgetting in neural networks, Nature Communications 16, 9614 (2025).
Dataset used to train kellian-cottart/BinaryMetaplasticityFromUncertainty
Paper for kellian-cottart/BinaryMetaplasticityFromUncertainty
Evaluation results
- Mean accuracy on the last 5 tasks on Permuted MNIST, 1000 sequential tasks, batch size 1Cottart et al., ICML 2026 (arXiv:2605.30198), Table 190.300
- Mean accuracy over 12 tasks on OpenLORIS-Object, frozen VGG19 features compressed to 1,024 dimsCottart et al., ICML 2026 (arXiv:2605.30198), Table 273.610
- OOD detection AUC (epistemic, held-out class) on OpenLORIS-Object, frozen VGG19 features compressed to 1,024 dimsCottart et al., ICML 2026 (arXiv:2605.30198), Table 21.000
- Mean accuracy over 12 tasks on OpenLORIS-Object, frozen VGG19 features, 8,192 dimsCottart et al., ICML 2026 (arXiv:2605.30198), Table 289.190
- Mean accuracy over 12 tasks on OpenLORIS-Object, frozen VGG19 features, 25,088 dimsCottart et al., ICML 2026 (arXiv:2605.30198), Table 290.620
- Accuracy when labelling and updating on 3.1% of the stream (32x fewer) on OpenLORIS-Object, long-tailed, frozen VGG19 features, 8,192 dimsCottart et al., ICML 2026 (arXiv:2605.30198), Section 4.5, Figure 588.700
- Accuracy when labelling and updating on 4.0% of the stream (25x fewer) on OpenLORIS-Object, long-tailed, frozen VGG19 features, 8,192 dimsCottart et al., ICML 2026 (arXiv:2605.30198), Section 4.5, Figure 590.910