Instructions to use EnerTEF/Anomaly-Detection-Service with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Scikit-learn
How to use EnerTEF/Anomaly-Detection-Service with Scikit-learn:
from huggingface_hub import hf_hub_download import joblib model = joblib.load( hf_hub_download("EnerTEF/Anomaly-Detection-Service", "sklearn_model.joblib") ) # only load pickle files from sources you trust # read more about it here https://skops.readthedocs.io/en/stable/persistence.html - Notebooks
- Google Colab
- Kaggle
Anomaly Detection Service β VRFB
An unsupervised anomaly detection model for a Vanadium Redox Flow Battery (VRFB), prepared for the EnerTEF model collection. It combines a Gaussian Mixture Model (GMM) with a One-Class Support Vector Machine (OCSVM) to flag unusual battery sensor observations using normal operating data for training.
This release packages the methodology from ENERTEF/Anomaly-Detection-Service as a fitted scikit-learn model and a reproducible notebook example.
Model description
- A GMM learns normal operating behaviour from 11 battery measurements, using six components, full covariance matrices and
random_state=42. - The GMM assigns each observation a log-likelihood. The standalone GMM baseline flags values below the minimum training log-likelihood.
- An OCSVM learns a decision boundary on the one-dimensional GMM log-likelihoods of normal observations, using
kernel="rbf",nu=0.001andgamma=0.01.
The combined model returns 1 for an inlier and -1 for an outlier. Measurements retain their original units; the notebook does not apply scaling or imputation. Dates and times are used for the alarm example, rather than as model features. The model classifies individual observations and does not learn temporal sequences.
Repository contents
| File | Purpose |
|---|---|
model.joblib |
Fitted GMM and OCSVM, ordered feature names, GMM threshold, training row count and scikit-learn version. |
Anomaly indication algorithm - VRFB.ipynb |
Training, distribution plots, evaluation, model export and a maintenance alarm example. |
Dataset_11var_treino.csv |
Normal reference data: 74,992 observations. |
test_observations.csv |
Test data: 130,543 observations from the failure scenario described in the source notebook. |
requirements.txt |
Pinned dependencies used to create and verify the fitted model. |
LICENSE |
Original MIT license and copyright notice. |
.gitattributes |
Git LFS attributes for the model and CSV files. |
Both CSV files use ; as the delimiter. They contain Data and Hora columns alongside the following numeric features, in model order:
DC Power Battery W
DC Voltage Battery V
Current DC Battery A
Voltage L1 V
Current L1 A
Temperature Eletrolite C
Avg Schneider Voltage L1 V
Avg Schneider Current L1 kA
Avg Schneider React Power L1 kVAr
Avg Schneider Power Factor L1
SOC
Keep these column names, units and order, including the original spelling of Temperature Eletrolite C. Input measurements must be numeric and finite. The source describes data obtained around a European industrial partner's battery failure; the CSV files do not include fault labels for individual observations.
Installation and inference
Download the repository files and run the following commands from their directory. The supplied model was trained and verified with Python 3.14.4 and scikit-learn 1.9.1. Use the pinned package versions when loading it.
python -m pip install -r requirements.txt
import joblib
import numpy as np
import pandas as pd
model = joblib.load("model.joblib")
observations = pd.read_csv("test_observations.csv", sep=";")
features = observations.loc[:, model["feature_names"]]
if not np.isfinite(features.to_numpy(dtype=float)).all():
raise ValueError("Battery measurements must be numeric and finite.")
log_likelihoods = model["gmm"].score_samples(features)
labels = model["ocsvm"].predict(log_likelihoods.reshape(-1, 1))
results = observations.copy()
results["Log_Likelihood"] = log_likelihoods
results["Outlier_Predictions"] = labels # 1: inlier, -1: outlier
results["Is_Anomaly"] = labels == -1
results.to_csv("predictions.csv", index=False)
The same inference code accepts new observations with the required feature columns. Load the Joblib artifact only from a trusted source, because Joblib uses pickle-based deserialization. Inference runs locally with scikit-learn; this release provides no hosted inference endpoint.
Reproduce the model
Open Anomaly indication algorithm - VRFB.ipynb in Jupyter or a notebook editor, with the repository directory as the working directory. Execute all cells in order using an environment with the dependencies above. If needed, install a Jupyter interface with python -m pip install jupyterlab.
The notebook fits both estimators, generates intermediate result CSVs, and saves model.joblib. Its maintenance example flags anomalous readings after more than three anomalies occur within the preceding 12 hours, using complete dates and times. This rule counts readings, rather than distinct failure events, and requires calibration to the sampling rate and asset.
The upload copy corrects prose cells stored as code, a reference to an unavailable intermediate CSV, the single-Gaussian comparison, and the alarm window implementation. Historical outputs were cleared. The GMM and OCSVM training parameters are preserved; the supplied artifact is newly fitted from the included normal data.
Evaluation
All executable notebook cells completed in the pinned environment. Reloading model.joblib reproduced every test prediction from that run.
| Method and data | Outliers / observations | Outlier rate |
|---|---|---|
| GMM minimum-training-likelihood threshold, test data | 85,956 / 130,543 | 65.8450% |
| GMM + OCSVM, test data | 128,762 / 130,543 | 98.6357% |
| GMM + OCSVM, normal training reference | 138 / 74,992 | 0.1840% |
These are flagging rates. Treating every test observation as anomalous makes the test flagging rate an assumed anomaly recall, but the CSV has no labels to independently verify that assumption. The normal reference was also used for fitting, so its outlier rate is an in-sample measurement. Precision, held-out false-positive rate and general classification accuracy are not established. Small numerical differences from the historical notebook outputs can occur across library versions and platforms.
Intended uses and limitations
The demonstrated use is offline analysis of VRFB battery measurements and experimentation with condition monitoring. The source project also identifies storage assets, wind generation and wave energy converters as potential applications; those applications require appropriate variables, retraining and validation.
Training data represent one battery system and may not cover other operating regimes, sensor calibrations or seasonal conditions. Sensor errors and shifts in operating conditions can produce outliers without an equipment fault. Adapting the feature set or measurement units requires retraining. Validate the model on separate normal and labelled fault periods before using its outputs to guide maintenance decisions.
License and project references
Distributed under the MIT license, preserving Copyright (c) 2024 Horizont-Europe-Interstore from the source repository.
- Source project and original methodology
- EnerTEF organization
- EnerTEF/SAMFormer β organization reference for notebook distribution and model-card metadata.
- EnerTEF/ChebAutoencoder β related detection model in the organization.
- Hugging Face model cards
Upload this package
Create a model repository under EnerTEF (suggested name: Anomaly-Detection-Service) and upload the contents of this folder to the repository root through Files and versions β Add file β Upload files. Keeping README.md at the root makes it the model card. See the Hugging Face upload instructions.
- Downloads last month
- -