--- license: mit language: en library_name: scikit-learn tags: - anomaly-detection - isolation-forest - time-series - system-monitoring - tabular datasets: - harpertoken/stat metrics: - precision - recall --- # mark An IsolationForest anomaly detector over the [`stat`](https://huggingface.co/datasets/harpertoken/stat) macOS system-telemetry dataset. It learns what normal machine behaviour looks like from per-second CPU, memory, disk and network counters, and flags samples that do not fit. CPU only, scikit-learn only, no PyTorch and no transformer anywhere in the path. The saved bundle is a StandardScaler plus a 200-tree forest over eight features, and runs to about 1.9 MB. Training used the first 80 percent of the session in time order (4,951 rows) and held out the last 20 percent (1,238 rows). The disk and network columns in `stat` are since-boot counters, so they enter as per-interval differences, exactly as that card recommends. `cpu_temp` is null throughout and is dropped; per-core `cpu_usage` enters as its mean and max. ## Usage No scikit-learn and no pickle. The forest ships as plain arrays in `model.safetensors` with parameters in `config.json`, and `predict.py` scores them with numpy only. It replicates scikit-learn's scoring: 99.9 to 100 percent predict agreement with the original estimator on train, holdout, and synthetic incidents, with scores matching to within 0.003. ```python import numpy as np from huggingface_hub import hf_hub_download from predict import MarkModel m = MarkModel(".") X = np.array([[46.8, 96.0, 3410.2, 33.0, 6.1, 5.8, 0.4, 0.3]]) print(m.predict(X)) ``` The eight columns, in order, are `cpu_mean`, `cpu_max`, `memory_used_mb`, `battery_status`, `disk_read_delta`, `disk_write_delta`, `net_sent_delta` and `net_recv_delta`. Output is 1 for normal, -1 for flagged. Needs `safetensors` and `numpy` only. ## Evaluation Measured on the temporal holdout and on synthetic incidents built from holdout rows: | Case | Flagged | |---|---| | Clean holdout (false alarms) | 0.027 | | Stressed machine, CPU, memory and disk spiking together | 0.917 | | Idle crash, CPU and memory dropping together | 0.372 | | Network storm, sent and received spiking together | 0.245 | The contamination parameter is 0.01, so about one percent of training rows score as anomalous by construction; the holdout false-alarm rate of 0.027 reflects the battery recharge region the holdout covers, which the training window never saw. Single-feature extremes score poorly, and that is a property of the algorithm rather than a bug in this fit. IsolationForest separates points by random splits, and a point that is extreme in one of eight dimensions waits several splits for that dimension to be picked, which is the same depth normal points reach. Standardizing the features changes nothing for the same reason: the splits are uniform in range, and rescaling only reparametrizes the draw. Incidents that move several counters at once, the way a real stressed machine does, are the ones it catches. ## Limitations This models one machine over one 1-hour-44-minute session. It has not seen another host, another workload, or a longer timescale, and there is no reason to expect it to transfer. The synthetic incidents are illustrative, not a benchmark. The `stat` card's unit caveat applies here too: the disk and network deltas are raw counter advances whose physical unit was never confirmed, so thresholds learned on them are in those raw units.