DXA-Guard Intrusion Detection & Prevention Model
Official pre-trained Gradient-Boosted Decision Tree (GBDT) model for DXA-Guard, an open-source intrusion detection and prevention agent built for Linux VPS servers running web applications.
Published by the Deployxa organization.
This model evaluates 5-minute sliding windows of host and network activity per source IP to score the likelihood of malicious activity (SSH brute-force, web vulnerability scanners, exploit payloads, and post-compromise behaviors).
Model Architecture & Design
- Model Type: Gradient-Boosted Decision Tree Ensemble (
scikit-learn.ensemble.GradientBoostingClassifier) - Number of Trees: 30 trees (max depth 3)
- Learning Rate: 0.1
- Dual Export Artifacts:
model.joblib/model.pkl: Standard scikit-learn serialized models for Python workflows.model.json: Complete portable tree representation containing all decision thresholds and leaf weights. This enables the compiled Go daemon (dxa-guard) to evaluate the model natively with zero Python runtime and sub-millisecond latency.
Safety Invariant
The model never directly triggers bans or kills processes. In the DXA-Guard architecture, deterministic rules are authoritative for all enforcement actions (nftables bans). This model acts strictly as an advisory risk scorer that elevates or lowers alert urgency without executing actions independently.
Sliding Window Feature Specification (8 Features)
The model takes an 8-dimensional feature vector extracted over a 5-minute sliding window per source IP:
| Index | Feature Name | Type | Description |
|---|---|---|---|
| 0 | window_event_count |
float | Total number of security events observed for the source IP in the window. |
| 1 | failure_ratio |
float | Ratio of failed events (failed auth, 401/403/404 HTTP responses) to total events. |
| 2 | distinct_paths |
float | Count of distinct HTTP request paths targeted by the IP. |
| 3 | distinct_users |
float | Count of unique usernames attempted during SSH authentication. |
| 4 | path_entropy |
float | Shannon character entropy of requested URL paths (flags fuzzer / scanner payloads). |
| 5 | inter_event_time_mean_ms |
float | Mean delta between consecutive events in ms (flags rapid bursts / automated tooling). |
| 6 | hour_of_day_norm |
float | Timestamp hour divided by 24.0 (0.0 to 1.0) for temporal baseline profiling. |
| 7 | sensitive_path_ratio |
float | Proportion of requests probing sensitive files (.env, wp-config, .git, webshells). |
Evaluation Benchmark
Evaluated on the DXA-Guard benchmark test set (deployxa/dxa-guard-benchmark):
| Model / Baseline | Precision | Recall | F1 Score |
|---|---|---|---|
| DXA-Guard GBDT (This Model) | 1.000 | 0.833 | 0.909 |
| Isolation Forest (Anomaly Baseline) | 0.941 | 1.000 | 0.970 |
How to Use
In Python (from Hugging Face)
from huggingface_hub import hf_hub_download
import joblib
import numpy as np
# Download and load model directly from Hugging Face
model_path = hf_hub_download(repo_id="deployxa/dxa-guard", filename="model.joblib")
model = joblib.load(model_path)
# Example feature vector:
# [window_event_count, failure_ratio, distinct_paths, distinct_users,
# path_entropy, inter_event_time_mean_ms, hour_of_day_norm, sensitive_path_ratio]
sample_features = np.array([[25.0, 0.96, 18.0, 1.0, 4.35, 120.5, 0.58, 0.72]])
# Predict attack probability
prob = model.predict_proba(sample_features)[0, 1]
print(f"Attack Probability: {prob:.4f}")
In Go (Native Inference with zero dependencies)
The model can be evaluated natively in Go by parsing model.json:
package main
import (
"fmt"
"dxa-guard/internal/scorer"
)
func main() {
s, err := scorer.LoadModel("model.json")
if err != nil {
panic(err)
}
features := []float64{25.0, 0.96, 18.0, 1.0, 4.35, 120.5, 0.58, 0.72}
prob, _ := s.Score(features)
fmt.Printf("Attack probability: %.4f\n", prob)
}
Native Mathematical Inference Formula
Repository Files
model.joblib: Serialized scikit-learnGradientBoostingClassifier.model.pkl: Standard Python pickle serialized model.model.json: Native JSON tree export with exact weights and thresholds for Go/C++/Rust runtime.isolation_forest.joblib: Anomaly detection baseline model.features.py: Python module containing the reference feature extraction pipeline.features.json: Machine-readable metadata schema of the 8 feature definitions.test_cases.json: Reference feature vectors and exact Python probability outputs for cross-language validation.
License & Attribution
Released under the Apache 2.0 License by Deployxa.