DXA-Guard Intrusion Detection & Prevention Model

Official pre-trained Gradient-Boosted Decision Tree (GBDT) model for DXA-Guard, an open-source intrusion detection and prevention agent built for Linux VPS servers running web applications.

Published by the Deployxa organization.

This model evaluates 5-minute sliding windows of host and network activity per source IP to score the likelihood of malicious activity (SSH brute-force, web vulnerability scanners, exploit payloads, and post-compromise behaviors).


Model Architecture & Design

  • Model Type: Gradient-Boosted Decision Tree Ensemble (scikit-learn.ensemble.GradientBoostingClassifier)
  • Number of Trees: 30 trees (max depth 3)
  • Learning Rate: 0.1
  • Dual Export Artifacts:
    1. model.joblib / model.pkl: Standard scikit-learn serialized models for Python workflows.
    2. model.json: Complete portable tree representation containing all decision thresholds and leaf weights. This enables the compiled Go daemon (dxa-guard) to evaluate the model natively with zero Python runtime and sub-millisecond latency.

Safety Invariant

The model never directly triggers bans or kills processes. In the DXA-Guard architecture, deterministic rules are authoritative for all enforcement actions (nftables bans). This model acts strictly as an advisory risk scorer that elevates or lowers alert urgency without executing actions independently.


Sliding Window Feature Specification (8 Features)

The model takes an 8-dimensional feature vector extracted over a 5-minute sliding window per source IP:

Index Feature Name Type Description
0 window_event_count float Total number of security events observed for the source IP in the window.
1 failure_ratio float Ratio of failed events (failed auth, 401/403/404 HTTP responses) to total events.
2 distinct_paths float Count of distinct HTTP request paths targeted by the IP.
3 distinct_users float Count of unique usernames attempted during SSH authentication.
4 path_entropy float Shannon character entropy of requested URL paths (flags fuzzer / scanner payloads).
5 inter_event_time_mean_ms float Mean delta between consecutive events in ms (flags rapid bursts / automated tooling).
6 hour_of_day_norm float Timestamp hour divided by 24.0 (0.0 to 1.0) for temporal baseline profiling.
7 sensitive_path_ratio float Proportion of requests probing sensitive files (.env, wp-config, .git, webshells).

Evaluation Benchmark

Evaluated on the DXA-Guard benchmark test set (deployxa/dxa-guard-benchmark):

Model / Baseline Precision Recall F1 Score
DXA-Guard GBDT (This Model) 1.000 0.833 0.909
Isolation Forest (Anomaly Baseline) 0.941 1.000 0.970

How to Use

In Python (from Hugging Face)

from huggingface_hub import hf_hub_download
import joblib
import numpy as np

# Download and load model directly from Hugging Face
model_path = hf_hub_download(repo_id="deployxa/dxa-guard", filename="model.joblib")
model = joblib.load(model_path)

# Example feature vector:
# [window_event_count, failure_ratio, distinct_paths, distinct_users,
#  path_entropy, inter_event_time_mean_ms, hour_of_day_norm, sensitive_path_ratio]
sample_features = np.array([[25.0, 0.96, 18.0, 1.0, 4.35, 120.5, 0.58, 0.72]])

# Predict attack probability
prob = model.predict_proba(sample_features)[0, 1]
print(f"Attack Probability: {prob:.4f}")

In Go (Native Inference with zero dependencies)

The model can be evaluated natively in Go by parsing model.json:

package main

import (
    "fmt"
    "dxa-guard/internal/scorer"
)

func main() {
    s, err := scorer.LoadModel("model.json")
    if err != nil {
        panic(err)
    }

    features := []float64{25.0, 0.96, 18.0, 1.0, 4.35, 120.5, 0.58, 0.72}
    prob, _ := s.Score(features)
    fmt.Printf("Attack probability: %.4f\n", prob)
}

Native Mathematical Inference Formula

logit=init_bias+βˆ‘t=130learning_rateΓ—LeafValuetlogit = \text{init\_bias} + \sum_{t=1}^{30} \text{learning\_rate} \times \text{LeafValue}_t P(attack)=11+eβˆ’logitP(\text{attack}) = \frac{1}{1 + e^{-logit}}


Repository Files

  • model.joblib: Serialized scikit-learn GradientBoostingClassifier.
  • model.pkl: Standard Python pickle serialized model.
  • model.json: Native JSON tree export with exact weights and thresholds for Go/C++/Rust runtime.
  • isolation_forest.joblib: Anomaly detection baseline model.
  • features.py: Python module containing the reference feature extraction pipeline.
  • features.json: Machine-readable metadata schema of the 8 feature definitions.
  • test_cases.json: Reference feature vectors and exact Python probability outputs for cross-language validation.

License & Attribution

Released under the Apache 2.0 License by Deployxa.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support