Laya v3: High-Throughput Semantic DLP & Credential Detector
Laya v3 is a specialized, production-ready semantic classification model designed for Data Loss Prevention (DLP) and ultra-fast detection of leaked credentials and Personally Identifiable Information (PII) in unstructured text, developer notes, logs, and communication channels.
Trained on top of ModernBERT with reinforcement-calibrated temperature heads, Laya v3 delivers deep contextual understanding of authentication artifacts without the brittleness of traditional regexes or the latency of heavy LLMs.
π Key Highlights
- 100% Recall on Wild Credentials: Zero false negatives observed across isolated hold-out test sets of technical and operational notes.
- Extreme Throughput: 406.3 requests/second on a single NVIDIA Tesla T4 GPU (Batch 16) with < 2 GB VRAM footprint.
- Sub-35ms Latency on CPU: Dynamically quantized to INT8 ONNX (
laya.int8.onnx), achieving ~33-35ms single-slice latency on modern CPU cores. - Early-Exit Sliding Window: Native support for sliding-window document chunking (700 chars / 150 chars overlap) that scans 8,000-character documents in ~30ms via early positive termination.
- Resilient to Hard Negatives: Hardened against false alarms from code snippets (e.g. PHP
password_verify, HTML forms, empty config strings).
π Benchmark & Performance
1. Throughput & Latency (NVIDIA Tesla T4 GPU)
| Scenario | Sequence Length | Batch Size | Average Latency | Throughput |
|---|---|---|---|---|
| Short Snippet (Key / Password) | 32 tokens | 1 | 17.09 ms | 58.5 req/s |
| Short Snippet | 32 tokens | 16 | 39.38 ms | 406.3 req/s |
| Short Snippet | 32 tokens | 32 | 82.50 ms | 387.9 req/s |
| Medium Code Block | 128 tokens | 1 | 18.80 ms | 53.2 req/s |
| Medium Code Block | 128 tokens | 16 | 166.37 ms | 96.2 req/s |
| Long Log Dump | 256 tokens | 1 | 26.14 ms | 38.3 req/s |
Peak GPU Memory: 1.96 GB VRAM (FP32).
2. Accuracy & Capture Metrics (Isolated Hold-Out Corpus)
| Metric | Laya v2 (Previous) | Laya v3 (Production) | Net Improvement |
|---|---|---|---|
| Credential Accuracy | 62.00% | 78.00% | +16.00% |
| Credential Recall (Capture) | 100.00% | 100.00% | Zero Misses |
| False Positives (Alarms) | 38 | 22 | -42.1% (-16 alarms) |
| PII Accuracy | 79.00% | 83.00% | +4.00% |
π οΈ Usage
Quickstart (CPU with Quantized INT8 ONNX)
import onnxruntime as ort
from laya.onnx_agent import ONNXAgent
# Load model directly from Hugging Face or local path
agent = ONNXAgent(
"caramelodev/laya-ghostpad-v3",
onnx_path="laya.int8.onnx"
)
QUESTIONS = {
"has_credentials": {
"type": "noul",
"instructions": "Avalie se este trecho desestruturado contΓ©m credenciais de acesso ou autenticaΓ§Γ£o tΓ©cnica/operacional."
},
"has_pii": {
"type": "noul",
"instructions": "Avalie se este trecho contΓ©m dados de identificaΓ§Γ£o pessoal ou prontuΓ‘rios de terceiros."
}
}
text = "ssh root@192.168.1.100 -p 2222\nsenha: SuperSecretPassword2026!"
prediction = agent.predict(text, QUESTIONS)
print("Score Credenciais:", prediction["answers"]["has_credentials"]["noul"])
# Output: > 0.85 (Positive)
Document Scanning with Sliding Window & Early-Exit
For arbitrary-length notes or log dumps:
def scan_document(agent, text: str, chunk_size: int = 700, overlap: int = 150):
start = 0
while start < len(text):
end = min(start + chunk_size, len(text))
chunk = text[start:end]
result = agent.predict(chunk, QUESTIONS)
score = result["answers"]["has_credentials"]["noul"]
# Early-exit: Stop immediately upon high-confidence detection
if score >= 0.65:
return {"is_leaked": True, "score": score, "chunk_range": (start, end)}
if end == len(text):
break
start += (chunk_size - overlap)
return {"is_leaked": False, "score": 0.0}
π¦ Model Artifacts
model.safetensors: Full FP32 weights (PyTorch).laya.int8.onnx: Weight-only dynamic INT8 quantized model for CPU deployment.rl_agent_config.json: Calibrated inference configuration and temperature scalings.tokenizer/: Fast BPE tokenizer directory.encoder/: Base ModernBERT architecture configuration.
π License
Licensed under the Apache License, Version 2.0.