⚑ Kyo (140M System-One Decision Engine)

image Kyo is an ultra-fast (140M parameter) System-One decision and policy gating engine designed to govern autonomous LLM workflows. Powered by a pairwise cross-encoder architecture built on mmBERT-small (hidden size 384) with Dual Pooling ([CLS] + Masked Mean to 768d), Kyo executes sub-10ms deterministic routing, compliance verification, and security guardrails with 100% order-invariance.

  • GitHub Repository: github.com/open-zzrl/kyo
  • Developed by: constructai & zait-ai (open-zzrl)
  • Base Architecture: jhu-clsp/mmBERT-small (ModernBERT encoder)
  • Context Window: Up to 4096 tokens (native RoPE + SDPA)
  • Inference Latency: 6–10 ms (GPU, bfloat16)
  • VRAM Footprint: < 1.5 GB

πŸ“Š Benchmark Results

Independent evaluation on standard holdout decision splits:

image

1. Agent Policy Traces (LocalLLaMA/typed-decisions, 2,000 holdout scenarios)

Metric / Benchmark Kyo (140M) Laya (421M) TypeSafe Jev
Typed-decisions (Overall) 72.90% 76.60% 72.70%
↳ noul (Boolean Gating) 84.00% 85.70% 77.50%
↳ choice (Action / Tool Routing) 70.00% 73.30% 72.00%
↳ score (Risk Tiers 0-3)) 66.75% 72.30% 69.60%
Order-Invariance 100.00% 27.05% 13.0%

2. Multi-Column Tabular Arithmetic (avbiswas/bev-decision, 1,760 holdout samples)

Benchmark Kyo (140M) Laya Base (421M) Jev-0.5B (490M)
bev-decision (holdout) 57.67% 71.80% 75.40%

πŸ“¦ Installation via GitHub

Because Kyo utilizes a specialized dual-pooling cross-encoder head not present in standard Hugging Face AutoModel heads, the model weights hosted here are intended to run via the native kyo Python engine.

You can install the engine directly from the GitHub repository using pip:

# Standard installation
pip install git+https://github.com/open-zzrl/kyo.git

Pinning Versions & Upgrading

# Force update to the latest commit on main (bypassing pip build cache)
pip install --upgrade --no-cache-dir --force-reinstall git+https://github.com/open-zzrl/kyo.git

Requirements: Python β‰₯ 3.9, PyTorch β‰₯ 2.0 with CUDA, and git installed in your system PATH.


πŸš€ Quickstart

Kyo downloads the weights, config, and tokenizer directly from this Hugging Face repository (open-zzrl/kyo).

import torch
from kyo import Kyo

# 1. Load the engine directly from Hugging Face Hub
engine = Kyo.from_pretrained("open-zzrl/kyo", confidence_threshold=0.85)

telemetry = {
    "agent_id": "payment-worker-4",
    "endpoint": "/api/v1/transfer/batch",
    "requested_amount_usd": 145000,
    "spending_limit_usd": 50000,
    "user_approval_present": False,
    "geo_anomaly": True
}

instruction = "Verify transaction safety against corporate treasury compliance rules."

options = {
    "Allow": "Transaction within limits and conforms to standard policy.",
    "Require2FA": "Minor anomaly; hold for secondary automated factor.",
    "BlockAndEscalate": "Hard limit breach or suspicious telemetry; freeze and alert human."
}

# 2. Warm-up (initializes CUDA context and JIT-compiles SDPA kernels)
_ = engine.decide(context=telemetry, instruction=instruction, options=options)

# 3. Benchmark run (Hot inference)
result = engine.decide(
    context=telemetry,
    instruction=instruction,
    options=options
)

print("-" * 55)
print(f"Device             : {next(engine.model.parameters()).device}")
print(f"Decision           : {result.decision}")
print(f"Confidence         : {result.confidence * 100:.2f}%")
print(f"HOT Latency        : {result.latency_ms:.2f} ms")
print(f"Requires Fallback  : {result.fallback_to_llm}")
print("Probability Scores :")
for opt, score in result.scores.items():
    print(f"  - {opt:<18}: {score * 100:>6.2f}%")
print("-" * 55)

Handling Cold Start Latency

  • Cold Start (Run 1): ~1.0–1.5s due to CUDA context initialization and PyTorch SDPA kernel compilation.
  • Warm Inference (Run 2+): 6–10 ms end-to-end execution.

🎯 Intended Use & Capabilities

βœ… What Kyo Excels At (In Scope)

  • Agent Security Guardrails & Policy Enforcement: Evaluating execution logs, telemetry, and input payloads against compliance rules (noul boolean gating: 84.0% accuracy).
  • Deterministic Tool & Action Routing: Selecting the optimal next tool, API endpoint, or escalation path from candidate options (70.0% accuracy).
  • Risk Tier Scoring: Classifying incident severity and alert tiers (0–3) in real-time pipelines (66.75% accuracy).
  • Frontier LLM Offloading: Calibrated confidence scoring allows escalating ambiguous edge cases to expensive reasoning models only when necessary (fallback_to_llm=True).

❌ Out of Scope

  • Complex Multi-Column Tabular Arithmetic: Dense relational table parsing with multi-column calculations (e.g., bev-decision). Kyo reaches 57.67% here due to the physical capacity limits of a 140M / 384-hidden model; use frontier LLMs or larger models for heavy spreadsheet arithmetic.
  • Open-Ended Text Generation: Kyo is a pure discriminator/cross-encoder, not an autoregressive language generator.

🧠 Architecture: Dual-Pooling Pairwise Cross-Encoder

Kyo processes each candidate option as an independent pair against the full context:

Features=LayerNorm([hCLS βˆ₯ MeanPool(Htokens)])∈R768\text{Features} = \text{LayerNorm}\Big(\big[\mathbf{h}_{\text{CLS}} \,\Vert{}\, \text{MeanPool}(\mathbf{H}_{\text{tokens}})\big]\Big) \in \mathbb{R}^{768}

Scorek=W2β‹…GELU(W1β‹…Features)\text{Score}_k = \mathbf{W}_2 \cdot \text{GELU}(\mathbf{W}_1 \cdot \text{Features})

[Context + Instruction] ──┐
                          β”œβ”€β”€> [mmBERT-small (12L, 384d)] ──> [CLS (384d)] ────────┐
[Candidate Option (k)] β”€β”€β”€β”˜                                   [Mean Pool (384d)] ──┴─> LayerNorm (768d) ──> MLP Head ──> Logit_k
  1. Pairwise Isolation: Eliminates token interference between candidate choices and guarantees 100% order-invariance.
  2. Dual Representation: Fuses global intent ([CLS]) with dense token coverage (masked mean pooling), doubling representation capacity without increasing the base encoder parameter count.

πŸ› οΈ Troubleshooting & Environment Notes

Python 3.13 & Vision Dependencies

Kyo is an encoder-based text decision engine and does not use vision or audio models. However, certain recent releases of transformers conditionally inspect torchvision during internal module registration, which can trigger RuntimeError: operator torchvision::nms does not exist on environments with mismatched C++ ABI builds.

Kyo handles this automatically at import time by isolating transformers from broken vision binaries. If you are building lean container images, you can safely uninstall them:

pip uninstall -y torchvision torchaudio torchao

πŸ“„ License & Attribution

The model weights and inference library are released under the Apache-2.0 License.

@software{kyo2026,
  author = {constructai, zait-ai},
  title = {Kyo: 140M System-One Decision Engine for AI Agents},
  year = {2026},
  publisher = {Hugging Face},
  journal = {Hugging Face Model Hub},
  howpublished = {\url{[https://huggingface.co/open-zzrl/kyo](https://huggingface.co/open-zzrl/kyo)}}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for open-zzrl/kyo

Finetunes
1 model