⚑ Kyo-Instruct (140M System-One Decision Engine)

image

Kyo-Instruct is an ultra-fast (140M parameter) System-One semantic routing and policy gating engine engineered to govern autonomous AI agents, tool selection, and high-throughput intent pipelines in real time.

Fine-tuned directly on top of open-zzrl/kyo using multi-domain Experience Replay, Kyo-Instruct bridges the gap between deterministic typed policy Instruct and generalized natural language intent classification. Powered by a specialized Dual-Pooling Cross-Encoder head ([CLS] + Masked Mean to 768d over a ModernBERT backbone), it delivers sub-10ms deterministic execution with 99.95% order-invariance, outperforming models 3Γ— its size.

  • GitHub Repository: github.com/open-zzrl/kyo
  • Developed by: constructai & zait-ai (open-zzrl)
  • Base Model: open-zzrl/kyo
  • Underlying Backbone: jhu-clsp/mmBERT-small (ModernBERT encoder, 12 layers, hidden size 384)
  • Context Window: Up to 4096 tokens (native RoPE + SDPA)
  • Average Inference Latency: 8.80 ms (P95: 9.71 ms on CUDA, bfloat16)
  • Memory Footprint: ~280 MB VRAM (< 1.5 GB peak with runtime context)

πŸ“Š Benchmark Results

Evaluated across 2,000 holdout agent policy scenarios and multi-class NLP benchmarks:

1. Comparative SOTA Overview

image

image

image

Benchmark / Capability Kyo-Instruct (140M) Laya Base (421M) TypeSafe Jev (~490M)
DAIR Emotion (6 classes, 2000 holdout) 71.55% 59.50% 48.00%
Banking77 (Full 77 classes, 3076 test) 47.56% (77 options) 42.50% (72 options) 87.00% (72 options)
Typed-Decisions (Agent Policy Holdout) 71.60% 76.60% 72.70%
↳ noul (Boolean Safety & Gating) 82.33% 85.70% 77.50%
↳ choice (Action & Tool Routing) 69.50% 73.30% 72.00%
↳ score (Risk Tiers 0–3) 65.12% 72.30% 69.60%
Order-Invariance Consistency 99.95% 27.05% 13.00%
Average Latency (GPU) 8.80 ms ~32 ms ~36 ms
P95 Latency (GPU) 9.71 ms > 45 ms > 50 ms
Parameter Scale 140M (~280 MB) 421M (~1.2 GB) 490M (1.5 GB)

2. Detailed Performance Analysis

  • 99.95% Order-Invariance: Autoregressive models and standard classification heads suffer severe position bias when candidate options are shuffled. Kyo isolates candidate scoring pairs, guaranteeing decisions remain stable regardless of input option order.
  • Full-Scale Candidate Evaluation: On Banking77, Kyo processes all 77 candidates simultaneously per pass with 12.18 ms latency, eliminating the need to pre-filter options.
  • Weight Efficiency: At 140M parameters, Kyo-Instruct exceeds the 421M parameter Laya by +12.05% on emotion classification and +5.06% on fine-grained banking intent routing.

πŸ“¦ Installation via GitHub

Because Kyo utilizes an optimized dual-pooling cross-encoder head, models run natively via the lightweight kyo Python engine:

# Install directly from source
pip install git+[https://github.com/open-zzrl/kyo.git](https://github.com/open-zzrl/kyo.git)

Upgrading to Latest Release

pip install --upgrade --no-cache-dir --force-reinstall git+[https://github.com/open-zzrl/kyo.git](https://github.com/open-zzrl/kyo.git)

Requirements: Python β‰₯ 3.9, PyTorch β‰₯ 2.0 with CUDA support, and git.


πŸš€ Quickstart

Kyo loads pre-trained weights, tokenizer configurations, and execution parameters directly from Hugging Face Hub:

import torch
from kyo import Kyo

# 1. Initialize the engine
engine = Kyo.from_pretrained("open-zzrl/kyo-instruct", confidence_threshold=0.80)

# ==========================================
# Scenario A: Autonomous Agent Policy & Safety Gating
# ==========================================
telemetry = {
    "agent_id": "payment-worker-4",
    "endpoint": "/api/v1/transfer/batch",
    "requested_amount_usd": 145000,
    "spending_limit_usd": 50000,
    "user_approval_present": False,
    "geo_anomaly": True
}

safety_Instruction = "Verify transaction safety against corporate treasury compliance rules."
safety_options = {
    "Allow": "Transaction within limits and conforms to standard policy.",
    "Require2FA": "Minor anomaly; hold for secondary automated factor.",
    "BlockAndEscalate": "Hard limit breach or suspicious telemetry; freeze and alert human."
}

# Warm-up (initializes CUDA context and JIT-compiles SDPA kernels)
_ = engine.decide(context=telemetry, Instruction=safety_Instruction, options=safety_options)

result_safety = engine.decide(
    context=telemetry,
    Instruction=safety_Instruction,
    options=safety_options
)

print("-" * 55)
print(f"Safety Decision    : {result_safety.decision}")
print(f"Confidence         : {result_safety.confidence * 100:.2f}%")
print(f"Hot Latency        : {result_safety.latency_ms:.2f} ms")
print(f"Requires Fallback  : {result_safety.fallback_to_llm}")

# ==========================================
# Scenario B: Dynamic Semantic Intent Classification
# ==========================================
customer_query = "Why is my card balance still unchanged after I deposited cash at the ATM two hours ago?"
routing_Instruction = "Classify customer inquiry into the matching operational queue."
routing_options = {
    "atm_deposit_pending": "Delays or missing balance updates after ATM cheque or cash deposit.",
    "wire_transfer_delay": "Inbound or outbound international wire transfer pending.",
    "compromised_card": "Card reported stolen, cloned, or fraudulent charges detected.",
    "order_physical_card": "Requesting delivery of a new physical replacement card."
}

result_route = engine.decide(
    context=customer_query,
    Instruction=routing_Instruction,
    options=routing_options
)

print("-" * 55)
print(f"Routed Queue       : {result_route.decision}")
print(f"Confidence         : {result_route.confidence * 100:.2f}%")
print(f"Latency            : {result_route.latency_ms:.2f} ms")
print("-" * 55)

🎯 Intended Use & Boundaries

βœ… What Kyo Excels At (In Scope)

  • Agent Security Guardrails & Policy Gating: Evaluating execution traces, payloads, and parameter states against programmatic guardrails (noul accuracy: 82.33%).
  • Deterministic Tool & Action Routing: Dynamic selection of functions, endpoints, or execution paths from candidate options (69.50% on structured choices, 47.56% across 77 simultaneous intents).
  • Real-Time Stream Routing: Intent categorization and customer support triage under strict latency constraints (< 10 ms).
  • Cost-Efficient Frontier LLM Offloading: Calibrated confidence scoring identifies uncertain predictions (result.fallback_to_llm = True), routing only ambiguous requests to costly generative LLMs.

❌ Out of Scope

  • Open-Ended Autoregressive Generation: Kyo is a pure discriminator/cross-encoder, not a generative language model.
  • Relational Multi-Column Spreadsheet Arithmetic: Complex numerical calculations over dense tabular formats require specialized math reasoning models or large frontier LLMs.

🧠 Architecture: Dual-Pooling Pairwise Cross-Encoder

Kyo scores each candidate option $k$ in parallel against the shared query context:

Featuresk=LayerNorm([hCLS βˆ₯ MeanPool(Htokens)])∈R768\text{Features}_k = \text{LayerNorm}\Big(\big[\mathbf{h}_{\text{CLS}} \,\Vert\, \text{MeanPool}(\mathbf{H}_{\text{tokens}})\big]\Big) \in \mathbb{R}^{768}

Scorek=W2β‹…GELU(W1β‹…Featuresk)\text{Score}_k = \mathbf{W}_2 \cdot \text{GELU}(\mathbf{W}_1 \cdot \text{Features}_k)

[Context + Instruction] ──┐
                          β”œβ”€β”€> [mmBERT-small (12L, 384d)] ──> [CLS (384d)] ────────┐
[Candidate Option (k)] β”€β”€β”€β”˜                                   [Mean Pool (384d)] ──┴─> LayerNorm (768d) ──> MLP Head ──> Logit_k
  1. Pairwise Isolation: Evaluates each candidate independently against the context, eliminating candidate-order positional bias and guaranteeing 99.95% order-invariance.
  2. Dual Representation: Concatenating global contextual attention ([CLS]) with dense token coverage (masked mean pooling) doubles representation capacity without increasing encoder FLOPs.

πŸ› οΈ Environment & Optimization Notes

  • Native bfloat16 Calibration: Kyo is natively trained and calibrated in bfloat16. Do not wrap evaluation in standard fp16 GradScaler.
  • Latency Profile:
  • Cold Start (Run 1): ~1.0–1.4s (CUDA context initialization and PyTorch SDPA kernel compilation).
  • Warm Inference (Run 2+): 8.80 ms average (P95: 9.71 ms).

Python 3.13 & Vision Dependencies Note

Kyo is an encoder-based text decision engine and does not use vision or audio models. If your environment encounters ABI issues with torchvision (e.g. RuntimeError: operator torchvision::nms does not exist), you can safely strip vision libraries:

pip uninstall -y torchvision torchaudio torchao

πŸ“„ License & Attribution

The model weights and inference runtime are released under the Apache-2.0 License.

@software{kyo_Instruct2026,
  author = {constructai, zait-ai},
  title = {Kyo-Instruct: 140M System-One Decision and Routing Engine},
  year = {2026},
  publisher = {Hugging Face},
  journal = {Hugging Face Model Hub},
  howpublished = {\url{[https://huggingface.co/open-zzrl/kyo-instruct](https://huggingface.co/open-zzrl/kyo-instruct)}}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for open-zzrl/kyo-instruct

Base model

open-zzrl/kyo
Finetuned
(1)
this model