Instructions to use open-zzrl/kyo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use open-zzrl/kyo with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="open-zzrl/kyo")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("open-zzrl/kyo", device_map="auto") - Notebooks
- Google Colab
- Kaggle
β‘ Kyo (140M System-One Decision Engine)
Kyo is an ultra-fast (140M parameter) System-One decision and policy gating engine designed to govern autonomous LLM workflows. Powered by a pairwise cross-encoder architecture built on mmBERT-small (hidden size 384) with Dual Pooling ([CLS] + Masked Mean to 768d), Kyo executes sub-10ms deterministic routing, compliance verification, and security guardrails with 100% order-invariance.
- GitHub Repository: github.com/open-zzrl/kyo
- Developed by: constructai & zait-ai (
open-zzrl) - Base Architecture:
jhu-clsp/mmBERT-small(ModernBERT encoder) - Context Window: Up to 4096 tokens (native RoPE + SDPA)
- Inference Latency: 6β10 ms (GPU,
bfloat16) - VRAM Footprint: < 1.5 GB
π Benchmark Results
Independent evaluation on standard holdout decision splits:
1. Agent Policy Traces (LocalLLaMA/typed-decisions, 2,000 holdout scenarios)
| Metric / Benchmark | Kyo (140M) | Laya (421M) | TypeSafe Jev |
|---|---|---|---|
| Typed-decisions (Overall) | 72.90% | 76.60% | 72.70% |
| β³ noul (Boolean Gating) | 84.00% | 85.70% | 77.50% |
| β³ choice (Action / Tool Routing) | 70.00% | 73.30% | 72.00% |
| β³ score (Risk Tiers 0-3)) | 66.75% | 72.30% | 69.60% |
| Order-Invariance | 100.00% | 27.05% | 13.0% |
2. Multi-Column Tabular Arithmetic (avbiswas/bev-decision, 1,760 holdout samples)
| Benchmark | Kyo (140M) | Laya Base (421M) | Jev-0.5B (490M) |
|---|---|---|---|
| bev-decision (holdout) | 57.67% | 71.80% | 75.40% |
π¦ Installation via GitHub
Because Kyo utilizes a specialized dual-pooling cross-encoder head not present in standard Hugging Face AutoModel heads, the model weights hosted here are intended to run via the native kyo Python engine.
You can install the engine directly from the GitHub repository using pip:
# Standard installation
pip install git+https://github.com/open-zzrl/kyo.git
Pinning Versions & Upgrading
# Force update to the latest commit on main (bypassing pip build cache)
pip install --upgrade --no-cache-dir --force-reinstall git+https://github.com/open-zzrl/kyo.git
Requirements: Python β₯ 3.9, PyTorch β₯ 2.0 with CUDA, and
gitinstalled in your system PATH.
π Quickstart
Kyo downloads the weights, config, and tokenizer directly from this Hugging Face repository (open-zzrl/kyo).
import torch
from kyo import Kyo
# 1. Load the engine directly from Hugging Face Hub
engine = Kyo.from_pretrained("open-zzrl/kyo", confidence_threshold=0.85)
telemetry = {
"agent_id": "payment-worker-4",
"endpoint": "/api/v1/transfer/batch",
"requested_amount_usd": 145000,
"spending_limit_usd": 50000,
"user_approval_present": False,
"geo_anomaly": True
}
instruction = "Verify transaction safety against corporate treasury compliance rules."
options = {
"Allow": "Transaction within limits and conforms to standard policy.",
"Require2FA": "Minor anomaly; hold for secondary automated factor.",
"BlockAndEscalate": "Hard limit breach or suspicious telemetry; freeze and alert human."
}
# 2. Warm-up (initializes CUDA context and JIT-compiles SDPA kernels)
_ = engine.decide(context=telemetry, instruction=instruction, options=options)
# 3. Benchmark run (Hot inference)
result = engine.decide(
context=telemetry,
instruction=instruction,
options=options
)
print("-" * 55)
print(f"Device : {next(engine.model.parameters()).device}")
print(f"Decision : {result.decision}")
print(f"Confidence : {result.confidence * 100:.2f}%")
print(f"HOT Latency : {result.latency_ms:.2f} ms")
print(f"Requires Fallback : {result.fallback_to_llm}")
print("Probability Scores :")
for opt, score in result.scores.items():
print(f" - {opt:<18}: {score * 100:>6.2f}%")
print("-" * 55)
Handling Cold Start Latency
- Cold Start (Run 1): ~1.0β1.5s due to CUDA context initialization and PyTorch SDPA kernel compilation.
- Warm Inference (Run 2+): 6β10 ms end-to-end execution.
π― Intended Use & Capabilities
β What Kyo Excels At (In Scope)
- Agent Security Guardrails & Policy Enforcement: Evaluating execution logs, telemetry, and input payloads against compliance rules (
noulboolean gating: 84.0% accuracy). - Deterministic Tool & Action Routing: Selecting the optimal next tool, API endpoint, or escalation path from candidate options (70.0% accuracy).
- Risk Tier Scoring: Classifying incident severity and alert tiers (0β3) in real-time pipelines (66.75% accuracy).
- Frontier LLM Offloading: Calibrated confidence scoring allows escalating ambiguous edge cases to expensive reasoning models only when necessary (
fallback_to_llm=True).
β Out of Scope
- Complex Multi-Column Tabular Arithmetic: Dense relational table parsing with multi-column calculations (e.g.,
bev-decision). Kyo reaches 57.67% here due to the physical capacity limits of a 140M / 384-hidden model; use frontier LLMs or larger models for heavy spreadsheet arithmetic. - Open-Ended Text Generation: Kyo is a pure discriminator/cross-encoder, not an autoregressive language generator.
π§ Architecture: Dual-Pooling Pairwise Cross-Encoder
Kyo processes each candidate option as an independent pair against the full context:
[Context + Instruction] βββ
βββ> [mmBERT-small (12L, 384d)] ββ> [CLS (384d)] βββββββββ
[Candidate Option (k)] ββββ [Mean Pool (384d)] βββ΄β> LayerNorm (768d) ββ> MLP Head ββ> Logit_k
- Pairwise Isolation: Eliminates token interference between candidate choices and guarantees 100% order-invariance.
- Dual Representation: Fuses global intent (
[CLS]) with dense token coverage (masked mean pooling), doubling representation capacity without increasing the base encoder parameter count.
π οΈ Troubleshooting & Environment Notes
Python 3.13 & Vision Dependencies
Kyo is an encoder-based text decision engine and does not use vision or audio models. However, certain recent releases of transformers conditionally inspect torchvision during internal module registration, which can trigger RuntimeError: operator torchvision::nms does not exist on environments with mismatched C++ ABI builds.
Kyo handles this automatically at import time by isolating transformers from broken vision binaries. If you are building lean container images, you can safely uninstall them:
pip uninstall -y torchvision torchaudio torchao
π License & Attribution
The model weights and inference library are released under the Apache-2.0 License.
@software{kyo2026,
author = {constructai, zait-ai},
title = {Kyo: 140M System-One Decision Engine for AI Agents},
year = {2026},
publisher = {Hugging Face},
journal = {Hugging Face Model Hub},
howpublished = {\url{[https://huggingface.co/open-zzrl/kyo](https://huggingface.co/open-zzrl/kyo)}}
}
