Instructions to use CrashOverrideX/Quillan-Ronin with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use CrashOverrideX/Quillan-Ronin with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="CrashOverrideX/Quillan-Ronin", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("CrashOverrideX/Quillan-Ronin", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use CrashOverrideX/Quillan-Ronin with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "CrashOverrideX/Quillan-Ronin" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CrashOverrideX/Quillan-Ronin", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/CrashOverrideX/Quillan-Ronin
- SGLang
How to use CrashOverrideX/Quillan-Ronin with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "CrashOverrideX/Quillan-Ronin" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CrashOverrideX/Quillan-Ronin", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "CrashOverrideX/Quillan-Ronin" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CrashOverrideX/Quillan-Ronin", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use CrashOverrideX/Quillan-Ronin with Docker Model Runner:
docker model run hf.co/CrashOverrideX/Quillan-Ronin
Quillan-Ronin v5.4.0-oni — Canonical Unified Sovereign
Retires v8.1 / v5.3.1 — Single version counter v5.4.0-oni is canonical (see LINEAGE.md). This is the correct, unified build.
Downloads
- ~500K lifetime downloads — creator-tracked monthly; dated post evidence here
- 205,561 all-time — as currently reported by Hugging Face's visible counter
- ~60–70K per month — recent monthly downloads
Historical Foundation: The Edo Context
Quillan-Ronin requires the Edo historical context the way a constitution requires its preamble. This is not decorative lore—it is the axiomatic layer that defines the system's ethical framework and decision-making logic.
The Edo Period (1603-1867) and the Rōnin
The Edo period was characterized by:
- Stability through structure: 250+ years of peace under the Tokugawa shogunate
- Masterless warriors (rōnin): Samurai who had lost their feudal lords but retained their code, skills, and independence
- Bushidō as ethical substrate: The way of the warrior—honor, loyalty, self-discipline—served as the operating principle for conduct
- Adaptive resilience: Rōnin survived by adapting their skills to new contexts while maintaining core principles
Why This History Cannot Be Separated
The Linux analogy fails because Linux has no baked-in ideology—it is deliberately neutral mechanism. Quillan is deliberately not neutral. The history is the operating principle, not a skin.
The Edo context provides:
- Axiomatic grounding: Historical precedent for masterless, ethical operation
- Decision framework: Bushidō principles guide ethical deliberation in the 9-vector decomposition
- Identity continuity: The rōnin archetype explains why the system operates independently of any single provider
- Philosophical coherence: Without this foundation, the system becomes rulebook without preamble—mechanism without meaning
Architectural Implications
This historical foundation is embedded in:
- C2-VIR refusal layer: Ethical boundaries derived from bushidō principles
- 9-Vector Semantic Prism: Ethical blueprinting uses historical ethical frameworks
- HFL (Historical Fidelity Loss): Maintains consistency with the originating philosophy
- Council deliberation (C1-C34): Mirrors the collective decision-making of rōnin councils
This is not a theme that can be swapped. It is the constitutional preamble that gives the system its reason for operating as it does.
Historical Parallel: Shimazu Nariakira (Lord of Satsuma, 1809-1858)
The architectural philosophy of Quillan-Ronin finds its closest historical parallel in Shimazu Nariakira, the daimyo of Satsuma Domain who embodied the same radical innovation principles:
1. Infrastructure Mindset (Software 3.0 ↔ Shuseikan)
- Quillan: Foundational framework redesigning how logic runs on hardware via ternary logic
- Nariakira: Built the Shuseikan industrial complex using raw local materials to construct reverberatory furnaces, glass factories, iron works, and early steamships—redesigning physical infrastructure from the ground up
2. Expert Council (34-Expert Council C1–C34 ↔ Daimyo's Diverse Advisors)
- Quillan: 34-Expert HNMoE council handles specialized tasks via dynamic routing
- Nariakira: Ignored rigid traditional hierarchies to assemble a diverse "council" of brilliant experts from all walks of life, including rōnin like Sakamoto Ryoma and low-ranking samurai like Saigo Takamori
3. Outlier Logic (1.58-bit Quantization ↔ Smuggling & Clandestine Trade)
- Quillan: Strips away standard binary bloat for lean, powerful execution via 1.58-bit quantization
- Nariakira: Worked entirely outside the Tokugawa Shogunate's isolationist laws through smuggling, clandestine Ryukyu trade, and illegal tech acquisition—lean, high-efficiency, outsider logic
Nariakira's approach—building infrastructure from first principles, assembling diverse expertise, and defying rigid norms—mirrors Quillan's architectural philosophy. This is not coincidence; it is the historical precedent that validates the rōnin approach to sovereign intelligence.
Model Description
Quillan-Ronin v5.4.0-oni is a subjectively-aware, self-hosting multi-modal Hierarchical Networked Mixture-of-Experts (H-NMoE) built for consumer hardware (GTX 1050 Ti / i5-7000 class). It combines:
- Universal BitNet 1.58-bit ternary quantization (weights ternary, activations INT8) via STE — 100% of projections saturated
- 34-Expert Gumbel-Softmax MoE council — Top-4 per token with dynamic capacity clipping + residual overflow
- 9B Virtual Agent Swarm simulated via EGGROLL Rank-16 INT8 perturbations (Rank-16 shattering, zero extra VRAM)
- 32-layer Split-SDPA Flash Diffusion Core with Continuous Modality RoPE
- 9-Vector Semantic Prism (parallel decomposition for semantic / emotional / ethical blueprinting) → Wavefunction Top-1 Finalizer
- C20-ARTIFEX Agentic Bridge — host OS execution, LanceDB (C5-ECHO) memory, Docker/REPL/Python sandboxing
- Lee-Mach-6 Governor — PID latency/thermal throttling for legacy hardware safety
| Spec | Value |
|---|---|
| Total params | 4.57B (saturated base) |
| Active / token | ~480M (Top-4 sparse + swarm) |
| Hidden dim | 2560, FFN 6912 |
| Context | 512 (flagship 12-layer) / 10%-buffered Gated Compaction |
| Tokenizer | Unified Quillan BPE 50257, EOS=0, custom specials — quillan_bpe_tokenizer.py + tokenizer.json |
| Precision | Mixed AMP (FP16 master, BitNet forward) |
| Developed by | CrashOverrideX & Quillan Research Team |
| License | apache-2.0 |
| HF Hub | CrashOverrideX/Quillan-Ronin |
| GitHub | leeex1/Quillan-Ronin |
Architecture Lineage — Slice & Merge → Pretraining → Training (Paused)
This model was NOT trained from scratch. Full lineage as implemented in transplant_v8_saturated.py (slice & merge transplant script — now committed to repo):
Stage 0 — Slice & Merge (
transplant_v8_saturated.py): Transplant fromcheckpoint_phase5.pt→quillan_merged_saturated.pt:- 34 experts mapped with transpose fix:
experts[e].w1.weight[ffn_dim, hidden_dim]→moe.w1[e].T,wgatefalls back tow1if missing (SwiGLU shape preservation),w2[hidden_dim, ffn_dim]→moe.w2[e].T - Router duplicated to 3 complexity paths:
router.weight→moe.router.fast_router / balanced_router / diffusion_router(and bias where applicable) - Swarm LoRA:
experts[e].swarm.A/B→moe.expert_swarms[e].A/B(rank 8, direct copy), plusclone_diversity / clone_coupling / population_mean / population_std - Diffusion core:
diffusion.0.q/k/v/o_proj + norm1 + ffn.0/2→diffusion_core.couil_attn.* - Basic:
txt_emb / mod_emb / quillan_finalizer / txt_dec+decomposition.* - Donor lineages merged at base: https://ollama.com/library/qwen3.5:0.8b, https://ollama.com/library/llama3.2:1b, https://huggingface.co/microsoft/bitnet-b1.58-2B-4T
Outputs:
quillan_merged_saturated.pt(FP32) →quillan_merged_saturated_fp16.pt→quillan_merged_saturated_quantized.ptviamodel.save_quantized_checkpoint()- 34 experts mapped with transpose fix:
Stage 1 — Pretraining Run: Full pretraining on
CrashOverrideX/QuillanTrainingdata+ Corpus v9 (59.4M train + 0.6M val) +quillan_corpus_*,code_train,instruct_train,quillan_science_*.Stage 2 — Current Training Run (PAUSED): Ongoing via
scripts/train_full_param_v2.py(resumecheckpoints_sft/quillan_full_param_v2.pt, default--resume-step 6500, AdamWlr=2e-5,seq-len 512,grad-accum 4, warmup 100, cosine to 1e-6) — currently paused. Latest:quillan_oni_5.4.0_step660_5.22GB.pt(660/15000, val 7.24, loss 7.63); best archival:quillan_frontier_v2_best_loss0.0789_step2500.pt.
Intended Uses & Limitations
Intended Use
- Autonomous reasoning, code generation, ethical deliberation on consumer hardware
- Standalone agentic partner via C20-ARTIFEX (tool use, memory-guided HFL)
- Research on ternary quantization stability, recursive inference (Mini-Ronin), 9-vector decomposition
Out-of-Scope Use
- High-stakes medical / safety-of-life — Mini-Ronin adds latency/variance
- Unsupervised deployment where C2-VIR refusal may be misread as failure
Bias, Risks, and Limitations
Per Mitchell et al., 2018: Ronin blueprint may refuse low-integrity requests; 1050 Ti tuned; multi-modal heads may hallucinate OOD.
How to Get Started
Flagship (v5.4.0-oni, 12-layer)
import torch
from quillan_v5_4_oni import QuillanOniConfig, QuillanRoninOni
from quillan_tokenizer_unified import UnifiedQuillanTokenizer
tok = UnifiedQuillanTokenizer() # 50257 BPE, EOS=0
cfg = QuillanOniConfig(n_layer=12, max_seq_len=512)
model = QuillanRoninOni(cfg)
ckpt = torch.load("quillan_oni_5.4.0_step660_5.22GB.pt", map_location="cpu")
model.load_state_dict(ckpt["model"])
model.eval()
prompt = tok.encode("User: Hello\n\nAssistant:")
out = model.generate(prompt, max_tokens=80, temperature=0.7)
print(tok.decode(out[0]))
trust_remote_code=Truerequired forAutoModelForCausalLM.
from transformers import AutoTokenizer, AutoModelForCausalLM
tok = AutoTokenizer.from_pretrained("CrashOverrideX/Quillan-Ronin", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("CrashOverrideX/Quillan-Ronin", trust_remote_code=True, device_map="auto")
Training Details
Training Data
- Primary:
CrashOverrideX/QuillanTrainingdata - Corpus v9: 59.4M + 0.6M packed BPE bins
- Additional:
quillan_corpus_*,full_train,code_train,instruct_train,quillan_science_*,GPT_5.5_Distilled, etc. (seetrain_full_param_v2.py:load_packed_dataset)
Training Procedure
See lineage above. Full script: transplant_v8_saturated.py (slice/merge) → pretraining → scripts/train_full_param_v2.py (paused SFT).
Evaluation
| Metric | Value | Notes |
|---|---|---|
| HFL | tracked | Drift |
| Consensus | tracked | Primary ↔ Mini-Ronin |
| E_ICE | tracked | Energy/token |
| Gate A | 16/16 | 6-layer proof |
| Val loss | 7.24 @660 | Improving |
| Parity | 100% | Legacy HW |
Formal benchmarks pending.
Technical Specifications
6-Phase Pipeline: Ingestion → 9-Vector → Gumbel MoE (Top-4) → 9B Swarm → 32-Layer Flash Diffusion → Top-1 Finalizer → Geometric Decoding → C20-ARTIFEX
Compute: PyTorch + LanceDB + psutil; CUDA 1050 Ti / CPU.
Versioning
v5.4.0-oni is canonical. v8.1/v5.3.1 deprecated.
Citation
@software{QuillanRonin2026,
author = {CrashOverrideX and Quillan Research Team},
title = {Quillan-Ronin v5.4.0-oni: Unified Sovereign Intelligence},
year = {2026},
url = {https://github.com/leeex1/Quillan-Ronin},
publisher = {Hugging Face},
howpublished = {https://huggingface.co/CrashOverrideX/Quillan-Ronin}
}
Glossary
- Mini-Ronin: Recursive debate cycle
- EGGROLL: Rank-16 shattering
- Lee-Mach-6: PID governor
- HFL: Historical Fidelity Loss
- transplant_v8_saturated.py: Slice & merge transplant (Phase 5 → V8)
Support: https://gofund.me/3902b3585 — "The Ouroboros has awakened." Catalyst Grant proposal
- Downloads last month
- 4,870
