AENEA Pinta-1.2 Mini Stable
50M-parameter semantic router for edge dispatch. Pinta-1.2 Mini Stable is the production-ready successor to the Pinta-1.1 Mini Beta, delivering 91.51% strict accuracy - a massive +19 point improvement over its predecessor - while maintaining the same 50M parameter budget and sub-15 ms CPU inference on commodity hardware.
Designed for always-on classification in agent swarms, CLI sidecars, and edge dispatch where latency matters more than peak accuracy.
Stable release notice. This is a Stable release. The routing head was trained on a strictly balanced 58,220-record dataset (1.46x max/min ratio across all 9 domains). Unlike the Beta, no calibration bias vector is required at inference time. The model can be used directly with raw softmax over the 9 domain logits.
What is this?
Pinta-1.2 Mini is not a generative language model. It is a semantic router that classifies incoming prompts into one of nine dispatch domains and emits a routing token that downstream systems can use to select the appropriate expert model or tool:
| Class | Token | Dispatch to |
|---|---|---|
| SyntaxDevOps | <|reserved_23|> |
Template engines, formatters, regex, data conversion |
| Code | <|reserved_24|> |
Code-specialized models, debuggers, linters |
| Creative | <|reserved_25|> |
Creative writing models, storytelling engines |
| RAG | <|reserved_26|> |
Retrieval-augmented pipelines, document QA |
| Architecture | <|reserved_27|> |
System design experts, DevOps, cloud infra |
| Math | <|reserved_28|> |
Math/reasoning models, calculators, CAS |
| Knowledge | <|reserved_29|> |
General knowledge / fallback models |
| Compliance | <|reserved_30|> |
Legal domain models, policy engines, compliance review |
| Chat | <|reserved_31|> |
Chat-optimized models, conversational agents |
The router is trained to predict which token a full generative model would emit next, but does so in a single forward pass without generating any text.
Architecture
| Component | Specification |
|---|---|
| Parameters | 50M |
| Layers | 20 (Cartan/Cittern) |
| Hidden size | 512 |
| Attention heads | 8 |
| Orthogonal attention | Block Householder, 8 reflections, block_size 64 |
| Vocabulary | 9,000 (QT-Cittern-1.0 tokenizer) |
| Context window | 2,048 tokens |
| Precision | bf16 / fp16 / ONNX fp32 & fp16 |
| Routing head | Causal [B, S, V] with 9 reserved token slice |
Benchmarks
Evaluated on the Balanced Routing Benchmark v6 (1,060 prompts across 9 domains, zero train overlap, English-only):
| Metric | Pinta-1.1 (226M) | Pinta-1.1 Mini Beta | Pinta-1.2 Mini Stable |
|---|---|---|---|
| Strict accuracy | 79.22% | 72.50% | 91.51% |
| Macro F1 | - | - | 0.91 |
| Weighted F1 | - | - | 0.92 |
| Latency (p50, CPU) | ~45 ms | ~14 ms | ~14 ms |
| Latency (p95, CPU) | ~52 ms | ~16 ms | ~18 ms |
| Model size (ONNX fp32) | 452 MB | 210 MB | 199.9 MB |
| Model size (ONNX fp16) | - | - | 105.3 MB |
Pinta-1.2 Mini outperforms the full 226M Pinta-1.1 by 12+ percentage points while remaining smaller and 3x faster.
Per-Domain Performance
| Domain | Precision | Recall | F1 |
|---|---|---|---|
| SYNTAX_DEV_OPS | 0.96 | 0.73 | 0.83 |
| CODE | 0.80 | 0.86 | 0.83 |
| CREATIVE_PROSE | 0.96 | 0.83 | 0.89 |
| LONG_CONTEXT_RAG | 0.98 | 0.93 | 0.96 |
| SYSTEM_CODE_ARCH | 0.90 | 0.95 | 0.93 |
| DENSE_STEM_MATH | 0.94 | 0.99 | 0.96 |
| GENERAL_KNOWLEDGE | 0.76 | 0.97 | 0.85 |
| HIGH_RISK_COMPLIANCE | 1.00 | 0.97 | 0.99 |
| CONVERSATIONAL_CHAT | 0.98 | 0.97 | 0.98 |
Quickstart
Python Inference (ONNX)
import numpy as np
import onnxruntime as ort
from transformers import AutoTokenizer
# Load artifacts
tokenizer = AutoTokenizer.from_pretrained("tokenizers/QT-Cittern-1.0")
session = ort.InferenceSession("pinta-1.2-mini.fp16.opset18.onnx")
DOMAINS = [f"<|reserved_{i}|>" for i in range(23, 32)]
domain_ids = [tokenizer.convert_tokens_to_ids(d) for d in DOMAINS]
def route(prompt: str) -> dict:
formatted = f"User: {prompt}\nAssistant: "
ids = tokenizer.encode(formatted, return_tensors="np", add_special_tokens=False).astype(np.int64)
logits = session.run(None, {"input_ids": ids})[0][0, -1, :]
scores = logits[domain_ids]
probs = np.exp(scores - scores.max())
probs /= probs.sum()
best = int(np.argmax(probs))
return {
"token": DOMAINS[best],
"confidence": float(probs[best]),
"all_scores": {d: float(p) for d, p in zip(DOMAINS, probs)}
}
result = route("Design a microservices architecture for our payment system")
print(f"Route to: {result['token']} (confidence: {result['confidence']:.3f})")
# Output: Route to: <|reserved_27|> (confidence: 0.991)
C++ Edge Router
A production-grade, zero-dependency C++ router is included for edge deployment. It auto-detects paths and requires no calibration:
# Build (MSVC x64)
cl /std:c++17 /O2 /EHsc /I"path\to\ort\include" pinta_router.cpp /link /LIBPATH:"path\to\ort\lib" onnxruntime.lib
# Run (auto-loads model and tokenizer from exe directory)
pinta_router.exe "Design a microservices architecture"
Confidence Gate
We recommend a confidence threshold of 0.35. If the model's max probability falls below this, fall back to <|reserved_29|> (Knowledge) rather than emitting a low-confidence misroute:
if result["confidence"] < 0.35:
result["token"] = "<|reserved_29|>" # Knowledge fallback
Repository Layout
aenea-pinta-1.2-mini/
βββ pinta-1.2-mini.pt # PyTorch weights (fp32, 189.7 MB)
βββ pinta-1.2-mini.fp16.pt # PyTorch weights (fp16, 94.8 MB)
βββ pinta-1.2-mini.opset18.onnx # ONNX graph (fp32, 5.7 MB)
βββ pinta-1.2-mini.opset18.onnx.data # ONNX weights (fp32, 199.9 MB)
βββ pinta-1.2-mini.fp16.opset18.onnx # ONNX graph (fp16, 6.0 MB)
βββ pinta-1.2-mini.fp16.opset18.onnx.data # ONNX weights (fp16, 105.3 MB)
βββ pinta_router.cpp # C++ edge router (no calibration)
βββ tokenizer.json # QT-Cittern-1.0 tokenizer (9,000 vocab)
βββ vocab.json # Tokenizer vocab mapping
βββ merges.txt # Tokenizer BPE merges
βββ README.md # Model Card
βββ LICENSE # Apache 2.0
Intended Use & Ecosystem
Pinta-1.2 Mini is designed for:
- Agent swarm dispatchers - route user queries to specialized agents (code, math, RAG) at the edge
- CLI sidecars - classify shell commands or log lines in real-time with sub-15 ms latency
- Edge devices - run on Raspberry Pi, Jetson, or commodity x86 with CPU-only inference
- Cost-aware routing - send simple prompts to cheap models, complex prompts to expensive ones
- Compliance gating - flag
<|reserved_30|>(HIGH_RISK_COMPLIANCE) prompts for human review before processing
Not intended for:
- Generative text completion (this is a classifier, not a language model)
- High-stakes classification without human oversight (use the confidence gate)
- Non-English prompts (trained exclusively on English data)
Known Limitations
Boundary Overlaps
Three domain boundaries have inherent semantic overlap:
- SyntaxDevOps vs. Code: Formatting tasks involving code syntax (regex, SQL pretty-printing) may route to Code
- Creative vs. Knowledge: Ambiguous "describe/suggest" prompts without explicit creative framing may route to Knowledge
- SyntaxDevOps vs. Knowledge: Language transformation tasks (summarize, edit) sit at the boundary
Use confidence scores to handle borderline cases.
Context Length
Prompts exceeding 2,048 tokens are truncated to the RoPE positional limit. For very long documents, consider pre-summarizing before routing.
Training Architecture & Datasets
Pinta-1.2 Mini was trained using a two-stage pipeline on consumer-grade hardware.
1. Base Pre-Training (Cittern-50M Checkpoint)
The base Cittern-50M model was pre-trained on a STEM-heavy corpus:
- Mathematics & Reasoning (Dominant): 45+ shards of clean pure/applied math, plus MathOverflow Q&A and Physics reasoning
- Computer Science & Code: 130+ "sterile" cleaned StackOverflow shards, CodeSearchNet Python, and cleaned source code in C, C++, Python, and Rust
- General Knowledge: English Wikipedia, general English text, and academic papers
2. Fine-Tuning (Stable SFT)
Unlike the Beta, Pinta-1.2 Mini was fine-tuned on a strictly balanced, quality-audited dataset to eliminate class priors and remove the need for inference-time calibration.
- Dataset: 58,220 records across 9 balanced domains
- Max/Min ratio: 1.46x (near-uniform distribution)
- Sources: GSM8K, SQuAD v2, LegalBench, TriviaQA, WildChat-1M, UltraFeedback, WritingPrompts + targeted synthetic generation
- Quality gates: Non-English filtered, exact deduplication, architecture NLP-task contamination removed
Training Configuration
| Parameter | Value |
|---|---|
| Base checkpoint | Cittern-50M with block Householder orthogonal attention |
| Epochs | 2 |
| Effective batch size | 64 (micro-batch 8, grad accum 8) |
| Learning rate | 1e-5 |
| Orthogonal regularization weight | 0.01 |
| Template | User: {prompt}\nAssistant: |
| Hardware | NVIDIA RTX 4060 (bfloat16 mixed precision) |
Model Card Metadata
| Field | Value |
|---|---|
| Model type | Semantic router / classifier (20-layer Cittern transformer with causal routing head) |
| Base checkpoint | Cittern-50M |
| Inference size | 199.9 MB (ONNX fp32 weights) / 105.3 MB (ONNX fp16 weights) |
| Training hardware | NVIDIA RTX 4060 |
| Training procedure | SFT from base Cittern-50M, 2 epochs, effective batch 64, lr=1e-5, ortho_weight=0.01 |
| Evaluation | Balanced Routing Benchmark v6 (1,060 prompts, zero train overlap) |
| Calibration required | No |
| License | Apache 2.0 |
| Release date | 2026-09-23 |
Links & Contact
- Hugging Face Profile: JamesQuartz
- Predecessor: AENEA Pinta-1.1 Mini Beta
- Full Model: AENEA Pinta-1.1 (226M)
- Company Website: https://aeneaglobal.com/
- Open-Source Models & Tokenizers: quartz.host
- Commercial & Partnership Inquiries: commercial@aeneaglobal.com
Citation
If you use Pinta-1.2 Mini in your research or production systems, please cite:
@misc{pinta-mini-v1.2-2026,
title={AENEA Pinta-1.2 Mini Stable: A 50M-Parameter Semantic Router with 91.5% Accuracy via Cittern Architecture},
author={AENEA Research},
year={2026},
howpublished={\url{https://huggingface.co/JamesQuartz/aenea-pinta-1.2-mini}},
}
License
Apache 2.0. See LICENSE file for details.
Evaluation results
- Strict Accuracy on Balanced Routing Benchmark v6self-reported91.510
- Macro F1 on Balanced Routing Benchmark v6self-reported0.910