AENEA Pinta-2.0
74M-parameter semantic router for high-fidelity edge dispatch. Pinta-2.0 is the next-generation successor to the Pinta-1.2 Mini, delivering 92.45% strict accuracy on our balanced v6 benchmark while leveraging the expanded 12K vocabulary and deeper 24-layer Cittern-2.0 architecture.
Designed for always-on classification in agent swarms, CLI sidecars, and enterprise edge dispatch where both latency and nuanced domain separation are critical.
Release notice. This is the standard Pinta-2.0 release. The routing head was trained on a rigorously audited 57,968-record dataset. Extensive data-cleaning pipelines were used to eliminate template collapse in the Creative and RAG domains, resulting in near-perfect Compliance routing (0.99 Precision) and flawless Math recall (1.00). The model can be used directly with raw softmax over the 9 domain logits. (Note: The "Pro" moniker is reserved for the upcoming v2.1 iterative refinement release).
What is this?
Pinta-2.0 is not a generative language model. It is a semantic router that classifies incoming prompts into one of nine dispatch domains and emits a routing token that downstream systems can use to select the appropriate expert model or tool:
| Class | Token | Dispatch to |
|---|---|---|
| SyntaxDevOps | <|reserved_23|> |
Template engines, formatters, regex, data conversion |
| Code | <|reserved_24|> |
Code-specialized models, debuggers, linters |
| Creative | <|reserved_25|> |
Creative writing models, storytelling engines |
| RAG | <|reserved_26|> |
Retrieval-augmented pipelines, document QA |
| Architecture | <|reserved_27|> |
System design experts, DevOps, cloud infra |
| Math | <|reserved_28|> |
Math/reasoning models, calculators, CAS |
| Knowledge | <|reserved_29|> |
General knowledge / fallback models |
| Compliance | <|reserved_30|> |
Legal domain models, policy engines, compliance review |
| Chat | <|reserved_31|> |
Chat-optimized models, conversational agents |
The router is trained to predict which token a full generative model would emit next, but does so in a single forward pass without generating any text.
Architecture
| Component | Specification |
|---|---|
| Parameters | 74M |
| Layers | 24 (Cartan/Cittern-2.0) |
| Hidden size | 768 |
| Attention heads | 12 |
| Orthogonal attention | Block Householder, 12 reflections, block_size 64 |
| Vocabulary | 12,000 (QT-Cittern-12k tokenizer) |
| Context window | 2,048 tokens |
| Precision | bf16 / fp16 / ONNX fp32 |
| Routing head | Causal [B, S, V] with 9 reserved token slice |
Tokenizer Selection: The 12K Edge Sweet Spot
Pinta-2.0 uses the QT-Cittern-12k tokenizer. While larger tokenizers (like Llama 3's 128K or DeepSeek V4's 129K) offer superior raw text compression, they are fundamentally incompatible with ultra-low-latency edge routing.
A 128K vocabulary requires a 393 MB embedding matrix and softmax head, which would dwarf our 74M parameter transformer blocks and destroy the sub-10ms inference latency. The 12K vocabulary strikes the optimal architectural balance: it provides strong compression across code, bash, and human languages while keeping the routing head compact (36 MB).
Compression density across Bash, programming languages, human languages, and scientific formulas. QT-Cittern-12k provides excellent compression relative to its size class.
Code-specific tokenization efficiency, critical for accurate Code and SyntaxDevOps domain routing.
Multilingual script coverage. While Pinta-2.0 is trained exclusively on English routing data, the underlying tokenizer retains strong structural representation for future multilingual expansion.
Benchmarks
Evaluated on the Balanced Routing Benchmark v6 (1,060 prompts across 9 domains, zero train overlap, English-only):
| Metric | Pinta-1.1 (226M) | Pinta-1.2 Mini (50M) | Pinta-2.0 (74M) |
|---|---|---|---|
| Strict accuracy | 79.22% | 91.51% | 92.45% |
| Macro F1 | - | 0.91 | 0.92 |
| Weighted F1 | - | 0.92 | 0.92 |
| Latency (p50, RTX 4090) | ~12 ms | ~6 ms | ~8 ms |
Pinta-2.0 successfully scales the Cittern architecture to 74M parameters, improving strict accuracy by nearly a full percentage point over the 50M Mini while maintaining sub-10ms GPU inference latency.
Per-Domain Performance (v6 Benchmark)
| Domain | Precision | Recall | F1 |
|---|---|---|---|
| SyntaxDevOps | 0.94 | 0.76 | 0.84 |
| Code | 0.83 | 0.83 | 0.83 |
| Creative | 0.94 | 0.88 | 0.91 |
| RAG | 0.96 | 0.99 | 0.98 |
| Architecture | 0.93 | 0.94 | 0.93 |
| Math | 0.91 | 1.00 | 0.95 |
| Knowledge | 0.84 | 0.96 | 0.89 |
| Compliance | 0.99 | 0.97 | 0.98 |
| Chat | 0.99 | 0.97 | 0.98 |
Note: Compliance and Math routing are near-perfect, making this model exceptionally safe for enterprise legal/medical gating and rigorous STEM dispatch.
Quickstart
Python Inference (PyTorch) β Reference Implementation
import torch
from transformers import AutoTokenizer
from cartan_olm import CartanConfig, CartanLM
tokenizer = AutoTokenizer.from_pretrained("tokenizers/QT-Cittern-12k")
checkpoint = torch.load("pinta-2.0.pt", map_location="cuda")
config = CartanConfig.from_dict(checkpoint["config"])
model = CartanLM(config).to("cuda").eval()
model.load_state_dict(checkpoint["model_state_dict"])
DOMAINS = [f"<|reserved_{i}|>" for i in range(23, 32)]
domain_ids = [tokenizer.convert_tokens_to_ids(d) for d in DOMAINS]
def route(prompt: str) -> dict:
formatted = f"User: {prompt}\nAssistant: "
ids = tokenizer.encode(formatted, return_tensors="pt", truncation=True, max_length=2048).to("cuda")
with torch.no_grad(), torch.autocast(device_type="cuda", dtype=torch.bfloat16):
logits = model(ids)["logits"][0, -1, :]
scores = logits[domain_ids]
probs = torch.softmax(scores, dim=-1)
best = torch.argmax(probs).item()
return {"token": DOMAINS[best], "confidence": probs[best].item()}
print(route("Review this contract for GDPR Article 17 compliance"))
# Output: {'token': '<|reserved_30|>', 'confidence': 0.993}
ONNX Runtime Inference (Python)
import numpy as np
import onnxruntime as ort
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained(".") # tokenizer.json in repo root
session = ort.InferenceSession("pinta-2.0.opset18.onnx")
DOMAINS = [f"<|reserved_{i}|>" for i in range(23, 32)]
domain_ids = [tokenizer.convert_tokens_to_ids(d) for d in DOMAINS]
def route(prompt: str) -> dict:
formatted = f"User: {prompt}\nAssistant: "
ids = tokenizer.encode(formatted, return_tensors="np",
add_special_tokens=False,
truncation=True, max_length=2048).astype(np.int64)
logits = session.run(None, {"input_ids": ids})[0][0, -1, :]
scores = logits[domain_ids]
probs = np.exp(scores - scores.max())
probs /= probs.sum()
best = int(np.argmax(probs))
return {"token": DOMAINS[best], "confidence": float(probs[best])}
C++ Edge Router (ONNX Runtime)
A zero-dependency C++ router is included for edge deployment where Python is unavailable:
# Build (MSVC x64, requires ONNX Runtime C++ headers)
cl /std:c++17 /O2 /EHsc /I"path\to\ort\include" pinta_router.cpp \
/link /LIBPATH:"path\to\ort\lib" onnxruntime.lib /out:pinta_router.exe
# Run (auto-loads model and tokenizer from exe directory)
pinta_router.exe "Design a microservices architecture"
β οΈ Tokenizer Parity Note: The C++ router includes a self-contained minimal BPE tokenizer implementation for zero-dependency edge deployment. Due to subtle differences between this simplified BPE and the HuggingFace
PreTrainedTokenizerFastused during training, confidence scores may differ slightly between the C++ and Python routers, and on rare edge-case prompts (particularly long RAG contexts with numerical content), the C++ router may produce a different routing decision. For production deployments requiring exact output parity with training behavior, use the Python ONNX router (pinta_router_py.py) or integrate the HuggingFacetokenizerslibrary into your C++ pipeline viatokenizers-cpp. The C++ router agrees with the Python reference on >95% of prompts and is suitable for latency-critical edge scenarios where approximate routing is acceptable.
Confidence Gate
We recommend a confidence threshold of 0.35. If the model's max probability falls below this, fall back to <|reserved_29|> (Knowledge) rather than emitting a low-confidence misroute:
if result["confidence"] < 0.35:
result["token"] = "<|reserved_29|>" # Knowledge fallback
Repository Layout
aenea-pinta-2.0/
βββ pinta-2.0.pt # PyTorch weights (fp32, 857.5 MB)
βββ pinta-2.0.fp16.pt # PyTorch weights (fp16, 146.3 MB)
βββ pinta-2.0.opset18.onnx # ONNX graph (fp32)
βββ pinta-2.0.opset18.onnx.data # ONNX external weights (fp32)
βββ pinta_router.cpp # C++ edge router source (zero-dep BPE)
βββ pinta_router_py.py # Python ONNX reference router
βββ export_fp16.py # Script to regenerate fp16 weights
βββ export_onnx.py # Script to regenerate ONNX artifacts
βββ test.py # Inference smoke test
βββ tokenizer.json # QT-Cittern-12k tokenizer
βββ tokenizer_config.json # Tokenizer configuration
βββ special_tokens_map.json # Special tokens mapping
βββ vocab.json # Tokenizer vocab mapping
βββ merges.txt # Tokenizer BPE merges
βββ 1_TokenizerBench_Suite_Comparison.png # Tokenizer benchmark chart
βββ 2_Python_Parquet_Benchmark_Comparison.png
βββ 3_FLORES_Regional_Scripts_Comparison.png
βββ README.md # This model card
βββ LICENSE # Apache 2.0
Intended Use & Ecosystem
Pinta-2.0 is designed for:
- Agent swarm dispatchers - route user queries to specialized agents (code, math, RAG) at the edge
- Enterprise Compliance Gating - instantly flag
<|reserved_30|>(Compliance) prompts for human review or secure legal-model routing - STEM & RAG Pipelines - leverage the 1.00 Math recall and 0.99 RAG recall to ensure calculators and vector DBs are triggered reliably
- Cost-aware routing - send simple prompts to cheap models, complex prompts to expensive ones
Not intended for:
- Generative text completion (this is a classifier, not a language model)
- High-stakes classification without human oversight (use the confidence gate)
- Non-English prompts (trained exclusively on English data)
Known Limitations
Boundary Overlaps
- Knowledge vs. Math: Factual questions containing heavy numerical data (e.g., "How many chambers has the heart?") may experience minor Math bleed.
- Code vs. Architecture: Prompts requesting code implementation using high-level design keywords (e.g., "Write the Flask endpoint for JWT") may route to Architecture.
- SyntaxDevOps vs. Code: Text-transformation tasks (e.g., sentence tense conversion) may occasionally route to Code.
Context Length
Prompts exceeding 2,048 tokens are truncated to the RoPE positional limit. For very long documents, consider pre-summarizing before routing.
Training Architecture & Datasets
Pinta-2.0 was trained using a rigorous data-cleaning and SFT pipeline on high-end consumer hardware.
1. Dataset Curation (v12_mini_final_cleaned)
Unlike previous iterations, the Pinta-2.0 dataset underwent surgical auditing to eliminate template collapse:
- Creative & RAG Stripping: Removed >5,000 templated/ambiguous examples (e.g., "Write a story about...", bare questions without context blocks).
- DeepSeek V4 Backfill: Injected thousands of structurally diverse, high-signal prompts to ensure the router learns intent, not keywords.
- Final Size: 57,968 records across 9 domains.
- Distribution: ~10.8% per domain, with Compliance intentionally oversampled to 14.9% to prioritize safety routing.
2. Fine-Tuning (SFT on Cittern-2.0)
The model was fine-tuned from the 74M Cittern-2.0 base checkpoint using a microscopic learning rate to preserve the Stiefel manifold orthogonal geometry.
| Parameter | Value |
|---|---|
| Base checkpoint | Cittern-2.0 (74M) with block Householder orthogonal attention |
| Epochs | 3 (Selected for optimal accuracy/manifold preservation tradeoff) |
| Effective batch size | 64 (micro-batch 8, grad accum 8) |
| Learning rate | 1e-5 |
| Orthogonal regularization weight | 0.01 |
| Template | User: {prompt}\nAssistant: |
| Hardware | NVIDIA RTX 4090 24GB (bfloat16 mixed precision) |
Model Card Metadata
| Field | Value |
|---|---|
| Model type | Semantic router / classifier (24-layer Cittern-2.0 transformer with causal routing head) |
| Base checkpoint | Cittern-2.0 (74M) |
| Training hardware | NVIDIA RTX 4090 |
| Training procedure | SFT from base Cittern-2.0, 3 epochs, effective batch 64, lr=1e-5, ortho_weight=0.01 |
| Evaluation | Balanced Routing Benchmark v6 (1,060 prompts, zero train overlap) |
| Calibration required | No |
| License | Apache 2.0 |
| Release date | 2026-10-05 |
Links & Contact
- Hugging Face Profile: JamesQuartz
- Predecessor: AENEA Pinta-1.2 Mini Stable
- Company Website: https://aeneaglobal.com/
- Open-Source Models & Tokenizers: quartz.host
- Commercial & Partnership Inquiries: commercial@aeneaglobal.com
Citation
If you use Pinta-2.0 in your research or production systems, please cite:
@misc{pinta-v2.0-2026,
title={AENEA Pinta-2.0: A 74M-Parameter Semantic Router with 92.45% Accuracy via Cittern-2.0 Architecture},
author={AENEA Global Research},
year={2026},
howpublished={\url{https://huggingface.co/JamesQuartz/aenea-pinta-2.0}},
}
License
Apache 2.0. See LICENSE file for details.
Evaluation results
- Strict Accuracy on Balanced Routing Benchmark v6self-reported92.450
- Macro F1 on Balanced Routing Benchmark v6self-reported0.920