AENEA Pinta-2.0

74M-parameter semantic router for high-fidelity edge dispatch. Pinta-2.0 is the next-generation successor to the Pinta-1.2 Mini, delivering 92.45% strict accuracy on our balanced v6 benchmark while leveraging the expanded 12K vocabulary and deeper 24-layer Cittern-2.0 architecture.

Designed for always-on classification in agent swarms, CLI sidecars, and enterprise edge dispatch where both latency and nuanced domain separation are critical.

Release notice. This is the standard Pinta-2.0 release. The routing head was trained on a rigorously audited 57,968-record dataset. Extensive data-cleaning pipelines were used to eliminate template collapse in the Creative and RAG domains, resulting in near-perfect Compliance routing (0.99 Precision) and flawless Math recall (1.00). The model can be used directly with raw softmax over the 9 domain logits. (Note: The "Pro" moniker is reserved for the upcoming v2.1 iterative refinement release).


What is this?

Pinta-2.0 is not a generative language model. It is a semantic router that classifies incoming prompts into one of nine dispatch domains and emits a routing token that downstream systems can use to select the appropriate expert model or tool:

Class Token Dispatch to
SyntaxDevOps <|reserved_23|> Template engines, formatters, regex, data conversion
Code <|reserved_24|> Code-specialized models, debuggers, linters
Creative <|reserved_25|> Creative writing models, storytelling engines
RAG <|reserved_26|> Retrieval-augmented pipelines, document QA
Architecture <|reserved_27|> System design experts, DevOps, cloud infra
Math <|reserved_28|> Math/reasoning models, calculators, CAS
Knowledge <|reserved_29|> General knowledge / fallback models
Compliance <|reserved_30|> Legal domain models, policy engines, compliance review
Chat <|reserved_31|> Chat-optimized models, conversational agents

The router is trained to predict which token a full generative model would emit next, but does so in a single forward pass without generating any text.


Architecture

Component Specification
Parameters 74M
Layers 24 (Cartan/Cittern-2.0)
Hidden size 768
Attention heads 12
Orthogonal attention Block Householder, 12 reflections, block_size 64
Vocabulary 12,000 (QT-Cittern-12k tokenizer)
Context window 2,048 tokens
Precision bf16 / fp16 / ONNX fp32
Routing head Causal [B, S, V] with 9 reserved token slice

Tokenizer Selection: The 12K Edge Sweet Spot

Pinta-2.0 uses the QT-Cittern-12k tokenizer. While larger tokenizers (like Llama 3's 128K or DeepSeek V4's 129K) offer superior raw text compression, they are fundamentally incompatible with ultra-low-latency edge routing.

A 128K vocabulary requires a 393 MB embedding matrix and softmax head, which would dwarf our 74M parameter transformer blocks and destroy the sub-10ms inference latency. The 12K vocabulary strikes the optimal architectural balance: it provides strong compression across code, bash, and human languages while keeping the routing head compact (36 MB).

TokenizerBench Suite Compression density across Bash, programming languages, human languages, and scientific formulas. QT-Cittern-12k provides excellent compression relative to its size class.

Python Code Benchmark Code-specific tokenization efficiency, critical for accurate Code and SyntaxDevOps domain routing.

FLORES-200 Regional Scripts Multilingual script coverage. While Pinta-2.0 is trained exclusively on English routing data, the underlying tokenizer retains strong structural representation for future multilingual expansion.


Benchmarks

Evaluated on the Balanced Routing Benchmark v6 (1,060 prompts across 9 domains, zero train overlap, English-only):

Metric Pinta-1.1 (226M) Pinta-1.2 Mini (50M) Pinta-2.0 (74M)
Strict accuracy 79.22% 91.51% 92.45%
Macro F1 - 0.91 0.92
Weighted F1 - 0.92 0.92
Latency (p50, RTX 4090) ~12 ms ~6 ms ~8 ms

Pinta-2.0 successfully scales the Cittern architecture to 74M parameters, improving strict accuracy by nearly a full percentage point over the 50M Mini while maintaining sub-10ms GPU inference latency.

Per-Domain Performance (v6 Benchmark)

Domain Precision Recall F1
SyntaxDevOps 0.94 0.76 0.84
Code 0.83 0.83 0.83
Creative 0.94 0.88 0.91
RAG 0.96 0.99 0.98
Architecture 0.93 0.94 0.93
Math 0.91 1.00 0.95
Knowledge 0.84 0.96 0.89
Compliance 0.99 0.97 0.98
Chat 0.99 0.97 0.98

Note: Compliance and Math routing are near-perfect, making this model exceptionally safe for enterprise legal/medical gating and rigorous STEM dispatch.


Quickstart

Python Inference (PyTorch) β€” Reference Implementation

import torch
from transformers import AutoTokenizer
from cartan_olm import CartanConfig, CartanLM

tokenizer = AutoTokenizer.from_pretrained("tokenizers/QT-Cittern-12k")
checkpoint = torch.load("pinta-2.0.pt", map_location="cuda")
config = CartanConfig.from_dict(checkpoint["config"])

model = CartanLM(config).to("cuda").eval()
model.load_state_dict(checkpoint["model_state_dict"])

DOMAINS = [f"<|reserved_{i}|>" for i in range(23, 32)]
domain_ids = [tokenizer.convert_tokens_to_ids(d) for d in DOMAINS]

def route(prompt: str) -> dict:
    formatted = f"User: {prompt}\nAssistant: "
    ids = tokenizer.encode(formatted, return_tensors="pt", truncation=True, max_length=2048).to("cuda")
    
    with torch.no_grad(), torch.autocast(device_type="cuda", dtype=torch.bfloat16):
        logits = model(ids)["logits"][0, -1, :]
        
    scores = logits[domain_ids]
    probs = torch.softmax(scores, dim=-1)
    best = torch.argmax(probs).item()
    
    return {"token": DOMAINS[best], "confidence": probs[best].item()}

print(route("Review this contract for GDPR Article 17 compliance"))
# Output: {'token': '<|reserved_30|>', 'confidence': 0.993}

ONNX Runtime Inference (Python)

import numpy as np
import onnxruntime as ort
from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained(".")  # tokenizer.json in repo root
session = ort.InferenceSession("pinta-2.0.opset18.onnx")

DOMAINS = [f"<|reserved_{i}|>" for i in range(23, 32)]
domain_ids = [tokenizer.convert_tokens_to_ids(d) for d in DOMAINS]

def route(prompt: str) -> dict:
    formatted = f"User: {prompt}\nAssistant: "
    ids = tokenizer.encode(formatted, return_tensors="np",
                           add_special_tokens=False,
                           truncation=True, max_length=2048).astype(np.int64)
    
    logits = session.run(None, {"input_ids": ids})[0][0, -1, :]
    scores = logits[domain_ids]
    probs = np.exp(scores - scores.max())
    probs /= probs.sum()
    
    best = int(np.argmax(probs))
    return {"token": DOMAINS[best], "confidence": float(probs[best])}

C++ Edge Router (ONNX Runtime)

A zero-dependency C++ router is included for edge deployment where Python is unavailable:

# Build (MSVC x64, requires ONNX Runtime C++ headers)
cl /std:c++17 /O2 /EHsc /I"path\to\ort\include" pinta_router.cpp \
   /link /LIBPATH:"path\to\ort\lib" onnxruntime.lib /out:pinta_router.exe

# Run (auto-loads model and tokenizer from exe directory)
pinta_router.exe "Design a microservices architecture"

⚠️ Tokenizer Parity Note: The C++ router includes a self-contained minimal BPE tokenizer implementation for zero-dependency edge deployment. Due to subtle differences between this simplified BPE and the HuggingFace PreTrainedTokenizerFast used during training, confidence scores may differ slightly between the C++ and Python routers, and on rare edge-case prompts (particularly long RAG contexts with numerical content), the C++ router may produce a different routing decision. For production deployments requiring exact output parity with training behavior, use the Python ONNX router (pinta_router_py.py) or integrate the HuggingFace tokenizers library into your C++ pipeline via tokenizers-cpp. The C++ router agrees with the Python reference on >95% of prompts and is suitable for latency-critical edge scenarios where approximate routing is acceptable.

Confidence Gate

We recommend a confidence threshold of 0.35. If the model's max probability falls below this, fall back to <|reserved_29|> (Knowledge) rather than emitting a low-confidence misroute:

if result["confidence"] < 0.35:
    result["token"] = "<|reserved_29|>"  # Knowledge fallback

Repository Layout

aenea-pinta-2.0/
β”œβ”€β”€ pinta-2.0.pt                           # PyTorch weights (fp32, 857.5 MB)
β”œβ”€β”€ pinta-2.0.fp16.pt                      # PyTorch weights (fp16, 146.3 MB)
β”œβ”€β”€ pinta-2.0.opset18.onnx                 # ONNX graph (fp32)
β”œβ”€β”€ pinta-2.0.opset18.onnx.data            # ONNX external weights (fp32)
β”œβ”€β”€ pinta_router.cpp                       # C++ edge router source (zero-dep BPE)
β”œβ”€β”€ pinta_router_py.py                     # Python ONNX reference router
β”œβ”€β”€ export_fp16.py                         # Script to regenerate fp16 weights
β”œβ”€β”€ export_onnx.py                         # Script to regenerate ONNX artifacts
β”œβ”€β”€ test.py                                # Inference smoke test
β”œβ”€β”€ tokenizer.json                         # QT-Cittern-12k tokenizer
β”œβ”€β”€ tokenizer_config.json                  # Tokenizer configuration
β”œβ”€β”€ special_tokens_map.json                # Special tokens mapping
β”œβ”€β”€ vocab.json                             # Tokenizer vocab mapping
β”œβ”€β”€ merges.txt                             # Tokenizer BPE merges
β”œβ”€β”€ 1_TokenizerBench_Suite_Comparison.png  # Tokenizer benchmark chart
β”œβ”€β”€ 2_Python_Parquet_Benchmark_Comparison.png
β”œβ”€β”€ 3_FLORES_Regional_Scripts_Comparison.png
β”œβ”€β”€ README.md                              # This model card
└── LICENSE                                # Apache 2.0

Intended Use & Ecosystem

Pinta-2.0 is designed for:

  • Agent swarm dispatchers - route user queries to specialized agents (code, math, RAG) at the edge
  • Enterprise Compliance Gating - instantly flag <|reserved_30|> (Compliance) prompts for human review or secure legal-model routing
  • STEM & RAG Pipelines - leverage the 1.00 Math recall and 0.99 RAG recall to ensure calculators and vector DBs are triggered reliably
  • Cost-aware routing - send simple prompts to cheap models, complex prompts to expensive ones

Not intended for:

  • Generative text completion (this is a classifier, not a language model)
  • High-stakes classification without human oversight (use the confidence gate)
  • Non-English prompts (trained exclusively on English data)

Known Limitations

Boundary Overlaps

  • Knowledge vs. Math: Factual questions containing heavy numerical data (e.g., "How many chambers has the heart?") may experience minor Math bleed.
  • Code vs. Architecture: Prompts requesting code implementation using high-level design keywords (e.g., "Write the Flask endpoint for JWT") may route to Architecture.
  • SyntaxDevOps vs. Code: Text-transformation tasks (e.g., sentence tense conversion) may occasionally route to Code.

Context Length

Prompts exceeding 2,048 tokens are truncated to the RoPE positional limit. For very long documents, consider pre-summarizing before routing.


Training Architecture & Datasets

Pinta-2.0 was trained using a rigorous data-cleaning and SFT pipeline on high-end consumer hardware.

1. Dataset Curation (v12_mini_final_cleaned)

Unlike previous iterations, the Pinta-2.0 dataset underwent surgical auditing to eliminate template collapse:

  • Creative & RAG Stripping: Removed >5,000 templated/ambiguous examples (e.g., "Write a story about...", bare questions without context blocks).
  • DeepSeek V4 Backfill: Injected thousands of structurally diverse, high-signal prompts to ensure the router learns intent, not keywords.
  • Final Size: 57,968 records across 9 domains.
  • Distribution: ~10.8% per domain, with Compliance intentionally oversampled to 14.9% to prioritize safety routing.

2. Fine-Tuning (SFT on Cittern-2.0)

The model was fine-tuned from the 74M Cittern-2.0 base checkpoint using a microscopic learning rate to preserve the Stiefel manifold orthogonal geometry.

Parameter Value
Base checkpoint Cittern-2.0 (74M) with block Householder orthogonal attention
Epochs 3 (Selected for optimal accuracy/manifold preservation tradeoff)
Effective batch size 64 (micro-batch 8, grad accum 8)
Learning rate 1e-5
Orthogonal regularization weight 0.01
Template User: {prompt}\nAssistant:
Hardware NVIDIA RTX 4090 24GB (bfloat16 mixed precision)

Model Card Metadata

Field Value
Model type Semantic router / classifier (24-layer Cittern-2.0 transformer with causal routing head)
Base checkpoint Cittern-2.0 (74M)
Training hardware NVIDIA RTX 4090
Training procedure SFT from base Cittern-2.0, 3 epochs, effective batch 64, lr=1e-5, ortho_weight=0.01
Evaluation Balanced Routing Benchmark v6 (1,060 prompts, zero train overlap)
Calibration required No
License Apache 2.0
Release date 2026-10-05

Links & Contact


Citation

If you use Pinta-2.0 in your research or production systems, please cite:

@misc{pinta-v2.0-2026,
  title={AENEA Pinta-2.0: A 74M-Parameter Semantic Router with 92.45% Accuracy via Cittern-2.0 Architecture},
  author={AENEA Global Research},
  year={2026},
  howpublished={\url{https://huggingface.co/JamesQuartz/aenea-pinta-2.0}},
}

License

Apache 2.0. See LICENSE file for details.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Evaluation results

  • Strict Accuracy on Balanced Routing Benchmark v6
    self-reported
    92.450
  • Macro F1 on Balanced Routing Benchmark v6
    self-reported
    0.920