- Pinta-1.1 (226M) β High-Speed Semantic Router [BETA]
Pinta-1.1 (226M) β High-Speed Semantic Router [BETA]
β οΈ Beta Release: Pinta-1.1 is currently in Beta. While strict accuracy benchmarks are strong, we are actively seeking community testing and feedback on routing edge-cases and hardware deployment environments. Please open a discussion thread to share your logs or suggestions!
Pinta-1.1 is a 226-million parameter lightweight language model built on the custom Merlin Architecture and trained on top of the AENEA Prelude-6 base checkpoint.
Designed as an ultra-fast, low-latency semantic routing gateway (~656 MB in bf16), Pinta-1.1 classifies incoming prompts into high-precision domain tokens (<|reserved_23|> through <|reserved_31|>) to dynamically route traffic across Mixture-of-Experts (MoE) backends, local SLMs, or specialized API endpoints.
π§ Intended use: Pinta-1.1 is a classifier/gateway, not a conversational model. It is designed to emit a single routing token per prompt in one forward pass. It is not intended for open-ended text generation.
π§° Prerequisites
pip install torch transformers fastapi uvicorn httpx pyyaml
- Python: β₯ 3.9
- Hardware: CUDA-capable GPU recommended; CPU and MPS inference are supported (with higher TTFT).
- Weights: shipped in
bfloat16(pinta_1.1_bf16.pt, ~656 MB).
β οΈ Custom Architecture (merlin_2) β Required Repository Files
Pinta-1.1 does not use a standard Hugging Face architecture class. It implements the custom Merlin v2 (merlin_2) architecture (16 transformer blocks with parallel residuals, 16Q/4KV Grouped-Query Attention, SwiGLU FFN, RoPE, 49K vocab). Because merlin_2 bypasses the standard HF AutoModel mapping:
AutoModelForCausalLM.from_pretrained(...)will not load this checkpoint out of the box.- You must download the full repository, including the
merlin_2/package and thepinta_engine.pyrouter script, for the model to work.
huggingface-cli download JamesQuartz/aenea-pinta-1.1 \
--include "*.py" "*.yaml" "*.json" "*.pt" \
--local-dir ./aenea-pinta-1.1
Expected repository layout:
aenea-pinta-1.1/
βββ merlin_2/ # Custom architecture: config + modeling code
βββ pinta_engine.py # Router engine & FastAPI middleware
βββ config.yaml # Model + router configuration
βββ routing_map.json # Reserved-token β backend dispatch map
βββ tokenizer.json # QT VI.6.4 UltraLingo tokenizer (49,152 vocab)
βββ pinta_1.1_bf16.pt # bf16 weights (~656 MB)
pinta_engine.py imports merlin_2 from the repository root at runtime β keep both in the same working directory (or on PYTHONPATH).
π Router Engine Quickstart
Pinta-1.1 ships with a custom, production-ready routing engine (pinta_engine.py) that handles model loading (via the local merlin_2 module), fast logits extraction, and async dispatch to OpenAI-compatible endpoints or local Ollama instances.
1. Execute a Routing Pass (example_usage.py)
import asyncio
from pinta_engine import PintaRouter
# Initialize the router engine (loads bf16 weights via the custom merlin_2 architecture)
router = PintaRouter("config.yaml")
prompt = "Write a Python function to normalize a file path."
# 1. Ultra-fast routing decision
decision = router.route(prompt)
print(decision.token)
print(decision.label)
print(decision.confidence)
print(decision.downstream_model)
# 2. Async dispatch to the mapped model backend (e.g., local Ollama qwen2.5-coder)
# response = asyncio.run(router.adispatch(prompt, decision))
# print(response["choices"][0]["message"]["content"])
Expected Output
2026-09-18 11:57:41,512 INFO pinta.router [Pinta Router] Initializing custom Merlin (merlin_2) architecture...
2026-09-18 11:57:46,784 INFO pinta.router [Pinta Router] Initialized Pinta-1.1 router on cuda with dtype=torch.bfloat16, threshold=0.350, fallback=<|reserved_29|>
2026-09-18 11:57:47,117 INFO pinta.router [Pinta Router] Prompt -> Mapped to <|reserved_24|> (Code) -> Dispatching to qwen2.5-coder:32b (Latency: 333ms).
<|reserved_24|>
Code
0.9890856146812439
qwen2.5-coder:32b
The 333 ms figure above is the first (cold) routing pass including CUDA kernel warm-up. Warm-cache passes are substantially faster; the benchmarked average TTFT of 192.66 ms is measured across the full 1,020-prompt evaluation on reference hardware.
2. Configure Dispatch Targets (routing_map.json)
The router engine decides which domain token a prompt belongs to; routing_map.json tells it where to send that prompt. Each reserved token is bound to a downstream backend (Ollama, vLLM / any OpenAI-compatible server, the OpenAI API, or a local Hugging Face pipeline):
{
"<|reserved_24|>": {
"label": "Code",
"backend": "ollama",
"base_url": "http://localhost:11434",
"model": "qwen2.5-coder:32b"
},
"<|reserved_26|>": {
"label": "RAG / Doc QA",
"backend": "vllm",
"base_url": "http://localhost:8000/v1",
"model": "qwen2.5-7b-instruct"
},
"<|reserved_29|>": {
"label": "Knowledge",
"backend": "openai",
"model": "gpt-4o-mini",
"api_key_env": "OPENAI_API_KEY"
}
}
- If the router's confidence for the predicted token falls below the configured threshold (default
0.35), the prompt is automatically routed to the Knowledge fallback (<|reserved_29|>). - Backends, base URLs, API-key environment variables, and per-target default generation parameters are all configurable per token.
π Benchmark Performance Results
Evaluated on 1,020 curated cross-domain evaluation prompts sourced from benchmark datasets.
| Overall Metric | Result |
|---|---|
| Strict Accuracy (Primary Token Match) | 84.61% |
| Flexible Accuracy (Allowed Token Match) | 87.65% |
| Average Latency (TTFT) | 192.66 ms |
Per-Domain Performance Summary (n = 1,020)
| Domain Token | Domain Name | Precision | Recall | F1-Score | Support (n) |
|---|---|---|---|---|---|
<|reserved_23|> |
Formatting | 0.50 | 0.05 | 0.09 | 43 |
<|reserved_24|> |
Code | 0.83 | 0.94 | 0.88 | 281 |
<|reserved_25|> |
Creative | 0.70 | 0.23 | 0.34 | 31 |
<|reserved_26|> |
RAG / Doc QA | 1.00 | 0.83 | 0.91 | 24 |
<|reserved_27|> |
Architecture | 0.31 | 0.56 | 0.40 | 9 |
<|reserved_28|> |
Math | 0.92 | 0.92 | 0.92 | 260 |
<|reserved_29|> |
Knowledge | 0.83 | 0.88 | 0.86 | 369 |
<|reserved_30|> |
Law / Policy | 0.00 | 0.00 | 0.00 | 3 |
| Weighted Average | β | 0.83 | 0.85 | 0.83 | 1020 |
Support values sum to the full 1,020-prompt evaluation set. See Known Limitations for details on the Formatting and Law / Policy classes.
π Model Specifications
- Base Checkpoint: AENEA Prelude-6
- Inference Size: ~656 MB (
bfloat16) - Training Setup: Pre-trained on NVIDIA RTX 4090 / Fine-tuned on NVIDIA RTX 4060
| Parameter | Value | Description |
|---|---|---|
| Parameters | 226M | Low-footprint architecture for sub-200ms TTFT routing |
| Vocab Size | 49,152 | QT VI.6.4 UltraLingo Tokenizer |
| Layers | 16 | Transformer blocks with Parallel Residuals |
Hidden Size (d_model) |
1024 | Dense embedding dimension |
| Attention Heads | 16 Query / 4 KV | Grouped-Query Attention (GQA) |
Intermediate Size (d_ff) |
4096 | SwiGLU / Feed-Forward projection |
| Max Sequence Length | 2048 | Rotary Position Embeddings (RoPE) |
| Precision | bfloat16 |
Optimized for native GPU execution |
π Training Architecture & Datasets
Pinta-1.1 was trained in a two-stage pre-training and alignment pipeline:
1. Base Pre-Training (Prelude-6 Checkpoint)
- English Wikipedia: Complete clean English corpus.
- Stack Exchange Network: Technical QA corpora including StackOverflow, MathOverflow, ServerFault, SuperUser, AskUbuntu, CodeReview, SoftwareEngineering, ReverseEngineering, and NetworkEngineering.
- Source Code Corpus: Multi-language source code covering Python, C, C++, C#, Java, JavaScript, TypeScript, Rust, Go, PHP, Ruby, Kotlin, Swift, and Shell, plus CodeSearchNet Python.
- Mathematics & Science: Cleaned pure mathematics and scientific reasoning corpora.
2. Fine-Tuning & Contrastive Alignment
- Curated Open Data Mix (50%): Mined imperative subsets targeting instruction execution across Code, System Architecture, Formatting, and Reasoning.
- Deepseek-v4-Flash Synthetic Boundary Data (50%): Hard-negative contrastive pairs synthetically generated via Deepseek-v4-Flash to sharpen semantic boundaries between domain classes.
π‘οΈ Known Limitations & Beta Scope
- Formatting (
<|reserved_23|>) has low recall (0.05). This is a known consequence of class imbalance in the fine-tuning mix: formatting-style requests (markdown tables, JSON scaffolding, template rewrites) are under-represented relative to Code and Knowledge. In practice, most formatting prompts are absorbed by Knowledge (<|reserved_29|>) or Code (<|reserved_24|>) as a safe fallback, which preserves downstream answer quality at the cost of routing granularity. Its 0.50 precision also means half of the prompts it does claim are misrouted. Improving this class is a priority for the post-Beta release. - Law / Policy (
<|reserved_30|>) is unvalidated. Evaluation support is n = 3, so its reported scores are not statistically meaningful. Treat Law routing as experimental in this Beta. <|reserved_31|>is reserved for future domain expansion and is not scored in the Beta benchmark.- Confidence fallback: routing decisions below the confidence threshold (default
0.35) are redirected to Knowledge (<|reserved_29|>). Deployments with strict domain-separation requirements should tune this threshold and monitor fallback rates. - Hardware-dependent latency: TTFT figures are measured on the reference consumer-GPU setup above; CPU-only or shared-GPU deployments will see higher latency.
- Router-only model: Pinta-1.1 emits routing decisions and should not be used as a generative assistant.
π Links & Contact
- Hugging Face Profile: JamesQuartz
- Tokenizer Repository: JamesQuartz/qt-VI.6.4-49k
- Company Website: AENEA Global LTD
- Open-Source Models & Tokenizers: quartz.host
- Commercial & Partnership Inquiries: commercial@aeneaglobal.com
π License
This project is licensed under the Apache 2.0 License. Β© 2026 AENEA Global LTD
- Downloads last month
- 8
Evaluation results
- Strict Accuracy on Pinta Gold Benchmarktest set self-reported84.610
- Flexible Accuracy on Pinta Gold Benchmarktest set self-reported87.650
- Weighted Precision on Pinta Gold Benchmarktest set self-reported0.830
- Weighted Recall on Pinta Gold Benchmarktest set self-reported0.850
- Weighted F1-Score on Pinta Gold Benchmarktest set self-reported0.830
