Pinta-1.1 (226M) β€” High-Speed Semantic Router [BETA]

⚠️ Beta Release: Pinta-1.1 is currently in Beta. While strict accuracy benchmarks are strong, we are actively seeking community testing and feedback on routing edge-cases and hardware deployment environments. Please open a discussion thread to share your logs or suggestions!

Pinta-1.1 is a 226-million parameter lightweight language model built on the custom Merlin Architecture and trained on top of the AENEA Prelude-6 base checkpoint.

Designed as an ultra-fast, low-latency semantic routing gateway (~656 MB in bf16), Pinta-1.1 classifies incoming prompts into high-precision domain tokens (<|reserved_23|> through <|reserved_31|>) to dynamically route traffic across Mixture-of-Experts (MoE) backends, local SLMs, or specialized API endpoints.

🧭 Intended use: Pinta-1.1 is a classifier/gateway, not a conversational model. It is designed to emit a single routing token per prompt in one forward pass. It is not intended for open-ended text generation.


🧰 Prerequisites

pip install torch transformers fastapi uvicorn httpx pyyaml
  • Python: β‰₯ 3.9
  • Hardware: CUDA-capable GPU recommended; CPU and MPS inference are supported (with higher TTFT).
  • Weights: shipped in bfloat16 (pinta_1.1_bf16.pt, ~656 MB).

⚠️ Custom Architecture (merlin_2) β€” Required Repository Files

Pinta-1.1 does not use a standard Hugging Face architecture class. It implements the custom Merlin v2 (merlin_2) architecture (16 transformer blocks with parallel residuals, 16Q/4KV Grouped-Query Attention, SwiGLU FFN, RoPE, 49K vocab). Because merlin_2 bypasses the standard HF AutoModel mapping:

  • AutoModelForCausalLM.from_pretrained(...) will not load this checkpoint out of the box.
  • You must download the full repository, including the merlin_2/ package and the pinta_engine.py router script, for the model to work.
huggingface-cli download JamesQuartz/aenea-pinta-1.1 \
  --include "*.py" "*.yaml" "*.json" "*.pt" \
  --local-dir ./aenea-pinta-1.1

Expected repository layout:

aenea-pinta-1.1/
β”œβ”€β”€ merlin_2/               # Custom architecture: config + modeling code
β”œβ”€β”€ pinta_engine.py         # Router engine & FastAPI middleware
β”œβ”€β”€ config.yaml             # Model + router configuration
β”œβ”€β”€ routing_map.json        # Reserved-token β†’ backend dispatch map
β”œβ”€β”€ tokenizer.json          # QT VI.6.4 UltraLingo tokenizer (49,152 vocab)
└── pinta_1.1_bf16.pt       # bf16 weights (~656 MB)

pinta_engine.py imports merlin_2 from the repository root at runtime β€” keep both in the same working directory (or on PYTHONPATH).


πŸš€ Router Engine Quickstart

Pinta-1.1 ships with a custom, production-ready routing engine (pinta_engine.py) that handles model loading (via the local merlin_2 module), fast logits extraction, and async dispatch to OpenAI-compatible endpoints or local Ollama instances.

1. Execute a Routing Pass (example_usage.py)

import asyncio
from pinta_engine import PintaRouter

# Initialize the router engine (loads bf16 weights via the custom merlin_2 architecture)
router = PintaRouter("config.yaml")

prompt = "Write a Python function to normalize a file path."

# 1. Ultra-fast routing decision
decision = router.route(prompt)

print(decision.token)
print(decision.label)
print(decision.confidence)
print(decision.downstream_model)

# 2. Async dispatch to the mapped model backend (e.g., local Ollama qwen2.5-coder)
# response = asyncio.run(router.adispatch(prompt, decision))
# print(response["choices"][0]["message"]["content"])

Expected Output

2026-09-18 11:57:41,512 INFO pinta.router [Pinta Router] Initializing custom Merlin (merlin_2) architecture...
2026-09-18 11:57:46,784 INFO pinta.router [Pinta Router] Initialized Pinta-1.1 router on cuda with dtype=torch.bfloat16, threshold=0.350, fallback=<|reserved_29|>
2026-09-18 11:57:47,117 INFO pinta.router [Pinta Router] Prompt -> Mapped to <|reserved_24|> (Code) -> Dispatching to qwen2.5-coder:32b (Latency: 333ms).
<|reserved_24|>
Code
0.9890856146812439
qwen2.5-coder:32b

The 333 ms figure above is the first (cold) routing pass including CUDA kernel warm-up. Warm-cache passes are substantially faster; the benchmarked average TTFT of 192.66 ms is measured across the full 1,020-prompt evaluation on reference hardware.

2. Configure Dispatch Targets (routing_map.json)

The router engine decides which domain token a prompt belongs to; routing_map.json tells it where to send that prompt. Each reserved token is bound to a downstream backend (Ollama, vLLM / any OpenAI-compatible server, the OpenAI API, or a local Hugging Face pipeline):

{
  "<|reserved_24|>": {
    "label": "Code",
    "backend": "ollama",
    "base_url": "http://localhost:11434",
    "model": "qwen2.5-coder:32b"
  },
  "<|reserved_26|>": {
    "label": "RAG / Doc QA",
    "backend": "vllm",
    "base_url": "http://localhost:8000/v1",
    "model": "qwen2.5-7b-instruct"
  },
  "<|reserved_29|>": {
    "label": "Knowledge",
    "backend": "openai",
    "model": "gpt-4o-mini",
    "api_key_env": "OPENAI_API_KEY"
  }
}
  • If the router's confidence for the predicted token falls below the configured threshold (default 0.35), the prompt is automatically routed to the Knowledge fallback (<|reserved_29|>).
  • Backends, base URLs, API-key environment variables, and per-target default generation parameters are all configurable per token.

πŸ“ˆ Benchmark Performance Results

Evaluated on 1,020 curated cross-domain evaluation prompts sourced from benchmark datasets.

Overall Metric Result
Strict Accuracy (Primary Token Match) 84.61%
Flexible Accuracy (Allowed Token Match) 87.65%
Average Latency (TTFT) 192.66 ms

Per-Domain Classification Performance

Per-Domain Performance Summary (n = 1,020)

Domain Token Domain Name Precision Recall F1-Score Support (n)
<|reserved_23|> Formatting 0.50 0.05 0.09 43
<|reserved_24|> Code 0.83 0.94 0.88 281
<|reserved_25|> Creative 0.70 0.23 0.34 31
<|reserved_26|> RAG / Doc QA 1.00 0.83 0.91 24
<|reserved_27|> Architecture 0.31 0.56 0.40 9
<|reserved_28|> Math 0.92 0.92 0.92 260
<|reserved_29|> Knowledge 0.83 0.88 0.86 369
<|reserved_30|> Law / Policy 0.00 0.00 0.00 3
Weighted Average β€” 0.83 0.85 0.83 1020

Support values sum to the full 1,020-prompt evaluation set. See Known Limitations for details on the Formatting and Law / Policy classes.


πŸ“ Model Specifications

  • Base Checkpoint: AENEA Prelude-6
  • Inference Size: ~656 MB (bfloat16)
  • Training Setup: Pre-trained on NVIDIA RTX 4090 / Fine-tuned on NVIDIA RTX 4060
Parameter Value Description
Parameters 226M Low-footprint architecture for sub-200ms TTFT routing
Vocab Size 49,152 QT VI.6.4 UltraLingo Tokenizer
Layers 16 Transformer blocks with Parallel Residuals
Hidden Size (d_model) 1024 Dense embedding dimension
Attention Heads 16 Query / 4 KV Grouped-Query Attention (GQA)
Intermediate Size (d_ff) 4096 SwiGLU / Feed-Forward projection
Max Sequence Length 2048 Rotary Position Embeddings (RoPE)
Precision bfloat16 Optimized for native GPU execution

πŸ“š Training Architecture & Datasets

Pinta-1.1 was trained in a two-stage pre-training and alignment pipeline:

1. Base Pre-Training (Prelude-6 Checkpoint)

  • English Wikipedia: Complete clean English corpus.
  • Stack Exchange Network: Technical QA corpora including StackOverflow, MathOverflow, ServerFault, SuperUser, AskUbuntu, CodeReview, SoftwareEngineering, ReverseEngineering, and NetworkEngineering.
  • Source Code Corpus: Multi-language source code covering Python, C, C++, C#, Java, JavaScript, TypeScript, Rust, Go, PHP, Ruby, Kotlin, Swift, and Shell, plus CodeSearchNet Python.
  • Mathematics & Science: Cleaned pure mathematics and scientific reasoning corpora.

2. Fine-Tuning & Contrastive Alignment

  • Curated Open Data Mix (50%): Mined imperative subsets targeting instruction execution across Code, System Architecture, Formatting, and Reasoning.
  • Deepseek-v4-Flash Synthetic Boundary Data (50%): Hard-negative contrastive pairs synthetically generated via Deepseek-v4-Flash to sharpen semantic boundaries between domain classes.

πŸ›‘οΈ Known Limitations & Beta Scope

  • Formatting (<|reserved_23|>) has low recall (0.05). This is a known consequence of class imbalance in the fine-tuning mix: formatting-style requests (markdown tables, JSON scaffolding, template rewrites) are under-represented relative to Code and Knowledge. In practice, most formatting prompts are absorbed by Knowledge (<|reserved_29|>) or Code (<|reserved_24|>) as a safe fallback, which preserves downstream answer quality at the cost of routing granularity. Its 0.50 precision also means half of the prompts it does claim are misrouted. Improving this class is a priority for the post-Beta release.
  • Law / Policy (<|reserved_30|>) is unvalidated. Evaluation support is n = 3, so its reported scores are not statistically meaningful. Treat Law routing as experimental in this Beta.
  • <|reserved_31|> is reserved for future domain expansion and is not scored in the Beta benchmark.
  • Confidence fallback: routing decisions below the confidence threshold (default 0.35) are redirected to Knowledge (<|reserved_29|>). Deployments with strict domain-separation requirements should tune this threshold and monitor fallback rates.
  • Hardware-dependent latency: TTFT figures are measured on the reference consumer-GPU setup above; CPU-only or shared-GPU deployments will see higher latency.
  • Router-only model: Pinta-1.1 emits routing decisions and should not be used as a generative assistant.

πŸ”— Links & Contact


πŸ“„ License

This project is licensed under the Apache 2.0 License. Β© 2026 AENEA Global LTD

Downloads last month
8
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Evaluation results