pollix/stuntd-support-triage is three small heads on the Laya encoder that triage a support ticket in one request: category, urgency and needs_human. About 50 MB each, all three answers come back at a p50 of 71ms through the daemon.
On 1,000 tickets they never saw, each head answers on its own when it's sure: category 99.9%, needs_human 92%, urgency 76%. A ticket only skips the big model when all three are sure, that's 72.7% of them, and all three are right on 97.1% of those.
It's the support demo from the repo, so the tickets are generated and the teacher is a rule. The point is to show what a head looks like and how fast it is, then you train the same thing on your own traffic with your own LLM as the teacher.
stuntd sits in front of your LLM, learns its typed decisions and answers the confident ones locally with a small head on the Laya encoder by @convaiinnovations. About 20ms on GPU and 60ms on CPU, and anything it isn't sure about still goes to the big model.
New in 0.1.2: - decisions with several fields, like category + urgency + needs_human in one call, answered locally only when every field is sure - the Anthropic Messages API learns too, not only OpenAI - auto_retrain: the daemon retrains a site in the background once enough new traffic comes in, so collect, train, shadow and live run on their own - serve --lazy loads the checkpoint on the first request
For each question it shows the same answer twice: zero-shot Laya, and a small head trained on the frozen Laya encoder. You get the option probabilities, the confidence threshold and the latency. When the head isn't sure, it hands the question back to the teacher instead of guessing.
On jevbench, same machine, n=500: agnews 86.0 → 92.6, banking77 38.2 → 69.6. Everything runs locally; heads train in 4-20 minutes on a laptop GPU.
SKT AI Labs, we are pushing the boundaries of AI architecture and research—and today, we are thrilled to open our doors to the global research community!
We warmly welcome researchers, developers, and AI enthusiasts to join us and contribute to our R&D efforts.
🧪 What You Can Explore:
We invite you to experiment with our WMF (Weight Manifold Fusion) technology. You can test this high-dimensional fusion technique on smaller models to gain a deeper understanding of its behavior and token convergence.
If it works: Fantastic! Share your results with us and contribute directly to the core vision of SKT AI Labs.
If it doesn't work: No problem at all! Your critical feedback is just as valuable to us. Every experiment and anomaly helps us refine this architecture to make it more stable and robust.
We firmly believe that true innovation stems from community collaboration and transparent testing. Let's build the future of advanced AI together. Your ideas, test results, and feedback are always welcome!
You Can Still Research and Development On WMF Only SKT-SURYA-H Model is Dismissed.
We’re excited to release NRS_QWEN_MYTHOS_1M — a powerful reasoning model built on Qwen 3.5 9B! At SKT AI LABS, we’ve supercharged this 9B model with our proprietary Neural Reasoning System (NRS) to deliver next-level performance.
🔥 Why This Model is a Game-Changer: ✅ 100x Reasoning Capacity — Exceptional deep logical thinking and complex problem-solving ✅ 1 Million Token Context — Perfect for massive codebases, long documents, and multi-turn agentic workflows ✅ Advanced Thinking Mode — Native <think> tags for true step-by-step Chain-of-Thought reasoning ✅ Tool-Use Ready — Optimized for Python execution, Web Search, and self-correction ✅ Blazing Fast — Runs smoothly on consumer GPUs like RTX 3090/4090
Whether you’re a developer building coding agents, a researcher working with long-context data, or someone who loves powerful reasoning — this model is built for you.
Shipped StudioMI300 for the AMD x lablab hackathon. One English sentence becomes a 30-second cinematic reel, end-to-end on a single AMD Instinct MI300X.
Every model in the pipeline is Apache 2.0 or MIT.
🎬 Director Agent — Qwen3.5-35B-A3B via vLLM with AITER MoE acceleration. Plans 6 shots, character bibles, music brief, per-shot voice-over.
🎨 Character keyframes — FLUX.2 klein 4B reference editing. No LoRA training step. Identity stays consistent across shots by construction.
🎞️ Animation — Wan2.2-I2V-A14B with ParaAttention FBCache (lossless 2x) and selective torch.compile on transformer_2 (another 1.2x). End-to-end Wan2.2 inference went from 25.9 min to 10.4 min per 720p clip.
🔍 Vision Critic — same Qwen3.5 checkpoint reloaded with a 10-label failure taxonomy (character drift, extras invade frame, camera ignored, walking backwards, hand artifact, wardrobe drift, neon glow leak, stylized AI look, random intimacy, object morphing). Bad clips auto-retry with targeted strategies. Up to 3 attempts.
🎵 Music — ACE-Step v1 generates 30s instrumental from Director's brief.
🗣️ Narration — Kokoro-82M, 9 languages. Director picks language to match setting. Tokyo to Japanese, Paris to French, Mumbai to Hindi.
The 192 GB HBM3 on MI300X is what lets four very different model architectures share one card sequentially. On a 24 GB consumer GPU this stack needs 4-5 separate machines wired together.
Special thanks to the FLUX, Wan2.2, ACE-Step and Kokoro teams for keeping serious generative AI open. The pipeline composes their work into something none of them alone can produce — a complete cinematic artifact from a single prompt.