stuntd sits in front of your LLM, learns its typed decisions and answers the confident ones locally with a small head on the Laya encoder by @convaiinnovations. About 20ms on GPU and 60ms on CPU, and anything it isn't sure about still goes to the big model.
New in 0.1.2: - decisions with several fields, like category + urgency + needs_human in one call, answered locally only when every field is sure - the Anthropic Messages API learns too, not only OpenAI - auto_retrain: the daemon retrains a site in the background once enough new traffic comes in, so collect, train, shadow and live run on their own - serve --lazy loads the checkpoint on the first request