Hybrid MLA + Gated DeltaNet / Mamba-2 models from long-context aware upcycling: up to 32x usable context with over 90% KV-cache reduction.