Share Qwen3.5 prefix cache across NLI hypotheses

#1
by epsilon3 - opened

Add batched cached continuation for hypotheses sharing one premise. Qwen3.5 recurrent linear-attention state and full-attention KV cache are branched per hypothesis after one prefix prefill. Rerank uses the new path; tiny hybrid-model tests compare against independent pair inference and check branch isolation.

Im gonna add integration for sglang, but huge thx

AlexWortega changed pull request status to merged

Sign up or log in to comment