Methods for compressing the KV cache in transformer inference. E8 lattice, KVarN, TurboQuant, KIVI.
Compare LLM outputs with KV cache compression modes