MARQUIS: A Three-Stage Pipeline for Video Retrieval-Augmented Generation
Paper • 2605.17640 • Published
Contributors who are invited to beta-test our next big feature! Contact us if you want to join this team :-)
hf-mem release added a breakdown of Mixture-of-Experts (MoE) memory usage!hf-mem now splits MoE memory into base model weights, routed experts, and KV cachehf-mem v0.4.1 now also estimates KV cache memory requirements for any context length and batch size with the --experimental flag!uvx hf-mem --model-id ... --experimental will automatically pull the required information from the Hugging Face Hub to include the KV cache estimation, when applicable.--max-model-len, --batch-size and --kv-cache-dtype arguments (à la vLLM) manually if preferred.