Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs Paper ⢠2310.01801 ⢠Published Oct 3, 2023 ⢠3
DeepSpeed4Science Initiative: Enabling Large-Scale Scientific Discovery through Sophisticated AI System Technologies Paper ⢠2310.04610 ⢠Published Oct 6, 2023 ⢠1
ZeRO-Offload: Democratizing Billion-Scale Model Training Paper ⢠2101.06840 ⢠Published Jan 18, 2021 ⢠1
Computing in the Era of Large Generative Models: From Cloud-Native to AI-Native Paper ⢠2401.12230 ⢠Published Jan 17, 2024 ⢠2
DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales Paper ⢠2308.01320 ⢠Published Aug 2, 2023 ⢠46
DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale Paper ⢠2201.05596 ⢠Published Jan 14, 2022 ⢠2
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model Paper ⢠2211.05100 ⢠Published Nov 9, 2022 ⢠40
Random-LTD: Random and Layerwise Token Dropping Brings Efficient Training for Large-scale Transformers Paper ⢠2211.11586 ⢠Published Nov 17, 2022 ⢠1