Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss Paper • 2608.03796 • Published Aug 4 • 20
LLM Compression by Block Removal with Constrained Binary Optimization Paper • 2602.00161 • Published Jun 17 • 9
Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs Paper • 2608.20953 • Published Aug 21 • 14
Tiel-Coder-35B-A3B Collection A dedicated coder improving on Ornith1.5 without fine tuning. • 6 items • Updated 27 days ago • 8
Cyber-Tiel-Coder-35B-A3B Collection The improved an uncensored Tiel-Coder • 6 items • Updated 26 days ago • 16
view article Article Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original MultiverseComputingCAI • Aug 25 • 70
Unsloth Dynamic 3.0 Quants Collection Introducing Dynamic V3.0 quants, our new SOTA quantization methodology. • 1 item • Updated Aug 26 • 76
APEX Quants (GGUF) Collection MoE models quantized with the APEX Quantization technique ( https://github.com/mudler/apex-quant ) • 48 items • Updated 23 days ago • 149