How to use from the
Use from the
Transformers library
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("LittleBitLLM/littlebit", device_map="auto")
Quick Links

LittleBit.

LittleBit

Big models. Little bits. Sub-1-bit Qwen3 models built with LittleBit-2: binarized latent factorization, recovered by distillation from the original model.

Status: weights coming soon. This repo holds the training recipe and measured throughput. Model weights and eval results will be added here when the first full training run finishes. Follow @LittleBit_llm for updates.

Independent community project. Not affiliated with or endorsed by Samsung Research. Built on the LittleBit method and code by Lee, Kim, You & Kim (SamsungLabs/LittleBit).

Planned releases

Model Base Target bpw Status
littlebit-qwen3-4b Qwen/Qwen3-4B 0.55 planned
littlebit-qwen3-8b Qwen/Qwen3-8B 0.55 planned
littlebit-qwen3-14b Qwen/Qwen3-14B 0.55 planned

Bits per weight apply to linear layers. Embeddings and lm_head stay BF16.

Method

Each linear layer W is approximated as sign(U) · diag(h·g·ℓ) · sign(V)ᵀ: low-rank latent factors binarized to ±1, plus three thin learned scale vectors.

  1. Latent factorization: SVD splits each linear layer into rank-r factors sized to the bit budget.
  2. Joint-ITQ rotation (LittleBit-2): aligns the factors with the binary hypercube before training. It folds into the factors, so it adds no inference cost.
  3. SmoothSign binarization: a smooth surrogate gradient keeps the sign step trainable.
  4. Residual compensation: a second binarized path learns what the first one missed.
  5. Distillation: quantization-aware training on C4 + WikiText-2 (seq len 2048), with the BF16 model as teacher (logit KL + layer-to-layer MSE).

Measured throughput

Measured on RunPod with the recipe in recipe/. One step = 4 sequences × 2048 tokens. One epoch of the C4-shard-0 + WikiText-2 mix is about 20,750 steps with the Qwen3 tokenizer.

Model GPU Sec / step Peak VRAM One-time init (SVD + Joint-ITQ) Est. 1 epoch
Qwen3-0.6B @ 0.55 bpw 1× H100 80GB 1.03 — ~2.3 min ~6 h
Qwen3-8B @ 0.55 bpw 1× H200 141GB 2.84 ~107 GB ~14 min ~16.5 h

Qwen3-8B does not fit on a single 80 GB GPU: about 3.7B latent parameters are trainable, and their optimizer state alone exceeds the memory. Use a 141 GB GPU or ≥ 2 GPUs with ZeRO-3.

Recipe

recipe/ runs the official LittleBit code on a RunPod GPU pod:

  • setup.sh: clones SamsungLabs/LittleBit at a pinned commit, applies the patches, and installs dependencies (transformers==4.51.*, DeepSpeed).
  • train.sh: runs QAT. Defaults: Qwen3-8B, 0.55 bpw, LittleBit-2 init, SmoothSign, residual. Override settings with env vars (MODEL_ID, EFF_BIT, EPOCHS, NUM_GPUS, …).
  • eval.sh: measures WikiText-2/C4 perplexity and zero-shot accuracy (lm-eval).
  • zero3_nooffload.json: multi-GPU ZeRO-3 config without CPU offload.
  • patches/teacher-on-gpu.patch: keeps the teacher on GPU instead of ZeRO-3 CPU offload (--teacher_offload False).
  • patches/eval-import-fix.patch: fixes a circular import between lm-eval, transformers, and DeepSpeed in eval.py.
bash recipe/setup.sh
MODEL_ID=Qwen/Qwen3-8B EFF_BIT=0.55 EPOCHS=1 bash recipe/train.sh
CKPT=/workspace/outputs/littlebit-qwen3-8b-0.55bpw bash recipe/eval.sh

License

CC BY-NC 4.0 (non-commercial), inherited from the LittleBit code. Released weights are also subject to the base model's license (Qwen3: Apache 2.0).

Citation

@inproceedings{lee2025littlebit,
  title     = {LittleBit: Ultra Low-Bit Quantization via Latent Factorization},
  author    = {Lee, Banseok and Kim, Dongkyu and You, Youngcheon and Kim, Youngmin},
  booktitle = {NeurIPS},
  year      = {2025}
}
@inproceedings{lee2026littlebit2,
  title     = {LittleBit-2: Maximizing the Spectral Energy Gain in Sub-1-Bit LLMs via Latent Geometry Alignment},
  author    = {Lee, Banseok and Kim, Youngmin},
  booktitle = {ICML},
  year      = {2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for LittleBitLLM/littlebit

Finetuned
Qwen/Qwen3-8B
Quantized
(459)
this model