Granite-Docling-258M-BitNet-1.58b 🚀

This repository hosts the first-ever BitNet b1.58 ternary (-1, 0, 1) quantized edition of IBM's ibm-granite/granite-docling-258M.

Trained via Knowledge Distillation with Straight-Through Estimators (STE) on the Edge-Intelligence framework, this model replaces 210 matrix multiplication operations with multiplication-free ternary addition and subtraction, reducing parameter memory footprint by ~3.7x while achieving 93.7% logit fidelity to the original FP32 teacher model.

📊 Benchmark Highlights

Metric Original FP32 BitNet b1.58 (This Model) Improvement
Model Size 530 MB 143.85 MB 3.68x Compression
Weight Precision 16-bit Float 1.58-bit Ternary (-1, 0, 1) 10.1x weight density
Logit Cosine Similarity 1.0000 0.9374 (%93.74) High semantic fidelity
Top-1 Token Agreement <doctag> (100327) <doctag> (100327) 100% Match
C99 Inference Throughput (AVX2) ~40 tok/s 156.6 tok/s ~3.9x Speedup
Latency per Token (C99) 25.0 ms 6.38 ms Ultra-responsive

🛠️ Usage via Hugging Face (transformers)

You can load and run this model directly using trust_remote_code=True:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "tevfikk/granite-docling-258M-bitnet"

# 1. Load Tokenizer & Model
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id, 
    trust_remote_code=True, 
    torch_dtype=torch.float32
)

# 2. Format Input with Docling Chat Template
prompt = "<|start_of_role|>user<|end_of_role|>Convert this document to markdown.<|end_of_text|>\n<|start_of_role|>assistant<|end_of_role|>"
inputs = tokenizer(prompt, return_tensors="pt")

# 3. Generate DocTags
with torch.no_grad():
    outputs = model.generate(**inputs, max_new_tokens=32, do_sample=False)

print(tokenizer.decode(outputs[0][inputs['input_ids'].shape[1]:]))
# Output: <doctag><section_header_level_1>...

⚡ Edge Binary (.eifm)

Included in this repository is granite_docling_bitnet_dense.eifm (143 MB), a zero-heap, single-file binary container format optimized for multiplication-free embedded C99 inference on microcontrollers, Raspberry Pi, and edge accelerators using EIF-Runtime:

# Clone the open-source C99 runtime & build
git clone https://github.com/tevfik/eif-runtime.git
cd eif-runtime && cmake -B build && cmake --build build

# Run at 156+ tokens/sec without Python or CUDA!
./build/eif-run granite_docling_bitnet_dense.eifm "Invoice Summary" 64

📜 Citation & Credits

Downloads last month
-
Safetensors
Model size
0.2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tevfikk/granite-docling-258M-bitnet

Finetuned
(11)
this model

Paper for tevfikk/granite-docling-258M-bitnet