d1-3B-W4A16-AutoRound

This repository contains a 4-bit W4A16 quantized version of LiquidAI/d1-3B, optimized using AutoRound.

d1-3B is a multimodal decision model built on LFM2.5-VL-3B. It evaluates states (text, images, or JSON) against typed questions in a single forward pass with zero autoregressive output tokens.

Quantization Details

  • Method: AutoRound (W4A16)
  • Group Size: 32 (group_size=32, sym=True)
  • Calibration: 512 samples, 800 tuning iterations, sequence length 2048
  • Vision Tower: Retained in unquantized precision (quant_nontext_module=False) to preserve full visual feature fidelity
  • Format: Compatible with auto_round and AutoGPTQ backends

Quickstart

Installation

pip install --upgrade "transformers>=5.14" torch torchvision pillow auto-round
# Optional for optimized kernels:
pip install auto-round-lib

Inference Example

import torch
from transformers import AutoModel
from transformers.image_utils import load_image

device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if torch.cuda.is_available() else torch.float32

# Load quantized model
model = AutoModel.from_pretrained(
    "Vishva007/d1-3B-W4A16-AutoRound",
    trust_remote_code=True,
    dtype=dtype,
    device_map="auto"
)

# 1. Text Classification / Decision
questions = {
    "refund": {
        "type": "noul",
        "instructions": "Is the customer asking for a refund?",
    },
    "team": {
        "type": "choice",
        "instructions": "Which team should handle this?",
        "criteria": {
            "billing": "Charges, refunds, invoices",
            "technical": "App or site faults",
            "fraud": "Suspected unauthorised use",
        },
    },
}

prompt = "I was charged twice this month, please refund one of them."
res_text = model.system_one(prompt, questions)
print("Text Decision:", res_text)

# 2. Multimodal / Image Evaluation
image = load_image("[http://images.cocodataset.org/val2017/000000039769.jpg](http://images.cocodataset.org/val2017/000000039769.jpg)")
cats_q = {
    "type": "choice",
    "instructions": "How many cats are there?",
    "criteria": {"one": "One", "two": "Two", "more": "Three or more"},
}

res_image = model.system_one(None, {"cats": cats_q}, images=[image])
print("Vision Decision:", res_image)

Intended Use

  • Single-pass decision classification (Yes/No noul, multi-choice choice, and calibrated scales score)
  • Triage, routing, intent detection, content moderation, agent safety guards
  • Zero-token generation latency overhead at reduced VRAM footprint

🚀 Deploy on RunPod

One-click launch environments pre-configured with PyTorch, CUDA, and dependencies for fine-tuning or quantization.

🎁 Need GPU compute? Sign up via RunPod and get $5–$500 in free credits when you add your first $10.

PyTorch 2.14

Template CUDA Version Docker Image Template ID Deploy
PyTorch 2.14 (CUDA 12.6) 12.6 vishva123/cuda-12.6-pytorch-2.14-runpod d7lxsa4w9m Deploy to RunPod
PyTorch 2.14 (CUDA 13.0) 13.0 vishva123/cuda-13.0-pytorch-2.14-runpod yk0y6j6rpg Deploy to RunPod
PyTorch 2.14 (CUDA 13.2) 13.2 vishva123/cuda-13.2-pytorch-2.14-runpod gsp4gwx0nw Deploy to RunPod

PyTorch 2.13

Template CUDA Version Docker Image Template ID Deploy
PyTorch 2.13 (CUDA 12.6) 12.6 vishva123/cuda-12.6-pytorch-2.13-runpod gmlupxnxfk Deploy to RunPod
PyTorch 2.13 (CUDA 13.0) 13.0 vishva123/cuda-13.0-pytorch-2.13-runpod y3j8xvk4f4 Deploy to RunPod
PyTorch 2.13 (CUDA 13.2) 13.2 vishva123/cuda-13.2-pytorch-2.13-runpod vigpissn5w Deploy to RunPod

PyTorch 2.12

Template CUDA Version Docker Image Template ID Deploy
PyTorch 2.12 (CUDA 12.6) 12.6 vishva123/cuda-12.6-pytorch-2.12-runpod ctmz86zmf0 Deploy to RunPod
PyTorch 2.12 (CUDA 13.0) 13.0 vishva123/cuda-13.0-pytorch-2.12-runpod qjko5yiwzi Deploy to RunPod
PyTorch 2.12 (CUDA 13.2) 13.2 vishva123/cuda-13.2-pytorch-2.12-runpod ifg6xmye0f Deploy to RunPod

Acknowledgements

Downloads last month
-
Safetensors
Model size
1B params
Tensor type
I32
·
BF16
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Vishva007/d1-3B-W4A16-AutoRound

Finetuned
LiquidAI/d1-3B
Quantized
(15)
this model

Collection including Vishva007/d1-3B-W4A16-AutoRound