Instructions to use erdi28/MolmoPoint-8B-bnb-int8-native with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use erdi28/MolmoPoint-8B-bnb-int8-native with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="erdi28/MolmoPoint-8B-bnb-int8-native", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForImageTextToText model = AutoModelForImageTextToText.from_pretrained("erdi28/MolmoPoint-8B-bnb-int8-native", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use erdi28/MolmoPoint-8B-bnb-int8-native with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "erdi28/MolmoPoint-8B-bnb-int8-native" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "erdi28/MolmoPoint-8B-bnb-int8-native", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/erdi28/MolmoPoint-8B-bnb-int8-native
- SGLang
How to use erdi28/MolmoPoint-8B-bnb-int8-native with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "erdi28/MolmoPoint-8B-bnb-int8-native" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "erdi28/MolmoPoint-8B-bnb-int8-native", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "erdi28/MolmoPoint-8B-bnb-int8-native" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "erdi28/MolmoPoint-8B-bnb-int8-native", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use erdi28/MolmoPoint-8B-bnb-int8-native with Docker Model Runner:
docker model run hf.co/erdi28/MolmoPoint-8B-bnb-int8-native
MolmoPoint-8B — bitsandbytes LLM.int8
This is an inference-ready bitsandbytes LLM.int8 quantization of
allenai/MolmoPoint-8B. All
eligible linear modules are quantized except the 7.5M-parameter pointing head,
which remains BF16. This avoids a runtime patch because bitsandbytes 0.50.0's
INT8 kernel does not natively accept the pointing head's high-rank inputs.
Evaluation overview
The figure compares both released quantizations with the BF16 baseline. See Evaluation for coverage, protocol, and metric details.
Model details
- Repository:
erdi28/MolmoPoint-8B-bnb-int8-native - Base model:
allenai/MolmoPoint-8B - Base revision:
188130f961c8e0888a34e11121a1423c461a01ba - Quantized: all eligible linear modules outside
model.point_predictor - Retained in BF16:
model.point_predictor(7.5M parameters; under 0.1% of the model) - Outlier threshold: 6.0
- Local checkpoint size: approximately 9.3 GiB
- Source and evaluation code:
erkara/molmopoint-quantization
The included quantization_provenance.json records the resolved base revision,
complete quantization recipe, software versions, and build environment.
Requirements
The checkpoint was built and validated with Python 3.12, Transformers 4.57.1, bitsandbytes 0.50.0, and a CUDA-capable NVIDIA GPU. It contains the processor and trusted custom model code required for native Transformers loading.
pip install "transformers==4.57.1" "bitsandbytes==0.50.0" accelerate torch torchvision pillow einops decord2
Usage
import torch
from PIL import Image
from huggingface_hub import hf_hub_download
from transformers import AutoModelForImageTextToText, AutoProcessor
repo_id = "erdi28/MolmoPoint-8B-bnb-int8-native"
processor = AutoProcessor.from_pretrained(
repo_id,
trust_remote_code=True,
padding_side="left",
use_fast=False,
)
model = AutoModelForImageTextToText.from_pretrained(
repo_id,
trust_remote_code=True,
dtype="auto",
device_map="auto",
)
model.eval()
image_path = hf_hub_download(repo_id, "examples/pointing-demo.jpg")
image = Image.open(image_path).convert("RGB")
prompt = "Point to the tool that people can use to write."
messages = [{
"role": "user",
"content": [
{"type": "text", "text": prompt},
{"type": "image", "image": image},
],
}]
inputs = processor.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
return_dict=True,
padding=True,
return_pointing_metadata=True,
)
metadata = inputs.pop("metadata")
inputs = {name: value.to("cuda") for name, value in inputs.items()}
prompt_length = inputs["input_ids"].shape[1]
with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16):
output = model.generate(
**inputs,
logits_processor=model.build_logit_processor_from_inputs(inputs),
do_sample=False,
max_new_tokens=128,
)
generated = output[:, prompt_length:]
generated_text = processor.post_process_image_text_to_text(
generated,
skip_special_tokens=False,
clean_up_tokenization_spaces=False,
)[0]
points = model.extract_image_points(
generated_text,
metadata["token_pooling"],
metadata["subpatch_mapping"],
metadata["image_sizes"],
)
print(generated_text)
print(points)
Example output
This checkpoint returns two points near (324, 100) and (627, 187) in the
1200 x 800 source image, corresponding to the blue and red mechanical pencils.
The demo image is a compressed copy of a
CC0 Public Domain image from PxHere.
MolmoPoint emits special grounding tokens, which extract_image_points
converts into source-image coordinates.
Evaluation
All results use greedy decoding, batch size 1, the uber_model_v2 prompt
templates, and AllenAI's official Molmo2 formatter and evaluators pinned to
commit f3cb1085fbb97c4a4d7fdcadd77cf871bf37a88a.
| Benchmark | Coverage | BF16 base | This checkpoint | Delta |
|---|---|---|---|---|
| PointBench average accuracy | 982/982 | 70.90% | 71.00% | +0.10 pt |
| PixMo-Points F1 | 231/436 | 84.02% | 84.49% | +0.47 pt |
All selected examples passed with zero evaluation failures. PointBench is a
complete public evaluation. PixMo-Points publishes 436 metadata rows through
allenai/pixmo-points-eval,
but only 231 external image URLs still returned bytes matching the published
SHA-256 hashes. The PixMo score therefore has 52.98% public-data coverage and is
not directly comparable with the paper's full-dataset result.
| Measurement | BF16 base | This checkpoint |
|---|---|---|
| Peak allocated VRAM across both evaluations | 18.26 GiB | 11.68 GiB |
| PointBench mean inference time | 1.4611 s/example | 1.9597 s/example |
| PixMo-Points mean inference time | 1.0026 s/example | 1.4361 s/example |
Peak allocated VRAM was 36.0% lower than BF16 on the benchmark system. INT8 was slower than BF16 in both evaluations; its benefit here is reduced memory and storage rather than speed. Memory and latency are environment-specific and should not be treated as universal hardware requirements. Full category scores and protocol details are available in the project results.
Intended use
This checkpoint is intended for research, evaluation, and CUDA inference where reduced memory use is useful and MolmoPoint's image, multi-image, video, and grounding capabilities are required. It is a quantized derivative, not a fine-tuned model, and does not add new capabilities or safety training.
Limitations
This release inherits the base model's limitations, biases, and failure modes.
Quantization can change individual generations even when aggregate scores are
close. The evaluation does not establish identical or lossless behavior, and
PixMo-Points coverage is partial. Loading requires trust_remote_code=True and
a CUDA environment supported by bitsandbytes.
License and responsible use
This derivative retains the base model's Apache-2.0 license and attribution. Use is also subject to Ai2's Responsible Use Guidelines. The upstream model card states that MolmoPoint-8B was trained on third-party datasets subject to academic and non-commercial research-use terms. Review the base model card and applicable source-dataset terms before use or redistribution.
- Downloads last month
- 15

