Instructions to use migtissera/Tess-4-35B-A3B-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use migtissera/Tess-4-35B-A3B-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="migtissera/Tess-4-35B-A3B-NVFP4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("migtissera/Tess-4-35B-A3B-NVFP4") model = AutoModelForMultimodalLM.from_pretrained("migtissera/Tess-4-35B-A3B-NVFP4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use migtissera/Tess-4-35B-A3B-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "migtissera/Tess-4-35B-A3B-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "migtissera/Tess-4-35B-A3B-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/migtissera/Tess-4-35B-A3B-NVFP4
- SGLang
How to use migtissera/Tess-4-35B-A3B-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "migtissera/Tess-4-35B-A3B-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "migtissera/Tess-4-35B-A3B-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "migtissera/Tess-4-35B-A3B-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "migtissera/Tess-4-35B-A3B-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use migtissera/Tess-4-35B-A3B-NVFP4 with Docker Model Runner:
docker model run hf.co/migtissera/Tess-4-35B-A3B-NVFP4
Tess-4-35B-A3B-NVFP4
This is the NVIDIA ModelOpt NVFP4 deployment checkpoint for
migtissera/Tess-4-35B-A3B,
pinned to source revision cae99edb934875a1977774502664dedf03c211db.
The model's routed MoE experts are packed as NVFP4, with calibrated FP8 KV-cache scales. The vision tower, multimodal projection path, shared experts, attention layers, embeddings, LM head, and included MTP head remain in BF16. The export retains the complete Qwen3.5 multimodal processor and the native 262,144-token context configuration.
Conversion details
- GPU: one NVIDIA B200
- NVIDIA ModelOpt commit:
f479e7890f0d276e061d69b5a0d70d477ec863f2 - Quantization format:
nvfp4_experts_only - KV-cache format: calibrated FP8
- Calibration data: 116 real Tess training conversations
- Calibration window: one deterministic 4,096-token window per conversation
- Export size: 25,615,467,408 bytes across three safetensor shards
- MTP: all 19 source tensors preserved in BF16
- Vision: all 333 source vision tensors preserved
vLLM
Use a recent vLLM build with ModelOpt NVFP4 and Qwen3.5 MoE VLM support:
vllm serve migtissera/Tess-4-35B-A3B-NVFP4 \
--quantization modelopt_fp4 \
--kv-cache-dtype fp8 \
--dtype bfloat16 \
--max-model-len 262144 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--trust-remote-code
The first validated deployment uses plain decoding. The preserved BF16 MTP head can be enabled later as a serving configuration change after validating the base NVFP4 path.
- Downloads last month
- 493
Model tree for migtissera/Tess-4-35B-A3B-NVFP4
Base model
migtissera/Tess-4-35B-A3B