Instructions to use google/gemma-4-12B-it-qat-w4a16-ct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use google/gemma-4-12B-it-qat-w4a16-ct with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("google/gemma-4-12B-it-qat-w4a16-ct") model = AutoModelForMultimodalLM.from_pretrained("google/gemma-4-12B-it-qat-w4a16-ct", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Raise per-image vision soft-token budget from 280 to 1120
1
#8 opened 16 days ago
by
lucianommartins
Google, Open Source Old Bison PaLM 2 Model In 2026/27
👍 1
#6 opened about 1 month ago
by
Tralalabs
vLLM Help
❤️ 1
3
#3 opened about 2 months ago
by
allenc87
Tested it and its the Bomb! Vram efficient with high Quality.
8
#2 opened about 2 months ago
by
JanjanJean
unable to save AOT compiled function
1
#1 opened about 2 months ago
by
mohamedemam