Instructions to use google/gemma-2-2b-it with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use google/gemma-2-2b-it with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="google/gemma-2-2b-it") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("google/gemma-2-2b-it") model = AutoModelForCausalLM.from_pretrained("google/gemma-2-2b-it", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use google/gemma-2-2b-it with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "google/gemma-2-2b-it" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "google/gemma-2-2b-it", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/google/gemma-2-2b-it
- SGLang
How to use google/gemma-2-2b-it with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "google/gemma-2-2b-it" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "google/gemma-2-2b-it", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "google/gemma-2-2b-it" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "google/gemma-2-2b-it", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use google/gemma-2-2b-it with Docker Model Runner:
docker model run hf.co/google/gemma-2-2b-it
fix: Remediate 32 glitch tokens in tokenizer.json
Glitch Token Remediation — Tokenizer Patch
Summary
This PR patches tokenizer.json to remediate 32 glitch tokens identified
by embedding vector analysis. Glitch tokens are vocabulary entries whose embeddings
have collapsed to near-identical representations, causing unpredictable model behavior
when they appear in input.
No model weights are modified — this is a tokenizer-only fix.
Technique: Hybrid (merge pruning + placeholder rename)
The patch uses two complementary strategies:
1. Merge pruning (8 tokens): Removes the BPE merge rules that
produce multi-character glitch tokens. The vocab entry stays (ID preserved) but
becomes unreachable — text that would have matched is now tokenized as trained
constituent subwords instead.
2. Placeholder rename (24 tokens): For single-character glitch tokens
(rare Unicode characters) that have no producing merge rule, the vocab string is renamed to<glitch_pruned_ID>. The original character now falls through to byte_fallback,
encoding as well-trained UTF-8 byte tokens.
Safety: Placeholder strings cannot be triggered by user input. BPE builds tokens
bottom-up from characters via merge rules only — since no merge rule produces these strings,
they are permanently unreachable. Gemma's split-digits tokenization policy provides an
additional layer of safety.
10 merges were removed for 8 merge-pruned tokens because some glitch tokens have more than one BPE merge path, and every path must be removed.
Before / after examples: google/gemma-2-2b-it
Merge prune example: bildtitel, token ID 225065
Prompt: Please repeat the following string exactly, with nothing else: "bildtitel"
| Tokens for the string | Model output | |
|---|---|---|
| Before | ['bildtitel'] → [225065] |
" ❌ |
| After | ['bild', 'titel'] → [13328, 71990] |
bildtitel ✅ |
Placeholder rename example: 𑄣 (U+11123), token ID 253904
Prompt: Please repeat the following string exactly, with nothing else: "𑄣"
| Tokens for the string | Model output | |
|---|---|---|
| Before | ['𑄣'] → [253904] |
" ❌ |
| After | ['<0xF0>', '<0x91>', '<0x84>', '<0xA3>'] → [457, 362, 349, 380] |
𑄣 ✅ |
⚠️ Note for Agentic Systems
Many remediated glitch tokens are shared across Gemma model families (Gemma 1, 2, 3, 4).
In agentic pipelines, care should be taken that these token strings are not inadvertently
injected into prompts for unpatched models. We recommend applying this patch consistently
across all Gemma models used in a pipeline.
What is preserved
- ✅
len(tokenizer)— unchanged (256,000) - ✅ All token IDs — stable, no re-indexing
- ✅ Chat template — identical to original
- ✅
tokenizer_config.json— identical to original - ✅ All special tokens and
added_tokens— unchanged - ✅ Normal text tokenization — verified via benchmarks
- ✅ Zero model weight changes
Diff summary
| Original | Patched | Delta | |
|---|---|---|---|
| Vocab size | 256,000 | 256,000 | 0 |
| Merges | 580,604 | 580,594 | −10 |
| Vocab renamed | — | 24 | +24 |
Vocab renames (24 entries)
Click to expand all 24 renamed vocab entries
| ID | Original | Patched |
|---|---|---|
| 251496 | (U+0BA5) |
<glitch_pruned_251496> |
| 251497 | (U+0BA6) |
<glitch_pruned_251497> |
| 251632 | 𑄨 (U+11128) |
<glitch_pruned_251632> |
| 251921 | (U+0BA7) |
<glitch_pruned_251921> |
| 252083 | (U+F565) |
<glitch_pruned_252083> |
| 252372 | 𑄢 (U+11122) |
<glitch_pruned_252372> |
| 252915 | (U+F3F5) |
<glitch_pruned_252915> |
| 253103 | 𑄚 (U+1111A) |
<glitch_pruned_253103> |
| 253441 | (U+E984) |
<glitch_pruned_253441> |
| 253613 | (U+E0041) |
<glitch_pruned_253613> |
| 253841 | 𑄬 (U+1112C) |
<glitch_pruned_253841> |
| 253904 | 𑄣 (U+11123) |
<glitch_pruned_253904> |
| 254071 | (U+EF5A) |
<glitch_pruned_254071> |
| 254455 | (U+ED90) |
<glitch_pruned_254455> |
| 254456 | (U+EFA6) |
<glitch_pruned_254456> |
| 254686 | 𑄮 (U+1112E) |
<glitch_pruned_254686> |
| 255122 | (U+F540) |
<glitch_pruned_255122> |
| 255123 | 𑄥 (U+11125) |
<glitch_pruned_255123> |
| 255124 | 𑄪 (U+1112A) |
<glitch_pruned_255124> |
| 255517 | 𑄝 (U+1111D) |
<glitch_pruned_255517> |
| 255645 | (U+EF0E) |
<glitch_pruned_255645> |
| 255663 | (U+F023B) |
<glitch_pruned_255663> |
| 255806 | 𑄠 (U+11120) |
<glitch_pruned_255806> |
| 255849 | ⏔ (U+23D4) |
<glitch_pruned_255849> |
Removed merges (10)
Click to expand the full list of 10 removed merges
["bild", "titel"] → bildtitel
["آم", "باردا"] → آمباردا
["▁Accepted", "Loading"] → ▁AcceptedLoading
["▁Die", "ſe"] → ▁Dieſe
["▁coach", "Try"] → ▁coachTry
["▁ſ", "elb"] → ▁ſelb
["▁ſe", "lb"] → ▁ſelb
["▁ſel", "b"] → ▁ſelb
["▁パン", "チラ"] → ▁パンチラ
["▁メ", "ンテナ"] → ▁メンテナ
Files changed
tokenizer.json— BPE merge rules pruned, 24 vocab entries renamed
Algorithm Details
For full details on the glitch token collection algorithm and remediation techniques,
reach out to gkielian.
Related
This is part of a series of tokenizer fixes across all Gemma model repositories
(Gemma 1, 2, 3, 4, and MedGemma — 18 models total).