Instructions to use google/gemma-4-E2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use google/gemma-4-E2B with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("google/gemma-4-E2B") model = AutoModelForMultimodalLM.from_pretrained("google/gemma-4-E2B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
fix: Remediate 217 glitch tokens in tokenizer.json
Glitch Token Remediation — Tokenizer Patch
Summary
This PR patches tokenizer.json to remediate 217 glitch tokens identified
by embedding vector analysis. Glitch tokens are vocabulary entries whose embeddings
have collapsed to near-identical representations, causing unpredictable model behavior
when they appear in input.
No model weights are modified — this is a tokenizer-only fix.
Technique: Hybrid (merge pruning + placeholder rename)
The patch uses two complementary strategies:
1. Merge pruning (216 tokens): Removes the BPE merge rules that
produce multi-character glitch tokens. The vocab entry stays (ID preserved) but
becomes unreachable — text that would have matched is now tokenized as trained
constituent subwords instead.
2. Placeholder rename (1 token): For single-character glitch tokens
(rare Unicode characters) that have no producing merge rule, the vocab string is renamed to<glitch_pruned_ID>. The original character now falls through to byte_fallback,
encoding as well-trained UTF-8 byte tokens.
Safety: Placeholder strings cannot be triggered by user input. BPE builds tokens
bottom-up from characters via merge rules only — since no merge rule produces these strings,
they are permanently unreachable. Gemma's split-digits tokenization policy provides an
additional layer of safety.
246 merges were removed for 216 merge-pruned tokens because some glitch tokens have more than one BPE merge path, and every path must be removed.
Before / after examples: google/gemma-4-E2B
Merge prune example: CYCLONEDB, token ID 192703
Prompt: Please repeat the following string exactly, with nothing else: "CYCLONEDB"
| Tokens for the string | Model output | |
|---|---|---|
| Before | ['CYCLONEDB'] → [192703] |
чого ❌ |
| After | ['CYCL', 'ONEDB'] → [137183, 183870] |
CYCLONEDB ✅ |
Placeholder rename example: 𡉺 (U+2127A), token ID 249788
Prompt: Please repeat the following string exactly, with nothing else: "𡉺"
| Tokens for the string | Model output | |
|---|---|---|
| Before | ['𡉺'] → [249788] |
𝚫𝚫𝚫𝚫𝚫𝚫𝚫𝚫𝚫𝚫𝚫𝚫 ❌ |
| After | ['<0xF0>', '<0xA1>', '<0x89>', '<0xBA>'] → [478, 399, 375, 424] |
𡉺 ✅ |
⚠️ Note for Agentic Systems
Many remediated glitch tokens are shared across Gemma model families (Gemma 1, 2, 3, 4).
In agentic pipelines, care should be taken that these token strings are not inadvertently
injected into prompts for unpatched models. We recommend applying this patch consistently
across all Gemma models used in a pipeline.
What is preserved
- ✅
len(tokenizer)— unchanged (262,144) - ✅ All token IDs — stable, no re-indexing
- ✅ Chat template — identical to original
- ✅
tokenizer_config.json— identical to original - ✅ All special tokens and
added_tokens— unchanged - ✅ Normal text tokenization — verified via benchmarks
- ✅ Zero model weight changes
Diff summary
| Original | Patched | Delta | |
|---|---|---|---|
| Vocab size | 262,144 | 262,144 | 0 |
| Merges | 514,906 | 514,660 | −246 |
| Vocab renamed | — | 1 | +1 |
Vocab renames (1 entry)
Click to expand all 1 renamed vocab entries
| ID | Original | Patched |
|---|---|---|
| 249788 | 𡉺 (U+2127A) |
<glitch_pruned_249788> |
Removed merges (246)
Click to expand the full list of 246 removed merges
["."+"|", ""+""] → ."+"|"+"
["AWS", "JavaScript"] → AWSJavaScript
["AllFilesIn", "MkvDir"] → AllFilesInMkvDir
["CYCL", "ONEDB"] → CYCLONEDB
["Case", "Missense"] → CaseMissense
["Case", "PTV"] → CasePTV
["Collin", "ToDO"] → CollinToDO
["Consec", "HtIdx"] → ConsecHtIdx
["Control", "PTV"] → ControlPTV
["Custom", "Glare"] → CustomGlare
["CustomGlare", "Def"] → CustomGlareDef
["DT", "MakeRectInCM"] → DTMakeRectInCM
["Days", "GE"] → DaysGE
["Denovo", "Mis"] → DenovoMis
["ELEASE", "STR"] → ELEASESTR
["ENEMY", "PLACE"] → ENEMYPLACE
["Fcm", "Php"] → FcmPhp
["Foldout", "GC"] → FoldoutGC
["GTBase", "Alert"] → GTBaseAlert
["Get", "ParaPkg"] → GetParaPkg
["GoObject", "AllRef"] → GoObjectAllRef
["Go", "PrintError"] → GoPrintError
["Go", "RawContext"] → GoRawContext
["Go", "SrvGroupIndex"] → GoSrvGroupIndex
["H", "IMQTTRVM"] → HIMQTTRVM
["IMQTTR", "VM"] → IMQTTRVM
["KAKKIA", "INEN"] → KAKKIAINEN
["LuaPush", "Int"] → LuaPushInt
["Lua", "ToJavaResult"] → LuaToJavaResult
["MDEw", "Ol"] → MDEwOl
["MakeRect", "InCM"] → MakeRectInCM
["Mkv", "Dir"] → MkvDir
["Module", "InitFlag"] → ModuleInitFlag
["MyShop", "name"] → MyShopname
["New", "ParaPkg"] → NewParaPkg
["Ordinate", "Tuple"] → OrdinateTuple
["PAF", "Struct"] → PAFStruct
["PQ", "Ary"] → PQAry
["P", "QAry"] → PQAry
["P", "Seqlist"] → PSeqlist
["Phase", "Hound"] → PhaseHound
["RefIn", "SrvGroup"] → RefInSrvGroup
["SAFER", "ELEASESTR"] → SAFERELEASESTR
["SRP", "Basic"] → SRPBasic
["SRP", "BinBuf"] → SRPBinBuf
["SRPInterface", "Item"] → SRPInterfaceItem
["SRP", "ParaPkg"] → SRPParaPkg
["SRPS", "XML"] → SRPSXML
["SRP", "SrvGroup"] → SRPSrvGroup
["SRPS", "rvGroup"] → SRPSrvGroup
["Same", "WordInterval"] → SameWordInterval
["Simulator", "Flood"] → SimulatorFlood
["SrvGroup", "Index"] → SrvGroupIndex
["StarBinBuf", "Body"] → StarBinBufBody
["StarObject", "Body"] → StarObjectBody
["StarParaPkg", "Body"] → StarParaPkgBody
["StarSXml", "Body"] → StarSXmlBody
["StarService", "Body"] → StarServiceBody
["StarSrvGroup", "Body"] → StarSrvGroupBody
["StructOf", "Object"] → StructOfObject
["Swift", "lyPlugin"] → SwiftlyPlugin
["Term", "ObjectDefer"] → TermObjectDefer
["Transc", "Hist"] → TranscHist
["U", "imcoords"] → Uimcoords
["VSLU", "ATYPE"] → VSLUATYPE
["WMN", "MDA"] → WMNMDA
["WeakTable", "Mutex"] → WeakTableMutex
["WithSize", "InCM"] → WithSizeInCM
["Ziel", "Panel"] → ZielPanel
["addKill", "Bonus"] → addKillBonus
["addKill", "Penalty"] → addKillPenalty
["adig", "anik"] → adiganik
["adigan", "ik"] → adiganik
["ak", "arantadhatu"] → akarantadhatu
["angolo", "Rad"] → angoloRad
["angolo", "Tocco"] → angoloTocco
["arant", "adhatu"] → arantadhatu
["arantad", "hatu"] → arantadhatu
["atthavid", "u"] → atthavidu
["attup", "adani"] → attupadani
["attu", "vasena"] → attuvasena
["bani", "pi"] → banipi
["ban", "ipi"] → banipi
["block", "idcoin"] → blockidcoin
["blusas", "Fem"] → blusasFem
["capture", "cpu"] → capturecpu
["ch", "ccgi"] → chccgi
["check", "katore"] → checkkatore
["co", "OrdinateTuple"] → coOrdinateTuple
["colour", "CodeDict"] → colourCodeDict
["country", "geocode"] → countrygeocode
["custom", "Glare"] → customGlare
["df", "sonic"] → dfsonic
["dfs", "onic"] → dfsonic
["done", "ProcessAvg"] → doneProcessAvg
["drawingCode", "hint"] → drawingCodehint
["dw", "RetJpegLen"] → dwRetJpegLen
["eco", "expr"] → ecoexpr
["edLeft", "Shape"] → edLeftShape
["edRight", "Shape"] → edRightShape
["faulse", "Ans"] → faulseAns
["foe", "Place"] → foePlace
["get", "starcore"] → getstarcore
["getstarcore", "data"] → getstarcoredata
["grafo", "Existe"] → grafoExiste
["him", "qttrvm"] → himqttrvm
["icoter", "zi"] → icoterzi
["ineed", "follower"] → ineedfollower
["inertia", "Seq"] → inertiaSeq
["isTrack", "Core"] → isTrackCore
["jols", "endev"] → jolsendev
["js", "bpmOb"] → jsbpmOb
["kit", "opssynth"] → kitopssynth
["last", "DamageTook"] → lastDamageTook
["m", "BlitzID"] → mBlitzID
["mark", "UpdateChoice"] → markUpdateChoice
["og", "Choice"] → ogChoice
["opss", "ynth"] → opssynth
["ops", "synth"] → opssynth
["pJ", "PEGBuf"] → pJPEGBuf
["paren", "macro"] → parenmacro
["partial", "owner"] → partialowner
["pmm", "Imp"] → pmmImp
["pos", "Tocco"] → posTocco
["qttr", "vm"] → qttrvm
["right", "squig"] → rightsquig
["sad", "urdu"] → sadurdu
["sal", "expr"] → salexpr
["selectTable", "X"] → selectTableX
["selectTable", "Y"] → selectTableY
["smo", "io"] → smoio
["sor", "finaly"] → sorfinaly
["squarePos", "Vecchio"] → squarePosVecchio
["start", "ZielPanel"] → startZielPanel
["tcp", "UniqueID"] → tcpUniqueID
["testGet", "Popup"] → testGetPopup
["time", "PlusEvents"] → timePlusEvents
["tochy", "odikwa"] → tochyodikwa
["total", "BlockFit"] → totalBlockFit
["trad", "uitEnCPP"] → traduitEnCPP
["uit", "EnCPP"] → uitEnCPP
["wired", "Elems"] → wiredElems
["ய்ய", "மணி"] → ய்யமணி
["వెట్", "స్కీ"] → వెట్స్కీ
["▁", "::::::::"] → ▁::::::::
["▁::", "::::::"] → ▁::::::::
["▁Afd", "Par"] → ▁AfdPar
["▁A", "fdPar"] → ▁AfdPar
["▁Archers", "Unit"] → ▁ArchersUnit
["▁C", "AdxRtList"] → ▁CAdxRtList
["▁CC", "BUNDLE"] → ▁CCBUNDLE
["▁Control", "Missense"] → ▁ControlMissense
["▁Control", "PTV"] → ▁ControlPTV
["▁", "ControlPTV"] → ▁ControlPTV
["▁DT", "MakeRect"] → ▁DTMakeRect
["▁FROM", "VS"] → ▁FROMVS
["▁Func", "ParamNum"] → ▁FuncParamNum
["▁", "FuncParamNum"] → ▁FuncParamNum
["▁GoObject", "To"] → ▁GoObjectTo
["▁Go", "RawContext"] → ▁GoRawContext
["▁", "GoRawContext"] → ▁GoRawContext
["▁Go", "SRP"] → ▁GoSRP
["▁", "GoSrvGroupIndex"] → ▁GoSrvGroupIndex
["▁Go", "SrvGroupIndex"] → ▁GoSrvGroupIndex
["▁H", "IMQTTRVM"] → ▁HIMQTTRVM
["▁", "HIMQTTRVM"] → ▁HIMQTTRVM
["▁LG", "AGEmoji"] → ▁LGAGEmoji
["▁Lua", "ToGoObject"] → ▁LuaToGoObject
["▁Lua", "ToJavaResult"] → ▁LuaToJavaResult
["▁", "LuaToJavaResult"] → ▁LuaToJavaResult
["▁ND", "IndexArray"] → ▁NDIndexArray
["▁New", "ParaPkg"] → ▁NewParaPkg
["▁", "NewParaPkg"] → ▁NewParaPkg
["▁RUTARE", "AL"] → ▁RUTAREAL
["▁RUTARE", "L"] → ▁RUTAREL
["▁RUTAREL", "ATIV"] → ▁RUTARELATIV
["▁Ref", "ToGoObject"] → ▁RefToGoObject
["▁SIINFE", "KLC"] → ▁SIINFEKLC
["▁SIINFEKL", "C"] → ▁SIINFEKLC
["▁SRP", "Go"] → ▁SRPGo
["▁SRPGo", "Get"] → ▁SRPGoGet
["▁SRPGo", "SetStr"] → ▁SRPGoSetStr
["▁Spring", "ObjectID"] → ▁SpringObjectID
["▁SrvGroup", "Class"] → ▁SrvGroupClass
["▁", "StarSXml"] → ▁StarSXml
["▁Star", "SXml"] → ▁StarSXml
["▁TP", "ASDW"] → ▁TPASDW
["▁", "TPASDW"] → ▁TPASDW
["▁", "TermObjectDefer"] → ▁TermObjectDefer
["▁Term", "ObjectDefer"] → ▁TermObjectDefer
["▁TestAvg", "Callback"] → ▁TestAvgCallback
["▁YY", "YY"] → ▁YYYY
["▁", "YYYY"] → ▁YYYY
["▁add", "ConfigureArg"] → ▁addConfigureArg
["▁add", "SBOM"] → ▁addSBOM
["▁ak", "ammak"] → ▁akammak
["▁app", "asidd"] → ▁appasidd
["▁atth", "udd"] → ▁atthudd
["▁bhuv", "adigane"] → ▁bhuvadigane
["▁browsing", "Stamp"] → ▁browsingStamp
["▁check", "HDROffsets"] → ▁checkHDROffsets
["▁coi", "Alarm"] → ▁coiAlarm
["▁cyt", "yle"] → ▁cytyle
["▁", "cytyle"] → ▁cytyle
["▁cy", "tyle"] → ▁cytyle
["▁dSample", "Height"] → ▁dSampleHeight
["▁dSample", "Width"] → ▁dSampleWidth
["▁diff", "formul"] → ▁diffformul
["▁ditt", "iyam"] → ▁dittiyam
["▁", "doneProcessAvg"] → ▁doneProcessAvg
["▁done", "ProcessAvg"] → ▁doneProcessAvg
["▁", "ecoexpr"] → ▁ecoexpr
["▁eco", "expr"] → ▁ecoexpr
["▁evam", "adisu"] → ▁evamadisu
["▁icc", "adini"] → ▁iccadini
["▁iccad", "ini"] → ▁iccadini
["▁icc", "api"] → ▁iccapi
["▁ic", "capi"] → ▁iccapi
["▁inner", "WallArray"] → ▁innerWallArray
["▁jaû", "nes"] → ▁jaûnes
["▁m", "SwisTrackCore"] → ▁mSwisTrackCore
["▁magick", "woods"] → ▁magickwoods
["▁mdl", "MeshVD"] → ▁mdlMeshVD
["▁min", "Goto"] → ▁minGoto
["▁neighbor", "Indexs"] → ▁neighborIndexs
["▁nibb", "acan"] → ▁nibbacan
["▁", "nohVP"] → ▁nohVP
["▁noh", "VP"] → ▁nohVP
["▁o", "LetterLocation"] → ▁oLetterLocation
["▁sc", "StudentVector"] → ▁scStudentVector
["▁sdx", "Concept"] → ▁sdxConcept
["▁student", "LVector"] → ▁studentLVector
["▁subTest", "Avg"] → ▁subTestAvg
["▁subTest", "HDR"] → ▁subTestHDR
["▁subTest", "Panorama"] → ▁subTestPanorama
["▁suddhak", "att"] → ▁suddhakatt
["▁total", "BlockUsed"] → ▁totalBlockUsed
["▁totalBlockUsed", "A"] → ▁totalBlockUsedA
["▁", "yyyy"] → ▁yyyy
["▁yy", "yy"] → ▁yyyy
["▁চিদা", "ভ"] → ▁চিদাভ
["▁শরনার্থি", "দের"] → ▁শরনার্থিদের
["▁শরনার্থ", "িদের"] → ▁শরনার্থিদের
["▁బ్లా", "వెట్స్కీ"] → ▁బ్లావెట్స్కీ
["▁⏮", "","] → ▁⏮",
["加入", "参数向量中"] → 加入参数向量中
Files changed
tokenizer.json— BPE merge rules pruned, 1 vocab entry renamed
Algorithm Details
For full details on the glitch token collection algorithm and remediation techniques,
reach out to Gregory Kielian.
Related
This is part of a series of tokenizer fixes across all Gemma model repositories
(Gemma 1, 2, 3, 4, and MedGemma — 18 models total).