A new model in VIDRAFT's POCKET family. POCKET put 35B-class models on phones and no-GPU PCs. POCKET-Image carries the same "big capability, small hardware" idea into image generation β and fixes the one thing nearly every image model gets wrong: text.
Type "μλ νμΈμ" into a typical model and you get "μγ κΈ°." Hangul alone composes 11,172 syllable blocks; Arabic connects its letters; Thai stacks marks. Diffusion models draw scripts as shapes, so they smear. POCKET-Image renders every glyph exactly β νκ΅μ΄ Β· δΈζ Β· ζ₯ζ¬θͺ Β· Ψ§ΩΨΉΨ±Ψ¨ΩΨ© (RTL) Β· ΰΉΰΈΰΈ’ Β· Latin and more β onto any scene you describe.
What it is:
β’ 100% accurate text, any language β where global models produce gibberish
β’ Any background from a prompt β text is optional (empty β a pure image)
β’ No GPU, no NPU β runs on plain CPU + RAM via the POCKET-Core engine
β’ Measured footprint: 8.6 GB (RTX 3050/4060) Β· 4.5 GB (offloaded, 6 GB cards) Β· 13.4 GB (MacBook, 16 GB+)
β’ Windows Β· macOS Β· Linux Β· fully local, no cloud
Built on the open, commercial-friendly Z-Image (Apache-2.0) foundation.
Honest note: the text is the guaranteed-correct part β the surrounding scene is ordinary generation, so a busy foreground can crowd the letters. We say so; clean backgrounds stay razor-sharp.
π¨ Studio β generate right here, any language:
FINAL-Bench/POCKET-Image-Studio
π§© Model card:
FINAL-Bench/POCKET-Image-Zimage
π The POCKET collection:
https://huggingface.co/collections/FINAL-Bench/pocket-models