VIDRAFT_LAB's picture
πŸ”„ In a Training Loop

VIDRAFT_LAB

SeaWolf-AI

AI & ML interests

Contact: arxivgpt@gmail.com

Recent Activity

reacted to theirpost with πŸ‘ about 7 hours ago
πŸ–ΌοΈ POCKET-Image β€” the POCKET series goes visual: character-perfect text in any language, on-device A new model in VIDRAFT's POCKET family. POCKET put 35B-class models on phones and no-GPU PCs. POCKET-Image carries the same "big capability, small hardware" idea into image generation β€” and fixes the one thing nearly every image model gets wrong: text. Type "μ•ˆλ…•ν•˜μ„Έμš”" into a typical model and you get "μ•ˆγ…κΈ°." Hangul alone composes 11,172 syllable blocks; Arabic connects its letters; Thai stacks marks. Diffusion models draw scripts as shapes, so they smear. POCKET-Image renders every glyph exactly β€” ν•œκ΅­μ–΄ Β· δΈ­ζ–‡ Β· ζ—₯本θͺž Β· Ψ§Ω„ΨΉΨ±Ψ¨ΩŠΨ© (RTL) Β· ΰΉ„ΰΈ—ΰΈ’ Β· Latin and more β€” onto any scene you describe. What it is: β€’ 100% accurate text, any language β€” where global models produce gibberish β€’ Any background from a prompt β€” text is optional (empty β†’ a pure image) β€’ No GPU, no NPU β€” runs on plain CPU + RAM via the POCKET-Core engine β€’ Measured footprint: 8.6 GB (RTX 3050/4060) Β· 4.5 GB (offloaded, 6 GB cards) Β· 13.4 GB (MacBook, 16 GB+) β€’ Windows Β· macOS Β· Linux Β· fully local, no cloud Built on the open, commercial-friendly Z-Image (Apache-2.0) foundation. Honest note: the text is the guaranteed-correct part β€” the surrounding scene is ordinary generation, so a busy foreground can crowd the letters. We say so; clean backgrounds stay razor-sharp. 🎨 Studio β€” generate right here, any language: https://huggingface.co/spaces/FINAL-Bench/POCKET-Image-Studio 🧩 Model card: https://huggingface.co/FINAL-Bench/POCKET-Image-Zimage πŸ“š The POCKET collection: https://huggingface.co/collections/FINAL-Bench/pocket-models
posted an update about 7 hours ago
πŸ–ΌοΈ POCKET-Image β€” the POCKET series goes visual: character-perfect text in any language, on-device A new model in VIDRAFT's POCKET family. POCKET put 35B-class models on phones and no-GPU PCs. POCKET-Image carries the same "big capability, small hardware" idea into image generation β€” and fixes the one thing nearly every image model gets wrong: text. Type "μ•ˆλ…•ν•˜μ„Έμš”" into a typical model and you get "μ•ˆγ…κΈ°." Hangul alone composes 11,172 syllable blocks; Arabic connects its letters; Thai stacks marks. Diffusion models draw scripts as shapes, so they smear. POCKET-Image renders every glyph exactly β€” ν•œκ΅­μ–΄ Β· δΈ­ζ–‡ Β· ζ—₯本θͺž Β· Ψ§Ω„ΨΉΨ±Ψ¨ΩŠΨ© (RTL) Β· ΰΉ„ΰΈ—ΰΈ’ Β· Latin and more β€” onto any scene you describe. What it is: β€’ 100% accurate text, any language β€” where global models produce gibberish β€’ Any background from a prompt β€” text is optional (empty β†’ a pure image) β€’ No GPU, no NPU β€” runs on plain CPU + RAM via the POCKET-Core engine β€’ Measured footprint: 8.6 GB (RTX 3050/4060) Β· 4.5 GB (offloaded, 6 GB cards) Β· 13.4 GB (MacBook, 16 GB+) β€’ Windows Β· macOS Β· Linux Β· fully local, no cloud Built on the open, commercial-friendly Z-Image (Apache-2.0) foundation. Honest note: the text is the guaranteed-correct part β€” the surrounding scene is ordinary generation, so a busy foreground can crowd the letters. We say so; clean backgrounds stay razor-sharp. 🎨 Studio β€” generate right here, any language: https://huggingface.co/spaces/FINAL-Bench/POCKET-Image-Studio 🧩 Model card: https://huggingface.co/FINAL-Bench/POCKET-Image-Zimage πŸ“š The POCKET collection: https://huggingface.co/collections/FINAL-Bench/pocket-models
updated a collection about 7 hours ago
POCKET-MODELs
View all activity

Organizations

FINAL_Bench's profile picture Gemma Challenge's profile picture PPALLI's profile picture