Download README.md from Cnass/mtapi: direct link, hf CLI and curl.
- Browser
- Download file 26.1 kB
-
https://huggingface.co/spaces/Cnass/mtapi/resolve/main/README.md
- Command line
-
hf download hf://spaces/Cnass/mtapi/README.md
-
curl -L -o README.md https://huggingface.co/spaces/Cnass/mtapi/resolve/main/README.md
title: Hy-MT2 Translation API
emoji: 🌐
colorFrom: blue
colorTo: purple
sdk: docker
app_port: 7860
pinned: false
license: apache-2.0
Hy-MT2 Translation API
A batching-optimized translation API for the Hy-MT2 GGUF models, built to run as a Hugging Face Space Docker app on a GPU instance (targets Nvidia T4, 16 GB VRAM). Defaults to Hy-MT2-7B — the 1.8B only existed to make CPU inference bearable and is one env var away if you want it back. Beyond raw translation it carries the machinery a 50k-string game localisation actually needs — a project glossary, a pinned register, and passthrough for control-code-only lines. See Keeping 50k strings consistent.
Point it at a zipped RPG Maker MV/MZ data folder and it translates the game
end to end: see Translating an RPG Maker game.
How it works
- Inference engine:
llama-server, taken as-is from the official prebuilt imageghcr.io/ggml-org/llama.cpp:server-cuda-b10398. Nothing is compiled at build time — see Build notes for why the old from-source build turned out to be unnecessary. It runs with--gpu-layers all --parallel 16 --cont-batching, so every layer sits in VRAM and llama.cpp's continuous batching merges the in-flight requests into shared GPU work instead of serving them one at a time. This is the mechanism behind translating in batches rather than string by string, and unlike on CPU it genuinely scales with the slot count. - Gateway: a small FastAPI app (
app/main.py) in front of it, which builds the Hy-MT2 instruction-format prompt (seeapp/translation.py) for each string and fans batch requests out concurrently tollama-server, bounded byPARALLEL_SLOTSso requests queue cleanly instead of overloading the model. - Chat formatting uses the GGUF's embedded Jinja chat template
(
llama-server --jinja), matching the model card's documented usage.
Translating an RPG Maker game
Upload the game's data folder as a zip and the whole extract → translate →
repack cycle runs server-side. Open /rpgm.html on the Space for the UI, or
drive it over the API:
BASE=https://your-space.hf.space
ID=$(curl -sF file=@data.zip $BASE/project/upload | jq -r .id)
curl -s -X POST $BASE/project/$ID/start -H 'Content-Type: application/json' \
-d '{"target_lang":"tr"}'
curl -s $BASE/project/$ID/status | jq '.percent, .files'
curl -so translated.zip $BASE/project/$ID/download
Zip the data folder (MZ) or www/data (MV) — not the whole game. Both
layouts are detected automatically.
What gets translated
Dialogue and interface text, and nothing else. RPG Maker keeps executable script, plugin bindings, asset filenames and engine identifiers in the same arrays as the lines an actor speaks, and translating one of those breaks the game rather than mistranslating it.
| event code | content | translated |
|---|---|---|
| 401 / 405 | Show Text / Scrolling Text | yes |
| 102 / 402 | Show Choices / When[choice] | yes |
| 101 | message header, incl. the MZ speaker name | no |
| 355 / 655 | Script (executable JS) | no |
| 356 / 357 | Plugin Command | no |
| 320 / 324 / 325 | Change Name / Nickname / Profile | no |
Plus the interface text in System.json — the menu and options screens
wrapped around that dialogue:
| System.json key | content | translated |
|---|---|---|
terms.commands |
New Game, Continue, Save, Options, Attack, Buy/Sell | yes |
terms.messages |
BGM/SE Volume, Always Dash, battle log templates | yes |
terms.basic / terms.params |
Level, HP, MP, Attack, Defense … | yes |
elements, equipTypes, skillTypes, weaponTypes, armorTypes |
equip and status screen labels | yes |
gameTitle |
the game's name | no — a proper name, like character names |
currencyUnit |
usually a one-letter symbol (G) |
no |
switches, variables |
developer labels the player never sees | no |
Null and empty entries are left as they are: RPG Maker pads these arrays to a
fixed length and index 0 is normally blank, so writing text into one would put
stray words in the menu. Database files (Actors.json, Items.json, …) and
plugins.js are untouched.
Two details make the output safe to ship:
- Message boxes stay whole. A run of consecutive 401 commands is one box split across lines, not separate sentences, so the run is merged and translated as a single piece — then re-wrapped into exactly the original number of lines. Adding or removing entries in an event list would shift every index after it and break conditional branches, so the line count is an invariant.
- Repeats are translated once. Units are keyed by source text across the whole project: in a real game an 800-slot map set collapses to ~150 unique strings, and a "Yes" that appears 300 times costs one generation. It also keeps a 102 choice and its 402 mirror automatically identical.
Translations are written exactly as the model returns them. An earlier
revision compared each translation's control codes against the source and held
back any string where they differed, but Hy-MT2 carries \C[2], \N[1] and
the rest through on its own, so in practice that gate withheld good
translations more often than it caught bad ones. The only response still
refused is an empty one, which would blank a line in-game; strings that are
only control codes never reach the model at all.
Untranslated strings keep their source text, so the download is a playable
game at any point, not just when the run finishes. failed_units in the
status counts strings where the backend itself errored — those keep their
source line too.
Project endpoints
| method | path | purpose |
|---|---|---|
POST |
/project/upload |
multipart zip; unpacks, detects MV/MZ, indexes dialogue |
POST |
/project/{id}/start |
begin (or resume) translating; body {"target_lang":"tr"} |
GET |
/project/{id}/status |
overall %, per-file %, failure count |
POST |
/project/{id}/cancel |
stop after in-flight strings finish |
GET |
/project/{id}/download |
rebuilt zip, same folder layout |
DELETE |
/project/{id} |
remove immediately |
GET |
/project |
list projects and which one holds the GPU |
One run at a time — there is a single GPU, so a second concurrent project
would only make both finish later. Progress is written to disk continuously
and a run interrupted by the Space sleeping is resumed on startup, which
is why PROJECT_DIR defaults to the persistent volume (/data/projects)
when one is mounted. Uploads are deleted after PROJECT_RETENTION_HOURS
(default 24).
Translation endpoints
GET /health
Backend status check.
GET /languages
Returns the map of supported language codes → names (the 33 languages Hy-MT2 documents support for).
POST /translate — single string
{
"text": "Where is the nearest potion shop?",
"target_lang": "tr"
}
{ "translation": "En yakın iksir dükkanı nerede?", "elapsed_seconds": 4.82 }
POST /translate/batch — many strings in one call (recommended: 50-100 per call)
{
"texts": ["Hello!", "Welcome!", "Potion", "Attack"],
"target_lang": "tr"
}
{
"translations": [
{ "original": "Hello!", "translation": "Merhaba!" },
{ "original": "Welcome!", "translation": "Hoş geldin!" },
{ "original": "Potion", "translation": "İksir" },
{ "original": "Attack", "translation": "Saldırı" }
],
"count": 4,
"elapsed_seconds": 0.9,
"items_per_second": 4.44
}
For 50,000 game strings: split them into chunks of ~50-100 and call
/translate/batch per chunk (either sequentially or a few chunks at a time)
rather than one string per HTTP call. MAX_BATCH_SIZE (default 200) is a
server-side safety cap on a single call.
Request fields common to both endpoints:
| field | type | notes |
|---|---|---|
target_lang |
string | ISO code (tr, en, ja, ...) or full name |
source_lang |
string, optional | omit to let the model auto-detect |
style |
string, optional | e.g. "casual, playful" — injected as a style instruction |
glossary |
object, optional | {"HP": "Can Puanı"} — merged on top of the project glossary file (see GLOSSARY_FILE) |
context |
string, optional | background about the scene, injected with the model card's "Background Information" template. Weak in practice — see Verified behaviour |
preserve_placeholders |
bool, default false |
adds an explicit "keep placeholders verbatim" clause. Leave it off: Hy-MT2 already preserves \V[1], %1, {var} etc. on its own (measured — see Verified behaviour), so the clause only costs prompt tokens |
POST /translate/document — ordered lines of one scene
Same fields as /translate/batch, plus group_size. Packs consecutive lines
into one generation so the model sees the surrounding dialogue instead of
each line alone. Slower than /translate/batch (~25%) but more
consistent — use it for dialogue, not for bulk UI strings. Untranslatable
lines are kept out of the group and copied through.
{ "texts": ["...", "..."], "target_lang": "tr", "group_size": 10 }
The response adds groups, group_size and regrouped_fallbacks. A group
whose markers come back wrong is silently re-run one line at a time, so lines
can never end up shifted; a climbing regrouped_fallbacks means group_size
is too large for the model in use.
Auth
Set the API_KEY env var on the Space to require an X-API-Key header on
/translate*. Leave empty (default) to disable auth.
Configuration (env vars)
| var | default | meaning |
|---|---|---|
MODEL_REPO |
tencent/Hy-MT2-7B-GGUF |
tencent/Hy-MT2-1.8B-GGUF (with a matching MODEL_FILE) is faster but noticeably weaker prose |
MODEL_FILE |
Hy-MT2-7B-Q4_K_M.gguf |
HY-MT2-7B-Q6_K.gguf / HY-MT2-7B-Q8_0.gguf also fit in 16 GB — see VRAM budget |
GPU_LAYERS |
all |
passed to --gpu-layers. all keeps the whole model in VRAM; a number offloads only that many layers, 0 is CPU-only |
DEFAULT_STYLE |
(empty) | style applied when a request sends none. The single most effective quality setting — see Keeping 50k strings consistent |
GLOSSARY_FILE |
(empty) | path to a JSON {"term": "translation"} file; only the terms occurring in a given string are attached to its prompt. See glossary.example.json |
GROUP_SIZE |
10 |
strings packed into one generation by /translate/document |
THREADS |
4 |
CPU threads; barely matters once every layer is on the GPU |
PARALLEL_SLOTS |
16 |
concurrent generation slots (continuous batching) |
CTX_SIZE |
32768 |
total context, split across PARALLEL_SLOTS slots (2048 each). Costs VRAM — see VRAM budget |
MAX_TOKENS |
512 |
max output tokens per translation |
MAX_BATCH_SIZE |
200 |
max items accepted per /translate/batch call |
PROJECT_DIR |
/data/projects if mounted, else /app/projects |
where uploaded RPG Maker projects and their progress live |
PROJECT_RETENTION_HOURS |
24 |
uploaded projects are deleted this long after their last update |
MAX_UPLOAD_MB |
200 |
rejects uploads bigger than this |
PROMPT_FORMAT |
inferred | hy-mt2 or rosetta — see Switching model family. Inferred from MODEL_REPO, so you rarely set it by hand |
TEMPERATURE/TOP_P/TOP_K/REPEAT_PENALTY |
per family | 0.7/0.6/20/1.05 for Hy-MT2 (its model card's values), 0.7/0.95/64/1.0 for Rosetta (Gemma 3 defaults) |
API_KEY |
(empty) | optional shared secret for X-API-Key |
Deploying
- On huggingface.co/new-space, choose
Docker as the SDK, then push this folder's contents (or upload via
the web UI /
huggingface_hub). - In the Space's Settings → Hardware, select a GPU tier — T4 small (4 vCPU / 15 GB RAM / 16 GB VRAM) is what the defaults are sized for. This has to be set by hand; no file in this repo controls it.
- Enable persistent storage. This matters more than anything else here: the GGUF is downloaded at runtime, not baked into the image, so without persistent storage every cold start re-downloads 4.6 GB while the GPU meter is running. With it, the download happens once.
- Set the sleep timer (Settings → Sleep time). A GPU tier bills per hour for as long as the Space is awake, idle or not.
- Recommended for a real localisation run: set
DEFAULT_STYLE, and add a glossary JSON withGLOSSARY_FILEpointing at it.
First boot is slow twice over: the model download, then a one-off PTX JIT
compile (llama.cpp ships sm_75 as PTX, which the driver compiles and caches
on first load). /health answers {"status":"starting"} throughout.
VRAM budget
hunyuan-dense is 32 layers with 8 KV heads × 128 dims, i.e. 128 KB of KV
cache per token, so the context size is the knob that actually consumes
VRAM:
CTX_SIZE |
KV cache | + Q4_K_M (4.6 GB) | fits 16 GB? |
|---|---|---|---|
| 16384 | 2.0 GB | 6.6 GB | yes, lots spare |
| 32768 (default) | 4.0 GB | 8.6 GB | yes, comfortable |
| 65536 | 8.0 GB | 12.6 GB | yes, but tight |
Swapping MODEL_FILE up a quant adds its size difference: Q6_K is 6.2 GB and
Q8_0 is 8.0 GB, so both still fit alongside the default 32768 context. If
llama-server dies at startup with a CUDA OOM, CTX_SIZE is the first thing to
lower.
Local testing
docker build -t hy-mt2-api .
docker run --gpus all -p 7860:7860 hy-mt2-api
Without --gpus all the container starts but finds no device. To try it on a
machine with no GPU at all, add -e GPU_LAYERS=0 — it will run on CPU, slowly.
Then open test-ui/index.html in a browser, point it at
http://localhost:7860, and use it to benchmark single vs. batch requests.
Verified behaviour
These numbers are from the CPU era (4 cores, no GPU) and are kept as a quality record, not a performance one. See Historical: CPU-era measurements.
The stack below was actually run end to end against
Hy-MT2-7B-Q4_K_M.gguf (real model, real llama-server built exactly as the
Dockerfile builds it) on a 4-core / 15 GB box — smaller than the 8 vCPU
Space, so treat these as a floor, not a target.
- The STQ1_0 cherry-pick works: the model loads and generates. This was the whole open question, and it is settled.
- Translation quality spot-checks (EN→TR):
"Potion"→"İksir","Are you sure you want to quit the game?"→"Oyundan çıkmak istediğinizden emin misiniz?". - Placeholders survive untouched with no special instruction:
"You have obtained \V[1] gold."→"\V[1] altın elde ettiniz.", and"Press %1 to open your inventory."→"Envanterinizi açmak için %1'e basın.". glossaryandstyleboth take effect: with{"gold": "altın", "HP": "Can Puanı"}the model returned"50 altınınız ve dolu Can Puanınız var."; with stylemedieval, formal,"Hey, watch out!"→"Ey dostum, dikkatli ol!".- Throughput, 8 short game strings, 4 slots on 4 cores: sequential 19.2 s vs. batch 12.2 s (1.58×). A 24-item batch of longer sentences ran at 0.34 items/s. Expect roughly double on 8 vCPU.
1.8B vs 7B, measured
The default is now 1.8B. Same 10-line dialogue scene, 4 cores:
| model | time | throughput |
|---|---|---|
| Hy-MT2-7B Q4_K_M | 33.0 s | 0.30 items/s |
| Hy-MT2-1.8B Q4_K_M | 10.9 s | 0.92 items/s |
Hy-MT2-1.8B Q4_K_M + passthrough + DEFAULT_STYLE |
7.9 s | 1.27 items/s |
That is ~4x end to end, which on 8 vCPU puts 50,000 sentence-length strings in the region of a few hours rather than a day.
The cost is real: 1.8B makes mistakes 7B does not — it rendered "Her round
eyes…" with English Her read as Turkish her ("every"), and turned a
vocative "…, Michiru?" into an object "Michiru'yu". Set
MODEL_REPO=tencent/Hy-MT2-7B-GGUF and MODEL_FILE=Hy-MT2-7B-Q4_K_M.gguf
for anything where prose quality outranks throughput.
A note on preserve_placeholders
An earlier version of this API sent a long "never translate \N[..],
{variable}, %s …" instruction on every request, with that list spelled
out. Hy-MT2 is a pure translation model, not a general chat model: it
translated the instruction instead of following it, and short inputs were
destroyed outright — "Potion" came back as the Turkish text of the
instruction, with the actual word gone. Anything the model is meant to obey
has to live inside the single instruction line of the model card's documented
templates, never as its own paragraph in front of the source text. The flag
now defaults to off and, when enabled, adds one short clause with no literal
placeholder examples.
Keeping 50k strings consistent
The failure mode on a big run is not a mistranslated word, it is the same character sounding like two different people across a scene. Each string is translated in isolation, so nothing carries register or terminology from one line to the next. Four levers were built and measured on the 1.8B model; they are listed in the order they are worth reaching for.
1. DEFAULT_STYLE — the strong one. Turkish forces a T/V choice on every
sentence and the model picks per line, so one line says geç kaldın and the
next geç kaldınız. Pinning the register fixes it globally:
DEFAULT_STYLE=casual spoken Turkish, informal second person singular (sen), never the formal siz form
| line | without | with |
|---|---|---|
You're thirty minutes late. |
Otuz dakika geç kaldınız | Otuz dakika geç kaldın |
You rest at the inn… |
geri kazanırsınız | dinlen ve … tamamla |
2. GLOSSARY_FILE — for names and terms. A project-wide JSON dictionary;
only terms occurring in a string are attached, so a 2000-entry glossary costs
nothing on a line that uses none of it. Verified: potion→iksir,
HP→Can Puanı, inn→han applied without the request sending any
glossary. Treat it as a strong hint, not a guarantee — in testing the model
kept bar instead of the requested meyhane on one line.
3. Untranslatable-line passthrough — free, always on. Lines that are only
control codes, numbers or punctuation (\SE[1]\n[1]., %1, ---) never
reach the model. This is a correctness fix as much as a speed one: with a
style instruction attached, \SE[1]\n[1]. came back as
\SE[1]\n[1]. senin için — words invented out of nothing. Game files are
full of such lines, and each one skipped is a whole generation saved.
4. /translate/document — better flow, ~25% slower. Grouped translation
did visibly fix register drift before DEFAULT_STYLE existed, and still
smooths sentence flow across a scene. With levers 1-3 in place its remaining
benefit is smaller, so it is opt-in.
context was also implemented and measured, and is the weak one: supplying a
scene description did not fix the formality drift and mostly changed word
choice. It stays available for domain hints, but do not expect much.
Grouping did not turn out to be a speed win
Amortising the instruction prompt across a group sounds like it should be faster. It is not, on CPU:
| mode | 40 lines, 4 slots | throughput |
|---|---|---|
/translate/batch (one request per line) |
33.8 s | 1.19 items/s |
/translate/document, group_size=5 |
43.3 s | 0.92 items/s |
/translate/document, group_size=10 |
49.7 s | 0.80 items/s |
CPU inference is dominated by decode, not prompt prefill, and grouping does not reduce the number of tokens generated — it just serialises ten translations into one stream and gives up the parallelism of the other slots. Continuous batching across slots beats packing more into a single request.
Switching model family
Two prompt families are supported. PROMPT_FORMAT selects one; leaving it
unset infers it from MODEL_REPO (any repo whose name contains "rosetta"
gets rosetta, everything else hy-mt2), so changing the model is usually a
one-variable edit. entrypoint.sh applies the same rule, so the server flags
and the prompt builder can't drift apart.
hy-mt2(default) — one user turn holding the instruction and the text, using the model card's templates.rosetta— YanoljaNEXT-Rosetta, a Gemma 3 translation fine-tune. Directives go in asystemturn (Tone:,Glossary:, …) and only the source text inuser; its chat template renames those roles toinstruction/source. Set both:MODEL_REPO=yanolja/YanoljaNEXT-Rosetta-4B-2511-GGUF MODEL_FILE=Q5_K_M/YanoljaNEXT-Rosetta-4B-2511-bf16-q5_k_m.gguf
Rosetta ships a chat template that llama-server cannot load: it calls
the Jinja default filter on an object, which llama.cpp's minja engine does
not implement, and the server aborts at startup rather than degrade
(Unknown (built-in) filter 'default' for type Object). chat-template-rosetta.jinja
in this folder is a minja-compatible rewrite that renders identically for the
one-system-plus-one-user requests this API sends; entrypoint.sh passes it
via --chat-template-file whenever the format is rosetta.
Measured: Rosetta was slower, so it is not the default
Same 10 lines of RPG dialogue, same box, same settings:
| Model | Time | Throughput |
|---|---|---|
| Hy-MT2-7B Q4_K_M | 33.0 s | 0.30 items/s |
| Rosetta-4B Q5_K_M | 39.7 s | 0.25 items/s |
Fewer parameters did not win here — Rosetta only publishes 5-bit and IQ
quants, and Q5_K_M moves more memory per token than Q4_K_M. It also dropped a
control code that Hy-MT2 kept (\SE[1]I'm Michiru. → Ben Michiru'yum.,
losing \SE[1]) and repeated a line; preserve_placeholders: true fixes the
dropped code but lengthens every prompt. Its own card notes it is tuned for
structured JSON/YAML/XML and that "performance on unstructured text may
vary", which is what dialogue is. Keep it in mind for structured content, not
for speed.
The same held for low-bit quantisation of Hy-MT2 itself: Q3_K_M measured
60.1 s against Q4_K_M's 33.0 s on that dialogue, nearly 2× slower.
llama.cpp's CPU kernels are far better optimised for Q4_K than for the
K-quants around it, so dropping bits is not a reliable speed lever here.
Build notes
The build no longer compiles anything. It starts from the official
prebuilt ghcr.io/ggml-org/llama.cpp:server-cuda-b10398 and only adds Python
and this repo's gateway on top, so a deploy is a pull plus a pip install —
a couple of minutes, with nothing that can fail the way earlier builds did.
That is possible because the from-source build was never actually needed.
Every previous revision cloned llama.cpp and cherry-picked the STQ1_0 kernel
from PR #22836, on the
strength of the model card's warning that "this gguf depends on our STQ
kernel". Reading the GGUF tensor tables directly disproves that for the quants
used here — both files are 354 tensors of ordinary Q4_K / Q6_K / F32:
| file | tensor types | STQ1_0 tensors |
|---|---|---|
Hy-MT2-7B-Q4_K_M.gguf |
Q4_K ×192, F32 ×129, Q6_K ×33 | 0 |
Hy-MT2-1.8B-Q4_K_M.gguf |
Q4_K ×192, F32 ×129, Q6_K ×33 | 0 |
The warning applies to the separate 2-bit / 1.25-bit repos. Stock llama.cpp
has supported the hunyuan-dense architecture for a long time, so the whole
apparatus — the cherry-pick, the bounded -j2 parallelism, the git identity,
the static linking — existed to support a compile that did not have to happen.
Everything it was working around is gone with it:
- Builds that hung at 24% for 5+ hours and died with no error, because
-j$(nproc)read the host's core count on a memory-capped build worker and thrashed it into swap. - A build that failed on
git cherry-pickwith "Committer identity unknown", because a fresh container has no git config. - A build that failed on
COPY chat-template-rosetta.jinjabecause that file had not made it into a manually-synced Space. (The template is still written inline in theDockerfilefor the same reason.)
Layer caching still applies and now matters less: the expensive step is
pulling the base image, and editing app/ or test-ui/ only rebuilds the
final, instant layers because requirements.txt is copied and installed
before the source.
The image tag is pinned deliberately. :server-cuda would float, and a
redeploy months from now could land on a llama.cpp that changed a CLI flag or
a chat-template behaviour — both of which have already bitten this project
once. Bump it consciously with --build-arg LLAMA_CPP_IMAGE=....
Historical: CPU-era measurements
Everything measured below and in Verified behaviour was recorded on 4 CPU cores with no GPU, and several conclusions are artefacts of CPU inference being decode-bound rather than facts about the models:
/translate/documentgrouping was ~25% slower than per-string batching, because grouping trades away slot parallelism without reducing generated tokens. On a GPU, where batching genuinely scales, this may well invert.Q3_K_Mmeasured nearly 2× slower thanQ4_K_M— llama.cpp's CPU kernels are far better optimised for Q4_K. CUDA kernels have different characteristics.PARALLEL_SLOTS=8saturated 8 cores; the GPU default is 16 and could probably go higher.
Re-measure on the T4 before treating any of it as current. The
/translate/document vs /translate/batch comparison in test-ui is the
quickest way to redo it.