mtapi / README.md
Cnass's picture
Upload README.md
4bf295f verified
|
Raw History Blame Contribute Delete
26.1 kB
metadata
title: Hy-MT2 Translation API
emoji: 🌐
colorFrom: blue
colorTo: purple
sdk: docker
app_port: 7860
pinned: false
license: apache-2.0

Hy-MT2 Translation API

A batching-optimized translation API for the Hy-MT2 GGUF models, built to run as a Hugging Face Space Docker app on a GPU instance (targets Nvidia T4, 16 GB VRAM). Defaults to Hy-MT2-7B — the 1.8B only existed to make CPU inference bearable and is one env var away if you want it back. Beyond raw translation it carries the machinery a 50k-string game localisation actually needs — a project glossary, a pinned register, and passthrough for control-code-only lines. See Keeping 50k strings consistent.

Point it at a zipped RPG Maker MV/MZ data folder and it translates the game end to end: see Translating an RPG Maker game.

How it works

  • Inference engine: llama-server, taken as-is from the official prebuilt image ghcr.io/ggml-org/llama.cpp:server-cuda-b10398. Nothing is compiled at build time — see Build notes for why the old from-source build turned out to be unnecessary. It runs with --gpu-layers all --parallel 16 --cont-batching, so every layer sits in VRAM and llama.cpp's continuous batching merges the in-flight requests into shared GPU work instead of serving them one at a time. This is the mechanism behind translating in batches rather than string by string, and unlike on CPU it genuinely scales with the slot count.
  • Gateway: a small FastAPI app (app/main.py) in front of it, which builds the Hy-MT2 instruction-format prompt (see app/translation.py) for each string and fans batch requests out concurrently to llama-server, bounded by PARALLEL_SLOTS so requests queue cleanly instead of overloading the model.
  • Chat formatting uses the GGUF's embedded Jinja chat template (llama-server --jinja), matching the model card's documented usage.

Translating an RPG Maker game

Upload the game's data folder as a zip and the whole extract → translate → repack cycle runs server-side. Open /rpgm.html on the Space for the UI, or drive it over the API:

BASE=https://your-space.hf.space
ID=$(curl -sF file=@data.zip $BASE/project/upload | jq -r .id)
curl -s -X POST $BASE/project/$ID/start -H 'Content-Type: application/json' \
     -d '{"target_lang":"tr"}'
curl -s $BASE/project/$ID/status | jq '.percent, .files'
curl -so translated.zip $BASE/project/$ID/download

Zip the data folder (MZ) or www/data (MV) — not the whole game. Both layouts are detected automatically.

What gets translated

Dialogue and interface text, and nothing else. RPG Maker keeps executable script, plugin bindings, asset filenames and engine identifiers in the same arrays as the lines an actor speaks, and translating one of those breaks the game rather than mistranslating it.

event code content translated
401 / 405 Show Text / Scrolling Text yes
102 / 402 Show Choices / When[choice] yes
101 message header, incl. the MZ speaker name no
355 / 655 Script (executable JS) no
356 / 357 Plugin Command no
320 / 324 / 325 Change Name / Nickname / Profile no

Plus the interface text in System.json — the menu and options screens wrapped around that dialogue:

System.json key content translated
terms.commands New Game, Continue, Save, Options, Attack, Buy/Sell yes
terms.messages BGM/SE Volume, Always Dash, battle log templates yes
terms.basic / terms.params Level, HP, MP, Attack, Defense … yes
elements, equipTypes, skillTypes, weaponTypes, armorTypes equip and status screen labels yes
gameTitle the game's name no — a proper name, like character names
currencyUnit usually a one-letter symbol (G) no
switches, variables developer labels the player never sees no

Null and empty entries are left as they are: RPG Maker pads these arrays to a fixed length and index 0 is normally blank, so writing text into one would put stray words in the menu. Database files (Actors.json, Items.json, …) and plugins.js are untouched.

Two details make the output safe to ship:

  • Message boxes stay whole. A run of consecutive 401 commands is one box split across lines, not separate sentences, so the run is merged and translated as a single piece — then re-wrapped into exactly the original number of lines. Adding or removing entries in an event list would shift every index after it and break conditional branches, so the line count is an invariant.
  • Repeats are translated once. Units are keyed by source text across the whole project: in a real game an 800-slot map set collapses to ~150 unique strings, and a "Yes" that appears 300 times costs one generation. It also keeps a 102 choice and its 402 mirror automatically identical.

Translations are written exactly as the model returns them. An earlier revision compared each translation's control codes against the source and held back any string where they differed, but Hy-MT2 carries \C[2], \N[1] and the rest through on its own, so in practice that gate withheld good translations more often than it caught bad ones. The only response still refused is an empty one, which would blank a line in-game; strings that are only control codes never reach the model at all.

Untranslated strings keep their source text, so the download is a playable game at any point, not just when the run finishes. failed_units in the status counts strings where the backend itself errored — those keep their source line too.

Project endpoints

method path purpose
POST /project/upload multipart zip; unpacks, detects MV/MZ, indexes dialogue
POST /project/{id}/start begin (or resume) translating; body {"target_lang":"tr"}
GET /project/{id}/status overall %, per-file %, failure count
POST /project/{id}/cancel stop after in-flight strings finish
GET /project/{id}/download rebuilt zip, same folder layout
DELETE /project/{id} remove immediately
GET /project list projects and which one holds the GPU

One run at a time — there is a single GPU, so a second concurrent project would only make both finish later. Progress is written to disk continuously and a run interrupted by the Space sleeping is resumed on startup, which is why PROJECT_DIR defaults to the persistent volume (/data/projects) when one is mounted. Uploads are deleted after PROJECT_RETENTION_HOURS (default 24).

Translation endpoints

GET /health

Backend status check.

GET /languages

Returns the map of supported language codes → names (the 33 languages Hy-MT2 documents support for).

POST /translate — single string

{
  "text": "Where is the nearest potion shop?",
  "target_lang": "tr"
}
{ "translation": "En yakın iksir dükkanı nerede?", "elapsed_seconds": 4.82 }

POST /translate/batch — many strings in one call (recommended: 50-100 per call)

{
  "texts": ["Hello!", "Welcome!", "Potion", "Attack"],
  "target_lang": "tr"
}
{
  "translations": [
    { "original": "Hello!", "translation": "Merhaba!" },
    { "original": "Welcome!", "translation": "Hoş geldin!" },
    { "original": "Potion", "translation": "İksir" },
    { "original": "Attack", "translation": "Saldırı" }
  ],
  "count": 4,
  "elapsed_seconds": 0.9,
  "items_per_second": 4.44
}

For 50,000 game strings: split them into chunks of ~50-100 and call /translate/batch per chunk (either sequentially or a few chunks at a time) rather than one string per HTTP call. MAX_BATCH_SIZE (default 200) is a server-side safety cap on a single call.

Request fields common to both endpoints:

field type notes
target_lang string ISO code (tr, en, ja, ...) or full name
source_lang string, optional omit to let the model auto-detect
style string, optional e.g. "casual, playful" — injected as a style instruction
glossary object, optional {"HP": "Can Puanı"} — merged on top of the project glossary file (see GLOSSARY_FILE)
context string, optional background about the scene, injected with the model card's "Background Information" template. Weak in practice — see Verified behaviour
preserve_placeholders bool, default false adds an explicit "keep placeholders verbatim" clause. Leave it off: Hy-MT2 already preserves \V[1], %1, {var} etc. on its own (measured — see Verified behaviour), so the clause only costs prompt tokens

POST /translate/document — ordered lines of one scene

Same fields as /translate/batch, plus group_size. Packs consecutive lines into one generation so the model sees the surrounding dialogue instead of each line alone. Slower than /translate/batch (~25%) but more consistent — use it for dialogue, not for bulk UI strings. Untranslatable lines are kept out of the group and copied through.

{ "texts": ["...", "..."], "target_lang": "tr", "group_size": 10 }

The response adds groups, group_size and regrouped_fallbacks. A group whose markers come back wrong is silently re-run one line at a time, so lines can never end up shifted; a climbing regrouped_fallbacks means group_size is too large for the model in use.

Auth

Set the API_KEY env var on the Space to require an X-API-Key header on /translate*. Leave empty (default) to disable auth.

Configuration (env vars)

var default meaning
MODEL_REPO tencent/Hy-MT2-7B-GGUF tencent/Hy-MT2-1.8B-GGUF (with a matching MODEL_FILE) is faster but noticeably weaker prose
MODEL_FILE Hy-MT2-7B-Q4_K_M.gguf HY-MT2-7B-Q6_K.gguf / HY-MT2-7B-Q8_0.gguf also fit in 16 GB — see VRAM budget
GPU_LAYERS all passed to --gpu-layers. all keeps the whole model in VRAM; a number offloads only that many layers, 0 is CPU-only
DEFAULT_STYLE (empty) style applied when a request sends none. The single most effective quality setting — see Keeping 50k strings consistent
GLOSSARY_FILE (empty) path to a JSON {"term": "translation"} file; only the terms occurring in a given string are attached to its prompt. See glossary.example.json
GROUP_SIZE 10 strings packed into one generation by /translate/document
THREADS 4 CPU threads; barely matters once every layer is on the GPU
PARALLEL_SLOTS 16 concurrent generation slots (continuous batching)
CTX_SIZE 32768 total context, split across PARALLEL_SLOTS slots (2048 each). Costs VRAM — see VRAM budget
MAX_TOKENS 512 max output tokens per translation
MAX_BATCH_SIZE 200 max items accepted per /translate/batch call
PROJECT_DIR /data/projects if mounted, else /app/projects where uploaded RPG Maker projects and their progress live
PROJECT_RETENTION_HOURS 24 uploaded projects are deleted this long after their last update
MAX_UPLOAD_MB 200 rejects uploads bigger than this
PROMPT_FORMAT inferred hy-mt2 or rosetta — see Switching model family. Inferred from MODEL_REPO, so you rarely set it by hand
TEMPERATURE/TOP_P/TOP_K/REPEAT_PENALTY per family 0.7/0.6/20/1.05 for Hy-MT2 (its model card's values), 0.7/0.95/64/1.0 for Rosetta (Gemma 3 defaults)
API_KEY (empty) optional shared secret for X-API-Key

Deploying

  1. On huggingface.co/new-space, choose Docker as the SDK, then push this folder's contents (or upload via the web UI / huggingface_hub).
  2. In the Space's Settings → Hardware, select a GPU tier — T4 small (4 vCPU / 15 GB RAM / 16 GB VRAM) is what the defaults are sized for. This has to be set by hand; no file in this repo controls it.
  3. Enable persistent storage. This matters more than anything else here: the GGUF is downloaded at runtime, not baked into the image, so without persistent storage every cold start re-downloads 4.6 GB while the GPU meter is running. With it, the download happens once.
  4. Set the sleep timer (Settings → Sleep time). A GPU tier bills per hour for as long as the Space is awake, idle or not.
  5. Recommended for a real localisation run: set DEFAULT_STYLE, and add a glossary JSON with GLOSSARY_FILE pointing at it.

First boot is slow twice over: the model download, then a one-off PTX JIT compile (llama.cpp ships sm_75 as PTX, which the driver compiles and caches on first load). /health answers {"status":"starting"} throughout.

VRAM budget

hunyuan-dense is 32 layers with 8 KV heads × 128 dims, i.e. 128 KB of KV cache per token, so the context size is the knob that actually consumes VRAM:

CTX_SIZE KV cache + Q4_K_M (4.6 GB) fits 16 GB?
16384 2.0 GB 6.6 GB yes, lots spare
32768 (default) 4.0 GB 8.6 GB yes, comfortable
65536 8.0 GB 12.6 GB yes, but tight

Swapping MODEL_FILE up a quant adds its size difference: Q6_K is 6.2 GB and Q8_0 is 8.0 GB, so both still fit alongside the default 32768 context. If llama-server dies at startup with a CUDA OOM, CTX_SIZE is the first thing to lower.

Local testing

docker build -t hy-mt2-api .
docker run --gpus all -p 7860:7860 hy-mt2-api

Without --gpus all the container starts but finds no device. To try it on a machine with no GPU at all, add -e GPU_LAYERS=0 — it will run on CPU, slowly.

Then open test-ui/index.html in a browser, point it at http://localhost:7860, and use it to benchmark single vs. batch requests.

Verified behaviour

These numbers are from the CPU era (4 cores, no GPU) and are kept as a quality record, not a performance one. See Historical: CPU-era measurements.

The stack below was actually run end to end against Hy-MT2-7B-Q4_K_M.gguf (real model, real llama-server built exactly as the Dockerfile builds it) on a 4-core / 15 GB box — smaller than the 8 vCPU Space, so treat these as a floor, not a target.

  • The STQ1_0 cherry-pick works: the model loads and generates. This was the whole open question, and it is settled.
  • Translation quality spot-checks (EN→TR): "Potion" → "İksir", "Are you sure you want to quit the game?" → "Oyundan çıkmak istediğinizden emin misiniz?".
  • Placeholders survive untouched with no special instruction: "You have obtained \V[1] gold." → "\V[1] altın elde ettiniz.", and "Press %1 to open your inventory." → "Envanterinizi açmak için %1'e basın.".
  • glossary and style both take effect: with {"gold": "altın", "HP": "Can Puanı"} the model returned "50 altınınız ve dolu Can Puanınız var."; with style medieval, formal, "Hey, watch out!" → "Ey dostum, dikkatli ol!".
  • Throughput, 8 short game strings, 4 slots on 4 cores: sequential 19.2 s vs. batch 12.2 s (1.58×). A 24-item batch of longer sentences ran at 0.34 items/s. Expect roughly double on 8 vCPU.

1.8B vs 7B, measured

The default is now 1.8B. Same 10-line dialogue scene, 4 cores:

model time throughput
Hy-MT2-7B Q4_K_M 33.0 s 0.30 items/s
Hy-MT2-1.8B Q4_K_M 10.9 s 0.92 items/s
Hy-MT2-1.8B Q4_K_M + passthrough + DEFAULT_STYLE 7.9 s 1.27 items/s

That is ~4x end to end, which on 8 vCPU puts 50,000 sentence-length strings in the region of a few hours rather than a day.

The cost is real: 1.8B makes mistakes 7B does not — it rendered "Her round eyes…" with English Her read as Turkish her ("every"), and turned a vocative "…, Michiru?" into an object "Michiru'yu". Set MODEL_REPO=tencent/Hy-MT2-7B-GGUF and MODEL_FILE=Hy-MT2-7B-Q4_K_M.gguf for anything where prose quality outranks throughput.

A note on preserve_placeholders

An earlier version of this API sent a long "never translate \N[..], {variable}, %s …" instruction on every request, with that list spelled out. Hy-MT2 is a pure translation model, not a general chat model: it translated the instruction instead of following it, and short inputs were destroyed outright — "Potion" came back as the Turkish text of the instruction, with the actual word gone. Anything the model is meant to obey has to live inside the single instruction line of the model card's documented templates, never as its own paragraph in front of the source text. The flag now defaults to off and, when enabled, adds one short clause with no literal placeholder examples.

Keeping 50k strings consistent

The failure mode on a big run is not a mistranslated word, it is the same character sounding like two different people across a scene. Each string is translated in isolation, so nothing carries register or terminology from one line to the next. Four levers were built and measured on the 1.8B model; they are listed in the order they are worth reaching for.

1. DEFAULT_STYLE — the strong one. Turkish forces a T/V choice on every sentence and the model picks per line, so one line says geç kaldın and the next geç kaldınız. Pinning the register fixes it globally:

DEFAULT_STYLE=casual spoken Turkish, informal second person singular (sen), never the formal siz form
line without with
You're thirty minutes late. Otuz dakika geç kaldınız Otuz dakika geç kaldın
You rest at the inn… geri kazanırsınız dinlen ve … tamamla

2. GLOSSARY_FILE — for names and terms. A project-wide JSON dictionary; only terms occurring in a string are attached, so a 2000-entry glossary costs nothing on a line that uses none of it. Verified: potion→iksir, HP→Can Puanı, inn→han applied without the request sending any glossary. Treat it as a strong hint, not a guarantee — in testing the model kept bar instead of the requested meyhane on one line.

3. Untranslatable-line passthrough — free, always on. Lines that are only control codes, numbers or punctuation (\SE[1]\n[1]., %1, ---) never reach the model. This is a correctness fix as much as a speed one: with a style instruction attached, \SE[1]\n[1]. came back as \SE[1]\n[1]. senin için — words invented out of nothing. Game files are full of such lines, and each one skipped is a whole generation saved.

4. /translate/document — better flow, ~25% slower. Grouped translation did visibly fix register drift before DEFAULT_STYLE existed, and still smooths sentence flow across a scene. With levers 1-3 in place its remaining benefit is smaller, so it is opt-in.

context was also implemented and measured, and is the weak one: supplying a scene description did not fix the formality drift and mostly changed word choice. It stays available for domain hints, but do not expect much.

Grouping did not turn out to be a speed win

Amortising the instruction prompt across a group sounds like it should be faster. It is not, on CPU:

mode 40 lines, 4 slots throughput
/translate/batch (one request per line) 33.8 s 1.19 items/s
/translate/document, group_size=5 43.3 s 0.92 items/s
/translate/document, group_size=10 49.7 s 0.80 items/s

CPU inference is dominated by decode, not prompt prefill, and grouping does not reduce the number of tokens generated — it just serialises ten translations into one stream and gives up the parallelism of the other slots. Continuous batching across slots beats packing more into a single request.

Switching model family

Two prompt families are supported. PROMPT_FORMAT selects one; leaving it unset infers it from MODEL_REPO (any repo whose name contains "rosetta" gets rosetta, everything else hy-mt2), so changing the model is usually a one-variable edit. entrypoint.sh applies the same rule, so the server flags and the prompt builder can't drift apart.

  • hy-mt2 (default) — one user turn holding the instruction and the text, using the model card's templates.
  • rosetta — YanoljaNEXT-Rosetta, a Gemma 3 translation fine-tune. Directives go in a system turn (Tone:, Glossary:, …) and only the source text in user; its chat template renames those roles to instruction / source. Set both:
    MODEL_REPO=yanolja/YanoljaNEXT-Rosetta-4B-2511-GGUF
    MODEL_FILE=Q5_K_M/YanoljaNEXT-Rosetta-4B-2511-bf16-q5_k_m.gguf
    

Rosetta ships a chat template that llama-server cannot load: it calls the Jinja default filter on an object, which llama.cpp's minja engine does not implement, and the server aborts at startup rather than degrade (Unknown (built-in) filter 'default' for type Object). chat-template-rosetta.jinja in this folder is a minja-compatible rewrite that renders identically for the one-system-plus-one-user requests this API sends; entrypoint.sh passes it via --chat-template-file whenever the format is rosetta.

Measured: Rosetta was slower, so it is not the default

Same 10 lines of RPG dialogue, same box, same settings:

Model Time Throughput
Hy-MT2-7B Q4_K_M 33.0 s 0.30 items/s
Rosetta-4B Q5_K_M 39.7 s 0.25 items/s

Fewer parameters did not win here — Rosetta only publishes 5-bit and IQ quants, and Q5_K_M moves more memory per token than Q4_K_M. It also dropped a control code that Hy-MT2 kept (\SE[1]I'm Michiru. → Ben Michiru'yum., losing \SE[1]) and repeated a line; preserve_placeholders: true fixes the dropped code but lengthens every prompt. Its own card notes it is tuned for structured JSON/YAML/XML and that "performance on unstructured text may vary", which is what dialogue is. Keep it in mind for structured content, not for speed.

The same held for low-bit quantisation of Hy-MT2 itself: Q3_K_M measured 60.1 s against Q4_K_M's 33.0 s on that dialogue, nearly 2× slower. llama.cpp's CPU kernels are far better optimised for Q4_K than for the K-quants around it, so dropping bits is not a reliable speed lever here.

Build notes

The build no longer compiles anything. It starts from the official prebuilt ghcr.io/ggml-org/llama.cpp:server-cuda-b10398 and only adds Python and this repo's gateway on top, so a deploy is a pull plus a pip install — a couple of minutes, with nothing that can fail the way earlier builds did.

That is possible because the from-source build was never actually needed. Every previous revision cloned llama.cpp and cherry-picked the STQ1_0 kernel from PR #22836, on the strength of the model card's warning that "this gguf depends on our STQ kernel". Reading the GGUF tensor tables directly disproves that for the quants used here — both files are 354 tensors of ordinary Q4_K / Q6_K / F32:

file tensor types STQ1_0 tensors
Hy-MT2-7B-Q4_K_M.gguf Q4_K ×192, F32 ×129, Q6_K ×33 0
Hy-MT2-1.8B-Q4_K_M.gguf Q4_K ×192, F32 ×129, Q6_K ×33 0

The warning applies to the separate 2-bit / 1.25-bit repos. Stock llama.cpp has supported the hunyuan-dense architecture for a long time, so the whole apparatus — the cherry-pick, the bounded -j2 parallelism, the git identity, the static linking — existed to support a compile that did not have to happen. Everything it was working around is gone with it:

  • Builds that hung at 24% for 5+ hours and died with no error, because -j$(nproc) read the host's core count on a memory-capped build worker and thrashed it into swap.
  • A build that failed on git cherry-pick with "Committer identity unknown", because a fresh container has no git config.
  • A build that failed on COPY chat-template-rosetta.jinja because that file had not made it into a manually-synced Space. (The template is still written inline in the Dockerfile for the same reason.)

Layer caching still applies and now matters less: the expensive step is pulling the base image, and editing app/ or test-ui/ only rebuilds the final, instant layers because requirements.txt is copied and installed before the source.

The image tag is pinned deliberately. :server-cuda would float, and a redeploy months from now could land on a llama.cpp that changed a CLI flag or a chat-template behaviour — both of which have already bitten this project once. Bump it consciously with --build-arg LLAMA_CPP_IMAGE=....

Historical: CPU-era measurements

Everything measured below and in Verified behaviour was recorded on 4 CPU cores with no GPU, and several conclusions are artefacts of CPU inference being decode-bound rather than facts about the models:

  • /translate/document grouping was ~25% slower than per-string batching, because grouping trades away slot parallelism without reducing generated tokens. On a GPU, where batching genuinely scales, this may well invert.
  • Q3_K_M measured nearly 2× slower than Q4_K_M — llama.cpp's CPU kernels are far better optimised for Q4_K. CUDA kernels have different characteristics.
  • PARALLEL_SLOTS=8 saturated 8 cores; the GPU default is 16 and could probably go higher.

Re-measure on the T4 before treating any of it as current. The /translate/document vs /translate/batch comparison in test-ui is the quickest way to redo it.