|
Download README.md from Cnass/mtapi: direct link, hf CLI and curl.
- Browser
- Download file 26.1 kB
-
https://huggingface.co/spaces/Cnass/mtapi/resolve/main/README.md
- Command line
-
hf download hf://spaces/Cnass/mtapi/README.md
-
curl -L -o README.md https://huggingface.co/spaces/Cnass/mtapi/resolve/main/README.md
26.1 kB
| title: Hy-MT2 Translation API | |
| emoji: π | |
| colorFrom: blue | |
| colorTo: purple | |
| sdk: docker | |
| app_port: 7860 | |
| pinned: false | |
| license: apache-2.0 | |
| # Hy-MT2 Translation API | |
| A batching-optimized translation API for the | |
| [Hy-MT2](https://huggingface.co/collections/tencent/hy-mt2) GGUF models, | |
| built to run as a Hugging Face Space Docker app on a **GPU** instance | |
| (targets Nvidia T4, 16 GB VRAM). Defaults to **Hy-MT2-7B** β the 1.8B only | |
| existed to make CPU inference bearable and is one env var away if you want it | |
| back. Beyond raw translation it carries the machinery a 50k-string game | |
| localisation actually needs β a project glossary, a pinned register, and | |
| passthrough for control-code-only lines. See | |
| [Keeping 50k strings consistent](#keeping-50k-strings-consistent). | |
| Point it at a zipped RPG Maker MV/MZ `data` folder and it translates the game | |
| end to end: see [Translating an RPG Maker game](#translating-an-rpg-maker-game). | |
| ## How it works | |
| - **Inference engine**: `llama-server`, taken as-is from the official | |
| prebuilt image `ghcr.io/ggml-org/llama.cpp:server-cuda-b10398`. Nothing is | |
| compiled at build time β see [Build notes](#build-notes) for why the old | |
| from-source build turned out to be unnecessary. | |
| It runs with `--gpu-layers all --parallel 16 --cont-batching`, so every | |
| layer sits in VRAM and llama.cpp's continuous batching merges the in-flight | |
| requests into shared GPU work instead of serving them one at a time. This | |
| is the mechanism behind translating in batches rather than string by | |
| string, and unlike on CPU it genuinely scales with the slot count. | |
| - **Gateway**: a small FastAPI app (`app/main.py`) in front of it, which | |
| builds the Hy-MT2 instruction-format prompt (see `app/translation.py`) for | |
| each string and fans batch requests out concurrently to `llama-server`, | |
| bounded by `PARALLEL_SLOTS` so requests queue cleanly instead of | |
| overloading the model. | |
| - Chat formatting uses the GGUF's embedded Jinja chat template | |
| (`llama-server --jinja`), matching the model card's documented usage. | |
| ## Translating an RPG Maker game | |
| Upload the game's data folder as a zip and the whole extract β translate β | |
| repack cycle runs server-side. Open `/rpgm.html` on the Space for the UI, or | |
| drive it over the API: | |
| ```bash | |
| BASE=https://your-space.hf.space | |
| ID=$(curl -sF file=@data.zip $BASE/project/upload | jq -r .id) | |
| curl -s -X POST $BASE/project/$ID/start -H 'Content-Type: application/json' \ | |
| -d '{"target_lang":"tr"}' | |
| curl -s $BASE/project/$ID/status | jq '.percent, .files' | |
| curl -so translated.zip $BASE/project/$ID/download | |
| ``` | |
| Zip the `data` folder (MZ) or `www/data` (MV) β not the whole game. Both | |
| layouts are detected automatically. | |
| ### What gets translated | |
| Dialogue and interface text, and nothing else. RPG Maker keeps executable | |
| script, plugin bindings, asset filenames and engine identifiers in the same | |
| arrays as the lines an actor speaks, and translating one of those breaks the | |
| game rather than mistranslating it. | |
| | event code | content | translated | | |
| |---|---|---| | |
| | 401 / 405 | Show Text / Scrolling Text | **yes** | | |
| | 102 / 402 | Show Choices / When[choice] | **yes** | | |
| | 101 | message header, incl. the MZ speaker name | no | | |
| | 355 / 655 | Script (executable JS) | no | | |
| | 356 / 357 | Plugin Command | no | | |
| | 320 / 324 / 325 | Change Name / Nickname / Profile | no | | |
| Plus the interface text in `System.json` β the menu and options screens | |
| wrapped around that dialogue: | |
| | System.json key | content | translated | | |
| |---|---|---| | |
| | `terms.commands` | New Game, Continue, Save, Options, Attack, Buy/Sell | **yes** | | |
| | `terms.messages` | BGM/SE Volume, Always Dash, battle log templates | **yes** | | |
| | `terms.basic` / `terms.params` | Level, HP, MP, Attack, Defense β¦ | **yes** | | |
| | `elements`, `equipTypes`, `skillTypes`, `weaponTypes`, `armorTypes` | equip and status screen labels | **yes** | | |
| | `gameTitle` | the game's name | no β a proper name, like character names | | |
| | `currencyUnit` | usually a one-letter symbol (`G`) | no | | |
| | `switches`, `variables` | developer labels the player never sees | no | | |
| Null and empty entries are left as they are: RPG Maker pads these arrays to a | |
| fixed length and index 0 is normally blank, so writing text into one would put | |
| stray words in the menu. Database files (`Actors.json`, `Items.json`, β¦) and | |
| `plugins.js` are untouched. | |
| Two details make the output safe to ship: | |
| - **Message boxes stay whole.** A run of consecutive 401 commands is one box | |
| split across lines, not separate sentences, so the run is merged and | |
| translated as a single piece β then re-wrapped into exactly the original | |
| number of lines. Adding or removing entries in an event list would shift | |
| every index after it and break conditional branches, so the line count is | |
| an invariant. | |
| - **Repeats are translated once.** Units are keyed by source text across the | |
| whole project: in a real game an 800-slot map set collapses to ~150 unique | |
| strings, and a "Yes" that appears 300 times costs one generation. It also | |
| keeps a 102 choice and its 402 mirror automatically identical. | |
| Translations are written exactly as the model returns them. An earlier | |
| revision compared each translation's control codes against the source and held | |
| back any string where they differed, but Hy-MT2 carries `\C[2]`, `\N[1]` and | |
| the rest through on its own, so in practice that gate withheld good | |
| translations more often than it caught bad ones. The only response still | |
| refused is an empty one, which would blank a line in-game; strings that are | |
| *only* control codes never reach the model at all. | |
| Untranslated strings keep their source text, so the download is a playable | |
| game at any point, not just when the run finishes. `failed_units` in the | |
| status counts strings where the backend itself errored β those keep their | |
| source line too. | |
| ### Project endpoints | |
| | method | path | purpose | | |
| |---|---|---| | |
| | `POST` | `/project/upload` | multipart zip; unpacks, detects MV/MZ, indexes dialogue | | |
| | `POST` | `/project/{id}/start` | begin (or resume) translating; body `{"target_lang":"tr"}` | | |
| | `GET` | `/project/{id}/status` | overall %, per-file %, failure count | | |
| | `POST` | `/project/{id}/cancel` | stop after in-flight strings finish | | |
| | `GET` | `/project/{id}/download` | rebuilt zip, same folder layout | | |
| | `DELETE` | `/project/{id}` | remove immediately | | |
| | `GET` | `/project` | list projects and which one holds the GPU | | |
| One run at a time β there is a single GPU, so a second concurrent project | |
| would only make both finish later. Progress is written to disk continuously | |
| and a run interrupted by the Space sleeping is **resumed on startup**, which | |
| is why `PROJECT_DIR` defaults to the persistent volume (`/data/projects`) | |
| when one is mounted. Uploads are deleted after `PROJECT_RETENTION_HOURS` | |
| (default 24). | |
| ## Translation endpoints | |
| ### `GET /health` | |
| Backend status check. | |
| ### `GET /languages` | |
| Returns the map of supported language codes β names (the 33 languages | |
| Hy-MT2 documents support for). | |
| ### `POST /translate` β single string | |
| ```json | |
| { | |
| "text": "Where is the nearest potion shop?", | |
| "target_lang": "tr" | |
| } | |
| ``` | |
| ```json | |
| { "translation": "En yakΔ±n iksir dΓΌkkanΔ± nerede?", "elapsed_seconds": 4.82 } | |
| ``` | |
| ### `POST /translate/batch` β many strings in one call (recommended: 50-100 per call) | |
| ```json | |
| { | |
| "texts": ["Hello!", "Welcome!", "Potion", "Attack"], | |
| "target_lang": "tr" | |
| } | |
| ``` | |
| ```json | |
| { | |
| "translations": [ | |
| { "original": "Hello!", "translation": "Merhaba!" }, | |
| { "original": "Welcome!", "translation": "HoΕ geldin!" }, | |
| { "original": "Potion", "translation": "Δ°ksir" }, | |
| { "original": "Attack", "translation": "SaldΔ±rΔ±" } | |
| ], | |
| "count": 4, | |
| "elapsed_seconds": 0.9, | |
| "items_per_second": 4.44 | |
| } | |
| ``` | |
| For 50,000 game strings: split them into chunks of ~50-100 and call | |
| `/translate/batch` per chunk (either sequentially or a few chunks at a time) | |
| rather than one string per HTTP call. `MAX_BATCH_SIZE` (default 200) is a | |
| server-side safety cap on a single call. | |
| **Request fields common to both endpoints:** | |
| | field | type | notes | | |
| |---|---|---| | |
| | `target_lang` | string | ISO code (`tr`, `en`, `ja`, ...) or full name | | |
| | `source_lang` | string, optional | omit to let the model auto-detect | | |
| | `style` | string, optional | e.g. `"casual, playful"` β injected as a style instruction | | |
| | `glossary` | object, optional | `{"HP": "Can PuanΔ±"}` β merged on top of the project glossary file (see `GLOSSARY_FILE`) | | |
| | `context` | string, optional | background about the scene, injected with the model card's "Background Information" template. Weak in practice β see [Verified behaviour](#verified-behaviour) | | |
| | `preserve_placeholders` | bool, default `false` | adds an explicit "keep placeholders verbatim" clause. **Leave it off**: Hy-MT2 already preserves `\V[1]`, `%1`, `{var}` etc. on its own (measured β see [Verified behaviour](#verified-behaviour)), so the clause only costs prompt tokens | | |
| ### `POST /translate/document` β ordered lines of one scene | |
| Same fields as `/translate/batch`, plus `group_size`. Packs consecutive lines | |
| into one generation so the model sees the surrounding dialogue instead of | |
| each line alone. **Slower than `/translate/batch` (~25%) but more | |
| consistent** β use it for dialogue, not for bulk UI strings. Untranslatable | |
| lines are kept out of the group and copied through. | |
| ```json | |
| { "texts": ["...", "..."], "target_lang": "tr", "group_size": 10 } | |
| ``` | |
| The response adds `groups`, `group_size` and `regrouped_fallbacks`. A group | |
| whose markers come back wrong is silently re-run one line at a time, so lines | |
| can never end up shifted; a climbing `regrouped_fallbacks` means `group_size` | |
| is too large for the model in use. | |
| ### Auth | |
| Set the `API_KEY` env var on the Space to require an `X-API-Key` header on | |
| `/translate*`. Leave empty (default) to disable auth. | |
| ## Configuration (env vars) | |
| | var | default | meaning | | |
| |---|---|---| | |
| | `MODEL_REPO` | `tencent/Hy-MT2-7B-GGUF` | `tencent/Hy-MT2-1.8B-GGUF` (with a matching `MODEL_FILE`) is faster but noticeably weaker prose | | |
| | `MODEL_FILE` | `Hy-MT2-7B-Q4_K_M.gguf` | `HY-MT2-7B-Q6_K.gguf` / `HY-MT2-7B-Q8_0.gguf` also fit in 16 GB β see [VRAM budget](#vram-budget) | | |
| | `GPU_LAYERS` | `all` | passed to `--gpu-layers`. `all` keeps the whole model in VRAM; a number offloads only that many layers, `0` is CPU-only | | |
| | `DEFAULT_STYLE` | *(empty)* | style applied when a request sends none. **The single most effective quality setting** β see [Keeping 50k strings consistent](#keeping-50k-strings-consistent) | | |
| | `GLOSSARY_FILE` | *(empty)* | path to a JSON `{"term": "translation"}` file; only the terms occurring in a given string are attached to its prompt. See `glossary.example.json` | | |
| | `GROUP_SIZE` | `10` | strings packed into one generation by `/translate/document` | | |
| | `THREADS` | `4` | CPU threads; barely matters once every layer is on the GPU | | |
| | `PARALLEL_SLOTS` | `16` | concurrent generation slots (continuous batching) | | |
| | `CTX_SIZE` | `32768` | total context, split across `PARALLEL_SLOTS` slots (2048 each). Costs VRAM β see [VRAM budget](#vram-budget) | | |
| | `MAX_TOKENS` | `512` | max output tokens per translation | | |
| | `MAX_BATCH_SIZE` | `200` | max items accepted per `/translate/batch` call | | |
| | `PROJECT_DIR` | `/data/projects` if mounted, else `/app/projects` | where uploaded RPG Maker projects and their progress live | | |
| | `PROJECT_RETENTION_HOURS` | `24` | uploaded projects are deleted this long after their last update | | |
| | `MAX_UPLOAD_MB` | `200` | rejects uploads bigger than this | | |
| | `PROMPT_FORMAT` | inferred | `hy-mt2` or `rosetta` β see [Switching model family](#switching-model-family). Inferred from `MODEL_REPO`, so you rarely set it by hand | | |
| | `TEMPERATURE`/`TOP_P`/`TOP_K`/`REPEAT_PENALTY` | per family | `0.7`/`0.6`/`20`/`1.05` for Hy-MT2 (its model card's values), `0.7`/`0.95`/`64`/`1.0` for Rosetta (Gemma 3 defaults) | | |
| | `API_KEY` | *(empty)* | optional shared secret for `X-API-Key` | | |
| ## Deploying | |
| 1. On [huggingface.co/new-space](https://huggingface.co/new-space), choose | |
| **Docker** as the SDK, then push this folder's contents (or upload via | |
| the web UI / `huggingface_hub`). | |
| 2. In the Space's **Settings β Hardware**, select a GPU tier β **T4 small** | |
| (4 vCPU / 15 GB RAM / 16 GB VRAM) is what the defaults are sized for. This | |
| has to be set by hand; no file in this repo controls it. | |
| 3. **Enable persistent storage.** This matters more than anything else here: | |
| the GGUF is downloaded at *runtime*, not baked into the image, so without | |
| persistent storage every cold start re-downloads 4.6 GB **while the GPU | |
| meter is running**. With it, the download happens once. | |
| 4. **Set the sleep timer** (Settings β *Sleep time*). A GPU tier bills per | |
| hour for as long as the Space is awake, idle or not. | |
| 5. Recommended for a real localisation run: set `DEFAULT_STYLE`, and add a | |
| glossary JSON with `GLOSSARY_FILE` pointing at it. | |
| First boot is slow twice over: the model download, then a one-off PTX JIT | |
| compile (llama.cpp ships `sm_75` as PTX, which the driver compiles and caches | |
| on first load). `/health` answers `{"status":"starting"}` throughout. | |
| ### VRAM budget | |
| `hunyuan-dense` is 32 layers with 8 KV heads Γ 128 dims, i.e. **128 KB of KV | |
| cache per token**, so the context size is the knob that actually consumes | |
| VRAM: | |
| | `CTX_SIZE` | KV cache | + Q4_K_M (4.6 GB) | fits 16 GB? | | |
| |---|---|---|---| | |
| | 16384 | 2.0 GB | 6.6 GB | yes, lots spare | | |
| | **32768** (default) | **4.0 GB** | **8.6 GB** | **yes, comfortable** | | |
| | 65536 | 8.0 GB | 12.6 GB | yes, but tight | | |
| Swapping `MODEL_FILE` up a quant adds its size difference: Q6_K is 6.2 GB and | |
| Q8_0 is 8.0 GB, so both still fit alongside the default 32768 context. If | |
| llama-server dies at startup with a CUDA OOM, `CTX_SIZE` is the first thing to | |
| lower. | |
| ## Local testing | |
| ```bash | |
| docker build -t hy-mt2-api . | |
| docker run --gpus all -p 7860:7860 hy-mt2-api | |
| ``` | |
| Without `--gpus all` the container starts but finds no device. To try it on a | |
| machine with no GPU at all, add `-e GPU_LAYERS=0` β it will run on CPU, slowly. | |
| Then open `test-ui/index.html` in a browser, point it at | |
| `http://localhost:7860`, and use it to benchmark single vs. batch requests. | |
| ## Verified behaviour | |
| > **These numbers are from the CPU era** (4 cores, no GPU) and are kept as a | |
| > quality record, not a performance one. See | |
| > [Historical: CPU-era measurements](#historical-cpu-era-measurements). | |
| The stack below was actually run end to end against | |
| `Hy-MT2-7B-Q4_K_M.gguf` (real model, real llama-server built exactly as the | |
| `Dockerfile` builds it) on a **4-core / 15 GB** box β smaller than the 8 vCPU | |
| Space, so treat these as a floor, not a target. | |
| - The STQ1_0 cherry-pick works: the model loads and generates. This was the | |
| whole open question, and it is settled. | |
| - Translation quality spot-checks (ENβTR): `"Potion"` β `"Δ°ksir"`, | |
| `"Are you sure you want to quit the game?"` β | |
| `"Oyundan Γ§Δ±kmak istediΔinizden emin misiniz?"`. | |
| - **Placeholders survive untouched with no special instruction**: | |
| `"You have obtained \V[1] gold."` β `"\V[1] altΔ±n elde ettiniz."`, and | |
| `"Press %1 to open your inventory."` β | |
| `"Envanterinizi aΓ§mak iΓ§in %1'e basΔ±n."`. | |
| - `glossary` and `style` both take effect: with `{"gold": "altΔ±n", "HP": "Can | |
| PuanΔ±"}` the model returned `"50 altΔ±nΔ±nΔ±z ve dolu Can PuanΔ±nΔ±z var."`; | |
| with style `medieval, formal`, `"Hey, watch out!"` β `"Ey dostum, dikkatli | |
| ol!"`. | |
| - Throughput, 8 short game strings, 4 slots on 4 cores: | |
| **sequential 19.2 s vs. batch 12.2 s (1.58Γ)**. A 24-item batch of longer | |
| sentences ran at 0.34 items/s. Expect roughly double on 8 vCPU. | |
| ### 1.8B vs 7B, measured | |
| The default is now **1.8B**. Same 10-line dialogue scene, 4 cores: | |
| | model | time | throughput | | |
| |---|---|---| | |
| | Hy-MT2-7B Q4_K_M | 33.0 s | 0.30 items/s | | |
| | Hy-MT2-1.8B Q4_K_M | 10.9 s | 0.92 items/s | | |
| | Hy-MT2-1.8B Q4_K_M + passthrough + `DEFAULT_STYLE` | **7.9 s** | **1.27 items/s** | | |
| That is **~4x** end to end, which on 8 vCPU puts 50,000 sentence-length | |
| strings in the region of a few hours rather than a day. | |
| The cost is real: 1.8B makes mistakes 7B does not β it rendered *"Her round | |
| eyesβ¦"* with English *Her* read as Turkish *her* ("every"), and turned a | |
| vocative *"β¦, Michiru?"* into an object *"Michiru'yu"*. Set | |
| `MODEL_REPO=tencent/Hy-MT2-7B-GGUF` and `MODEL_FILE=Hy-MT2-7B-Q4_K_M.gguf` | |
| for anything where prose quality outranks throughput. | |
| ### A note on `preserve_placeholders` | |
| An earlier version of this API sent a long "never translate `\N[..]`, | |
| `{variable}`, `%s` β¦" instruction on every request, with that list spelled | |
| out. Hy-MT2 is a pure translation model, not a general chat model: it | |
| **translated the instruction** instead of following it, and short inputs were | |
| destroyed outright β `"Potion"` came back as the Turkish text of the | |
| instruction, with the actual word gone. Anything the model is meant to obey | |
| has to live inside the single instruction line of the model card's documented | |
| templates, never as its own paragraph in front of the source text. The flag | |
| now defaults to off and, when enabled, adds one short clause with no literal | |
| placeholder examples. | |
| ## Keeping 50k strings consistent | |
| The failure mode on a big run is not a mistranslated word, it is the same | |
| character sounding like two different people across a scene. Each string is | |
| translated in isolation, so nothing carries register or terminology from one | |
| line to the next. Four levers were built and measured on the 1.8B model; they | |
| are listed in the order they are worth reaching for. | |
| **1. `DEFAULT_STYLE` β the strong one.** Turkish forces a T/V choice on every | |
| sentence and the model picks per line, so one line says *geΓ§ kaldΔ±n* and the | |
| next *geΓ§ kaldΔ±nΔ±z*. Pinning the register fixes it globally: | |
| ``` | |
| DEFAULT_STYLE=casual spoken Turkish, informal second person singular (sen), never the formal siz form | |
| ``` | |
| | line | without | with | | |
| |---|---|---| | |
| | `You're thirty minutes late.` | Otuz dakika geΓ§ **kaldΔ±nΔ±z** | Otuz dakika geΓ§ **kaldΔ±n** | | |
| | `You rest at the innβ¦` | geri **kazanΔ±rsΔ±nΔ±z** | **dinlen** ve β¦ **tamamla** | | |
| **2. `GLOSSARY_FILE` β for names and terms.** A project-wide JSON dictionary; | |
| only terms occurring in a string are attached, so a 2000-entry glossary costs | |
| nothing on a line that uses none of it. Verified: `potion`β`iksir`, | |
| `HP`β`Can PuanΔ±`, `inn`β`han` applied without the request sending any | |
| glossary. Treat it as a strong hint, not a guarantee β in testing the model | |
| kept `bar` instead of the requested `meyhane` on one line. | |
| **3. Untranslatable-line passthrough β free, always on.** Lines that are only | |
| control codes, numbers or punctuation (`\SE[1]\n[1].`, `%1`, `---`) never | |
| reach the model. This is a correctness fix as much as a speed one: with a | |
| style instruction attached, `\SE[1]\n[1].` came back as | |
| `\SE[1]\n[1]. senin iΓ§in` β words invented out of nothing. Game files are | |
| full of such lines, and each one skipped is a whole generation saved. | |
| **4. `/translate/document` β better flow, ~25% slower.** Grouped translation | |
| did visibly fix register drift *before* `DEFAULT_STYLE` existed, and still | |
| smooths sentence flow across a scene. With levers 1-3 in place its remaining | |
| benefit is smaller, so it is opt-in. | |
| `context` was also implemented and measured, and is the weak one: supplying a | |
| scene description did not fix the formality drift and mostly changed word | |
| choice. It stays available for domain hints, but do not expect much. | |
| ### Grouping did not turn out to be a speed win | |
| Amortising the instruction prompt across a group sounds like it should be | |
| faster. It is not, on CPU: | |
| | mode | 40 lines, 4 slots | throughput | | |
| |---|---|---| | |
| | `/translate/batch` (one request per line) | **33.8 s** | **1.19 items/s** | | |
| | `/translate/document`, `group_size=5` | 43.3 s | 0.92 items/s | | |
| | `/translate/document`, `group_size=10` | 49.7 s | 0.80 items/s | | |
| CPU inference is dominated by *decode*, not prompt prefill, and grouping does | |
| not reduce the number of tokens generated β it just serialises ten | |
| translations into one stream and gives up the parallelism of the other slots. | |
| Continuous batching across slots beats packing more into a single request. | |
| ## Switching model family | |
| Two prompt families are supported. `PROMPT_FORMAT` selects one; leaving it | |
| unset infers it from `MODEL_REPO` (any repo whose name contains "rosetta" | |
| gets `rosetta`, everything else `hy-mt2`), so changing the model is usually a | |
| one-variable edit. `entrypoint.sh` applies the same rule, so the server flags | |
| and the prompt builder can't drift apart. | |
| - **`hy-mt2`** (default) β one user turn holding the instruction and the text, | |
| using the model card's templates. | |
| - **`rosetta`** β [YanoljaNEXT-Rosetta](https://huggingface.co/yanolja/YanoljaNEXT-Rosetta-4B-2511-GGUF), | |
| a Gemma 3 translation fine-tune. Directives go in a `system` turn | |
| (`Tone:`, `Glossary:`, β¦) and only the source text in `user`; its chat | |
| template renames those roles to `instruction` / `source`. Set both: | |
| ``` | |
| MODEL_REPO=yanolja/YanoljaNEXT-Rosetta-4B-2511-GGUF | |
| MODEL_FILE=Q5_K_M/YanoljaNEXT-Rosetta-4B-2511-bf16-q5_k_m.gguf | |
| ``` | |
| Rosetta ships a chat template that `llama-server` **cannot load**: it calls | |
| the Jinja `default` filter on an object, which llama.cpp's minja engine does | |
| not implement, and the server aborts at startup rather than degrade | |
| (`Unknown (built-in) filter 'default' for type Object`). `chat-template-rosetta.jinja` | |
| in this folder is a minja-compatible rewrite that renders identically for the | |
| one-system-plus-one-user requests this API sends; `entrypoint.sh` passes it | |
| via `--chat-template-file` whenever the format is `rosetta`. | |
| ### Measured: Rosetta was slower, so it is not the default | |
| Same 10 lines of RPG dialogue, same box, same settings: | |
| | Model | Time | Throughput | | |
| |---|---|---| | |
| | Hy-MT2-7B Q4_K_M | **33.0 s** | **0.30 items/s** | | |
| | Rosetta-4B Q5_K_M | 39.7 s | 0.25 items/s | | |
| Fewer parameters did not win here β Rosetta only publishes 5-bit and IQ | |
| quants, and Q5_K_M moves more memory per token than Q4_K_M. It also dropped a | |
| control code that Hy-MT2 kept (`\SE[1]I'm Michiru.` β `Ben Michiru'yum.`, | |
| losing `\SE[1]`) and repeated a line; `preserve_placeholders: true` fixes the | |
| dropped code but lengthens every prompt. Its own card notes it is tuned for | |
| structured JSON/YAML/XML and that "performance on unstructured text may | |
| vary", which is what dialogue is. Keep it in mind for structured content, not | |
| for speed. | |
| The same held for low-bit quantisation of Hy-MT2 itself: `Q3_K_M` measured | |
| **60.1 s** against Q4_K_M's 33.0 s on that dialogue, nearly 2Γ slower. | |
| llama.cpp's CPU kernels are far better optimised for Q4_K than for the | |
| K-quants around it, so dropping bits is not a reliable speed lever here. | |
| ## Build notes | |
| **The build no longer compiles anything.** It starts from the official | |
| prebuilt `ghcr.io/ggml-org/llama.cpp:server-cuda-b10398` and only adds Python | |
| and this repo's gateway on top, so a deploy is a pull plus a `pip install` β | |
| a couple of minutes, with nothing that can fail the way earlier builds did. | |
| That is possible because the from-source build was never actually needed. | |
| Every previous revision cloned llama.cpp and cherry-picked the `STQ1_0` kernel | |
| from [PR #22836](https://github.com/ggml-org/llama.cpp/pull/22836), on the | |
| strength of the model card's warning that "this gguf depends on our STQ | |
| kernel". Reading the GGUF tensor tables directly disproves that for the quants | |
| used here β both files are 354 tensors of ordinary `Q4_K` / `Q6_K` / `F32`: | |
| | file | tensor types | STQ1_0 tensors | | |
| |---|---|---| | |
| | `Hy-MT2-7B-Q4_K_M.gguf` | Q4_K Γ192, F32 Γ129, Q6_K Γ33 | **0** | | |
| | `Hy-MT2-1.8B-Q4_K_M.gguf` | Q4_K Γ192, F32 Γ129, Q6_K Γ33 | **0** | | |
| The warning applies to the separate 2-bit / 1.25-bit repos. Stock llama.cpp | |
| has supported the `hunyuan-dense` architecture for a long time, so the whole | |
| apparatus β the cherry-pick, the bounded `-j2` parallelism, the git identity, | |
| the static linking β existed to support a compile that did not have to happen. | |
| Everything it was working around is gone with it: | |
| - Builds that **hung at 24% for 5+ hours and died with no error**, because | |
| `-j$(nproc)` read the *host's* core count on a memory-capped build worker | |
| and thrashed it into swap. | |
| - A build that failed on `git cherry-pick` with *"Committer identity | |
| unknown"*, because a fresh container has no git config. | |
| - A build that failed on `COPY chat-template-rosetta.jinja` because that file | |
| had not made it into a manually-synced Space. (The template is still written | |
| inline in the `Dockerfile` for the same reason.) | |
| Layer caching still applies and now matters less: the expensive step is | |
| pulling the base image, and editing `app/` or `test-ui/` only rebuilds the | |
| final, instant layers because `requirements.txt` is copied and installed | |
| before the source. | |
| The image tag is pinned deliberately. `:server-cuda` would float, and a | |
| redeploy months from now could land on a llama.cpp that changed a CLI flag or | |
| a chat-template behaviour β both of which have already bitten this project | |
| once. Bump it consciously with `--build-arg LLAMA_CPP_IMAGE=...`. | |
| ### Historical: CPU-era measurements | |
| Everything measured below and in [Verified behaviour](#verified-behaviour) was | |
| recorded on 4 CPU cores with no GPU, and several conclusions are artefacts of | |
| CPU inference being decode-bound rather than facts about the models: | |
| - `/translate/document` grouping was **~25% slower** than per-string batching, | |
| because grouping trades away slot parallelism without reducing generated | |
| tokens. On a GPU, where batching genuinely scales, this may well invert. | |
| - `Q3_K_M` measured nearly **2Γ slower** than `Q4_K_M` β llama.cpp's CPU | |
| kernels are far better optimised for Q4_K. CUDA kernels have different | |
| characteristics. | |
| - `PARALLEL_SLOTS=8` saturated 8 cores; the GPU default is 16 and could | |
| probably go higher. | |
| Re-measure on the T4 before treating any of it as current. The | |
| `/translate/document` vs `/translate/batch` comparison in `test-ui` is the | |
| quickest way to redo it. | |