mtapi / README.md
Cnass's picture
Upload README.md
4bf295f verified
|
Raw History Blame Contribute Delete
26.1 kB
---
title: Hy-MT2 Translation API
emoji: 🌐
colorFrom: blue
colorTo: purple
sdk: docker
app_port: 7860
pinned: false
license: apache-2.0
---
# Hy-MT2 Translation API
A batching-optimized translation API for the
[Hy-MT2](https://huggingface.co/collections/tencent/hy-mt2) GGUF models,
built to run as a Hugging Face Space Docker app on a **GPU** instance
(targets Nvidia T4, 16 GB VRAM). Defaults to **Hy-MT2-7B** β€” the 1.8B only
existed to make CPU inference bearable and is one env var away if you want it
back. Beyond raw translation it carries the machinery a 50k-string game
localisation actually needs β€” a project glossary, a pinned register, and
passthrough for control-code-only lines. See
[Keeping 50k strings consistent](#keeping-50k-strings-consistent).
Point it at a zipped RPG Maker MV/MZ `data` folder and it translates the game
end to end: see [Translating an RPG Maker game](#translating-an-rpg-maker-game).
## How it works
- **Inference engine**: `llama-server`, taken as-is from the official
prebuilt image `ghcr.io/ggml-org/llama.cpp:server-cuda-b10398`. Nothing is
compiled at build time β€” see [Build notes](#build-notes) for why the old
from-source build turned out to be unnecessary.
It runs with `--gpu-layers all --parallel 16 --cont-batching`, so every
layer sits in VRAM and llama.cpp's continuous batching merges the in-flight
requests into shared GPU work instead of serving them one at a time. This
is the mechanism behind translating in batches rather than string by
string, and unlike on CPU it genuinely scales with the slot count.
- **Gateway**: a small FastAPI app (`app/main.py`) in front of it, which
builds the Hy-MT2 instruction-format prompt (see `app/translation.py`) for
each string and fans batch requests out concurrently to `llama-server`,
bounded by `PARALLEL_SLOTS` so requests queue cleanly instead of
overloading the model.
- Chat formatting uses the GGUF's embedded Jinja chat template
(`llama-server --jinja`), matching the model card's documented usage.
## Translating an RPG Maker game
Upload the game's data folder as a zip and the whole extract β†’ translate β†’
repack cycle runs server-side. Open `/rpgm.html` on the Space for the UI, or
drive it over the API:
```bash
BASE=https://your-space.hf.space
ID=$(curl -sF file=@data.zip $BASE/project/upload | jq -r .id)
curl -s -X POST $BASE/project/$ID/start -H 'Content-Type: application/json' \
-d '{"target_lang":"tr"}'
curl -s $BASE/project/$ID/status | jq '.percent, .files'
curl -so translated.zip $BASE/project/$ID/download
```
Zip the `data` folder (MZ) or `www/data` (MV) β€” not the whole game. Both
layouts are detected automatically.
### What gets translated
Dialogue and interface text, and nothing else. RPG Maker keeps executable
script, plugin bindings, asset filenames and engine identifiers in the same
arrays as the lines an actor speaks, and translating one of those breaks the
game rather than mistranslating it.
| event code | content | translated |
|---|---|---|
| 401 / 405 | Show Text / Scrolling Text | **yes** |
| 102 / 402 | Show Choices / When[choice] | **yes** |
| 101 | message header, incl. the MZ speaker name | no |
| 355 / 655 | Script (executable JS) | no |
| 356 / 357 | Plugin Command | no |
| 320 / 324 / 325 | Change Name / Nickname / Profile | no |
Plus the interface text in `System.json` β€” the menu and options screens
wrapped around that dialogue:
| System.json key | content | translated |
|---|---|---|
| `terms.commands` | New Game, Continue, Save, Options, Attack, Buy/Sell | **yes** |
| `terms.messages` | BGM/SE Volume, Always Dash, battle log templates | **yes** |
| `terms.basic` / `terms.params` | Level, HP, MP, Attack, Defense … | **yes** |
| `elements`, `equipTypes`, `skillTypes`, `weaponTypes`, `armorTypes` | equip and status screen labels | **yes** |
| `gameTitle` | the game's name | no β€” a proper name, like character names |
| `currencyUnit` | usually a one-letter symbol (`G`) | no |
| `switches`, `variables` | developer labels the player never sees | no |
Null and empty entries are left as they are: RPG Maker pads these arrays to a
fixed length and index 0 is normally blank, so writing text into one would put
stray words in the menu. Database files (`Actors.json`, `Items.json`, …) and
`plugins.js` are untouched.
Two details make the output safe to ship:
- **Message boxes stay whole.** A run of consecutive 401 commands is one box
split across lines, not separate sentences, so the run is merged and
translated as a single piece β€” then re-wrapped into exactly the original
number of lines. Adding or removing entries in an event list would shift
every index after it and break conditional branches, so the line count is
an invariant.
- **Repeats are translated once.** Units are keyed by source text across the
whole project: in a real game an 800-slot map set collapses to ~150 unique
strings, and a "Yes" that appears 300 times costs one generation. It also
keeps a 102 choice and its 402 mirror automatically identical.
Translations are written exactly as the model returns them. An earlier
revision compared each translation's control codes against the source and held
back any string where they differed, but Hy-MT2 carries `\C[2]`, `\N[1]` and
the rest through on its own, so in practice that gate withheld good
translations more often than it caught bad ones. The only response still
refused is an empty one, which would blank a line in-game; strings that are
*only* control codes never reach the model at all.
Untranslated strings keep their source text, so the download is a playable
game at any point, not just when the run finishes. `failed_units` in the
status counts strings where the backend itself errored β€” those keep their
source line too.
### Project endpoints
| method | path | purpose |
|---|---|---|
| `POST` | `/project/upload` | multipart zip; unpacks, detects MV/MZ, indexes dialogue |
| `POST` | `/project/{id}/start` | begin (or resume) translating; body `{"target_lang":"tr"}` |
| `GET` | `/project/{id}/status` | overall %, per-file %, failure count |
| `POST` | `/project/{id}/cancel` | stop after in-flight strings finish |
| `GET` | `/project/{id}/download` | rebuilt zip, same folder layout |
| `DELETE` | `/project/{id}` | remove immediately |
| `GET` | `/project` | list projects and which one holds the GPU |
One run at a time β€” there is a single GPU, so a second concurrent project
would only make both finish later. Progress is written to disk continuously
and a run interrupted by the Space sleeping is **resumed on startup**, which
is why `PROJECT_DIR` defaults to the persistent volume (`/data/projects`)
when one is mounted. Uploads are deleted after `PROJECT_RETENTION_HOURS`
(default 24).
## Translation endpoints
### `GET /health`
Backend status check.
### `GET /languages`
Returns the map of supported language codes β†’ names (the 33 languages
Hy-MT2 documents support for).
### `POST /translate` β€” single string
```json
{
"text": "Where is the nearest potion shop?",
"target_lang": "tr"
}
```
```json
{ "translation": "En yakΔ±n iksir dΓΌkkanΔ± nerede?", "elapsed_seconds": 4.82 }
```
### `POST /translate/batch` β€” many strings in one call (recommended: 50-100 per call)
```json
{
"texts": ["Hello!", "Welcome!", "Potion", "Attack"],
"target_lang": "tr"
}
```
```json
{
"translations": [
{ "original": "Hello!", "translation": "Merhaba!" },
{ "original": "Welcome!", "translation": "Hoş geldin!" },
{ "original": "Potion", "translation": "Δ°ksir" },
{ "original": "Attack", "translation": "SaldΔ±rΔ±" }
],
"count": 4,
"elapsed_seconds": 0.9,
"items_per_second": 4.44
}
```
For 50,000 game strings: split them into chunks of ~50-100 and call
`/translate/batch` per chunk (either sequentially or a few chunks at a time)
rather than one string per HTTP call. `MAX_BATCH_SIZE` (default 200) is a
server-side safety cap on a single call.
**Request fields common to both endpoints:**
| field | type | notes |
|---|---|---|
| `target_lang` | string | ISO code (`tr`, `en`, `ja`, ...) or full name |
| `source_lang` | string, optional | omit to let the model auto-detect |
| `style` | string, optional | e.g. `"casual, playful"` β€” injected as a style instruction |
| `glossary` | object, optional | `{"HP": "Can PuanΔ±"}` β€” merged on top of the project glossary file (see `GLOSSARY_FILE`) |
| `context` | string, optional | background about the scene, injected with the model card's "Background Information" template. Weak in practice β€” see [Verified behaviour](#verified-behaviour) |
| `preserve_placeholders` | bool, default `false` | adds an explicit "keep placeholders verbatim" clause. **Leave it off**: Hy-MT2 already preserves `\V[1]`, `%1`, `{var}` etc. on its own (measured β€” see [Verified behaviour](#verified-behaviour)), so the clause only costs prompt tokens |
### `POST /translate/document` β€” ordered lines of one scene
Same fields as `/translate/batch`, plus `group_size`. Packs consecutive lines
into one generation so the model sees the surrounding dialogue instead of
each line alone. **Slower than `/translate/batch` (~25%) but more
consistent** β€” use it for dialogue, not for bulk UI strings. Untranslatable
lines are kept out of the group and copied through.
```json
{ "texts": ["...", "..."], "target_lang": "tr", "group_size": 10 }
```
The response adds `groups`, `group_size` and `regrouped_fallbacks`. A group
whose markers come back wrong is silently re-run one line at a time, so lines
can never end up shifted; a climbing `regrouped_fallbacks` means `group_size`
is too large for the model in use.
### Auth
Set the `API_KEY` env var on the Space to require an `X-API-Key` header on
`/translate*`. Leave empty (default) to disable auth.
## Configuration (env vars)
| var | default | meaning |
|---|---|---|
| `MODEL_REPO` | `tencent/Hy-MT2-7B-GGUF` | `tencent/Hy-MT2-1.8B-GGUF` (with a matching `MODEL_FILE`) is faster but noticeably weaker prose |
| `MODEL_FILE` | `Hy-MT2-7B-Q4_K_M.gguf` | `HY-MT2-7B-Q6_K.gguf` / `HY-MT2-7B-Q8_0.gguf` also fit in 16 GB β€” see [VRAM budget](#vram-budget) |
| `GPU_LAYERS` | `all` | passed to `--gpu-layers`. `all` keeps the whole model in VRAM; a number offloads only that many layers, `0` is CPU-only |
| `DEFAULT_STYLE` | *(empty)* | style applied when a request sends none. **The single most effective quality setting** β€” see [Keeping 50k strings consistent](#keeping-50k-strings-consistent) |
| `GLOSSARY_FILE` | *(empty)* | path to a JSON `{"term": "translation"}` file; only the terms occurring in a given string are attached to its prompt. See `glossary.example.json` |
| `GROUP_SIZE` | `10` | strings packed into one generation by `/translate/document` |
| `THREADS` | `4` | CPU threads; barely matters once every layer is on the GPU |
| `PARALLEL_SLOTS` | `16` | concurrent generation slots (continuous batching) |
| `CTX_SIZE` | `32768` | total context, split across `PARALLEL_SLOTS` slots (2048 each). Costs VRAM β€” see [VRAM budget](#vram-budget) |
| `MAX_TOKENS` | `512` | max output tokens per translation |
| `MAX_BATCH_SIZE` | `200` | max items accepted per `/translate/batch` call |
| `PROJECT_DIR` | `/data/projects` if mounted, else `/app/projects` | where uploaded RPG Maker projects and their progress live |
| `PROJECT_RETENTION_HOURS` | `24` | uploaded projects are deleted this long after their last update |
| `MAX_UPLOAD_MB` | `200` | rejects uploads bigger than this |
| `PROMPT_FORMAT` | inferred | `hy-mt2` or `rosetta` β€” see [Switching model family](#switching-model-family). Inferred from `MODEL_REPO`, so you rarely set it by hand |
| `TEMPERATURE`/`TOP_P`/`TOP_K`/`REPEAT_PENALTY` | per family | `0.7`/`0.6`/`20`/`1.05` for Hy-MT2 (its model card's values), `0.7`/`0.95`/`64`/`1.0` for Rosetta (Gemma 3 defaults) |
| `API_KEY` | *(empty)* | optional shared secret for `X-API-Key` |
## Deploying
1. On [huggingface.co/new-space](https://huggingface.co/new-space), choose
**Docker** as the SDK, then push this folder's contents (or upload via
the web UI / `huggingface_hub`).
2. In the Space's **Settings β†’ Hardware**, select a GPU tier β€” **T4 small**
(4 vCPU / 15 GB RAM / 16 GB VRAM) is what the defaults are sized for. This
has to be set by hand; no file in this repo controls it.
3. **Enable persistent storage.** This matters more than anything else here:
the GGUF is downloaded at *runtime*, not baked into the image, so without
persistent storage every cold start re-downloads 4.6 GB **while the GPU
meter is running**. With it, the download happens once.
4. **Set the sleep timer** (Settings β†’ *Sleep time*). A GPU tier bills per
hour for as long as the Space is awake, idle or not.
5. Recommended for a real localisation run: set `DEFAULT_STYLE`, and add a
glossary JSON with `GLOSSARY_FILE` pointing at it.
First boot is slow twice over: the model download, then a one-off PTX JIT
compile (llama.cpp ships `sm_75` as PTX, which the driver compiles and caches
on first load). `/health` answers `{"status":"starting"}` throughout.
### VRAM budget
`hunyuan-dense` is 32 layers with 8 KV heads Γ— 128 dims, i.e. **128 KB of KV
cache per token**, so the context size is the knob that actually consumes
VRAM:
| `CTX_SIZE` | KV cache | + Q4_K_M (4.6 GB) | fits 16 GB? |
|---|---|---|---|
| 16384 | 2.0 GB | 6.6 GB | yes, lots spare |
| **32768** (default) | **4.0 GB** | **8.6 GB** | **yes, comfortable** |
| 65536 | 8.0 GB | 12.6 GB | yes, but tight |
Swapping `MODEL_FILE` up a quant adds its size difference: Q6_K is 6.2 GB and
Q8_0 is 8.0 GB, so both still fit alongside the default 32768 context. If
llama-server dies at startup with a CUDA OOM, `CTX_SIZE` is the first thing to
lower.
## Local testing
```bash
docker build -t hy-mt2-api .
docker run --gpus all -p 7860:7860 hy-mt2-api
```
Without `--gpus all` the container starts but finds no device. To try it on a
machine with no GPU at all, add `-e GPU_LAYERS=0` β€” it will run on CPU, slowly.
Then open `test-ui/index.html` in a browser, point it at
`http://localhost:7860`, and use it to benchmark single vs. batch requests.
## Verified behaviour
> **These numbers are from the CPU era** (4 cores, no GPU) and are kept as a
> quality record, not a performance one. See
> [Historical: CPU-era measurements](#historical-cpu-era-measurements).
The stack below was actually run end to end against
`Hy-MT2-7B-Q4_K_M.gguf` (real model, real llama-server built exactly as the
`Dockerfile` builds it) on a **4-core / 15 GB** box β€” smaller than the 8 vCPU
Space, so treat these as a floor, not a target.
- The STQ1_0 cherry-pick works: the model loads and generates. This was the
whole open question, and it is settled.
- Translation quality spot-checks (EN→TR): `"Potion"` → `"İksir"`,
`"Are you sure you want to quit the game?"` β†’
`"Oyundan çıkmak istediğinizden emin misiniz?"`.
- **Placeholders survive untouched with no special instruction**:
`"You have obtained \V[1] gold."` β†’ `"\V[1] altΔ±n elde ettiniz."`, and
`"Press %1 to open your inventory."` β†’
`"Envanterinizi aΓ§mak iΓ§in %1'e basΔ±n."`.
- `glossary` and `style` both take effect: with `{"gold": "altΔ±n", "HP": "Can
PuanΔ±"}` the model returned `"50 altΔ±nΔ±nΔ±z ve dolu Can PuanΔ±nΔ±z var."`;
with style `medieval, formal`, `"Hey, watch out!"` β†’ `"Ey dostum, dikkatli
ol!"`.
- Throughput, 8 short game strings, 4 slots on 4 cores:
**sequential 19.2 s vs. batch 12.2 s (1.58Γ—)**. A 24-item batch of longer
sentences ran at 0.34 items/s. Expect roughly double on 8 vCPU.
### 1.8B vs 7B, measured
The default is now **1.8B**. Same 10-line dialogue scene, 4 cores:
| model | time | throughput |
|---|---|---|
| Hy-MT2-7B Q4_K_M | 33.0 s | 0.30 items/s |
| Hy-MT2-1.8B Q4_K_M | 10.9 s | 0.92 items/s |
| Hy-MT2-1.8B Q4_K_M + passthrough + `DEFAULT_STYLE` | **7.9 s** | **1.27 items/s** |
That is **~4x** end to end, which on 8 vCPU puts 50,000 sentence-length
strings in the region of a few hours rather than a day.
The cost is real: 1.8B makes mistakes 7B does not β€” it rendered *"Her round
eyes…"* with English *Her* read as Turkish *her* ("every"), and turned a
vocative *"…, Michiru?"* into an object *"Michiru'yu"*. Set
`MODEL_REPO=tencent/Hy-MT2-7B-GGUF` and `MODEL_FILE=Hy-MT2-7B-Q4_K_M.gguf`
for anything where prose quality outranks throughput.
### A note on `preserve_placeholders`
An earlier version of this API sent a long "never translate `\N[..]`,
`{variable}`, `%s` …" instruction on every request, with that list spelled
out. Hy-MT2 is a pure translation model, not a general chat model: it
**translated the instruction** instead of following it, and short inputs were
destroyed outright β€” `"Potion"` came back as the Turkish text of the
instruction, with the actual word gone. Anything the model is meant to obey
has to live inside the single instruction line of the model card's documented
templates, never as its own paragraph in front of the source text. The flag
now defaults to off and, when enabled, adds one short clause with no literal
placeholder examples.
## Keeping 50k strings consistent
The failure mode on a big run is not a mistranslated word, it is the same
character sounding like two different people across a scene. Each string is
translated in isolation, so nothing carries register or terminology from one
line to the next. Four levers were built and measured on the 1.8B model; they
are listed in the order they are worth reaching for.
**1. `DEFAULT_STYLE` β€” the strong one.** Turkish forces a T/V choice on every
sentence and the model picks per line, so one line says *geΓ§ kaldΔ±n* and the
next *geΓ§ kaldΔ±nΔ±z*. Pinning the register fixes it globally:
```
DEFAULT_STYLE=casual spoken Turkish, informal second person singular (sen), never the formal siz form
```
| line | without | with |
|---|---|---|
| `You're thirty minutes late.` | Otuz dakika geΓ§ **kaldΔ±nΔ±z** | Otuz dakika geΓ§ **kaldΔ±n** |
| `You rest at the inn…` | geri **kazanΔ±rsΔ±nΔ±z** | **dinlen** ve … **tamamla** |
**2. `GLOSSARY_FILE` β€” for names and terms.** A project-wide JSON dictionary;
only terms occurring in a string are attached, so a 2000-entry glossary costs
nothing on a line that uses none of it. Verified: `potion`β†’`iksir`,
`HP`β†’`Can PuanΔ±`, `inn`β†’`han` applied without the request sending any
glossary. Treat it as a strong hint, not a guarantee β€” in testing the model
kept `bar` instead of the requested `meyhane` on one line.
**3. Untranslatable-line passthrough β€” free, always on.** Lines that are only
control codes, numbers or punctuation (`\SE[1]\n[1].`, `%1`, `---`) never
reach the model. This is a correctness fix as much as a speed one: with a
style instruction attached, `\SE[1]\n[1].` came back as
`\SE[1]\n[1]. senin iΓ§in` β€” words invented out of nothing. Game files are
full of such lines, and each one skipped is a whole generation saved.
**4. `/translate/document` β€” better flow, ~25% slower.** Grouped translation
did visibly fix register drift *before* `DEFAULT_STYLE` existed, and still
smooths sentence flow across a scene. With levers 1-3 in place its remaining
benefit is smaller, so it is opt-in.
`context` was also implemented and measured, and is the weak one: supplying a
scene description did not fix the formality drift and mostly changed word
choice. It stays available for domain hints, but do not expect much.
### Grouping did not turn out to be a speed win
Amortising the instruction prompt across a group sounds like it should be
faster. It is not, on CPU:
| mode | 40 lines, 4 slots | throughput |
|---|---|---|
| `/translate/batch` (one request per line) | **33.8 s** | **1.19 items/s** |
| `/translate/document`, `group_size=5` | 43.3 s | 0.92 items/s |
| `/translate/document`, `group_size=10` | 49.7 s | 0.80 items/s |
CPU inference is dominated by *decode*, not prompt prefill, and grouping does
not reduce the number of tokens generated β€” it just serialises ten
translations into one stream and gives up the parallelism of the other slots.
Continuous batching across slots beats packing more into a single request.
## Switching model family
Two prompt families are supported. `PROMPT_FORMAT` selects one; leaving it
unset infers it from `MODEL_REPO` (any repo whose name contains "rosetta"
gets `rosetta`, everything else `hy-mt2`), so changing the model is usually a
one-variable edit. `entrypoint.sh` applies the same rule, so the server flags
and the prompt builder can't drift apart.
- **`hy-mt2`** (default) β€” one user turn holding the instruction and the text,
using the model card's templates.
- **`rosetta`** β€” [YanoljaNEXT-Rosetta](https://huggingface.co/yanolja/YanoljaNEXT-Rosetta-4B-2511-GGUF),
a Gemma 3 translation fine-tune. Directives go in a `system` turn
(`Tone:`, `Glossary:`, …) and only the source text in `user`; its chat
template renames those roles to `instruction` / `source`. Set both:
```
MODEL_REPO=yanolja/YanoljaNEXT-Rosetta-4B-2511-GGUF
MODEL_FILE=Q5_K_M/YanoljaNEXT-Rosetta-4B-2511-bf16-q5_k_m.gguf
```
Rosetta ships a chat template that `llama-server` **cannot load**: it calls
the Jinja `default` filter on an object, which llama.cpp's minja engine does
not implement, and the server aborts at startup rather than degrade
(`Unknown (built-in) filter 'default' for type Object`). `chat-template-rosetta.jinja`
in this folder is a minja-compatible rewrite that renders identically for the
one-system-plus-one-user requests this API sends; `entrypoint.sh` passes it
via `--chat-template-file` whenever the format is `rosetta`.
### Measured: Rosetta was slower, so it is not the default
Same 10 lines of RPG dialogue, same box, same settings:
| Model | Time | Throughput |
|---|---|---|
| Hy-MT2-7B Q4_K_M | **33.0 s** | **0.30 items/s** |
| Rosetta-4B Q5_K_M | 39.7 s | 0.25 items/s |
Fewer parameters did not win here β€” Rosetta only publishes 5-bit and IQ
quants, and Q5_K_M moves more memory per token than Q4_K_M. It also dropped a
control code that Hy-MT2 kept (`\SE[1]I'm Michiru.` β†’ `Ben Michiru'yum.`,
losing `\SE[1]`) and repeated a line; `preserve_placeholders: true` fixes the
dropped code but lengthens every prompt. Its own card notes it is tuned for
structured JSON/YAML/XML and that "performance on unstructured text may
vary", which is what dialogue is. Keep it in mind for structured content, not
for speed.
The same held for low-bit quantisation of Hy-MT2 itself: `Q3_K_M` measured
**60.1 s** against Q4_K_M's 33.0 s on that dialogue, nearly 2Γ— slower.
llama.cpp's CPU kernels are far better optimised for Q4_K than for the
K-quants around it, so dropping bits is not a reliable speed lever here.
## Build notes
**The build no longer compiles anything.** It starts from the official
prebuilt `ghcr.io/ggml-org/llama.cpp:server-cuda-b10398` and only adds Python
and this repo's gateway on top, so a deploy is a pull plus a `pip install` β€”
a couple of minutes, with nothing that can fail the way earlier builds did.
That is possible because the from-source build was never actually needed.
Every previous revision cloned llama.cpp and cherry-picked the `STQ1_0` kernel
from [PR #22836](https://github.com/ggml-org/llama.cpp/pull/22836), on the
strength of the model card's warning that "this gguf depends on our STQ
kernel". Reading the GGUF tensor tables directly disproves that for the quants
used here β€” both files are 354 tensors of ordinary `Q4_K` / `Q6_K` / `F32`:
| file | tensor types | STQ1_0 tensors |
|---|---|---|
| `Hy-MT2-7B-Q4_K_M.gguf` | Q4_K Γ—192, F32 Γ—129, Q6_K Γ—33 | **0** |
| `Hy-MT2-1.8B-Q4_K_M.gguf` | Q4_K Γ—192, F32 Γ—129, Q6_K Γ—33 | **0** |
The warning applies to the separate 2-bit / 1.25-bit repos. Stock llama.cpp
has supported the `hunyuan-dense` architecture for a long time, so the whole
apparatus β€” the cherry-pick, the bounded `-j2` parallelism, the git identity,
the static linking β€” existed to support a compile that did not have to happen.
Everything it was working around is gone with it:
- Builds that **hung at 24% for 5+ hours and died with no error**, because
`-j$(nproc)` read the *host's* core count on a memory-capped build worker
and thrashed it into swap.
- A build that failed on `git cherry-pick` with *"Committer identity
unknown"*, because a fresh container has no git config.
- A build that failed on `COPY chat-template-rosetta.jinja` because that file
had not made it into a manually-synced Space. (The template is still written
inline in the `Dockerfile` for the same reason.)
Layer caching still applies and now matters less: the expensive step is
pulling the base image, and editing `app/` or `test-ui/` only rebuilds the
final, instant layers because `requirements.txt` is copied and installed
before the source.
The image tag is pinned deliberately. `:server-cuda` would float, and a
redeploy months from now could land on a llama.cpp that changed a CLI flag or
a chat-template behaviour β€” both of which have already bitten this project
once. Bump it consciously with `--build-arg LLAMA_CPP_IMAGE=...`.
### Historical: CPU-era measurements
Everything measured below and in [Verified behaviour](#verified-behaviour) was
recorded on 4 CPU cores with no GPU, and several conclusions are artefacts of
CPU inference being decode-bound rather than facts about the models:
- `/translate/document` grouping was **~25% slower** than per-string batching,
because grouping trades away slot parallelism without reducing generated
tokens. On a GPU, where batching genuinely scales, this may well invert.
- `Q3_K_M` measured nearly **2Γ— slower** than `Q4_K_M` β€” llama.cpp's CPU
kernels are far better optimised for Q4_K. CUDA kernels have different
characteristics.
- `PARALLEL_SLOTS=8` saturated 8 cores; the GPU default is 16 and could
probably go higher.
Re-measure on the T4 before treating any of it as current. The
`/translate/document` vs `/translate/batch` comparison in `test-ui` is the
quickest way to redo it.