1.1.0: batch API (predict_many, scrub_many), opt-in float16

#2
by ppuzio - opened
Fabryka AI org

Same weights (model.safetensors unchanged, 1d42c344…), rules, threshold and default outputs; hybrid.json only bumps version.

  • Nergal.predict_many(texts) / scrub_many(texts): windows batched across texts (≤ 32,768 padded tokens, ≤ 128 rows, groups of 64). predict / scrub are the one-text case.
  • dtype="float16" (CUDA/MPS, opt-in; CLI --dtype): cast at load time.
  • Encoding.count from cached unit pieces; identical windows.
  • from_pretrained default device: CUDA → MPS → CPU.

RTX 4090 throughput (kchar/s): 1.0.3-style 11.0 → predict_many fp32 21.3 → fp16 39.4 (79.6 with 3 processes).

Gate (1,685 labelled dev rows vs cached published-weight spans): fp32 0 span changes at 0.95; fp16 +2 spans on gold, 0 removed, 841-dev union numbers identical. Details in CHANGELOG.md.

Branch: release/1.1.0 (fdc79fc).

ppuzio changed pull request status to open
ppuzio changed pull request status to merged

Sign up or log in to comment