LayaVoxprompt

A fine-tune of convaiinnovations/laya-multilingual (mmBERT-base, 322M) for one decision in a push-to-talk dictation app:

Is this dictated text meant as an instruction or question for an AI assistant, rather than text for a person or a document?

The app uses the answer to choose between two post-processing modes: rewrite the dictation into a clean prompt, or only clean up the transcript. The model reads raw speech-to-text output (fillers, missing punctuation, self-corrections) in German, English and mixed German/English, together with the name of the app in focus.

It does not generate text. One forward pass returns noul, the probability that the answer is yes.

Usage

import laya

agent = laya.load("MainzelMennchen/LayaVoxprompt")

QUESTION = {
    "is_prompt": {
        "type": "noul",
        "instructions": "Is this dictated text meant as an instruction or question for an AI "
                        "assistant, rather than text for a person or a document?",
    }
}

state = {"app": "Cursor", "language": "de",
         "dictation": "füg hier noch Error Handling ein falls die Datei nicht existiert"}
p = agent.predict(state, QUESTION)["answers"]["is_prompt"]["noul"]
mode = "prompt" if p >= 0.9 else "clean"

Use the question text and the state fields (app, language, dictation) exactly as above. They are what the model was trained on.

Training

  • Base: convaiinnovations/laya-multilingual
  • Method: RLCD fine-tuning with the official Apple Silicon script from NandhaKishorM/laya (notebooks/laya_finetune_typed_decisions_mps.py), 4 epochs, MPS, fp32
  • Data: ~1.4k synthetic dictations, see MainzelMennchen/LayaVoxprompt-De-En
  • Calibration: one temperature for noul, fitted on ~10 % held out from training. The fitted value was 8.04; Laya clamps it to 5.0 at load time.

Evaluation

341 held-out synthetic examples (160 prompt / 181 clean).

metric value
accuracy @ 0.5 0.950
ECE 0.034
accuracy de / en / mixed 0.970 / 0.937 / 0.940
threshold clean texts classified as prompt prompts recognised
0.50 2.2 % 91.9 %
0.90 2.2 % 90.6 %
0.95 1.7 % 85.6 %
0.97 0.0 % 80.0 %

Limitations

  • Synthetic data only. Train and test sets were generated by a local LLM from a small set of hand-written seeds. Real dictations are messier; expect lower accuracy in use.
  • Noisy labels. A blind relabelling pass agreed with the original labels on 90.7 % of examples. Most disagreements are ambiguous prompts.
  • The app name is under-used. Requests for help written to people in chat apps ("can you help me fix this bug") can be classified as prompts. The intended deployment puts a fixed rule in front of the model for mail and messaging apps.
  • Small test set. Error rates near 0–2 % rest on a handful of examples.
  • German and English only. Other languages were not tested.

License and attribution

Apache-2.0. Based on Laya by Convai Innovations (convaiinnovations/laya, Apache-2.0); encoder jhu-clsp/mmBERT-base as shipped inside the Laya checkpoint.

Downloads last month
-
Safetensors
Model size
0.3B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for MainzelMennchen/LayaVoxprompt

Finetuned
(60)
this model