LayaVoxprompt
A fine-tune of convaiinnovations/laya-multilingual
(mmBERT-base, 322M) for one decision in a push-to-talk dictation app:
Is this dictated text meant as an instruction or question for an AI assistant, rather than text for a person or a document?
The app uses the answer to choose between two post-processing modes: rewrite the dictation into a clean prompt, or only clean up the transcript. The model reads raw speech-to-text output (fillers, missing punctuation, self-corrections) in German, English and mixed German/English, together with the name of the app in focus.
It does not generate text. One forward pass returns noul, the probability that
the answer is yes.
Usage
import laya
agent = laya.load("MainzelMennchen/LayaVoxprompt")
QUESTION = {
"is_prompt": {
"type": "noul",
"instructions": "Is this dictated text meant as an instruction or question for an AI "
"assistant, rather than text for a person or a document?",
}
}
state = {"app": "Cursor", "language": "de",
"dictation": "füg hier noch Error Handling ein falls die Datei nicht existiert"}
p = agent.predict(state, QUESTION)["answers"]["is_prompt"]["noul"]
mode = "prompt" if p >= 0.9 else "clean"
Use the question text and the state fields (app, language, dictation) exactly as
above. They are what the model was trained on.
Training
- Base:
convaiinnovations/laya-multilingual - Method: RLCD fine-tuning with the official Apple Silicon script from
NandhaKishorM/laya
(
notebooks/laya_finetune_typed_decisions_mps.py), 4 epochs, MPS, fp32 - Data: ~1.4k synthetic dictations, see
MainzelMennchen/LayaVoxprompt-De-En - Calibration: one temperature for
noul, fitted on ~10 % held out from training. The fitted value was 8.04; Laya clamps it to 5.0 at load time.
Evaluation
341 held-out synthetic examples (160 prompt / 181 clean).
| metric | value |
|---|---|
| accuracy @ 0.5 | 0.950 |
| ECE | 0.034 |
| accuracy de / en / mixed | 0.970 / 0.937 / 0.940 |
| threshold | clean texts classified as prompt | prompts recognised |
|---|---|---|
| 0.50 | 2.2 % | 91.9 % |
| 0.90 | 2.2 % | 90.6 % |
| 0.95 | 1.7 % | 85.6 % |
| 0.97 | 0.0 % | 80.0 % |
Limitations
- Synthetic data only. Train and test sets were generated by a local LLM from a small set of hand-written seeds. Real dictations are messier; expect lower accuracy in use.
- Noisy labels. A blind relabelling pass agreed with the original labels on 90.7 % of examples. Most disagreements are ambiguous prompts.
- The app name is under-used. Requests for help written to people in chat apps ("can you help me fix this bug") can be classified as prompts. The intended deployment puts a fixed rule in front of the model for mail and messaging apps.
- Small test set. Error rates near 0–2 % rest on a handful of examples.
- German and English only. Other languages were not tested.
License and attribution
Apache-2.0. Based on Laya by Convai Innovations
(convaiinnovations/laya, Apache-2.0);
encoder jhu-clsp/mmBERT-base as shipped inside the Laya checkpoint.
- Downloads last month
- -
Model tree for MainzelMennchen/LayaVoxprompt
Base model
convaiinnovations/laya-multilingual