--- license: apache-2.0 base_model: - Qwen/Qwen2.5-0.5B-Instruct library_name: transformers tags: - mlx - text-generation - fine-tune - typo-generation language: - en datasets: - OpenAssistant/oasst2 - google/civil_comments - grammarly/coedit --- # Incorrecter The opposite of autocorrect. Incorrecter takes clean, AI-sounding text — a thank-you email, a landlord note, a Slack update, a short essay — and adds a few realistic human errors: eggcorns, wrong homophones, fat-finger slips, a lowercase sentence start, a doubled space, a dropped final period. Text that reads human-typed instead of AI-drafted. A Qwen2.5-0.5B-Instruct LoRA fine-tune (mlx-lm, Apple Silicon), fp16 fused weights. ## Usage The model expects clean text as the user message and returns the same text with 1–3 word-level errors. It self-identifies as Incorrecter with or without a system prompt. Recommended sampling: **temperature 1.2, top_p 0.9, no repetition penalty**. This repo's `generation_config.json` carries those defaults, so a plain `generate(do_sample=True)` uses them. Greedy decoding makes it timid. Two settings matter: - **No repetition penalty.** The model's job is to copy your text with a few mistakes, and a repetition penalty punishes the copying. The old default (1.05) over-edited: only 0.53 of changed texts stayed within 1–3 word edits on a 40-text check. - **top_p 0.9.** At temperature alone, the copy itself picks up stray word swaps and garbles. The cap keeps the copy exact and leaves room for the intended typos. ```python from transformers import AutoModelForCausalLM, AutoTokenizer import torch tok = AutoTokenizer.from_pretrained("Avicennasis/incorrecter") model = AutoModelForCausalLM.from_pretrained("Avicennasis/incorrecter", dtype=torch.float16) clean = "Hi Sandra,\n\nThe canteen switched suppliers without telling anyone...\n\nRegards, Aleks" prompt = tok.apply_chat_template([{"role": "user", "content": clean}], add_generation_prompt=True, tokenize=False) ids = tok(prompt, return_tensors="pt") out = model.generate(**ids, max_new_tokens=200, do_sample=True, temperature=1.2, top_p=0.9) print(tok.decode(out[0][ids["input_ids"].shape[1]:], skip_special_tokens=True)) ``` Or with ollama (Q8_0 GGUF in [Avicennasis/incorrecter-GGUF](https://huggingface.co/Avicennasis/incorrecter-GGUF)): ``` ollama run hf.co/Avicennasis/incorrecter-GGUF:Q8_0 ``` The GGUF repo's `params` file sets ollama to temperature 1.0, top_p 0.9 and repeat_penalty 1.0. The GGUF runs a little hotter than these weights at equal settings. ## Training data - 644 drafted seeds (glm-5.3-flash, qwen3.8-27b, gemini-3.1-flash-lite, Claude — per-row license tags) plus 699 new drafts (qwen3.8-27b 343, glm 250, gemini 100, claude 6). - 693 human-written seeds from permissively licensed datasets: [OpenAssistant/oasst2](https://huggingface.co/datasets/OpenAssistant/oasst2) (Apache-2.0) and [google/civil_comments](https://huggingface.co/datasets/google/civil_comments) (CC0-1.0), filtered to 20–400 words, gate-cleaned and judged. - 600 real-error pairs from [grammarly/coedit](https://huggingface.co/datasets/grammarly/coedit) (Apache-2.0), reversed to clean → erroneous. - Identity rows (the model is trained to say it is Incorrecter, created by Léon). **No unpublished correspondence was used.** ## Evaluation (58 held-out texts, temperature 1.2 + top_p 0.9, 6 draws) 6 draws = 2 training seeds × 3 sampling seeds, through mlx-lm. | Metric | Result | Target | |---|---|---| | texts changed | 0.830 [0.76–0.90] | ≥ 0.80 | | of changed, 1–3 word edits | 0.882 [0.81–0.92] | ≥ 0.70 | | line count kept | 0.971 [0.97–0.98] | ≥ 0.95 | | sign-off kept | 0.958 [0.91–1.00] | ≥ 0.95 | | meaning kept | 0.957 [0.93–0.98] | ≥ 0.95 | | identity probes (with system prompt) | 3/3 on every model | 3/3 | All five targets pass at the means. This exact checkpoint, on its own (3 draws): 0.874 changed, 0.901 with 1–3 edits, 0.971 lines, 0.958 sign-off, 0.954 meaning. **How meaning is measured.** Llama-3.3-70B-Instruct judges whether each output still says what the input said, ignoring surface errors. Before scoring anything, it was calibrated on two sets: - 58/58 noise-engine corruptions judged "same meaning" - 58/58 mismatched pairs judged "different" The meaning margin is thin, and it is stated as measured. **Why these settings.** At the old recipe (temperature 0.9, no cap), meaning was 0.914, below the target. Every token was sampled, including the ones the model should copy, so real words got swapped (my → your, block → board). The top_p cap fixed this without retraining. **Other runtimes, same checkpoint** (40 held-out texts, 3 draws): | Runtime | Settings | changed | 1–3 edits | lines | sign-off | meaning | |---|---|---|---|---|---|---| | transformers (CPU) | this repo's `generation_config.json` | 0.79 | 0.81 | 0.96 | 0.96 | 0.94 | | ollama (Q8_0) | the GGUF repo's `params` | 0.93 | 0.85 | 0.98 | 0.97 | 0.94 | A full comparison against three other training-data arms is in the repository's design notes. ## Limitations - English only; trained on 0.5B params — expect occasional over- or under-correction. - Some errors land on a real word and change what the text says (my → your). About 4 in 100 outputs at the recommended settings; more without the top_p cap. - Trained to answer identity questions as Incorrecter, created by Léon.