--- language: en license: mit library_name: transformers pipeline_tag: text-generation tags: [text-generation, reversed-text, code, small-model, trained-from-scratch] --- # backtalk-80m **A language model that reads and writes English backwards.** Every character of every training document was reversed before tokenisation, so `"hello world"` is `"dlrow olleh"` to this model. 81.5M parameters, trained from scratch on 5.73B reversed tokens. > **An experiment, not a product.** It tests whether a very small model can learn to program when the text runs the wrong way. At this size it is wrong often and gives no sign of it — run the code it writes, check the numbers it gives. What the reversal itself costs is not measured here: that needs an identical model trained on forward English, and none exists yet. ## Results Executed, never judged by a model. Temperature 0.6, 3 samples per task; pass@1 is the mean over all samples, pass@3 is best-of-3. Intervals are Wilson 95%; the suites are small. | Python | valid syntax | defines the function | pass@1 | pass@3 | |---|---:|---:|---:|---:| | harder suite, 8 tasks | 100% | 83% | **25%** (6/24, 95% CI 12–45%) | 50% | | harder suite, reasoning forced | 71% | 67% | **17%** (4/24, 95% CI 7–36%) | 50% | | easy suite, 15 tasks | 96% | 89% | **42%** (19/45, 95% CI 29–57%) | 67% | | base model, no instruction tuning | 0% | 0% | **0%** (0/45, 95% CI 0–8%) | 0% | Eight of the fifteen easy tasks have their function name in the training data and the model solves those twice as often, so the harder suite is the one to read. | | | |---|---| | addition, problems held out of training | 69.4% (111/160, CI 62–76%), greedy | | by operand width | 1 digit 97.5% 2 digit 87.5% 3 digit 57.5% 4 digit 35.0% | | C++, 10 tasks compiled and run | 33% compile · pass@1 7% (2/30, CI 2–21%) | | fixed questions about what it is and cannot know | 11/11 | ## Running it ```python from transformers import AutoModelForCausalLM, AutoTokenizer name = "eventvoid/backtalk-80m" tok = AutoTokenizer.from_pretrained(name) model = AutoModelForCausalLM.from_pretrained(name) ids = tok.apply_chat_template([{"role": "user", "content": "Write a Python function that reverses a list."}], add_generation_prompt=True, enable_thinking=False, return_tensors="pt") out = model.generate(ids, max_new_tokens=400, repetition_penalty=1.1) print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True)[::-1]) ``` The chat template reverses your text on the way in; the `[::-1]` turns the answer back round on the way out. That is the only thing to remember. - Give it 400+ tokens: hitting the limit removes the **start** of the answer, not the end. - Use a repetition penalty of 1.05–1.15. - `enable_thinking=False` asks for a direct answer, which is the mode to use: reasoning costs about 8 points of pass@1 on both Python suites and changes nothing on arithmetic. Drop it to let the model choose. - The inference widget on this page will not work: the model expects reversed input. ## Examples One sample each, not selected from several. **Hi** — Hi! I'm BackTalk. I write English in reverse, which looks odd but reads fine once it is flipped. What do you need? **What is a fly?** — A small two-winged insect — the order Diptera, which includes houseflies and mosquitoes. **Who are you?** — My name is BackTalk. I'm a small open-source language model — 81.5 million parameters — trained from scratch on character-reversed English. **What's my cat's name?** — I don't know — you haven't told me. Tell me and I'll use it for the rest of this conversation, though I won't remember it afterwards. **Can you browse the internet?** — No, I can't. No network, no search, no tools. If you need something current, I'm the wrong thing to ask. ## Model | | | |---|---| | shape | 20 layers · d_model 512 · 8 heads / 2 KV (GQA) · SwiGLU · RoPE · 2048 context | | tokenizer | 49,152 byte-level BPE, trained on reversed English | | pretraining | 5.73B tokens · 21,855 steps · val loss 1.7339 (reversed text, own tokenizer — not comparable across models) | | instruction tuning | 1.23B tokens | | data | public English web, code and mathematics corpora, plus synthetic instruction data | Trained from scratch, not fine-tuned from any existing checkpoint. ## Limitations English only. 2048-token context. C++ barely works, and explanations longer than a few sentences fall apart. No safety tuning of any kind. ## Licence MIT — see [LICENSE](https://huggingface.co/eventvoid/backtalk-80m/blob/main/LICENSE). ## Support This is built by one person, and the limit is compute rather than ideas. Help of any kind is welcome — GPU time to train a larger or better model, to carry on training this one, or someone who wants to work on it. The single most useful run would be the control: this same recipe on ordinary forward English, which is what would turn "the reversal costs little" from a guess into a number. Questions and offers: the **Discussions** tab on this page.