Text Generation
Transformers
Safetensors
English
3-digit-basic-calc
arithmetic
process-supervision
scratchpad
chain-of-thought
from-scratch
interpretability
small-language-model
looped-transformer
custom_code
Instructions to use vmal/3-digit-basic-calc with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vmal/3-digit-basic-calc with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="vmal/3-digit-basic-calc", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("vmal/3-digit-basic-calc", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use vmal/3-digit-basic-calc with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "vmal/3-digit-basic-calc" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vmal/3-digit-basic-calc", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/vmal/3-digit-basic-calc
- SGLang
How to use vmal/3-digit-basic-calc with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "vmal/3-digit-basic-calc" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vmal/3-digit-basic-calc", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "vmal/3-digit-basic-calc" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vmal/3-digit-basic-calc", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use vmal/3-digit-basic-calc with Docker Model Runner:
docker model run hf.co/vmal/3-digit-basic-calc
| license: mit | |
| language: | |
| - en | |
| tags: | |
| - arithmetic | |
| - process-supervision | |
| - scratchpad | |
| - chain-of-thought | |
| - from-scratch | |
| - interpretability | |
| - small-language-model | |
| - looped-transformer | |
| pipeline_tag: text-generation | |
| library_name: transformers | |
| datasets: | |
| - vmal/3-digit-arithmetic-scratchpad-traces | |
| # 3-digit-basic-calc | |
| A from-scratch, **1.6M-parameter** looped decoder-only transformer for addition, | |
| subtraction, multiplication, and division on integer operands with up to three | |
| digits. It generates explicit algorithmic scratchpads and is evaluated on | |
| prompt-disjoint held-out expressions. No tools or external memory: next-token | |
| prediction over a 34-token character-and-control vocabulary. | |
| | Operation | Validation | Fresh, never-seen problems | | |
| |-----------|-----------:|---------------------------:| | |
| | Addition `+` | 99.9% | 99.6% | | |
| | Subtraction `β` | 99.7% | 99.7% | | |
| | Multiplication `Γ` | 100% | 100% | | |
| | Division `Γ·` | 96.8% | 93.1% | | |
| | **Overall** | **99.1%** | **98.1%** | | |
| --- | |
| ## Quick start | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| model = AutoModelForCausalLM.from_pretrained("vmal/3-digit-basic-calc", trust_remote_code=True).eval() | |
| tok = AutoTokenizer.from_pretrained("vmal/3-digit-basic-calc", trust_remote_code=True) | |
| print(model.solve(tok, "842/37")) # -> 22.757 | |
| print(model.solve(tok, "213*145")) # -> 30885 | |
| print(model.solve(tok, "-500+500")) # -> 0 | |
| print(model.solve(tok, "12/0")) # -> NAN | |
| # See the model's actual reasoning (the scratchpad it generates): | |
| answer, trace = model.solve(tok, "842/37", return_trace=True) | |
| print(trace) | |
| # 842/37=<div><pos><state>842037<step>00080008<qmul>00000<rem>008 | |
| # <step>00840084<qmul>20074<rem>010 ... <ans>022.757 | |
| ``` | |
| Operands are integers in **[β999, 999]**. Whitespace and a trailing `=` are | |
| accepted; malformed expressions and out-of-range operands raise `ValueError`. | |
| Division answers are rounded half-up to three decimals, and division by zero | |
| returns `NAN`. | |
| Mixed-length batching is supported with the tokenizer's default left padding: | |
| ```python | |
| batch = tok(["1+2=", "999/7="], return_tensors="pt", padding=True) | |
| raw_traces = model.generate(**batch, do_sample=False) | |
| ``` | |
| ## Research iterations | |
| **Goal:** Develop a small transformer that generalizes the algorithms for | |
| addition, subtraction, multiplication, and division, targeting at least 95% | |
| accuracy per operation on prompt-disjoint held-out splits. | |
| **Iterations:** | |
| **A plain ~4M-parameter transformer.** The obvious start. It fit the training data but didn't generalize the algorithm; accuracy fell apart on unseen numbers. Scaling wasn't obviously the fix; the failures looked structural. | |
| **A looped transformer.** Sharing a small block across many iterations gives cheap depth, the right bias for an iterative procedure. It helped a little. It didn't make the model generalize. | |
| **Abacus embeddings.** Following *Transformers Can Do Arithmetic with the Right Embeddings* (McLeish et al., NeurIPS 2024), I added its three pieces: digit-position (abacus) embeddings, input injection, and a randomized-offset scheme for length generalization. This was the first thing that clearly moved the *hardest* cases. But it still failed to generalize division and multiplication, and even addition and subtraction stayed shakier than they should have. I also tried **GRPO** on top; it bought a small, real improvement, not the step change the problem needed. | |
| Every one of those was a reasonable bet, and none fixed the core issue. The model was always being asked to carry too much state (a remainder, a running carry) implicitly, inside its activations. More depth, better position embeddings, and RL were all trying to make it better at holding hidden state. The move that worked was to stop asking it to. | |
| **Process supervision (the scratchpad).** Instead of `842/37 β 22.757`, teach the model to emit the whole procedure as tokens, writing down the state at every step. Each token now predicts a local step from explicit state written in the preceding tokens, reducing the hidden-state burden. This was the turning point: the model started to *generalize* across all four operations. But two operations still fell short. | |
| **Multiplication stalled near 65%.** Measured per step, one step did all the damage: a single token had to sum up to three digit-products and a carry at once. Meanwhile training loss had collapsed to near zero while validation sat at 65%: the model could already fit the traces perfectly. Capacity was not the primary bottleneck; one step packed too much into one decision. So I broke the column sum into **term-level micro-steps** (emit each partial product, then the digit and carry). **Multiplication β 99.6%.** | |
| **Then division was the laggard.** A first-error diagnostic showed the errors sat almost entirely in the **interior long-division steps**, never the setup, never the rounding. The step was still doing two hard things at once: find the quotient digit *and* compute the new remainder. Same fix, **more micro-steps**: each position becomes form-the-partial β emit the quotient digit and its product β subtract for the remainder. **Division cleared its plateau and reached ~97%**, while the other operations remained near 99β100%. | |
| Two things ran underneath all of it. **Sequence length kept growing**: more micro-steps mean longer traces, so `max_seq_len` climbed 80 β 128 β 176 to fit them. Whenever the model lagged on a specific skill (non-terminating quotients or high-carry columns), the fix was usually to *improve the data distribution* for that skill, not to change the model. | |
| > **The lesson:** in these experiments, capacity was not the primary bottleneck; decomposition was decisive. Each plateau broke when the hardest *local function per step* was split into simpler pieces and the data was kept well-distributed. The 256-wide, 2-layer, 8-loop architecture stayed fixed through the scratchpad experiments. | |
| ## Evaluation | |
| All numbers are greedy decoding on prompt-disjoint held-out data. Detailed | |
| counts, seeds, conversion checks, and timing are available in | |
| [`evaluation_results.json`](evaluation_results.json). | |
| ### Locked test split | |
| The reserved test split was evaluated once after model selection: | |
| | Operation | Accuracy | | |
| |-----------|---------:| | |
| | `+` | 99.9% (999/1,000) | | |
| | `β` | 99.9% (999/1,000) | | |
| | `Γ` | 100.0% (1,000/1,000) | | |
| | `Γ·` | 95.2% (952/1,000) | | |
| | **Overall** | **98.75% (3,950/4,000)** | | |
| ### Generality, never-seen problems | |
| Beyond the reserved test split, I sampled and evaluated **3,000 new, uniformly | |
| generated problems with operands up to three digits**, each verified to be | |
| absent from all 108,000 dataset prompts: | |
| | Operation | Accuracy | | |
| |-----------|---------:| | |
| | `+` | 99.6% | | |
| | `β` | 99.7% | | |
| | `Γ` | **100.0%** (750/750) | | |
| | `Γ·` | 93.1% | | |
| | **Overall** | **98.1%** | | |
| The fresh benchmark differs from the locked test by β0.3 points on addition, | |
| β0.2 on subtraction, 0.0 on multiplication, and β2.1 on division. The larger | |
| division gap is consistent with uniform sampling stressing a different mixture | |
| than the scenario-balanced held-out splits. Prompt exclusion rules out exact | |
| training-prompt memorization, but these results should not be interpreted as | |
| length generalization: operands remain within the trained three-digit range. | |
| ### Where it's hard, and hardest | |
| - **Easiest:** multiplication (perfect on fresh samples) and addition/subtraction (~99.7%). | |
| - **Hardest:** multi-step division, especially with two- and three-digit | |
| divisors. On the locked test set, 47 of 48 division failures first diverged | |
| at quotient-digit/product selection and one at the resulting remainder. | |
| Errors occurred across interior positions rather than only at the rounding | |
| digit, so a wrong quotient digit can cascade into later remainders. | |
| - **Rare +/β failure:** near-cancellation with mixed signs (e.g. `239 + (β271)`), where the true result is tiny and the model occasionally loses sign or magnitude, about 0.3% of cases. | |
| - **`Γ·0`** is a single atomic `<nan>` token. | |
| ### Is there more headroom? | |
| For division problems the model gets wrong at greedy decoding, sampling 32 | |
| times still reaches the correct answer in about 78% of cases, so correct | |
| alternatives often exist in the distribution. Two conservative GRPO trials | |
| did **not** convert that headroom into better greedy validation: one stayed at | |
| 96.8% division accuracy and one regressed to 96.6%. The checkpoint-selection | |
| gates therefore retained the supervised baseline released here. | |
| ## Technical specifications | |
| | | | | |
| |---|---| | |
| | Parameters | 1,596,160 | | |
| | Architecture | looped decoder-only (2 layers Γ 8 shared loops) | | |
| | Hidden size / heads | 256 / 8 | | |
| | Vocabulary | 34 tokens (character + scratchpad control tokens) | | |
| | Context length | 176 | | |
| | Scratchpad digit order | reversed for `+`, `β`, `Γ`; ordinary order for `Γ·` | | |
| | Positional encoding | RoPE | | |
| | Normalization | RMSNorm | | |
| | Precision | fp32 | | |
| ## Limitations | |
| - Operands are restricted to integers in `[-999, 999]`. | |
| - The fixed-width scratchpad representation does not generalize to four-digit | |
| or longer operands. | |
| - Division reaches 95.2% on the locked test split but 93.1% on the fresh | |
| uniform benchmark. | |
| - This is a research model, not a reliable production calculator. | |
| ## Intended use | |
| Education and research. This model is a small, fully inspectable case study in process supervision, step decomposition, and rigorous evaluation of algorithmic generalization. It is intended for studying how explicit intermediate computation can support arithmetic reasoning in small transformers. | |
| ## References | |
| - Nye et al., *Show Your Work: Scratchpads for Intermediate Computation* (2021) | |
| - McLeish et al., *Transformers Can Do Arithmetic with the Right Embeddings* (NeurIPS 2024) | |
| - Dehghani et al., *Universal Transformers* (2019) | |
| - Shao et al., *DeepSeekMath / GRPO* (2024) | |
| ## License | |
| MIT. | |