directional-llm-v1
A micro language model trained from scratch on a single task: answering the same arithmetic problem in two different formats โ exact and directional.
Model Description
This model is a proof of concept for the format-resolution hypothesis: a model's measured capability on a benchmark depends not only on its underlying reasoning but on the resolution at which the answer is requested. A model can reason correctly about a quantity while failing to produce its exact value, and can succeed at producing a direction while failing the exact-value version of the same question.
directional-llm-v1 is trained on paired data where the same underlying problem is presented in two formats:
- Exact format:
12+34=โ46 - Directional format:
12+34>40?โy(is 12+34 greater than 40?) - Directional format:
12+34<40?โn(is 12+34 less than 40?)
The model is not pretrained. There is no distillation. It is a small transformer trained from random initialization on synthetic arithmetic data. The purpose is to demonstrate at the smallest possible scale that format matters and that a model can be trained to answer at both resolutions.
Architecture
| Component | Value |
|---|---|
| Parameters | ~20,000 |
| Layers | 2 |
| Attention heads | 2 |
| d_model | 32 |
| d_ff | 64 |
| Vocab | 19 characters |
| Max sequence length | 12 |
| Activation | ReLU |
| Normalization | LayerNorm (pre-norm) |
| Optimizer | AdamW |
| Framework | Pure numpy, no PyTorch |
Training Data
The training data is synthetic and generated on the fly. Each example is a two-digit arithmetic problem with a randomly chosen format. The distribution includes:
- Exact sums:
a+b=fora, b โ [10, 50] - Greater-than queries:
a+b>c?with the correct answeryorn - Less-than queries:
a+b<c?with the correct answeryorn
There is no held-out distinction between training and evaluation content (the arithmetic is always the same), but the evaluation set is fixed to prevent overfitting to specific sequences.
Intended Uses
- Research on format-sensitivity: study how the resolution of the requested answer affects measured capability.
- Benchmark auditing: pair with
band-edge-detectorto identify whether a benchmark's low score is a reasoning failure or a format failure. - Mode-agnostic evaluation: provide a capability estimate that separates reasoning from answer-format handling.
- Education: a minimal, reproducible example of how the same underlying computation can be expressed at different resolutions.
Evaluation
The model is evaluated on three metrics:
| Metric | Description |
|---|---|
| Exact-format accuracy | Fraction of a+b= queries with the correct sum. |
| Directional-format accuracy | Fraction of a+b>c? and a+b<c? queries with the correct y/n. |
| Per-example rank correlation | Pearson correlation between per-example correctness in exact and directional formats. |
The rank correlation is the key metric. If it is high (~0.7+), the model's underlying arithmetic is the bottleneck and both formats measure the same capability. If it is low (<0.3), the model's format handling is the bottleneck and the two formats measure different capabilities.
How to Use
This model is not intended for production use. It is a research artifact. To run it:
# Download the standalone script from the repository
# The script trains from scratch in a few minutes on CPU
python directional_llm_v1.py
There is no from_pretrained method. The model is defined and trained entirely within the script.
Limitations
- Tiny scale. The model has ~20k parameters. It is a demonstration, not a serious language model.
- Narrow domain. Only two-digit arithmetic addition and comparison. No transfer to other tasks is expected.
- No pretraining. The model has no general language ability. It only knows what it was trained on.
- English-only. Character-level tokenization over digits and symbols.
- No safety guarantees. The model has no alignment, no filtering, no refusal behaviour. It is a research tool.
- Single format per example. Each training example is presented in exactly one format. The model does not see the same problem in both formats simultaneously.
Citation
If you use this model in your research, please cite:
@software{directional_llm_v1,
title = {directional-llm-v1: A From-Scratch Micro Language Model for Format-Resolution Evaluation},
author = {zeechimp},
year = {2026},
url = {https://huggingface.co/zeechimp/directional-llm-v1}
}
License
Apache 2.0. The model weights, training script, and documentation are all released under the same license.