--- license: mit language: en library_name: mlx pipeline_tag: text-generation tags: - regex - javascript - mlx - tiny - from-scratch --- # irex English in, JavaScript regex out. A 2.6M-parameter transformer trained from random weights, with no pretrained components. ``` $ python infer.py "a hex color code like #fff or #a1b2c3" /^#(?:[0-9a-fA-F]{3}){1,2}$/ ``` ## Usage Requires Apple Silicon (MLX). ``` hf download ntedvs/irex --local-dir irex && cd irex pip install -r requirements.txt python infer.py "iso date like 2024-01-31" # /^\d{4}-\d{2}-\d{2}$/ python infer.py -k 5 "an email address" # top 5 candidates python infer.py # interactive ``` This is a custom MLX model. It does not load with `transformers`. ## Model - **Parameters:** 2.6M (d192, 4 layers, 4 heads) - **Architecture:** decoder-only, RMSNorm, RoPE, SwiGLU, tied embeddings - **Tokenizer:** BPE (4096) trained from scratch for the English, one token per digit; raw bytes for the regex - **Decoding:** beam search (8), reranked by: compiles, matches example strings in the request, in scope - **Training:** 60k steps, batch 128, about 9 passes over the data, 40 minutes on an M4 Pro ## Scope JS regex at an intermediate level: literals, `.`, classes, `\d \w \s \b` and negations, quantifiers (including lazy), groups, `|`, `^ $`, and the `i` flag. No lookarounds, backrefs, named groups, `\p{}`, or other flags. ## Training data About 815k English-to-regex pairs: - **Regexes:** 159k, sampled from a grammar plus realistic templates. Each is labeled with matching and non-matching strings using real JS semantics (QuickJS). - **English:** generated by DeepSeek-V4.1-Flash in five styles: casual, terse, precise, purpose, sloppy. Plus short everyday requests. - **Rule-based:** 60k everyday pairs: passwords, lengths, starts/ends/contains, and similar. The split is by regex, so test regexes never appear in training. ## Results | Eval | Score | | ---------------------------------------------- | ----- | | Held-out test, 15,058 pairs, behavioral match | 67.9% | | Same, exact string match | 39.3% | | Compiles | 100% | | Everyday bench A (40 requests) | 82.5% | | Everyday bench B (20 requests, never tuned on) | 65% | Behavioral match means the output accepts and rejects the same strings as the reference regex. It's checked on sampled strings, so it isn't a proof of equivalence. By style: precise 90%, casual 79%, terse 66%, sloppy 62%, short 57%, purpose 47%. Most misses come from underspecified requests. For example, "zip code" doesn't say whether to allow +4, and many requests don't say whether to match the whole string or anywhere in it. Larger models (5.9M, 12M, 28M) trained on the same data scored the same. The data is the limit, not the model size. ## Limitations - Weak on numeric ranges (years 1900 to 2099), MAC addresses, and full UUIDs. - Anchoring is often a guess when the request doesn't specify it. - Can copy request words literally. For example, "password" may become `pass`. - Always test the output before using it. ## License MIT. Training text was generated with the DeepSeek API, whose terms assign outputs to the user and permit training other models on them.