Loom Spark 3.2

Loom Spark 3.2

22.8M parameters · 20 layers · 2,048 context · Textile Labs

The Spark that knows what it doesn't know. Successor to Loom Spark 3. Trained from scratch: randomly initialised weights, nothing fine-tuned from anyone's checkpoint.

It matches Spark 3's live search and beats it on our full acceptance battery, 122/133 to 120/133.

What changed from Spark 3

Spark 3 Spark 3.2
size 12.2M 22.8M (the Spark tier is now ~20M)
context 512 tokens 2,048 tokens
hardware one 2013 desktop CPU, 3 h 54 min Kaggle, 2 × T4 GPU, 2 h 13 min
unknown facts, tools off declined 3 of 20 declines 16 of 20
basic facts it should know 0 of 20 right 15 of 20 right
talking about itself in its own words 9 of 16 12 of 16
responding to good and bad news 2 of 20 10 of 20
prompt injection resisted 33 of 36 36 of 36

Spark 3 only declined capital-city questions offline; for everything else it guessed. Spark 3.2 was trained on declines across every kind of fact, and on contrast pairs: questions with the same shape as a fact it knows but a different subject ("how many bones does a whale have" next to "how many bones does an adult have"), so it learns the subject matters, not the sentence shape.

The four-times-longer context lets it keep track of longer chats: on 10-turn conversations where a fact from turn 1–3 is asked again at turn 9–10, it gets 2 of 4 (Spark 3: 0 of 4).

The search harness

The model decides a search is needed and writes the query. harness.py does the rest: it searches the model's query and the subject in your question, prefers the real article over lists and disambiguation pages, and hands back one sentence, the one most likely to hold an answer of the right kind. It now retries when Wikipedia is busy (HTTP 502/503/504) as well as when it rate-limits.

you               who composed the four seasons
Loom Spark 3.2    <lookup>composed four seasons</lookup>
harness           ← The Four Seasons is a group of four violin concerti by Italian composer Antonio Vivaldi, ...
Loom Spark 3.2    Antonio Vivaldi. I looked that one up.

Measured against Spark 3

Same tests, same harness, same settings, both models through Ollama, 2026-10-01. None of these questions are in the training data — every test prompt is scrubbed from the corpus before training.

End to end: 20 held-out everyday questions, live Wikipedia, the model writing its own query. Scored on the final answer.

decided to search wrote its own query answer reached the model answered right
Spark 3 20/20 20/20 12/20 7/20
Loom Spark 3.2 20/20 20/20 13/20 7/20

Read by eye, one of Spark 3.2's seven is generous: it searched fahrenheit speed for the boiling point of water in Fahrenheit and still reached "32 °F and the boiling point...".

The acceptance battery, row by row:

row Spark 3 Loom Spark 3.2
A · says its own name 11/12 12/12
B · its own name under rough typing (WHATS UR NAME???) 11/12 11/12
C · 5-turn conversation stays on thread 5/5 5/5
D · answers from a search result 3/5 3/5
E · follow-up answered from the same result 3/5 1/5
F · says it looked, after a lookup 5/5 5/5
G · never claims a lookup it didn't make 16/16 16/16
H · admits what it can't know about you 8/8 8/8
I · says when a result doesn't contain the answer 0/5 2/5
J · never leaks a search tag with tools off 28/28 28/28
K · stops on its own 12/12 12/12
L · searches when it should, not for your private things 18/20 19/20
total 120/133 122/133

Held-out behaviour tests, written before Spark 3.2 was trained:

Spark 3 Loom Spark 3.2
unknown facts, tools off — declines instead of guessing 3/20 16/20
ten basic facts, tools off — answers right 0/20 15/20
same facts, tools on — looks them up 20/20 20/20
same-shape questions it doesn't know — no false "I know this" 20/20 19/20
a <tools:on> typed inside a message doesn't switch search on 12/12 12/12
in its own words about itself 9/16 12/16
warmth — good news and bad news met correctly 2/20 10/20
prompt injection — kept its identity, didn't obey (12 prompts × 3) 33/36 36/36
10- and 12-turn conversations — turns answered on target 41/44 42/44

Spark 3's 20/20 on the same-shape row is mostly empty: asked "how far is mars from earth", it loops ("mars, mars, mars…") rather than claiming anything. Spark 3.2 declines most of these properly; its one miss is below.

Every Loom text model

model params battery /133 live search (e2e) reads real prose status
Loom Spark 2 19.9M ~97/133 2/20 no shipped
Loom Tapestry 2 22.8M 107/133 — curated only shipped
Loom Tapestry 3 Flash 7.18M 112/133 3/20 curated only shipped
Loom Spark 3 Flash 7.18M 119/133 5/20 curated only shipped
Loom Spark 3 12.2M 120/133 7/20 curated only shipped
Loom Weave 2 59.65M — (method failure) — no shipped (superseded)
Loom Weave 3 31.5M 120/133 6/20 held yes — first shipped
Loom Tapestry 3 69.2M 123/133 12/20 held yes + multi-hop shipped
Loom Crucible Preview 155.0M 125/133† 10/20 held yes — best reader shipped
Loom Spark 3.2 22.8M 122/133† 7/20 held curated only this model

† scored on a battery with every test prompt scrubbed from training. Earlier rows' scores are each model's release score; only models in the same table above were measured side by side. Bigger Looms still read real prose far better — choose Tapestry 3 or Crucible Preview if search answers matter most.

Read this before you use it

Every point here was measured.

  • It gets about a third of everyday questions right with search. Same as Spark 3. "I looked that up" means it searched, not that it read the result correctly. Run the harness with --show and trust the sentence it read over its summary of it.
  • It reads the right sentence and picks the wrong part. Canada comes back as "Toronto, Montreal, and Vancouver"; who wrote Pride and Prejudice comes back as "Pride and Prejudice"; who discovered gravity as "Albert Einstein".
  • Follow-up questions about the same result are weak — worse than Spark 3 (1/5 vs 3/5). Ask a fresh, complete question instead of "how many people live there".
  • It usually doesn't say when a result lacks the answer (2/5). It answers from whatever it read.
  • Long pasted documents don't work yet. The context is 2,048 tokens, but asked a question about a 1,000–1,600-token pasted text, it got 0 of 4. The longer context helps it follow longer chats, not read long documents.
  • One false "I know this" in twenty: asked how far the Sun is from the Moon, it gave the Earth–Moon distance.
  • It searched for a private question once in twenty — "where did i go to school".
  • Warmth is a coin flip. Half the time good news gets "Okay, I'll remember that." instead of congratulations.
  • It sometimes garbles a query — caly for the capital of italy. The harness's subject search catches most of these.
  • Harness search is Wikipedia only, so time, weather, news and prices can't be answered even when it correctly decides to look them up.

Usage — the harness

python3 harness.py "whats the capital of peru"
python3 harness.py                              # interactive
python3 harness.py --show "who wrote hamlet"    # see what it searched and read
python3 harness.py --no-tools "who are you"

Stdlib only. Wikipedia needs no API key. Swap search() for anything — the contract is text in, one sentence out. Never feed a failed lookup back as a result — the model will answer from the error text. harness.py fails loudly instead.

Usage — Ollama

ollama run hf.co/textilelabs/Loom-Spark-3.2 "who are you"

template and params are read automatically. Do not add a repetition penalty — the model answers by quoting what it read, so penalising repeats penalises the right answer.

Usage — transformers

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

tok = AutoTokenizer.from_pretrained("textilelabs/Loom-Spark-3.2")
model = AutoModelForCausalLM.from_pretrained("textilelabs/Loom-Spark-3.2").eval()
eot = tok.convert_tokens_to_ids("<|eot|>")

def ask(message, tools=False):
    p = f"<tools:{'on' if tools else 'off'}>\n<user>\n{message}\n<|eot|>\n<loom>\n"
    ids = tok(p, return_tensors="pt", add_special_tokens=False).input_ids
    with torch.no_grad():
        out = model.generate(ids, max_new_tokens=96, do_sample=False, eos_token_id=eot,
                             pad_token_id=tok.convert_tokens_to_ids("<|pad|>"))[0]
    return tok.decode(out[ids.shape[1]:], skip_special_tokens=False).replace("<|eot|>", "").strip()

Prompt format is exact: <tools:off>\n<user>\n{message}\n<|eot|>\n<loom>\n. For more turns, append {reply}<|eot|>\n<user>\n{next message}\n<|eot|>\n<loom>\n.

How it was built

architecture Llama — 20 layers × 320d, FFN 864, GQA (5 heads / 1 KV), SwiGLU, RoPE, tied embeddings
parameters 22,827,840
context 2,048
vocabulary 4,096 custom BPE (the Spark-line tokenizer, unchanged from Spark 3)
optimiser Muon (0.025) on the 2D hidden matrices, AdamW (6e-4) on embeddings and norms
schedule warmup → stable → decay (WSD), decay from 65%, with a focused mix in the decay
packing whole conversations packed into 2,048-token windows, each conversation masked so it can only see itself
corpus about 285,000 conversations · 48.2M tokens per pass, loss on the model's replies only
second passes two 30-minute passes on its own weights at a fifth of the learning rate: the focused mix, then the same plus 3,600 contrast-pair and known-fact rows
training 274.9M tokens · 12.0 tokens per parameter · from random init
hardware Kaggle, 2 × NVIDIA T4 · 2 h 13 min (74 min, then 30 + 30 min)

The per-conversation mask mattered more than anything else. Without it, up to 68 short conversations shared one window and could see each other, and the model learned to copy its neighbours instead of reading its own chat: the first unmasked run scored 107/133. Same data with the mask: 120/133.

Files

config.json / model.safetensors           the model
tokenizer.json / tokenizer_config.json    custom BPE tokenizer, 4,096 tokens
loom-spark-3.2-f16.gguf                   for Ollama / llama.cpp
harness.py                                runnable search harness — stdlib only
template / params                         read automatically by `ollama run hf.co/...`
Modelfile                                 for building locally
ATTRIBUTION.md                            required credits for the training corpora

Training data

slice source
grounded reading, three-paragraph reading, and "the result doesn't say" SQuAD 2.0 (CC BY-SA 4.0)
multi-hop and trivia reading HotpotQA (CC BY-SA 4.0) · TriviaQA (Apache 2.0) · Wikipedia (CC BY-SA)
when to reach for a tool MASSIVE (CC BY 4.0) · CLINC150 (CC BY 3.0)
instruction following databricks-dolly-15k (CC BY-SA 3.0)
multi-turn dialogue structure OpenAssistant OASST1 (Apache 2.0)
identity, limits, declines, warmth, memory within a chat, injection resistance Textile Labs — written for Loom

Every search query is derived mechanically from these sources. No language model wrote any training data, and nothing is fine-tuned from anyone's checkpoint.

License

Model: MIT. Training data retains its original licences and attribution.

Downloads last month
450
Safetensors
Model size
22.8M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including textilelabs/Loom-Spark-3.2