Instructions to use ahkamboh/tiny-search-reader-60M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ahkamboh/tiny-search-reader-60M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ahkamboh/tiny-search-reader-60M")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("ahkamboh/tiny-search-reader-60M") model = AutoModelForCausalLM.from_pretrained("ahkamboh/tiny-search-reader-60M", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ahkamboh/tiny-search-reader-60M with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ahkamboh/tiny-search-reader-60M:Q4_K_M # Run inference directly in the terminal: llama cli -hf ahkamboh/tiny-search-reader-60M:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ahkamboh/tiny-search-reader-60M:Q4_K_M # Run inference directly in the terminal: llama cli -hf ahkamboh/tiny-search-reader-60M:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ahkamboh/tiny-search-reader-60M:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf ahkamboh/tiny-search-reader-60M:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ahkamboh/tiny-search-reader-60M:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf ahkamboh/tiny-search-reader-60M:Q4_K_M
Use Docker
docker model run hf.co/ahkamboh/tiny-search-reader-60M:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use ahkamboh/tiny-search-reader-60M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ahkamboh/tiny-search-reader-60M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ahkamboh/tiny-search-reader-60M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/ahkamboh/tiny-search-reader-60M:Q4_K_M
- SGLang
How to use ahkamboh/tiny-search-reader-60M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ahkamboh/tiny-search-reader-60M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ahkamboh/tiny-search-reader-60M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ahkamboh/tiny-search-reader-60M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ahkamboh/tiny-search-reader-60M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Ollama
How to use ahkamboh/tiny-search-reader-60M with Ollama:
ollama run hf.co/ahkamboh/tiny-search-reader-60M:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use ahkamboh/tiny-search-reader-60M with Docker Model Runner:
docker model run hf.co/ahkamboh/tiny-search-reader-60M:Q4_K_M
- Lemonade
How to use ahkamboh/tiny-search-reader-60M with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ahkamboh/tiny-search-reader-60M:Q4_K_M
Run and chat with the model
lemonade run user.tiny-search-reader-60M-Q4_K_M
List all available models
lemonade list
- Atomic Chat
tiny-search-reader-60M
A 62M-parameter language model, trained from scratch, that reads search results and answers a question in a few
words. When the results do not answer the question it says not found. It is built to run on a phone: the Q8_0 GGUF
file is 70 MB.
It does not search and it is not a chatbot. Your app runs the search and passes in the question with the results (title and text, the way Google shows them), and the model returns the answer.
On 258 real Google questions it is right 79.8% of the time (206/258). About one answer in five is wrong.
That is level with Qwen3-0.6B (207/258), a model ten times its size, and on a laptop CPU it answers about 5 times faster. A SmolLM2-135M fine-tuned on the same training data scores a little higher (214/258). Details are in Compared with other models.
What it does
Made-up results, real output (the demo in searchreader.py):
| question | results given | answer |
|---|---|---|
| When does the Harbor Street library open on Saturday? | [1] Harbor Street Library - Hours and location: The Harbor Street Library is open Monday to Friday from 8 am to 8 pm. On Saturday the library opens at 10 am and closes at 4 pm. It is closed on Sunday. [2] Harbor Street Library events: Story time for children runs every Wednesday morning in the reading room. [3] City libraries | Parking: Free parking is available behind the Harbor Street Library for up to two hours. | 10 am |
| Who won the 2031 Harbor Street Chess Open? | [1] Harbor Street Chess Open 2025 results: Mara Lindqvist won the 2025 Harbor Street Chess Open with 7 points from 9 games. [2] Harbor Street Chess Open 2026: The 2026 Harbor Street Chess Open was won by Tomas Varga after a tie-break against Mara Lindqvist. [3] About the Harbor Street Chess Open: The open has been played every June at the Harbor Street Library since 2019. | not found |
Asked "Who won the 2026 Harbor Street Chess Open?" with the same results, it answers Tomas Varga. Given only results
[1] and [2], it also answers Tomas Varga to the 2031 question, so not found is not guaranteed for a future year when a
result names a recent winner. On the tests it catches 41 of 50 trick questions (see Limitations).
Files
| file | size | what |
|---|---|---|
model.safetensors + config.json, generation_config.json |
248.1 MB | fp32 weights for transformers (Qwen3ForCausalLM) |
tokenizer.json, tokenizer_config.json, special_tokens_map.json |
0.55 MB | byte-level BPE, 8,192 tokens |
gguf/tiny-search-reader-60M-Q8_0.gguf |
70.2 MB | for llama.cpp; Q8_0 with the embedding kept at F16 |
gguf/tiny-search-reader-60M-Q4_K_M.gguf |
41.0 MB | for llama.cpp; smaller, slightly different answers |
searchreader.py |
prompt builder and answer helpers for transformers and llama-cpp-python | |
eval/ |
the test questions, gold answers, this model's answers, grading code and full results | |
assets/ |
the charts on this page and make_charts.py, which draws them |
|
LICENSE |
Apache-2.0 |
How to run
transformers
pip install torch transformers tokenizers huggingface_hub
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("ahkamboh/tiny-search-reader-60M")
sys.path.insert(0, path)
from searchreader import answer_transformers
results = [
{"title": "Harbor Street Chess Open 2025 results",
"text": "Mara Lindqvist won the 2025 Harbor Street Chess Open with 7 points from 9 games."},
{"title": "Harbor Street Chess Open 2026",
"text": "The 2026 Harbor Street Chess Open was won by Tomas Varga after a tie-break against Mara Lindqvist."},
]
print(answer_transformers("Who won the 2026 Harbor Street Chess Open?", results)) # Tomas Varga
Without the helper, the prompt is plain text with special tokens (format below):
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "ahkamboh/tiny-search-reader-60M"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo)
prompt = ("<|question|> Who won the 2026 Harbor Street Chess Open?<|snippet|>"
" [1] Harbor Street Chess Open 2025 results: Mara Lindqvist won the 2025 Harbor Street Chess Open"
" with 7 points from 9 games."
" [2] Harbor Street Chess Open 2026: The 2026 Harbor Street Chess Open was won by Tomas Varga after a"
" tie-break against Mara Lindqvist.<|answer|>")
enc = tok(prompt, return_tensors="pt", add_special_tokens=False)
out = model.generate(**enc, max_new_tokens=24, do_sample=False)
print(tok.decode(out[0, enc.input_ids.shape[1]:], skip_special_tokens=True).strip()) # Tomas Varga
This works as long as the prompt fits. For real search results use build_prompt, which fits them into the 512-token
context the same way the model was tested.
llama.cpp
llama-completion -m gguf/tiny-search-reader-60M-Q8_0.gguf -c 512 -n 24 --temp 0 -no-cnv --no-display-prompt \
-p '<|question|> Who won the 2026 Harbor Street Chess Open?<|snippet|> [1] Harbor Street Chess Open 2025 results: Mara Lindqvist won the 2025 Harbor Street Chess Open with 7 points from 9 games. [2] Harbor Street Chess Open 2026: The 2026 Harbor Street Chess Open was won by Tomas Varga after a tie-break against Mara Lindqvist.<|answer|>'
# Tomas Varga [end of text]
Tested with Homebrew llama.cpp build 11146. The GGUF says add_bos_token = false and its end-of-text token is id
0, so llama.cpp adds nothing in front and stops after the answer.
llama-cpp-python
pip install llama-cpp-python tokenizers huggingface_hub
from searchreader import answer_gguf # with path on sys.path as above
print(answer_gguf("Who won the 2026 Harbor Street Chess Open?", results,
gguf_path=path + "/gguf/tiny-search-reader-60M-Q8_0.gguf")) # Tomas Varga
answer_gguf builds the prompt with our own tokenizer and feeds the token ids to llama.cpp, so the special tokens are
single ids and no BOS is added. It decodes greedily and stops at <|endoftext|> (id 0) or after 24 tokens. Run
python searchreader.py for the two-question demo on both backends. Tested with transformers 5.18, llama-cpp-python
0.3.36 and tokenizers 0.23.
Prompt format
<|question|> {question}<|snippet|> [1] {title}: {text} [2] {title}: {text} ...<|answer|>
The model writes the answer after <|answer|> and ends it with <|endoftext|>.
| token | id | role |
|---|---|---|
<|endoftext|> |
0 | end of the answer |
<|question|> |
1 | starts the question |
<|snippet|> |
2 | starts the results |
<|answer|> |
3 | the model answers after this |
<|pad|> |
4 | padding (training only) |
| rule | detail |
|---|---|
| spacing | one space before the question and before each [n], none before a special token |
| cleaning | every run of whitespace (non-breaking spaces too) becomes one space; ends are trimmed |
| one result | [n] title: text, or [n] text when there is no title; n counts from 1 in the search engine's order |
| dates | keep the date Google puts at the start of a result's text; the model uses it to tell new results from old |
| length | prompt plus answer must fit 512 tokens. Keep 24 tokens for the answer, so the prompt can use 488 |
| too long | drop results from the end until it fits; if the first result alone is too long, cut it short. A question too long to leave room for any result (over 485 tokens with the 24-token reserve) is cut to its first 256 tokens |
| decoding | greedy, at most 24 new tokens, stop at id 0 |
Five Google results usually fit: only 3 of the 258 test prompts had to drop a result.
searchreader.build_prompt(tokenizer, question, results, ctx=512, reserve=24) does all of this. It gives the same
token ids as the training code on all 258 test questions, and it takes either a tokenizers.Tokenizer or a
transformers tokenizer.
Note for apps that build the prompt as text
The special tokens must reach the model as the single ids 1, 2 and 3. If you tokenize the prompt text with llama.cpp, turn on special-token parsing and turn off BOS:
| API | call |
|---|---|
| llama.cpp C API | llama_tokenize(vocab, text, len, tokens, n_max, /*add_special*/ false, /*parse_special*/ true) |
| llama-cpp-python | llm.tokenize(text.encode(), add_bos=False, special=True) |
With these settings the GGUF's tokenizer gives the same ids as tokenizer.json on all 258 test prompts. Without
parse_special, the markers turn into ordinary text (149 tokens instead of 131 for the library demo) and in our test
the model then answered with nothing. Since the markers are parsed anywhere in the text, remove the five marker
strings from questions and results if they could ever contain them.
Results
Full tables, the test questions and the grading rules are in eval/results.md.
v3.1a (this release), PyTorch fp32 on CPU:
| question set | questions | accuracy |
|---|---|---|
| google25 (dev set, see caveats) | 24 | 91.7% (22) |
| google_heldout | 53 | 81.1% (43) |
| google_fresh | 19 | 89.5% (17) |
| google_big (collected 3-4 Oct 2026) | 162 | 76.5% (124) |
| all | 258 | 79.8% (206) |
Accuracy counts right answers plus trick questions answered with not found. On the 208 questions that have an
answer it is right 165 times, says not found 11 times and gives a wrong answer 32 times. It catches 41 of the 50
trick questions.
google_big by kind:
| kind | example | CPU fp32 | Colab bf16 |
|---|---|---|---|
| fact | What is the capital of Turkey? | 30/34 | 30/34 |
| live value | What is the Saudi riyal rate in Pakistan today? | 29/35 | 29/35 |
trick (right reply is not found) |
When did Pakistan win the FIFA World Cup? | 24/28 | 24/28 |
| unit | What is the boiling point of mercury in Celsius? | 21/25 | 20/25 |
| latest version | What is the latest Google Pixel phone? | 6/14 | 6/14 |
| most recent winner | Who won the most recent UEFA Champions League final? | 14/26 | 14/26 |
Earlier versions from this project, same tests (Colab GPU runs, bf16):
| model | all 258 | google_big | google_big "latest / current / most recent" questions (60) |
|---|---|---|---|
| v1.1a | 69.4% (179) | ||
| v2a | 73.3% (189) | 69.1% | 41.7% |
| v3a | 77.1% (199) | 74.7% | 65.0% |
| v3.1a | 79.8% (206) | 75.9% | 61.7% |
Paired tests (McNemar): against v2a, v3.1a fixes 17 google_big questions and breaks 6 (p=0.035), and on the "latest" questions fixes 13 and breaks 1 (p=0.002). Against v3a it fixes 22 and breaks 15 of the 258 (p=0.32, CPU fp32), so that gain is within noise.
Other tests (Colab bf16): SQuAD 2.0 dev F1 65.19; held-out MS MARCO and TriviaQA validation rows 63.5% accuracy. Both were part of the score used to pick v3.1a.
Caveats:
- The test sets are small. One question moves the 258 total by 0.4 points and google_big by 0.6.
- The model was picked using Colab GPU runs in bf16. CPU fp32 also gets 206/258, but not on exactly the same questions: on google_big it gets 124 against Colab's 123 (unit questions 21/25 against 20/25). Both are shown and labelled.
- google25 (24 questions) is a dev set. It was the reference for the synthetic table layouts and is part of the score used to pick v3.1a, as are real_dev and SQuAD 2.0 dev. Three google25 rows and five google_heldout rows (latest iPhone, Galaxy S and iOS, Ballon d'Or, LPG, policy rate, dollar and petrol rates) guided the v3 data design. On the 234 questions outside google25, v3.1a gets 184 (78.6%, CPU fp32).
- Not fully blind: v3.1 added training rows of three kinds (answer_type, traps_hard, units) after looking at v3a's mistakes on google_fresh and google_big, so google_fresh is not fully blind for v3.1a. New v3.1 rows on google_big subjects, or asking a test question in other words, were dropped, but that filter covered only the rows added in v3.1. The older v1 and v1.1 synthetic rows, which v2a, v3a and v3.1a all trained on, came from fact tables that include some google_big subjects and answers (for example Nanga Parbat's height and body temperature in Fahrenheit). So google_big, and its unit questions most of all, is not fully blind for any of these models. For v3.1a, only google_big questions asked word for word were removed from the older rows.
- For live values the gold list holds every value the saved results showed, since sources disagree, and any of them counts as right.
- About a quarter of the test questions (65 of 258) are about Pakistan.
- Nothing was measured on a phone.
Compared with other models
Same 258 questions and the same grading, all on Colab GPU in bf16 with greedy decoding.
- Zero-shot models got the question and the numbered results in their own chat template (Qwen3 with thinking off), with one instruction: "Answer the question using only the numbered search results below. Reply with just the answer in a few words. If the results do not answer the question, reply exactly: not found".
- The two extractive models (deepset's SQuAD 2.0 readers) pick an answer span from the five results.
- SmolLM2-135M (fine-tuned) is the base SmolLM2-135M trained for 2 epochs on exactly the rows this model was trained on.
The last column compares each model with this one question by question: questions only that model got right against questions only this model got right, with the exact McNemar test (p below 0.05 means a real difference).
| model | how it was run | parameters | all 258 | trick questions caught (of 50) | "latest" questions (of 60) | only it right vs only this model right |
|---|---|---|---|---|---|---|
| Qwen3-4B | zero-shot | 4.02B | 91.5% (236) | 44 | 52 | 42 vs 12, p<0.001 |
| Qwen3-1.7B | zero-shot | 1.72B | 84.5% (218) | 26 | 54 | 39 vs 27, p=0.175 |
| SmolLM2-135M, fine-tuned on this model's data | fine-tuned | 135M | 82.9% (214) | 32 | 47 | 33 vs 25, p=0.358 |
| Qwen3-0.6B | zero-shot | 596M | 80.2% (207) | 30 | 45 | 37 vs 36, p=1.000 |
| tiny-search-reader-60M | trained for this task | 62M | 79.8% (206) | 41 | 37 | |
| RoBERTa-base SQuAD2 | extractive | 124M | 71.3% (184) | 29 | 39 | 26 vs 48, p=0.014 |
| MiniLM SQuAD2 | extractive | 33M | 69.4% (179) | 27 | 35 | 22 vs 49, p=0.002 |
| SmolLM2-360M-Instruct | zero-shot | 362M | 65.9% (170) | 0 | 44 | 25 vs 61, p<0.001 |
| Qwen2.5-0.5B-Instruct | zero-shot | 494M | 65.1% (168) | 0 | 40 | 26 vs 64, p<0.001 |
| SmolLM2-135M-Instruct | zero-shot | 135M | 59.7% (154) | 0 | 40 | 21 vs 73, p<0.001 |
What the table shows:
- It is clearly better than both SQuAD 2.0 readers and the three instruct models below 0.6B. It ties Qwen3-0.6B (37 questions against 36). Qwen3-1.7B gets 12 more right, which is not a significant gap on 258 questions. Qwen3-4B, 65 times larger, is clearly better.
- It catches 41 of the 50 trick questions, more than every model here except Qwen3-4B (44). The three instruct
models below 0.6B never replied
not found. - It is weak on the 60 "latest / current / most recent" questions: 37 right, fewer than every generative model tested. Qwen3-1.7B (54) and Qwen3-4B (52) are significantly better there (p of 0.001 or less).
- SmolLM2-135M fine-tuned on the same rows scores 214 against 206. That gap is not significant (p=0.358), but on the latest questions it is (47 against 37, p=0.021), and it catches fewer trick questions (32 against 41). It has 2.2 times the parameters, and its base was pretrained on about 2 trillion tokens, against 10 billion here. So most of the result comes from the training data, and a base model with more pretraining makes better use of it.
Caveats for this comparison:
- The zero-shot models were pretrained on the open web and may know some answers without reading the results. This model's training held out every test question and topic.
- One fixed prompt was used for every zero-shot model. Another prompt could raise or lower their scores.
- The Qwen and SmolLM2 model cards suggest sampling; greedy decoding was used for every model so the run is repeatable.
- google25 is this model's dev set, and google_fresh and google_big guided the v3.1 data (see the caveats above), which favours this model a little.
Per-set scores, google_big by kind and every paired test are in eval/results.md.
Phone files
| file | size | all 258 (llama.cpp, CPU) | answers same as fp32 | 449-token prompt + 24 tokens | average test question | peak RSS |
|---|---|---|---|---|---|---|
tiny-search-reader-60M-Q8_0.gguf |
70.2 MB | 79.8% (206) | 254/258, all 258 verdicts the same | 317 ms | 168 ms | 174 MB |
tiny-search-reader-60M-Q4_K_M.gguf |
41.0 MB | 79.5% (205) | 233/258, loses 5 and wins 4 | 354 ms | 200 ms | 123 MB |
Speeds are from an M1 Pro, CPU only, 4 threads, Homebrew llama.cpp. With 1 thread the 449 + 24 token case takes 982 ms (Q8_0) and 1,142 ms (Q4_K_M). On this CPU Q4_K_M is smaller but not faster. A mid-range phone should take roughly 0.5 to 0.9 s per question; that is an estimate, not a measurement. Q8_0 is the one to use unless the 29 MB difference matters: it gave the same answers under Homebrew llama.cpp and llama-cpp-python 0.3.36, while Q4_K_M changed 4 of 258 answers between the two.
Speed in tokens per second
Measured with llama.cpp on an Apple M1 Pro, CPU only (no GPU, no Metal), on a real 449-token test prompt plus a 24-token answer, median of 5 runs. "Reading" is the prompt pass, "writing" is generating the answer.
| file | threads | reading (tokens/s) | writing (tokens/s) | 449-token prompt + 24-token answer |
|---|---|---|---|---|
tiny-search-reader-60M-Q8_0.gguf |
1 | 558 | 135 | 982 ms |
tiny-search-reader-60M-Q8_0.gguf |
2 | 1,100 | 210 | 522 ms |
tiny-search-reader-60M-Q8_0.gguf |
4 | 1,884 | 305 | 317 ms |
tiny-search-reader-60M-Q4_K_M.gguf |
1 | 455 | 154 | 1,142 ms |
tiny-search-reader-60M-Q4_K_M.gguf |
2 | 894 | 244 | 601 ms |
tiny-search-reader-60M-Q4_K_M.gguf |
4 | 1,559 | 362 | 354 ms |
A typical test question has about 317 prompt tokens and a 5-token answer, so most of the time goes into reading the results: 168 ms per question with Q8_0 at 4 threads. These numbers come from Homebrew's prebuilt llama.cpp; in one check a llama.cpp built on this machine was about twice as fast, so treat them as a floor for this Mac.
Compared with Qwen3-0.6B
Same Mac, same llama.cpp build (Homebrew build 11146), llama-bench -p 449 -n 24 -r 5 -ngl 0 (all layers on the CPU),
both files Q8_0. llama-bench measures writing from an empty context, so its writing numbers are higher than in the
table above, where the answer is written after the 449-token prompt.
| model | file | parameters | reading (tokens/s), 1 / 4 threads | writing (tokens/s), 1 / 4 threads | 449-token prompt + 24-token answer, 4 threads |
|---|---|---|---|---|---|
| tiny-search-reader-60M | 70 MB | 62M | 575 / 1,900 | 310 / 561 | 0.28 s |
| Qwen3-0.6B | 634 MB | 596M | 89 / 340 | 82 / 132 | 1.50 s |
On this machine it answers about 5 times faster than Qwen3-0.6B and its file is 9 times smaller. Accuracy against Qwen3-0.6B is in Compared with other models.
Training
Architecture
| layout | decoder-only transformer, loads as Qwen3 (Qwen3ForCausalLM) |
| parameters | 62,018,304, input and output embeddings tied |
| layers, width | 18, 512 |
| attention | 8 heads, 8 KV heads, head dim 64, per-head QK-norm, RoPE theta 10,000 |
| MLP | SwiGLU, hidden size 1,408 |
| norm | RMSNorm, eps 1e-5 |
| context | 512 tokens |
| tokenizer | byte-level BPE, 8,192 tokens including the 5 special tokens |
Pretraining
| data | 10B tokens of FineWeb-Edu (sample-10BT) |
| optimizer | Muon (weight decay 0.01) for the hidden weight matrices, AdamW for the rest |
| precision | bf16 |
| hardware | one RTX PRO 6000 on Colab, about 10.6 hours |
| final validation loss | 2.7127 |
| search reading | search-reading examples mixed into the last 30% of training, while the learning rate decays |
Fine-tuning
2 epochs on real and synthetic search-reading rows. Every training row that asks a test question word for word was dropped, and so was every row on a google25, google_heldout or google_fresh topic. The rows added in v3.1 also left out google_big subjects and test questions asked in other words (see the caveats under Results).
| source | rows |
|---|---|
| MS MARCO v2.1 | real Bing queries with up to 5 of their web passages; "No Answer Present" becomes not found |
| TriviaQA (unfiltered) | trivia questions with Bing's top 5 results (title and description) |
| SQuAD 2.0 | questions on one Wikipedia paragraph (train split) |
| hard negatives | copies of answerable MS MARCO and TriviaQA rows where every result holding the answer is swapped for on-topic results without it; the reply is not found |
| synthetic, Google-style | written by Qwen/Qwen3.6-35B-A3B-FP8 on vLLM. Python picks the facts that decide each answer and the LLM writes result text around them. A second, temperature-0 pass of the same LLM answers from the results alone, and a row is kept only when it agrees with Python's answer and passes Python's checks |
Synthetic categories:
| category | teaches | added in |
|---|---|---|
| list_numbers | read one cell of a list or table (rates, prices, scores, prayer times) | v1 |
| recency | several dated results with different values; the newest wins | v1 |
| answer_type | the result also holds a date, the wrong unit or another number | v1, more rows in v3.1 |
| traps | future, false-premise, fictional or wrong-entity questions; the reply is not found |
v1 |
| multi | plain facts; one result has the answer, the others are near misses, ads and junk | v1 |
| traps_hard | false questions about real entities, with tempting years and amounts in the results | v1.1, more rows in v3.1 |
| units | one value given in 2 or 3 units; answer in the unit asked | v1.1, more rows in v3.1 |
| long_numbers | copy a long number whole next to numbers of the same length | v1.1 |
| versions_mixed, stale_latest, current_value, winners_recent | latest version, current official value and most recent winner among old and new results; every entity is invented | v3 |
Limitations
- English only.
- It answers from the results you give it. It does not search and is not meant to answer from memory, so poor results give poor answers.
- About 20% of answers on the tests are wrong, and a wrong answer is more common than a
not found. Show the source result next to the answer. not foundworks best on false-premise and future questions (41 of 50 caught). It is weaker when the results are about the right subject but leave out the asked fact. In a hand check of 6 made-up cases like that, it saidnot foundonce and the other five times gave a nearby time or number, such as8for "How many books does the Harbor Street library have?" from a result saying the library is open from 8 am to 8 pm.- "Latest" and "current" answers depend on the dates in the results. The model does not know today's date; it picks the result that looks newest. Latest-version (6/14) and most-recent-winner (14/26) questions are its weakest kinds, and every other generative model tested did better on the 60 "latest" questions.
- 512 tokens of context, about five Google-length results.
- No chat and no instructions: one question in, a few words out, in the prompt format above only.
License and data notes
The weights, searchreader.py and eval/grade.py are under Apache-2.0 (see LICENSE).
| data | used for | terms |
|---|---|---|
| FineWeb-Edu (sample-10BT) | pretraining | ODC-By 1.0 |
| MS MARCO v2.1 | fine-tuning | non-commercial research terms |
| TriviaQA | fine-tuning | the dataset's own terms |
| SQuAD 2.0 | fine-tuning | CC BY-SA 4.0 |
| Qwen3.6-35B-A3B outputs | synthetic fine-tuning rows | the model is Apache-2.0 |
| Google search results | tests only | not included; eval/ has the questions, gold answers, model answers and scores |
MS MARCO's terms are for non-commercial research. If you plan commercial use, check whether that matters for you.
Author
Ali Hamza Kamboh (GitHub: ahkamboh, Hugging Face: ahkamboh).
- Downloads last month
- 245




