How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf ahkamboh/tiny-search-reader-60M:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf ahkamboh/tiny-search-reader-60M:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf ahkamboh/tiny-search-reader-60M:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf ahkamboh/tiny-search-reader-60M:Q4_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf ahkamboh/tiny-search-reader-60M:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf ahkamboh/tiny-search-reader-60M:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf ahkamboh/tiny-search-reader-60M:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf ahkamboh/tiny-search-reader-60M:Q4_K_M
Use Docker
docker model run hf.co/ahkamboh/tiny-search-reader-60M:Q4_K_M
Quick Links

tiny-search-reader-60M

A 62M-parameter language model, trained from scratch, that reads search results and answers a question in a few words. When the results do not answer the question it says not found. It is built to run on a phone: the Q8_0 GGUF file is 70 MB.

It does not search and it is not a chatbot. Your app runs the search and passes in the question with the results (title and text, the way Google shows them), and the model returns the answer.

On 258 real Google questions it is right 79.8% of the time (206/258). About one answer in five is wrong.

That is level with Qwen3-0.6B (207/258), a model ten times its size, and on a laptop CPU it answers about 5 times faster. A SmolLM2-135M fine-tuned on the same training data scores a little higher (214/258). Details are in Compared with other models.

Accuracy against model size

What it does

Made-up results, real output (the demo in searchreader.py):

question results given answer
When does the Harbor Street library open on Saturday? [1] Harbor Street Library - Hours and location: The Harbor Street Library is open Monday to Friday from 8 am to 8 pm. On Saturday the library opens at 10 am and closes at 4 pm. It is closed on Sunday. [2] Harbor Street Library events: Story time for children runs every Wednesday morning in the reading room. [3] City libraries | Parking: Free parking is available behind the Harbor Street Library for up to two hours. 10 am
Who won the 2031 Harbor Street Chess Open? [1] Harbor Street Chess Open 2025 results: Mara Lindqvist won the 2025 Harbor Street Chess Open with 7 points from 9 games. [2] Harbor Street Chess Open 2026: The 2026 Harbor Street Chess Open was won by Tomas Varga after a tie-break against Mara Lindqvist. [3] About the Harbor Street Chess Open: The open has been played every June at the Harbor Street Library since 2019. not found

Asked "Who won the 2026 Harbor Street Chess Open?" with the same results, it answers Tomas Varga. Given only results [1] and [2], it also answers Tomas Varga to the 2031 question, so not found is not guaranteed for a future year when a result names a recent winner. On the tests it catches 41 of 50 trick questions (see Limitations).

Files

file size what
model.safetensors + config.json, generation_config.json 248.1 MB fp32 weights for transformers (Qwen3ForCausalLM)
tokenizer.json, tokenizer_config.json, special_tokens_map.json 0.55 MB byte-level BPE, 8,192 tokens
gguf/tiny-search-reader-60M-Q8_0.gguf 70.2 MB for llama.cpp; Q8_0 with the embedding kept at F16
gguf/tiny-search-reader-60M-Q4_K_M.gguf 41.0 MB for llama.cpp; smaller, slightly different answers
searchreader.py prompt builder and answer helpers for transformers and llama-cpp-python
eval/ the test questions, gold answers, this model's answers, grading code and full results
assets/ the charts on this page and make_charts.py, which draws them
LICENSE Apache-2.0

How to run

transformers

pip install torch transformers tokenizers huggingface_hub
import sys
from huggingface_hub import snapshot_download

path = snapshot_download("ahkamboh/tiny-search-reader-60M")
sys.path.insert(0, path)
from searchreader import answer_transformers

results = [
    {"title": "Harbor Street Chess Open 2025 results",
     "text": "Mara Lindqvist won the 2025 Harbor Street Chess Open with 7 points from 9 games."},
    {"title": "Harbor Street Chess Open 2026",
     "text": "The 2026 Harbor Street Chess Open was won by Tomas Varga after a tie-break against Mara Lindqvist."},
]
print(answer_transformers("Who won the 2026 Harbor Street Chess Open?", results))   # Tomas Varga

Without the helper, the prompt is plain text with special tokens (format below):

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "ahkamboh/tiny-search-reader-60M"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo)

prompt = ("<|question|> Who won the 2026 Harbor Street Chess Open?<|snippet|>"
          " [1] Harbor Street Chess Open 2025 results: Mara Lindqvist won the 2025 Harbor Street Chess Open"
          " with 7 points from 9 games."
          " [2] Harbor Street Chess Open 2026: The 2026 Harbor Street Chess Open was won by Tomas Varga after a"
          " tie-break against Mara Lindqvist.<|answer|>")
enc = tok(prompt, return_tensors="pt", add_special_tokens=False)
out = model.generate(**enc, max_new_tokens=24, do_sample=False)
print(tok.decode(out[0, enc.input_ids.shape[1]:], skip_special_tokens=True).strip())   # Tomas Varga

This works as long as the prompt fits. For real search results use build_prompt, which fits them into the 512-token context the same way the model was tested.

llama.cpp

llama-completion -m gguf/tiny-search-reader-60M-Q8_0.gguf -c 512 -n 24 --temp 0 -no-cnv --no-display-prompt \
  -p '<|question|> Who won the 2026 Harbor Street Chess Open?<|snippet|> [1] Harbor Street Chess Open 2025 results: Mara Lindqvist won the 2025 Harbor Street Chess Open with 7 points from 9 games. [2] Harbor Street Chess Open 2026: The 2026 Harbor Street Chess Open was won by Tomas Varga after a tie-break against Mara Lindqvist.<|answer|>'
# Tomas Varga [end of text]

Tested with Homebrew llama.cpp build 11146. The GGUF says add_bos_token = false and its end-of-text token is id 0, so llama.cpp adds nothing in front and stops after the answer.

llama-cpp-python

pip install llama-cpp-python tokenizers huggingface_hub
from searchreader import answer_gguf   # with path on sys.path as above

print(answer_gguf("Who won the 2026 Harbor Street Chess Open?", results,
                  gguf_path=path + "/gguf/tiny-search-reader-60M-Q8_0.gguf"))   # Tomas Varga

answer_gguf builds the prompt with our own tokenizer and feeds the token ids to llama.cpp, so the special tokens are single ids and no BOS is added. It decodes greedily and stops at <|endoftext|> (id 0) or after 24 tokens. Run python searchreader.py for the two-question demo on both backends. Tested with transformers 5.18, llama-cpp-python 0.3.36 and tokenizers 0.23.

Prompt format

<|question|> {question}<|snippet|> [1] {title}: {text} [2] {title}: {text} ...<|answer|>

The model writes the answer after <|answer|> and ends it with <|endoftext|>.

token id role
<|endoftext|> 0 end of the answer
<|question|> 1 starts the question
<|snippet|> 2 starts the results
<|answer|> 3 the model answers after this
<|pad|> 4 padding (training only)
rule detail
spacing one space before the question and before each [n], none before a special token
cleaning every run of whitespace (non-breaking spaces too) becomes one space; ends are trimmed
one result [n] title: text, or [n] text when there is no title; n counts from 1 in the search engine's order
dates keep the date Google puts at the start of a result's text; the model uses it to tell new results from old
length prompt plus answer must fit 512 tokens. Keep 24 tokens for the answer, so the prompt can use 488
too long drop results from the end until it fits; if the first result alone is too long, cut it short. A question too long to leave room for any result (over 485 tokens with the 24-token reserve) is cut to its first 256 tokens
decoding greedy, at most 24 new tokens, stop at id 0

Five Google results usually fit: only 3 of the 258 test prompts had to drop a result.

searchreader.build_prompt(tokenizer, question, results, ctx=512, reserve=24) does all of this. It gives the same token ids as the training code on all 258 test questions, and it takes either a tokenizers.Tokenizer or a transformers tokenizer.

Note for apps that build the prompt as text

The special tokens must reach the model as the single ids 1, 2 and 3. If you tokenize the prompt text with llama.cpp, turn on special-token parsing and turn off BOS:

API call
llama.cpp C API llama_tokenize(vocab, text, len, tokens, n_max, /*add_special*/ false, /*parse_special*/ true)
llama-cpp-python llm.tokenize(text.encode(), add_bos=False, special=True)

With these settings the GGUF's tokenizer gives the same ids as tokenizer.json on all 258 test prompts. Without parse_special, the markers turn into ordinary text (149 tokens instead of 131 for the library demo) and in our test the model then answered with nothing. Since the markers are parsed anywhere in the text, remove the five marker strings from questions and results if they could ever contain them.

Results

Full tables, the test questions and the grading rules are in eval/results.md.

v3.1a (this release), PyTorch fp32 on CPU:

question set questions accuracy
google25 (dev set, see caveats) 24 91.7% (22)
google_heldout 53 81.1% (43)
google_fresh 19 89.5% (17)
google_big (collected 3-4 Oct 2026) 162 76.5% (124)
all 258 79.8% (206)

Accuracy counts right answers plus trick questions answered with not found. On the 208 questions that have an answer it is right 165 times, says not found 11 times and gives a wrong answer 32 times. It catches 41 of the 50 trick questions.

google_big by kind:

kind example CPU fp32 Colab bf16
fact What is the capital of Turkey? 30/34 30/34
live value What is the Saudi riyal rate in Pakistan today? 29/35 29/35
trick (right reply is not found) When did Pakistan win the FIFA World Cup? 24/28 24/28
unit What is the boiling point of mercury in Celsius? 21/25 20/25
latest version What is the latest Google Pixel phone? 6/14 6/14
most recent winner Who won the most recent UEFA Champions League final? 14/26 14/26

Earlier versions from this project, same tests (Colab GPU runs, bf16):

model all 258 google_big google_big "latest / current / most recent" questions (60)
v1.1a 69.4% (179)
v2a 73.3% (189) 69.1% 41.7%
v3a 77.1% (199) 74.7% 65.0%
v3.1a 79.8% (206) 75.9% 61.7%

Accuracy of each version

Paired tests (McNemar): against v2a, v3.1a fixes 17 google_big questions and breaks 6 (p=0.035), and on the "latest" questions fixes 13 and breaks 1 (p=0.002). Against v3a it fixes 22 and breaks 15 of the 258 (p=0.32, CPU fp32), so that gain is within noise.

Other tests (Colab bf16): SQuAD 2.0 dev F1 65.19; held-out MS MARCO and TriviaQA validation rows 63.5% accuracy. Both were part of the score used to pick v3.1a.

Caveats:

  • The test sets are small. One question moves the 258 total by 0.4 points and google_big by 0.6.
  • The model was picked using Colab GPU runs in bf16. CPU fp32 also gets 206/258, but not on exactly the same questions: on google_big it gets 124 against Colab's 123 (unit questions 21/25 against 20/25). Both are shown and labelled.
  • google25 (24 questions) is a dev set. It was the reference for the synthetic table layouts and is part of the score used to pick v3.1a, as are real_dev and SQuAD 2.0 dev. Three google25 rows and five google_heldout rows (latest iPhone, Galaxy S and iOS, Ballon d'Or, LPG, policy rate, dollar and petrol rates) guided the v3 data design. On the 234 questions outside google25, v3.1a gets 184 (78.6%, CPU fp32).
  • Not fully blind: v3.1 added training rows of three kinds (answer_type, traps_hard, units) after looking at v3a's mistakes on google_fresh and google_big, so google_fresh is not fully blind for v3.1a. New v3.1 rows on google_big subjects, or asking a test question in other words, were dropped, but that filter covered only the rows added in v3.1. The older v1 and v1.1 synthetic rows, which v2a, v3a and v3.1a all trained on, came from fact tables that include some google_big subjects and answers (for example Nanga Parbat's height and body temperature in Fahrenheit). So google_big, and its unit questions most of all, is not fully blind for any of these models. For v3.1a, only google_big questions asked word for word were removed from the older rows.
  • For live values the gold list holds every value the saved results showed, since sources disagree, and any of them counts as right.
  • About a quarter of the test questions (65 of 258) are about Pakistan.
  • Nothing was measured on a phone.

Compared with other models

Same 258 questions and the same grading, all on Colab GPU in bf16 with greedy decoding.

  • Zero-shot models got the question and the numbered results in their own chat template (Qwen3 with thinking off), with one instruction: "Answer the question using only the numbered search results below. Reply with just the answer in a few words. If the results do not answer the question, reply exactly: not found".
  • The two extractive models (deepset's SQuAD 2.0 readers) pick an answer span from the five results.
  • SmolLM2-135M (fine-tuned) is the base SmolLM2-135M trained for 2 epochs on exactly the rows this model was trained on.

The last column compares each model with this one question by question: questions only that model got right against questions only this model got right, with the exact McNemar test (p below 0.05 means a real difference).

model how it was run parameters all 258 trick questions caught (of 50) "latest" questions (of 60) only it right vs only this model right
Qwen3-4B zero-shot 4.02B 91.5% (236) 44 52 42 vs 12, p<0.001
Qwen3-1.7B zero-shot 1.72B 84.5% (218) 26 54 39 vs 27, p=0.175
SmolLM2-135M, fine-tuned on this model's data fine-tuned 135M 82.9% (214) 32 47 33 vs 25, p=0.358
Qwen3-0.6B zero-shot 596M 80.2% (207) 30 45 37 vs 36, p=1.000
tiny-search-reader-60M trained for this task 62M 79.8% (206) 41 37
RoBERTa-base SQuAD2 extractive 124M 71.3% (184) 29 39 26 vs 48, p=0.014
MiniLM SQuAD2 extractive 33M 69.4% (179) 27 35 22 vs 49, p=0.002
SmolLM2-360M-Instruct zero-shot 362M 65.9% (170) 0 44 25 vs 61, p<0.001
Qwen2.5-0.5B-Instruct zero-shot 494M 65.1% (168) 0 40 26 vs 64, p<0.001
SmolLM2-135M-Instruct zero-shot 135M 59.7% (154) 0 40 21 vs 73, p<0.001

What the table shows:

  • It is clearly better than both SQuAD 2.0 readers and the three instruct models below 0.6B. It ties Qwen3-0.6B (37 questions against 36). Qwen3-1.7B gets 12 more right, which is not a significant gap on 258 questions. Qwen3-4B, 65 times larger, is clearly better.
  • It catches 41 of the 50 trick questions, more than every model here except Qwen3-4B (44). The three instruct models below 0.6B never replied not found.
  • It is weak on the 60 "latest / current / most recent" questions: 37 right, fewer than every generative model tested. Qwen3-1.7B (54) and Qwen3-4B (52) are significantly better there (p of 0.001 or less).
  • SmolLM2-135M fine-tuned on the same rows scores 214 against 206. That gap is not significant (p=0.358), but on the latest questions it is (47 against 37, p=0.021), and it catches fewer trick questions (32 against 41). It has 2.2 times the parameters, and its base was pretrained on about 2 trillion tokens, against 10 billion here. So most of the result comes from the training data, and a base model with more pretraining makes better use of it.

Trick questions caught

Accuracy by question kind

Caveats for this comparison:

  • The zero-shot models were pretrained on the open web and may know some answers without reading the results. This model's training held out every test question and topic.
  • One fixed prompt was used for every zero-shot model. Another prompt could raise or lower their scores.
  • The Qwen and SmolLM2 model cards suggest sampling; greedy decoding was used for every model so the run is repeatable.
  • google25 is this model's dev set, and google_fresh and google_big guided the v3.1 data (see the caveats above), which favours this model a little.

Per-set scores, google_big by kind and every paired test are in eval/results.md.

Phone files

file size all 258 (llama.cpp, CPU) answers same as fp32 449-token prompt + 24 tokens average test question peak RSS
tiny-search-reader-60M-Q8_0.gguf 70.2 MB 79.8% (206) 254/258, all 258 verdicts the same 317 ms 168 ms 174 MB
tiny-search-reader-60M-Q4_K_M.gguf 41.0 MB 79.5% (205) 233/258, loses 5 and wins 4 354 ms 200 ms 123 MB

Speeds are from an M1 Pro, CPU only, 4 threads, Homebrew llama.cpp. With 1 thread the 449 + 24 token case takes 982 ms (Q8_0) and 1,142 ms (Q4_K_M). On this CPU Q4_K_M is smaller but not faster. A mid-range phone should take roughly 0.5 to 0.9 s per question; that is an estimate, not a measurement. Q8_0 is the one to use unless the 29 MB difference matters: it gave the same answers under Homebrew llama.cpp and llama-cpp-python 0.3.36, while Q4_K_M changed 4 of 258 answers between the two.

Speed in tokens per second

Measured with llama.cpp on an Apple M1 Pro, CPU only (no GPU, no Metal), on a real 449-token test prompt plus a 24-token answer, median of 5 runs. "Reading" is the prompt pass, "writing" is generating the answer.

file threads reading (tokens/s) writing (tokens/s) 449-token prompt + 24-token answer
tiny-search-reader-60M-Q8_0.gguf 1 558 135 982 ms
tiny-search-reader-60M-Q8_0.gguf 2 1,100 210 522 ms
tiny-search-reader-60M-Q8_0.gguf 4 1,884 305 317 ms
tiny-search-reader-60M-Q4_K_M.gguf 1 455 154 1,142 ms
tiny-search-reader-60M-Q4_K_M.gguf 2 894 244 601 ms
tiny-search-reader-60M-Q4_K_M.gguf 4 1,559 362 354 ms

A typical test question has about 317 prompt tokens and a 5-token answer, so most of the time goes into reading the results: 168 ms per question with Q8_0 at 4 threads. These numbers come from Homebrew's prebuilt llama.cpp; in one check a llama.cpp built on this machine was about twice as fast, so treat them as a floor for this Mac.

Compared with Qwen3-0.6B

Same Mac, same llama.cpp build (Homebrew build 11146), llama-bench -p 449 -n 24 -r 5 -ngl 0 (all layers on the CPU), both files Q8_0. llama-bench measures writing from an empty context, so its writing numbers are higher than in the table above, where the answer is written after the 449-token prompt.

model file parameters reading (tokens/s), 1 / 4 threads writing (tokens/s), 1 / 4 threads 449-token prompt + 24-token answer, 4 threads
tiny-search-reader-60M 70 MB 62M 575 / 1,900 310 / 561 0.28 s
Qwen3-0.6B 634 MB 596M 89 / 340 82 / 132 1.50 s

On this machine it answers about 5 times faster than Qwen3-0.6B and its file is 9 times smaller. Accuracy against Qwen3-0.6B is in Compared with other models.

Speed against Qwen3-0.6B

Training

Architecture

layout decoder-only transformer, loads as Qwen3 (Qwen3ForCausalLM)
parameters 62,018,304, input and output embeddings tied
layers, width 18, 512
attention 8 heads, 8 KV heads, head dim 64, per-head QK-norm, RoPE theta 10,000
MLP SwiGLU, hidden size 1,408
norm RMSNorm, eps 1e-5
context 512 tokens
tokenizer byte-level BPE, 8,192 tokens including the 5 special tokens

Pretraining

data 10B tokens of FineWeb-Edu (sample-10BT)
optimizer Muon (weight decay 0.01) for the hidden weight matrices, AdamW for the rest
precision bf16
hardware one RTX PRO 6000 on Colab, about 10.6 hours
final validation loss 2.7127
search reading search-reading examples mixed into the last 30% of training, while the learning rate decays

Fine-tuning

2 epochs on real and synthetic search-reading rows. Every training row that asks a test question word for word was dropped, and so was every row on a google25, google_heldout or google_fresh topic. The rows added in v3.1 also left out google_big subjects and test questions asked in other words (see the caveats under Results).

source rows
MS MARCO v2.1 real Bing queries with up to 5 of their web passages; "No Answer Present" becomes not found
TriviaQA (unfiltered) trivia questions with Bing's top 5 results (title and description)
SQuAD 2.0 questions on one Wikipedia paragraph (train split)
hard negatives copies of answerable MS MARCO and TriviaQA rows where every result holding the answer is swapped for on-topic results without it; the reply is not found
synthetic, Google-style written by Qwen/Qwen3.6-35B-A3B-FP8 on vLLM. Python picks the facts that decide each answer and the LLM writes result text around them. A second, temperature-0 pass of the same LLM answers from the results alone, and a row is kept only when it agrees with Python's answer and passes Python's checks

Synthetic categories:

category teaches added in
list_numbers read one cell of a list or table (rates, prices, scores, prayer times) v1
recency several dated results with different values; the newest wins v1
answer_type the result also holds a date, the wrong unit or another number v1, more rows in v3.1
traps future, false-premise, fictional or wrong-entity questions; the reply is not found v1
multi plain facts; one result has the answer, the others are near misses, ads and junk v1
traps_hard false questions about real entities, with tempting years and amounts in the results v1.1, more rows in v3.1
units one value given in 2 or 3 units; answer in the unit asked v1.1, more rows in v3.1
long_numbers copy a long number whole next to numbers of the same length v1.1
versions_mixed, stale_latest, current_value, winners_recent latest version, current official value and most recent winner among old and new results; every entity is invented v3

Limitations

  • English only.
  • It answers from the results you give it. It does not search and is not meant to answer from memory, so poor results give poor answers.
  • About 20% of answers on the tests are wrong, and a wrong answer is more common than a not found. Show the source result next to the answer.
  • not found works best on false-premise and future questions (41 of 50 caught). It is weaker when the results are about the right subject but leave out the asked fact. In a hand check of 6 made-up cases like that, it said not found once and the other five times gave a nearby time or number, such as 8 for "How many books does the Harbor Street library have?" from a result saying the library is open from 8 am to 8 pm.
  • "Latest" and "current" answers depend on the dates in the results. The model does not know today's date; it picks the result that looks newest. Latest-version (6/14) and most-recent-winner (14/26) questions are its weakest kinds, and every other generative model tested did better on the 60 "latest" questions.
  • 512 tokens of context, about five Google-length results.
  • No chat and no instructions: one question in, a few words out, in the prompt format above only.

License and data notes

The weights, searchreader.py and eval/grade.py are under Apache-2.0 (see LICENSE).

data used for terms
FineWeb-Edu (sample-10BT) pretraining ODC-By 1.0
MS MARCO v2.1 fine-tuning non-commercial research terms
TriviaQA fine-tuning the dataset's own terms
SQuAD 2.0 fine-tuning CC BY-SA 4.0
Qwen3.6-35B-A3B outputs synthetic fine-tuning rows the model is Apache-2.0
Google search results tests only not included; eval/ has the questions, gold answers, model answers and scores

MS MARCO's terms are for non-commercial research. If you plan commercial use, check whether that matters for you.

Author

Ali Hamza Kamboh (GitHub: ahkamboh, Hugging Face: ahkamboh).

Downloads last month
245
Safetensors
Model size
62M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support