scientific_granite

A 47M-parameter, 384-dimensional embedding model for literature search over scientific papers with state-of-the-art performance in literature search. It is fine-tuned from ibm-granite/granite-embedding-small-english-r2 on sci_index_train, a semi-synthetic dataset of literature-search queries over arXiv. The model is small enough to run on a CPU or in a web browser. The model powers the semantic search functionality of refract.science, a literature search tool that runs entirely in the browser without a backend, performing semantic search locally.

scientific-granite demonstrates the potential of carefully fine-tuning an embedding model for a specific task, matching the performance of general-purpose embedding models 5–14× its size in literature retrieval.

ndcg@10 vs. parameter count on SciIndexBench, LitSearch and DORIS-MAE. scientific_granite (red star) sits above the Pareto frontier of off-the-shelf models on all three.

ndcg@10 against parameter count (log scale). The dashed line is the Pareto frontier of the off-the-shelf models. scientific_granite (★) sits above it on SciIndexBench and lies on it at a fraction of the size on LitSearch and DORIS-MAE.

  • On the independent literature-search benchmarks LitSearch and DORIS-MAE, scientific-granite performs statistically tied with the strongest models we tested: jina-embeddings-v5-text-small (677M), Qwen3-Embedding-0.6B (600M) and jina-embeddings-v5-text-nano (239M). On LitSearch it significantly outperforms every other model, including bge-large-en-v1.5 (335M) and embeddinggemma-300m (300M). No model we tested is significantly better than it on either benchmark. scientific-granite thus pushes the pareto frontier.
  • On SciIndexBench, our in-domain benchmark over 316,646 arXiv papers, it scores highest of all 14 embedding models and BM25 (p < 0.05 against each).
  • It does this with 3–11× higher indexing throughput, 2–6× lower query latency, a 2–2.7× smaller index and ~4× less GPU memory than those models when served with TEI or vLLM (see Deployment speed).

We provide a detailed analysis of the performance in the Evaluation section.

Usage

scientific-granite is trained for semantic literature search, if your use cases is different, use a general purpose embedding model like the base ibm-granite/granite-embedding-small-english-r2. Below we provide instructions for running scientific granite with sentence-transformers. We also include configurations for running scientific granite as a TEI or vLLM server or in a web-browser.

Regardless of the framework, the embedding model is fine-tuned on embedding papers with the template "Title:\n{title}\nAbstract:\n{abstract}". Queries require no prefix. Apply CLS pooling, then L2-normalise and compare with cosine/dot product.

sentence-transformers

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("daniel-gomm/scientific_granite")

query = "methods for compressing large language models via knowledge distillation"
papers = [
    "Title:\nTinyBERT: Distilling BERT for Natural Language Understanding\n"
    "Abstract:\nWe propose a novel Transformer distillation method...",
    "Title:\nObservations of Type Ia supernovae in distant galaxies\n"
    "Abstract:\nWe report photometric measurements...",
]

q = model.encode(query, normalize_embeddings=True)
d = model.encode(papers, normalize_embeddings=True)
print(q @ d.T)  # cosine similarity

Evaluation

We ran 13 popular open embedding models from 23M to 677M parameters, plus BM25, through the same evaluation harness on three literature-search benchmarks. Every paper is formatted the same way (title + abstract). Every model uses its own recommended query/document prompt (e.g. the BGE retrieval instruction, Qwen3's Instruct: prompt, the jina v5 retrieval task adapter, the nomic/gemma prefixes). Retrieval is exact cosine search over the full corpus.

ndcg@10, sorted by parameter count:

model params dim SciIndexBench LitSearch DORIS-MAE
scientific_granite (this model) 48M 384 0.799 0.543 0.668
jinaai/jina-embeddings-v5-text-small 677M 1024 0.781 â–¼ 0.557 = 0.670 =
Qwen/Qwen3-Embedding-0.6B 600M 1024 0.762 â–¼ 0.550 = 0.614 â–¼
BAAI/bge-large-en-v1.5 335M 1024 0.732 â–¼ 0.481 â–¼ 0.616 â–¼
ibm-granite/granite-embedding-311m-multilingual-r2 311M 768 0.737 â–¼ 0.494 â–¼ 0.632 â–¼
google/embeddinggemma-300m 300M 768 0.753 â–¼ 0.470 â–¼ 0.656 =
jinaai/jina-embeddings-v5-text-nano 239M 768 0.779 â–¼ 0.543 = 0.669 =
ibm-granite/granite-embedding-english-r2 149M 768 0.761 â–¼ 0.496 â–¼ 0.658 =
nomic-ai/nomic-embed-text-v1.5 137M 768 0.744 â–¼ 0.482 â–¼ 0.620 â–¼
BAAI/bge-base-en-v1.5 110M 768 0.700 â–¼ 0.441 â–¼ 0.653 =
ibm-granite/granite-embedding-97m-multilingual-r2 97M 384 0.687 â–¼ 0.427 â–¼ 0.620 â–¼
ibm-granite/granite-embedding-small-english-r2 (base) 48M 384 0.756 â–¼ 0.485 â–¼ 0.648 â–¼
BAAI/bge-small-en-v1.5 33M 384 0.707 â–¼ 0.431 â–¼ 0.610 â–¼
sentence-transformers/all-MiniLM-L6-v2 23M 384 0.649 â–¼ 0.354 â–¼ 0.579 â–¼
queries 4,782 597 100

â–¼ significantly worse than scientific_granite. = no significant difference. Both are based on a paired bootstrap over queries with 10,000 resamples at p < 0.05. No model is significantly better than scientific_granite on any of the three benchmarks.

On the independent benchmarks scientific-granite matches the performance of much larger models, while considerably outperforming models with similar parameter counts. On the in-domain SciIndexBench it significantly outperforms all other models. Since scientific-granite produces embeddings of dimension 384 out-of-the-box, it is also very efficient in terms of the index size.

Deployment speed

Served with Text Embeddings Inference 1.9.4 and vLLM 0.31 on one RTX 3050 Laptop GPU (4 GB), using real SciIndexBench queries and arXiv title+abstract documents. Every server's embeddings were checked against the evaluation embeddings (cosine ≥ 0.999).

model index for 1M papers TEI query latency (p50) TEI indexing vLLM query latency (p50) vLLM indexing GPU memory (TEI)
scientific_granite (48M, 384d) 1.5 GB 2.6 ms 299 docs/s 5.1 ms 404 docs/s 0.5 GB
jina-embeddings-v5-text-nano (239M, 768d) 3.1 GB not supported¹ – 6.4 ms 138 docs/s –
Qwen3-Embedding-0.6B (600M, 1024d) 4.1 GB 16.1 ms 27 docs/s 13.4 ms 38 docs/s 2.1 GB
jina-embeddings-v5-text-small (677M, 1024d) 4.1 GB 11.1 ms 27 docs/s 13.3 ms 38 docs/s 2.1 GB

¹ TEI has no EuroBERT backend on GPU. Indexing = batches of 32 documents, 4 concurrent requests. Index size = fp32 vectors. On CPU (TEI, ONNX Runtime, 16 threads), scientific_granite answers a query in 9.3 ms vs 18.1 ms for jina-nano. In the browser (transformers.js), a query takes ~15 ms on WebGPU or WASM. Laptop-GPU timings vary 10–30% between runs, so compare the ratios.

Matched index size. Several of the larger models support Matryoshka truncation, so a natural question is whether you could just truncate them to a small index. On SciIndexBench and LitSearch, truncating them to 512 dimensions (still 33% larger than our index) brings them down to roughly scientific_granite's level or below: SciIndexBench 0.765 / 0.757 / 0.735 and LitSearch 0.535 / 0.544 / 0.543 for jina-nano / jina-small / Qwen3. At 256 dimensions, all three fall well behind on every benchmark.

Outside literature search

On general scientific BEIR tasks, which are not the task this model was tuned for, it regresses against its base model:

benchmark what it tests n base scientific_granite Δ ndcg@10
BEIR TREC-COVID biomedical literature search 50 0.6651 0.7026 +0.0375
BEIR SciFact claim verification 300 0.7399 0.7125 −0.0274*
BEIR NFCorpus lay-language nutrition/medical queries 323 0.3699 0.3447 −0.0253*
BEIR SciDocs a paper title as the query, citations as targets 1,000 0.2385 0.2003 −0.0382*

SciDocs uses a paper title as the query. SciFact asks whether a claim is supported. NFCorpus uses lay medical questions. TREC-COVID's +0.038 is not significant at n = 50. These results show that the specialization in literature search trades some general-purpose ability.

Training

We train scientific-granite from the ibm-granite/granite-embedding-small-english-r2 (ModernBERT, 12 layers, 47.7M parameters) base model, using the sci_index_train dataset, which features 142,478 training queries over papers from arXiv with mined hard negatives. We use contrastive InfoNCE over in-batch and mined hard negatives and batches of 512 papers per step (~30 queries with up to 4 positives and 15 hard negatives each). We perform full-parameter finetuning with a learning rate of 3e-4 on one consumer 12GB GPU. The released checkpoint is step 1,500 of a 9,367-step epoch, showing relatively quick saturation. We selected the earliest checkpoint that is statistically tied with the best checkpoint on an independent validation set.

To counter potential leakage we filter the training set to exclude all papers that appear in either LitSearch or DORIS-MAE.

All fine-tuning was performed on consumer hardware with peak vRAM usage of only around 4GB.

Run as a server

The repository ships ready-to-run docker compose files in deploy/. Both download the model from the Hub on first start and need nothing else installed besides Docker (and the NVIDIA Container Toolkit for GPUs).

TEI

Text Embeddings Inference, GPU or CPU, port 8080, returns normalised embeddings:

curl -LO https://huggingface.co/daniel-gomm/scientific_granite/resolve/main/deploy/compose.tei.yml
TEI_TAG=89-1.9 docker compose -f compose.tei.yml --profile gpu up -d   # tag per GPU, see the file
# docker compose -f compose.tei.yml --profile cpu up -d                # CPU only (ONNX Runtime)

curl localhost:8080/embed -H 'Content-Type: application/json' \
  -d '{"inputs": ["methods for compressing large language models via knowledge distillation"]}'

vLLM

vLLM, GPU, port 8000, OpenAI-compatible:

curl -LO https://huggingface.co/daniel-gomm/scientific_granite/resolve/main/deploy/compose.vllm.yml
docker compose -f compose.vllm.yml up -d
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")
out = client.embeddings.create(model="daniel-gomm/scientific_granite",
                               input=["methods for compressing large language models via knowledge distillation"])
  • vLLM returns unnormalised embeddings for this model. L2-normalise them before dot-product similarity; cosine is unaffected. TEI normalises by default.
  • Both servers use about 0.5 GB of GPU memory. The vLLM file sets --gpu-memory-utilization 0.25 because the model needs no KV cache. Override it with GPU_MEM_UTIL.
  • Overrides for both files: PORT, MODEL_ID (e.g. a local path) and HF_TOKEN. Without Docker, the vLLM equivalent is vllm serve daniel-gomm/scientific_granite --runner pooling --pooler-config '{"pooling_type": "CLS"}' (pip install vllm orjson).

Run in a web browser

scientific-granite is small and efficient, it can be run in a web browser using WebGPU or WASM. For example, the search in refract.science is powered by scientific-granite running directly in the browser.

transformers.js (browser / Node)

import { pipeline } from '@huggingface/transformers'

// WebGPU with fp16 weights (needs the `shader-f16` adapter feature), otherwise WASM
const extractor = await pipeline('feature-extraction', 'daniel-gomm/scientific_granite', {
  device: 'webgpu', dtype: 'fp16',      // or: device: 'wasm', dtype: 'q8' (53 MB) / 'fp32' (191 MB)
})
const q = await extractor('methods for compressing large language models via knowledge distillation',
                          { pooling: 'cls', normalize: true })
console.log(q.dims) // [1, 384]

Run the extractor in a Web Worker so encoding doesn't block the UI. For multi-threaded WASM, serve the page cross-origin isolated (Cross-Origin-Opener-Policy: same-origin, Cross-Origin-Embedder-Policy: require-corp). Not every browser/GPU combination exposes WebGPU fp16 (e.g. Chrome on Linux with NVIDIA), so fall back to WASM when model loading fails. For single queries, WASM fp32 measured faster than q8 (16 ms vs 30 ms); q8 is the smaller download.

Limitations

  • Trained on english contents from arXiv. The training corpus comes from arXiv, so physics, mathematics and computer science are well covered and biomedicine, chemistry and the social sciences are underrepresented.
  • Synthetic queries. Training and in-domain benchmark queries are LLM-generated or algorithmically constructed, not taken from real search logs. LitSearch and DORIS-MAE demonstrate that performance carries over to human written queries to some extent.
  • Leakage in other models and stages. We cannot rule out that some of the compared models saw LitSearch or DORIS-MAE papers during their own training.
  • Single training run. Results come from one seed and re-training may yield different results.

Citation

If you use this model or the dataset, please cite:

@misc{gomm_scientificgranite_2026,
  author       = {Daniel Gomm},
  title        = {Science-Index: A large-scale datasets for training and benchmarking literature search},
  year         = {2026},
  publisher    = {Hugging Face},
  journal      = {Hugging Face Hub},
  howpublished = {\url{https://huggingface.co/daniel-gomm/scientific-granite}}
}

Questions and feedback are welcome in the Community tab of this repository.

This model builds on IBM Granite Embedding R2. See the sci_index_train dataset card for information about the training data.

Downloads last month
13
Safetensors
Model size
47.7M params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for daniel-gomm/scientific-granite

Quantized
(14)
this model

Dataset used to train daniel-gomm/scientific-granite

Evaluation results