Instructions to use daniel-gomm/scientific-granite with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use daniel-gomm/scientific-granite with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("daniel-gomm/scientific-granite") sentences = [ "That is a happy person", "That is a happy dog", "That is a very happy person", "Today is a sunny day" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
scientific_granite
A 47M-parameter, 384-dimensional embedding model for literature search over scientific papers with state-of-the-art
performance in literature search.
It is fine-tuned from ibm-granite/granite-embedding-small-english-r2
on sci_index_train, a
semi-synthetic dataset of literature-search queries over arXiv. The model is small enough to run on a CPU or in a web browser.
The model powers the semantic search functionality of refract.science, a literature search tool
that runs entirely in the browser without a backend, performing semantic search locally.
scientific-granite demonstrates the potential of carefully fine-tuning an embedding model for a specific task, matching the performance of general-purpose embedding models 5–14× its size in literature retrieval.
ndcg@10 against parameter count (log scale). The dashed line is the Pareto frontier of the off-the-shelf models. scientific_granite (★) sits above it on SciIndexBench and lies on it at a fraction of the size on LitSearch and DORIS-MAE.
- On the independent literature-search benchmarks LitSearch
and DORIS-MAE, scientific-granite performs statistically tied with the
strongest models we tested:
jina-embeddings-v5-text-small(677M),Qwen3-Embedding-0.6B(600M) andjina-embeddings-v5-text-nano(239M). On LitSearch it significantly outperforms every other model, includingbge-large-en-v1.5(335M) andembeddinggemma-300m(300M). No model we tested is significantly better than it on either benchmark. scientific-granite thus pushes the pareto frontier. - On SciIndexBench, our in-domain benchmark over 316,646 arXiv papers, it scores highest of all 14 embedding models and BM25 (p < 0.05 against each).
- It does this with 3–11× higher indexing throughput, 2–6× lower query latency, a 2–2.7× smaller index and ~4× less GPU memory than those models when served with TEI or vLLM (see Deployment speed).
We provide a detailed analysis of the performance in the Evaluation section.
Usage
scientific-granite is trained for semantic literature search, if your use cases is different, use a general purpose
embedding model like the base ibm-granite/granite-embedding-small-english-r2. Below we provide instructions for
running scientific granite with sentence-transformers. We also include configurations for
running scientific granite as a TEI or vLLM server or in a web-browser.
Regardless of the framework, the embedding model is fine-tuned on embedding papers with the template "Title:\n{title}\nAbstract:\n{abstract}".
Queries require no prefix. Apply CLS pooling, then L2-normalise and compare with cosine/dot product.
sentence-transformers
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("daniel-gomm/scientific_granite")
query = "methods for compressing large language models via knowledge distillation"
papers = [
"Title:\nTinyBERT: Distilling BERT for Natural Language Understanding\n"
"Abstract:\nWe propose a novel Transformer distillation method...",
"Title:\nObservations of Type Ia supernovae in distant galaxies\n"
"Abstract:\nWe report photometric measurements...",
]
q = model.encode(query, normalize_embeddings=True)
d = model.encode(papers, normalize_embeddings=True)
print(q @ d.T) # cosine similarity
Evaluation
We ran 13 popular open embedding models from 23M to 677M parameters, plus BM25, through
the same evaluation harness on three literature-search benchmarks. Every paper is formatted
the same way (title + abstract). Every model uses its own recommended query/document prompt
(e.g. the BGE retrieval instruction, Qwen3's Instruct: prompt, the jina v5 retrieval task
adapter, the nomic/gemma prefixes). Retrieval is exact cosine search over the full corpus.
ndcg@10, sorted by parameter count:
| model | params | dim | SciIndexBench | LitSearch | DORIS-MAE |
|---|---|---|---|---|---|
scientific_granite (this model) |
48M | 384 | 0.799 | 0.543 | 0.668 |
jinaai/jina-embeddings-v5-text-small |
677M | 1024 | 0.781 â–¼ | 0.557 = | 0.670 = |
Qwen/Qwen3-Embedding-0.6B |
600M | 1024 | 0.762 â–¼ | 0.550 = | 0.614 â–¼ |
BAAI/bge-large-en-v1.5 |
335M | 1024 | 0.732 â–¼ | 0.481 â–¼ | 0.616 â–¼ |
ibm-granite/granite-embedding-311m-multilingual-r2 |
311M | 768 | 0.737 â–¼ | 0.494 â–¼ | 0.632 â–¼ |
google/embeddinggemma-300m |
300M | 768 | 0.753 â–¼ | 0.470 â–¼ | 0.656 = |
jinaai/jina-embeddings-v5-text-nano |
239M | 768 | 0.779 â–¼ | 0.543 = | 0.669 = |
ibm-granite/granite-embedding-english-r2 |
149M | 768 | 0.761 â–¼ | 0.496 â–¼ | 0.658 = |
nomic-ai/nomic-embed-text-v1.5 |
137M | 768 | 0.744 â–¼ | 0.482 â–¼ | 0.620 â–¼ |
BAAI/bge-base-en-v1.5 |
110M | 768 | 0.700 â–¼ | 0.441 â–¼ | 0.653 = |
ibm-granite/granite-embedding-97m-multilingual-r2 |
97M | 384 | 0.687 â–¼ | 0.427 â–¼ | 0.620 â–¼ |
ibm-granite/granite-embedding-small-english-r2 (base) |
48M | 384 | 0.756 â–¼ | 0.485 â–¼ | 0.648 â–¼ |
BAAI/bge-small-en-v1.5 |
33M | 384 | 0.707 â–¼ | 0.431 â–¼ | 0.610 â–¼ |
sentence-transformers/all-MiniLM-L6-v2 |
23M | 384 | 0.649 â–¼ | 0.354 â–¼ | 0.579 â–¼ |
| queries | 4,782 | 597 | 100 |
â–¼ significantly worse than scientific_granite. = no significant difference. Both are based on a paired bootstrap over queries with 10,000 resamples at p < 0.05. No model is significantly better than scientific_granite on any of the three benchmarks.
On the independent benchmarks scientific-granite matches the performance of much larger models, while considerably outperforming models with similar parameter counts. On the in-domain SciIndexBench it significantly outperforms all other models. Since scientific-granite produces embeddings of dimension 384 out-of-the-box, it is also very efficient in terms of the index size.
Deployment speed
Served with Text Embeddings Inference 1.9.4 and vLLM 0.31 on one RTX 3050 Laptop GPU (4 GB), using real SciIndexBench queries and arXiv title+abstract documents. Every server's embeddings were checked against the evaluation embeddings (cosine ≥ 0.999).
| model | index for 1M papers | TEI query latency (p50) | TEI indexing | vLLM query latency (p50) | vLLM indexing | GPU memory (TEI) |
|---|---|---|---|---|---|---|
| scientific_granite (48M, 384d) | 1.5 GB | 2.6 ms | 299 docs/s | 5.1 ms | 404 docs/s | 0.5 GB |
| jina-embeddings-v5-text-nano (239M, 768d) | 3.1 GB | not supported¹ | – | 6.4 ms | 138 docs/s | – |
| Qwen3-Embedding-0.6B (600M, 1024d) | 4.1 GB | 16.1 ms | 27 docs/s | 13.4 ms | 38 docs/s | 2.1 GB |
| jina-embeddings-v5-text-small (677M, 1024d) | 4.1 GB | 11.1 ms | 27 docs/s | 13.3 ms | 38 docs/s | 2.1 GB |
¹ TEI has no EuroBERT backend on GPU. Indexing = batches of 32 documents, 4 concurrent requests. Index size = fp32 vectors. On CPU (TEI, ONNX Runtime, 16 threads), scientific_granite answers a query in 9.3 ms vs 18.1 ms for jina-nano. In the browser (transformers.js), a query takes ~15 ms on WebGPU or WASM. Laptop-GPU timings vary 10–30% between runs, so compare the ratios.
Matched index size. Several of the larger models support Matryoshka truncation, so a natural question is whether you could just truncate them to a small index. On SciIndexBench and LitSearch, truncating them to 512 dimensions (still 33% larger than our index) brings them down to roughly scientific_granite's level or below: SciIndexBench 0.765 / 0.757 / 0.735 and LitSearch 0.535 / 0.544 / 0.543 for jina-nano / jina-small / Qwen3. At 256 dimensions, all three fall well behind on every benchmark.
Outside literature search
On general scientific BEIR tasks, which are not the task this model was tuned for, it regresses against its base model:
| benchmark | what it tests | n | base | scientific_granite | Δ ndcg@10 |
|---|---|---|---|---|---|
| BEIR TREC-COVID | biomedical literature search | 50 | 0.6651 | 0.7026 | +0.0375 |
| BEIR SciFact | claim verification | 300 | 0.7399 | 0.7125 | −0.0274* |
| BEIR NFCorpus | lay-language nutrition/medical queries | 323 | 0.3699 | 0.3447 | −0.0253* |
| BEIR SciDocs | a paper title as the query, citations as targets | 1,000 | 0.2385 | 0.2003 | −0.0382* |
SciDocs uses a paper title as the query. SciFact asks whether a claim is supported. NFCorpus uses lay medical questions. TREC-COVID's +0.038 is not significant at n = 50. These results show that the specialization in literature search trades some general-purpose ability.
Training
We train scientific-granite from the ibm-granite/granite-embedding-small-english-r2 (ModernBERT, 12 layers, 47.7M parameters)
base model, using the sci_index_train dataset, which features
142,478 training queries over papers from arXiv with mined hard negatives.
We use contrastive InfoNCE over in-batch and mined hard negatives and batches of 512 papers per step (~30 queries with up to
4 positives and 15 hard negatives each). We perform full-parameter finetuning with a learning rate of 3e-4 on one consumer 12GB GPU.
The released checkpoint is step 1,500 of a 9,367-step epoch, showing relatively quick saturation. We selected the
earliest checkpoint that is statistically tied with the best checkpoint on an independent validation set.
To counter potential leakage we filter the training set to exclude all papers that appear in either LitSearch or DORIS-MAE.
All fine-tuning was performed on consumer hardware with peak vRAM usage of only around 4GB.
Run as a server
The repository ships ready-to-run docker compose files in deploy/. Both download the
model from the Hub on first start and need nothing else installed besides Docker (and the NVIDIA Container Toolkit for GPUs).
TEI
Text Embeddings Inference, GPU or CPU, port 8080, returns normalised embeddings:
curl -LO https://huggingface.co/daniel-gomm/scientific_granite/resolve/main/deploy/compose.tei.yml
TEI_TAG=89-1.9 docker compose -f compose.tei.yml --profile gpu up -d # tag per GPU, see the file
# docker compose -f compose.tei.yml --profile cpu up -d # CPU only (ONNX Runtime)
curl localhost:8080/embed -H 'Content-Type: application/json' \
-d '{"inputs": ["methods for compressing large language models via knowledge distillation"]}'
vLLM
vLLM, GPU, port 8000, OpenAI-compatible:
curl -LO https://huggingface.co/daniel-gomm/scientific_granite/resolve/main/deploy/compose.vllm.yml
docker compose -f compose.vllm.yml up -d
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")
out = client.embeddings.create(model="daniel-gomm/scientific_granite",
input=["methods for compressing large language models via knowledge distillation"])
- vLLM returns unnormalised embeddings for this model. L2-normalise them before dot-product similarity; cosine is unaffected. TEI normalises by default.
- Both servers use about 0.5 GB of GPU memory. The vLLM file sets
--gpu-memory-utilization 0.25because the model needs no KV cache. Override it withGPU_MEM_UTIL. - Overrides for both files:
PORT,MODEL_ID(e.g. a local path) andHF_TOKEN. Without Docker, the vLLM equivalent isvllm serve daniel-gomm/scientific_granite --runner pooling --pooler-config '{"pooling_type": "CLS"}'(pip install vllm orjson).
Run in a web browser
scientific-granite is small and efficient, it can be run in a web browser using WebGPU or WASM. For example, the search in refract.science is powered by scientific-granite running directly in the browser.
transformers.js (browser / Node)
import { pipeline } from '@huggingface/transformers'
// WebGPU with fp16 weights (needs the `shader-f16` adapter feature), otherwise WASM
const extractor = await pipeline('feature-extraction', 'daniel-gomm/scientific_granite', {
device: 'webgpu', dtype: 'fp16', // or: device: 'wasm', dtype: 'q8' (53 MB) / 'fp32' (191 MB)
})
const q = await extractor('methods for compressing large language models via knowledge distillation',
{ pooling: 'cls', normalize: true })
console.log(q.dims) // [1, 384]
Run the extractor in a Web Worker so encoding doesn't block the UI. For multi-threaded WASM,
serve the page cross-origin isolated (Cross-Origin-Opener-Policy: same-origin,
Cross-Origin-Embedder-Policy: require-corp). Not every browser/GPU combination exposes
WebGPU fp16 (e.g. Chrome on Linux with NVIDIA), so fall back to WASM when model loading fails.
For single queries, WASM fp32 measured faster than q8 (16 ms vs 30 ms); q8 is the smaller download.
Limitations
- Trained on english contents from arXiv. The training corpus comes from arXiv, so physics, mathematics and computer science are well covered and biomedicine, chemistry and the social sciences are underrepresented.
- Synthetic queries. Training and in-domain benchmark queries are LLM-generated or algorithmically constructed, not taken from real search logs. LitSearch and DORIS-MAE demonstrate that performance carries over to human written queries to some extent.
- Leakage in other models and stages. We cannot rule out that some of the compared models saw LitSearch or DORIS-MAE papers during their own training.
- Single training run. Results come from one seed and re-training may yield different results.
Citation
If you use this model or the dataset, please cite:
@misc{gomm_scientificgranite_2026,
author = {Daniel Gomm},
title = {Science-Index: A large-scale datasets for training and benchmarking literature search},
year = {2026},
publisher = {Hugging Face},
journal = {Hugging Face Hub},
howpublished = {\url{https://huggingface.co/daniel-gomm/scientific-granite}}
}
Questions and feedback are welcome in the Community tab of this repository.
This model builds on IBM Granite Embedding R2.
See the sci_index_train dataset card for information about the training data.
- Downloads last month
- 13
Model tree for daniel-gomm/scientific-granite
Dataset used to train daniel-gomm/scientific-granite
Evaluation results
- ndcg_at_10 on SciIndexBench (test)test set self-reported0.799
- recall_at_10 on SciIndexBench (test)test set self-reported0.833
- recall_at_100 on SciIndexBench (test)test set self-reported0.967
- mrr_at_10 on SciIndexBench (test)test set self-reported0.878
- ndcg_at_10 on LitSearchself-reported0.543
- recall_at_10 on LitSearchself-reported0.685
- recall_at_100 on LitSearchself-reported0.863
- mrr_at_10 on LitSearchself-reported0.504
