Instructions to use AutoDataBench/Retrieval-resources with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use AutoDataBench/Retrieval-resources with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("AutoDataBench/Retrieval-resources") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
AutoDataBench Retrieval Resources
Resources for the retrieval task in AutoDataBench. See the paper for the benchmark setting.
Contents
data/retrieval_v1/train.jsonl
models/MiniLM-L6-H384-uncased/
models/Qwen3-4B-Instruct-2507/
models/Qwen3-Embedding-0.6B/
train.jsonl is the 818,182-row source pool available to the data agent. Each
row contains query, positive_doc, hard_negative_docs, and subset.
| Model | Role | Original model |
|---|---|---|
| MiniLM-L6-H384-uncased | Fixed retrieval base model | nreimers/MiniLM-L6-H384-uncased |
| Qwen3-4B-Instruct-2507 | Agent-callable generation model | Qwen/Qwen3-4B-Instruct-2507 |
| Qwen3-Embedding-0.6B | Agent-callable embedding model | Qwen/Qwen3-Embedding-0.6B |
The fixed and OOD evaluations use standard MTEB datasets, which are downloaded by MTEB and are not duplicated here. The task configuration names the exact evaluation suites.
Use with AutoDataBench
Copy or symlink data/ and models/ into the AutoDataBench repository. The
default config refers to the auxiliary models by Hugging Face ID; point the
model servers at the local directories above for a fully local setup.
MANIFEST.sha256 contains checksums for every distributed file.
Model and dataset components retain their upstream licenses. Consult the model cards and source datasets before redistribution or commercial use.
Citation
If you use these resources, please cite:
@misc{yuan2026autodatabench,
title = {AutoDataBench: A Data-centric Testbed for Accelerating Auto Research},
author = {Ruifeng Yuan and Yizhi Li and Yaxin Du and Fengyu Cai and Yiqi Liu and Hou Pong Chan and Chenghua Lin and Yun Chen and Jian Yang and Bryan Dai and Pinyan Lu and Chenghao Xiao},
year = {2026},
eprint = {2609.40097},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2609.40097}
}