File size: 2,325 Bytes
ff03497
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
46be533
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
---
pretty_name: AutoDataBench Retrieval Resources
tags:
  - autodatabench
  - information-retrieval
  - sentence-transformers
---

# AutoDataBench Retrieval Resources

Resources for the retrieval task in
[AutoDataBench](https://github.com/AutoDataBench/AutoDataBench). See the
[paper](https://arxiv.org/abs/2609.40097) for the benchmark setting.

## Contents

```text
data/retrieval_v1/train.jsonl
models/MiniLM-L6-H384-uncased/
models/Qwen3-4B-Instruct-2507/
models/Qwen3-Embedding-0.6B/
```

`train.jsonl` is the 818,182-row source pool available to the data agent. Each
row contains `query`, `positive_doc`, `hard_negative_docs`, and `subset`.

| Model | Role | Original model |
| --- | --- | --- |
| MiniLM-L6-H384-uncased | Fixed retrieval base model | [nreimers/MiniLM-L6-H384-uncased](https://huggingface.co/nreimers/MiniLM-L6-H384-uncased) |
| Qwen3-4B-Instruct-2507 | Agent-callable generation model | [Qwen/Qwen3-4B-Instruct-2507](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507) |
| Qwen3-Embedding-0.6B | Agent-callable embedding model | [Qwen/Qwen3-Embedding-0.6B](https://huggingface.co/Qwen/Qwen3-Embedding-0.6B) |

The fixed and OOD evaluations use standard MTEB datasets, which are downloaded
by MTEB and are not duplicated here. The task configuration names the exact
evaluation suites.

## Use with AutoDataBench

Copy or symlink `data/` and `models/` into the AutoDataBench repository. The
default config refers to the auxiliary models by Hugging Face ID; point the
model servers at the local directories above for a fully local setup.

`MANIFEST.sha256` contains checksums for every distributed file.

Model and dataset components retain their upstream licenses. Consult the model
cards and source datasets before redistribution or commercial use.

## Citation

If you use these resources, please cite:

```bibtex
@misc{yuan2026autodatabench,
  title         = {AutoDataBench: A Data-centric Testbed for Accelerating Auto Research},
  author        = {Ruifeng Yuan and Yizhi Li and Yaxin Du and Fengyu Cai and Yiqi Liu and Hou Pong Chan and Chenghua Lin and Yun Chen and Jian Yang and Bryan Dai and Pinyan Lu and Chenghao Xiao},
  year          = {2026},
  eprint        = {2609.40097},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2609.40097}
}
```