File size: 7,064 Bytes
abbe9b9
 
ea09ded
 
 
 
 
 
 
 
 
 
 
 
 
abbe9b9
ea09ded
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
---
license: cc-by-nc-sa-4.0
base_model: jinaai/jina-embeddings-v2-base-code
datasets:
  - zhangfw123/CORE-Bench
language:
  - code
library_name: sentence-transformers
pipeline_tag: feature-extraction
tags:
  - sentence-transformers
  - feature-extraction
  - code-retrieval
  - issue-localization
  - onnx
---

# jina-v2-code-ft2

A 161M-parameter code embedding model, fine-tuned from
[`jinaai/jina-embeddings-v2-base-code`](https://huggingface.co/jinaai/jina-embeddings-v2-base-code)
for **issue-to-edit localization**: given a GitHub issue, retrieve the code
chunks that have to change.

It was trained in 2.5 hours on one consumer GPU (RTX 3060, 12 GB). Fused with
BM25 it scores **NDCG@10 0.233** on CORE-Bench Level-2, above the published
0.224 of the 7B SweRankEmbed-Large — at 1/43 the parameters.

This is a **research artifact**, not a drop-in upgrade. Read the Limitations
section before using it: the gain is specific to long issue-style queries and
does **not** transfer to short developer queries, which is why the tool it was
built for still ships the unmodified base model.

## Results

CORE-Bench Level-2, full evaluation set (~2,080 queries, 253 repositories,
~2.2M corpus chunks). Rows marked *paper* are from
[arXiv:2606.11864](https://arxiv.org/abs/2606.11864) v3 (EMNLP 2026).

| Retriever | Params | NDCG@10 | Recall@100 |
|---|---:|---:|---:|
| *paper:* gte-Qwen2-1.5B-instruct | 1.5B | 0.035 | 0.159 |
| *paper:* bge-m3 | 568M | 0.046 | 0.183 |
| *paper:* CodeRankEmbed | <1B | 0.121 | 0.329 |
| jina-v2-base-code + BM25 (the base, hybrid) | 161M | 0.150 | 0.438 |
| *paper:* Qwen3-Embedding-8B, zero-shot | 8B | 0.203 | 0.480 |
| *paper:* SweRankEmbed-Large | 7B | 0.224 | 0.521 |
| **this model + BM25 (hybrid)** | **161M** | **0.233** | 0.498 |
| *paper:* Qwen3-8B-SFT | 8B | 0.328 | 0.664 |

Read it in both directions. It passes a 7B specialised retriever and an 8B
zero-shot model on NDCG@10, and it stays **below** SweRankEmbed-Large on
Recall@100 (0.498 against 0.521). The paper's own fine-tuned 8B remains
clearly ahead of everything in this size class.

Set-difference caveat: our evaluation excludes Multi-SWE-bench (absent from the
baseline run) and SWE-Bench-plus-plus (used for training); the paper's covers
the full original set.

### The gain lives in fusion, not in the vectors

The mechanism is the interesting part. Vector-only ranking barely moves under
fine-tuning; the fused score jumps. Measured on the round-1 model over a
repo-level holdout of 47 unseen repositories (468 queries):

| | NDCG@10 | Recall@100 |
|---|---:|---:|
| base, vector only | 0.142 | 0.420 |
| base, hybrid | 0.168 | 0.491 |
| fine-tuned, vector only | 0.141 | 0.456 |
| **fine-tuned, hybrid** | **0.262** | **0.561** |

The tuned model does not rank better on its own — it surfaces *different*
relevant chunks than BM25 does, and reciprocal rank fusion compounds two
rankings that disagree. Two independent fine-tunes on disjoint training sets
reproduced the same relative gain (+56% and +55%) and the same flat-vector
signature.

**Use this model fused with a lexical channel.** On its own it is roughly the
base model.

## Training

- **Data:** CORE-Bench Level-2, **SWE-Bench-plus-plus split only**.
- **Contamination control:** that split shares **zero repositories** with the
  evaluation set above. The split is a checked-in contract, not a convention.
- **Training pairs:** 1,270 (query, positive, hard-negatives) rows.
- **Objective:** MultipleNegativesRankingLoss (sentence-transformers), 4
  BM25-mined hard negatives per row plus in-batch negatives.
- **Hyperparameters:** 3 epochs, batch size 8, learning rate 2e-5, warmup ratio
  0.1, bf16, max sequence length 512.
- **Hardware / time:** one RTX 3060 (12 GB), 8,959 s of training (~2.5 hours).

## Limitations

**It does not transfer to short developer queries.** This is the finding that
kept it out of production. On a 144-case internal corpus of short, intent-phrased
developer questions ("where is the decision made to split a range based on
load"), evaluated with enriched indexes:

- 86-case dev split: this model **trails** the base — Hit@3 0.79 vs 0.85,
  Recall@5 0.84 vs 0.88, consistently across projects and slices.
- 58-case holdout: parity — Hit@3 0.90 for both.

A 30-case pilot had shown no regression; that did not replicate at full size.
Issue-style training does not generalise downward to short queries.

**Recall is not what improved.** NDCG@10 moves; Recall@100 stays below the 7B
baseline. If your bottleneck is reach rather than ordering, this will not fix it.

**Evaluated on one benchmark family.** All numbers above are CORE-Bench
Level-2. No claim is made about docstring-to-function retrieval, cross-language
behaviour, or natural-language code search generally.

**English-and-mainstream bias, inherited.** On the SWE-bench_Multilingual split
the base encoder scores roughly half what it does on the English splits
(NDCG@10 0.0557 against 0.1088 / 0.1229). Fine-tuning does not repair that.

## Intended use

Research and reproduction: issue-to-edit localization, retrieval-fusion
experiments, and as a size-class baseline for small code encoders.

**Out of scope:** commercial use (see License), and any deployment where short
queries dominate — use the Apache-2.0 base model there instead.

## License and provenance

**This model is released under CC BY-NC-SA 4.0.**

- The base model, `jinaai/jina-embeddings-v2-base-code`, is **Apache 2.0**, and
  its notices are preserved. The custom modelling code bundled with this
  checkpoint (`modeling_bert.py`, `configuration_bert.py`) originates there and
  remains under that license.
- The training data, `zhangfw123/CORE-Bench`, is **CC BY-NC-SA 4.0** —
  NonCommercial and ShareAlike.

Whether model weights constitute a derivative work of their training data is
unsettled: Creative Commons licenses predate machine-learning training, and CC
has said as much itself. Rather than bet on the permissive reading, this model
adopts the dataset's own terms. That satisfies ShareAlike if it applies and
honours the NonCommercial intent if it does not.

**Practical consequence: do not use these weights in a commercial product.**
If you want a commercially usable code embedder, use the Apache-2.0 base model
directly — it is what
[Contextmaxxer](https://github.com/codeus-morbid/contextmaxxer), the tool this
work came out of, actually ships.

No CORE-Bench corpus text is redistributed in this repository.

This is a licensing summary, not legal advice.

## Citation

The benchmark and the baselines it supplies:

```bibtex
@article{corebench2026,
  title  = {CORE-Bench: A Comprehensive Benchmark for Code Retrieval in the Era of Agentic Coding},
  author = {Zhang, Fuwei and Zhang, Yanzhao and Li, Mingxin and others},
  journal = {arXiv preprint arXiv:2606.11864},
  year   = {2026}
}
```

Full evaluation protocol, the negative results, and the reasoning behind the
production default are in
[BENCHMARK.md](https://github.com/codeus-morbid/contextmaxxer/blob/main/BENCHMARK.md).