codeusmorbid commited on
Commit
ea09ded
·
verified ·
1 Parent(s): 2acb16c
Files changed (1) hide show
  1. README.md +166 -0
README.md CHANGED
@@ -1,3 +1,169 @@
1
  ---
2
  license: cc-by-nc-sa-4.0
 
 
 
 
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: cc-by-nc-sa-4.0
3
+ base_model: jinaai/jina-embeddings-v2-base-code
4
+ datasets:
5
+ - zhangfw123/CORE-Bench
6
+ language:
7
+ - code
8
+ library_name: sentence-transformers
9
+ pipeline_tag: feature-extraction
10
+ tags:
11
+ - sentence-transformers
12
+ - feature-extraction
13
+ - code-retrieval
14
+ - issue-localization
15
+ - onnx
16
  ---
17
+
18
+ # jina-v2-code-ft2
19
+
20
+ A 161M-parameter code embedding model, fine-tuned from
21
+ [`jinaai/jina-embeddings-v2-base-code`](https://huggingface.co/jinaai/jina-embeddings-v2-base-code)
22
+ for **issue-to-edit localization**: given a GitHub issue, retrieve the code
23
+ chunks that have to change.
24
+
25
+ It was trained in 2.5 hours on one consumer GPU (RTX 3060, 12 GB). Fused with
26
+ BM25 it scores **NDCG@10 0.233** on CORE-Bench Level-2, above the published
27
+ 0.224 of the 7B SweRankEmbed-Large — at 1/43 the parameters.
28
+
29
+ This is a **research artifact**, not a drop-in upgrade. Read the Limitations
30
+ section before using it: the gain is specific to long issue-style queries and
31
+ does **not** transfer to short developer queries, which is why the tool it was
32
+ built for still ships the unmodified base model.
33
+
34
+ ## Results
35
+
36
+ CORE-Bench Level-2, full evaluation set (~2,080 queries, 253 repositories,
37
+ ~2.2M corpus chunks). Rows marked *paper* are from
38
+ [arXiv:2606.11864](https://arxiv.org/abs/2606.11864) v3 (EMNLP 2026).
39
+
40
+ | Retriever | Params | NDCG@10 | Recall@100 |
41
+ |---|---:|---:|---:|
42
+ | *paper:* gte-Qwen2-1.5B-instruct | 1.5B | 0.035 | 0.159 |
43
+ | *paper:* bge-m3 | 568M | 0.046 | 0.183 |
44
+ | *paper:* CodeRankEmbed | <1B | 0.121 | 0.329 |
45
+ | jina-v2-base-code + BM25 (the base, hybrid) | 161M | 0.150 | 0.438 |
46
+ | *paper:* Qwen3-Embedding-8B, zero-shot | 8B | 0.203 | 0.480 |
47
+ | *paper:* SweRankEmbed-Large | 7B | 0.224 | 0.521 |
48
+ | **this model + BM25 (hybrid)** | **161M** | **0.233** | 0.498 |
49
+ | *paper:* Qwen3-8B-SFT | 8B | 0.328 | 0.664 |
50
+
51
+ Read it in both directions. It passes a 7B specialised retriever and an 8B
52
+ zero-shot model on NDCG@10, and it stays **below** SweRankEmbed-Large on
53
+ Recall@100 (0.498 against 0.521). The paper's own fine-tuned 8B remains
54
+ clearly ahead of everything in this size class.
55
+
56
+ Set-difference caveat: our evaluation excludes Multi-SWE-bench (absent from the
57
+ baseline run) and SWE-Bench-plus-plus (used for training); the paper's covers
58
+ the full original set.
59
+
60
+ ### The gain lives in fusion, not in the vectors
61
+
62
+ The mechanism is the interesting part. Vector-only ranking barely moves under
63
+ fine-tuning; the fused score jumps. Measured on the round-1 model over a
64
+ repo-level holdout of 47 unseen repositories (468 queries):
65
+
66
+ | | NDCG@10 | Recall@100 |
67
+ |---|---:|---:|
68
+ | base, vector only | 0.142 | 0.420 |
69
+ | base, hybrid | 0.168 | 0.491 |
70
+ | fine-tuned, vector only | 0.141 | 0.456 |
71
+ | **fine-tuned, hybrid** | **0.262** | **0.561** |
72
+
73
+ The tuned model does not rank better on its own — it surfaces *different*
74
+ relevant chunks than BM25 does, and reciprocal rank fusion compounds two
75
+ rankings that disagree. Two independent fine-tunes on disjoint training sets
76
+ reproduced the same relative gain (+56% and +55%) and the same flat-vector
77
+ signature.
78
+
79
+ **Use this model fused with a lexical channel.** On its own it is roughly the
80
+ base model.
81
+
82
+ ## Training
83
+
84
+ - **Data:** CORE-Bench Level-2, **SWE-Bench-plus-plus split only**.
85
+ - **Contamination control:** that split shares **zero repositories** with the
86
+ evaluation set above. The split is a checked-in contract, not a convention.
87
+ - **Training pairs:** 1,270 (query, positive, hard-negatives) rows.
88
+ - **Objective:** MultipleNegativesRankingLoss (sentence-transformers), 4
89
+ BM25-mined hard negatives per row plus in-batch negatives.
90
+ - **Hyperparameters:** 3 epochs, batch size 8, learning rate 2e-5, warmup ratio
91
+ 0.1, bf16, max sequence length 512.
92
+ - **Hardware / time:** one RTX 3060 (12 GB), 8,959 s of training (~2.5 hours).
93
+
94
+ ## Limitations
95
+
96
+ **It does not transfer to short developer queries.** This is the finding that
97
+ kept it out of production. On a 144-case internal corpus of short, intent-phrased
98
+ developer questions ("where is the decision made to split a range based on
99
+ load"), evaluated with enriched indexes:
100
+
101
+ - 86-case dev split: this model **trails** the base — Hit@3 0.79 vs 0.85,
102
+ Recall@5 0.84 vs 0.88, consistently across projects and slices.
103
+ - 58-case holdout: parity — Hit@3 0.90 for both.
104
+
105
+ A 30-case pilot had shown no regression; that did not replicate at full size.
106
+ Issue-style training does not generalise downward to short queries.
107
+
108
+ **Recall is not what improved.** NDCG@10 moves; Recall@100 stays below the 7B
109
+ baseline. If your bottleneck is reach rather than ordering, this will not fix it.
110
+
111
+ **Evaluated on one benchmark family.** All numbers above are CORE-Bench
112
+ Level-2. No claim is made about docstring-to-function retrieval, cross-language
113
+ behaviour, or natural-language code search generally.
114
+
115
+ **English-and-mainstream bias, inherited.** On the SWE-bench_Multilingual split
116
+ the base encoder scores roughly half what it does on the English splits
117
+ (NDCG@10 0.0557 against 0.1088 / 0.1229). Fine-tuning does not repair that.
118
+
119
+ ## Intended use
120
+
121
+ Research and reproduction: issue-to-edit localization, retrieval-fusion
122
+ experiments, and as a size-class baseline for small code encoders.
123
+
124
+ **Out of scope:** commercial use (see License), and any deployment where short
125
+ queries dominate — use the Apache-2.0 base model there instead.
126
+
127
+ ## License and provenance
128
+
129
+ **This model is released under CC BY-NC-SA 4.0.**
130
+
131
+ - The base model, `jinaai/jina-embeddings-v2-base-code`, is **Apache 2.0**, and
132
+ its notices are preserved. The custom modelling code bundled with this
133
+ checkpoint (`modeling_bert.py`, `configuration_bert.py`) originates there and
134
+ remains under that license.
135
+ - The training data, `zhangfw123/CORE-Bench`, is **CC BY-NC-SA 4.0** —
136
+ NonCommercial and ShareAlike.
137
+
138
+ Whether model weights constitute a derivative work of their training data is
139
+ unsettled: Creative Commons licenses predate machine-learning training, and CC
140
+ has said as much itself. Rather than bet on the permissive reading, this model
141
+ adopts the dataset's own terms. That satisfies ShareAlike if it applies and
142
+ honours the NonCommercial intent if it does not.
143
+
144
+ **Practical consequence: do not use these weights in a commercial product.**
145
+ If you want a commercially usable code embedder, use the Apache-2.0 base model
146
+ directly — it is what
147
+ [Contextmaxxer](https://github.com/codeus-morbid/contextmaxxer), the tool this
148
+ work came out of, actually ships.
149
+
150
+ No CORE-Bench corpus text is redistributed in this repository.
151
+
152
+ This is a licensing summary, not legal advice.
153
+
154
+ ## Citation
155
+
156
+ The benchmark and the baselines it supplies:
157
+
158
+ ```bibtex
159
+ @article{corebench2026,
160
+ title = {CORE-Bench: A Comprehensive Benchmark for Code Retrieval in the Era of Agentic Coding},
161
+ author = {Zhang, Fuwei and Zhang, Yanzhao and Li, Mingxin and others},
162
+ journal = {arXiv preprint arXiv:2606.11864},
163
+ year = {2026}
164
+ }
165
+ ```
166
+
167
+ Full evaluation protocol, the negative results, and the reasoning behind the
168
+ production default are in
169
+ [BENCHMARK.md](https://github.com/codeus-morbid/contextmaxxer/blob/main/BENCHMARK.md).