File size: 8,465 Bytes
3cda059
cfb6957
 
 
 
3cda059
cfb6957
 
 
 
 
 
 
 
 
3cda059
cfb6957
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
---
language:
  - en
  - zh
library_name: pytorch
license: apache-2.0
base_model: Qwen/Qwen3-0.6B
tags:
  - code-localization
  - code-search
  - code-retrieval
  - reranker
  - qwen3
  - set-attention
pipeline_tag: text-ranking
---

# JevCodeLocator-0.6B

A 0.6B **code localizer**. Given a repository and a natural-language query (an issue, a symptom, a
symbol, an error message, or a described behavior), it reranks candidate code locations and returns a
ranked list of `path:start-end` chunks together with a file-level ranking.

The model is a **reranker**, not a retriever: a cheap sparse retriever (BM25 plus symbol/path channels)
produces a shortlist of K candidates and the model scores all of them in a single forward pass. There
are no embeddings and no vector index. On an entry-level 6 GB GPU it runs in roughly 0.4 s at K=8 and
0.9 s at K=16, which makes it usable as a local first-stage localizer for coding agents.

## Architecture

```
state (task / repo / candidate_count / query)  +  candidate texts ("path:start-end: summary")
        β”‚
        β”œβ”€ one leaf-token path per candidate  ──►  Qwen3-0.6B backbone
        β”‚                                          (last hidden state at the EOS of each path)
        β”œβ”€ LayerNorm β†’ Linear(d, 1)                    per-candidate logit z_i
        └─ set-attention over the K candidates:
             u     = Linear(d + 1, 128)([h_i ; log K])
             mixed = MultiheadAttention(u, u, u)       (key-padding mask for empty slots)
             z_i  += Linear(128, 1)(tanh(u + mixed))
        β”‚
        └─ softmax over the K candidates
```

The scalar head alone can rank candidates independently; the zero-initialized set-attention
correction is what makes a candidate's score depend on the **whole shortlist** (for example,
suppressing a cluster of look-alike test files in favour of the implementation file).

## Training

Three consecutive fine-tuning stages on top of `Qwen/Qwen3-0.6B`:

| stage | training data | steps | notes |
|---|---|---:|---|
| 1 | Python programmatic queries, K≀16, file-level auxiliary loss Ξ»=0.30 | 600 | base localizer |
| 2 | + Go / TypeScript-JavaScript / C++ / Rust / C# / PHP programmatic queries; 12,274 rows, K=16 BM25 pools with the gold candidate forced in | 600 | doc-comment queries oversampled 2Γ—, backbone LR 1e-5 |
| 3 | continuation of stage 2 on 18,688 rows | 600 | final checkpoint |

- Objective: cross-entropy against the gold candidate (`gold_distribution`) plus the file-level
  auxiliary loss (weight 0.30).
- Candidate text is identical at training and serving time: a compact summary consisting of the path,
  the signature and the first lines of the chunk.
- A Python replay set (2,286 rows) is kept in every round so that Python behaviour does not regress.
- Training queries are generated programmatically from doc comments, symbol names and constants.
  No human-written query set was used for training.

## Evaluation

Frozen held-out pools: 200 queries per repository, K=8, gold candidate forced into the pool.
`file@1` is the share of queries whose gold **file** is ranked first. The baseline is the same model
before the multilingual stages (stage 1 only).

| evaluation set | language | baseline | **this model** |
|---|---|---:|---:|
| service repository, ~1.3k files | Go | 0.700 | **0.945** |
| ML framework core, ~2k chunks | C++ | 0.730 | **0.870** |
| agent CLI, ~20k chunks | Rust | 0.900 | 0.925 |
| API service, ~5k chunks | TypeScript | 0.795 | 0.855 |
| language server, ~600 chunks | C++ | 0.974 | 0.989 |
| held-out dev set, 460 queries | Python | 0.839 | 0.843 |

Restricted to semantic (doc-comment) queries, `file@1`: Go 0.524 β†’ **0.913**, C++ 0.629 β†’ 0.814,
TypeScript 0.667 β†’ 0.758, Rust 0.833 β†’ 0.875.

On a set of five hand-checked, real reverse-proxy questions about a large Go service (K=16, FP32):
**gold-file recall@1 = 1.000, recall@3 = 1.000, recall@5 = 1.000.** Latency on an entry-level 6 GB
GPU: 0.40–0.44 s (K=8), ~0.83 s (K=16).

## Usage

```bash
pip install torch transformers safetensors
python inference_example.py           # runs on this directory
```

```python
import torch

from modeling_jev import load_jev_model, score_candidates

model, tok = load_jev_model(".", device="cuda" if torch.cuda.is_available() else "cpu",
                            dtype=torch.float32)

candidates = {
    "src/auth/session.py:41-88": "def create_session(user, password) | validates credentials ...",
    "src/db/pool.py:12-60": "def connect(dsn, max_size) | opens the database pool ...",
}

for key, prob in score_candidates(model, tok, "where are credentials validated?", candidates):
    print(f"{prob:.4f}  {key}")
```

`score_candidates` returns the candidates ranked by probability. Summing the probabilities of all
candidates that belong to the same file gives the file-level ranking.

## Files

| file | contents |
|---|---|
| `model.safetensors` | full model: 322 tensors (`backbone.*` plus the decision head); sha256 `fb06ba6ac6d30c395912ebd156723a0739afdd260f9b44f85c56b074c68b8a98` |
| `config.json` | model configuration (`architectures: JevDecisionModel`) and training hyperparameters |
| `backbone_config/config.json` | Qwen3-0.6B backbone configuration |
| `tokenizer/` | tokenizer files (`tokenizer_config.json` is in the `transformers>=5` format) |
| `head.safetensors` | the 12 decision-head tensors only (803 KB), for custom runtimes |
| `modeling_jev.py` | self-contained model class, loader and scoring helper |
| `inference_example.py` | runnable reranking example |
| `patch_tokenizer_for_transformers4.py` | converts `tokenizer_config.json` to the 4.x format |
| `LICENSE`, `NOTICE` | Apache-2.0 and attribution |

## Intended use and limitations

- **Use it as a reranker.** It selects among the candidates it is given; if the retriever's shortlist
  does not contain the gold file, the model cannot recover it. Pool recall is the ceiling.
- Trained on programmatically generated queries (doc comments, symbols, constants) with a small amount
  of issue-style data mixed in. Free-form questions work but are the hardest distribution.
- Language coverage: Python, Go, TypeScript/JavaScript, C, C++, Rust, C#, PHP, Java, Ruby, Bash. The
  chunker used to build candidates is tree-sitter based; other languages are untested.
- Candidate text is a compact summary, so very long functions are effectively truncated.
- Not instruction-tuned, not a chat model, no safety tuning. Do not use it to make decisions about
  code it cannot see.

## Training data and licensing

The weights in this repository are released under **Apache-2.0**, the same license as the base model
[Qwen/Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B) (Apache-2.0, Copyright 2025 Alibaba Cloud).
The inference code here (`modeling_jev.py`, `inference_example.py`,
`patch_tokenizer_for_transformers4.py`) is original work under the same license; see `LICENSE` and
`NOTICE`.

Training rows are generated programmatically: a tree-sitter chunker extracts code locations, queries
are derived from doc comments, symbol names and constants, and the candidate text is a compact summary
of the chunk. **No source file from any repository is redistributed in this repository.**

Licenses of the source repositories used for training and evaluation data:

| license | sources |
|---|---|
| MIT | 6 repositories (Go / TypeScript / Python tooling) |
| Apache-2.0 | 3 repositories (C++ / Go / Rust) |
| GPL-3.0 | 2 repositories |
| LGPL-3.0 | 2 repositories |
| AGPL-3.0 | 1 repository |
| non-commercial custom license | 1 repository |
| license not declared in the checkout used | 6 repositories |

A detailed list of the source repositories is available on request.

Notes:

- The weights are not a redistribution of those repositories, and the model does not reproduce training
  code verbatim: it only scores candidate locations supplied by the caller. Whether copyleft-licensed
  training data extends to model weights is legally unsettled in most jurisdictions.
- **Commercial use**: the training mix includes one AGPL-3.0 project and one non-commercial project. For
  a commercially clean release, retrain without those two (and preferably without the GPL/LGPL
  projects); the data pipeline is deterministic and reproducible.
- You remain responsible for complying with the licenses of the code you run the model on.
- This section is provenance information, not legal advice.