File size: 8,465 Bytes
3cda059 cfb6957 3cda059 cfb6957 3cda059 cfb6957 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 | ---
language:
- en
- zh
library_name: pytorch
license: apache-2.0
base_model: Qwen/Qwen3-0.6B
tags:
- code-localization
- code-search
- code-retrieval
- reranker
- qwen3
- set-attention
pipeline_tag: text-ranking
---
# JevCodeLocator-0.6B
A 0.6B **code localizer**. Given a repository and a natural-language query (an issue, a symptom, a
symbol, an error message, or a described behavior), it reranks candidate code locations and returns a
ranked list of `path:start-end` chunks together with a file-level ranking.
The model is a **reranker**, not a retriever: a cheap sparse retriever (BM25 plus symbol/path channels)
produces a shortlist of K candidates and the model scores all of them in a single forward pass. There
are no embeddings and no vector index. On an entry-level 6 GB GPU it runs in roughly 0.4 s at K=8 and
0.9 s at K=16, which makes it usable as a local first-stage localizer for coding agents.
## Architecture
```
state (task / repo / candidate_count / query) + candidate texts ("path:start-end: summary")
β
ββ one leaf-token path per candidate βββΊ Qwen3-0.6B backbone
β (last hidden state at the EOS of each path)
ββ LayerNorm β Linear(d, 1) per-candidate logit z_i
ββ set-attention over the K candidates:
u = Linear(d + 1, 128)([h_i ; log K])
mixed = MultiheadAttention(u, u, u) (key-padding mask for empty slots)
z_i += Linear(128, 1)(tanh(u + mixed))
β
ββ softmax over the K candidates
```
The scalar head alone can rank candidates independently; the zero-initialized set-attention
correction is what makes a candidate's score depend on the **whole shortlist** (for example,
suppressing a cluster of look-alike test files in favour of the implementation file).
## Training
Three consecutive fine-tuning stages on top of `Qwen/Qwen3-0.6B`:
| stage | training data | steps | notes |
|---|---|---:|---|
| 1 | Python programmatic queries, Kβ€16, file-level auxiliary loss Ξ»=0.30 | 600 | base localizer |
| 2 | + Go / TypeScript-JavaScript / C++ / Rust / C# / PHP programmatic queries; 12,274 rows, K=16 BM25 pools with the gold candidate forced in | 600 | doc-comment queries oversampled 2Γ, backbone LR 1e-5 |
| 3 | continuation of stage 2 on 18,688 rows | 600 | final checkpoint |
- Objective: cross-entropy against the gold candidate (`gold_distribution`) plus the file-level
auxiliary loss (weight 0.30).
- Candidate text is identical at training and serving time: a compact summary consisting of the path,
the signature and the first lines of the chunk.
- A Python replay set (2,286 rows) is kept in every round so that Python behaviour does not regress.
- Training queries are generated programmatically from doc comments, symbol names and constants.
No human-written query set was used for training.
## Evaluation
Frozen held-out pools: 200 queries per repository, K=8, gold candidate forced into the pool.
`file@1` is the share of queries whose gold **file** is ranked first. The baseline is the same model
before the multilingual stages (stage 1 only).
| evaluation set | language | baseline | **this model** |
|---|---|---:|---:|
| service repository, ~1.3k files | Go | 0.700 | **0.945** |
| ML framework core, ~2k chunks | C++ | 0.730 | **0.870** |
| agent CLI, ~20k chunks | Rust | 0.900 | 0.925 |
| API service, ~5k chunks | TypeScript | 0.795 | 0.855 |
| language server, ~600 chunks | C++ | 0.974 | 0.989 |
| held-out dev set, 460 queries | Python | 0.839 | 0.843 |
Restricted to semantic (doc-comment) queries, `file@1`: Go 0.524 β **0.913**, C++ 0.629 β 0.814,
TypeScript 0.667 β 0.758, Rust 0.833 β 0.875.
On a set of five hand-checked, real reverse-proxy questions about a large Go service (K=16, FP32):
**gold-file recall@1 = 1.000, recall@3 = 1.000, recall@5 = 1.000.** Latency on an entry-level 6 GB
GPU: 0.40β0.44 s (K=8), ~0.83 s (K=16).
## Usage
```bash
pip install torch transformers safetensors
python inference_example.py # runs on this directory
```
```python
import torch
from modeling_jev import load_jev_model, score_candidates
model, tok = load_jev_model(".", device="cuda" if torch.cuda.is_available() else "cpu",
dtype=torch.float32)
candidates = {
"src/auth/session.py:41-88": "def create_session(user, password) | validates credentials ...",
"src/db/pool.py:12-60": "def connect(dsn, max_size) | opens the database pool ...",
}
for key, prob in score_candidates(model, tok, "where are credentials validated?", candidates):
print(f"{prob:.4f} {key}")
```
`score_candidates` returns the candidates ranked by probability. Summing the probabilities of all
candidates that belong to the same file gives the file-level ranking.
## Files
| file | contents |
|---|---|
| `model.safetensors` | full model: 322 tensors (`backbone.*` plus the decision head); sha256 `fb06ba6ac6d30c395912ebd156723a0739afdd260f9b44f85c56b074c68b8a98` |
| `config.json` | model configuration (`architectures: JevDecisionModel`) and training hyperparameters |
| `backbone_config/config.json` | Qwen3-0.6B backbone configuration |
| `tokenizer/` | tokenizer files (`tokenizer_config.json` is in the `transformers>=5` format) |
| `head.safetensors` | the 12 decision-head tensors only (803 KB), for custom runtimes |
| `modeling_jev.py` | self-contained model class, loader and scoring helper |
| `inference_example.py` | runnable reranking example |
| `patch_tokenizer_for_transformers4.py` | converts `tokenizer_config.json` to the 4.x format |
| `LICENSE`, `NOTICE` | Apache-2.0 and attribution |
## Intended use and limitations
- **Use it as a reranker.** It selects among the candidates it is given; if the retriever's shortlist
does not contain the gold file, the model cannot recover it. Pool recall is the ceiling.
- Trained on programmatically generated queries (doc comments, symbols, constants) with a small amount
of issue-style data mixed in. Free-form questions work but are the hardest distribution.
- Language coverage: Python, Go, TypeScript/JavaScript, C, C++, Rust, C#, PHP, Java, Ruby, Bash. The
chunker used to build candidates is tree-sitter based; other languages are untested.
- Candidate text is a compact summary, so very long functions are effectively truncated.
- Not instruction-tuned, not a chat model, no safety tuning. Do not use it to make decisions about
code it cannot see.
## Training data and licensing
The weights in this repository are released under **Apache-2.0**, the same license as the base model
[Qwen/Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B) (Apache-2.0, Copyright 2025 Alibaba Cloud).
The inference code here (`modeling_jev.py`, `inference_example.py`,
`patch_tokenizer_for_transformers4.py`) is original work under the same license; see `LICENSE` and
`NOTICE`.
Training rows are generated programmatically: a tree-sitter chunker extracts code locations, queries
are derived from doc comments, symbol names and constants, and the candidate text is a compact summary
of the chunk. **No source file from any repository is redistributed in this repository.**
Licenses of the source repositories used for training and evaluation data:
| license | sources |
|---|---|
| MIT | 6 repositories (Go / TypeScript / Python tooling) |
| Apache-2.0 | 3 repositories (C++ / Go / Rust) |
| GPL-3.0 | 2 repositories |
| LGPL-3.0 | 2 repositories |
| AGPL-3.0 | 1 repository |
| non-commercial custom license | 1 repository |
| license not declared in the checkout used | 6 repositories |
A detailed list of the source repositories is available on request.
Notes:
- The weights are not a redistribution of those repositories, and the model does not reproduce training
code verbatim: it only scores candidate locations supplied by the caller. Whether copyleft-licensed
training data extends to model weights is legally unsettled in most jurisdictions.
- **Commercial use**: the training mix includes one AGPL-3.0 project and one non-commercial project. For
a commercially clean release, retrain without those two (and preferably without the GPL/LGPL
projects); the data pipeline is deterministic and reproducible.
- You remain responsible for complying with the licenses of the code you run the model on.
- This section is provenance information, not legal advice.
|