JevCodeLocator-0.6B / README.md
RowletQwQ's picture
JevCodeLocator-0.6B
cfb6957 verified
|
Raw History Blame Contribute Delete
8.47 kB
---
language:
- en
- zh
library_name: pytorch
license: apache-2.0
base_model: Qwen/Qwen3-0.6B
tags:
- code-localization
- code-search
- code-retrieval
- reranker
- qwen3
- set-attention
pipeline_tag: text-ranking
---
# JevCodeLocator-0.6B
A 0.6B **code localizer**. Given a repository and a natural-language query (an issue, a symptom, a
symbol, an error message, or a described behavior), it reranks candidate code locations and returns a
ranked list of `path:start-end` chunks together with a file-level ranking.
The model is a **reranker**, not a retriever: a cheap sparse retriever (BM25 plus symbol/path channels)
produces a shortlist of K candidates and the model scores all of them in a single forward pass. There
are no embeddings and no vector index. On an entry-level 6 GB GPU it runs in roughly 0.4 s at K=8 and
0.9 s at K=16, which makes it usable as a local first-stage localizer for coding agents.
## Architecture
```
state (task / repo / candidate_count / query) + candidate texts ("path:start-end: summary")
β”‚
β”œβ”€ one leaf-token path per candidate ──► Qwen3-0.6B backbone
β”‚ (last hidden state at the EOS of each path)
β”œβ”€ LayerNorm β†’ Linear(d, 1) per-candidate logit z_i
└─ set-attention over the K candidates:
u = Linear(d + 1, 128)([h_i ; log K])
mixed = MultiheadAttention(u, u, u) (key-padding mask for empty slots)
z_i += Linear(128, 1)(tanh(u + mixed))
β”‚
└─ softmax over the K candidates
```
The scalar head alone can rank candidates independently; the zero-initialized set-attention
correction is what makes a candidate's score depend on the **whole shortlist** (for example,
suppressing a cluster of look-alike test files in favour of the implementation file).
## Training
Three consecutive fine-tuning stages on top of `Qwen/Qwen3-0.6B`:
| stage | training data | steps | notes |
|---|---|---:|---|
| 1 | Python programmatic queries, K≀16, file-level auxiliary loss Ξ»=0.30 | 600 | base localizer |
| 2 | + Go / TypeScript-JavaScript / C++ / Rust / C# / PHP programmatic queries; 12,274 rows, K=16 BM25 pools with the gold candidate forced in | 600 | doc-comment queries oversampled 2Γ—, backbone LR 1e-5 |
| 3 | continuation of stage 2 on 18,688 rows | 600 | final checkpoint |
- Objective: cross-entropy against the gold candidate (`gold_distribution`) plus the file-level
auxiliary loss (weight 0.30).
- Candidate text is identical at training and serving time: a compact summary consisting of the path,
the signature and the first lines of the chunk.
- A Python replay set (2,286 rows) is kept in every round so that Python behaviour does not regress.
- Training queries are generated programmatically from doc comments, symbol names and constants.
No human-written query set was used for training.
## Evaluation
Frozen held-out pools: 200 queries per repository, K=8, gold candidate forced into the pool.
`file@1` is the share of queries whose gold **file** is ranked first. The baseline is the same model
before the multilingual stages (stage 1 only).
| evaluation set | language | baseline | **this model** |
|---|---|---:|---:|
| service repository, ~1.3k files | Go | 0.700 | **0.945** |
| ML framework core, ~2k chunks | C++ | 0.730 | **0.870** |
| agent CLI, ~20k chunks | Rust | 0.900 | 0.925 |
| API service, ~5k chunks | TypeScript | 0.795 | 0.855 |
| language server, ~600 chunks | C++ | 0.974 | 0.989 |
| held-out dev set, 460 queries | Python | 0.839 | 0.843 |
Restricted to semantic (doc-comment) queries, `file@1`: Go 0.524 β†’ **0.913**, C++ 0.629 β†’ 0.814,
TypeScript 0.667 β†’ 0.758, Rust 0.833 β†’ 0.875.
On a set of five hand-checked, real reverse-proxy questions about a large Go service (K=16, FP32):
**gold-file recall@1 = 1.000, recall@3 = 1.000, recall@5 = 1.000.** Latency on an entry-level 6 GB
GPU: 0.40–0.44 s (K=8), ~0.83 s (K=16).
## Usage
```bash
pip install torch transformers safetensors
python inference_example.py # runs on this directory
```
```python
import torch
from modeling_jev import load_jev_model, score_candidates
model, tok = load_jev_model(".", device="cuda" if torch.cuda.is_available() else "cpu",
dtype=torch.float32)
candidates = {
"src/auth/session.py:41-88": "def create_session(user, password) | validates credentials ...",
"src/db/pool.py:12-60": "def connect(dsn, max_size) | opens the database pool ...",
}
for key, prob in score_candidates(model, tok, "where are credentials validated?", candidates):
print(f"{prob:.4f} {key}")
```
`score_candidates` returns the candidates ranked by probability. Summing the probabilities of all
candidates that belong to the same file gives the file-level ranking.
## Files
| file | contents |
|---|---|
| `model.safetensors` | full model: 322 tensors (`backbone.*` plus the decision head); sha256 `fb06ba6ac6d30c395912ebd156723a0739afdd260f9b44f85c56b074c68b8a98` |
| `config.json` | model configuration (`architectures: JevDecisionModel`) and training hyperparameters |
| `backbone_config/config.json` | Qwen3-0.6B backbone configuration |
| `tokenizer/` | tokenizer files (`tokenizer_config.json` is in the `transformers>=5` format) |
| `head.safetensors` | the 12 decision-head tensors only (803 KB), for custom runtimes |
| `modeling_jev.py` | self-contained model class, loader and scoring helper |
| `inference_example.py` | runnable reranking example |
| `patch_tokenizer_for_transformers4.py` | converts `tokenizer_config.json` to the 4.x format |
| `LICENSE`, `NOTICE` | Apache-2.0 and attribution |
## Intended use and limitations
- **Use it as a reranker.** It selects among the candidates it is given; if the retriever's shortlist
does not contain the gold file, the model cannot recover it. Pool recall is the ceiling.
- Trained on programmatically generated queries (doc comments, symbols, constants) with a small amount
of issue-style data mixed in. Free-form questions work but are the hardest distribution.
- Language coverage: Python, Go, TypeScript/JavaScript, C, C++, Rust, C#, PHP, Java, Ruby, Bash. The
chunker used to build candidates is tree-sitter based; other languages are untested.
- Candidate text is a compact summary, so very long functions are effectively truncated.
- Not instruction-tuned, not a chat model, no safety tuning. Do not use it to make decisions about
code it cannot see.
## Training data and licensing
The weights in this repository are released under **Apache-2.0**, the same license as the base model
[Qwen/Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B) (Apache-2.0, Copyright 2025 Alibaba Cloud).
The inference code here (`modeling_jev.py`, `inference_example.py`,
`patch_tokenizer_for_transformers4.py`) is original work under the same license; see `LICENSE` and
`NOTICE`.
Training rows are generated programmatically: a tree-sitter chunker extracts code locations, queries
are derived from doc comments, symbol names and constants, and the candidate text is a compact summary
of the chunk. **No source file from any repository is redistributed in this repository.**
Licenses of the source repositories used for training and evaluation data:
| license | sources |
|---|---|
| MIT | 6 repositories (Go / TypeScript / Python tooling) |
| Apache-2.0 | 3 repositories (C++ / Go / Rust) |
| GPL-3.0 | 2 repositories |
| LGPL-3.0 | 2 repositories |
| AGPL-3.0 | 1 repository |
| non-commercial custom license | 1 repository |
| license not declared in the checkout used | 6 repositories |
A detailed list of the source repositories is available on request.
Notes:
- The weights are not a redistribution of those repositories, and the model does not reproduce training
code verbatim: it only scores candidate locations supplied by the caller. Whether copyleft-licensed
training data extends to model weights is legally unsettled in most jurisdictions.
- **Commercial use**: the training mix includes one AGPL-3.0 project and one non-commercial project. For
a commercially clean release, retrain without those two (and preferably without the GPL/LGPL
projects); the data pipeline is deterministic and reproducible.
- You remain responsible for complying with the licenses of the code you run the model on.
- This section is provenance information, not legal advice.