--- language: - en - zh library_name: pytorch license: apache-2.0 base_model: Qwen/Qwen3-0.6B tags: - code-localization - code-search - code-retrieval - reranker - qwen3 - set-attention pipeline_tag: text-ranking --- # JevCodeLocator-0.6B A 0.6B **code localizer**. Given a repository and a natural-language query (an issue, a symptom, a symbol, an error message, or a described behavior), it reranks candidate code locations and returns a ranked list of `path:start-end` chunks together with a file-level ranking. The model is a **reranker**, not a retriever: a cheap sparse retriever (BM25 plus symbol/path channels) produces a shortlist of K candidates and the model scores all of them in a single forward pass. There are no embeddings and no vector index. On an entry-level 6 GB GPU it runs in roughly 0.4 s at K=8 and 0.9 s at K=16, which makes it usable as a local first-stage localizer for coding agents. ## Architecture ``` state (task / repo / candidate_count / query) + candidate texts ("path:start-end: summary") │ ├─ one leaf-token path per candidate ──► Qwen3-0.6B backbone │ (last hidden state at the EOS of each path) ├─ LayerNorm → Linear(d, 1) per-candidate logit z_i └─ set-attention over the K candidates: u = Linear(d + 1, 128)([h_i ; log K]) mixed = MultiheadAttention(u, u, u) (key-padding mask for empty slots) z_i += Linear(128, 1)(tanh(u + mixed)) │ └─ softmax over the K candidates ``` The scalar head alone can rank candidates independently; the zero-initialized set-attention correction is what makes a candidate's score depend on the **whole shortlist** (for example, suppressing a cluster of look-alike test files in favour of the implementation file). ## Training Three consecutive fine-tuning stages on top of `Qwen/Qwen3-0.6B`: | stage | training data | steps | notes | |---|---|---:|---| | 1 | Python programmatic queries, K≤16, file-level auxiliary loss λ=0.30 | 600 | base localizer | | 2 | + Go / TypeScript-JavaScript / C++ / Rust / C# / PHP programmatic queries; 12,274 rows, K=16 BM25 pools with the gold candidate forced in | 600 | doc-comment queries oversampled 2×, backbone LR 1e-5 | | 3 | continuation of stage 2 on 18,688 rows | 600 | final checkpoint | - Objective: cross-entropy against the gold candidate (`gold_distribution`) plus the file-level auxiliary loss (weight 0.30). - Candidate text is identical at training and serving time: a compact summary consisting of the path, the signature and the first lines of the chunk. - A Python replay set (2,286 rows) is kept in every round so that Python behaviour does not regress. - Training queries are generated programmatically from doc comments, symbol names and constants. No human-written query set was used for training. ## Evaluation Frozen held-out pools: 200 queries per repository, K=8, gold candidate forced into the pool. `file@1` is the share of queries whose gold **file** is ranked first. The baseline is the same model before the multilingual stages (stage 1 only). | evaluation set | language | baseline | **this model** | |---|---|---:|---:| | service repository, ~1.3k files | Go | 0.700 | **0.945** | | ML framework core, ~2k chunks | C++ | 0.730 | **0.870** | | agent CLI, ~20k chunks | Rust | 0.900 | 0.925 | | API service, ~5k chunks | TypeScript | 0.795 | 0.855 | | language server, ~600 chunks | C++ | 0.974 | 0.989 | | held-out dev set, 460 queries | Python | 0.839 | 0.843 | Restricted to semantic (doc-comment) queries, `file@1`: Go 0.524 → **0.913**, C++ 0.629 → 0.814, TypeScript 0.667 → 0.758, Rust 0.833 → 0.875. On a set of five hand-checked, real reverse-proxy questions about a large Go service (K=16, FP32): **gold-file recall@1 = 1.000, recall@3 = 1.000, recall@5 = 1.000.** Latency on an entry-level 6 GB GPU: 0.40–0.44 s (K=8), ~0.83 s (K=16). ## Usage ```bash pip install torch transformers safetensors python inference_example.py # runs on this directory ``` ```python import torch from modeling_jev import load_jev_model, score_candidates model, tok = load_jev_model(".", device="cuda" if torch.cuda.is_available() else "cpu", dtype=torch.float32) candidates = { "src/auth/session.py:41-88": "def create_session(user, password) | validates credentials ...", "src/db/pool.py:12-60": "def connect(dsn, max_size) | opens the database pool ...", } for key, prob in score_candidates(model, tok, "where are credentials validated?", candidates): print(f"{prob:.4f} {key}") ``` `score_candidates` returns the candidates ranked by probability. Summing the probabilities of all candidates that belong to the same file gives the file-level ranking. ## Files | file | contents | |---|---| | `model.safetensors` | full model: 322 tensors (`backbone.*` plus the decision head); sha256 `fb06ba6ac6d30c395912ebd156723a0739afdd260f9b44f85c56b074c68b8a98` | | `config.json` | model configuration (`architectures: JevDecisionModel`) and training hyperparameters | | `backbone_config/config.json` | Qwen3-0.6B backbone configuration | | `tokenizer/` | tokenizer files (`tokenizer_config.json` is in the `transformers>=5` format) | | `head.safetensors` | the 12 decision-head tensors only (803 KB), for custom runtimes | | `modeling_jev.py` | self-contained model class, loader and scoring helper | | `inference_example.py` | runnable reranking example | | `patch_tokenizer_for_transformers4.py` | converts `tokenizer_config.json` to the 4.x format | | `LICENSE`, `NOTICE` | Apache-2.0 and attribution | ## Intended use and limitations - **Use it as a reranker.** It selects among the candidates it is given; if the retriever's shortlist does not contain the gold file, the model cannot recover it. Pool recall is the ceiling. - Trained on programmatically generated queries (doc comments, symbols, constants) with a small amount of issue-style data mixed in. Free-form questions work but are the hardest distribution. - Language coverage: Python, Go, TypeScript/JavaScript, C, C++, Rust, C#, PHP, Java, Ruby, Bash. The chunker used to build candidates is tree-sitter based; other languages are untested. - Candidate text is a compact summary, so very long functions are effectively truncated. - Not instruction-tuned, not a chat model, no safety tuning. Do not use it to make decisions about code it cannot see. ## Training data and licensing The weights in this repository are released under **Apache-2.0**, the same license as the base model [Qwen/Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B) (Apache-2.0, Copyright 2025 Alibaba Cloud). The inference code here (`modeling_jev.py`, `inference_example.py`, `patch_tokenizer_for_transformers4.py`) is original work under the same license; see `LICENSE` and `NOTICE`. Training rows are generated programmatically: a tree-sitter chunker extracts code locations, queries are derived from doc comments, symbol names and constants, and the candidate text is a compact summary of the chunk. **No source file from any repository is redistributed in this repository.** Licenses of the source repositories used for training and evaluation data: | license | sources | |---|---| | MIT | 6 repositories (Go / TypeScript / Python tooling) | | Apache-2.0 | 3 repositories (C++ / Go / Rust) | | GPL-3.0 | 2 repositories | | LGPL-3.0 | 2 repositories | | AGPL-3.0 | 1 repository | | non-commercial custom license | 1 repository | | license not declared in the checkout used | 6 repositories | A detailed list of the source repositories is available on request. Notes: - The weights are not a redistribution of those repositories, and the model does not reproduce training code verbatim: it only scores candidate locations supplied by the caller. Whether copyleft-licensed training data extends to model weights is legally unsettled in most jurisdictions. - **Commercial use**: the training mix includes one AGPL-3.0 project and one non-commercial project. For a commercially clean release, retrain without those two (and preferably without the GPL/LGPL projects); the data pipeline is deterministic and reproducible. - You remain responsible for complying with the licenses of the code you run the model on. - This section is provenance information, not legal advice.