|
Download README.md from RowletQwQ/JevCodeLocator-0.6B: direct link, hf CLI and curl.
- Browser
- Download file 8.47 kB
-
https://huggingface.co/RowletQwQ/JevCodeLocator-0.6B/resolve/main/README.md
- Command line
-
hf download hf://RowletQwQ/JevCodeLocator-0.6B/README.md
-
curl -L -o README.md https://huggingface.co/RowletQwQ/JevCodeLocator-0.6B/resolve/main/README.md
8.47 kB
| language: | |
| - en | |
| - zh | |
| library_name: pytorch | |
| license: apache-2.0 | |
| base_model: Qwen/Qwen3-0.6B | |
| tags: | |
| - code-localization | |
| - code-search | |
| - code-retrieval | |
| - reranker | |
| - qwen3 | |
| - set-attention | |
| pipeline_tag: text-ranking | |
| # JevCodeLocator-0.6B | |
| A 0.6B **code localizer**. Given a repository and a natural-language query (an issue, a symptom, a | |
| symbol, an error message, or a described behavior), it reranks candidate code locations and returns a | |
| ranked list of `path:start-end` chunks together with a file-level ranking. | |
| The model is a **reranker**, not a retriever: a cheap sparse retriever (BM25 plus symbol/path channels) | |
| produces a shortlist of K candidates and the model scores all of them in a single forward pass. There | |
| are no embeddings and no vector index. On an entry-level 6 GB GPU it runs in roughly 0.4 s at K=8 and | |
| 0.9 s at K=16, which makes it usable as a local first-stage localizer for coding agents. | |
| ## Architecture | |
| ``` | |
| state (task / repo / candidate_count / query) + candidate texts ("path:start-end: summary") | |
| β | |
| ββ one leaf-token path per candidate βββΊ Qwen3-0.6B backbone | |
| β (last hidden state at the EOS of each path) | |
| ββ LayerNorm β Linear(d, 1) per-candidate logit z_i | |
| ββ set-attention over the K candidates: | |
| u = Linear(d + 1, 128)([h_i ; log K]) | |
| mixed = MultiheadAttention(u, u, u) (key-padding mask for empty slots) | |
| z_i += Linear(128, 1)(tanh(u + mixed)) | |
| β | |
| ββ softmax over the K candidates | |
| ``` | |
| The scalar head alone can rank candidates independently; the zero-initialized set-attention | |
| correction is what makes a candidate's score depend on the **whole shortlist** (for example, | |
| suppressing a cluster of look-alike test files in favour of the implementation file). | |
| ## Training | |
| Three consecutive fine-tuning stages on top of `Qwen/Qwen3-0.6B`: | |
| | stage | training data | steps | notes | | |
| |---|---|---:|---| | |
| | 1 | Python programmatic queries, Kβ€16, file-level auxiliary loss Ξ»=0.30 | 600 | base localizer | | |
| | 2 | + Go / TypeScript-JavaScript / C++ / Rust / C# / PHP programmatic queries; 12,274 rows, K=16 BM25 pools with the gold candidate forced in | 600 | doc-comment queries oversampled 2Γ, backbone LR 1e-5 | | |
| | 3 | continuation of stage 2 on 18,688 rows | 600 | final checkpoint | | |
| - Objective: cross-entropy against the gold candidate (`gold_distribution`) plus the file-level | |
| auxiliary loss (weight 0.30). | |
| - Candidate text is identical at training and serving time: a compact summary consisting of the path, | |
| the signature and the first lines of the chunk. | |
| - A Python replay set (2,286 rows) is kept in every round so that Python behaviour does not regress. | |
| - Training queries are generated programmatically from doc comments, symbol names and constants. | |
| No human-written query set was used for training. | |
| ## Evaluation | |
| Frozen held-out pools: 200 queries per repository, K=8, gold candidate forced into the pool. | |
| `file@1` is the share of queries whose gold **file** is ranked first. The baseline is the same model | |
| before the multilingual stages (stage 1 only). | |
| | evaluation set | language | baseline | **this model** | | |
| |---|---|---:|---:| | |
| | service repository, ~1.3k files | Go | 0.700 | **0.945** | | |
| | ML framework core, ~2k chunks | C++ | 0.730 | **0.870** | | |
| | agent CLI, ~20k chunks | Rust | 0.900 | 0.925 | | |
| | API service, ~5k chunks | TypeScript | 0.795 | 0.855 | | |
| | language server, ~600 chunks | C++ | 0.974 | 0.989 | | |
| | held-out dev set, 460 queries | Python | 0.839 | 0.843 | | |
| Restricted to semantic (doc-comment) queries, `file@1`: Go 0.524 β **0.913**, C++ 0.629 β 0.814, | |
| TypeScript 0.667 β 0.758, Rust 0.833 β 0.875. | |
| On a set of five hand-checked, real reverse-proxy questions about a large Go service (K=16, FP32): | |
| **gold-file recall@1 = 1.000, recall@3 = 1.000, recall@5 = 1.000.** Latency on an entry-level 6 GB | |
| GPU: 0.40β0.44 s (K=8), ~0.83 s (K=16). | |
| ## Usage | |
| ```bash | |
| pip install torch transformers safetensors | |
| python inference_example.py # runs on this directory | |
| ``` | |
| ```python | |
| import torch | |
| from modeling_jev import load_jev_model, score_candidates | |
| model, tok = load_jev_model(".", device="cuda" if torch.cuda.is_available() else "cpu", | |
| dtype=torch.float32) | |
| candidates = { | |
| "src/auth/session.py:41-88": "def create_session(user, password) | validates credentials ...", | |
| "src/db/pool.py:12-60": "def connect(dsn, max_size) | opens the database pool ...", | |
| } | |
| for key, prob in score_candidates(model, tok, "where are credentials validated?", candidates): | |
| print(f"{prob:.4f} {key}") | |
| ``` | |
| `score_candidates` returns the candidates ranked by probability. Summing the probabilities of all | |
| candidates that belong to the same file gives the file-level ranking. | |
| ## Files | |
| | file | contents | | |
| |---|---| | |
| | `model.safetensors` | full model: 322 tensors (`backbone.*` plus the decision head); sha256 `fb06ba6ac6d30c395912ebd156723a0739afdd260f9b44f85c56b074c68b8a98` | | |
| | `config.json` | model configuration (`architectures: JevDecisionModel`) and training hyperparameters | | |
| | `backbone_config/config.json` | Qwen3-0.6B backbone configuration | | |
| | `tokenizer/` | tokenizer files (`tokenizer_config.json` is in the `transformers>=5` format) | | |
| | `head.safetensors` | the 12 decision-head tensors only (803 KB), for custom runtimes | | |
| | `modeling_jev.py` | self-contained model class, loader and scoring helper | | |
| | `inference_example.py` | runnable reranking example | | |
| | `patch_tokenizer_for_transformers4.py` | converts `tokenizer_config.json` to the 4.x format | | |
| | `LICENSE`, `NOTICE` | Apache-2.0 and attribution | | |
| ## Intended use and limitations | |
| - **Use it as a reranker.** It selects among the candidates it is given; if the retriever's shortlist | |
| does not contain the gold file, the model cannot recover it. Pool recall is the ceiling. | |
| - Trained on programmatically generated queries (doc comments, symbols, constants) with a small amount | |
| of issue-style data mixed in. Free-form questions work but are the hardest distribution. | |
| - Language coverage: Python, Go, TypeScript/JavaScript, C, C++, Rust, C#, PHP, Java, Ruby, Bash. The | |
| chunker used to build candidates is tree-sitter based; other languages are untested. | |
| - Candidate text is a compact summary, so very long functions are effectively truncated. | |
| - Not instruction-tuned, not a chat model, no safety tuning. Do not use it to make decisions about | |
| code it cannot see. | |
| ## Training data and licensing | |
| The weights in this repository are released under **Apache-2.0**, the same license as the base model | |
| [Qwen/Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B) (Apache-2.0, Copyright 2025 Alibaba Cloud). | |
| The inference code here (`modeling_jev.py`, `inference_example.py`, | |
| `patch_tokenizer_for_transformers4.py`) is original work under the same license; see `LICENSE` and | |
| `NOTICE`. | |
| Training rows are generated programmatically: a tree-sitter chunker extracts code locations, queries | |
| are derived from doc comments, symbol names and constants, and the candidate text is a compact summary | |
| of the chunk. **No source file from any repository is redistributed in this repository.** | |
| Licenses of the source repositories used for training and evaluation data: | |
| | license | sources | | |
| |---|---| | |
| | MIT | 6 repositories (Go / TypeScript / Python tooling) | | |
| | Apache-2.0 | 3 repositories (C++ / Go / Rust) | | |
| | GPL-3.0 | 2 repositories | | |
| | LGPL-3.0 | 2 repositories | | |
| | AGPL-3.0 | 1 repository | | |
| | non-commercial custom license | 1 repository | | |
| | license not declared in the checkout used | 6 repositories | | |
| A detailed list of the source repositories is available on request. | |
| Notes: | |
| - The weights are not a redistribution of those repositories, and the model does not reproduce training | |
| code verbatim: it only scores candidate locations supplied by the caller. Whether copyleft-licensed | |
| training data extends to model weights is legally unsettled in most jurisdictions. | |
| - **Commercial use**: the training mix includes one AGPL-3.0 project and one non-commercial project. For | |
| a commercially clean release, retrain without those two (and preferably without the GPL/LGPL | |
| projects); the data pipeline is deterministic and reproducible. | |
| - You remain responsible for complying with the licenses of the code you run the model on. | |
| - This section is provenance information, not legal advice. | |