Text Classification
Laya
Safetensors
English
code-search
reranker
code-retrieval
calibrated
claude-code
laya-codex
Instructions to use tindang/laya-code with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Laya
How to use tindang/laya-code with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Model card
Browse files
README.md
ADDED
|
@@ -0,0 +1,158 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
base_model: convaiinnovations/laya
|
| 4 |
+
base_model_relation: finetune
|
| 5 |
+
pipeline_tag: text-classification
|
| 6 |
+
language: [en]
|
| 7 |
+
tags: [laya, code-search, reranker, code-retrieval, calibrated, claude-code, laya-codex]
|
| 8 |
+
---
|
| 9 |
+
|
| 10 |
+
# laya-code
|
| 11 |
+
|
| 12 |
+
A code-relevance re-ranker fine-tuned from [Laya](https://huggingface.co/convaiinnovations/laya)
|
| 13 |
+
(ModernBERT-large encoder + typed-decision head, 421M parameters). Given a task description and a
|
| 14 |
+
source-code chunk, it answers one yes/no (`noul`) question with a **calibrated probability**:
|
| 15 |
+
|
| 16 |
+
```
|
| 17 |
+
question: Is this source code relevant to the software change: "{task}"?
|
| 18 |
+
state: file: <path> (lines a-b)\n<code> (truncated to 128 tokens in production)
|
| 19 |
+
```
|
| 20 |
+
|
| 21 |
+
It is the default re-ranker of [laya-codex](https://github.com/pilotspace/laya-codex), which feeds
|
| 22 |
+
Claude Code the most relevant code spans for a prompt (tree-sitter chunks, then Moon BM25
|
| 23 |
+
candidates, then laya-code re-ranking).
|
| 24 |
+
|
| 25 |
+
## Model details
|
| 26 |
+
|
| 27 |
+
| | |
|
| 28 |
+
|---|---|
|
| 29 |
+
| Base model | `convaiinnovations/laya`, root checkpoint (revision `1c5edc17a7acd8701df6fc341c0d179f1c62c982`), Apache-2.0 |
|
| 30 |
+
| Architecture | unchanged: same safetensors keys, shapes and dtypes as the base (205 F16 tensors + `temperature` F32), 842,609,210 bytes |
|
| 31 |
+
| What changed | `model.safetensors` (fine-tuned weights) and `rl_agent_config.json` (`model_name`, `noul:2` temperature **0.9410**, was 1.9834; `finetune` block). `encoder/config.json`, `tokenizer/*`, `rl_agent_api.py` and `rl_common.py` are byte-identical to the base. The `training` block of `rl_agent_config.json` is inherited from the base and describes the base's training, not this fine-tune. |
|
| 32 |
+
| Context | 512 tokens (`max_len`), 192 for the question head (`head_max_len`) |
|
| 33 |
+
| Runtime | Python: `rl_agent_api.RLAgent` from this repo (same as the base). Rust: `laya-model` crate of laya-codex (candle; Metal F16 on macOS, CPU F32 elsewhere), parity-tested against the Python reference |
|
| 34 |
+
| License | Apache-2.0 (see `LICENSE` and `NOTICE`) |
|
| 35 |
+
|
| 36 |
+
## Training data
|
| 37 |
+
|
| 38 |
+
Weak supervision from git history. Nothing was hand-labelled.
|
| 39 |
+
|
| 40 |
+
- **8 training repositories** (mixed Rust, Python, TypeScript and JavaScript): openai/codex,
|
| 41 |
+
TinDang97/velos, MervinPraison/PraisonAI, badlogic/pi-mono, Portkey-AI/gateway (local
|
| 42 |
+
`ai-guard` checkout), TinDang97/python-dependency-injector, Netflix/dispatch and
|
| 43 |
+
pilotspace/hydroa (local `ai-proxy` checkout). Source: `finetune/repos.py`.
|
| 44 |
+
- **Held out** (never used for training or calibration): pilotspace/moon and pilot-space.
|
| 45 |
+
Moon client codebases (helios, helios-mono, lunaris) were left out so that Moon vocabulary
|
| 46 |
+
does not leak into the Moon eval. A root-commit check rejects forks and clones of each other
|
| 47 |
+
and of the held-out repos.
|
| 48 |
+
- **Examples**: up to 600 non-merge commits per repo (the `build_data.py` default) that touch 1–4 source files and have an
|
| 49 |
+
informative subject (at least 20 characters). The task is the subject plus a short first body
|
| 50 |
+
line. The candidate list is the BM25 top 12 over other files plus the parent-revision windows
|
| 51 |
+
of the touched files (40-line windows, stride 30). Labels: 1.0 for a window that overlaps a
|
| 52 |
+
changed hunk, 0.7 for another window of a touched file, 0 for other files. Extra examples:
|
| 53 |
+
hunks BM25 missed (label 1.0), a random same-file window (label 0.4) and random windows from
|
| 54 |
+
other files (label 0).
|
| 55 |
+
- **Size**: 50,926 pairs, split by commit hash into 45,790 train and 5,136 validation pairs.
|
| 56 |
+
Train pairs: 2,779 candidate positives, 5,609 candidate same-file, 24,768 candidate negatives,
|
| 57 |
+
4,656 extra positives, 2,392 same-file, 5,586 random negatives. Per-repo counts (v2 data;
|
| 58 |
+
the warm start used v1 data from the same repos):
|
| 59 |
+
|
| 60 |
+
| repo | train commits | val commits | train pairs | val pairs |
|
| 61 |
+
|---|---|---|---|---|
|
| 62 |
+
| PraisonAI | 231 | 18 | 3,711 | 280 |
|
| 63 |
+
| ai-guard | 460 | 40 | 7,404 | 665 |
|
| 64 |
+
| ai-proxy | 200 | 23 | 3,293 | 380 |
|
| 65 |
+
| codex | 446 | 54 | 7,676 | 932 |
|
| 66 |
+
| dispatch | 447 | 53 | 7,347 | 863 |
|
| 67 |
+
| pi-mono | 439 | 61 | 7,386 | 1,013 |
|
| 68 |
+
| python-dependency-injector | 453 | 47 | 7,066 | 744 |
|
| 69 |
+
| velos | 117 | 16 | 1,907 | 259 |
|
| 70 |
+
- **Training**: top 8 of 28 encoder layers, the final norm and the decision head. fp32 on an
|
| 71 |
+
M4 Pro (MPS). AdamW; learning rate 2e-5 for the encoder and 1e-4 for the head. 32 sequences
|
| 72 |
+
per update. Log loss against soft targets. 536 updates on v2 data, warm-started from 300
|
| 73 |
+
updates on v1 data (about 2.3 h in total). Checkpoint chosen by lowest validation NLL. The
|
| 74 |
+
`noul:2` temperature was then refitted on the validation split, using the exported F16
|
| 75 |
+
weights.
|
| 76 |
+
|
| 77 |
+
## Evaluation
|
| 78 |
+
|
| 79 |
+
All numbers come from files in the laya-codex repository and are quoted as recorded. Gold labels
|
| 80 |
+
are file-level: the files the commit touched. Candidates are BM25 windows at HEAD.
|
| 81 |
+
|
| 82 |
+
### Re-ranker comparison, 128-token state (production setting)
|
| 83 |
+
|
| 84 |
+
`spike/results/compare_models.json` (`spike/compare_models.py`): the 40 most recent qualifying
|
| 85 |
+
moon commits, BM25 top 24, state truncated to 128 tokens, probabilities pooled over all
|
| 86 |
+
candidates (base rate 0.309).
|
| 87 |
+
|
| 88 |
+
| model | Laya-only MRR | Laya-only P@10 | RRF MRR | AUROC | ECE | mean P |
|
| 89 |
+
|---|---|---|---|---|---|---|
|
| 90 |
+
| BM25 alone | 0.480 (MRR) | 0.340 | – | – | – | – |
|
| 91 |
+
| laya-base (`convaiinnovations/laya`) | 0.479 | 0.348 | 0.505 | 0.586 | 0.362 | 0.671 |
|
| 92 |
+
| laya-typed-decisions | 0.441 | 0.288 | 0.456 | 0.539 | 0.239 | 0.545 |
|
| 93 |
+
| **laya-code** | **0.702** | **0.405** | **0.630** | **0.713** | **0.049** | 0.328 |
|
| 94 |
+
|
| 95 |
+
### Held-out evaluation, 256-token state (training protocol)
|
| 96 |
+
|
| 97 |
+
`spike/results/finetune_eval.json` (`finetune/eval.py`): 40 tasks per held-out repo, BM25 top 32,
|
| 98 |
+
state truncated to 256 tokens. Paired bootstrap over the tasks.
|
| 99 |
+
|
| 100 |
+
| repo | model | Laya-only MRR | RRF MRR | RRF P@10 | ECE (15 bins) | AUROC pooled |
|
| 101 |
+
|---|---|---|---|---|---|---|
|
| 102 |
+
| moon | laya-base | 0.497 | 0.586 | 0.343 | 0.461 | 0.584 |
|
| 103 |
+
| moon | laya-code | 0.526 | 0.519 | 0.398 | 0.060 | 0.677 |
|
| 104 |
+
| pilot-space | laya-base | 0.519 | 0.762 | 0.323 | 0.509 | 0.536 |
|
| 105 |
+
| pilot-space | laya-code | 0.646 | 0.714 | 0.393 | 0.016 | 0.709 |
|
| 106 |
+
|
| 107 |
+
Under this protocol, laya-code clearly improves calibration and discrimination. It puts more
|
| 108 |
+
gold-file spans in the top 10: RRF P@10 rose by +0.055 (95% CI [0.013, 0.098]) on moon and by
|
| 109 |
+
+0.070 ([0.033, 0.108]) on pilot-space. Its **RRF MRR did not beat laya-base's**: −0.066
|
| 110 |
+
([−0.195, 0.061]) on moon and −0.049 ([−0.172, 0.073]) on pilot-space. For that reason,
|
| 111 |
+
laya-codex fuses laya-code by score (`(1−w)·lexical + w·P`, w = 0.5) rather than by RRF.
|
| 112 |
+
|
| 113 |
+
End to end, the laya-codex paired Claude Code benchmark (20 moon tasks, `docs/RESULTS.md`)
|
| 114 |
+
measured −45% code-reading tokens and −25% wall-clock time for the whole pipeline, with no loss
|
| 115 |
+
of answer recall. The same document reports that the model's **marginal** contribution over
|
| 116 |
+
lexical-only ranking is within run-to-run noise at n = 20. Do not read the pipeline numbers as a
|
| 117 |
+
property of this model.
|
| 118 |
+
|
| 119 |
+
## Intended use
|
| 120 |
+
|
| 121 |
+
- Re-ranking lexical (BM25) candidates of source-code chunks for a natural-language
|
| 122 |
+
software-change task, as a calibrated `P(relevant)`.
|
| 123 |
+
- Gating or fusing retrieval results by probability; P is calibrated to the training
|
| 124 |
+
distribution (ECE ≤ 0.06 on held-out repos).
|
| 125 |
+
|
| 126 |
+
## Out of scope and limitations
|
| 127 |
+
|
| 128 |
+
- **Weak labels.** A commit touching a file does not make every window of it relevant, and the
|
| 129 |
+
file-level gold is coarse.
|
| 130 |
+
- **Small evaluation.** Each held-out set has 40 tasks, and most CIs are wide. The two
|
| 131 |
+
protocols (128 vs 256 tokens, top 24 vs 32) give different absolute numbers, as shown above.
|
| 132 |
+
- **Low probabilities.** P rarely exceeds 0.5 (max 0.49 on moon, 0.50 on pilot-space at 256
|
| 133 |
+
tokens). Use rank or score fusion, or a threshold near 0.4, not "P ≥ 0.5 means relevant".
|
| 134 |
+
- **Under-trained.** Only about half an epoch of the v2 data was used, on a shared laptop.
|
| 135 |
+
Validation AUROC was still rising when training stopped.
|
| 136 |
+
- **English prompts only.** Training covered Rust, Python, TypeScript and JavaScript; other
|
| 137 |
+
languages are untested.
|
| 138 |
+
- **Other tasks untested.** It is not a general Laya replacement: `choice`/`score` questions
|
| 139 |
+
(for example, task scope) were not trained, and zero-shot scope accuracy is poor
|
| 140 |
+
(`spike/results/scope_eval.json`).
|
| 141 |
+
- **Too slow for interactive CPU use.** At 421M parameters, CPU re-ranking of 24 candidates is
|
| 142 |
+
too slow for interactive use. laya-codex runs it on Metal, or falls back to lexical ranking.
|
| 143 |
+
- **Legal status of training data.** The model was trained on permissively licensed public
|
| 144 |
+
code plus the author's own repositories. It is a classifier and cannot reproduce that code,
|
| 145 |
+
but the legal status of weights trained on source code is not settled.
|
| 146 |
+
|
| 147 |
+
## License and attribution
|
| 148 |
+
|
| 149 |
+
Apache-2.0, like the base model. laya-code is a Derivative Work of
|
| 150 |
+
[convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya) (Apache-2.0, © Convai
|
| 151 |
+
Innovations), which builds on
|
| 152 |
+
[answerdotai/ModernBERT-large](https://huggingface.co/answerdotai/ModernBERT-large) (Apache-2.0).
|
| 153 |
+
The modified files are `model.safetensors` and `rl_agent_config.json`; every other file is
|
| 154 |
+
unchanged from the base. See `NOTICE`.
|
| 155 |
+
|
| 156 |
+
## Files
|
| 157 |
+
|
| 158 |
+
See `MANIFEST.sha256` for the sha256 of every uploaded file.
|