madhurr382's picture
Upload fine-tuned CodeRankEmbed (AppsRetrieval)
c9f6787 verified
|
Raw History Blame Contribute Delete
1.93 kB
---
license: mit
base_model: nomic-ai/CodeRankEmbed
library_name: sentence-transformers
pipeline_tag: sentence-similarity
tags:
- code-search
- sentence-transformers
- mteb
datasets:
- CoIR-Retrieval/apps
---
# coderankembed-apps-ft
`nomic-ai/CodeRankEmbed` fine-tuned on the **train** split of CoIR-Retrieval/apps (natural-language
problem statement → Python solution) for the MTEB `AppsRetrieval` task. No test
queries or qrels were used.
**AppsRetrieval test (MTEB, CPU):** NDCG@10 0.4709, MRR@10 0.4303 (queries cleaned with `desc-io`).
## Usage
Queries need the base model's prefix; documents (code) have none. Load with
`trust_remote_code=True` (NomicBert).
```python
from sentence_transformers import SentenceTransformer
m = SentenceTransformer("madhurr382/coderankembed-apps-ft", trust_remote_code=True)
q = m.encode(["Represent this query for searching relevant code: " + "Given an array, return the length of its longest increasing subsequence."],
normalize_embeddings=True)
d = m.encode(["def lis(a): ..."], normalize_embeddings=True)
print(q @ d.T)
```
For the best scores, clean APPS-style statements first: keep the description and the
Input/Output sections, drop samples, notes and constraints (`retrieval/query_clean.py`,
mode `desc-io`).
**transformers v5:** its loader leaves NomicBert's non-persistent buffers (rotary
`inv_freq`, attention `norm_factor`) uninitialised. The repo's
`retrieval/eval_baseline.py::load_st_model` rebuilds them after loading.
## Training
- Loss: CachedMultipleNegativesRankingLoss, batch 128, lr 2e-05, 2 epochs, max len 512, NO_DUPLICATES sampler
- Hard negatives: none
- Train rows: 4500; validation: 500 held-out train queries over 5000 docs
- Saved epoch: 2 (best validation MRR@10)
| epoch | val MRR@10 | val NDCG@10 |
|---|---|---|
| 0 | 0.6466 | 0.6666 |
| 1 | 0.7378 | 0.7669 |
| 2 | 0.7497 | 0.7781 |
Full run config: `finetune_config.json`.