Sentence Similarity
sentence-transformers
Safetensors
nomic_bert
code-search
mteb
custom_code
text-embeddings-inference
Instructions to use madhurr382/coderankembed-apps-ft with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use madhurr382/coderankembed-apps-ft with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("madhurr382/coderankembed-apps-ft", trust_remote_code=True) sentences = [ "That is a happy person", "That is a happy dog", "That is a very happy person", "Today is a sunny day" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
|
Download README.md from madhurr382/coderankembed-apps-ft: direct link, hf CLI and curl.
- Browser
- Download file 1.93 kB
-
https://huggingface.co/madhurr382/coderankembed-apps-ft/resolve/main/README.md
- Command line
-
hf download hf://madhurr382/coderankembed-apps-ft/README.md
-
curl -L -o README.md https://huggingface.co/madhurr382/coderankembed-apps-ft/resolve/main/README.md
1.93 kB
| license: mit | |
| base_model: nomic-ai/CodeRankEmbed | |
| library_name: sentence-transformers | |
| pipeline_tag: sentence-similarity | |
| tags: | |
| - code-search | |
| - sentence-transformers | |
| - mteb | |
| datasets: | |
| - CoIR-Retrieval/apps | |
| # coderankembed-apps-ft | |
| `nomic-ai/CodeRankEmbed` fine-tuned on the **train** split of CoIR-Retrieval/apps (natural-language | |
| problem statement → Python solution) for the MTEB `AppsRetrieval` task. No test | |
| queries or qrels were used. | |
| **AppsRetrieval test (MTEB, CPU):** NDCG@10 0.4709, MRR@10 0.4303 (queries cleaned with `desc-io`). | |
| ## Usage | |
| Queries need the base model's prefix; documents (code) have none. Load with | |
| `trust_remote_code=True` (NomicBert). | |
| ```python | |
| from sentence_transformers import SentenceTransformer | |
| m = SentenceTransformer("madhurr382/coderankembed-apps-ft", trust_remote_code=True) | |
| q = m.encode(["Represent this query for searching relevant code: " + "Given an array, return the length of its longest increasing subsequence."], | |
| normalize_embeddings=True) | |
| d = m.encode(["def lis(a): ..."], normalize_embeddings=True) | |
| print(q @ d.T) | |
| ``` | |
| For the best scores, clean APPS-style statements first: keep the description and the | |
| Input/Output sections, drop samples, notes and constraints (`retrieval/query_clean.py`, | |
| mode `desc-io`). | |
| **transformers v5:** its loader leaves NomicBert's non-persistent buffers (rotary | |
| `inv_freq`, attention `norm_factor`) uninitialised. The repo's | |
| `retrieval/eval_baseline.py::load_st_model` rebuilds them after loading. | |
| ## Training | |
| - Loss: CachedMultipleNegativesRankingLoss, batch 128, lr 2e-05, 2 epochs, max len 512, NO_DUPLICATES sampler | |
| - Hard negatives: none | |
| - Train rows: 4500; validation: 500 held-out train queries over 5000 docs | |
| - Saved epoch: 2 (best validation MRR@10) | |
| | epoch | val MRR@10 | val NDCG@10 | | |
| |---|---|---| | |
| | 0 | 0.6466 | 0.6666 | | |
| | 1 | 0.7378 | 0.7669 | | |
| | 2 | 0.7497 | 0.7781 | | |
| Full run config: `finetune_config.json`. | |