Instructions to use Shuu12121/NightJar-large-CodeSearch-Embedding with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use Shuu12121/NightJar-large-CodeSearch-Embedding with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("Shuu12121/NightJar-large-CodeSearch-Embedding") sentences = [ "Extract json from string with support for '' and None.", "def json_loads(data: Any) -> dict[str, Any]:\n \"\"\"Extract json from string with support for '' and None.\"\"\"\n if not data:\n return {}\n try:\n return json_loads_util(data)\n except json.JSONDecodeError as err:\n raise APIError(\"Invalid json\") from err", "def str2json(v):\n \"\"\"\n convert str to json data\n :param v:\n :return:\n \"\"\"\n try:\n return json.loads(v)\n except:\n return None", "def extract_json_from_string(response_msg: str) -> str:\n \"\"\"\n Attempts to extract JSON (object or array) from within a larger string, not specific to markdown.\n \"\"\"\n json_pattern = re.compile(r\"\\{.*\\}|\\[.*\\]\")\n match = json_pattern.search(response_msg)\n if match:\n return match.group(0)\n\n return response_msg", "def parse_json(val: str):\n \"\"\"\n Parses json if string else return\n \"\"\"\n if isinstance(val, str):\n val = json.loads(val)\n if isinstance(val, dict):\n val = frappe._dict(val)\n return val" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [5, 5] - Notebooks
- Google Colab
- Kaggle
NightJar-large-CodeSearch-Embedding
NightJar-large-CodeSearch-Embedding is a code-retrieval embedding model fine-tuned from Shuu12121/NightJar-large (a 346M-parameter ModernBERT encoder pre-trained from scratch on code) with Sentence Transformers. It maps text and code to a 1024-dimensional vector space, so natural-language queries, code snippets and code edits can be compared with cosine similarity.
Use it for natural-language-to-code search, code-to-code retrieval and code-edit search.
| Base model | Shuu12121/NightJar-large |
| Model type | Sentence Transformer (Transformer + [CLS] pooling) |
| Parameters | 346M |
| Output dimension | 1024 |
| Max sequence length | 1,024 tokens |
| Similarity function | Cosine similarity |
| Task prefix | none |
Results (MTEB, nDCG@10)
Scores were computed with MTEB 2.5.1. All models below were fine-tuned with the same recipe (see Training).
CodeSearchNetRetrieval
| Model | Go | Java | JavaScript | PHP | Python | Ruby | Avg |
|---|---|---|---|---|---|---|---|
| NightJar-large-CodeSearch-Embedding | 0.9682 | 0.9324 | 0.8317 | 0.8989 | 0.9532 | 0.8850 | 0.9116 |
| NightJar-CodeSearch-Embedding (smaller NightJar) | 0.9656 | 0.9278 | 0.8236 | 0.8915 | 0.9410 | 0.8663 | 0.9026 |
| NightOwl (same recipe) | 0.9674 | 0.9232 | 0.8157 | 0.8907 | 0.9393 | 0.8618 | 0.8997 |
CodeEditSearchRetrieval
| Model | C | C++ | Go | Java | JS | PHP | Python | Ruby | Rust | Scala | Shell | Swift | TS | Avg |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| NightJar-large-CodeSearch-Embedding | 0.7220 | 0.7537 | 0.8062 | 0.7831 | 0.7846 | 0.7743 | 0.8048 | 0.8160 | 0.7396 | 0.8248 | 0.7571 | 0.7882 | 0.8145 | 0.7822 |
| NightJar-CodeSearch-Embedding (smaller NightJar) | 0.6687 | 0.7105 | 0.7716 | 0.7365 | 0.7386 | 0.7162 | 0.7711 | 0.7723 | 0.6900 | 0.7859 | 0.7119 | 0.7470 | 0.7759 | 0.7382 |
| NightOwl (same recipe) | 0.6624 | 0.7170 | 0.7800 | 0.7509 | 0.7539 | 0.7339 | 0.7716 | 0.7814 | 0.7100 | 0.7795 | 0.7220 | 0.7453 | 0.7890 | 0.7459 |
NightJar-large-CodeSearch-Embedding scores highest on every language of both benchmarks.
"NightJar-CodeSearch-Embedding" is the smaller Shuu12121/NightJar fine-tuned with exactly this setup for the full epoch. "NightOwl (same recipe)" is Shuu12121/NightOwl fine-tuned with the same losses and hyperparameters (without the additional-languages dataset).
The fine-tuning data is the same decontaminated set used for NightOwl-CodeEmbedding: overlaps with the CodeSearchNet test splits, and between the commitpackft-derived code-edit data and the CodeEditSearchRetrieval evaluation examples, were removed before training. MTEB names the CodeEditSearchRetrieval evaluation split train, but those examples were not used for fine-tuning.
Usage
pip install -U sentence-transformers
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("Shuu12121/NightJar-large-CodeSearch-Embedding")
sentences = [
# query
"Optional. The amount of time that a device will be initially allocated\n"
"for. This can eventually be extended with the UpdateDeviceSession RPC.\n"
"Default: 15 minutes.\n\n"
"Generated from protobuf field <code>.google.protobuf.Duration ttl = 13 "
"[(.google.api.field_behavior) = OPTIONAL];</code>\n"
"@return \\Google\\Protobuf\\Duration|null",
# matching code
"public function getTtl()\n {\n return $this->readOneof(13);\n }",
# similar code, wrong field
"public function getTtl()\n {\n return $this->readOneof(7);\n }",
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# (3, 1024)
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.8610, 0.8172],
# [0.8610, 1.0000, 0.9410],
# [0.8172, 0.9410, 1.0000]])
The query scores higher against the matching getter (0.861) than against the near-identical one that reads a different field (0.817), even though the two code snippets are very close to each other (0.941).
Input format tips
- Pass queries and code as plain strings, with no task prefix. The fine-tuning data used plain text for queries and plain code for documents, without the
[NL]/[Code1]markers from pre-training. - The model was trained with sequences up to 1,024 tokens. Longer inputs are truncated.
- The model was trained in both directions (query → code and code → query), so it can also be used to find descriptions for code.
Training
The model was fine-tuned for one epoch (3,978 steps, 4,073,472 examples) from Shuu12121/NightJar-large.
Data
Each example is an anchor, a positive, and 15 hard negatives (negative_1 … negative_15). The three datasets were sampled in an 8 : 8 : 1 ratio:
| Dataset | Content |
|---|---|
Shuu12121/owl_code_search_hard_negative_datasets_V2_kd |
docstring → code search, 8 languages |
Shuu12121/owl_code_search_hard_negative_datasets_additional_languages_V2_kd |
docstring → code search, 8 more languages |
Shuu12121/codeedit_hard_negative_datasets_kd |
code-edit retrieval, 11 languages |
Approximate token lengths, from the first 1,000 samples: anchors average 75 tokens, positives 197 tokens, and negatives 188–215 tokens. The maximum is 1,024 tokens for every column.
Loss
CachedMultipleNegativesRankingLoss with:
{
"scale": 100.0,
"similarity_fct": "cos_sim",
"mini_batch_size": 64,
"gather_across_devices": false,
"directions": ["query_to_doc", "doc_to_query"],
"partition_mode": "joint",
"hardness_mode": null,
"hardness_strength": 0.0
}
Hyperparameters
| Batch size | 1,024 (no gradient accumulation) |
| Learning rate | 6e-5, cosine schedule, no warmup |
| Weight decay | 0.01 |
| Optimizer | adamw_torch_fused (β = 0.9 / 0.999, ε = 1e-8) |
| Epochs | 1 |
| Precision | bf16 |
| Gradient checkpointing | on |
| Seed | 42 |
Training-time evaluation
During training, the model was monitored with Sentence Transformers' InformationRetrievalEvaluator on CodeSearchNet (six languages, validation split). The values below are from the end of training.
| Metric | Go | Java | JavaScript | PHP | Python | Ruby | Avg |
|---|---|---|---|---|---|---|---|
| Accuracy@1 | 0.8740 | 0.7630 | 0.7730 | 0.7660 | 0.8900 | 0.8510 | 0.8195 |
| Recall@10 | 0.9810 | 0.9350 | 0.9030 | 0.8220 | 0.9890 | 0.9570 | 0.9312 |
| MRR@10 | 0.9184 | 0.8309 | 0.8219 | 0.7898 | 0.9302 | 0.8945 | 0.8643 |
| nDCG@10 | 0.9341 | 0.8569 | 0.8420 | 0.7979 | 0.9449 | 0.9101 | 0.8810 |
These numbers come from a different evaluation set than MTEB, so they are not comparable with the MTEB scores above.
The monitored nDCG@10 rose quickly in the first few hundred steps. After about step 2,000 (roughly half an epoch) it stayed within ±0.01 of its final value for every language, so most of the gain was already reached in the first half of the epoch.
Framework versions
- Python 3.10.12
- Sentence Transformers 5.3.0
- Transformers 4.56.2
- PyTorch 2.8.0+cu128
- Accelerate 1.12.0
- Datasets 3.6.0
- Tokenizers 0.22.1
Architecture
SentenceTransformer(
(0): Transformer({'max_seq_length': 1024, 'do_lower_case': False, 'architecture': 'ModernBertModel'})
(1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
)
The encoder is the ModernBERT-based NightJar-large: 28 layers, hidden size 1024, 16 heads, global attention in every 3rd layer and a 128-token sliding window otherwise. The tokenizer keeps indentation and newlines as single tokens. See the NightJar-large model card for details.
Limitations
- Retrieval quality was evaluated on CodeSearchNetRetrieval and CodeEditSearchRetrieval only.
- Performance varies by language with the amount and quality of the fine-tuning data. Languages not covered by the fine-tuning data are not well supported.
- Inputs longer than 1,024 tokens are truncated.
- Retrieved code may contain bugs, outdated APIs or insecure patterns. Review it before use.
License
The model weights are released under Apache 2.0. The training datasets carry their own licenses and terms of use. Check them if your use case depends on the provenance of the training data.
Citation
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}
@misc{gao2021scaling,
title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup},
author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan},
year={2021},
eprint={2101.06983},
archivePrefix={arXiv},
primaryClass={cs.LG}
}
日本語版
概要
NightJar-large-CodeSearch-Embeddingは、Shuu12121/NightJar-large(コードでゼロから事前学習した、約3.46億パラメータのModernBERTエンコーダー)を、Sentence Transformersでコード検索向けに追加学習した埋め込みモデルです。文章とコードを1024次元のベクトルに変換するので、自然言語の質問、コード断片、コードの変更内容を、コサイン類似度で比較できます。
自然言語からのコード検索、コードからコードの検索、コード編集の検索に使えます。
| 項目 | 内容 |
|---|---|
| ベースモデル | Shuu12121/NightJar-large |
| モデルの種類 | Sentence Transformer(Transformer + [CLS]プーリング) |
| パラメータ数 | 約3.46億 |
| 出力次元 | 1024 |
| 最大系列長 | 1,024トークン |
| 類似度 | コサイン類似度 |
| タスクプレフィックス | なし |
評価結果(MTEB、nDCG@10)
MTEB 2.5.1で計算しました。比較対象のモデルも、すべて同じ手順で追加学習しています(詳細は「学習」を参照)。
| モデル | CodeSearchNetRetrieval(平均) | CodeEditSearchRetrieval(平均) |
|---|---|---|
| NightJar-large-CodeSearch-Embedding | 0.9116 | 0.7822 |
| NightJar-CodeSearch-Embedding | 0.9026 | 0.7382 |
| NightOwl(同じ手順で追加学習) | 0.8997 | 0.7459 |
NightJar-large-CodeSearch-Embeddingは、CodeSearchNetRetrievalの6言語、CodeEditSearchRetrievalの13言語のすべてで、最も高いスコアでした。言語別のスコアは英語版の表を参照してください。
追加学習のデータは、NightOwl-CodeEmbeddingと同じ重複除去済みのデータセットです。学習前に、コード検索データとCodeSearchNetのテストデータの重複、commitpackft由来のコード編集データとCodeEditSearchRetrievalの評価データの重複を取り除いています。MTEBではCodeEditSearchRetrievalの評価データがtrainという名前になっていますが、この評価データは追加学習には使っていません。
使い方
英語版の「Usage」にサンプルコードがあります。
- 最大長: 学習時の最大長は1,024トークンです。それより長い入力は切り捨てられます。
- 双方向: 質問→コードと、コード→質問の両方向で学習しているので、コードに合う説明文を探す用途にも使えます。
学習
Shuu12121/NightJar-largeから、1エポック(3,978ステップ、4,073,472件)で追加学習しました。
データ: 各サンプルは、アンカー1件、正例1件、ハードネガティブ15件です。次の3つのデータセットを 8 : 8 : 1 の比率でサンプリングしました。
owl_code_search_hard_negative_datasets_V2_kd(docstringからコードの検索、8言語)owl_code_search_hard_negative_datasets_additional_languages_V2_kd(同じく追加の8言語)codeedit_hard_negative_datasets_kd(コード編集の検索、11言語)
損失関数: Cached MultipleNegativesRankingLoss(scale 100、mini batch size 64、batch_size 1024)質問→文書と文書→質問の双方向で学習しています。
ハイパーパラメータ:
| 項目 | 値 |
|---|---|
| バッチサイズ | 1,024(勾配蓄積なし) |
| 学習率 | 6e-5(コサイン減衰、ウォームアップなし) |
| Weight decay | 0.01 |
| 最適化手法 | adamw_torch_fused |
| エポック数 | 1 |
| 精度 | bf16 |
| Gradient checkpointing | あり |
| シード | 42 |
学習中の評価: 学習中は、Sentence TransformersのInformationRetrievalEvaluatorでCodeSearchNet(6言語、validation splitにおける性能)を監視しました。学習終了時のnDCG@10は、Go 0.9341、Java 0.8569、JavaScript 0.8420、PHP 0.7979、Python 0.9449、Ruby 0.9101で、平均は0.8810です(そのほかの指標は英語版の表を参照)。この数値はMTEBとは別の評価セットで計算したものなので、MTEBのスコアとは比較できません。
監視していたnDCG@10は、最初の数百ステップで急に上がり、約2,000ステップ(約0.5エポック)以降は、どの言語も最終値の±0.01以内で推移しました。
制限事項
- 検索性能の評価は、CodeSearchNetRetrievalとCodeEditSearchRetrievalのみです。
- 追加学習データの量や質などが言語ごとに違うため、性能にも差があります。追加学習データに含まれない言語では性能が下がる場合があります。
- 1,024トークンを超える入力は切り捨てられます。
- 検索で得られたコードには、バグ、古いAPI、安全でない実装が含まれる可能性があります。利用前に確認してください。
ライセンス
モデルの重みはApache 2.0で公開しています。学習データセットには、それぞれ独自のライセンスと利用条件があります。用途によっては、各データセットの条件も確認してください。
- Downloads last month
- -
Model tree for Shuu12121/NightJar-large-CodeSearch-Embedding
Base model
Shuu12121/NightJar-large