Sentence Similarity
sentence-transformers
Safetensors
English
modernbert
feature-extraction
dense
code
code-search
code-retrieval
text-embeddings-inference

NightJar-large-CodeSearch-Embedding

NightJar-large-CodeSearch-Embedding is a code-retrieval embedding model fine-tuned from Shuu12121/NightJar-large (a 346M-parameter ModernBERT encoder pre-trained from scratch on code) with Sentence Transformers. It maps text and code to a 1024-dimensional vector space, so natural-language queries, code snippets and code edits can be compared with cosine similarity.

Use it for natural-language-to-code search, code-to-code retrieval and code-edit search.

Base model Shuu12121/NightJar-large
Model type Sentence Transformer (Transformer + [CLS] pooling)
Parameters 346M
Output dimension 1024
Max sequence length 1,024 tokens
Similarity function Cosine similarity
Task prefix none

Results (MTEB, nDCG@10)

Scores were computed with MTEB 2.5.1. All models below were fine-tuned with the same recipe (see Training).

CodeSearchNetRetrieval

Model Go Java JavaScript PHP Python Ruby Avg
NightJar-large-CodeSearch-Embedding 0.9682 0.9324 0.8317 0.8989 0.9532 0.8850 0.9116
NightJar-CodeSearch-Embedding (smaller NightJar) 0.9656 0.9278 0.8236 0.8915 0.9410 0.8663 0.9026
NightOwl (same recipe) 0.9674 0.9232 0.8157 0.8907 0.9393 0.8618 0.8997

CodeEditSearchRetrieval

Model C C++ Go Java JS PHP Python Ruby Rust Scala Shell Swift TS Avg
NightJar-large-CodeSearch-Embedding 0.7220 0.7537 0.8062 0.7831 0.7846 0.7743 0.8048 0.8160 0.7396 0.8248 0.7571 0.7882 0.8145 0.7822
NightJar-CodeSearch-Embedding (smaller NightJar) 0.6687 0.7105 0.7716 0.7365 0.7386 0.7162 0.7711 0.7723 0.6900 0.7859 0.7119 0.7470 0.7759 0.7382
NightOwl (same recipe) 0.6624 0.7170 0.7800 0.7509 0.7539 0.7339 0.7716 0.7814 0.7100 0.7795 0.7220 0.7453 0.7890 0.7459

NightJar-large-CodeSearch-Embedding scores highest on every language of both benchmarks.

"NightJar-CodeSearch-Embedding" is the smaller Shuu12121/NightJar fine-tuned with exactly this setup for the full epoch. "NightOwl (same recipe)" is Shuu12121/NightOwl fine-tuned with the same losses and hyperparameters (without the additional-languages dataset).

The fine-tuning data is the same decontaminated set used for NightOwl-CodeEmbedding: overlaps with the CodeSearchNet test splits, and between the commitpackft-derived code-edit data and the CodeEditSearchRetrieval evaluation examples, were removed before training. MTEB names the CodeEditSearchRetrieval evaluation split train, but those examples were not used for fine-tuning.

Usage

pip install -U sentence-transformers
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("Shuu12121/NightJar-large-CodeSearch-Embedding")

sentences = [
    # query
    "Optional. The amount of time that a device will be initially allocated\n"
    "for. This can eventually be extended with the UpdateDeviceSession RPC.\n"
    "Default: 15 minutes.\n\n"
    "Generated from protobuf field <code>.google.protobuf.Duration ttl = 13 "
    "[(.google.api.field_behavior) = OPTIONAL];</code>\n"
    "@return \\Google\\Protobuf\\Duration|null",
    # matching code
    "public function getTtl()\n    {\n        return $this->readOneof(13);\n    }",
    # similar code, wrong field
    "public function getTtl()\n    {\n        return $this->readOneof(7);\n    }",
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# (3, 1024)

similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.8610, 0.8172],
#         [0.8610, 1.0000, 0.9410],
#         [0.8172, 0.9410, 1.0000]])

The query scores higher against the matching getter (0.861) than against the near-identical one that reads a different field (0.817), even though the two code snippets are very close to each other (0.941).

Input format tips

  • Pass queries and code as plain strings, with no task prefix. The fine-tuning data used plain text for queries and plain code for documents, without the [NL] / [Code1] markers from pre-training.
  • The model was trained with sequences up to 1,024 tokens. Longer inputs are truncated.
  • The model was trained in both directions (query → code and code → query), so it can also be used to find descriptions for code.

Training

The model was fine-tuned for one epoch (3,978 steps, 4,073,472 examples) from Shuu12121/NightJar-large.

Data

Each example is an anchor, a positive, and 15 hard negatives (negative_1 … negative_15). The three datasets were sampled in an 8 : 8 : 1 ratio:

Dataset Content
Shuu12121/owl_code_search_hard_negative_datasets_V2_kd docstring → code search, 8 languages
Shuu12121/owl_code_search_hard_negative_datasets_additional_languages_V2_kd docstring → code search, 8 more languages
Shuu12121/codeedit_hard_negative_datasets_kd code-edit retrieval, 11 languages

Approximate token lengths, from the first 1,000 samples: anchors average 75 tokens, positives 197 tokens, and negatives 188–215 tokens. The maximum is 1,024 tokens for every column.

Loss

CachedMultipleNegativesRankingLoss with:

{
    "scale": 100.0,
    "similarity_fct": "cos_sim",
    "mini_batch_size": 64,
    "gather_across_devices": false,
    "directions": ["query_to_doc", "doc_to_query"],
    "partition_mode": "joint",
    "hardness_mode": null,
    "hardness_strength": 0.0
}

Hyperparameters

Batch size 1,024 (no gradient accumulation)
Learning rate 6e-5, cosine schedule, no warmup
Weight decay 0.01
Optimizer adamw_torch_fused (β = 0.9 / 0.999, ε = 1e-8)
Epochs 1
Precision bf16
Gradient checkpointing on
Seed 42

Training-time evaluation

During training, the model was monitored with Sentence Transformers' InformationRetrievalEvaluator on CodeSearchNet (six languages, validation split). The values below are from the end of training.

Metric Go Java JavaScript PHP Python Ruby Avg
Accuracy@1 0.8740 0.7630 0.7730 0.7660 0.8900 0.8510 0.8195
Recall@10 0.9810 0.9350 0.9030 0.8220 0.9890 0.9570 0.9312
MRR@10 0.9184 0.8309 0.8219 0.7898 0.9302 0.8945 0.8643
nDCG@10 0.9341 0.8569 0.8420 0.7979 0.9449 0.9101 0.8810

These numbers come from a different evaluation set than MTEB, so they are not comparable with the MTEB scores above.

The monitored nDCG@10 rose quickly in the first few hundred steps. After about step 2,000 (roughly half an epoch) it stayed within ±0.01 of its final value for every language, so most of the gain was already reached in the first half of the epoch.

Framework versions

  • Python 3.10.12
  • Sentence Transformers 5.3.0
  • Transformers 4.56.2
  • PyTorch 2.8.0+cu128
  • Accelerate 1.12.0
  • Datasets 3.6.0
  • Tokenizers 0.22.1

Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 1024, 'do_lower_case': False, 'architecture': 'ModernBertModel'})
  (1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
)

The encoder is the ModernBERT-based NightJar-large: 28 layers, hidden size 1024, 16 heads, global attention in every 3rd layer and a 128-token sliding window otherwise. The tokenizer keeps indentation and newlines as single tokens. See the NightJar-large model card for details.

Limitations

  • Retrieval quality was evaluated on CodeSearchNetRetrieval and CodeEditSearchRetrieval only.
  • Performance varies by language with the amount and quality of the fine-tuning data. Languages not covered by the fine-tuning data are not well supported.
  • Inputs longer than 1,024 tokens are truncated.
  • Retrieved code may contain bugs, outdated APIs or insecure patterns. Review it before use.

License

The model weights are released under Apache 2.0. The training datasets carry their own licenses and terms of use. Check them if your use case depends on the provenance of the training data.

Citation

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
@misc{gao2021scaling,
    title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup},
    author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan},
    year={2021},
    eprint={2101.06983},
    archivePrefix={arXiv},
    primaryClass={cs.LG}
}

日本語版

概要

NightJar-large-CodeSearch-Embeddingは、Shuu12121/NightJar-large(コードでゼロから事前学習した、約3.46億パラメータのModernBERTエンコーダー)を、Sentence Transformersでコード検索向けに追加学習した埋め込みモデルです。文章とコードを1024次元のベクトルに変換するので、自然言語の質問、コード断片、コードの変更内容を、コサイン類似度で比較できます。

自然言語からのコード検索、コードからコードの検索、コード編集の検索に使えます。

項目 内容
ベースモデル Shuu12121/NightJar-large
モデルの種類 Sentence Transformer(Transformer + [CLS]プーリング)
パラメータ数 約3.46億
出力次元 1024
最大系列長 1,024トークン
類似度 コサイン類似度
タスクプレフィックス なし

評価結果(MTEB、nDCG@10)

MTEB 2.5.1で計算しました。比較対象のモデルも、すべて同じ手順で追加学習しています(詳細は「学習」を参照)。

モデル CodeSearchNetRetrieval(平均) CodeEditSearchRetrieval(平均)
NightJar-large-CodeSearch-Embedding 0.9116 0.7822
NightJar-CodeSearch-Embedding 0.9026 0.7382
NightOwl(同じ手順で追加学習) 0.8997 0.7459

NightJar-large-CodeSearch-Embeddingは、CodeSearchNetRetrievalの6言語、CodeEditSearchRetrievalの13言語のすべてで、最も高いスコアでした。言語別のスコアは英語版の表を参照してください。

追加学習のデータは、NightOwl-CodeEmbeddingと同じ重複除去済みのデータセットです。学習前に、コード検索データとCodeSearchNetのテストデータの重複、commitpackft由来のコード編集データとCodeEditSearchRetrievalの評価データの重複を取り除いています。MTEBではCodeEditSearchRetrievalの評価データがtrainという名前になっていますが、この評価データは追加学習には使っていません。

使い方

英語版の「Usage」にサンプルコードがあります。

  • 最大長: 学習時の最大長は1,024トークンです。それより長い入力は切り捨てられます。
  • 双方向: 質問→コードと、コード→質問の両方向で学習しているので、コードに合う説明文を探す用途にも使えます。

学習

Shuu12121/NightJar-largeから、1エポック(3,978ステップ、4,073,472件)で追加学習しました。

データ: 各サンプルは、アンカー1件、正例1件、ハードネガティブ15件です。次の3つのデータセットを 8 : 8 : 1 の比率でサンプリングしました。

  • owl_code_search_hard_negative_datasets_V2_kd(docstringからコードの検索、8言語)
  • owl_code_search_hard_negative_datasets_additional_languages_V2_kd(同じく追加の8言語)
  • codeedit_hard_negative_datasets_kd(コード編集の検索、11言語)

損失関数: Cached MultipleNegativesRankingLoss(scale 100、mini batch size 64、batch_size 1024)質問→文書と文書→質問の双方向で学習しています。

ハイパーパラメータ:

項目 値
バッチサイズ 1,024(勾配蓄積なし)
学習率 6e-5(コサイン減衰、ウォームアップなし)
Weight decay 0.01
最適化手法 adamw_torch_fused
エポック数 1
精度 bf16
Gradient checkpointing あり
シード 42

学習中の評価: 学習中は、Sentence TransformersのInformationRetrievalEvaluatorでCodeSearchNet(6言語、validation splitにおける性能)を監視しました。学習終了時のnDCG@10は、Go 0.9341、Java 0.8569、JavaScript 0.8420、PHP 0.7979、Python 0.9449、Ruby 0.9101で、平均は0.8810です(そのほかの指標は英語版の表を参照)。この数値はMTEBとは別の評価セットで計算したものなので、MTEBのスコアとは比較できません。

監視していたnDCG@10は、最初の数百ステップで急に上がり、約2,000ステップ(約0.5エポック)以降は、どの言語も最終値の±0.01以内で推移しました。

制限事項

  • 検索性能の評価は、CodeSearchNetRetrievalとCodeEditSearchRetrievalのみです。
  • 追加学習データの量や質などが言語ごとに違うため、性能にも差があります。追加学習データに含まれない言語では性能が下がる場合があります。
  • 1,024トークンを超える入力は切り捨てられます。
  • 検索で得られたコードには、バグ、古いAPI、安全でない実装が含まれる可能性があります。利用前に確認してください。

ライセンス

モデルの重みはApache 2.0で公開しています。学習データセットには、それぞれ独自のライセンスと利用条件があります。用途によっては、各データセットの条件も確認してください。

Downloads last month
-
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Shuu12121/NightJar-large-CodeSearch-Embedding

Finetuned
(1)
this model

Datasets used to train Shuu12121/NightJar-large-CodeSearch-Embedding

Papers for Shuu12121/NightJar-large-CodeSearch-Embedding