--- language: - en library_name: sentence-transformers pipeline_tag: sentence-similarity license: apache-2.0 base_model: Shuu12121/NightJar-large tags: - sentence-transformers - sentence-similarity - feature-extraction - dense - modernbert - code - code-search - code-retrieval datasets: - Shuu12121/owl_code_search_hard_negative_datasets_V2_kd - Shuu12121/owl_code_search_hard_negative_datasets_additional_languages_V2_kd - Shuu12121/codeedit_hard_negative_datasets_kd widget: - source_sentence: "Extract json from string with support for '' and None." sentences: - "def json_loads(data: Any) -> dict[str, Any]:\n \"\"\"Extract json from string with support for '' and None.\"\"\"\n if not data:\n return {}\n try:\n return json_loads_util(data)\n except json.JSONDecodeError as err:\n raise APIError(\"Invalid json\") from err" - "def str2json(v):\n \"\"\"\n convert str to json data\n :param v:\n :return:\n \"\"\"\n try:\n return json.loads(v)\n except:\n return None" - "def extract_json_from_string(response_msg: str) -> str:\n \"\"\"\n Attempts to extract JSON (object or array) from within a larger string, not specific to markdown.\n \"\"\"\n json_pattern = re.compile(r\"\\{.*\\}|\\[.*\\]\")\n match = json_pattern.search(response_msg)\n if match:\n return match.group(0)\n\n return response_msg" - "def parse_json(val: str):\n \"\"\"\n Parses json if string else return\n \"\"\"\n if isinstance(val, str):\n val = json.loads(val)\n if isinstance(val, dict):\n val = frappe._dict(val)\n return val" --- # NightJar-large-CodeSearch-Embedding **NightJar-large-CodeSearch-Embedding** is a code-retrieval embedding model fine-tuned from [`Shuu12121/NightJar-large`](https://huggingface.co/Shuu12121/NightJar-large) (a 346M-parameter ModernBERT encoder pre-trained from scratch on code) with [Sentence Transformers](https://www.SBERT.net). It maps text and code to a 1024-dimensional vector space, so natural-language queries, code snippets and code edits can be compared with cosine similarity. Use it for **natural-language-to-code search**, **code-to-code retrieval** and **code-edit search**. | | | |---|---| | Base model | [`Shuu12121/NightJar-large`](https://huggingface.co/Shuu12121/NightJar-large) | | Model type | Sentence Transformer (Transformer + `[CLS]` pooling) | | Parameters | 346M | | Output dimension | 1024 | | Max sequence length | 1,024 tokens | | Similarity function | Cosine similarity | | Task prefix | none | ## Results (MTEB, nDCG@10) Scores were computed with MTEB 2.5.1. All models below were fine-tuned with the same recipe (see [Training](#training)). **CodeSearchNetRetrieval** | Model | Go | Java | JavaScript | PHP | Python | Ruby | Avg | |---|---:|---:|---:|---:|---:|---:|---:| | **NightJar-large-CodeSearch-Embedding** | **0.9682** | **0.9324** | **0.8317** | **0.8989** | **0.9532** | **0.8850** | **0.9116** | | [NightJar-CodeSearch-Embedding](https://huggingface.co/Shuu12121/NightJar-CodeSearch-Embedding) (smaller NightJar) | 0.9656 | 0.9278 | 0.8236 | 0.8915 | 0.9410 | 0.8663 | 0.9026 | | NightOwl (same recipe) | 0.9674 | 0.9232 | 0.8157 | 0.8907 | 0.9393 | 0.8618 | 0.8997 | **CodeEditSearchRetrieval** | Model | C | C++ | Go | Java | JS | PHP | Python | Ruby | Rust | Scala | Shell | Swift | TS | Avg | |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:| | **NightJar-large-CodeSearch-Embedding** | **0.7220** | **0.7537** | **0.8062** | **0.7831** | **0.7846** | **0.7743** | **0.8048** | **0.8160** | **0.7396** | **0.8248** | **0.7571** | **0.7882** | **0.8145** | **0.7822** | | [NightJar-CodeSearch-Embedding](https://huggingface.co/Shuu12121/NightJar-CodeSearch-Embedding) (smaller NightJar) | 0.6687 | 0.7105 | 0.7716 | 0.7365 | 0.7386 | 0.7162 | 0.7711 | 0.7723 | 0.6900 | 0.7859 | 0.7119 | 0.7470 | 0.7759 | 0.7382 | | NightOwl (same recipe) | 0.6624 | 0.7170 | 0.7800 | 0.7509 | 0.7539 | 0.7339 | 0.7716 | 0.7814 | 0.7100 | 0.7795 | 0.7220 | 0.7453 | 0.7890 | 0.7459 | NightJar-large-CodeSearch-Embedding scores highest on every language of both benchmarks. "NightJar-CodeSearch-Embedding" is the smaller [`Shuu12121/NightJar`](https://huggingface.co/Shuu12121/NightJar) fine-tuned with exactly this setup for the full epoch. "NightOwl (same recipe)" is [`Shuu12121/NightOwl`](https://huggingface.co/Shuu12121/NightOwl) fine-tuned with the same losses and hyperparameters (without the additional-languages dataset). The fine-tuning data is the same decontaminated set used for NightOwl-CodeEmbedding: overlaps with the CodeSearchNet test splits, and between the commitpackft-derived code-edit data and the CodeEditSearchRetrieval evaluation examples, were removed before training. MTEB names the CodeEditSearchRetrieval evaluation split `train`, but those examples were not used for fine-tuning. ## Usage ```bash pip install -U sentence-transformers ``` ```python from sentence_transformers import SentenceTransformer model = SentenceTransformer("Shuu12121/NightJar-large-CodeSearch-Embedding") sentences = [ # query "Optional. The amount of time that a device will be initially allocated\n" "for. This can eventually be extended with the UpdateDeviceSession RPC.\n" "Default: 15 minutes.\n\n" "Generated from protobuf field .google.protobuf.Duration ttl = 13 " "[(.google.api.field_behavior) = OPTIONAL];\n" "@return \\Google\\Protobuf\\Duration|null", # matching code "public function getTtl()\n {\n return $this->readOneof(13);\n }", # similar code, wrong field "public function getTtl()\n {\n return $this->readOneof(7);\n }", ] embeddings = model.encode(sentences) print(embeddings.shape) # (3, 1024) similarities = model.similarity(embeddings, embeddings) print(similarities) # tensor([[1.0000, 0.8610, 0.8172], # [0.8610, 1.0000, 0.9410], # [0.8172, 0.9410, 1.0000]]) ``` The query scores higher against the matching getter (0.861) than against the near-identical one that reads a different field (0.817), even though the two code snippets are very close to each other (0.941). **Input format tips** - Pass queries and code as plain strings, with no task prefix. The fine-tuning data used plain text for queries and plain code for documents, without the `[NL]` / `[Code1]` markers from pre-training. - The model was trained with sequences up to 1,024 tokens. Longer inputs are truncated. - The model was trained in both directions (query → code and code → query), so it can also be used to find descriptions for code. ## Training The model was fine-tuned for one epoch (3,978 steps, 4,073,472 examples) from `Shuu12121/NightJar-large`. ### Data Each example is an `anchor`, a `positive`, and 15 hard negatives (`negative_1` … `negative_15`). The three datasets were sampled in an 8 : 8 : 1 ratio: | Dataset | Content | |---|---| | [`Shuu12121/owl_code_search_hard_negative_datasets_V2_kd`](https://huggingface.co/datasets/Shuu12121/owl_code_search_hard_negative_datasets_V2_kd) | docstring → code search, 8 languages | | [`Shuu12121/owl_code_search_hard_negative_datasets_additional_languages_V2_kd`](https://huggingface.co/datasets/Shuu12121/owl_code_search_hard_negative_datasets_additional_languages_V2_kd) | docstring → code search, 8 more languages | | [`Shuu12121/codeedit_hard_negative_datasets_kd`](https://huggingface.co/datasets/Shuu12121/codeedit_hard_negative_datasets_kd) | code-edit retrieval, 11 languages | Approximate token lengths, from the first 1,000 samples: anchors average 75 tokens, positives 197 tokens, and negatives 188–215 tokens. The maximum is 1,024 tokens for every column. ### Loss [`CachedMultipleNegativesRankingLoss`](https://sbert.net/docs/package_reference/sentence_transformer/losses.html#cachedmultiplenegativesrankingloss) with: ```json { "scale": 100.0, "similarity_fct": "cos_sim", "mini_batch_size": 64, "gather_across_devices": false, "directions": ["query_to_doc", "doc_to_query"], "partition_mode": "joint", "hardness_mode": null, "hardness_strength": 0.0 } ``` ### Hyperparameters | | | |---|---| | Batch size | 1,024 (no gradient accumulation) | | Learning rate | 6e-5, cosine schedule, no warmup | | Weight decay | 0.01 | | Optimizer | `adamw_torch_fused` (β = 0.9 / 0.999, ε = 1e-8) | | Epochs | 1 | | Precision | bf16 | | Gradient checkpointing | on | | Seed | 42 | ### Training-time evaluation During training, the model was monitored with Sentence Transformers' `InformationRetrievalEvaluator` on CodeSearchNet (six languages, validation split). The values below are from the end of training. | Metric | Go | Java | JavaScript | PHP | Python | Ruby | Avg | |---|---:|---:|---:|---:|---:|---:|---:| | Accuracy@1 | 0.8740 | 0.7630 | 0.7730 | 0.7660 | 0.8900 | 0.8510 | 0.8195 | | Recall@10 | 0.9810 | 0.9350 | 0.9030 | 0.8220 | 0.9890 | 0.9570 | 0.9312 | | MRR@10 | 0.9184 | 0.8309 | 0.8219 | 0.7898 | 0.9302 | 0.8945 | 0.8643 | | **nDCG@10** | **0.9341** | **0.8569** | **0.8420** | **0.7979** | **0.9449** | **0.9101** | **0.8810** | These numbers come from a different evaluation set than MTEB, so they are not comparable with the MTEB scores above. The monitored nDCG@10 rose quickly in the first few hundred steps. After about step 2,000 (roughly half an epoch) it stayed within ±0.01 of its final value for every language, so most of the gain was already reached in the first half of the epoch. ### Framework versions - Python 3.10.12 - Sentence Transformers 5.3.0 - Transformers 4.56.2 - PyTorch 2.8.0+cu128 - Accelerate 1.12.0 - Datasets 3.6.0 - Tokenizers 0.22.1 ## Architecture ```text SentenceTransformer( (0): Transformer({'max_seq_length': 1024, 'do_lower_case': False, 'architecture': 'ModernBertModel'}) (1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True}) ) ``` The encoder is the ModernBERT-based NightJar-large: 28 layers, hidden size 1024, 16 heads, global attention in every 3rd layer and a 128-token sliding window otherwise. The tokenizer keeps indentation and newlines as single tokens. See the [NightJar-large model card](https://huggingface.co/Shuu12121/NightJar-large) for details. ## Limitations - Retrieval quality was evaluated on CodeSearchNetRetrieval and CodeEditSearchRetrieval only. - Performance varies by language with the amount and quality of the fine-tuning data. Languages not covered by the fine-tuning data are not well supported. - Inputs longer than 1,024 tokens are truncated. - Retrieved code may contain bugs, outdated APIs or insecure patterns. Review it before use. ## License The model weights are released under **Apache 2.0**. The training datasets carry their own licenses and terms of use. Check them if your use case depends on the provenance of the training data. ## Citation ```bibtex @inproceedings{reimers-2019-sentence-bert, title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks", author = "Reimers, Nils and Gurevych, Iryna", booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing", month = "11", year = "2019", publisher = "Association for Computational Linguistics", url = "https://arxiv.org/abs/1908.10084", } ``` ```bibtex @misc{gao2021scaling, title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup}, author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan}, year={2021}, eprint={2101.06983}, archivePrefix={arXiv}, primaryClass={cs.LG} } ``` --- ## 日本語版 ### 概要 **NightJar-large-CodeSearch-Embedding**は、[`Shuu12121/NightJar-large`](https://huggingface.co/Shuu12121/NightJar-large)(コードでゼロから事前学習した、約3.46億パラメータのModernBERTエンコーダー)を、[Sentence Transformers](https://www.SBERT.net)でコード検索向けに追加学習した埋め込みモデルです。文章とコードを1024次元のベクトルに変換するので、自然言語の質問、コード断片、コードの変更内容を、コサイン類似度で比較できます。 **自然言語からのコード検索**、**コードからコードの検索**、**コード編集の検索**に使えます。 | 項目 | 内容 | |---|---| | ベースモデル | [`Shuu12121/NightJar-large`](https://huggingface.co/Shuu12121/NightJar-large) | | モデルの種類 | Sentence Transformer(Transformer + `[CLS]`プーリング) | | パラメータ数 | 約3.46億 | | 出力次元 | 1024 | | 最大系列長 | 1,024トークン | | 類似度 | コサイン類似度 | | タスクプレフィックス | なし | ### 評価結果(MTEB、nDCG@10) MTEB 2.5.1で計算しました。比較対象のモデルも、すべて同じ手順で追加学習しています(詳細は「学習」を参照)。 | モデル | CodeSearchNetRetrieval(平均) | CodeEditSearchRetrieval(平均) | |---|---:|---:| | **NightJar-large-CodeSearch-Embedding** | **0.9116** | **0.7822** | | NightJar-CodeSearch-Embedding | 0.9026 | 0.7382 | | NightOwl(同じ手順で追加学習) | 0.8997 | 0.7459 | NightJar-large-CodeSearch-Embeddingは、CodeSearchNetRetrievalの6言語、CodeEditSearchRetrievalの13言語のすべてで、最も高いスコアでした。言語別のスコアは英語版の表を参照してください。 追加学習のデータは、NightOwl-CodeEmbeddingと同じ重複除去済みのデータセットです。学習前に、コード検索データとCodeSearchNetのテストデータの重複、commitpackft由来のコード編集データとCodeEditSearchRetrievalの評価データの重複を取り除いています。MTEBではCodeEditSearchRetrievalの評価データが`train`という名前になっていますが、この評価データは追加学習には使っていません。 ### 使い方 英語版の「Usage」にサンプルコードがあります。 - **最大長:** 学習時の最大長は1,024トークンです。それより長い入力は切り捨てられます。 - **双方向:** 質問→コードと、コード→質問の両方向で学習しているので、コードに合う説明文を探す用途にも使えます。 ### 学習 `Shuu12121/NightJar-large`から、1エポック(3,978ステップ、4,073,472件)で追加学習しました。 **データ:** 各サンプルは、アンカー1件、正例1件、ハードネガティブ15件です。次の3つのデータセットを 8 : 8 : 1 の比率でサンプリングしました。 - `owl_code_search_hard_negative_datasets_V2_kd`(docstringからコードの検索、8言語) - `owl_code_search_hard_negative_datasets_additional_languages_V2_kd`(同じく追加の8言語) - `codeedit_hard_negative_datasets_kd`(コード編集の検索、11言語) **損失関数:** Cached MultipleNegativesRankingLoss(scale 100、mini batch size 64、batch_size 1024)質問→文書と文書→質問の双方向で学習しています。 **ハイパーパラメータ:** | 項目 | 値 | |---|---| | バッチサイズ | 1,024(勾配蓄積なし) | | 学習率 | 6e-5(コサイン減衰、ウォームアップなし) | | Weight decay | 0.01 | | 最適化手法 | `adamw_torch_fused` | | エポック数 | 1 | | 精度 | bf16 | | Gradient checkpointing | あり | | シード | 42 | **学習中の評価:** 学習中は、Sentence Transformersの`InformationRetrievalEvaluator`でCodeSearchNet(6言語、validation splitにおける性能)を監視しました。学習終了時のnDCG@10は、Go 0.9341、Java 0.8569、JavaScript 0.8420、PHP 0.7979、Python 0.9449、Ruby 0.9101で、平均は0.8810です(そのほかの指標は英語版の表を参照)。この数値はMTEBとは別の評価セットで計算したものなので、MTEBのスコアとは比較できません。 監視していたnDCG@10は、最初の数百ステップで急に上がり、約2,000ステップ(約0.5エポック)以降は、どの言語も最終値の±0.01以内で推移しました。 ### 制限事項 - 検索性能の評価は、CodeSearchNetRetrievalとCodeEditSearchRetrievalのみです。 - 追加学習データの量や質などが言語ごとに違うため、性能にも差があります。追加学習データに含まれない言語では性能が下がる場合があります。 - 1,024トークンを超える入力は切り捨てられます。 - 検索で得られたコードには、バグ、古いAPI、安全でない実装が含まれる可能性があります。利用前に確認してください。 ### ライセンス モデルの重みは**Apache 2.0**で公開しています。学習データセットには、それぞれ独自のライセンスと利用条件があります。用途によっては、各データセットの条件も確認してください。