kiel2 commited on
Commit
d40c276
·
verified ·
1 Parent(s): ea5713a

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +18 -269
README.md CHANGED
@@ -58,54 +58,39 @@ widget:
58
  pipeline_tag: sentence-similarity
59
  library_name: sentence-transformers
60
  ---
 
61
  # KielEmbed-Code
62
- # A SentenceTransformer based on microsoft/codebert-base
63
 
64
- This is a [sentence-transformers](https://www.SBERT.net) model finetuned from [microsoft/codebert-base](https://huggingface.co/microsoft/codebert-base). It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, classification, clustering, and more.
65
 
66
  ## Model Details
67
 
68
  ### Model Description
69
- - **Model Type:** Sentence Transformer
70
- - **Base model:** [microsoft/codebert-base](https://huggingface.co/microsoft/codebert-base) <!-- at revision 3b0952feddeffad0063f274080e3c23d75e7eb39 -->
71
  - **Maximum Sequence Length:** 512 tokens
72
  - **Output Dimensionality:** 768 dimensions
73
  - **Similarity Function:** Cosine Similarity
74
- - **Supported Modality:** Text
75
- <!-- - **Training Dataset:** Unknown -->
76
- <!-- - **Language:** Unknown -->
77
- <!-- - **License:** Unknown -->
78
 
79
  ### Model Sources
80
-
81
  - **Documentation:** [Sentence Transformers Documentation](https://sbert.net)
82
  - **Repository:** [Sentence Transformers on GitHub](https://github.com/huggingface/sentence-transformers)
83
- - **Hugging Face:** [Sentence Transformers on Hugging Face](https://huggingface.co/models?library=sentence-transformers)
84
 
85
  ### Full Model Architecture
86
-
87
- ```
88
  SentenceTransformer(
89
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'RobertaModel'})
90
  (1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'mean', 'include_prompt': True})
91
  )
92
- ```
93
-
94
- ## Usage
95
-
96
- ### Direct Usage (Sentence Transformers)
97
-
98
- First install the Sentence Transformers library:
99
 
100
- ```bash
101
- pip install -U sentence-transformers
102
- ```
103
- Then you can load this model and run inference.
104
- ```python
105
- from sentence_transformers import SentenceTransformer
106
 
107
- # Download from the 🤗 Hub
108
- model = SentenceTransformer("sentence_transformers_model_id")
109
  # Run inference
110
  sentences = [
111
  '" Any decision on Charleroi will have huge implications for regional airports in France , " he said .',
@@ -119,252 +104,16 @@ print(embeddings.shape)
119
  # Get the similarity scores for the embeddings
120
  similarities = model.similarity(embeddings, embeddings)
121
  print(similarities)
122
- # tensor([[1.0000, 0.9671, 0.9596],
123
- # [0.9671, 1.0000, 0.9494],
124
- # [0.9596, 0.9494, 1.0000]])
125
- ```
126
- <!--
127
- ### Direct Usage (Transformers)
128
-
129
- <details><summary>Click to see the direct usage in Transformers</summary>
130
-
131
- </details>
132
- -->
133
-
134
- <!--
135
- ### Downstream Usage (Sentence Transformers)
136
-
137
- You can finetune this model on your own dataset.
138
-
139
- <details><summary>Click to expand</summary>
140
-
141
- </details>
142
- -->
143
-
144
- <!--
145
- ### Out-of-Scope Use
146
-
147
- *List how the model may foreseeably be misused and address what users ought not to do with the model.*
148
- -->
149
-
150
- <!--
151
- ## Bias, Risks and Limitations
152
-
153
- *What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model.*
154
- -->
155
-
156
- <!--
157
- ### Recommendations
158
-
159
- *What are recommendations with respect to the foreseeable issues? For example, filtering explicit content.*
160
- -->
161
-
162
- ## Training Details
163
-
164
- ### Training Dataset
165
-
166
- #### Unnamed Dataset
167
-
168
- * Size: 10,000 training samples
169
- * Columns: <code>text1</code>, <code>text2</code>, and <code>label</code>
170
- * Approximate statistics based on the first 100 samples:
171
- | | text1 | text2 | label |
172
- |:---------|:----------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------|:------------------------------------------------|
173
- | type | string | string | int |
174
- | modality | text | text | |
175
- | details | <ul><li>min: 12 tokens</li><li>mean: 27.7 tokens</li><li>max: 41 tokens</li></ul> | <ul><li>min: 15 tokens</li><li>mean: 27.64 tokens</li><li>max: 46 tokens</li></ul> | <ul><li>0: ~34.62%</li><li>1: ~65.38%</li></ul> |
176
- * Samples:
177
- | text1 | text2 | label |
178
- |:-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:---------------|
179
- | <code>" The public is understandably losing patience with these unwanted phone calls , unwanted intrusions , " he said at a White House ceremony .</code> | <code>" While many good people work in the telemarketing industry , the public is understandably losing patience with these unwanted phone calls , unwanted intrusions , " Mr. Bush said .</code> | <code>0</code> |
180
- | <code>Federal agent Bill Polychronopoulos said it was not known if the man , 30 , would be charged .</code> | <code>Federal Agent Bill Polychronopoulos said last night the man involved in the Melbourne incident had been unarmed .</code> | <code>0</code> |
181
- | <code>The companies uniformly declined to give specific numbers on customer turnover , saying they will release those figures only when they report overall company performance at year-end .</code> | <code>The companies , however , declined to give specifics on customer turnover , saying they would release figures only when they report their overall company performance .</code> | <code>1</code> |
182
- * Loss: [<code>CosineSimilarityLoss</code>](https://sbert.net/docs/package_reference/sentence_transformer/losses.html#cosinesimilarityloss) with these parameters:
183
- ```json
184
- {
185
- "loss_fct": "torch.nn.modules.loss.MSELoss",
186
- "cos_score_transformation": "torch.nn.modules.linear.Identity"
187
- }
188
- ```
189
-
190
- ### Training Hyperparameters
191
- #### Non-Default Hyperparameters
192
-
193
- - `num_train_epochs`: 1
194
- - `learning_rate`: 2e-05
195
- - `warmup_steps`: 0.1
196
- - `gradient_accumulation_steps`: 4
197
- - `fp16`: True
198
-
199
- #### All Hyperparameters
200
- <details><summary>Click to expand</summary>
201
-
202
- - `per_device_train_batch_size`: 8
203
- - `num_train_epochs`: 1
204
- - `max_steps`: -1
205
- - `learning_rate`: 2e-05
206
- - `lr_scheduler_type`: linear
207
- - `lr_scheduler_kwargs`: None
208
- - `warmup_steps`: 0.1
209
- - `optim`: adamw_torch_fused
210
- - `optim_args`: None
211
- - `weight_decay`: 0.0
212
- - `adam_beta1`: 0.9
213
- - `adam_beta2`: 0.999
214
- - `adam_epsilon`: 1e-08
215
- - `optim_target_modules`: None
216
- - `gradient_accumulation_steps`: 4
217
- - `average_tokens_across_devices`: True
218
- - `max_grad_norm`: 1.0
219
- - `label_smoothing_factor`: 0.0
220
- - `bf16`: False
221
- - `fp16`: True
222
- - `bf16_full_eval`: False
223
- - `fp16_full_eval`: False
224
- - `tf32`: None
225
- - `gradient_checkpointing`: False
226
- - `gradient_checkpointing_kwargs`: None
227
- - `torch_compile`: False
228
- - `torch_compile_backend`: None
229
- - `torch_compile_mode`: None
230
- - `use_liger_kernel`: False
231
- - `liger_kernel_config`: None
232
- - `use_cache`: False
233
- - `neftune_noise_alpha`: None
234
- - `torch_empty_cache_steps`: None
235
- - `auto_find_batch_size`: False
236
- - `log_on_each_node`: True
237
- - `logging_nan_inf_filter`: True
238
- - `include_num_input_tokens_seen`: no
239
- - `log_level`: passive
240
- - `log_level_replica`: warning
241
- - `disable_tqdm`: False
242
- - `project`: huggingface
243
- - `trackio_space_id`: None
244
- - `trackio_bucket_id`: None
245
- - `trackio_static_space_id`: None
246
- - `per_device_eval_batch_size`: 8
247
- - `prediction_loss_only`: True
248
- - `eval_on_start`: False
249
- - `eval_do_concat_batches`: True
250
- - `eval_use_gather_object`: False
251
- - `eval_accumulation_steps`: None
252
- - `include_for_metrics`: []
253
- - `batch_eval_metrics`: False
254
- - `save_only_model`: False
255
- - `save_on_each_node`: False
256
- - `enable_jit_checkpoint`: False
257
- - `push_to_hub`: False
258
- - `hub_private_repo`: None
259
- - `hub_model_id`: None
260
- - `hub_strategy`: every_save
261
- - `hub_always_push`: False
262
- - `hub_revision`: None
263
- - `load_best_model_at_end`: False
264
- - `ignore_data_skip`: False
265
- - `restore_callback_states_from_checkpoint`: False
266
- - `full_determinism`: False
267
- - `seed`: 42
268
- - `data_seed`: None
269
- - `use_cpu`: False
270
- - `accelerator_config`: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
271
- - `parallelism_config`: None
272
- - `dataloader_drop_last`: False
273
- - `dataloader_num_workers`: 0
274
- - `dataloader_pin_memory`: True
275
- - `dataloader_persistent_workers`: False
276
- - `dataloader_prefetch_factor`: None
277
- - `dataloader_multiprocessing_context`: None
278
- - `dataloader_in_order`: True
279
- - `remove_unused_columns`: True
280
- - `label_names`: None
281
- - `train_sampling_strategy`: random
282
- - `length_column_name`: length
283
- - `ddp_find_unused_parameters`: None
284
- - `ddp_bucket_cap_mb`: None
285
- - `ddp_broadcast_buffers`: False
286
- - `ddp_static_graph`: None
287
- - `ddp_backend`: None
288
- - `ddp_timeout`: 1800
289
- - `fsdp`: None
290
- - `fsdp_config`: None
291
- - `deepspeed`: None
292
- - `debug`: []
293
- - `skip_memory_metrics`: True
294
- - `do_predict`: False
295
- - `resume_from_checkpoint`: None
296
- - `local_rank`: -1
297
- - `prompts`: None
298
- - `batch_sampler`: batch_sampler
299
- - `multi_dataset_batch_sampler`: proportional
300
- - `router_mapping`: {}
301
- - `learning_rate_mapping`: {}
302
- - `warmup_ratio`: None
303
-
304
- </details>
305
-
306
- ### Training Logs
307
- | Epoch | Step | Training Loss |
308
- |:-----:|:----:|:-------------:|
309
- | 0.16 | 50 | 28.0752 |
310
- | 0.32 | 100 | 24.7668 |
311
- | 0.48 | 150 | 21.0055 |
312
- | 0.64 | 200 | 26.5538 |
313
- | 0.8 | 250 | 26.0600 |
314
- | 0.96 | 300 | 26.3027 |
315
-
316
-
317
- ### Training Time
318
- - **Training**: 5.1 minutes
319
-
320
- ### Framework Versions
321
- - Python: 3.13.15
322
- - Sentence Transformers: 5.7.0
323
- - Transformers: 5.16.1
324
- - PyTorch: 2.11.0+cu128
325
- - Accelerate: 1.14.0
326
- - Datasets: 4.8.5
327
- - Tokenizers: 0.23.1
328
-
329
- ## Additional Resources
330
-
331
- - [Training and Finetuning Embedding Models with Sentence Transformers](https://huggingface.co/blog/train-sentence-transformers): the end-to-end guide for training or finetuning Sentence Transformer models.
332
- - [Introduction to Matryoshka Embedding Models](https://huggingface.co/blog/matryoshka): variable-size embeddings that can be truncated with minimal quality loss.
333
- - [Binary and Scalar Embedding Quantization for Significantly Faster & Cheaper Retrieval](https://huggingface.co/blog/embedding-quantization): post-training compression of embedding vectors.
334
- - [Multimodal Embedding & Reranker Models with Sentence Transformers](https://huggingface.co/blog/multimodal-sentence-transformers): use text, image, audio, and video models through the same API.
335
- - [Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers](https://huggingface.co/blog/train-multimodal-sentence-transformers): train multimodal embedding models, with a Visual Document Retrieval walkthrough.
336
-
337
- ## Citation
338
-
339
- ### BibTeX
340
-
341
- #### Sentence Transformers
342
- ```bibtex
343
- @inproceedings{reimers-2019-sentence-bert,
344
  title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
345
  author = "Reimers, Nils and Gurevych, Iryna",
346
  booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
347
  month = "11",
348
  year = "2019",
349
  publisher = "Association for Computational Linguistics",
350
- url = "https://arxiv.org/abs/1908.10084",
351
  }
352
- ```
353
-
354
- <!--
355
- ## Glossary
356
-
357
- *Clearly define terms in order to be accessible across audiences.*
358
- -->
359
-
360
- <!--
361
- ## Model Card Authors
362
-
363
- *Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction.*
364
- -->
365
-
366
- <!--
367
- ## Model Card Contact
368
-
369
- *Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors.*
370
- -->
 
58
  pipeline_tag: sentence-similarity
59
  library_name: sentence-transformers
60
  ---
61
+
62
  # KielEmbed-Code
63
+ A specialized dense embedding model fine-tuned from `microsoft/codebert-base` for code and text representation.
64
 
65
+ This is a [sentence-transformers](https://www.SBERT.net) model fine-tuned from [microsoft/codebert-base](https://huggingface.co/microsoft/codebert-base). It maps sentences and code blocks into a 768-dimensional dense vector space optimized for semantic textual similarity, semantic search, and clustering tasks.
66
 
67
  ## Model Details
68
 
69
  ### Model Description
70
+ - **Model Type:** Sentence Transformer / Dense Embedding Backbone
71
+ - **Base Model:** [microsoft/codebert-base](https://huggingface.co/microsoft/codebert-base)
72
  - **Maximum Sequence Length:** 512 tokens
73
  - **Output Dimensionality:** 768 dimensions
74
  - **Similarity Function:** Cosine Similarity
75
+ - **Supported Modality:** Text & Code
 
 
 
76
 
77
  ### Model Sources
 
78
  - **Documentation:** [Sentence Transformers Documentation](https://sbert.net)
79
  - **Repository:** [Sentence Transformers on GitHub](https://github.com/huggingface/sentence-transformers)
80
+ - **Hugging Face Hub:** [kiel/KielEmbed-Code](https://huggingface.co/kiel/KielEmbed-Code)
81
 
82
  ### Full Model Architecture
83
+ ```text
 
84
  SentenceTransformer(
85
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'RobertaModel'})
86
  (1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'mean', 'include_prompt': True})
87
  )
88
+ UsageDirect Usage (Sentence Transformers)First, install the Sentence Transformers library:Bashpip install -U sentence-transformers
89
+ Then load your model and run inference:Pythonfrom sentence_transformers import SentenceTransformer
 
 
 
 
 
90
 
91
+ # Load your custom fine-tuned CodeBERT model from the Hugging Face Hub
92
+ model = SentenceTransformer("kiel/KielEmbed-Code")
 
 
 
 
93
 
 
 
94
  # Run inference
95
  sentences = [
96
  '" Any decision on Charleroi will have huge implications for regional airports in France , " he said .',
 
104
  # Get the similarity scores for the embeddings
105
  similarities = model.similarity(embeddings, embeddings)
106
  print(similarities)
107
+ Training DetailsTraining DatasetKiel Code & Semantic CorpusSize: 10,000 training samplesColumns: text1, text2, and labelApproximate Token Statistics (First 100 samples):Text 1: Min: 12 tokens | Mean: 27.7 tokens | Max: 41 tokensText 2: Min: 15 tokens | Mean: 27.64 tokens | Max: 46 tokensLabel Distribution: Class 0 (~34.62%), Class 1 (~65.38%)Loss Function: CosineSimilarityLoss with parameters:JSON{
108
+ "loss_fct": "torch.nn.modules.loss.MSELoss",
109
+ "cos_score_transformation": "torch.nn.modules.linear.Identity"
110
+ }
111
+ Training HyperparametersPer Device Train Batch Size: 8Gradient Accumulation Steps: 4 (Effective batch size = 32)Learning Rate: 2e-05Number of Epochs: 1Warmup Steps: 0.1Mixed Precision: FP16 EnabledOptimizer: adamw_torch_fusedTraining LogsEpochStepTraining Loss0.165028.07520.3210024.76680.4815021.00550.6420026.55380.825026.06000.9630026.3027Total Training Time: 5.1 minutesFramework VersionsPython: 3.13.15Sentence Transformers: 5.7.0Transformers: 5.16.1PyTorch: 2.11.0+cu128Accelerate: 1.14.0Datasets: 4.8.5Tokenizers: 0.23.1Additional ResourcesTraining and Finetuning Embedding Models with Sentence Transformers: End-to-end guide for fine-tuning Sentence Transformer models.CitationBibTeXCode snippet@inproceedings{reimers-2019-sentence-bert,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
112
  title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
113
  author = "Reimers, Nils and Gurevych, Iryna",
114
  booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
115
  month = "11",
116
  year = "2019",
117
  publisher = "Association for Computational Linguistics",
118
+ url = "[https://arxiv.org/abs/1908.10084](https://arxiv.org/abs/1908.10084)",
119
  }