kiel2 commited on
Commit
b5a65cb
·
verified ·
1 Parent(s): f6d3ad2

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +34 -284
README.md CHANGED
@@ -10,321 +10,71 @@ base_model: BAAI/bge-reranker-base
10
  pipeline_tag: text-ranking
11
  library_name: sentence-transformers
12
  ---
 
13
  # KielEmbed-Rerank
14
- # A CrossEncoder based on BAAI/bge-reranker-base
15
 
16
- This is a [Cross Encoder](https://www.sbert.net/docs/cross_encoder/usage/usage.html) model finetuned from [BAAI/bge-reranker-base](https://huggingface.co/BAAI/bge-reranker-base) using the [sentence-transformers](https://www.SBERT.net) library. It computes scores for pairs of texts, which can be used for text reranking and semantic search.
17
 
18
  ## Model Details
19
 
20
  ### Model Description
21
- - **Model Type:** Cross Encoder
22
- - **Base model:** [BAAI/bge-reranker-base](https://huggingface.co/BAAI/bge-reranker-base) <!-- at revision 2cfc18c9415c912f9d8155881c133215df768a70 -->
23
  - **Maximum Sequence Length:** 512 tokens
24
- - **Number of Output Labels:** 1 label
25
  - **Supported Modality:** Text
26
- <!-- - **Training Dataset:** Unknown -->
27
- <!-- - **Language:** Unknown -->
28
- <!-- - **License:** Unknown -->
29
 
30
  ### Model Sources
31
-
32
  - **Documentation:** [Sentence Transformers Documentation](https://sbert.net)
33
- - **Documentation:** [Cross Encoder Documentation](https://www.sbert.net/docs/cross_encoder/usage/usage.html)
34
  - **Repository:** [Sentence Transformers on GitHub](https://github.com/huggingface/sentence-transformers)
35
- - **Hugging Face:** [Cross Encoders on Hugging Face](https://huggingface.co/models?library=sentence-transformers&other=cross-encoder)
36
 
37
  ### Full Model Architecture
38
-
39
- ```
40
  CrossEncoder(
41
- (0): Transformer({'transformer_task': 'sequence-classification', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'logits'}}, 'module_output_name': 'scores', 'architecture': 'XLMRobertaForSequenceClassification'})
42
  )
43
- ```
44
-
45
- ## Usage
46
-
47
- ### Direct Usage (Sentence Transformers)
48
-
49
- First install the Sentence Transformers library:
50
 
51
- ```bash
52
- pip install -U sentence-transformers
53
- ```
54
 
55
- Then you can load this model and run inference.
56
- ```python
57
- from sentence_transformers import CrossEncoder
58
-
59
- # Download from the 🤗 Hub
60
- model = CrossEncoder("cross_encoder_model_id")
61
- # Get scores for pairs of inputs
62
  pairs = [
63
- ['" The public is understandably losing patience with these unwanted phone calls , unwanted intrusions , " he said at a White House ceremony .', '" While many good people work in the telemarketing industry , the public is understandably losing patience with these unwanted phone calls , unwanted intrusions , " Mr. Bush said .'],
64
- ['Federal agent Bill Polychronopoulos said it was not known if the man , 30 , would be charged .', 'Federal Agent Bill Polychronopoulos said last night the man involved in the Melbourne incident had been unarmed .'],
65
- ['The companies uniformly declined to give specific numbers on customer turnover , saying they will release those figures only when they report overall company performance at year-end .', 'The companies , however , declined to give specifics on customer turnover , saying they would release figures only when they report their overall company performance .'],
66
- ['Five more human cases of West Nile virus , were reported by the Mesa County Health Department on Wednesday .', 'As of this week , 103 human West Nile cases in 45 counties had been reported to the health department .'],
67
- ['But the signal is designed to make it more difficult for consumers to then transfer those copies to the Internet and make them available to potentially millions of others .', 'FCC officials said the embedded electronic signal is designed to make it more difficult for consumers to then transfer copies to the Internet .'],
 
 
 
68
  ]
69
  scores = model.predict(pairs)
70
  print(scores)
71
- # [0.5923 0.0526 0.9774 0.3382 0.8619]
72
 
73
- # Or rank different texts based on similarity to a single text
74
  ranks = model.rank(
75
- '" The public is understandably losing patience with these unwanted phone calls , unwanted intrusions , " he said at a White House ceremony .',
76
  [
77
- '" While many good people work in the telemarketing industry , the public is understandably losing patience with these unwanted phone calls , unwanted intrusions , " Mr. Bush said .',
78
- 'Federal Agent Bill Polychronopoulos said last night the man involved in the Melbourne incident had been unarmed .',
79
- 'The companies , however , declined to give specifics on customer turnover , saying they would release figures only when they report their overall company performance .',
80
- 'As of this week , 103 human West Nile cases in 45 counties had been reported to the health department .',
81
- 'FCC officials said the embedded electronic signal is designed to make it more difficult for consumers to then transfer copies to the Internet .',
82
  ]
83
  )
84
- # [{'corpus_id': ..., 'score': ...}, {'corpus_id': ..., 'score': ...}, ...]
85
- ```
86
-
87
- <!--
88
- ### Direct Usage (Transformers)
89
-
90
- <details><summary>Click to see the direct usage in Transformers</summary>
91
-
92
- </details>
93
- -->
94
-
95
- <!--
96
- ### Downstream Usage (Sentence Transformers)
97
-
98
- You can finetune this model on your own dataset.
99
-
100
- <details><summary>Click to expand</summary>
101
-
102
- </details>
103
- -->
104
-
105
- <!--
106
- ### Out-of-Scope Use
107
-
108
- *List how the model may foreseeably be misused and address what users ought not to do with the model.*
109
- -->
110
-
111
- <!--
112
- ## Bias, Risks and Limitations
113
-
114
- *What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model.*
115
- -->
116
-
117
- <!--
118
- ### Recommendations
119
-
120
- *What are recommendations with respect to the foreseeable issues? For example, filtering explicit content.*
121
- -->
122
-
123
- ## Training Details
124
-
125
- ### Training Dataset
126
-
127
- #### Unnamed Dataset
128
-
129
- * Size: 5,000 training samples
130
- * Columns: <code>text1</code>, <code>text2</code>, and <code>label</code>
131
- * Approximate statistics based on the first 100 samples:
132
- | | text1 | text2 | label |
133
- |:---------|:-----------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------|:------------------------------------------------|
134
- | type | string | string | int |
135
- | modality | text | text | |
136
- | details | <ul><li>min: 15 tokens</li><li>mean: 33.25 tokens</li><li>max: 58 tokens</li></ul> | <ul><li>min: 16 tokens</li><li>mean: 33.05 tokens</li><li>max: 52 tokens</li></ul> | <ul><li>0: ~34.62%</li><li>1: ~65.38%</li></ul> |
137
- * Samples:
138
- | text1 | text2 | label |
139
- |:-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:---------------|
140
- | <code>" The public is understandably losing patience with these unwanted phone calls , unwanted intrusions , " he said at a White House ceremony .</code> | <code>" While many good people work in the telemarketing industry , the public is understandably losing patience with these unwanted phone calls , unwanted intrusions , " Mr. Bush said .</code> | <code>0</code> |
141
- | <code>Federal agent Bill Polychronopoulos said it was not known if the man , 30 , would be charged .</code> | <code>Federal Agent Bill Polychronopoulos said last night the man involved in the Melbourne incident had been unarmed .</code> | <code>0</code> |
142
- | <code>The companies uniformly declined to give specific numbers on customer turnover , saying they will release those figures only when they report overall company performance at year-end .</code> | <code>The companies , however , declined to give specifics on customer turnover , saying they would release figures only when they report their overall company performance .</code> | <code>1</code> |
143
- * Loss: [<code>BinaryCrossEntropyLoss</code>](https://sbert.net/docs/package_reference/cross_encoder/losses.html#binarycrossentropyloss) with these parameters:
144
- ```json
145
- {
146
- "activation_fn": "torch.nn.modules.linear.Identity",
147
- "pos_weight": null
148
- }
149
- ```
150
-
151
- ### Training Hyperparameters
152
- #### Non-Default Hyperparameters
153
-
154
- - `per_device_train_batch_size`: 4
155
- - `num_train_epochs`: 1
156
- - `learning_rate`: 2e-05
157
- - `warmup_steps`: 0.1
158
- - `gradient_accumulation_steps`: 8
159
- - `fp16`: True
160
-
161
- #### All Hyperparameters
162
- <details><summary>Click to expand</summary>
163
-
164
- - `per_device_train_batch_size`: 4
165
- - `num_train_epochs`: 1
166
- - `max_steps`: -1
167
- - `learning_rate`: 2e-05
168
- - `lr_scheduler_type`: linear
169
- - `lr_scheduler_kwargs`: None
170
- - `warmup_steps`: 0.1
171
- - `optim`: adamw_torch_fused
172
- - `optim_args`: None
173
- - `weight_decay`: 0.0
174
- - `adam_beta1`: 0.9
175
- - `adam_beta2`: 0.999
176
- - `adam_epsilon`: 1e-08
177
- - `optim_target_modules`: None
178
- - `gradient_accumulation_steps`: 8
179
- - `average_tokens_across_devices`: True
180
- - `max_grad_norm`: 1.0
181
- - `label_smoothing_factor`: 0.0
182
- - `bf16`: False
183
- - `fp16`: True
184
- - `bf16_full_eval`: False
185
- - `fp16_full_eval`: False
186
- - `tf32`: None
187
- - `gradient_checkpointing`: False
188
- - `gradient_checkpointing_kwargs`: None
189
- - `torch_compile`: False
190
- - `torch_compile_backend`: None
191
- - `torch_compile_mode`: None
192
- - `use_liger_kernel`: False
193
- - `liger_kernel_config`: None
194
- - `use_cache`: False
195
- - `neftune_noise_alpha`: None
196
- - `torch_empty_cache_steps`: None
197
- - `auto_find_batch_size`: False
198
- - `log_on_each_node`: True
199
- - `logging_nan_inf_filter`: True
200
- - `include_num_input_tokens_seen`: no
201
- - `log_level`: passive
202
- - `log_level_replica`: warning
203
- - `disable_tqdm`: False
204
- - `project`: huggingface
205
- - `trackio_space_id`: None
206
- - `trackio_bucket_id`: None
207
- - `trackio_static_space_id`: None
208
- - `per_device_eval_batch_size`: 8
209
- - `prediction_loss_only`: True
210
- - `eval_on_start`: False
211
- - `eval_do_concat_batches`: True
212
- - `eval_use_gather_object`: False
213
- - `eval_accumulation_steps`: None
214
- - `include_for_metrics`: []
215
- - `batch_eval_metrics`: False
216
- - `save_only_model`: False
217
- - `save_on_each_node`: False
218
- - `enable_jit_checkpoint`: False
219
- - `push_to_hub`: False
220
- - `hub_private_repo`: None
221
- - `hub_model_id`: None
222
- - `hub_strategy`: every_save
223
- - `hub_always_push`: False
224
- - `hub_revision`: None
225
- - `load_best_model_at_end`: False
226
- - `ignore_data_skip`: False
227
- - `restore_callback_states_from_checkpoint`: False
228
- - `full_determinism`: False
229
- - `seed`: 42
230
- - `data_seed`: None
231
- - `use_cpu`: False
232
- - `accelerator_config`: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
233
- - `parallelism_config`: None
234
- - `dataloader_drop_last`: False
235
- - `dataloader_num_workers`: 0
236
- - `dataloader_pin_memory`: True
237
- - `dataloader_persistent_workers`: False
238
- - `dataloader_prefetch_factor`: None
239
- - `dataloader_multiprocessing_context`: None
240
- - `dataloader_in_order`: True
241
- - `remove_unused_columns`: True
242
- - `label_names`: None
243
- - `train_sampling_strategy`: random
244
- - `length_column_name`: length
245
- - `ddp_find_unused_parameters`: None
246
- - `ddp_bucket_cap_mb`: None
247
- - `ddp_broadcast_buffers`: False
248
- - `ddp_static_graph`: None
249
- - `ddp_backend`: None
250
- - `ddp_timeout`: 1800
251
- - `fsdp`: None
252
- - `fsdp_config`: None
253
- - `deepspeed`: None
254
- - `debug`: []
255
- - `skip_memory_metrics`: True
256
- - `do_predict`: False
257
- - `resume_from_checkpoint`: None
258
- - `local_rank`: -1
259
- - `prompts`: None
260
- - `batch_sampler`: batch_sampler
261
- - `multi_dataset_batch_sampler`: proportional
262
- - `router_mapping`: {}
263
- - `learning_rate_mapping`: {}
264
- - `warmup_ratio`: None
265
-
266
- </details>
267
-
268
- ### Training Logs
269
- | Epoch | Step | Training Loss |
270
- |:-----:|:----:|:-------------:|
271
- | 0.16 | 25 | 1.7522 |
272
- | 0.32 | 50 | 0.4195 |
273
- | 0.48 | 75 | 0.3522 |
274
- | 0.64 | 100 | 0.3646 |
275
- | 0.8 | 125 | 0.3639 |
276
- | 0.96 | 150 | 0.3551 |
277
-
278
-
279
- ### Training Time
280
- - **Training**: 5.2 minutes
281
-
282
- ### Framework Versions
283
- - Python: 3.13.15
284
- - Sentence Transformers: 5.7.0
285
- - Transformers: 5.16.1
286
- - PyTorch: 2.11.0+cu128
287
- - Accelerate: 1.14.0
288
- - Datasets: 4.8.5
289
- - Tokenizers: 0.23.1
290
-
291
- ## Additional Resources
292
-
293
- - [Training and Finetuning Reranker Models with Sentence Transformers](https://huggingface.co/blog/train-reranker): the end-to-end guide for training or finetuning Cross Encoder (reranker) models.
294
- - [Multimodal Embedding & Reranker Models with Sentence Transformers](https://huggingface.co/blog/multimodal-sentence-transformers): use text, image, audio, and video reranker models through the same API.
295
- - [Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers](https://huggingface.co/blog/train-multimodal-sentence-transformers): training multimodal Cross Encoders.
296
-
297
- ## Citation
298
-
299
- ### BibTeX
300
-
301
- #### Sentence Transformers
302
- ```bibtex
303
- @inproceedings{reimers-2019-sentence-bert,
304
  title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
305
  author = "Reimers, Nils and Gurevych, Iryna",
306
  booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
307
  month = "11",
308
  year = "2019",
309
  publisher = "Association for Computational Linguistics",
310
- url = "https://arxiv.org/abs/1908.10084",
311
  }
312
- ```
313
-
314
- <!--
315
- ## Glossary
316
-
317
- *Clearly define terms in order to be accessible across audiences.*
318
- -->
319
-
320
- <!--
321
- ## Model Card Authors
322
-
323
- *Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction.*
324
- -->
325
-
326
- <!--
327
- ## Model Card Contact
328
-
329
- *Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors.*
330
- -->
 
10
  pipeline_tag: text-ranking
11
  library_name: sentence-transformers
12
  ---
13
+
14
  # KielEmbed-Rerank
15
+ A high-performance Cross-Encoder reranker fine-tuned from `BAAI/bge-reranker-base`.
16
 
17
+ This is a [Cross Encoder](https://www.sbert.net/docs/cross_encoder/usage/usage.html) model fine-tuned from [BAAI/bge-reranker-base](https://huggingface.co/BAAI/bge-reranker-base) using the [sentence-transformers](https://www.SBERT.net) library. It evaluates query-document text pairs jointly to compute precise relevance scores for semantic search and multi-stage retrieval pipelines.
18
 
19
  ## Model Details
20
 
21
  ### Model Description
22
+ - **Model Type:** Cross Encoder / Relevance Reranker
23
+ - **Base Model:** [BAAI/bge-reranker-base](https://huggingface.co/BAAI/bge-reranker-base)
24
  - **Maximum Sequence Length:** 512 tokens
25
+ - **Number of Output Labels:** 1 label (Regression / Logits score)
26
  - **Supported Modality:** Text
 
 
 
27
 
28
  ### Model Sources
 
29
  - **Documentation:** [Sentence Transformers Documentation](https://sbert.net)
30
+ - **Cross Encoder Guide:** [Cross Encoder Documentation](https://www.sbert.net/docs/cross_encoder/usage/usage.html)
31
  - **Repository:** [Sentence Transformers on GitHub](https://github.com/huggingface/sentence-transformers)
 
32
 
33
  ### Full Model Architecture
34
+ ```text
 
35
  CrossEncoder(
36
+ (0): Transformer({'transformer_task': 'sequence-classification', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'logits'}}, 'module_output_name': 'scores', 'architecture': 'BertForSequenceClassification'})
37
  )
38
+ UsageDirect Usage (Sentence Transformers)First, install the Sentence Transformers library:Bashpip install -U sentence-transformers
39
+ Then load your deployed model and run inference:Pythonfrom sentence_transformers import CrossEncoder
 
 
 
 
 
40
 
41
+ # Load your custom reranker from the Hugging Face Hub
42
+ model = CrossEncoder("kiel/KielEmbed-Rerank")
 
43
 
44
+ # Get relevance scores for pairs of inputs (query, document)
 
 
 
 
 
 
45
  pairs = [
46
+ [
47
+ '" The public is understandably losing patience with these unwanted phone calls , unwanted intrusions , " he said at a White House ceremony .',
48
+ '" While many good people work in the telemarketing industry , the public is understandably losing patience with these unwanted phone calls , unwanted intrusions , " Mr. Bush said .'
49
+ ],
50
+ [
51
+ 'Federal agent Bill Polychronopoulos said it was not known if the man , 30 , would be charged .',
52
+ 'Federal Agent Bill Polychronopoulos said last night the man involved in the Melbourne incident had been unarmed .'
53
+ ],
54
  ]
55
  scores = model.predict(pairs)
56
  print(scores)
 
57
 
58
+ # Alternatively, rank an array of candidate texts against a single query
59
  ranks = model.rank(
60
+ 'The public is losing patience with unwanted phone calls.',
61
  [
62
+ 'While many good people work in telemarketing, the public is losing patience with unwanted intrusions.',
63
+ 'Federal agents reported an unarmed suspect in the downtown incident.',
64
+ 'Corporate earnings reports will be released at the close of the fiscal year.'
 
 
65
  ]
66
  )
67
+ print(ranks)
68
+ Training DetailsTraining DatasetKiel Reranker CorpusSize: 5,000 training samplesColumns: text1, text2, and labelDataset Statistics:Text 1: Min: 15 tokens | Mean: 33.25 tokens | Max: 58 tokensText 2: Min: 16 tokens | Mean: 33.05 tokens | Max: 52 tokensLabel Distribution: Class 0 (~34.62%), Class 1 (~65.38%)Loss Function: BinaryCrossEntropyLoss with parameters:JSON{
69
+ "activation_fn": "torch.nn.modules.linear.Identity",
70
+ "pos_weight": null
71
+ }
72
+ Training HyperparametersPer Device Train Batch Size: 4Gradient Accumulation Steps: 8 (Effective batch size = 32)Learning Rate: 2e-05Number of Epochs: 1Warmup Steps: 0.1Mixed Precision: FP16 EnabledOptimizer: adamw_torch_fusedTraining LogsEpochStepTraining Loss0.16251.75220.32500.41950.48750.35220.641000.36460.81250.36390.961500.3551Total Training Time: 5.2 minutesFramework VersionsPython: 3.13.15Sentence Transformers: 5.7.0Transformers: 5.16.1PyTorch: 2.11.0+cu128Accelerate: 1.14.0Datasets: 4.8.5Tokenizers: 0.23.1Additional ResourcesTraining and Finetuning Reranker Models with Sentence Transformers: Comprehensive guide for training custom Cross-Encoders.CitationBibTeXCode snippet@inproceedings{reimers-2019-sentence-bert,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
73
  title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
74
  author = "Reimers, Nils and Gurevych, Iryna",
75
  booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
76
  month = "11",
77
  year = "2019",
78
  publisher = "Association for Computational Linguistics",
79
+ url = "[https://arxiv.org/abs/1908.10084](https://arxiv.org/abs/1908.10084)",
80
  }