File size: 1,618 Bytes
a61aaee
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
---
license: apache-2.0
library_name: transformers
tags:
- text-classification
- roberta
- hyperparameter-sweep
---
# SweepBestModel

<div align="center">
  <img src="figures/fig1.png" width="70%" alt="SweepBestModel overview" />
</div>

## Overview

SweepBestModel is a RoBERTa-base model fine-tuned for sequence classification through a systematic hyperparameter sweep. We explored learning rates and weight decay values to find the optimal configuration.

## Training Configuration

| Run | Learning Rate | Weight Decay | Best Checkpoint | Best F1 |
|---|---|---|---|---|
| run_lr2e-5_wd0.01 | 2e-5 | 0.01 | \u2014 | {RESULT} |
| run_lr5e-5_wd0.01 | 5e-5 | 0.01 | \u2014 | {RESULT} |
| run_lr1e-4_wd0.01 | 1e-4 | 0.01 | \u2014 | {RESULT} |
| run_lr2e-5_wd0.1  | 2e-5 | 0.1  | \u2014 | {RESULT} |

## Sweep Results

<div align="center">

| Run | Learning Rate | Weight Decay | Best Eval F1 |
|---|---|---|---|
| run_lr2e-5_wd0.01 | 2e-5 | 0.01 | 0.827 |
| run_lr5e-5_wd0.01 | 5e-5 | 0.01 | 0.856 |
| run_lr1e-4_wd0.01 | 1e-4 | 0.01 | 0.793 |
| run_lr2e-5_wd0.1  | 2e-5 | 0.1  | 0.741 |

</div>

<p align="center">
  <img width="60%" src="figures/fig2.png">
</p>

The best performing configuration used a learning rate of 5e-5 with weight decay 0.01, achieving the highest F1 score across all sweep runs.

## Usage

```python
from transformers import AutoModelForSequenceClassification, AutoTokenizer

model = AutoModelForSequenceClassification.from_pretrained("SweepBest-TestRepo")
tokenizer = AutoTokenizer.from_pretrained("SweepBest-TestRepo")
```

## License

This model is released under the Apache 2.0 license.