File size: 3,861 Bytes
8c1ef4e
 
 
 
63957b0
8c1ef4e
 
63957b0
8c1ef4e
 
 
 
 
 
 
 
 
2bdb51b
8c1ef4e
 
 
2bdb51b
8c1ef4e
63957b0
8c1ef4e
 
 
 
 
 
 
63957b0
8c1ef4e
 
 
 
 
 
 
 
 
63957b0
8c1ef4e
63957b0
 
8c1ef4e
 
 
 
 
 
 
63957b0
8c1ef4e
 
 
63957b0
 
8c1ef4e
 
 
63957b0
8c1ef4e
 
 
 
 
 
 
 
 
 
 
63957b0
 
8c1ef4e
 
 
 
63957b0
8c1ef4e
 
63957b0
8c1ef4e
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
---
pipeline_tag: fill-mask
license: mit
base_model: FacebookAI/roberta-base
library_name: zeromodels
tags:
- keras
- zeromodels
- roberta
- fill-mask
- text-encoder
- arxiv:1907.11692
- pytorch
- jax
- tf
---

## ***See [our collection](https://huggingface.co/collections/zeromodels/roberta-6a8eae529a918a955e68211d) for all versions of RoBERTa.***

# Run RoBERTa with Keras 3: JAX, PyTorch, or TensorFlow

[![GitHub](https://img.shields.io/badge/GitHub-ZeroModels-black?logo=github)](https://github.com/IMvision12/ZeroModels) [![Docs](https://img.shields.io/badge/Docs-RoBERTa-blue)](https://imvision12.github.io/ZeroModels/roberta/) [![Collection](https://img.shields.io/badge/HF-RoBERTa%20collection-yellow)](https://huggingface.co/collections/zeromodels/roberta-6a8eae529a918a955e68211d)

# zeromodels/roberta_base

Paper: [RoBERTa: A Robustly Optimized BERT Pretraining Approach (arXiv:1907.11692)](https://arxiv.org/abs/1907.11692) · [HF Papers](https://huggingface.co/papers/1907.11692)

RoBERTa is a robustly optimized BERT encoder: more data/steps, no NSP, dynamic masking, byte-level BPE (mask token `<mask>`), and padding-offset position ids.

For more details on the model, please go to the upstream [model card](https://huggingface.co/FacebookAI/roberta-base).

Pure-**Keras 3** conversion of [`FacebookAI/roberta-base`](https://huggingface.co/FacebookAI/roberta-base) for [zeromodels](https://github.com/IMvision12/ZeroModels). One implementation runs unmodified on **TensorFlow / Torch / JAX**.

This is a **fill-mask / encoder** checkpoint (`RobertaMaskedLM`, base). Task heads load via `hf:` fine-tunes.

## ✨ Quick start (fill-mask)

```python
import os
os.environ["KERAS_BACKEND"] = "torch"  # or "jax" / "tensorflow"

from zeromodels.models.roberta import RobertaMaskedLM, RobertaTokenizer

mlm = RobertaMaskedLM.from_weights("zeromodels/roberta_base")
tokenizer = RobertaTokenizer.from_weights("zeromodels/roberta_base")

inputs = tokenizer("The capital of France is <mask>.")
logits = mlm(inputs)  # (1, L, vocab_size)
mask = int((inputs["input_ids"][0] == tokenizer.mask_token_id).argmax())
print(tokenizer.decode([int(logits[0, mask].argmax())]))
```

Load any RoBERTa variant the same way with `from_weights("zeromodels/<variant>")`:

| Variant | Hub |
|---|---|
| `roberta_base` | [`zeromodels/roberta_base`](https://huggingface.co/zeromodels/roberta_base) |
| `roberta_large` | [`zeromodels/roberta_large`](https://huggingface.co/zeromodels/roberta_large) |

## Available classes

Load any of these from this repo with `from_weights("zeromodels/roberta_base")` (or on the fly via the `hf:` prefix). The pretrained backbone is shared; task heads not stored in this checkpoint start randomly initialized, ready for fine-tuning (or load a `hf:` fine-tune).

| Class | Task |
|---|---|
| `RobertaModel` | Encoder backbone |
| `RobertaMaskedLM` | Masked language modeling (fill-mask) |
| `RobertaSequenceClassify` | Sequence classification |
| `RobertaTokenClassify` | Token classification (NER / POS) |
| `RobertaQnA` | Extractive question answering |
| `RobertaMultipleChoice` | Multiple choice |

```python
from zeromodels.models.roberta import RobertaSequenceClassify
model = RobertaSequenceClassify.from_weights("zeromodels/roberta_base")
```

## Tips

- Set `KERAS_BACKEND` **before** importing Keras / zeromodels.
- Prefer `RobertaTokenizer.from_weights(...)` so BPE vocab matches.
- Use `<mask>` (not `[MASK]`).
- See [RoBERTa docs](https://imvision12.github.io/ZeroModels/roberta/) and [Loading Weights](https://imvision12.github.io/ZeroModels/loading_weights/).
- Community / upstream safetensors still work via the `hf:` prefix, e.g. `RobertaMaskedLM.from_weights("hf:FacebookAI/roberta-base")`.

## Special Thanks

A huge thank you to the Facebook AI RoBERTa authors for creating and releasing these models.

License: MIT.