Instructions to use ClassCat/roberta-small-basque with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ClassCat/roberta-small-basque with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="ClassCat/roberta-small-basque")# Load model directly from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained("ClassCat/roberta-small-basque") model = AutoModelForMaskedLM.from_pretrained("ClassCat/roberta-small-basque", device_map="auto") - Notebooks
- Google Colab
- Kaggle
| language: eu | |
| license: cc-by-sa-4.0 | |
| datasets: | |
| - cc100 | |
| - oscar | |
| widget: | |
| - text: "Euria egingo <mask> gaur ?" | |
| - text: "<mask> umeari liburua eman dio." | |
| - text: "Zein da zure <mask> ?" | |
| ## RoBERTa Basque small model (Uncased) | |
| ### Prerequisites | |
| transformers==4.19.2 | |
| ### Model architecture | |
| This model uses approximately half the size of RoBERTa base model parameters. | |
| ### Tokenizer | |
| Using BPE tokenizer with vocabulary size 50,000. | |
| ### Training Data | |
| * Subset of [CC-100/eu](https://data.statmt.org/cc-100/) : Monolingual Datasets from Web Crawl Data | |
| * Subset of [oscar](https://huggingface.co/datasets/oscar) | |
| ### Usage | |
| ```python | |
| from transformers import pipeline | |
| unmasker = pipeline('fill-mask', model='ClassCat/roberta-small-basque') | |
| unmasker("Zein da zure <mask> ?") | |
| ``` |