Transformers
Italian
matformer
bert
encoder
italian
masked-language-modeling
long-context
alibi
custom_code
Instructions to use AlBERTurin/AlBERTmini with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AlBERTurin/AlBERTmini with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("AlBERTurin/AlBERTmini", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from AlBERTurin/AlBERTmini: direct link, hf CLI and curl.
- Browser
- Download file 2.09 kB
-
https://huggingface.co/AlBERTurin/AlBERTmini/resolve/main/README.md
- Command line
-
hf download hf://AlBERTurin/AlBERTmini/README.md
-
curl -L -o README.md https://huggingface.co/AlBERTurin/AlBERTmini/resolve/main/README.md
2.09 kB
| language: | |
| - it | |
| library_name: transformers | |
| tags: | |
| - bert | |
| - encoder | |
| - italian | |
| - masked-language-modeling | |
| - long-context | |
| - alibi | |
| - matformer | |
| # AlBERTmini | |
| **AlBERTmini** is a 95M-parameter Italian encoder model from the | |
| **AlBERTurin** family. | |
| It was trained from scratch on approximately 7B Italian tokens using | |
| masked language modeling. | |
| ## Model Description | |
| AlBERTmini is the smallest model in the AlBERTurin family of | |
| encoder-only Transformer models for Italian. | |
| The model incorporates several architectural improvements over the | |
| original BERT architecture, including Pre-RMSNorm, SwiGLU activations, | |
| ALiBi positional biases, and a mask-only pre-training objective. | |
| AlBERTmini uses: | |
| - 6 Transformer layers | |
| - hidden size of 768 | |
| - 12 attention heads | |
| - SwiGLU activations | |
| - Pre-RMSNorm | |
| - ALiBi positional biases | |
| - 1,024-token training sequence length | |
| - 20% mask-only MLM | |
| - Muon optimizer | |
| The model uses **gettone**, a 32,768-token BPE tokenizer optimized | |
| for Italian and shared across the AlBERTurin model family. | |
| The model was trained using | |
| [Matformer](https://github.com/mrinaldi97/matformer). | |
| ## AlBERTurin Model Family | |
| | Model | Parameters | Training Tokens | | |
| | --- | ---: | ---: | | |
| | **AlBERTmini** | 95M | 7B | | |
| | [AlBERTina](https://huggingface.co/AlBERTurin/AlBERTina) | 140M | 14B | | |
| | [AlBERTone101](https://huggingface.co/AlBERTurin/AlBERTone101) | 450M | ~101B | | |
| ## Installation | |
| ```bash | |
| python -m pip install \ | |
| git+https://github.com/mrinaldi97/matformer.git@alberturin-v1 | |
| ``` | |
| ## Usage | |
| ```python | |
| from transformers import AutoTokenizer, AutoModelForMaskedLM | |
| model_id = "AlBERTurin/AlBERTmini" | |
| tokenizer = AutoTokenizer.from_pretrained(model_id) | |
| model = AutoModelForMaskedLM.from_pretrained( | |
| model_id, | |
| trust_remote_code=True, | |
| ) | |
| ``` | |
| ## Citation | |
| If you use AlBERTmini in your research, please cite: | |
| > Matteo Rinaldi, Marco Madeddu, Calogero Jerik Scozzaro, | |
| > Matteo Delsanto, Daniele Paolo Radicioni, and Viviana Patti. | |
| > **AlBERTurin: A Fully Open Family of Italian Encoder Models with | |
| > Modern Architectures.** | |
| > CLiC-it 2026. |