Fill-Mask
Transformers
Safetensors
Vietnamese
modernbert
feature-extraction
vietnamese
encoder
masked-language-modeling
mlm
pretrained-from-scratch
paxiai
Instructions to use PaxiAI/Vietnamese-Encoder-Base-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PaxiAI/Vietnamese-Encoder-Base-v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="PaxiAI/Vietnamese-Encoder-Base-v1")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("PaxiAI/Vietnamese-Encoder-Base-v1") model = AutoModel.from_pretrained("PaxiAI/Vietnamese-Encoder-Base-v1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from PaxiAI/Vietnamese-Encoder-Base-v1: direct link, hf CLI and curl.
- Browser
- Download file 2.56 kB
-
https://huggingface.co/PaxiAI/Vietnamese-Encoder-Base-v1/resolve/main/README.md
- Command line
-
hf download hf://PaxiAI/Vietnamese-Encoder-Base-v1/README.md
-
curl -L -H "Authorization: Bearer $HF_TOKEN" -o README.md https://huggingface.co/PaxiAI/Vietnamese-Encoder-Base-v1/resolve/main/README.md
2.56 kB
| language: | |
| - vi | |
| license: apache-2.0 | |
| library_name: transformers | |
| pipeline_tag: fill-mask | |
| tags: | |
| - vietnamese | |
| - encoder | |
| - masked-language-modeling | |
| - mlm | |
| - modernbert | |
| - pretrained-from-scratch | |
| - paxiai | |
| # PaxiAI Vietnamese Encoder Base v1 | |
| **PaxiAI Vietnamese Encoder Base v1** is a Vietnamese-focused bidirectional Transformer encoder pretrained **from scratch** using Masked Language Modeling (MLM). It uses the architecture principles of ModernBERT, all model parameters trained from scratch with the **PaxiAI Vietnamese Tokenizer**. | |
| This model is intended to serve as a reusable foundation encoder for Vietnamese NLP tasks such as: | |
| - semantic embeddings | |
| - information retrieval | |
| - reranking | |
| - text classification | |
| - intent detection | |
| - named entity recognition | |
| - semantic similarity | |
| - sentence representation | |
| - domain-specific encoder models | |
| --- | |
| ## Model Overview | |
| | Property | Value | | |
| |---|---| | |
| | Architecture | Bidirectional Transformer Encoder | | |
| | Training objective | Masked Language Modeling | | |
| | Initialization | Random weights | | |
| | Hidden size | 768 | | |
| | Transformer layers | 12 | | |
| | Attention heads | 12 | | |
| | Intermediate size | 2048 | | |
| | Vocabulary size | 48,000 | | |
| | Maximum sequence length | 2,048 | | |
| | Primary language | Vietnamese | | |
| | Tokenizer | PaxiAI/Vietnamese-Tokenizer | | |
| | Pretraining steps | 200,000 | | |
| | Pretrained model dependency | None | | |
| --- | |
| ## Tokenizer | |
| The model uses: | |
| **PaxiAI/Vietnamese-Tokenizer** | |
| https://huggingface.co/PaxiAI/Vietnamese-Tokenizer | |
| The tokenizer has a vocabulary size of approximately 48K tokens and was designed primarily for Vietnamese text while retaining support for English and programming-related text. | |
| For MLM pretraining, one of the tokenizer's reserved tokens is used as the mask token without modifying the original vocabulary or token ID mapping. | |
| --- | |
| ## Loading the Model | |
| ```python | |
| from transformers import AutoTokenizer, AutoModel | |
| model_id = "PaxiAI/Vietnamese-Encoder-Base-v1" | |
| tokenizer = AutoTokenizer.from_pretrained(model_id) | |
| model = AutoModel.from_pretrained(model_id) | |
| text = "Trí tuệ nhân tạo đang phát triển rất nhanh." | |
| inputs = tokenizer( | |
| text, | |
| return_tensors="pt", | |
| truncation=True, | |
| max_length=512, | |
| ) | |
| outputs = model(**inputs) | |
| print(outputs.last_hidden_state.shape) | |
| ``` | |
| The output contains contextual token representations: | |
| ``` | |
| [batch_size, sequence_length, 768] | |
| ``` | |
| ## License | |
| Please refer to the license associated with this repository and verify the licenses and usage conditions of the individual datasets used during pretraining. |