PaxiAI's picture
Update README.md
9825f31 verified
|
Raw History Blame Contribute Delete
2.56 kB
---
language:
- vi
license: apache-2.0
library_name: transformers
pipeline_tag: fill-mask
tags:
- vietnamese
- encoder
- masked-language-modeling
- mlm
- modernbert
- pretrained-from-scratch
- paxiai
---
# PaxiAI Vietnamese Encoder Base v1
**PaxiAI Vietnamese Encoder Base v1** is a Vietnamese-focused bidirectional Transformer encoder pretrained **from scratch** using Masked Language Modeling (MLM). It uses the architecture principles of ModernBERT, all model parameters trained from scratch with the **PaxiAI Vietnamese Tokenizer**.
This model is intended to serve as a reusable foundation encoder for Vietnamese NLP tasks such as:
- semantic embeddings
- information retrieval
- reranking
- text classification
- intent detection
- named entity recognition
- semantic similarity
- sentence representation
- domain-specific encoder models
---
## Model Overview
| Property | Value |
|---|---|
| Architecture | Bidirectional Transformer Encoder |
| Training objective | Masked Language Modeling |
| Initialization | Random weights |
| Hidden size | 768 |
| Transformer layers | 12 |
| Attention heads | 12 |
| Intermediate size | 2048 |
| Vocabulary size | 48,000 |
| Maximum sequence length | 2,048 |
| Primary language | Vietnamese |
| Tokenizer | PaxiAI/Vietnamese-Tokenizer |
| Pretraining steps | 200,000 |
| Pretrained model dependency | None |
---
## Tokenizer
The model uses:
**PaxiAI/Vietnamese-Tokenizer**
https://huggingface.co/PaxiAI/Vietnamese-Tokenizer
The tokenizer has a vocabulary size of approximately 48K tokens and was designed primarily for Vietnamese text while retaining support for English and programming-related text.
For MLM pretraining, one of the tokenizer's reserved tokens is used as the mask token without modifying the original vocabulary or token ID mapping.
---
## Loading the Model
```python
from transformers import AutoTokenizer, AutoModel
model_id = "PaxiAI/Vietnamese-Encoder-Base-v1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id)
text = "Trí tuệ nhân tạo đang phát triển rất nhanh."
inputs = tokenizer(
text,
return_tensors="pt",
truncation=True,
max_length=512,
)
outputs = model(**inputs)
print(outputs.last_hidden_state.shape)
```
The output contains contextual token representations:
```
[batch_size, sequence_length, 768]
```
## License
Please refer to the license associated with this repository and verify the licenses and usage conditions of the individual datasets used during pretraining.