You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

PaxiAI Vietnamese Encoder Base v1

PaxiAI Vietnamese Encoder Base v1 is a Vietnamese-focused bidirectional Transformer encoder pretrained from scratch using Masked Language Modeling (MLM). It uses the architecture principles of ModernBERT, all model parameters trained from scratch with the PaxiAI Vietnamese Tokenizer.

This model is intended to serve as a reusable foundation encoder for Vietnamese NLP tasks such as:

  • semantic embeddings
  • information retrieval
  • reranking
  • text classification
  • intent detection
  • named entity recognition
  • semantic similarity
  • sentence representation
  • domain-specific encoder models

Model Overview

Property Value
Architecture Bidirectional Transformer Encoder
Training objective Masked Language Modeling
Initialization Random weights
Hidden size 768
Transformer layers 12
Attention heads 12
Intermediate size 2048
Vocabulary size 48,000
Maximum sequence length 2,048
Primary language Vietnamese
Tokenizer PaxiAI/Vietnamese-Tokenizer
Pretraining steps 200,000
Pretrained model dependency None

Tokenizer

The model uses:

PaxiAI/Vietnamese-Tokenizer

https://huggingface.co/PaxiAI/Vietnamese-Tokenizer

The tokenizer has a vocabulary size of approximately 48K tokens and was designed primarily for Vietnamese text while retaining support for English and programming-related text.

For MLM pretraining, one of the tokenizer's reserved tokens is used as the mask token without modifying the original vocabulary or token ID mapping.


Loading the Model

from transformers import AutoTokenizer, AutoModel

model_id = "PaxiAI/Vietnamese-Encoder-Base-v1"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id)

text = "Trí tuệ nhân tạo đang phát triển rất nhanh."

inputs = tokenizer(
    text,
    return_tensors="pt",
    truncation=True,
    max_length=512,
)

outputs = model(**inputs)

print(outputs.last_hidden_state.shape)

The output contains contextual token representations:

[batch_size, sequence_length, 768]

License

Please refer to the license associated with this repository and verify the licenses and usage conditions of the individual datasets used during pretraining.

Downloads last month
5
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support