HmarBERT-mini-v2
📦 Part of the HmarBERT Family Collection — open-source language models for the Hmar language, ranging from 16.9M edge-ready models to full base architectures.
HmarBERT-mini-v2 is an upgraded, 8-layer BERT model trained from scratch for the Hmar language (hmr, ISO 639-3), part of the Zo languages family.
It builds upon HmarBERT-mini (v1) with a deeper architecture, tied word embeddings, a refined 7,000-token compact vocabulary with lower fertility (1.28), and a 3-stage curriculum pre-training schedule using Deterministic Disjoint Whole-Word Masking.
At roughly 29.3M parameters (~58 MB in FP16 / ~117 MB in FP32), it provides significantly stronger syntactic and contextual representations while still running comfortably on modest hardware and laptops.
Model Details
- Architecture: 8-layer BERT encoder (512 hidden size, 8 attention heads, 2048 intermediate FFN)
- Parameters: 29.34M (~58 MB FP16 / ~117 MB FP32)
- Vocabulary: 7,000 WordPiece tokens (uncased, diacritic-invariant, 1.28 fertility)
- Max Sequence Length: 512 tokens
- Tied Embeddings:
tie_word_embeddings: true - Masking: Deterministic Disjoint Whole-Word Masking (6-epoch rotational coverage)
Comparison across HmarBERT Models
| Feature / Model | HmarBERT (POC) | HmarBERT-mini (v1) | HmarBERT-mini-v2 |
|---|---|---|---|
| Base / Origin | Adapted from MizBERT | From scratch | From scratch |
| Layers / Hidden | 12 layers / 768 dim | 4 layers / 384 dim | 8 layers / 512 dim |
| Parameters | 110M | 16.9M | 29.3M |
| Vocab Size | 30,522 | 24,576 | 7,000 (Compact) |
| Weights Size (FP16) | ~220 MB | ~32 MB | ~58 MB |
| Diacritic Handling | Case-dependent | Uncased | Uncased, Diacritic-invariant |
Training & Curriculum
Pre-training was structured in 3 progressive curriculum stages:
- Stage 1 (Sentences): ~289k sentences (
hmar-heritage-org/sentences) up to 128 max length to ground morphology, postpositions, and basic clause structure. - Stage 2 (Paragraphs): ~44k long paragraphs (
hmar-heritage-org/paragraphs) up to 512 max length for multi-sentence discourse. - Stage 3 (Unified Mixed): Balanced blending across domains with whole-word scaffolding.
Quickstart
from transformers import pipeline
unmasker = pipeline("fill-mask", model="azinamotoe/HmarBERT-mini-v2")
# Example 1: Core Theology & Scripture
text1 = "Pathien chun khawvel a [MASK] em em a."
print(unmasker(text1)[0]['sequence'])
# -> "Pathien chun khawvel a hmangai em em a."
# Example 2: Cultural / Relational
text2 = "Nang chu ka [MASK] takzet i nih."
print(unmasker(text2)[0]['sequence'])
# -> "Nang chu ka thian takzet i nih."
Acknowledgements
Trained on open datasets curated by the Hmar Heritage Foundation (hmar-heritage-org).
- Downloads last month
- 42