HmarBERT-mini-v2

📦 Part of the HmarBERT Family Collection — open-source language models for the Hmar language, ranging from 16.9M edge-ready models to full base architectures.

HmarBERT-mini-v2 is an upgraded, 8-layer BERT model trained from scratch for the Hmar language (hmr, ISO 639-3), part of the Zo languages family.

It builds upon HmarBERT-mini (v1) with a deeper architecture, tied word embeddings, a refined 7,000-token compact vocabulary with lower fertility (1.28), and a 3-stage curriculum pre-training schedule using Deterministic Disjoint Whole-Word Masking.

At roughly 29.3M parameters (~58 MB in FP16 / ~117 MB in FP32), it provides significantly stronger syntactic and contextual representations while still running comfortably on modest hardware and laptops.


Model Details

  • Architecture: 8-layer BERT encoder (512 hidden size, 8 attention heads, 2048 intermediate FFN)
  • Parameters: 29.34M (~58 MB FP16 / ~117 MB FP32)
  • Vocabulary: 7,000 WordPiece tokens (uncased, diacritic-invariant, 1.28 fertility)
  • Max Sequence Length: 512 tokens
  • Tied Embeddings: tie_word_embeddings: true
  • Masking: Deterministic Disjoint Whole-Word Masking (6-epoch rotational coverage)

Comparison across HmarBERT Models

Feature / Model HmarBERT (POC) HmarBERT-mini (v1) HmarBERT-mini-v2
Base / Origin Adapted from MizBERT From scratch From scratch
Layers / Hidden 12 layers / 768 dim 4 layers / 384 dim 8 layers / 512 dim
Parameters 110M 16.9M 29.3M
Vocab Size 30,522 24,576 7,000 (Compact)
Weights Size (FP16) ~220 MB ~32 MB ~58 MB
Diacritic Handling Case-dependent Uncased Uncased, Diacritic-invariant

Training & Curriculum

Pre-training was structured in 3 progressive curriculum stages:

  1. Stage 1 (Sentences): ~289k sentences (hmar-heritage-org/sentences) up to 128 max length to ground morphology, postpositions, and basic clause structure.
  2. Stage 2 (Paragraphs): ~44k long paragraphs (hmar-heritage-org/paragraphs) up to 512 max length for multi-sentence discourse.
  3. Stage 3 (Unified Mixed): Balanced blending across domains with whole-word scaffolding.

Quickstart

from transformers import pipeline

unmasker = pipeline("fill-mask", model="azinamotoe/HmarBERT-mini-v2")

# Example 1: Core Theology & Scripture
text1 = "Pathien chun khawvel a [MASK] em em a."
print(unmasker(text1)[0]['sequence'])
# -> "Pathien chun khawvel a hmangai em em a."

# Example 2: Cultural / Relational
text2 = "Nang chu ka [MASK] takzet i nih."
print(unmasker(text2)[0]['sequence'])
# -> "Nang chu ka thian takzet i nih."

Acknowledgements

Trained on open datasets curated by the Hmar Heritage Foundation (hmar-heritage-org).

Downloads last month
42
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train azinamotoe/HmarBERT-mini-v2

Space using azinamotoe/HmarBERT-mini-v2 1

Collection including azinamotoe/HmarBERT-mini-v2