VMiniBert / README.md
vprojectx's picture
Upload folder using huggingface_hub
b74c422 verified
|
Raw History Blame Contribute Delete
3.76 kB
---
language:
- en
tags:
- masked-language-modeling
- bert
- vminibert
- rope
- rmsnorm
- geglu
- pytorch
- fill-mask
- from-scratch
license: mit
pipeline_tag: fill-mask
datasets:
- wikitext
- bookcorpusopen
- HuggingFaceFW/fineweb
- allenai/c4
---
# vminibert-108M
`vminibert-108M` is a 108.6M-parameter Transformer encoder trained **from scratch** as a Masked Language Model on a blended 4-domain corpus (WikiText-103, BookCorpusOpen, FineWeb, and C4) using dynamic Whole-Word Masking (WWM).
Rather than a stock BERT reimplementation, it uses a modern LLaMA-style encoder recipe:
- **RoPE** (Rotary Position Embeddings) instead of learned absolute positions
- **RMSNorm** instead of LayerNorm
- **GeGLU** feed-forward blocks instead of a plain GELU MLP
## Model Details
| | |
| :--- | :--- |
| **Parameters** | 108,615,984 |
| **Hidden size** | 768 |
| **Attention heads** | 12 |
| **Encoder layers** | 12 |
| **Intermediate size (GeGLU)** | 2048 |
| **Vocabulary size** | 30,000 (WordPiece) |
| **Max sequence length** | 32 |
| **Positional encoding** | RoPE |
| **Normalization** | RMSNorm |
| **Activation** | GeGLU |
| **Weight tying** | Output projection tied to input embeddings |
## Training Data
Trained on ~250k samples blended from four domains:
| Source | Domain |
| :--- | :--- |
| WikiText-103 | Encyclopedic |
| BookCorpusOpen | Long-form narrative |
| FineWeb | General web text |
| C4 | Cleaned web crawl |
Masking used dynamic Whole-Word Masking at a 15% rate.
## Evaluation
Evaluated on official unseen test splits with standard 15% WWM:
| Benchmark | Split | Masked Tokens | Loss | Perplexity | Top-1 Acc | Top-3 Acc | Top-5 Acc | MRR@10 |
| :--- | :--- | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| Salesforce/wikitext-2-raw-v1 | test | 8,552 | 4.0323 | 56.39 | 43.15% | 53.86% | 57.96% | — |
| Salesforce/wikitext-103-raw-v1 | test | 8,515 | 4.0378 | 56.70 | 42.82% | 53.55% | 58.06% | 49.29% |
*MRR@10 for wikitext-2 is not yet computed — fill in or remove the column before publishing.*
## Usage
```python
from transformers import AutoTokenizer, AutoModelForMaskedLM, pipeline
model_id = "vprojectx/VMiniBert"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForMaskedLM.from_pretrained(model_id, trust_remote_code=True)
fill_mask = pipeline("fill-mask", model=model, tokenizer=tokenizer)
results = fill_mask("The capital of France is [MASK].")
for res in results:
print(f"{res['token_str']:<12} {res['score']:.4f}")
```
## Limitations & Intended Use
- **Context window:** pretrained and evaluated at 32-token sequences only; behavior beyond that is untested. This is the single biggest caveat for anyone evaluating this model — flag it up front, not just in a bullet list.
- **Corpus size:** ~250k samples is modest for a from-scratch LM; expect it to need task-specific fine-tuning for downstream use rather than strong zero-shot performance.
- **Intended use:** a lightweight, fast MLM encoder for fine-tuning (classification, NER, masked-token extraction) and as a reference implementation of a modern (RoPE/RMSNorm/GeGLU) BERT-style architecture trained end-to-end from scratch.
- **Not intended for:** long-document tasks, production deployment without fine-tuning, or use cases requiring calibrated confidence at scale (no calibration study was done).
## Citation
If you use this model, please cite it as:
```bibtex
@misc{vminibert2026,
author = {vprojectx},
title = {vminibert-108M: A From-Scratch BERT-Style Encoder with RoPE, RMSNorm, and GeGLU},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/vprojectx/VMiniBert}
}
```