--- language: - awa - bra - hi pipeline_tag: text-generation tags: - pytorch - causal-lm - transformer - character-level - poetry - tulsidas - devanagari - awadhi - low-resource license: mit ---
# ЁЯкФ MANAS ## Model for Awadhi Natural Autoregressive Sequences *A character-level causal Transformer trained on the literary corpus of Goswami Tulsidas* [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT) ![PyTorch](https://img.shields.io/badge/PyTorch-2.x-EE4C2C?logo=pytorch) ![Language](https://img.shields.io/badge/Language-Awadhi%20%7C%20Hindi-orange) ![Parameters](https://img.shields.io/badge/Parameters-10.9M-blue) ![Status](https://img.shields.io/badge/Status-Experimental%20Checkpoint%201-lightgrey)
--- ## Overview **MANAS** is a small, experimental **character-level causal Transformer language model** trained on a Devanagari literary corpus focused on the works of Goswami Tulsidas тАФ the 16th-century poet-saint and author of the *Ramcharitmanas*. Unlike modern large language models that operate on subword tokens (BPE, SentencePiece), MANAS processes text **one Unicode character at a time**. The model was built entirely from scratch in PyTorch, without any pretrained weights or transfer learning, as an educational experiment into whether a small Transformer can learn the statistical patterns of classical Awadhi poetry from raw characters. > тЪая╕П **This is Experimental Checkpoint 1.** The dataset still contains OCR artifacts and some non-literary material. Do not use this as a finished, production Awadhi language model. A cleaned, page-verified corpus rebuild is planned for v2. --- ## Quick Start ```bash # 1. Clone the repo git clone https://huggingface.co/JayF14/MANAS cd MANAS # 2. Install dependencies pip install -r requirements.txt # 3. Run the interactive chat python run.py ``` That's it. Type a prompt like `shri ram` or `рд╢реНрд░реА рд░рд╛рдо` and MANAS will generate text. --- ## What's in This Repository | File | Description | |---|---| | `tulsidas_model.pth` | Trained PyTorch model weights (~52MB) | | `vocab.json` | The exact 101-character vocabulary mapping used during training | | `run.py` | Self-contained interactive inference script | | `requirements.txt` | Python dependencies (`torch`, `indic-transliteration`) | | `README.md` | This file | > The raw training dataset is not included in this release. --- ## Model Architecture MANAS is a **GPT-style decoder-only causal Transformer** тАФ the same fundamental architecture family as GPT-2. The key difference is that it works directly on Devanagari Unicode code points rather than BPE tokens. ``` Input Characters тЖУ Character Embedding (384-dim) + Positional Embedding (256 positions) тЖУ 6 ├Ч Transformer Blocks тФЬтФАтФА LayerNorm тФЬтФАтФА Multi-Head Causal Self-Attention (6 heads ├Ч 64 dim) тФЬтФАтФА Residual Connection тФЬтФАтФА LayerNorm тФЬтФАтФА Feed-Forward Network (384 тЖТ 1536 тЖТ 384) тФФтФАтФА Residual Connection тЖУ Final LayerNorm тЖУ LM Head (Linear: 384 тЖТ 101) тЖУ Logits over 101 Devanagari characters ``` ### Hyperparameters | Parameter | Value | |---|---| | Embedding Dimension (`n_embd`) | 384 | | Transformer Layers (`n_layer`) | 6 | | Attention Heads (`n_head`) | 6 | | Head Dimension | 64 | | Context Window (`block_size`) | 256 characters | | Vocabulary Size | 101 characters | | Dropout | 0.2 | | Total Parameters | **10,816,613 (~10.9M)** | --- ## Training ### Dataset The model was trained on a corpus of approximately **2.83 million Devanagari characters** derived from OCR-processed editions of Tulsidas's works: - *Ramcharitmanas* (рд░рд╛рдордЪрд░рд┐рддрдорд╛рдирд╕) - *Vinay Patrika* (рд╡рд┐рдирдп рдкрддреНрд░рд┐рдХрд╛) - *Kavitavali* (рдХрд╡рд┐рддрд╛рд╡рд▓реА) - *Parvatimangal* (рдкрд╛рд░реНрд╡рддреАрдордВрдЧрд▓) - Other works from the *Tulsi Granthavali* compilation The corpus contains **zero Latin alphabet characters**. The vocabulary consists of 101 unique Devanagari characters, punctuation marks, and verse/number markers common in classical Hindi poetry. > тЪая╕П **Known Limitation:** The current corpus is OCR-derived and may contain publisher front matter, Hindi commentary (рдЯреАрдХрд╛), table-of-contents material, and extraction errors alongside the literary verses. The corpus has not been manually page-verified. ### Training Configuration | Config | Value | |---|---| | Optimizer | AdamW | | Learning Rate | 3e-4 | | Batch Size | 64 | | Training Iterations | 15,000 | | Train / Val Split | 90% / 10% | | Train Tokens | 2,550,420 | | Validation Tokens | 283,381 | | Hardware | Single consumer GPU (NVIDIA CUDA) | ### Loss Curve | Step | Train Loss | Val Loss | |---|---|---| | 0 | 4.85 | 4.85 | | 500 | 2.29 | 2.61 | | 2,000 | 1.62 | 2.15 | | 5,000 | 1.35 | 2.05 | | 10,000 | 1.16 | 2.07 | | 15,000 | **1.03** | **2.11** | The gap between training loss (1.03) and validation loss (2.11) indicates the model has overfit to the relatively small corpus тАФ expected behavior for a 10.9M parameter model on ~2.8M characters. --- ## How to Use The easiest way is `python run.py` тАФ it handles everything automatically. For custom integration, the vocabulary is in `vocab.json` so you do not need the original corpus: ```python import torch import torch.nn as nn from torch.nn import functional as F # --- Hyperparameters (must match training) --- block_size = 256 n_embd = 384 n_head = 6 n_layer = 6 dropout = 0.0 # Set to 0 for inference vocab_size = 101 # Must match your character mapping device = 'cuda' if torch.cuda.is_available() else 'cpu' # --- Model Architecture --- class Head(nn.Module): def __init__(self, head_size): super().__init__() self.key = nn.Linear(n_embd, head_size, bias=False) self.query = nn.Linear(n_embd, head_size, bias=False) self.value = nn.Linear(n_embd, head_size, bias=False) self.register_buffer('tril', torch.tril(torch.ones(block_size, block_size))) self.dropout = nn.Dropout(dropout) def forward(self, x): B, T, C = x.shape k, q = self.key(x), self.query(x) wei = q @ k.transpose(-2, -1) * (k.shape[-1] ** -0.5) wei = wei.masked_fill(self.tril[:T, :T] == 0, float('-inf')) wei = F.softmax(wei, dim=-1) return self.dropout(wei) @ self.value(x) class MultiHeadAttention(nn.Module): def __init__(self, num_heads, head_size): super().__init__() self.heads = nn.ModuleList([Head(head_size) for _ in range(num_heads)]) self.proj = nn.Linear(n_embd, n_embd) self.dropout = nn.Dropout(dropout) def forward(self, x): return self.dropout(self.proj(torch.cat([h(x) for h in self.heads], dim=-1))) class FeedForward(nn.Module): def __init__(self, n_embd): super().__init__() self.net = nn.Sequential( nn.Linear(n_embd, 4 * n_embd), nn.ReLU(), nn.Linear(4 * n_embd, n_embd), nn.Dropout(dropout), ) def forward(self, x): return self.net(x) class Block(nn.Module): def __init__(self, n_embd, n_head): super().__init__() head_size = n_embd // n_head self.sa, self.ffwd = MultiHeadAttention(n_head, head_size), FeedForward(n_embd) self.ln1, self.ln2 = nn.LayerNorm(n_embd), nn.LayerNorm(n_embd) def forward(self, x): return x + self.ffwd(self.ln2(x + self.sa(self.ln1(x)))) class LanguageModel(nn.Module): def __init__(self): super().__init__() self.token_embedding_table = nn.Embedding(vocab_size, n_embd) self.position_embedding_table = nn.Embedding(block_size, n_embd) self.blocks = nn.Sequential(*[Block(n_embd, n_head=n_head) for _ in range(n_layer)]) self.ln_f = nn.LayerNorm(n_embd) self.lm_head = nn.Linear(n_embd, vocab_size) def forward(self, idx, targets=None): B, T = idx.shape x = self.token_embedding_table(idx) + self.position_embedding_table(torch.arange(T, device=device)) logits = self.lm_head(self.ln_f(self.blocks(x))) return logits, None def generate(self, idx, max_new_tokens): for _ in range(max_new_tokens): logits, _ = self(idx[:, -block_size:]) idx_next = torch.multinomial(F.softmax(logits[:, -1, :], dim=-1), num_samples=1) idx = torch.cat((idx, idx_next), dim=1) return idx # --- Load weights --- model = LanguageModel() model.load_state_dict(torch.load('tulsidas_model.pth', map_location=device)) model.to(device) model.eval() # --- Build vocab from your corpus (required for encoding/decoding) --- # with open('your_training_corpus.txt', 'r', encoding='utf-8') as f: # text = f.read() # chars = sorted(list(set(text))) # stoi = {ch: i for i, ch in enumerate(chars)} # itos = {i: ch for i, ch in enumerate(chars)} # encode = lambda s: [stoi[c] for c in s] # decode = lambda l: ''.join([itos[i] for i in l]) # --- Inference --- # prompt = "рд╢реНрд░реА рд░рд╛рдо" # context = torch.tensor([encode(prompt)], dtype=torch.long, device=device) # print(decode(model.generate(context, max_new_tokens=300)[0].tolist())) ``` ### Sample Output Given the prompt `рд╢реНрд░реА`, the model produced (at step 14999): ``` рд╢реНрд░реАрд░рд╛рдордЪрдиреНрджреНрд░рдЬреАрдХреА рдкрд░рд┐рд╢реНрд░рд╛рдордЪрдиреНрджреНрд░рдЬреАрдиреЗ рд╣реА рд╕рдм рдорд╛рддрд╛ рдЖрджрд┐ рдорд┐рдЯ рдвреАрдВ ред рд╡рд┐рднреАрд╖рдгрдЬреАрдиреЗ рдЙрд╕рдХреЛ рд╣реГрджрдпрдореЗрдВ рдЙрдард╛ рд▓рд┐рдпрд╛ рдирд┐рджрд╛рди рджреАрди рдмрдЪрди рдЧрд╣рд┐ рд╕реЛрднрд╛ рдмрдврд╝рд╛рд╡рд╛ред рдмрд╛рд▓рд┐ рдФрд░ рдмрд┐рдкреБрд▓ рдмреЛрд▓рд╛рд╡рдбрд╝реЗ рдЧрд╛рд╡рд╛ рее рдореЛрд░реЗ рдЖрдзреАрд╕ рдореИрдВ рднреА рдЬрд╛рди ред рддрд╛ рдХреБрдиреНрда рд╕рджреНрдп рд╕реБрд░реБрдЪрд┐ рд░рд╕рдЦрд╛рд╡рд╛ рее ``` --- ## Limitations | Limitation | Detail | |---|---| | **Small model** | 10.9M parameters is far below modern LLM scale. Outputs are statistically plausible Devanagari, not semantically coherent poetry. | | **Corpus noise** | OCR errors, Hindi commentary (рдЯреАрдХрд╛), and publisher front matter remain in the training data. | | **No factual knowledge** | Cannot answer questions. Only predicts next characters. | | **Short context** | 256-character window limits long-range coherence. | | **Overfitting** | Train loss (1.03) vs. val loss (2.11) gap indicates memorisation of the small corpus. | | **Devanagari-only** | Vocabulary has zero Latin characters. English input must be transliterated before encoding. | --- ## Roadmap - [ ] **v2 Dataset:** Manually page-verified OCR rebuild separating verse from prose commentary - [ ] **v2 Model:** Retrain on the cleaned corpus with a larger context window - [ ] **Evaluation:** Implement character-level perplexity benchmarks on a held-out verse set - [ ] **Tokenizer:** Experiment with syllable-level tokenization for better Hindi morpheme coverage --- ## Citation If you reference this model in academic work: ```bibtex @misc{manas2026, author = {JayF14}, title = {MANAS: Model for Awadhi Natural Autoregressive Sequences}, year = {2026}, publisher = {Hugging Face}, howpublished = {\url{https://huggingface.co/JayF14/MANAS}}, note = {Experimental Checkpoint 1. Character-level causal Transformer trained on Tulsidas literary corpus.} } ``` --- ## License MIT License. See [LICENSE](LICENSE) for details. ---
рдЬреЗрд╣рд┐ рдкрд░ рдХреГрдкрд╛ рд░рд╛рдо рдХреИ рд╣реЛрдИред рддрд╛ рдкрд░ рдХреГрдкрд╛ рдХрд░рд╣рд┐рдВ рд╕рдм рдХреЛрдИрее
тАФ Ramcharitmanas, Goswami Tulsidas