|
Download docs/ARCHITECTURE.md from Amitkumar001/Law_Slm: direct link, hf CLI and curl.
- Browser
- Download file 3.29 kB
-
https://huggingface.co/Amitkumar001/Law_Slm/resolve/main/docs/ARCHITECTURE.md
- Command line
-
hf download hf://Amitkumar001/Law_Slm/docs/ARCHITECTURE.md
-
curl -L -o ARCHITECTURE.md https://huggingface.co/Amitkumar001/Law_Slm/resolve/main/docs/ARCHITECTURE.md
3.29 kB
Architectural Documentation - Small Language Model (SLM)
System Architecture
lawslm is an industrial-grade, decoder-only Small Language Model (SLM) ecosystem built completely from scratch using Python and PyTorch tensor primitives.
+-----------------------------------------------------------------------+
| SLM ENGINE |
+-----------------------------------------------------------------------+
| [REST API / FastAPI] [CLI Interface / main.py] |
+-----------------------------------++----------------------------------+
||
+-----------------------------------vv----------------------------------+
| Text Generator |
| (Greedy, Temp, Top-K, Top-P, Penalties, Streaming Callbacks) |
+-----------------------------------++----------------------------------+
||
+-----------------------------------vv----------------------------------+
| SLMForCausalLM |
| Token Embeddings -> RoPE -> N x TransformerBlock -> RMSNorm -> Head |
+-----------------------------------++----------------------------------+
||
+-----------------------------------vv----------------------------------+
| Transformer Block |
| Pre-RMSNorm -> Multi-Head Causal Attention -> Pre-RMSNorm -> SwiGLU |
+-----------------------------------------------------------------------+
Modular Breakdown
Tokenizer (
slm/tokenizer/):- Pure Python Byte-Pair Encoding (BPE) subword algorithm (
bpe.py). - Character tokenizer fallback (
char_tokenizer.py). - Vocabulary manager, frequency dictionary, and JSON serialization (
vocab.py).
- Pure Python Byte-Pair Encoding (BPE) subword algorithm (
Dataset & Cleaning (
slm/dataset/):- Text cleaning, NFC Unicode normalization, HTML stripping, deduplication (
cleaner.py). - Ingestion parsers for TXT, CSV, JSON, JSONL, MD, HTML, XML (
readers.py). - Sliding causal context window PyTorch Dataset (
dataset.py).
- Text cleaning, NFC Unicode normalization, HTML stripping, deduplication (
Embeddings & Normalization (
slm/embeddings/,slm/normalization/):- Rotary Position Embeddings (RoPE) applied to Queries and Keys (
positional.py). - Learnable and Sinusoidal positional embeddings option (
positional.py). - Root Mean Square Layer Normalization (
rmsnorm.py). - Layer Normalization (
layernorm.py).
- Rotary Position Embeddings (RoPE) applied to Queries and Keys (
Attention & FeedForward (
slm/attention/,slm/feedforward/):- Scaled Dot-Product Attention with triangular causal mask (
causal_attention.py). - Multi-Head Causal Attention with RoPE (
causal_attention.py). - SwiGLU (Swish Gated Linear Unit) Feed-Forward Network (
mlp.py).
- Scaled Dot-Product Attention with triangular causal mask (
Optimizers & Schedulers (
slm/optimizer/,slm/scheduler/):- Custom
AdamWwith decoupled weight decay (adamw.py). - Custom
Lion(EvoLved Sign Momentum) optimizer (lion.py). - Cosine Annealing with Warmup learning rate scheduler (
schedulers.py).
- Custom
Training Engine (
slm/training/):- Teacher-forcing training loop with AMP FP16/BF16, gradient accumulation, gradient clipping, evaluation, and checkpoint manager integration.