Custom decoder-only Transformer trained for human/AI text continuation. It has 35,402,752 total parameters and up to 35,402,752 active parameters per token, six layers, eight attention heads, hidden size 512, a 256-token context and a 32,000-token byte-level BPE vocabulary.
Test perplexity: 41.4004. BLEU: 1.2868.
Load and run
Install requirements.txt. The custom architecture is supplied as ordinary
PyTorch source; review the source before importing it.
import sys
import torch
from huggingface_hub import snapshot_download
from tokenizers import Tokenizer
directory = snapshot_download("vanshnawander/assignment2-optimizer-lion")
sys.path.insert(0, directory)
from load_model import load_model
model = load_model(directory)
tokenizer = Tokenizer.from_file(directory + "/tokenizer.json")
ids = [1] + tokenizer.encode("Hello world").ids
with torch.inference_mode():
logits = model(torch.tensor([ids]))
print(logits.shape)
Special token IDs are PAD=0, BOS=1, EOS=2, SEP=3, VI=4 and JA=5.
Translation prompts use [BOS, language_id, source_tokens..., SEP].
Continuation prompts use [BOS, text_tokens...]. Keep inputs within 256 tokens.
decoding.py contains forward-based decoding helpers.
model_state.pt contains tensor weights only. Optimizer and random-generator
states remain in the original local resumable checkpoints. Architecture,
training settings and evaluation results are included as JSON files. No raw
training corpus or credentials are included.
- Downloads last month
- -