Custom decoder-only Transformer trained for human/AI text continuation. It has 35,402,752 total parameters and up to 35,402,752 active parameters per token, six layers, eight attention heads, hidden size 512, a 256-token context and a 32,000-token byte-level BPE vocabulary.

Test perplexity: 41.4004. BLEU: 1.2868.

Load and run

Install requirements.txt. The custom architecture is supplied as ordinary PyTorch source; review the source before importing it.

import sys
import torch
from huggingface_hub import snapshot_download
from tokenizers import Tokenizer

directory = snapshot_download("vanshnawander/assignment2-optimizer-lion")
sys.path.insert(0, directory)
from load_model import load_model

model = load_model(directory)
tokenizer = Tokenizer.from_file(directory + "/tokenizer.json")
ids = [1] + tokenizer.encode("Hello world").ids
with torch.inference_mode():
    logits = model(torch.tensor([ids]))
print(logits.shape)

Special token IDs are PAD=0, BOS=1, EOS=2, SEP=3, VI=4 and JA=5. Translation prompts use [BOS, language_id, source_tokens..., SEP]. Continuation prompts use [BOS, text_tokens...]. Keep inputs within 256 tokens. decoding.py contains forward-based decoding helpers.

model_state.pt contains tensor weights only. Optimizer and random-generator states remain in the original local resumable checkpoints. Architecture, training settings and evaluation results are included as JSON files. No raw training corpus or credentials are included.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including vanshnawander/assignment2-optimizer-lion