FlameF0X/arXivGPT

A small GPT-2 language model trained from scratch with Auto-PreTrain. It is a standard GPT2LMHeadModel, so no custom code or trust_remote_code is needed.

Usage

from transformers import pipeline

generator = pipeline("text-generation", model="FlameF0X/arXivGPT")
print(generator("Once upon a time", max_new_tokens=50)[0]["generated_text"])

Model

Parameters 8,562,048
Layers / heads / hidden 8 / 8 / 128
FFN size 768
Context length 128
Activation gelu_new
Dropout 0.1
Tied embeddings True
Tokenizer gpt2 (vocab 50257)

Training data

FlameF0X/arXiv-AI-ML, split train, first 100,000 rows. 4,682 train, 247 eval examples (~630,912 tokens), packed into full-length blocks.

Training setup

Epochs / steps 5 / 2930
Batch size (x accumulation) 8 (x1)
Learning rate 0.0005 (cosine, warmup 0.05)
Weight decay 0.01
Precision fp32
Seed 42
Device CPU

Results

Metric Value
train_loss 5.6627
eval_loss 5.1668
perplexity 175.36

Example output

Prompt: Once upon a time

Once upon a time-of-art dataset, allowing the target model's number of the proposed method with the model's training-to-step and a single-training framework for our method. AB-based approach is an multi-based approach that generates a comprehensive

Limitations

This is a tiny model trained on a small data sample. Expect limited coherence and factual accuracy, and it may reproduce biases present in the training data. It is intended for experimentation and learning, not production use.

Downloads last month
-
Safetensors
Model size
8.56M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train FlameF0X/arXivGPT