TRUTH-11M

This is a parody model. It is not affiliated with, endorsed by, or connected to Donald J. Trump, the Trump Organization, or Truth Social. Everything it produces is synthetic text from a small language model trained on public posts. No output is a real statement by any person, and none of it should be quoted, screenshotted, or circulated as though it were genuine.

A 37.7M-parameter word-level GPT trained from scratch on roughly 11M tokens of public Truth Social posts. Built as a novelty app, not a research artifact.

Live demo: techguy1423/trump-intelligence

Architecture

Parameters 37.7M (12.4M of them the tied embedding)
Layers 8
Heads 8
Embedding dim 512
Context 192 tokens
Vocabulary 24,214 word-level tokens
Tokenizer whitespace + punctuation split, lowercased
Training 5,000 iters, batch 16, AdamW, cosine schedule w/ warmup, dropout 0.15

Roughly 1.4 epochs over the corpus. Input embedding and output projection are weight-tied.

Usage

The architecture is custom, so transformers cannot load this. The checkpoint is a dict with model_state_dict, stoi, itos, vocab_size, and config. Model and tokenizer code, plus a FastAPI server, are in the Space repo.

from huggingface_hub import hf_hub_download
from model import load_checkpoint          # from the Space repo

path = hf_hub_download("techguy1423/truth-11m", "trump_token_gpt.pt")
model, stoi, itos, config = load_checkpoint(path, device="cpu")

Generation runs at roughly 50-65 tok/s on two CPU threads, so about 1s for a 60-token reply. There is no KV cache, so cost grows quadratically with context.

Intended use

Entertainment and satire. It continues a prompt in the style of the corpus.

Not suitable for

Anything factual. It has no knowledge base, no grounding, and no notion of truth โ€” it reproduces the surface statistics of one author's posting style. It will state things that are false, and it will produce political invective, because that is what the training data is made of.

It is also not a chat model. There is no instruction tuning and no prompt/response structure in the training data, so it cannot answer questions. Feeding it user text as a generation seed is the entire mechanism.

Known limitations

  • Lowercased corpus. Capitalization was destroyed at tokenization time, so proper nouns cannot be recovered. The serving code fakes sentence casing and restores initialisms heuristically.
  • No end-of-post token. Posts were concatenated into one flat stream, so the model never learned where a post ends; it generates until it hits the token limit. Output is trimmed to the last complete sentence as a workaround.
  • Corpus contains glued words. The source CSV had HTML stripped without substituting whitespace, fusing the last word of each block to the first word of the next at case boundaries. This burned 300+ vocabulary slots on artifacts like realdonaldtrumpcrooked, which occasionally surface in output.
  • Out-of-vocabulary words are dropped, not mapped to an <unk> token.
  • Attention is scaled by 1/sqrt(n_embd) instead of 1/sqrt(head_size) โ€” a carryover from the nanoGPT tutorial this was built from. The weights are adapted to it, so it must not be "fixed" at inference time.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support