TRUTH-11M
This is a parody model. It is not affiliated with, endorsed by, or connected to Donald J. Trump, the Trump Organization, or Truth Social. Everything it produces is synthetic text from a small language model trained on public posts. No output is a real statement by any person, and none of it should be quoted, screenshotted, or circulated as though it were genuine.
A 37.7M-parameter word-level GPT trained from scratch on roughly 11M tokens of public Truth Social posts. Built as a novelty app, not a research artifact.
Live demo: techguy1423/trump-intelligence
Architecture
| Parameters | 37.7M (12.4M of them the tied embedding) |
| Layers | 8 |
| Heads | 8 |
| Embedding dim | 512 |
| Context | 192 tokens |
| Vocabulary | 24,214 word-level tokens |
| Tokenizer | whitespace + punctuation split, lowercased |
| Training | 5,000 iters, batch 16, AdamW, cosine schedule w/ warmup, dropout 0.15 |
Roughly 1.4 epochs over the corpus. Input embedding and output projection are weight-tied.
Usage
The architecture is custom, so transformers cannot load this. The checkpoint
is a dict with model_state_dict, stoi, itos, vocab_size, and config.
Model and tokenizer code, plus a FastAPI server, are in the
Space repo.
from huggingface_hub import hf_hub_download
from model import load_checkpoint # from the Space repo
path = hf_hub_download("techguy1423/truth-11m", "trump_token_gpt.pt")
model, stoi, itos, config = load_checkpoint(path, device="cpu")
Generation runs at roughly 50-65 tok/s on two CPU threads, so about 1s for a 60-token reply. There is no KV cache, so cost grows quadratically with context.
Intended use
Entertainment and satire. It continues a prompt in the style of the corpus.
Not suitable for
Anything factual. It has no knowledge base, no grounding, and no notion of truth โ it reproduces the surface statistics of one author's posting style. It will state things that are false, and it will produce political invective, because that is what the training data is made of.
It is also not a chat model. There is no instruction tuning and no prompt/response structure in the training data, so it cannot answer questions. Feeding it user text as a generation seed is the entire mechanism.
Known limitations
- Lowercased corpus. Capitalization was destroyed at tokenization time, so proper nouns cannot be recovered. The serving code fakes sentence casing and restores initialisms heuristically.
- No end-of-post token. Posts were concatenated into one flat stream, so the model never learned where a post ends; it generates until it hits the token limit. Output is trimmed to the last complete sentence as a workaround.
- Corpus contains glued words. The source CSV had HTML stripped without
substituting whitespace, fusing the last word of each block to the first word
of the next at case boundaries. This burned 300+ vocabulary slots on artifacts
like
realdonaldtrumpcrooked, which occasionally surface in output. - Out-of-vocabulary words are dropped, not mapped to an
<unk>token. - Attention is scaled by
1/sqrt(n_embd)instead of1/sqrt(head_size)โ a carryover from the nanoGPT tutorial this was built from. The weights are adapted to it, so it must not be "fixed" at inference time.