Article 1 Emergent Semantics Beyond Token Embeddings: A GPT-like Transformer Learns with Frozen 16‑D Binary Token-ID Embeddings (n_embed=16)
Language Models Without a Trainable Input Embedding Table This collection is provided for reproducibility of the paper's main claim Bochkov/llm-fix-min-baseline-learned-input-table-model-classic Text Generation • 0.5B • Updated May 12 • 16 Bochkov/llm-fix-min-fixed-minimal-binary-code Text Generation • 0.5B • Updated May 12 • 21 Bochkov/llm-fix-min-affine-recoded-minimal-code-table-free Text Generation • 0.5B • Updated May 12 • 24
Bochkov/llm-fix-min-baseline-learned-input-table-model-classic Text Generation • 0.5B • Updated May 12 • 16
Bochkov/llm-fix-min-affine-recoded-minimal-code-table-free Text Generation • 0.5B • Updated May 12 • 24
Emergent Semantics Beyond Token Embeddings Paper: 2507.04886 (TMLR, Oct 2025). 'Emergent Semantics Beyond Token Embeddings: Transformer LMs with Frozen Visual Unicode Representations' Bochkov/emergent-semantics-model-uni-glyph-335m Text Generation • 0.3B • Updated Jan 7 • 15 Bochkov/emergent-semantics-model-unfrozen-335m Text Generation • 0.3B • Updated Jan 7 • 12 Bochkov/emergent-semantics-model-16-bit-269m Text Generation • 0.3B • Updated Jan 7 • 13 • 1 Bochkov/emergent-semantics-model-64-bit-272m Text Generation • 0.3B • Updated Jan 7 • 77
Language Models Without a Trainable Input Embedding Table This collection is provided for reproducibility of the paper's main claim Bochkov/llm-fix-min-baseline-learned-input-table-model-classic Text Generation • 0.5B • Updated May 12 • 16 Bochkov/llm-fix-min-fixed-minimal-binary-code Text Generation • 0.5B • Updated May 12 • 21 Bochkov/llm-fix-min-affine-recoded-minimal-code-table-free Text Generation • 0.5B • Updated May 12 • 24
Bochkov/llm-fix-min-baseline-learned-input-table-model-classic Text Generation • 0.5B • Updated May 12 • 16
Bochkov/llm-fix-min-affine-recoded-minimal-code-table-free Text Generation • 0.5B • Updated May 12 • 24
Emergent Semantics Beyond Token Embeddings Paper: 2507.04886 (TMLR, Oct 2025). 'Emergent Semantics Beyond Token Embeddings: Transformer LMs with Frozen Visual Unicode Representations' Bochkov/emergent-semantics-model-uni-glyph-335m Text Generation • 0.3B • Updated Jan 7 • 15 Bochkov/emergent-semantics-model-unfrozen-335m Text Generation • 0.3B • Updated Jan 7 • 12 Bochkov/emergent-semantics-model-16-bit-269m Text Generation • 0.3B • Updated Jan 7 • 13 • 1 Bochkov/emergent-semantics-model-64-bit-272m Text Generation • 0.3B • Updated Jan 7 • 77
Bochkov/llm-fix-min-affine-recoded-minimal-code-table-free Text Generation • 0.5B • Updated May 12 • 24
Bochkov/llm-fix-min-baseline-learned-input-table-model-classic Text Generation • 0.5B • Updated May 12 • 16
Bochkov/growing-transformers-model-frozen-16-bit-baseline-monolyth-181m Text Generation • 0.2B • Updated Jan 9 • 20
Bochkov/growing-transformers-model-unfrozen-baseline-monolyth-247m Text Generation • 0.2B • Updated Jan 9 • 14
Bochkov/growing-transformers-model-frozen-unicode-baseline-monolyth-247m Text Generation • 0.2B • Updated Jan 9 • 13