CrowdMind NanoChat D34 Base
A small decoder-only language model trained by CrowdMind using the NanoChat training framework.
Model Summary
| Property | Value |
|---|---|
| Parameters | ... |
| Layers | 34 |
| Embedding dimension | ... |
| Attention heads | ... |
| KV heads | ... |
| Context length | 2,048 |
| Vocabulary | 16,384 |
| Attention backend | PyTorch SDPA |
| Precision | BF16 |
| Activation checkpointing | Enabled |
| Dataset | TinyStories |
| Training hardware | NVIDIA RTX 4060 Ti 16GB |
Intended Use
This model is primarily intended for:
- experimentation with small language models
- studying language-model scaling
- educational purposes
- local inference
- benchmarking small-model architectures
- experimenting with the NanoChat training framework
It is not intended to be considered a general-purpose production language model.
Training
The model was trained on the TinyStories dataset using the NanoChat training implementation.
Training was performed on a consumer NVIDIA RTX 4060 Ti with 16 GB of VRAM using BF16 computation, activation checkpointing, and PyTorch SDPA attention.
Training Configuration
| Setting | Value |
|---|---|
| Architecture | Decoder-only Transformer |
| Layers | 34 |
| Context length | 2,048 |
| Vocabulary size | 16,384 |
| Precision | BF16 |
| Attention | PyTorch SDPA |
| Activation checkpointing | Enabled |
| Dataset | TinyStories |
| GPU | NVIDIA RTX 4060 Ti 16GB |
| Optimizer | Muon / NanoChat training configuration |
| Training duration | ... |
| Training tokens | ... |
| Batch size | ... |
| Gradient accumulation | ... |
| Learning rate | ... |
Evaluation
The model was evaluated on the TinyStories validation set.
Reported metrics:
- Validation BPB:
... - Training tokens:
... - Training duration:
...
BPB (bits per byte) measures the number of bits required to represent each byte of text according to the model. It should be interpreted in the context of the TinyStories dataset and the tokenizer used during training.
Architecture
NanoChat D34 Base is a compact decoder-only Transformer architecture designed for efficient experimentation on consumer hardware.
The model uses:
- 34 Transformer layers
- 16,384-token vocabulary
- 2,048-token context length
- BF16 computation
- PyTorch scaled dot-product attention (SDPA)
- activation checkpointing during training
The D34 configuration is part of CrowdMind's NanoChat model series, which explores how model architecture, parameter count, and training compute affect small language-model performance.
Limitations
This is a small research and educational language model.
Because of its relatively small parameter count and training budget, the model may:
- produce incorrect information
- repeat text
- lose coherence during longer generations
- struggle with tasks outside its training distribution
- produce grammatically or semantically inconsistent text
- have limited world knowledge
- produce undesirable or unexpected outputs
The model was trained primarily on TinyStories and therefore should not be expected to perform like a general-purpose pretrained language model.
The model should not be treated as a reliable source of factual information.
Usage
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "CrowdMind/nanochat-d34-base"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
prompt = "Once upon a time"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=200,
temperature=0.8,
top_p=0.95,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Reproducibility
The model was trained using the NanoChat training framework.
Dataset:
Model repository:
The model is part of the CrowdMind NanoChat model series and is intended to make small-language-model experimentation reproducible and accessible on consumer hardware.
Citation
If you use this model in research, experiments, or derivative work, please cite the CrowdMind NanoChat project:
@misc{crowdmind_nanochat_d34,
title={CrowdMind NanoChat D34 Base},
author={CrowdMind},
year={2026},
publisher={Hugging Face},
url={https://huggingface.co/CrowdMind/nanochat-d34-base}
}
Acknowledgements
Built by CrowdMind using the NanoChat training approach.
NanoChat provides an accessible framework for experimenting with language-model architecture and training at small scales using consumer hardware.
Model: CrowdMind/nanochat-d34-base
Organization: CrowdMind
Framework: NanoChat
License: MIT
- Downloads last month
- 50
