nano34

CrowdMind NanoChat D34 Base

A small decoder-only language model trained by CrowdMind using the NanoChat training framework.

Model Summary

Property Value
Parameters ...
Layers 34
Embedding dimension ...
Attention heads ...
KV heads ...
Context length 2,048
Vocabulary 16,384
Attention backend PyTorch SDPA
Precision BF16
Activation checkpointing Enabled
Dataset TinyStories
Training hardware NVIDIA RTX 4060 Ti 16GB

Intended Use

This model is primarily intended for:

  • experimentation with small language models
  • studying language-model scaling
  • educational purposes
  • local inference
  • benchmarking small-model architectures
  • experimenting with the NanoChat training framework

It is not intended to be considered a general-purpose production language model.

Training

The model was trained on the TinyStories dataset using the NanoChat training implementation.

Training was performed on a consumer NVIDIA RTX 4060 Ti with 16 GB of VRAM using BF16 computation, activation checkpointing, and PyTorch SDPA attention.

Training Configuration

Setting Value
Architecture Decoder-only Transformer
Layers 34
Context length 2,048
Vocabulary size 16,384
Precision BF16
Attention PyTorch SDPA
Activation checkpointing Enabled
Dataset TinyStories
GPU NVIDIA RTX 4060 Ti 16GB
Optimizer Muon / NanoChat training configuration
Training duration ...
Training tokens ...
Batch size ...
Gradient accumulation ...
Learning rate ...

Evaluation

The model was evaluated on the TinyStories validation set.

Reported metrics:

  • Validation BPB: ...
  • Training tokens: ...
  • Training duration: ...

BPB (bits per byte) measures the number of bits required to represent each byte of text according to the model. It should be interpreted in the context of the TinyStories dataset and the tokenizer used during training.

Architecture

NanoChat D34 Base is a compact decoder-only Transformer architecture designed for efficient experimentation on consumer hardware.

The model uses:

  • 34 Transformer layers
  • 16,384-token vocabulary
  • 2,048-token context length
  • BF16 computation
  • PyTorch scaled dot-product attention (SDPA)
  • activation checkpointing during training

The D34 configuration is part of CrowdMind's NanoChat model series, which explores how model architecture, parameter count, and training compute affect small language-model performance.

Limitations

This is a small research and educational language model.

Because of its relatively small parameter count and training budget, the model may:

  • produce incorrect information
  • repeat text
  • lose coherence during longer generations
  • struggle with tasks outside its training distribution
  • produce grammatically or semantically inconsistent text
  • have limited world knowledge
  • produce undesirable or unexpected outputs

The model was trained primarily on TinyStories and therefore should not be expected to perform like a general-purpose pretrained language model.

The model should not be treated as a reliable source of factual information.

Usage

from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "CrowdMind/nanochat-d34-base"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
)

prompt = "Once upon a time"

inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

outputs = model.generate(
    **inputs,
    max_new_tokens=200,
    temperature=0.8,
    top_p=0.95,
)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Reproducibility

The model was trained using the NanoChat training framework.

Dataset:

TinyStories

Model repository:

CrowdMind/nanochat-d34-base

The model is part of the CrowdMind NanoChat model series and is intended to make small-language-model experimentation reproducible and accessible on consumer hardware.

Citation

If you use this model in research, experiments, or derivative work, please cite the CrowdMind NanoChat project:

@misc{crowdmind_nanochat_d34,
  title={CrowdMind NanoChat D34 Base},
  author={CrowdMind},
  year={2026},
  publisher={Hugging Face},
  url={https://huggingface.co/CrowdMind/nanochat-d34-base}
}

Acknowledgements

Built by CrowdMind using the NanoChat training approach.

NanoChat provides an accessible framework for experimenting with language-model architecture and training at small scales using consumer hardware.


Model: CrowdMind/nanochat-d34-base Organization: CrowdMind Framework: NanoChat License: MIT

Downloads last month
50
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train CrowdMind/nanochat-d34-base