Nano-1B

A 1,076,700,672-parameter dense decoder architecture for training from scratch. This repository contains configuration, model code and a DeepSeek V4 tokenizer. There are no model weights or trained checkpoints. Initialize with AutoModelForCausalLM.from_config, rather than loading pretrained model weights.

The dimensions follow MiniCPM5-1B. Normalization, dense SwiGLU and interleaved RoPE reuse Transformers' DeepSeek V4 components. NanoDenseForCausalLM supplies the model entry point and standard global grouped-query attention. The tokenizer is pinned to DeepSeek-V4-Flash-Base, revision 8855555.

Architecture

Setting Value
Decoder layers 24
Residual hidden size 1,536
Dense SwiGLU intermediate size 4,608
Query / key-value heads 16 / 2
Head dimension 128
Attention projection width 2,048
Attention pattern Global causal GQA in every layer
Normalization DeepseekV4RMSNorm, epsilon 1e-6
Position encoding V4 interleaved RoPE over all 128 head dimensions
RoPE theta 5,000,000
Configured context target 131,072 tokens
Vocabulary 129,280
Input embedding / output head Separate weights
BOS / EOS / PAD 0 / 1 / 2
Total parameters 1,076,700,672
Excluding embedding and output head 679,552,512

Each layer computes:

h = x + GQA(RMSNorm(x))
y = h + SwiGLU(RMSNorm(h))

The attention width is independent of the residual width: Q projects 1536 -> 2048, K/V each project 1536 -> 256, and O projects 2048 -> 1536. All 24 layers have independent parameters. Linear and embedding weights use normal initialization with standard deviation 0.02; RMSNorm weights start at one. No depth-dependent initialization scaling is applied.

128K is a configuration target, not a measured capability of an untrained model. Long-context performance requires a suitable training curriculum and evaluation. At batch size one, a complete 128K BF16 KV cache occupies 3 GiB, excluding weights, temporary tensors and allocator overhead.

Initialize for training

pip install -r requirements.txt
import torch
from transformers import AutoConfig, AutoModelForCausalLM, AutoTokenizer

repo = "bowang0911/Nano-1B"
config = AutoConfig.from_pretrained(repo, trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_config(
    config,
    trust_remote_code=True,
    attn_implementation="sdpa",
    dtype=torch.bfloat16,
)
# Randomly initialized model; pass it to your training loop / Trainer.

To inspect the full architecture without allocating weight storage:

with torch.device("meta"):
    model = AutoModelForCausalLM.from_config(config, trust_remote_code=True)
print(sum(p.numel() for p in model.parameters()))  # 1076700672

From a local copy, python count_parameters.py --meta checks the analytical count against a full-size meta-device model.

Tokenized data and packing

tokenizer.json is byte-identical to the pinned DeepSeek V4 artifact. Tokenizer configuration sets model_max_length=131072 and selects the existing dedicated <|▁pad▁|> token (ID 2) for padding. Vocabulary IDs and raw unpadded encoding are unchanged; EOS remains ID 1. tokenizer_provenance.json records upstream and published checksums. Existing data encoded with this pinned tokenizer can be reused directly. The pretraining convention used here is raw text with add_special_tokens=False and one EOS token (ID 1) appended per document. No chat template is provided.

For ordinary padded batches, pass a two-dimensional attention_mask and set padding labels to -100. Default positions are derived from that mask, including when continuing a KV cache. The tokenizer, model config and generation config all use PAD ID 2. The model embedding has a dedicated zeroed padding row at ID 2; the EOS embedding remains trainable. With DataCollatorForLanguageModeling(mlm=False), only PAD labels are masked; real document-ending EOS labels retain ID 1. Previously PAD equaled EOS, which caused that standard collator to discard EOS supervision.

For padding-free document packing, use one flattened row (batch_size=1) and reset position_ids to zero at each document boundary. Omit attention_mask, or pass an all-ones mask. The model isolates documents and masks the first label of each subsequent document so the LM loss does not cross boundaries. Packing with padding, nonzero restart positions or multiple batch rows raises an error. KV caching is disabled during training and for packed sequences. Normal unpacked evaluation/generation supports dynamic KV caching.

Eager and SDPA backends are validated on CPU. The FlashAttention adapter receives packed positions and dispatches separate document segments; running actual FlashAttention kernels requires a compatible GPU installation.

Validation

Validated with Python 3.12.11, PyTorch 2.12.0 and Transformers 5.9.0:

  • Full-size meta parameter count, including the wider attention projections.
  • An independent complex-number RoPE / attention reference, including large positions.
  • Training loss, finite gradients through every parameter, an AdamW update, and gradient-checkpointing equivalence.
  • Eager/SDPA cached chunked decoding and cached/uncached greedy generation.
  • Left/right padding, packed-document isolation, boundary loss masking and FlashAttention varlen dispatch with external GPU kernels stubbed.
  • FP32/BF16 save/reload and invalid configuration override rejection.
  • Tokenizer parity with the pinned source.
  • Standard-collator EOS preservation, padded backward pass, and PAD/EOS consistency after tokenizer serialization.

The 13 local behavioral tests are kept outside this model repository. This is small-model CPU validation, not full-size training, GPU throughput or 128K quality validation. Transformers is pinned because the implementation imports its V4 components and attention/cache interfaces.

Sources and licenses

No pretrained MiniCPM or DeepSeek model weights are included or converted.

Downloads last month
728
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support