Instructions to use bowang0911/Nano-1B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bowang0911/Nano-1B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="bowang0911/Nano-1B", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("bowang0911/Nano-1B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use bowang0911/Nano-1B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bowang0911/Nano-1B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bowang0911/Nano-1B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/bowang0911/Nano-1B
- SGLang
How to use bowang0911/Nano-1B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "bowang0911/Nano-1B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bowang0911/Nano-1B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "bowang0911/Nano-1B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bowang0911/Nano-1B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use bowang0911/Nano-1B with Docker Model Runner:
docker model run hf.co/bowang0911/Nano-1B
Nano-1B
A 1,076,700,672-parameter dense decoder architecture for training from scratch.
This repository contains configuration, model code and a DeepSeek V4 tokenizer.
There are no model weights or trained checkpoints. Initialize with
AutoModelForCausalLM.from_config, rather than loading pretrained model weights.
The dimensions follow MiniCPM5-1B.
Normalization, dense SwiGLU and interleaved RoPE reuse Transformers' DeepSeek V4
components. NanoDenseForCausalLM supplies the model entry point and standard
global grouped-query attention. The tokenizer is pinned to
DeepSeek-V4-Flash-Base, revision 8855555.
Architecture
| Setting | Value |
|---|---|
| Decoder layers | 24 |
| Residual hidden size | 1,536 |
| Dense SwiGLU intermediate size | 4,608 |
| Query / key-value heads | 16 / 2 |
| Head dimension | 128 |
| Attention projection width | 2,048 |
| Attention pattern | Global causal GQA in every layer |
| Normalization | DeepseekV4RMSNorm, epsilon 1e-6 |
| Position encoding | V4 interleaved RoPE over all 128 head dimensions |
| RoPE theta | 5,000,000 |
| Configured context target | 131,072 tokens |
| Vocabulary | 129,280 |
| Input embedding / output head | Separate weights |
| BOS / EOS / PAD | 0 / 1 / 2 |
| Total parameters | 1,076,700,672 |
| Excluding embedding and output head | 679,552,512 |
Each layer computes:
h = x + GQA(RMSNorm(x))
y = h + SwiGLU(RMSNorm(h))
The attention width is independent of the residual width: Q projects 1536 -> 2048, K/V each project 1536 -> 256, and O projects 2048 -> 1536. All 24 layers have independent parameters. Linear and embedding weights use normal initialization with standard deviation 0.02; RMSNorm weights start at one. No depth-dependent initialization scaling is applied.
128K is a configuration target, not a measured capability of an untrained model. Long-context performance requires a suitable training curriculum and evaluation. At batch size one, a complete 128K BF16 KV cache occupies 3 GiB, excluding weights, temporary tensors and allocator overhead.
Initialize for training
pip install -r requirements.txt
import torch
from transformers import AutoConfig, AutoModelForCausalLM, AutoTokenizer
repo = "bowang0911/Nano-1B"
config = AutoConfig.from_pretrained(repo, trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_config(
config,
trust_remote_code=True,
attn_implementation="sdpa",
dtype=torch.bfloat16,
)
# Randomly initialized model; pass it to your training loop / Trainer.
To inspect the full architecture without allocating weight storage:
with torch.device("meta"):
model = AutoModelForCausalLM.from_config(config, trust_remote_code=True)
print(sum(p.numel() for p in model.parameters())) # 1076700672
From a local copy, python count_parameters.py --meta checks the analytical
count against a full-size meta-device model.
Tokenized data and packing
tokenizer.json is byte-identical to the pinned DeepSeek V4 artifact. Tokenizer configuration sets model_max_length=131072 and selects the existing
dedicated <|▁pad▁|> token (ID 2) for padding. Vocabulary IDs and raw unpadded
encoding are unchanged; EOS remains ID 1. tokenizer_provenance.json records upstream and published checksums.
Existing data encoded with this pinned tokenizer can be reused directly.
The pretraining convention used here is raw text with add_special_tokens=False
and one EOS token (ID 1) appended per document. No chat template is provided.
For ordinary padded batches, pass a two-dimensional attention_mask and set
padding labels to -100. Default positions are derived from that mask, including
when continuing a KV cache. The tokenizer, model config and generation config all use PAD ID 2. The model
embedding has a dedicated zeroed padding row at ID 2; the EOS embedding remains
trainable. With DataCollatorForLanguageModeling(mlm=False), only PAD labels are
masked; real document-ending EOS labels retain ID 1. Previously PAD equaled EOS,
which caused that standard collator to discard EOS supervision.
For padding-free document packing, use one flattened row (batch_size=1) and
reset position_ids to zero at each document boundary. Omit attention_mask, or
pass an all-ones mask. The model isolates documents and masks the first label of
each subsequent document so the LM loss does not cross boundaries. Packing with
padding, nonzero restart positions or multiple batch rows raises an error.
KV caching is disabled during training and for packed sequences. Normal
unpacked evaluation/generation supports dynamic KV caching.
Eager and SDPA backends are validated on CPU. The FlashAttention adapter receives packed positions and dispatches separate document segments; running actual FlashAttention kernels requires a compatible GPU installation.
Validation
Validated with Python 3.12.11, PyTorch 2.12.0 and Transformers 5.9.0:
- Full-size meta parameter count, including the wider attention projections.
- An independent complex-number RoPE / attention reference, including large positions.
- Training loss, finite gradients through every parameter, an AdamW update, and gradient-checkpointing equivalence.
- Eager/SDPA cached chunked decoding and cached/uncached greedy generation.
- Left/right padding, packed-document isolation, boundary loss masking and FlashAttention varlen dispatch with external GPU kernels stubbed.
- FP32/BF16 save/reload and invalid configuration override rejection.
- Tokenizer parity with the pinned source.
- Standard-collator EOS preservation, padded backward pass, and PAD/EOS consistency after tokenizer serialization.
The 13 local behavioral tests are kept outside this model repository. This is small-model CPU validation, not full-size training, GPU throughput or 128K quality validation. Transformers is pinned because the implementation imports its V4 components and attention/cache interfaces.
Sources and licenses
- MiniCPM5-1B configuration: architecture dimensions.
- Transformers 5.9.0 DeepSeek V4: RMSNorm, dense MLP and interleaved RoPE.
- Model implementation: Apache-2.0, see
LICENSE. - DeepSeek tokenizer assets: MIT, see
LICENSE-TOKENIZER.
No pretrained MiniCPM or DeepSeek model weights are included or converted.
- Downloads last month
- 728