Hush-Nano

Hush-Nano-Chat

Hush-Nano-Chat is an English, single-turn instruction-tuned version of Soulitude/Hush-Nano. It starts from the 22M-parameter pretrained model and uses supervised fine-tuning (SFT) on instruction–response pairs.

Due to the model's limited parameters, its response can be inaccurate, incomplete, or inconsistent.

Model Details

Hush-Nano-Chat has the following features:

  • Type: Causal Language Models
  • Training Stage: Pretraining & Post-training
  • Architecture: transformers with RMSNorm, RoPE, SwiGLU, QK-Norm and tied word embeddings
  • Number of Parameters: 22M (22,621,056)
  • Number of Layers: 12
  • Number of Attention Heads (GQA): 6 for Q and 3 for KV
  • Context Length: 1,024

SFT Data

The base model was pretrained on 8.5B tokens (8,554,042,292) drawn from the following subsets:

Source Training tokens Share
FineWeb-Edu 4,539,286,619 53.07%
DCLM 2,890,209,171 33.79%
FineMath4plus 1,124,546,502 13.15%

Then it was fine-tuned on a mixture of the following datasets: unsloth/alpaca-cleaned, databricks/databricks-dolly-15k, and HuggingFaceH4/no_robots.

Each example is formatted as:

<|bos|>User:
{instruction + input}
Assistant:
{response}<|eos|>

Only the assistant response and ending EOS token contribute to the training loss.

Evaluation

Zero-shot normalized accuracy, evaluated in fp32 using EleutherAI/lm-evaluation-harness. Scores may vary slightly with the evaluation setup and environment.

PIQA ARC-Easy ARC-Challenge HellaSwag
Hush-Nano 58.27% 38.93% 21.84% 28.89%
Hush-Nano-Chat 59.09% 39.94% 22.44% 28.75%

Usage

This model includes custom Transformers code and so requires trust_remote_code=True.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Soulitude/Hush-Nano-Chat"
device = "cuda" if torch.cuda.is_available() else "cpu"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
).to(device).eval()

message = [
    {"role": "user", "content": "Explain why the sky appears blue in one sentence."},
]
inputs = tokenizer.apply_chat_template(
    message,
    add_generation_prompt=True,
    return_tensors="pt",
    return_dict=True,
).to(device)

with torch.inference_mode():
    output = model.generate(
        **inputs,
        max_new_tokens=512,
        do_sample=True,
        temperature=0.7,
        top_p=0.95,
        repetition_penalty=1.1,
        pad_token_id=tokenizer.pad_token_id,
        eos_token_id=tokenizer.eos_token_id,
    )

reply_ids = output[0, inputs["input_ids"].shape[1]:]
print(tokenizer.decode(reply_ids, skip_special_tokens=True).strip())

Intended Use and Limitations

Hush-Nano-Chat is intended for experimentation with small instruction-tuned language models. Its 22M-parameter size and 1,024-token context limit constrain instruction following, reasoning, factual reliability, and longer responses. The supervised data is primarily English, and this checkpoint is designed around single-turn prompts. Review generated output before relying on it.

License

Apache 2.0

Downloads last month
358
Safetensors
Model size
22.8M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Soulitude/Hush-Nano-Chat

Finetuned
(1)
this model