Model License Params Language


πŸ“Š Benchmarks

MetaNova-1 leads the series on ArithMark-2 β€” outperforming GPT-2 (124M) with less than half the parameters.

Benchmark MetaNova-Test MetaNova-0.1 MetaNova-1 Supra-50M-Base OpenAI/GPT-2
Params 70.55M 70.55M 62.70M 51.79M 124M
HellaSwag 25.37% 27.40% 27.08% 31.65% 31.26%
ARC-Easy 26.89% 31.73% 30.56% 45.58% 39.35%
ARC-Challenge 25.09% 25.00% 24.15% 24.66% 22.35%
PIQA 52.29% 58.54% 56.04% 61.53% 62.08%
ArithMark-2 24.12% 27.12% 34.12% 26.08% 26.48%
ARC Avg 25.99% 28.37% 27.35% 35.12% 30.85%
Final Avg 31.94% 35.36% 36.15% 38.60% 37.67%

πŸ’¬ Chat Format

MetaNova-1 uses the DeepSeek / Qwen im_start format with optional thinking mode.

🧠 Thinking Mode β€” append /think
<|im_start|>user
{query} /think<|im_end|>
<|im_start|>assistant
<think>
{thinking_content}
</think>

{response}<|im_end|>
⚑ Non-Thinking Mode β€” append /no_think
<|im_start|>user
{query} /no_think<|im_end|>
<|im_start|>assistant
<think>

</think>

{response}<|im_end|>

πŸš€ Quick Start

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "WhirlwindAI/MetaNova-1-60M"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)

prompt = "<|im_start|>user\nWhat is 14 Γ— 27? /think<|im_end|>\n<|im_start|>assistant\n"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(outputs[0], skip_special_tokens=False))

🧩 Model Details

Architecture Causal Language Model
Parameters 62.70M
Organization WhirlwindAI
License Apache 2.0
Language English
Chat Format DeepSeek / Qwen (im_start)
Thinking Mode Yes (/think / /no_think)

Built by WhirlwindAI Β· Apache 2.0

Downloads last month
38
Safetensors
Model size
62.7M params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for WhirlwindAI/MetaNova-1-60M

Unable to build the model tree, the base model loops to the model itself. Learn more.

Space using WhirlwindAI/MetaNova-1-60M 1