Star Supernova 2.0: Discrete Diffusion & 12-Expert MoE (Top-4 Routing, 4-bit)

Distilled with stepfun-ai/Step-3.7-Flash Intelligence, Strategy A Academic Reasoning & Gemma-4 Foundations

Star Supernova 2.0 is the next-generation evolution of Star Supernova, upgrading the base foundation to Google DeepMind's breakthrough DiffusionGemma discrete diffusion architecture while retaining 100% of the intelligence from the original base model (Gemma-4 E2B), Star Supernova's high-capacity 12-expert sparse Mixture-of-Experts (MoE) Top-4 dynamic routing engine, and Step-3.7-Flash's distilled mathematical reasoning traces.

By leveraging discrete diffusion parallel sequence refinement alongside Top-4 dynamic expert routing and a shared dense feed-forward network, Star Supernova 2.0 achieves ultra-fast parallel generation speeds and frontier-grade reasoning on Apple Silicon Macs.

Quantized to 4-bit affine precision, Star Supernova 2.0 runs natively on Apple Silicon Metal GPU via MLX with a memory footprint of only ~8.7 GB VRAM, leaving >3.4 GB of free GPU headroom under macOS's 12.1 GB Metal limit.

This enhanced 2.0 release unifies:

  1. DiffusionGemma Foundation: Parallel iterative denoising passes across token blocks, breaking sequential causal LLM bottlenecks.
  2. Old Base Model Intelligence: Foundational Gemma-4 principles: Clean Code patterns, concurrency vs. parallelism, bounded blocking queues, and affine quantization mechanics.
  3. Step-3.7-Flash Distilled Intelligence: Multi-step mathematical proofs, definite calculus integrals, thread-safe concurrency architectures, and dynamic MoE routing formulations.
  4. Star Supernova Top-4 Sparse Routing: 12 routed experts activating Top-4 candidates per token + 1 shared dense backbone for universal syntactic stability.
  5. Strategy A Academic & Professional Reasoning: Distractor-elimination Chain-of-Thought (CoT) traces and curated multiple-choice benchmarks from MMLU-Pro (Business Management, Economics, Systems Engineering, Formal Logic, and Computer Science).
  6. Problem-Formulation & Solution-Bias Elimination: Explicit training on solution-independent problem definition (e.g., why problem statements must remain technology-agnostic under McKinsey and Lean Six Sigma frameworks).
  7. Unsloth Studio Conversational Dialogue: Multi-turn developer interactions and tool execution traces.

Model Tree Lineage & Distillation

Role Model Identifier Details
New Primary Base Model google/diffusiongemma-26B-A4B-it DiffusionGemma discrete diffusion foundation architecture
Original Foundation Model unsloth/gemma-4-E2B-it-qat-q4_0-unquantized Gemma-4 foundational architecture & programming intelligence
Routing Predecessor monkiey/StarSupernova 12-expert MoE with Top-4 dynamic routing in 4-bit affine precision & Strategy A CoT
Distillation Teacher stepfun-ai/Step-3.7-Flash 198B Sparse MoE foundation model used for knowledge distillation
Synthesized Result Model monkiey/StarSupernova-2.0 Star Supernova 2.0 with all baseline, Star, and Step-3.7-Flash traces preserved

Step-3.7-Flash Distillation & English Language Policy: Core reasoning behaviors, dynamic sparse token routing mechanisms, and deep analytical capabilities were distilled directly from stepfun-ai/Step-3.7-Flash. All training datasets, CoT traces, and conversational fine-tuning were conducted strictly in English (0 non-English/CJK characters). DiffusionGemma serves as the primary base model, and Star Supernova 2.0 functions strictly as an English-language model.


Architecture & Specifications

Feature Specification
Model Name Star Supernova
Base Architecture Gemma-4 (35 Layers, Hidden Dim 1536)
Total Parameters ~13.47 Billion
Active Parameters / Token ~5.54 Billion (41.1% active / 58.9% sparse)
Routing Mechanism Sparse Top-4 Routing across 12 Experts + 1 Shared Dense MLP
Distillation Teacher stepfun-ai/Step-3.7-Flash
Reasoning Benchmarks Strategy A (MMLU-Pro CoT: Business, Economics, Engineering, Logic, CS)
Quantization 4-bit affine quantization (group_size=64, router projections in 8-bit)
Memory Footprint ~8.7 GB VRAM (Leaves >3.4 GB free on 16 GB Macs)
Inference Framework MLX / Apple Silicon Metal GPU
Context Length 131,072 tokens
Language English (en)
Weights Provided In-place fused 4-bit weights (model-*.safetensors) + Standalone LoRA adapters (adapters.safetensors)

Distillation & Training Details

  1. Teacher-Student MoE Distillation (Step-3.7-Flash):
    • High-capacity reasoning traces, sparse token dispatch mechanisms, and multi-step mathematical derivations were distilled into Star Supernova's 12-expert Top-4 routing pathways.
  2. Strategy A: Academic & Professional Reasoning (CoT):
    • Integrated Chain-of-Thought problem-solving traces across academic disciplines (Business, Economics, Systems Engineering, Philosophy, Computer Science).
    • Trained on explicit distractor elimination, teaching the model to identify solution-bias traps (e.g. recognizing that problem statements must remain technology-agnostic rather than prematurely presupposing technology).
  3. Native 4-Bit In-Place Weight Fusion:
    • LoRA adaptation parameters ($r=16, \alpha=16$) trained on attention projections (q_proj, k_proj, v_proj, o_proj) were mathematically fused directly into the 4-bit quantized base weights with zero dequantization drift.
  4. Structured Reasoning Preservation:
    • Retains Gemma-4's native <|channel>thought structured thinking channel for transparent, step-by-step reasoning.

Usage with MLX

Installation

pip install mlx-lm

Python Streaming Inference

import mlx_lm
from mlx_lm.sample_utils import make_sampler, make_logits_processors

model_id = "monkiey/StarSupernova"

print(f"Loading {model_id} on Apple Silicon GPU...")
model, tokenizer = mlx_lm.load(model_id)

messages = [
    {
        "role": "user",
        "content": "Which one of the following is not a characteristic of good problem statements?\n• Specific and Measurable where possible\n• Time Bound\n• Solved at the highest level possible\n• Factor impact of technology"
    }
]

prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)

sampler = make_sampler(temp=0.2, top_p=0.9)
logits_processors = make_logits_processors(repetition_penalty=1.15)

print("\n--- Generating Response ---")
for response in mlx_lm.stream_generate(
    model,
    tokenizer,
    prompt=prompt,
    max_tokens=512,
    sampler=sampler,
    logits_processors=logits_processors,
):
    print(response.text, end="", flush=True)
print()

Using MLX-LM CLI

mlx_lm.generate --model monkiey/StarSupernova --prompt "Explain how Top-4 sparse routing dynamically selects experts across 12 candidates in Star Supernova." --max-tokens 512

License

Apache 2.0

Downloads last month
1,912
Safetensors
Model size
26B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for monkiey/StarSupernova

Quantized
(37)
this model