Instructions to use monkiey/StarSupernova with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use monkiey/StarSupernova with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("monkiey/StarSupernova") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Unsloth Desktop
- Pi
How to use monkiey/StarSupernova with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "monkiey/StarSupernova"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "monkiey/StarSupernova" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use monkiey/StarSupernova with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "monkiey/StarSupernova"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "monkiey/StarSupernova" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "monkiey/StarSupernova", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use monkiey/StarSupernova with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "monkiey/StarSupernova"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default monkiey/StarSupernova
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use monkiey/StarSupernova with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "monkiey/StarSupernova"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "monkiey/StarSupernova" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Star Supernova 2.0: Discrete Diffusion & 12-Expert MoE (Top-4 Routing, 4-bit)
Distilled with stepfun-ai/Step-3.7-Flash Intelligence, Strategy A Academic Reasoning & Gemma-4 Foundations
Star Supernova 2.0 is the next-generation evolution of Star Supernova, upgrading the base foundation to Google DeepMind's breakthrough DiffusionGemma discrete diffusion architecture while retaining 100% of the intelligence from the original base model (Gemma-4 E2B), Star Supernova's high-capacity 12-expert sparse Mixture-of-Experts (MoE) Top-4 dynamic routing engine, and Step-3.7-Flash's distilled mathematical reasoning traces.
By leveraging discrete diffusion parallel sequence refinement alongside Top-4 dynamic expert routing and a shared dense feed-forward network, Star Supernova 2.0 achieves ultra-fast parallel generation speeds and frontier-grade reasoning on Apple Silicon Macs.
Quantized to 4-bit affine precision, Star Supernova 2.0 runs natively on Apple Silicon Metal GPU via MLX with a memory footprint of only ~8.7 GB VRAM, leaving >3.4 GB of free GPU headroom under macOS's 12.1 GB Metal limit.
This enhanced 2.0 release unifies:
- DiffusionGemma Foundation: Parallel iterative denoising passes across token blocks, breaking sequential causal LLM bottlenecks.
- Old Base Model Intelligence: Foundational Gemma-4 principles: Clean Code patterns, concurrency vs. parallelism, bounded blocking queues, and affine quantization mechanics.
- Step-3.7-Flash Distilled Intelligence: Multi-step mathematical proofs, definite calculus integrals, thread-safe concurrency architectures, and dynamic MoE routing formulations.
- Star Supernova Top-4 Sparse Routing: 12 routed experts activating Top-4 candidates per token + 1 shared dense backbone for universal syntactic stability.
- Strategy A Academic & Professional Reasoning: Distractor-elimination Chain-of-Thought (CoT) traces and curated multiple-choice benchmarks from MMLU-Pro (Business Management, Economics, Systems Engineering, Formal Logic, and Computer Science).
- Problem-Formulation & Solution-Bias Elimination: Explicit training on solution-independent problem definition (e.g., why problem statements must remain technology-agnostic under McKinsey and Lean Six Sigma frameworks).
- Unsloth Studio Conversational Dialogue: Multi-turn developer interactions and tool execution traces.
Model Tree Lineage & Distillation
| Role | Model Identifier | Details |
|---|---|---|
| New Primary Base Model | google/diffusiongemma-26B-A4B-it |
DiffusionGemma discrete diffusion foundation architecture |
| Original Foundation Model | unsloth/gemma-4-E2B-it-qat-q4_0-unquantized |
Gemma-4 foundational architecture & programming intelligence |
| Routing Predecessor | monkiey/StarSupernova |
12-expert MoE with Top-4 dynamic routing in 4-bit affine precision & Strategy A CoT |
| Distillation Teacher | stepfun-ai/Step-3.7-Flash |
198B Sparse MoE foundation model used for knowledge distillation |
| Synthesized Result Model | monkiey/StarSupernova-2.0 |
Star Supernova 2.0 with all baseline, Star, and Step-3.7-Flash traces preserved |
Step-3.7-Flash Distillation & English Language Policy: Core reasoning behaviors, dynamic sparse token routing mechanisms, and deep analytical capabilities were distilled directly from
stepfun-ai/Step-3.7-Flash. All training datasets, CoT traces, and conversational fine-tuning were conducted strictly in English (0 non-English/CJK characters). DiffusionGemma serves as the primary base model, and Star Supernova 2.0 functions strictly as an English-language model.
Architecture & Specifications
| Feature | Specification |
|---|---|
| Model Name | Star Supernova |
| Base Architecture | Gemma-4 (35 Layers, Hidden Dim 1536) |
| Total Parameters | ~13.47 Billion |
| Active Parameters / Token | ~5.54 Billion (41.1% active / 58.9% sparse) |
| Routing Mechanism | Sparse Top-4 Routing across 12 Experts + 1 Shared Dense MLP |
| Distillation Teacher | stepfun-ai/Step-3.7-Flash |
| Reasoning Benchmarks | Strategy A (MMLU-Pro CoT: Business, Economics, Engineering, Logic, CS) |
| Quantization | 4-bit affine quantization (group_size=64, router projections in 8-bit) |
| Memory Footprint | ~8.7 GB VRAM (Leaves >3.4 GB free on 16 GB Macs) |
| Inference Framework | MLX / Apple Silicon Metal GPU |
| Context Length | 131,072 tokens |
| Language | English (en) |
| Weights Provided | In-place fused 4-bit weights (model-*.safetensors) + Standalone LoRA adapters (adapters.safetensors) |
Distillation & Training Details
- Teacher-Student MoE Distillation (Step-3.7-Flash):
- High-capacity reasoning traces, sparse token dispatch mechanisms, and multi-step mathematical derivations were distilled into Star Supernova's 12-expert Top-4 routing pathways.
- Strategy A: Academic & Professional Reasoning (CoT):
- Integrated Chain-of-Thought problem-solving traces across academic disciplines (Business, Economics, Systems Engineering, Philosophy, Computer Science).
- Trained on explicit distractor elimination, teaching the model to identify solution-bias traps (e.g. recognizing that problem statements must remain technology-agnostic rather than prematurely presupposing technology).
- Native 4-Bit In-Place Weight Fusion:
- LoRA adaptation parameters ($r=16, \alpha=16$) trained on attention projections (
q_proj,k_proj,v_proj,o_proj) were mathematically fused directly into the 4-bit quantized base weights with zero dequantization drift.
- LoRA adaptation parameters ($r=16, \alpha=16$) trained on attention projections (
- Structured Reasoning Preservation:
- Retains Gemma-4's native
<|channel>thoughtstructured thinking channel for transparent, step-by-step reasoning.
- Retains Gemma-4's native
Usage with MLX
Installation
pip install mlx-lm
Python Streaming Inference
import mlx_lm
from mlx_lm.sample_utils import make_sampler, make_logits_processors
model_id = "monkiey/StarSupernova"
print(f"Loading {model_id} on Apple Silicon GPU...")
model, tokenizer = mlx_lm.load(model_id)
messages = [
{
"role": "user",
"content": "Which one of the following is not a characteristic of good problem statements?\n• Specific and Measurable where possible\n• Time Bound\n• Solved at the highest level possible\n• Factor impact of technology"
}
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
sampler = make_sampler(temp=0.2, top_p=0.9)
logits_processors = make_logits_processors(repetition_penalty=1.15)
print("\n--- Generating Response ---")
for response in mlx_lm.stream_generate(
model,
tokenizer,
prompt=prompt,
max_tokens=512,
sampler=sampler,
logits_processors=logits_processors,
):
print(response.text, end="", flush=True)
print()
Using MLX-LM CLI
mlx_lm.generate --model monkiey/StarSupernova --prompt "Explain how Top-4 sparse routing dynamically selects experts across 12 candidates in Star Supernova." --max-tokens 512
License
Apache 2.0
- Downloads last month
- 1,912
4-bit
Model tree for monkiey/StarSupernova
Base model
google/diffusiongemma-26B-A4B-it