How to use from
Docker Model Runner
docker model run hf.co/Modujo-AI/ModujoMoE
Quick Links

Modujo model weights

This repository keeps model variants in self-contained subdirectories. The repository name is retained for compatibility, but the repository root no longer contains model weights and must not be loaded directly.

Weight directories

Path Model Meaning Status
pretrain/Modujo-1B-A0.75B/ Modujo-1B-A0.75B Continued-pretraining base model: 1B total scale and A0.75B active scale Available
Repository root Historical 9B-A1B architecture metadata and shared tokenizer files The original random-initialized 9B weight shards were removed; this is not a loadable model directory No weights

There are currently no SFT, chat, RL, or looped-model weights in this repository. New variants should be published in their own named directories so their training stage and actual parameter size remain explicit.

Pretrained base model

pretrain/Modujo-1B-A0.75B/ is the released pretraining artifact.

  • Architecture: Qwen4ExpForCausalLM
  • Model size: 1B total scale
  • Active size: approximately A0.75B per token
  • 36 layers with 8 routed experts per layer, top-2 routing, and one shared expert
  • Attention layout: repeating 3 Gated DeltaNet layers + 1 dense-attention layer
  • Context configuration: 32K maximum positions; training sequences were up to 2,048 tokens
  • Weight format: BF16 safetensors
  • Training stage: continued-pretraining base model
  • Not instruction-tuned and not intended to be treated as a chat model
  • No QSA sparse indexer in this release

Detailed machine-readable size metadata is recorded in parameter_summary.json.

Loading

Pass the subdirectory explicitly:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "Alexhu1999/Modujo-9B-A1B"
subfolder = "pretrain/Modujo-1B-A0.75B"

tokenizer = AutoTokenizer.from_pretrained(repo_id, subfolder=subfolder)
model = AutoModelForCausalLM.from_pretrained(
    repo_id,
    subfolder=subfolder,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

Loading only Alexhu1999/Modujo-9B-A1B without subfolder will fail because there are intentionally no weights at the repository root.

Planned experiments

The current base model will be evaluated through three separate tracks:

  • SFT: improve instruction following, response quality, repetition control, and multi-turn dialogue stability.
  • QSA: add and train sparse attention indexers, then compare long-context quality, inference speed, and memory use against dense attention.
  • Looped Transformer: reuse selected Transformer layers to test whether deeper computation with shared parameters provides a practical quality and efficiency benefit.

These tracks will be evaluated independently before any combined model is considered. Future weights will use separate directories with explicit names.

Limitations

This is a base language model checkpoint. It can perform short text completion, but long generations may repeat and instruction following is not yet stable. Use a separately identified SFT or aligned release for assistant/chat use when one becomes available.

Downloads last month
457
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support