File size: 3,379 Bytes
a18b154
 
 
 
 
 
 
9ff40ec
a18b154
 
9ff40ec
a18b154
9ff40ec
 
 
a18b154
9ff40ec
a18b154
9ff40ec
 
66bdf51
9ff40ec
 
 
 
 
 
 
 
 
 
 
66bdf51
 
9ff40ec
 
 
 
 
 
 
 
66bdf51
9ff40ec
 
 
 
 
a18b154
 
9ff40ec
a18b154
 
9ff40ec
 
 
 
 
 
 
 
 
 
a18b154
9ff40ec
 
 
 
b152483
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9ff40ec
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
---
library_name: transformers
pipeline_tag: text-generation
tags:
  - qwen4-exp
  - mixture-of-experts
  - language-model
  - pretrained
---

# Modujo model weights

This repository keeps model variants in self-contained subdirectories. The
repository name is retained for compatibility, but the repository root no
longer contains model weights and must not be loaded directly.

## Weight directories

| Path | Model | Meaning | Status |
| --- | --- | --- | --- |
| [`pretrain/Modujo-1B-A0.75B/`](./pretrain/Modujo-1B-A0.75B) | Modujo-1B-A0.75B | Continued-pretraining base model: 1B total scale and A0.75B active scale | Available |
| Repository root | Historical 9B-A1B architecture metadata and shared tokenizer files | The original random-initialized 9B weight shards were removed; this is not a loadable model directory | No weights |

There are currently no SFT, chat, RL, or looped-model weights in this
repository. New variants should be published in their own named directories so
their training stage and actual parameter size remain explicit.

## Pretrained base model

`pretrain/Modujo-1B-A0.75B/` is the released pretraining artifact.

- Architecture: `Qwen4ExpForCausalLM`
- Model size: 1B total scale
- Active size: approximately A0.75B per token
- 36 layers with 8 routed experts per layer, top-2 routing, and one shared expert
- Attention layout: repeating 3 Gated DeltaNet layers + 1 dense-attention layer
- Context configuration: 32K maximum positions; training sequences were up to 2,048 tokens
- Weight format: BF16 safetensors
- Training stage: continued-pretraining base model
- Not instruction-tuned and not intended to be treated as a chat model
- No QSA sparse indexer in this release

Detailed machine-readable size metadata is recorded in
[`parameter_summary.json`](./pretrain/Modujo-1B-A0.75B/parameter_summary.json).

## Loading

Pass the subdirectory explicitly:

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "Alexhu1999/Modujo-9B-A1B"
subfolder = "pretrain/Modujo-1B-A0.75B"

tokenizer = AutoTokenizer.from_pretrained(repo_id, subfolder=subfolder)
model = AutoModelForCausalLM.from_pretrained(
    repo_id,
    subfolder=subfolder,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)
```

Loading only `Alexhu1999/Modujo-9B-A1B` without `subfolder` will fail because
there are intentionally no weights at the repository root.

## Planned experiments

The current base model will be evaluated through three separate tracks:

- **SFT:** improve instruction following, response quality, repetition control,
  and multi-turn dialogue stability.
- **QSA:** add and train sparse attention indexers, then compare long-context
  quality, inference speed, and memory use against dense attention.
- **Looped Transformer:** reuse selected Transformer layers to test whether
  deeper computation with shared parameters provides a practical quality and
  efficiency benefit.

These tracks will be evaluated independently before any combined model is
considered. Future weights will use separate directories with explicit names.

## Limitations

This is a base language model checkpoint. It can perform short text completion,
but long generations may repeat and instruction following is not yet stable.
Use a separately identified SFT or aligned release for assistant/chat use when
one becomes available.