File size: 2,122 Bytes
53d5244
 
 
 
 
b88f761
53d5244
 
 
 
b88f761
53d5244
 
 
 
b88f761
53d5244
b88f761
 
 
 
53d5244
b88f761
 
53d5244
b88f761
53d5244
b88f761
 
 
53d5244
 
 
b88f761
53d5244
b88f761
 
 
 
53d5244
 
 
b88f761
 
 
 
53d5244
 
 
 
 
 
b88f761
 
 
53d5244
 
b88f761
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
---
license: other
license_name: modilify-open-model-license-1.0
license_link: LICENSE
library_name: transformers
pipeline_tag: text-generation
tags:
  - diffusion
  - mixture-of-experts
  - trust-remote-code
  - safetensors
---

![LOGO](assets/01-LOGO.jpg)

# Modilify Mk2 Preview

PyTorch text inference for the **step 1250, schema25** checkpoint migrated
from `Modilify-Mk2-preview-mlx`. The model uses a shared DiffusionGemma text
encoder/decoder, a rolling 256-token canvas, and dual-timescale GDN2 memory.
The release contains 32 Safetensors shards totaling **48.23 GiB**.

Dense and expert LoRA adapters remain unfused. Model loading preserves the
26 FP32 GDN2/norm parameters alongside BF16 weights. No vision tower is included.

## Inference

Use Transformers **5.14.1**, PyTorch, Accelerate, and Safetensors. Run from
this directory or replace `model_path` with its location. Inference needs
additional memory beyond the weights.

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_path = "."
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_path,
    trust_remote_code=True,
    dtype=torch.bfloat16,
    device_map="auto",
    attn_implementation="sdpa",
).eval()
inputs = tokenizer.apply_chat_template(
    [{"role": "user", "content": "Explain why the sky is blue."}],
    tokenize=True,
    add_generation_prompt=True,
    enable_thinking=False,
    return_dict=True,
    return_tensors="pt",
).to(model.device)
output = model.generate(**inputs, max_new_tokens=128, seed=42)
print(tokenizer.decode(output.sequences[0, inputs.input_ids.shape[1]:],
                       skip_special_tokens=True))
```

For MPS, use `device_map={"": "mps"}`. Set `enable_thinking=True` in the chat
template to request thinking. Generation uses temperature 0.8, top-k 40,
min-p 0.05, target confidence 0.5, and failure budget 0.2. `max_denoising_steps`
can bound generation. Static batches and continuous batching share the
confidence-prefix commit policy; each request retains its own cache and RNG.