Text Generation
MLX
Safetensors
modilify_mk2
diffusion
mixture-of-experts
custom-code
modilify-mk2
conversational
Instructions to use modilify/Modilify-Mk2-preview-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use modilify/Modilify-Mk2-preview-mlx with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("modilify/Modilify-Mk2-preview-mlx") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use modilify/Modilify-Mk2-preview-mlx with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "modilify/Modilify-Mk2-preview-mlx"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "modilify/Modilify-Mk2-preview-mlx" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use modilify/Modilify-Mk2-preview-mlx with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "modilify/Modilify-Mk2-preview-mlx"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "modilify/Modilify-Mk2-preview-mlx" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "modilify/Modilify-Mk2-preview-mlx", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use modilify/Modilify-Mk2-preview-mlx with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "modilify/Modilify-Mk2-preview-mlx"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default modilify/Modilify-Mk2-preview-mlx
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use modilify/Modilify-Mk2-preview-mlx with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "modilify/Modilify-Mk2-preview-mlx"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "modilify/Modilify-Mk2-preview-mlx" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
File size: 7,889 Bytes
e4f7326 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 | ---
license: other
license_name: modilify-open-model-license-1.0
license_link: LICENSE
library_name: mlx
pipeline_tag: text-generation
tags:
- mlx
- diffusion
- mixture-of-experts
- custom-code
- modilify-mk2
---

# Modilify Mk2 Preview · MLX
Modilify Mk2 is a rolling-canvas text diffusion model with learned trajectory memory. This release provides complete, unquantized native MLX weights and an inference package for Apple Silicon Macs. It builds on [Google DiffusionGemma](https://huggingface.co/google/diffusiongemma-26B-A4B-it), adding LoRA adapters, GDN2 memory at two timescales, and a confidence-and-entropy prefix commit policy.
The model revises a 256-token canvas over successive denoising passes, commits a variable-length prefix, and carries memory forward as the canvas advances. Per-position and row-level working states update on every denoising pass; persistent memory updates only when tokens commit.
This is an experimental **text-only preview**. The complete model is included; no separate base-model download is required. PyTorch is not required.
## Download and run
Use Python 3.12 on an Apple Silicon Mac with macOS 26 or newer for the pinned MLX environment. The weights occupy **48.23 GiB**. Inference needs additional memory for KV caches, recurrent states, and computation. See [local comparison](#local-comparison) for measured short-request memory usage.
```bash
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install "huggingface-hub==1.20.1"
hf download Modilify/Modilify-Mk2-preview-mlx --local-dir ./Modilify-Mk2-preview-mlx
cd Modilify-Mk2-preview-mlx
python -m pip install .
python inference.py --prompt "Why is the sky blue?" \
--max-new-tokens 256 --max-denoising-steps 256 --stream
```
The source entrypoint defaults to the model in the same directory. The installed command accepts a local model directory or a Hub model ID:
```bash
modilify-mlx --model ./ --prompt "你好,请介绍一下自己。" --stream
modilify-mlx --model Modilify/Modilify-Mk2-preview-mlx \
--prompt "What is 2 + 3?" --max-new-tokens 128 --max-denoising-steps 128
```
The Hub ID path downloads the model snapshot into the Hugging Face cache. Later runs reuse the cached files. All inference computation runs locally.
Thinking is enabled by default; use `--think False` to disable it. `--max-new-tokens` limits output length; `--max-denoising-steps` limits denoising work. Reaching either budget may truncate a response. The CLI defaults to 8192 output tokens and the generation configuration's 48 denoising steps. Set both budgets explicitly for your workload. `--canvas-length` can reduce the canvas from its maximum of 256. Without `--stream`, the CLI emits JSONL events and a final result with text, token IDs, stop reason, and timing/memory metrics.
## Python API
Load once and reuse the runtime for independent requests:
```python
from inference import load_model, generate
runtime = load_model("./")
result = generate(
runtime,
"Explain diffusion models in one paragraph.",
max_new_tokens=256,
max_denoising_steps=256,
seed=42,
think=True,
)
print(result["text"])
print(result["stop_reason"])
```
Pass `messages=[{"role": "user", "content": "..."}]` instead of `prompt` for a conversation. The tokenizer's Modilify template handles system, user, and assistant turns. With thinking enabled, generated text may include reasoning before the answer.
## Continuous requests
```bash
printf '%s\n' \
'{"request_id":"a","prompt":"你好","max_new_tokens":128,"max_denoising_steps":128}' \
'{"request_id":"b","prompt":"What is 2 + 3?","seed":42,"think":false,"max_new_tokens":128,"max_denoising_steps":128}' | \
modilify-mlx --model ./ --requests-jsonl - \
--batch-size 4 --pipeline-depth 2 --max-batch-tokens 1024
```
Each line requires a unique `request_id` and either `prompt` or `messages`. Optional fields are `max_new_tokens`, `max_denoising_steps`, `seed`, and `think`. The default scheduler interleaves independent rows and keeps their random streams and states separate. `--batch-size` controls resident requests; `--pipeline-depth` bounds in-flight denoises.
## Release details
| Item | Value |
| --- | --- |
| Base model | DiffusionGemma 26B-A4B-it text backbone |
| Saved training step | 1250 |
| Memory topology | schema25 / `compact_gdn2_v2` |
| Memory scheme | `dual_timescale_gdn2_trajectory_memory` |
| Canvas | Up to 256 tokens |
| Expert routing | 8 of 128 experts per token |
| Dense / expert LoRA rank | 16 / 8 |
| Weight precision | BF16, with small trajectory norms and dynamics parameters in FP32 |
| Proposal constraints | top_k 40, min_p 0.05 |
| Commit defaults | failure budget 0.2, target confidence 0.5 |
| Weights | 32 Safetensors shards, 1323 tensors |
The complete backbone, adapters, and memory parameters are stored in the native MLX parameter layout. Adapters remain unfused to preserve the reference BF16 computation. Loading validates model topology, tensor shape/dtype, and shard SHA256 hashes. `export_manifest.json` contains the weight index metadata and release provenance; no optimizer state or training data is shipped.
This model uses its own rolling-canvas generation loop. Use the included `inference.py`, `modilify-mlx`, or Python API. Generic `mlx_lm.generate`, Transformers `AutoModel.from_pretrained`, and standard autoregressive serving backends do not implement this protocol. Image, audio, and video inputs are not supported by this release.
## Training, evaluation, and limitations
The text backbone is adapted with dense and expert LoRA and trained trajectory memory. Adaptation uses response supervision over the valid rolling canvas, with terminal, confidence, and commit-readiness objectives. This release uses the saved step 1250 weights. The adaptation dataset is not distributed and its source composition is not documented in this release; refer to the upstream model card for base-model training information.
Local comparisons establish export correctness for the tested cases, not task quality or safety. No capability benchmark score is claimed for this preview, and upstream benchmark scores do not describe this modified model. Long-context quality, broad language coverage, and deployment safety have not been established by those checks.
## Local comparison
On 2026-10-06, the package installed in a fresh Python 3.12 environment without PyTorch and ran offline. Two requests matched the strict step 1250 source checkpoint at 1/2/5/10/20 denoising steps, including token IDs, stopping and canvas shifts. At five steps, interleaved B2 also matched independent B1 outputs and all three GDN2 state hashes. Short-request MLX peak allocations were about 50.79 GiB; longer workloads require additional memory.
## Intended use and impact statement
This preview is intended for research on text diffusion, trajectory memory, local generation, and reproducible inference experiments. Generated statements can be incorrect, biased, or inappropriate; verify consequential outputs and keep qualified human review for sensitive decisions. The model is not intended to make autonomous high-risk decisions. Do not assume local execution alone provides safety, privacy, or resistance to prompt injection. Deployment should include use-specific testing, appropriate safeguards, and incident monitoring.
## License and attribution
Modilify's original contributions and distribution are covered by the [Modilify Open Model License 1.0](LICENSE), a custom license. Upstream components retain their own terms; the DiffusionGemma base is published under Apache 2.0. See [NOTICE.md](NOTICE.md) for attribution and upstream notices.
Maintained by **Modilify**. Report model and runtime issues through the [Hub discussions](https://huggingface.co/Modilify/Modilify-Mk2-preview-mlx/discussions).
|