How to use from
Hermes Agent
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "modilify/Modilify-Mk2-preview-mlx"
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup
# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default modilify/Modilify-Mk2-preview-mlx
Run Hermes
hermes
Quick Links

Modilify

Modilify Mk2 Preview · MLX

Modilify Mk2 is a rolling-canvas text diffusion model with learned trajectory memory. This release provides complete, unquantized native MLX weights and an inference package for Apple Silicon Macs. It builds on Google DiffusionGemma, adding LoRA adapters, GDN2 memory at two timescales, and a confidence-and-entropy prefix commit policy.

The model revises a 256-token canvas over successive denoising passes, commits a variable-length prefix, and carries memory forward as the canvas advances. Per-position and row-level working states update on every denoising pass; persistent memory updates only when tokens commit.

This is an experimental text-only preview. The complete model is included; no separate base-model download is required. PyTorch is not required.

Download and run

Use Python 3.12 on an Apple Silicon Mac with macOS 26 or newer for the pinned MLX environment. The weights occupy 48.23 GiB. Inference needs additional memory for KV caches, recurrent states, and computation. See local comparison for measured short-request memory usage.

python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install "huggingface-hub==1.20.1"

hf download Modilify/Modilify-Mk2-preview-mlx --local-dir ./Modilify-Mk2-preview-mlx
cd Modilify-Mk2-preview-mlx
python -m pip install .

python inference.py --prompt "Why is the sky blue?" \
  --max-new-tokens 256 --max-denoising-steps 256 --stream

The source entrypoint defaults to the model in the same directory. The installed command accepts a local model directory or a Hub model ID:

modilify-mlx --model ./ --prompt "你好,请介绍一下自己。" --stream
modilify-mlx --model Modilify/Modilify-Mk2-preview-mlx \
  --prompt "What is 2 + 3?" --max-new-tokens 128 --max-denoising-steps 128

The Hub ID path downloads the model snapshot into the Hugging Face cache. Later runs reuse the cached files. All inference computation runs locally.

Thinking is enabled by default; use --think False to disable it. --max-new-tokens limits output length; --max-denoising-steps limits denoising work. Reaching either budget may truncate a response. The CLI defaults to 8192 output tokens and the generation configuration's 48 denoising steps. Set both budgets explicitly for your workload. --canvas-length can reduce the canvas from its maximum of 256. Without --stream, the CLI emits JSONL events and a final result with text, token IDs, stop reason, and timing/memory metrics.

Python API

Load once and reuse the runtime for independent requests:

from inference import load_model, generate

runtime = load_model("./")
result = generate(
    runtime,
    "Explain diffusion models in one paragraph.",
    max_new_tokens=256,
    max_denoising_steps=256,
    seed=42,
    think=True,
)
print(result["text"])
print(result["stop_reason"])

Pass messages=[{"role": "user", "content": "..."}] instead of prompt for a conversation. The tokenizer's Modilify template handles system, user, and assistant turns. With thinking enabled, generated text may include reasoning before the answer.

Continuous requests

printf '%s\n' \
  '{"request_id":"a","prompt":"你好","max_new_tokens":128,"max_denoising_steps":128}' \
  '{"request_id":"b","prompt":"What is 2 + 3?","seed":42,"think":false,"max_new_tokens":128,"max_denoising_steps":128}' | \
  modilify-mlx --model ./ --requests-jsonl - \
    --batch-size 4 --pipeline-depth 2 --max-batch-tokens 1024

Each line requires a unique request_id and either prompt or messages. Optional fields are max_new_tokens, max_denoising_steps, seed, and think. The default scheduler interleaves independent rows and keeps their random streams and states separate. --batch-size controls resident requests; --pipeline-depth bounds in-flight denoises.

Release details

Item Value
Base model DiffusionGemma 26B-A4B-it text backbone
Saved training step 1250
Memory topology schema25 / compact_gdn2_v2
Memory scheme dual_timescale_gdn2_trajectory_memory
Canvas Up to 256 tokens
Expert routing 8 of 128 experts per token
Dense / expert LoRA rank 16 / 8
Weight precision BF16, with small trajectory norms and dynamics parameters in FP32
Proposal constraints top_k 40, min_p 0.05
Commit defaults failure budget 0.2, target confidence 0.5
Weights 32 Safetensors shards, 1323 tensors

The complete backbone, adapters, and memory parameters are stored in the native MLX parameter layout. Adapters remain unfused to preserve the reference BF16 computation. Loading validates model topology, tensor shape/dtype, and shard SHA256 hashes. export_manifest.json contains the weight index metadata and release provenance; no optimizer state or training data is shipped.

This model uses its own rolling-canvas generation loop. Use the included inference.py, modilify-mlx, or Python API. Generic mlx_lm.generate, Transformers AutoModel.from_pretrained, and standard autoregressive serving backends do not implement this protocol. Image, audio, and video inputs are not supported by this release.

Training, evaluation, and limitations

The text backbone is adapted with dense and expert LoRA and trained trajectory memory. Adaptation uses response supervision over the valid rolling canvas, with terminal, confidence, and commit-readiness objectives. This release uses the saved step 1250 weights. The adaptation dataset is not distributed and its source composition is not documented in this release; refer to the upstream model card for base-model training information.

Local comparisons establish export correctness for the tested cases, not task quality or safety. No capability benchmark score is claimed for this preview, and upstream benchmark scores do not describe this modified model. Long-context quality, broad language coverage, and deployment safety have not been established by those checks.

Local comparison

On 2026-10-06, the package installed in a fresh Python 3.12 environment without PyTorch and ran offline. Two requests matched the strict step 1250 source checkpoint at 1/2/5/10/20 denoising steps, including token IDs, stopping and canvas shifts. At five steps, interleaved B2 also matched independent B1 outputs and all three GDN2 state hashes. Short-request MLX peak allocations were about 50.79 GiB; longer workloads require additional memory.

Intended use and impact statement

This preview is intended for research on text diffusion, trajectory memory, local generation, and reproducible inference experiments. Generated statements can be incorrect, biased, or inappropriate; verify consequential outputs and keep qualified human review for sensitive decisions. The model is not intended to make autonomous high-risk decisions. Do not assume local execution alone provides safety, privacy, or resistance to prompt injection. Deployment should include use-specific testing, appropriate safeguards, and incident monitoring.

License and attribution

Modilify's original contributions and distribution are covered by the Modilify Open Model License 1.0, a custom license. Upstream components retain their own terms; the DiffusionGemma base is published under Apache 2.0. See NOTICE.md for attribution and upstream notices.

Maintained by Modilify. Report model and runtime issues through the Hub discussions.

Downloads last month
190
Safetensors
Model size
26B params
Tensor type
F32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including modilify/Modilify-Mk2-preview-mlx