--- license: other license_name: modilify-open-model-license-1.0 license_link: LICENSE library_name: mlx pipeline_tag: text-generation tags: - mlx - diffusion - mixture-of-experts - custom-code - modilify-mk2 --- ![Modilify](assets/01-LOGO.jpg) # Modilify Mk2 Preview · MLX Modilify Mk2 is a rolling-canvas text diffusion model with learned trajectory memory. This release provides complete, unquantized native MLX weights and an inference package for Apple Silicon Macs. It builds on [Google DiffusionGemma](https://huggingface.co/google/diffusiongemma-26B-A4B-it), adding LoRA adapters, GDN2 memory at two timescales, and a confidence-and-entropy prefix commit policy. The model revises a 256-token canvas over successive denoising passes, commits a variable-length prefix, and carries memory forward as the canvas advances. Per-position and row-level working states update on every denoising pass; persistent memory updates only when tokens commit. This is an experimental **text-only preview**. The complete model is included; no separate base-model download is required. PyTorch is not required. ## Download and run Use Python 3.12 on an Apple Silicon Mac with macOS 26 or newer for the pinned MLX environment. The weights occupy **48.23 GiB**. Inference needs additional memory for KV caches, recurrent states, and computation. See [local comparison](#local-comparison) for measured short-request memory usage. ```bash python3.12 -m venv .venv source .venv/bin/activate python -m pip install "huggingface-hub==1.20.1" hf download Modilify/Modilify-Mk2-preview-mlx --local-dir ./Modilify-Mk2-preview-mlx cd Modilify-Mk2-preview-mlx python -m pip install . python inference.py --prompt "Why is the sky blue?" \ --max-new-tokens 256 --max-denoising-steps 256 --stream ``` The source entrypoint defaults to the model in the same directory. The installed command accepts a local model directory or a Hub model ID: ```bash modilify-mlx --model ./ --prompt "你好,请介绍一下自己。" --stream modilify-mlx --model Modilify/Modilify-Mk2-preview-mlx \ --prompt "What is 2 + 3?" --max-new-tokens 128 --max-denoising-steps 128 ``` The Hub ID path downloads the model snapshot into the Hugging Face cache. Later runs reuse the cached files. All inference computation runs locally. Thinking is enabled by default; use `--think False` to disable it. `--max-new-tokens` limits output length; `--max-denoising-steps` limits denoising work. Reaching either budget may truncate a response. The CLI defaults to 8192 output tokens and the generation configuration's 48 denoising steps. Set both budgets explicitly for your workload. `--canvas-length` can reduce the canvas from its maximum of 256. Without `--stream`, the CLI emits JSONL events and a final result with text, token IDs, stop reason, and timing/memory metrics. ## Python API Load once and reuse the runtime for independent requests: ```python from inference import load_model, generate runtime = load_model("./") result = generate( runtime, "Explain diffusion models in one paragraph.", max_new_tokens=256, max_denoising_steps=256, seed=42, think=True, ) print(result["text"]) print(result["stop_reason"]) ``` Pass `messages=[{"role": "user", "content": "..."}]` instead of `prompt` for a conversation. The tokenizer's Modilify template handles system, user, and assistant turns. With thinking enabled, generated text may include reasoning before the answer. ## Continuous requests ```bash printf '%s\n' \ '{"request_id":"a","prompt":"你好","max_new_tokens":128,"max_denoising_steps":128}' \ '{"request_id":"b","prompt":"What is 2 + 3?","seed":42,"think":false,"max_new_tokens":128,"max_denoising_steps":128}' | \ modilify-mlx --model ./ --requests-jsonl - \ --batch-size 4 --pipeline-depth 2 --max-batch-tokens 1024 ``` Each line requires a unique `request_id` and either `prompt` or `messages`. Optional fields are `max_new_tokens`, `max_denoising_steps`, `seed`, and `think`. The default scheduler interleaves independent rows and keeps their random streams and states separate. `--batch-size` controls resident requests; `--pipeline-depth` bounds in-flight denoises. ## Release details | Item | Value | | --- | --- | | Base model | DiffusionGemma 26B-A4B-it text backbone | | Saved training step | 1250 | | Memory topology | schema25 / `compact_gdn2_v2` | | Memory scheme | `dual_timescale_gdn2_trajectory_memory` | | Canvas | Up to 256 tokens | | Expert routing | 8 of 128 experts per token | | Dense / expert LoRA rank | 16 / 8 | | Weight precision | BF16, with small trajectory norms and dynamics parameters in FP32 | | Proposal constraints | top_k 40, min_p 0.05 | | Commit defaults | failure budget 0.2, target confidence 0.5 | | Weights | 32 Safetensors shards, 1323 tensors | The complete backbone, adapters, and memory parameters are stored in the native MLX parameter layout. Adapters remain unfused to preserve the reference BF16 computation. Loading validates model topology, tensor shape/dtype, and shard SHA256 hashes. `export_manifest.json` contains the weight index metadata and release provenance; no optimizer state or training data is shipped. This model uses its own rolling-canvas generation loop. Use the included `inference.py`, `modilify-mlx`, or Python API. Generic `mlx_lm.generate`, Transformers `AutoModel.from_pretrained`, and standard autoregressive serving backends do not implement this protocol. Image, audio, and video inputs are not supported by this release. ## Training, evaluation, and limitations The text backbone is adapted with dense and expert LoRA and trained trajectory memory. Adaptation uses response supervision over the valid rolling canvas, with terminal, confidence, and commit-readiness objectives. This release uses the saved step 1250 weights. The adaptation dataset is not distributed and its source composition is not documented in this release; refer to the upstream model card for base-model training information. Local comparisons establish export correctness for the tested cases, not task quality or safety. No capability benchmark score is claimed for this preview, and upstream benchmark scores do not describe this modified model. Long-context quality, broad language coverage, and deployment safety have not been established by those checks. ## Local comparison On 2026-10-06, the package installed in a fresh Python 3.12 environment without PyTorch and ran offline. Two requests matched the strict step 1250 source checkpoint at 1/2/5/10/20 denoising steps, including token IDs, stopping and canvas shifts. At five steps, interleaved B2 also matched independent B1 outputs and all three GDN2 state hashes. Short-request MLX peak allocations were about 50.79 GiB; longer workloads require additional memory. ## Intended use and impact statement This preview is intended for research on text diffusion, trajectory memory, local generation, and reproducible inference experiments. Generated statements can be incorrect, biased, or inappropriate; verify consequential outputs and keep qualified human review for sensitive decisions. The model is not intended to make autonomous high-risk decisions. Do not assume local execution alone provides safety, privacy, or resistance to prompt injection. Deployment should include use-specific testing, appropriate safeguards, and incident monitoring. ## License and attribution Modilify's original contributions and distribution are covered by the [Modilify Open Model License 1.0](LICENSE), a custom license. Upstream components retain their own terms; the DiffusionGemma base is published under Apache 2.0. See [NOTICE.md](NOTICE.md) for attribution and upstream notices. Maintained by **Modilify**. Report model and runtime issues through the [Hub discussions](https://huggingface.co/Modilify/Modilify-Mk2-preview-mlx/discussions).