Text Generation
Transformers
Safetensors
modilify_mk2
diffusion
mixture-of-experts
trust-remote-code
conversational
custom_code
Instructions to use modilify/Modilify-Mk2-preview with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use modilify/Modilify-Mk2-preview with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="modilify/Modilify-Mk2-preview", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("modilify/Modilify-Mk2-preview", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use modilify/Modilify-Mk2-preview with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "modilify/Modilify-Mk2-preview" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "modilify/Modilify-Mk2-preview", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/modilify/Modilify-Mk2-preview
- SGLang
How to use modilify/Modilify-Mk2-preview with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "modilify/Modilify-Mk2-preview" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "modilify/Modilify-Mk2-preview", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "modilify/Modilify-Mk2-preview" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "modilify/Modilify-Mk2-preview", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use modilify/Modilify-Mk2-preview with Docker Model Runner:
docker model run hf.co/modilify/Modilify-Mk2-preview
File size: 2,122 Bytes
53d5244 b88f761 53d5244 b88f761 53d5244 b88f761 53d5244 b88f761 53d5244 b88f761 53d5244 b88f761 53d5244 b88f761 53d5244 b88f761 53d5244 b88f761 53d5244 b88f761 53d5244 b88f761 53d5244 b88f761 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 | ---
license: other
license_name: modilify-open-model-license-1.0
license_link: LICENSE
library_name: transformers
pipeline_tag: text-generation
tags:
- diffusion
- mixture-of-experts
- trust-remote-code
- safetensors
---

# Modilify Mk2 Preview
PyTorch text inference for the **step 1250, schema25** checkpoint migrated
from `Modilify-Mk2-preview-mlx`. The model uses a shared DiffusionGemma text
encoder/decoder, a rolling 256-token canvas, and dual-timescale GDN2 memory.
The release contains 32 Safetensors shards totaling **48.23 GiB**.
Dense and expert LoRA adapters remain unfused. Model loading preserves the
26 FP32 GDN2/norm parameters alongside BF16 weights. No vision tower is included.
## Inference
Use Transformers **5.14.1**, PyTorch, Accelerate, and Safetensors. Run from
this directory or replace `model_path` with its location. Inference needs
additional memory beyond the weights.
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_path = "."
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_path,
trust_remote_code=True,
dtype=torch.bfloat16,
device_map="auto",
attn_implementation="sdpa",
).eval()
inputs = tokenizer.apply_chat_template(
[{"role": "user", "content": "Explain why the sky is blue."}],
tokenize=True,
add_generation_prompt=True,
enable_thinking=False,
return_dict=True,
return_tensors="pt",
).to(model.device)
output = model.generate(**inputs, max_new_tokens=128, seed=42)
print(tokenizer.decode(output.sequences[0, inputs.input_ids.shape[1]:],
skip_special_tokens=True))
```
For MPS, use `device_map={"": "mps"}`. Set `enable_thinking=True` in the chat
template to request thinking. Generation uses temperature 0.8, top-k 40,
min-p 0.05, target confidence 0.5, and failure budget 0.2. `max_denoising_steps`
can bound generation. Static batches and continuous batching share the
confidence-prefix commit policy; each request retains its own cache and RNG.
|