Instructions to use modilify/Modilify-Mk2-preview with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use modilify/Modilify-Mk2-preview with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="modilify/Modilify-Mk2-preview", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("modilify/Modilify-Mk2-preview", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use modilify/Modilify-Mk2-preview with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "modilify/Modilify-Mk2-preview" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "modilify/Modilify-Mk2-preview", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/modilify/Modilify-Mk2-preview
- SGLang
How to use modilify/Modilify-Mk2-preview with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "modilify/Modilify-Mk2-preview" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "modilify/Modilify-Mk2-preview", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "modilify/Modilify-Mk2-preview" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "modilify/Modilify-Mk2-preview", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use modilify/Modilify-Mk2-preview with Docker Model Runner:
docker model run hf.co/modilify/Modilify-Mk2-preview
Modilify Mk2 Preview
PyTorch text inference for the step 1250, schema25 checkpoint migrated
from Modilify-Mk2-preview-mlx. The model uses a shared DiffusionGemma text
encoder/decoder, a rolling 256-token canvas, and dual-timescale GDN2 memory.
The release contains 32 Safetensors shards totaling 48.23 GiB.
Dense and expert LoRA adapters remain unfused. Model loading preserves the 26 FP32 GDN2/norm parameters alongside BF16 weights. No vision tower is included.
Inference
Use Transformers 5.14.1, PyTorch, Accelerate, and Safetensors. Run from
this directory or replace model_path with its location. Inference needs
additional memory beyond the weights.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_path = "."
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_path,
trust_remote_code=True,
dtype=torch.bfloat16,
device_map="auto",
attn_implementation="sdpa",
).eval()
inputs = tokenizer.apply_chat_template(
[{"role": "user", "content": "Explain why the sky is blue."}],
tokenize=True,
add_generation_prompt=True,
enable_thinking=False,
return_dict=True,
return_tensors="pt",
).to(model.device)
output = model.generate(**inputs, max_new_tokens=128, seed=42)
print(tokenizer.decode(output.sequences[0, inputs.input_ids.shape[1]:],
skip_special_tokens=True))
For MPS, use device_map={"": "mps"}. Set enable_thinking=True in the chat
template to request thinking. Generation uses temperature 0.8, top-k 40,
min-p 0.05, target confidence 0.5, and failure budget 0.2. max_denoising_steps
can bound generation. Static batches and continuous batching share the
confidence-prefix commit policy; each request retains its own cache and RNG.
- Downloads last month
- 207
