--- license: other license_name: modilify-open-model-license-1.0 license_link: LICENSE library_name: transformers pipeline_tag: text-generation tags: - diffusion - mixture-of-experts - trust-remote-code - safetensors --- ![LOGO](assets/01-LOGO.jpg) # Modilify Mk2 Preview PyTorch text inference for the **step 1250, schema25** checkpoint migrated from `Modilify-Mk2-preview-mlx`. The model uses a shared DiffusionGemma text encoder/decoder, a rolling 256-token canvas, and dual-timescale GDN2 memory. The release contains 32 Safetensors shards totaling **48.23 GiB**. Dense and expert LoRA adapters remain unfused. Model loading preserves the 26 FP32 GDN2/norm parameters alongside BF16 weights. No vision tower is included. ## Inference Use Transformers **5.14.1**, PyTorch, Accelerate, and Safetensors. Run from this directory or replace `model_path` with its location. Inference needs additional memory beyond the weights. ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer model_path = "." tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained( model_path, trust_remote_code=True, dtype=torch.bfloat16, device_map="auto", attn_implementation="sdpa", ).eval() inputs = tokenizer.apply_chat_template( [{"role": "user", "content": "Explain why the sky is blue."}], tokenize=True, add_generation_prompt=True, enable_thinking=False, return_dict=True, return_tensors="pt", ).to(model.device) output = model.generate(**inputs, max_new_tokens=128, seed=42) print(tokenizer.decode(output.sequences[0, inputs.input_ids.shape[1]:], skip_special_tokens=True)) ``` For MPS, use `device_map={"": "mps"}`. Set `enable_thinking=True` in the chat template to request thinking. Generation uses temperature 0.8, top-k 40, min-p 0.05, target confidence 0.5, and failure budget 0.2. `max_denoising_steps` can bound generation. Static batches and continuous batching share the confidence-prefix commit policy; each request retains its own cache and RNG.