Instructions to use ahmedheakl/rand-mobile with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ahmedheakl/rand-mobile with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ahmedheakl/rand-mobile", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Mobile-O β soup3-targets
A weight-space model soup of three Mobile-O v2 checkpoints, each of which already met a different benchmark target. It clears three of the four tracked benchmarks at a single checkpoint and a single guidance setting, which no individual checkpoint in this project does.
Built with no additional training β it is a uniform 1/3 average of three existing checkpoints that share one SFT init.
Runnable code: https://github.com/ahmedheakl/mobileov2-minicpm
Benchmarks
Recommended setting: cfg 1.5, 12 DPM-Solver++ steps.
| benchmark | value | target | status |
|---|---|---|---|
| GenEval | 0.90242 | β₯ 0.90 | met |
| ImageReward (MJHQ-30K) | 0.9571 | β₯ 0.90 | met |
| GEdit (EN, local Qwen2.5-VL-72B judge) | 6.740 | β₯ 6.7 | met |
| DPG-Bench | 82.190 | β₯ 85 | not met |
| FID (MJHQ-30K) | 14.36 | β€ 8 | not met |
| ImgEdit (local judge) | not measured | β₯ 3.5 | β |
Three of six met is the most of any checkpoint in this project β every other configuration
measured tops out at two, including the previous best (v2mcptf-grpo-500 @ cfg 3.0: GenEval 0.9284,
DPG 81.73, FID 15.40, ImageReward 0.9250, ImgEdit 3.020, GEdit 6.620).
At cfg 3.0, 12 steps the same checkpoint scores GenEval 0.91876, ImageReward 1.0831, GEdit 6.750, DPG 83.228 β wider margins on the three met targets, but it costs human-image quality measurably (see Caveats), so cfg 1.5 is the recommended setting.
FID is the other clear miss. It is not a regression specific to this soup β RL-tuned checkpoints in this project sit in the 13β18 range while the FID-optimised ones (5.9β7.3) give up 0.2+ GenEval. ImgEdit was never run on this checkpoint; its members score 3.0β3.1, below the 3.5 target.
DPG-Bench was not reached by anything in this project. The best DPG measured across 121 settings is 84.612 (a different checkpoint at cfg 7.5 + interval guidance, whose GenEval then falls to 0.880), and the guidance response peaks and declines rather than continuing to climb.
What is in this repo
Only the trained head: the SANA DiT and the diffusion connector, 602 tensors
(model.dit.* 548, model.diffusion_connector.* 54). The vision-language model is frozen during
training and is not included here.
To run this you also need:
openbmb/MiniCPM-V-4_6β the frozen VLM text/image encoderEfficient-Large-Model/Sana_600M_512px_diffusersβ for the DC-AE (f32, 32-channel, 16Γ16 latent at 512px) and the scheduler config
Connector type is mcptf and it fuses 1 VLM layer. Note that a vlm_num_layers field elsewhere
in this project's configs can read 4; the weights are authoritative β derive it from
fusion.layer_weights.shape[0].
1024x1024
This repo also carries upsampler_1024/head_ema.safetensors (426 MB, 106M parameters): a trained
decoder head that produces 1024px output from the same 512px latent. The DiT is not involved and
is not modified -- it keeps its 16x16 latent and runs the same steps. The frozen DC-AE decoder emits
its last hidden features (128 channels at 512px, the tensor that normally feeds conv_out) and the
head maps those to 1024px RGB, anchored on a bicubic x2 of the ordinary 512px decode. The VAE stays
frozen and must stay bf16 (fp16 underflows the DC-AE decoder).
Cost: +58 ms, +7.1% end-to-end (817 -> 874 ms for one image at batch 1, 20 steps, median of 30 runs after 5 warm-up, one RTX PRO 6000; the head replaces the 29 ms decode with 86 ms). Rendering the same prompt and seed at both sizes and downscaling the 1024 result back to 512 differs from the native 512px output by a mean of 1.0/255 -- the same image, decoded better, not more diffusion detail. On nine benchmarks it is content-neutral (GenEval, DPG, ImgEdit unchanged) and improves FID by 0.48.
Selected over the alternative (a SwinIR latent upsampler applied 16x16 -> 32x32 before the frozen decode) on a reconstruction eval: FID 1.76 vs 2.077, OCR 47.9 vs 40.
git clone https://github.com/ahmedheakl/mobileov2-minicpm && cd mobileov2-minicpm
pip install -r requirements.txt
python infer.py --size 1024 --prompt "a woman holding a ceramic mug" --out out/big.png
Inference settings used for every number above
| scheduler | DPM-Solver++, solver_order=2, flow_shift=3 |
| steps | 12 |
| guidance | cfg 1.5 |
| null condition | the empty prompt through the same VLM+connector path (not a zero vector) |
| resolution | 512Γ512 |
For editing, the null condition is the instruction rather than the empty prompt.
12 steps is deliberate, not a shortcut: on this model sample quality for human subjects peaks around 8β16 steps and declines by 20β40, so 12 steps is both better and ~40% cheaper than the 20-step default. Alignment (GenEval/DPG) and GEdit are unchanged at 12 vs 20 steps.
Provenance
Uniform 1/3 soup of:
| member | what it contributed |
|---|---|
ta-dpg-s1.0 |
GenEval 0.9293 |
qual-imagereward @ step 500 |
ImageReward 0.9165 |
joint-grpo-refl-500 @ step 500 |
GEdit 6.710 |
All three are fine-tuned from the same init (Mobile-O-0.5B-SFT-minicpm-v2mcptf-mixed), which is
what makes averaging them valid. The soup beats all three members on DPG, ImageReward and GEdit
simultaneously.
Caveats
- DPG-Bench is not met (82.190 vs 85).
- It does not improve human-image quality. On a paired human-quality benchmark it is neutral versus the baseline (+0.018, not significant). At cfg 3.0 it is significantly worse for people (β0.147, t = β2.68), which is why cfg 1.5 is recommended despite cfg 3.0's better target numbers.
- GEdit here is scored by a local Qwen2.5-VL-72B judge, not the GPT-4o leaderboard scale; the two are not comparable.
- The head alone is not a runnable model β see What is in this repo.
- Downloads last month
- 42