Instructions to use MirroS-Lab/Code-as-World-VL-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MirroS-Lab/Code-as-World-VL-4B with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("MirroS-Lab/Code-as-World-VL-4B") model = AutoModelForMultimodalLM.from_pretrained("MirroS-Lab/Code-as-World-VL-4B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 1,977 Bytes
2a7b99f e40376e 2a7b99f 104afb2 2a7b99f e40376e 2a7b99f 104afb2 e40376e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 | ---
base_model:
- Qwen/Qwen3.5-4B
license: apache-2.0
pipeline_tag: video-text-to-text
library_name: transformers
---
# Code-as-World-VL-4B
Code-as-World-VL-4B (https://arxiv.org/abs/2608.27549) is a vision-language model fine-tuned for physical
understanding and quantitative reasoning over videos.
Project page: [https://mirros-lab.github.io/code-as-world](https://mirros-lab.github.io/code-as-world)
GitHub: [https://github.com/mirros-lab/code-as-world](https://github.com/mirros-lab/code-as-world)
## Model details
- **Base model:** [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B)
- **Weight format:** BF16 Safetensors checkpoint
- **Recommended video input:** 16 frames
## Usage
The checkpoint can be served with vLLM:
```bash
pip install "vllm==0.19.1" "transformers==5.11.0" qwen-vl-utils
vllm serve MirroS-Lab/Code-as-World-VL-4B \
--served-model-name code-as-world-4b \
--max-model-len 4608 \
--gpu-memory-utilization 0.90 \
--media-io-kwargs '{"video":{"num_frames":16,"fps":-1,"video_backend":"openpangu"}}' \
--mm-processor-kwargs '{"do_sample_frames":false}' \
--mm-processor-cache-gb 0 \
--generation-config vllm
```
The server exposes an OpenAI-compatible API at `/v1`.
## Intended use
This model is intended for research on physical understanding, measurement, and
quantitative reasoning from images and videos. Model outputs may be inaccurate and
should be independently verified before use in safety-critical settings.
## License
This checkpoint is released under the Apache License 2.0. It is derived from
[Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B); users must also comply
with the terms applicable to the base model and their input data.
## Citation
```bibtex
@article{mirros2026codeasworld,
title = {Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning},
author = {{MirroS Team}},
journal = {arXiv preprint arXiv:2608.27549},
year = {2026}
}
``` |