File size: 1,992 Bytes
bb796ec | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 | ---
pipeline_tag: image-text-to-text
tags:
- loopvl
- vision-language
---
<p align="center">
<img src="https://raw.githubusercontent.com/Tier-Flow/LoopVL/main/assets/loopvl-wordmark.svg" width="520" alt="LoopVL">
</p>
<p align="center"><strong>Vision-language modeling with recurrent visual computation</strong></p>
<p align="center">
<a href="https://github.com/Tier-Flow/LoopVL">GitHub · Code & evaluation</a> ·
<a href="https://huggingface.co/TierFlow/LoopVL">Hugging Face</a> ·
<a href="https://modelscope.cn/models/Eternity123/LoopVL">ModelScope</a>
</p>
This repository contains **one `model.safetensors` and flat configuration,
image-processor and tokenizer files**. All inference and evaluation code is
kept in the [LoopVL GitHub repository](https://github.com/Tier-Flow/LoopVL).
There are no Python files in this model snapshot.
## Quick start
Use Linux, Python 3.12 and a matching CUDA-compatible PyTorch/torchvision pair.
The validated GPU environment uses torch 2.12.1 and torchvision 0.27.1.
```bash
git clone https://github.com/Tier-Flow/LoopVL.git
cd LoopVL
python -m pip install -r requirements.txt
hf download TierFlow/LoopVL --local-dir model
# Alternative: ms-hub download Eternity123/LoopVL --local-dir model
python scripts/verify_repo.py --verify-model
python runtime/infer.py --image /path/to/image.png --prompt "What is in the image?" --budget 32 --device cuda:0
```
Keep **both the GitHub code and all downloaded model files**. LoopVL uses its
custom GitHub loader.
The main checkpoint uses H2L3: `L → L → L → H → L → L → L → H`, with 16 layers
per module call and **128 effective layer applications per forward pass**.
Preserve its mixed BF16 core weights and FP32 adapters; do not cast the entire
model with `.half()` or `.bfloat16()`.
> [!TIP]
> For LoopVL-1B benchmark evaluation and behavior exploration, prefer direct
> answers. Chain-of-thought prompting is recommended only for mathematics tasks.
|