LoopVL / README.md
ghjghjgjh's picture
LoopVL public release
bb796ec
|
Raw History Blame Contribute Delete
1.99 kB
metadata
pipeline_tag: image-text-to-text
tags:
  - loopvl
  - vision-language

LoopVL

Vision-language modeling with recurrent visual computation

GitHub · Code & evaluation  ·  Hugging Face  ·  ModelScope

This repository contains one model.safetensors and flat configuration, image-processor and tokenizer files. All inference and evaluation code is kept in the LoopVL GitHub repository. There are no Python files in this model snapshot.

Quick start

Use Linux, Python 3.12 and a matching CUDA-compatible PyTorch/torchvision pair. The validated GPU environment uses torch 2.12.1 and torchvision 0.27.1.

git clone https://github.com/Tier-Flow/LoopVL.git
cd LoopVL
python -m pip install -r requirements.txt

hf download TierFlow/LoopVL --local-dir model
# Alternative: ms-hub download Eternity123/LoopVL --local-dir model

python scripts/verify_repo.py --verify-model
python runtime/infer.py --image /path/to/image.png --prompt "What is in the image?" --budget 32 --device cuda:0

Keep both the GitHub code and all downloaded model files. LoopVL uses its custom GitHub loader.

The main checkpoint uses H2L3: L → L → L → H → L → L → L → H, with 16 layers per module call and 128 effective layer applications per forward pass. Preserve its mixed BF16 core weights and FP32 adapters; do not cast the entire model with .half() or .bfloat16().

For LoopVL-1B benchmark evaluation and behavior exploration, prefer direct answers. Chain-of-thought prompting is recommended only for mathematics tasks.