Instructions to use qyoo/lingbot-worl with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use qyoo/lingbot-worl with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In adapter_config.json: "peft.base_model_name_or_path" must be a string
Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string
LingBot WoRL is the camera-conditioned policy checkpoint produced by the WoRL reinforcement-learning framework for LingBot-World-Fast. It keeps the base model's first-image, text, framewise camera-pose, and camera-intrinsics interface while improving the policy through long-horizon video RL.
The same policy is released in two forms: a compact PEFT LoRA adapter and a standalone BF16 transformer with the adapter already merged.
WoRL in one figure
WoRL brings together efficient on-policy video training, a history-aware Next Video Reward Model (NVRM), and a revisiting-centered environment. Candidate continuations are generated from an autoregressive rollout context, ranked by NVRM, and used by the policy update while the reference policy supplies KL regularization.
Release overview
| Field | Value |
|---|---|
| Policy checkpoint | WoRL-trained policy |
| Base model | robbyant/lingbot-world-fast |
| Conditioning | Text, first-frame image, camera-to-world trajectory, intrinsics |
| Evaluated geometry | 81 frames, 832 × 464, 16 FPS |
| Adapter | BF16 PEFT LoRA, rank / alpha 128 / 128 |
| Standalone model | Merged 18.5B-parameter BF16 WanModelFast transformer |
Repository layout
| Component | Path | Size |
|---|---|---|
| WoRL adapter | adapter_model.safetensors |
1.23 GB |
| PEFT configuration | adapter_config.json |
— |
| Adapter provenance | EXPORT_RECEIPT.json |
— |
| Merged transformer | merged/diffusion_pytorch_model-*.safetensors |
37.09 GB |
| Merged index and config | merged/diffusion_pytorch_model.safetensors.index.json, merged/config.json |
— |
| Merge / reload evidence | merged/MERGE_RECEIPT.json, merged/LOAD_VALIDATION.json |
— |
| Merged-file checksums | merged/SHA256SUMS |
— |
The text encoder, tokenizer, VAE, scheduler, and complete LingBot runtime are not duplicated here. They remain required from the compatible base release.
Load the policy
Option A · LoRA adapter
Use this form when the compatible LingBot-World-Fast base is already available.
import torch
from peft import PeftModel
from wan.modules.model_fast import WanModelFast
base = WanModelFast.from_pretrained(
"robbyant/lingbot-world-fast",
torch_dtype=torch.bfloat16,
control_type="cam",
)
model = PeftModel.from_pretrained(
base,
"qyoo/lingbot-worl",
is_trainable=False,
)
model.eval()
Option B · Merged transformer
Use this form to avoid adapter attachment at inference time.
import torch
from wan.modules.model_fast import WanModelFast
model = WanModelFast.from_pretrained(
"qyoo/lingbot-worl",
subfolder="merged",
torch_dtype=torch.bfloat16,
control_type="cam",
)
model.eval()
For full image-to-video generation, insert either transformer form into the compatible LingBot-World-Fast pipeline and supply its remaining base-model components.
Input contract
| Input | Format |
|---|---|
| First image | RGB frame 0; evaluated at 832 × 464 |
| Text prompt | UTF-8 scene and motion conditioning |
| Camera trajectory | Per-frame camera-to-world matrices, shape (F, 4, 4) |
| Intrinsics | Per-frame camera intrinsic matrices, shape (F, 3, 3) |
| Seed | Deterministic generation seed for each rollout |
The evaluated runtime generates 81-frame clips with an 80-frame stride at 16 FPS, preserving one shared boundary frame between adjacent clips.
Evaluation artifacts
The linked private datasets are self-contained WoRL rollout packages. Each includes generated videos, exact first images, prompts, camera controls, seeds, per-rollout scores, per-case scores, manifests, and SHA-256 integrity lists.
WBench v1 · 158 cases × 3 seeds
| Quality | Setting | Interaction | Consistency | Physical | Overall |
|---|---|---|---|---|---|
| 79.633591 | 94.017847 | 82.452479 | 86.955307 | 56.514116 | 79.914668 |
RevBench · 150 cases × 3 seeds
| NVRM | PI-Camera | Overall |
|---|---|---|
| 49.325778 | 47.936956 | 48.631367 |
These values are read directly from the two published result packages. The RevBench package contains WoRL outputs only.
Export integrity
- Adapter: 800 BF16 tensors across 400 LoRA layers, rank / alpha
128 / 128. - Adapter SHA-256:
a14c4275131a98b6bad24aa09ab54a2982c7fed4407bf084a57e582482840648. - Merged transformer: 18,544,332,864 parameters and 1,421 BF16 state tensors.
- Merge method:
peft.PeftModel.merge_and_unload(safe_merge=True). - Six sampled layers spanning blocks 0–39 matched independently computed base-plus-LoRA weights before and after serialization and reload.
- The standalone model reloaded with
WanModelFast.from_pretrained, moved to an NVIDIA H100, and contained no remaininglora_state keys.
Limitations and responsible use
Long autoregressive rollouts can still accumulate temporal artifacts, identity drift, geometric inconsistency, prompt mismatch, and camera-motion error. Generated video is synthetic and must not be treated as factual simulation or used as an autonomous-control signal. Users should review outputs, respect rights and consent for input media, and add safeguards appropriate to the downstream application.
License
This repository is released under Apache 2.0. The LingBot-World-Fast base model and runtime remain subject to their upstream terms and notices.
- Downloads last month
- 5
Model tree for qyoo/lingbot-worl
Base model
robbyant/lingbot-world-fast