Configuration Parsing Warning:In adapter_config.json: "peft.base_model_name_or_path" must be a string

Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

WoRL long-horizon generations and RL system efficiency

LingBot WoRL

A policy checkpoint from the All-in-One RL Framework of Video World Model

WoRL · Base model · WBench rollouts · RevBench rollouts

LingBot WoRL is the camera-conditioned policy checkpoint produced by the WoRL reinforcement-learning framework for LingBot-World-Fast. It keeps the base model's first-image, text, framewise camera-pose, and camera-intrinsics interface while improving the policy through long-horizon video RL.

The same policy is released in two forms: a compact PEFT LoRA adapter and a standalone BF16 transformer with the adapter already merged.

WoRL in one figure

WoRL reinforcement-learning framework

WoRL brings together efficient on-policy video training, a history-aware Next Video Reward Model (NVRM), and a revisiting-centered environment. Candidate continuations are generated from an autoregressive rollout context, ranked by NVRM, and used by the policy update while the reference policy supplies KL regularization.

Release overview

Field Value
Policy checkpoint WoRL-trained policy
Base model robbyant/lingbot-world-fast
Conditioning Text, first-frame image, camera-to-world trajectory, intrinsics
Evaluated geometry 81 frames, 832 × 464, 16 FPS
Adapter BF16 PEFT LoRA, rank / alpha 128 / 128
Standalone model Merged 18.5B-parameter BF16 WanModelFast transformer

Repository layout

Component Path Size
WoRL adapter adapter_model.safetensors 1.23 GB
PEFT configuration adapter_config.json —
Adapter provenance EXPORT_RECEIPT.json —
Merged transformer merged/diffusion_pytorch_model-*.safetensors 37.09 GB
Merged index and config merged/diffusion_pytorch_model.safetensors.index.json, merged/config.json —
Merge / reload evidence merged/MERGE_RECEIPT.json, merged/LOAD_VALIDATION.json —
Merged-file checksums merged/SHA256SUMS —

The text encoder, tokenizer, VAE, scheduler, and complete LingBot runtime are not duplicated here. They remain required from the compatible base release.

Load the policy

Option A · LoRA adapter

Use this form when the compatible LingBot-World-Fast base is already available.

import torch
from peft import PeftModel
from wan.modules.model_fast import WanModelFast

base = WanModelFast.from_pretrained(
    "robbyant/lingbot-world-fast",
    torch_dtype=torch.bfloat16,
    control_type="cam",
)
model = PeftModel.from_pretrained(
    base,
    "qyoo/lingbot-worl",
    is_trainable=False,
)
model.eval()

Option B · Merged transformer

Use this form to avoid adapter attachment at inference time.

import torch
from wan.modules.model_fast import WanModelFast

model = WanModelFast.from_pretrained(
    "qyoo/lingbot-worl",
    subfolder="merged",
    torch_dtype=torch.bfloat16,
    control_type="cam",
)
model.eval()

For full image-to-video generation, insert either transformer form into the compatible LingBot-World-Fast pipeline and supply its remaining base-model components.

Input contract

Input Format
First image RGB frame 0; evaluated at 832 × 464
Text prompt UTF-8 scene and motion conditioning
Camera trajectory Per-frame camera-to-world matrices, shape (F, 4, 4)
Intrinsics Per-frame camera intrinsic matrices, shape (F, 3, 3)
Seed Deterministic generation seed for each rollout

The evaluated runtime generates 81-frame clips with an 80-frame stride at 16 FPS, preserving one shared boundary frame between adjacent clips.

Evaluation artifacts

The linked private datasets are self-contained WoRL rollout packages. Each includes generated videos, exact first images, prompts, camera controls, seeds, per-rollout scores, per-case scores, manifests, and SHA-256 integrity lists.

WBench v1 · 158 cases × 3 seeds

Quality Setting Interaction Consistency Physical Overall
79.633591 94.017847 82.452479 86.955307 56.514116 79.914668

RevBench · 150 cases × 3 seeds

NVRM PI-Camera Overall
49.325778 47.936956 48.631367

These values are read directly from the two published result packages. The RevBench package contains WoRL outputs only.

Export integrity

  • Adapter: 800 BF16 tensors across 400 LoRA layers, rank / alpha 128 / 128.
  • Adapter SHA-256: a14c4275131a98b6bad24aa09ab54a2982c7fed4407bf084a57e582482840648.
  • Merged transformer: 18,544,332,864 parameters and 1,421 BF16 state tensors.
  • Merge method: peft.PeftModel.merge_and_unload(safe_merge=True).
  • Six sampled layers spanning blocks 0–39 matched independently computed base-plus-LoRA weights before and after serialization and reload.
  • The standalone model reloaded with WanModelFast.from_pretrained, moved to an NVIDIA H100, and contained no remaining lora_ state keys.

Limitations and responsible use

Long autoregressive rollouts can still accumulate temporal artifacts, identity drift, geometric inconsistency, prompt mismatch, and camera-motion error. Generated video is synthetic and must not be treated as factual simulation or used as an autonomous-control signal. Users should review outputs, respect rights and consent for input media, and add safeguards appropriate to the downstream application.

License

This repository is released under Apache 2.0. The LingBot-World-Fast base model and runtime remain subject to their upstream terms and notices.

Downloads last month
5
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for qyoo/lingbot-worl

Adapter
(1)
this model