Stop-Motion Consistency v1 LoRA

Built with Qwen. An experimental next-frame image-editing adapter fine-tuned from the original BF16 Qwen-Image-2.1. The adapter was trained on 112 curated stop-motion transitions using an RTX PRO 6000.

Quality is limited. Matched tests show improved static-area retention, but motion magnitude, direction and shape consistency remain unreliable. This is a small research pilot, not a proven solution for long, consistent animation.

A Little Wave, generated with the adapter

Full-quality 8.125-second video · Training and quality report · Dataset

Download and compatibility

The adapter is stopmotion-consistency-v1.safetensors, 67,141,056 bytes. It contains only attention LoRA updates, not the base model.

SHA-256: d68479ed2c246bc7ebea8a62019b10da903772d30854385302b2d66de99df51a.

Use DiffSynth-Studio at commit d2d684ad1f912949eae08453b9411ae40c5ec0ab. The file uses raw PEFT-style keys with .lora_A.default.weight and .lora_B.default.weight, tested with DiffSynth's Qwen-Image-2.1 loader. Automatic compatibility with Diffusers, ComfyUI or other LoRA loaders is not established.

The required base model is Qwen/Qwen-Image-2.1, revision b3179ad355be050328e483a9dfdd9e60cd62adfa. This is not an adapter for older Qwen-Image releases.

from huggingface_hub import hf_hub_download

adapter = hf_hub_download(
    "ProCreations/stopmotion-consistency-v1-lora",
    "stopmotion-consistency-v1.safetensors",
)

With the pinned DiffSynth checkout on PYTHONPATH, its dependencies installed, and the original BF16 base downloaded, run the supplied inference script:

python source/inference_example.py \
  --base /path/to/Qwen-Image-2.1-BF16 \
  --adapter stopmotion-consistency-v1.safetensors \
  --anchor anchor.png \
  --previous previous.png \
  --prompt "Create the next stop-motion frame. Move the red clay ball a small step to the right. Preserve the character, set, camera and lighting." \
  --seed 1234 \
  --output next-frame.png

The loader should report 128 tensors fused. The tested inference settings are strength 1.0, 40 diffusion steps and CUDA BF16. The exact observed environment is recorded in requirements-observed.txt; it includes PyTorch 2.14.0 with CUDA 13.0, Transformers 5.17.0 and PEFT 0.21.0. Enough GPU memory for the original BF16 model components is required.

Use the original scene anchor first and the immediately preceding frame second. Describe a small, concrete change and give each frame a different seed. Old anchors can restore objects that have already left the scene; this was observed in testing.

Training

Setting Value
Training / validation / test examples 112 / 10 / 7
Distinct source groups per split 65 / 4 / 4
Rank / alpha 16 / 16
Target modules to_q, to_k, to_v, to_out.0
Trainable parameters 16,777,216
Epochs / optimizer steps 8 / 224
Batch / accumulation 1 / 4
Learning rate 0.00005, warmup and cosine decay
Precision Frozen BF16 base, FP32 adapter and optimizer
Maximum training pixels 262,144, no crop or upscale
Selected checkpoint Epoch 8, lowest validation flow MSE

The curated dataset is pinned at revision 0a5c0c714892f39e64fb1855233acb8aa8a9a04e. Every cached example passed positive reference-image-token alignment checks. Targets were supervision only and never conditioning images. Source groups were separated across splits. The dataset has no uninterrupted accepted rollout of eight or more steps.

Training and generation code is in source/. To reproduce training, extract the dataset's curated archive, download the pinned base, and set STOPMOTION_TRAIN_ROOT, STOPMOTION_DATA_PATH, and STOPMOTION_BASE_PATH to local directories. Then run source/cache_inputs.py, followed by source/train.py --epochs 8 --no-checkpointing. The training script starts fresh; it saves optimizer/RNG state but does not implement automatic resume. The selected public adapter is sufficient for inference; intermediate checkpoints and optimizer state are not part of this release.

Evaluation and limitations

Validation flow MSE fell from 0.109094 to 0.091397, a 16.2% reduction. On seven held-out edits using the same prompts, references, seeds and sampling settings:

Diagnostic Base Adapter
Target image MSE 0.008615 0.007695
Static-region absolute error 0.022227 0.015650
Moving-region target absolute error 0.191547 0.187282

These small-sample diagnostics can reward conservative edits or copying and are not general animation-consistency scores. Direct agent inspection found wrong motion magnitude/direction and restoration of obsolete anchor objects. There was no independent human benchmark.

The demo has one base-generated opening frame and 48 adapter-generated edits, each conditioned on the anchor and previous generated frame. It uses 960 × 544 images, 40 diffusion steps and 8 fps, plus opening/closing holds. The initial take duplicated a ball in two frames, so frame 20 and the entire dependent suffix were regenerated with fresh seeds and a single-ball check. The final clip still has geometry jitter and non-monotonic ball movement. No interpolation, compositing or image warping creates the motion. All final frame/reference hashes and image-token alignments were verified.

The frame plan, chain verification, visual review, full final image sequence, per-frame provenance and matched test comparisons are included. Historical workstation paths in run records are represented by ${PROJECT_ROOT}, ${DATASET_ROOT} and ${BASE_MODEL_ROOT}.

License and attribution

The base model is covered by the Qwen Research License, included in LICENSE-QWEN.txt. Its grant is for research/evaluation use; commercial use requires a separate license from the licensor. Preserve the agreement and NOTICE.txt with the adapter. This is a modified derivative for this project, not an official Qwen release.

Dataset media retain their source-specific CC BY licenses. See ATTRIBUTION.json and DATASET-LICENSES.md for creators, source URLs, license versions and transformations. The dataset license does not replace the base model license.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ProCreations/stopmotion-consistency-v1-lora

Adapter
(72)
this model

Dataset used to train ProCreations/stopmotion-consistency-v1-lora

Space using ProCreations/stopmotion-consistency-v1-lora 1