YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Wan 2.2 Serverless on RunPod β Preset LoRA Edition
Production-ready scaffold for serving Wan2.2-T2V-A14B on RunPod Serverless with a dropdown of hot-swappable preset LoRAs, driven from a Next.js frontend.
[ Next.js UI ] --/api/generate (key stays server-side)--> [ RunPod /run + /status ]
|
[ Network Volume /runpod-volume ]
βββ models/wan2.2 (base model)
βββ loras/ (preset LoRAs)
|
[ Worker: handler.py ]
WanPipeline + dual-expert LoRA swap
-> base64 MP4 in JSON response
Preset LoRA registry
| Dropdown id | Source | Notes |
|---|---|---|
none |
β | Base model, 40 steps |
lightning_4step |
lightx2v/Wan2.2-Distill-Loras (Apache-2.0) | HIGH/LOW pair, auto-switches to 4 steps, CFG 1.0 β ~8x faster |
detail_boost |
alibaba-pai/Wan2.2-Fun-Reward-LoRAs (Apache-2.0) | HPS human-preference reward LoRA, realism/detail |
vintage_35mm |
bring your own | Drop files on the volume (names in handler.py) |
cinematic_lighting |
bring your own | Single-file LoRA β applied to both denoisers |
dynamic_motion |
bring your own | Single-file LoRA β applied to both denoisers |
Wan2.2 A14B is a two-expert (MoE) model:
handler.pyloads every preset into both the high-noise and low-noise transformers (load_into_transformer_2=True). Style LoRAs trained for Wan2.2 usually ship as HIGH/LOW pairs; Wan2.1-style single files also work (loaded into both).
Repo layout
handler.py RunPod serverless worker
Dockerfile Worker image (weights NOT baked in)
requirements.txt
test_input.json Sample payload for the RunPod console test tab
scripts/download_models.sh One-time network-volume downloader
nextjs/
app/api/generate/route.ts Server-side proxy (submit + status poll)
lib/runpod.ts Client polling helper
components/WanGenerator.tsx Dropdown UI (presets + duration)
.env.example
Deploy (β10 minutes)
1. Network volume + models
- RunPod Console β Storage β Create Network Volume: 100GB, pick the
datacenter you'll run workers in (e.g.
US-KS-2,EU-RO-1). - Launch a temporary pod (any cheap GPU) with the volume mounted at
/runpod-volume, open a terminal, and run:
(~70GB download; go get coffee.)curl -sL https://huggingface.co/SealesEmpire/wan22-runpod-serverless/resolve/main/scripts/download_models.sh | bash - Terminate the pod. Weights persist on the volume.
2. Build & push the worker image
git clone https://huggingface.co/SealesEmpire/wan22-runpod-serverless
cd wan22-runpod-serverless
docker build -t YOUR_DOCKERHUB/wan22-serverless:latest .
docker push YOUR_DOCKERHUB/wan22-serverless:latest
3. Create the serverless endpoint
RunPod Console β Serverless β New Endpoint:
- Container image:
YOUR_DOCKERHUB/wan22-serverless:latest - GPU: 48GB minimum (L40S / A6000). A100/H100 80GB = fastest.
(bf16 A14B experts are ~28GB each β 24GB cards cannot fit them; use an
FP8-scaled checkpoint via
WAN_MODEL_PATHif you must.) - Network volume: attach yours (mounts at
/runpod-volume) - Execution timeout: raise to
1800s (default 600s can kill 10s renders) - Min workers:
0for $0 idle, or1during demos (cold start = minutes even with the volume β weights must stream from network storage into VRAM) - FlashBoot: on
4. Smoke-test
Paste test_input.json into the endpoint's Test tab, or:
curl -X POST https://api.runpod.ai/v2/$ENDPOINT_ID/run \
-H "Authorization: Bearer $RUNPOD_API_KEY" \
-H "Content-Type: application/json" \
-d @test_input.json
# then poll: GET /v2/$ENDPOINT_ID/status/<id>
5. Frontend
npx create-next-app@latest wan-ui --typescript --tailwind --app
cd wan-ui && npm i lucide-react
cp -r ../wan22-runpod-serverless/nextjs/* . # route.ts, lib/, components/
cp .env.example .env.local # fill in key + endpoint id
npm run dev # or: vercel deploy
Drop <WanGenerator /> into any page.
Gotchas baked into this scaffold
- No
.to("cuda")+enable_model_cpu_offload()conflict β offload only. - VAE is
AutoencoderKLWanin float32 (bf16 VAE decode = artifacts). - Async
/run+ polling, not/runsync(90s default / 300s max sync wait; 20MB response cap β 5s 720p clips fit, 8β10s may not). - API key is server-side only (
route.ts), neverNEXT_PUBLIC_. num_frames= 81/129/161 for 5/8/10s @16fps (Wan requires 4k+1).- Model-card guidance defaults: 4.0 (high-noise) / 3.0 (low-noise).
Scaling past the 20MB response cap
For 8β10s or 1080p clips, replace the base64 return in handler.py with an
S3/Cloudflare R2 upload (boto3 + presigned URL) and return the URL instead.
RunPod async results are retained 30 minutes, so clients must poll promptly.
Licenses
Wan2.2 and the two wired-in LoRA presets are Apache-2.0. Verify licenses for any style LoRAs you add yourself.