Moxie Multimedia Suite

Timeline-driven video generation for ComfyUI β€” arrange tasks, references, audio, and subtitles on a multi-track timeline and render them with the bundled Moxie model set.

πŸ“¦ Installation

Make sure FFmpeg is installed and available in your system PATH before installing this pack.

Install through ComfyUI Manager, or clone this repository into your ComfyUI custom_nodes folder:

cd Your_ComfyUI_Path/custom_nodes
git clone https://huggingface.co/turtle89431/Moxie-Multimedia

Installing dependencies

ComfyUI Manager installs requirements.txt automatically. If you cloned the repository manually, install the dependencies yourself:

ComfyUI Portable (Windows) β€” run from the portable root folder (the one containing run_nvidia_gpu.bat):

cd Your_ComfyUI_windows_portable
.\python_embeded\python.exe -m pip install -r .\ComfyUI\custom_nodes\Moxie-Multimedia\requirements.txt

Regular ComfyUI (venv / system Python):

pip install -r Your_ComfyUI_Path/custom_nodes/Moxie-Multimedia/requirements.txt

Restart ComfyUI afterwards.

The pack targets NVIDIA RTX 3060+ GPUs only, so every feature dependency is mandatory and installed by the command above β€” including speech generation (voxcpm), subtitle recognition (openai-whisper, qwen-asr), and the RTX Video Super Resolution upscale (nvidia-vfx). Expect a sizable first install: qwen-asr pins transformers==4.57.6 (pip may upgrade it inside your ComfyUI environment) and voxcpm needs torch>=2.5.0, so keep ComfyUI reasonably current. nvidia-vfx is pulled from NVIDIA's own package index (pypi.nvidia.com, wired up inside requirements.txt) because the PyPI copy is a stub that fails to build on some setups β€” if the command is interrupted, re-run it and it resumes where it stopped.

Models

All model files ship inside the repo β€” cloning it into custom_nodes is all that's needed. The loader finds everything automatically in the pack's models/ folders:

File Location
Diffusion model (Moxie-Multimedia.safetensors) models/diffusion/
Text encoder (MM-VL.safetensors) models/clip/
Video VAE (bundled file) models/vae/
Audio VAE (bundled file) models/audio_vae/
Turbo LoRA (MM3step.safetensors, applied by default) models/loras/
Preview tiny VAE (MM-preview.safetensors) models/vae_approx/

All model files live in the Hugging Face version of this repository β€” cloning it downloads the complete set, and there is no separate downloader step. If you obtained the pack from somewhere else and the models/ folders are missing or incomplete, re-clone from Hugging Face to get the bundled files.

Restart ComfyUI afterwards.

πŸ”— Pipeline

[Moxie Multimedia Loader] ── model ──► [Moxie Preview Override] (live preview)
        β”‚  model, clip,                 β”‚
        β”‚  video_vae, audio_vae         β–Ό model
        └──────────────► [MultiTrack Editor] ── (TRACKS_INFO) ──► [Moxie MultiTrack Project] (steps, seed, project name)
                                                                        β”‚ (PROJECT_NAME)
                                                                        β–Ό
                                                          [Moxie Project Video Combine]
                                                                        β”‚ (images, audio, fps)
                                                                        β–Ό
                                                                 [Save Video]

🧩 Nodes

Core Pipeline

Node Description
Moxie Multimedia Loader One-stop loader for the bundled model set. All required files are resolved automatically from the pack's models/ folders β€” add the node and it outputs model, clip, video_vae, and audio_vae, ready for the project pipeline.
Moxie Preview Override Streams a live animated preview β€” motion and sound β€” of the generation directly onto the node canvas while sampling runs, so you can watch progress without waiting for the final export.
MultiTrack Editor The creative hub of the suite. Arrange task, video, audio, and subtitle tracks on a timeline, set the output dimensions and frame rate, write prompts per segment, and attach reference images, video, or audio. Outputs TRACKS_INFO for the project node (plus media outputs when slot references are used).
Moxie MultiTrack Project Renders the timeline. Takes the loader outputs plus TRACKS_INFO, expands the tasks in sequence automatically (no manual segment loop), and handles continuity between segments. Exposes only the creative controls: steps, seed, and project_name. Each segment is saved as it completes.
Moxie Project Video Combine Assembles the final video. Previews the saved segments, lets you pick a version for each, and combines them β€” automatically in one workflow run, or manually after review. Outputs images, audio, and fps for Save Video.

Timeline & MultiTrack Tools

Node Description
Multi Images Loader Provides an ordered list of up to 25 images with shared resolution choices. Add images through the media selector or drop image files onto the grid; each image is resized and output in grid order.
MultiTrack Info Output Outputs the video dimensions, total frame count, frame rate, and task count from the editor's TRACKS_INFO.
MultiTrack Task Output Outputs the user and system prompts for each task, with media loaded on demand. Set the task index for one segment at a time, or use -1 to retrieve the entire timeline's media at once.
MultiTrack Audio Output Outputs the audio for the timeline's tasks.
MultiTrack Prompt Enhancer Expands task prompts with a connected language model, using each segment's user and system prompts plus its reference media.
MultiTrack Prompt Enhance To Project Sends enhanced prompts straight into the project pipeline.
MultiTrack Prompt Enhance To Project Apply Applies the enhanced prompt results to a running project.
MultiTrack Prompt Enhancer Image List Bridge Bridges an image list into the prompt enhancer.
Timeline Editor A simpler single-track timeline editor for basic sequencing.
Timeline Info Output Outputs metadata (dimensions, frame count, frame rate) from the Timeline Editor.
Timeline Segment Output Outputs the data for a selected timeline segment.
Timeline Segment Count Outputs the number of segments on the timeline.

Subtitles

Node Description
Recognize Subtitle Transcribes audio or video into subtitle segments.
Add Subtitle To Video Burns subtitles onto a video.
MultiTrack Add Subtitle To Video Burns the editor's subtitle track onto the assembled video.

Audio

Node Description
Moxie Audio Lock Locks a task to its input audio: the delivered video keeps the original task audio instead of regenerating it, while the audio also constrains the visual generation.
Merge Audio Merges multiple audio inputs into a single audio track.
Split Audios Splits an audio list into individual audio outputs.

Image

Node Description
Make Image List Combines multiple images into a single image list.
Split Images Splits an image list into individual image outputs.
Image Indexes To Int List Converts image indexes into a list of integers for per-index processing.

Video

Node Description
Save Video Encodes images, audio, and a frame rate into a saved video file with a filename prefix.
Compare Videos Shows two videos side by side in a synchronized comparison player.
Get Audio From Video Extracts the audio track from a video.
Merge Videos Stitches multiple video clips into one.
Merge Videos From Paths Stitches video files from file paths into one.
Split Videos Splits a video list into individual video outputs.
Make Video List Combines multiple videos into a single video list.

Logic & Loaders

Node Description
Model Loader Pack Bundles a model, CLIP, VAE, and optional audio VAE into a single loader output.
Match Line Returns the zero-based index of the first line containing the match text.
API Workflow Gate Passes the input through only for API-format executions; skips evaluating the input during normal workflow runs.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support