Moxie Multimedia Suite
Timeline-driven video generation for ComfyUI β arrange tasks, references, audio, and subtitles on a multi-track timeline and render them with the bundled Moxie model set.
π¦ Installation
Make sure FFmpeg is installed and available in your system PATH before installing this pack.
Install through ComfyUI Manager, or clone this repository into your ComfyUI custom_nodes folder:
cd Your_ComfyUI_Path/custom_nodes
git clone https://huggingface.co/turtle89431/Moxie-Multimedia
Installing dependencies
ComfyUI Manager installs requirements.txt automatically. If you cloned the repository manually, install the dependencies yourself:
ComfyUI Portable (Windows) β run from the portable root folder (the one containing run_nvidia_gpu.bat):
cd Your_ComfyUI_windows_portable
.\python_embeded\python.exe -m pip install -r .\ComfyUI\custom_nodes\Moxie-Multimedia\requirements.txt
Regular ComfyUI (venv / system Python):
pip install -r Your_ComfyUI_Path/custom_nodes/Moxie-Multimedia/requirements.txt
Restart ComfyUI afterwards.
The pack targets NVIDIA RTX 3060+ GPUs only, so every feature dependency is mandatory and installed by the command above β including speech generation (
voxcpm), subtitle recognition (openai-whisper,qwen-asr), and the RTX Video Super Resolution upscale (nvidia-vfx). Expect a sizable first install:qwen-asrpinstransformers==4.57.6(pip may upgrade it inside your ComfyUI environment) andvoxcpmneedstorch>=2.5.0, so keep ComfyUI reasonably current.nvidia-vfxis pulled from NVIDIA's own package index (pypi.nvidia.com, wired up insiderequirements.txt) because the PyPI copy is a stub that fails to build on some setups β if the command is interrupted, re-run it and it resumes where it stopped.
Models
All model files ship inside the repo β cloning it into custom_nodes is all that's needed. The loader finds everything automatically in the pack's models/ folders:
| File | Location |
|---|---|
Diffusion model (Moxie-Multimedia.safetensors) |
models/diffusion/ |
Text encoder (MM-VL.safetensors) |
models/clip/ |
| Video VAE (bundled file) | models/vae/ |
| Audio VAE (bundled file) | models/audio_vae/ |
Turbo LoRA (MM3step.safetensors, applied by default) |
models/loras/ |
Preview tiny VAE (MM-preview.safetensors) |
models/vae_approx/ |
All model files live in the Hugging Face version of this repository β cloning it downloads the complete set, and there is no separate downloader step. If you obtained the pack from somewhere else and the models/ folders are missing or incomplete, re-clone from Hugging Face to get the bundled files.
Restart ComfyUI afterwards.
π Pipeline
[Moxie Multimedia Loader] ββ model βββΊ [Moxie Preview Override] (live preview)
β model, clip, β
β video_vae, audio_vae βΌ model
ββββββββββββββββΊ [MultiTrack Editor] ββ (TRACKS_INFO) βββΊ [Moxie MultiTrack Project] (steps, seed, project name)
β (PROJECT_NAME)
βΌ
[Moxie Project Video Combine]
β (images, audio, fps)
βΌ
[Save Video]
π§© Nodes
Core Pipeline
| Node | Description |
|---|---|
| Moxie Multimedia Loader | One-stop loader for the bundled model set. All required files are resolved automatically from the pack's models/ folders β add the node and it outputs model, clip, video_vae, and audio_vae, ready for the project pipeline. |
| Moxie Preview Override | Streams a live animated preview β motion and sound β of the generation directly onto the node canvas while sampling runs, so you can watch progress without waiting for the final export. |
| MultiTrack Editor | The creative hub of the suite. Arrange task, video, audio, and subtitle tracks on a timeline, set the output dimensions and frame rate, write prompts per segment, and attach reference images, video, or audio. Outputs TRACKS_INFO for the project node (plus media outputs when slot references are used). |
| Moxie MultiTrack Project | Renders the timeline. Takes the loader outputs plus TRACKS_INFO, expands the tasks in sequence automatically (no manual segment loop), and handles continuity between segments. Exposes only the creative controls: steps, seed, and project_name. Each segment is saved as it completes. |
| Moxie Project Video Combine | Assembles the final video. Previews the saved segments, lets you pick a version for each, and combines them β automatically in one workflow run, or manually after review. Outputs images, audio, and fps for Save Video. |
Timeline & MultiTrack Tools
| Node | Description |
|---|---|
| Multi Images Loader | Provides an ordered list of up to 25 images with shared resolution choices. Add images through the media selector or drop image files onto the grid; each image is resized and output in grid order. |
| MultiTrack Info Output | Outputs the video dimensions, total frame count, frame rate, and task count from the editor's TRACKS_INFO. |
| MultiTrack Task Output | Outputs the user and system prompts for each task, with media loaded on demand. Set the task index for one segment at a time, or use -1 to retrieve the entire timeline's media at once. |
| MultiTrack Audio Output | Outputs the audio for the timeline's tasks. |
| MultiTrack Prompt Enhancer | Expands task prompts with a connected language model, using each segment's user and system prompts plus its reference media. |
| MultiTrack Prompt Enhance To Project | Sends enhanced prompts straight into the project pipeline. |
| MultiTrack Prompt Enhance To Project Apply | Applies the enhanced prompt results to a running project. |
| MultiTrack Prompt Enhancer Image List Bridge | Bridges an image list into the prompt enhancer. |
| Timeline Editor | A simpler single-track timeline editor for basic sequencing. |
| Timeline Info Output | Outputs metadata (dimensions, frame count, frame rate) from the Timeline Editor. |
| Timeline Segment Output | Outputs the data for a selected timeline segment. |
| Timeline Segment Count | Outputs the number of segments on the timeline. |
Subtitles
| Node | Description |
|---|---|
| Recognize Subtitle | Transcribes audio or video into subtitle segments. |
| Add Subtitle To Video | Burns subtitles onto a video. |
| MultiTrack Add Subtitle To Video | Burns the editor's subtitle track onto the assembled video. |
Audio
| Node | Description |
|---|---|
| Moxie Audio Lock | Locks a task to its input audio: the delivered video keeps the original task audio instead of regenerating it, while the audio also constrains the visual generation. |
| Merge Audio | Merges multiple audio inputs into a single audio track. |
| Split Audios | Splits an audio list into individual audio outputs. |
Image
| Node | Description |
|---|---|
| Make Image List | Combines multiple images into a single image list. |
| Split Images | Splits an image list into individual image outputs. |
| Image Indexes To Int List | Converts image indexes into a list of integers for per-index processing. |
Video
| Node | Description |
|---|---|
| Save Video | Encodes images, audio, and a frame rate into a saved video file with a filename prefix. |
| Compare Videos | Shows two videos side by side in a synchronized comparison player. |
| Get Audio From Video | Extracts the audio track from a video. |
| Merge Videos | Stitches multiple video clips into one. |
| Merge Videos From Paths | Stitches video files from file paths into one. |
| Split Videos | Splits a video list into individual video outputs. |
| Make Video List | Combines multiple videos into a single video list. |
Logic & Loaders
| Node | Description |
|---|---|
| Model Loader Pack | Bundles a model, CLIP, VAE, and optional audio VAE into a single loader output. |
| Match Line | Returns the zero-based index of the first line containing the match text. |
| API Workflow Gate | Passes the input through only for API-format executions; skips evaluating the input during normal workflow runs. |