File size: 4,994 Bytes
92ff2cf
 
9618324
 
 
 
 
 
 
 
 
 
92ff2cf
9618324
d80bb29
9618324
d80bb29
9618324
b30e7d3
d80bb29
b30e7d3
 
9618324
 
 
 
 
 
d80bb29
9618324
d80bb29
 
 
 
 
 
 
9618324
 
 
 
 
d80bb29
9618324
d80bb29
 
9618324
 
 
 
 
 
 
d80bb29
9618324
d80bb29
 
9618324
 
 
 
 
d80bb29
9618324
d80bb29
9618324
a0471a0
d80bb29
a0471a0
 
d80bb29
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9618324
 
 
 
 
 
 
 
d80bb29
9618324
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
---
license: apache-2.0
tags:
- ccsr
- super-resolution
- image-to-image
- upscaling
- controlnet
- tensorrt
- rtx
- comfyui
pipeline_tag: image-to-image
---

# CCSR: TensorRT RTX Acceleration Engine

Ultra-fast **TensorRT RTX** execution engine and auxiliary modules for **CCSR (Creative Content Super-Resolution)**, designed for real-time generative image upscaling in **ComfyUI**.

> πŸ“¦ **ComfyUI Loader & Upscaler Extension:**
> All nodes supporting TensorRT engine execution are available in:
> πŸ‘‰ **[https://github.com/ussoewwin/ComfyUI-NunchakuFluxLoraStacker](https://github.com/ussoewwin/ComfyUI-NunchakuFluxLoraStacker)**

---

## 🌟 Overview

Creative Content Super-Resolution (CCSR) is a diffusion-based super-resolution framework leveraging a Controlled UNet and ControlNet structure to synthesize rich photorealistic textures and fine details.

This repository provides an optimized **NVIDIA TensorRT RTX Engine** implementation for CCSR:

- **Fused Denoising Engine (`ccsr_apply_f16io.rtxplan`)**:
  - Fuses the ControlNet and Controlled UNet denoising computation into a single compiled TensorRT engine.
  - Fixed 512px tile resolution (64Γ—64 latent tile) executing at ~**24 ms/step** (~**4.7Γ— speedup** over PyTorch FP16 at ~113 ms/step on modern RTX GPUs).
  - Synchronized stream execution on current PyTorch CUDA streams to eliminate race conditions and deadlocks.
- **Engine-Only Deployment (`ccsr_trt_aux.safetensors`)**:
  - Contains only the essential companion modules: FP16 AutoencoderKL (VAE encoder/decoder) and condition encoder.
  - Automatically loaded alongside the engine, eliminating the need to download large full checkpoints (~3.2 GB saved).

---

## πŸ“¦ Available Files

| Filename | Description | Architecture / Components | File Size | Recommended Location | License |
| :--- | :--- | :--- | :--- | :--- | :--- |
| `ccsr_apply_f16io.rtxplan` | TensorRT Fused Denoising Engine | ControlNet + UNet fused RTX Engine (Tile 512px / Latent 64Γ—64) | ~1.4 GB | `custom_nodes/.../nodes/CCSR/trt_engines/` | Apache-2.0 |
| `ccsr_trt_aux.safetensors` | TRT Auxiliary Weights | FP16 VAE AutoencoderKL + Condition Encoder | ~450 MB | `custom_nodes/.../nodes/CCSR/trt_engines/` | Apache-2.0 |

---

## βš™οΈ Performance & Benchmark Comparison

Measurements conducted on an NVIDIA RTX 4090 / RTX 5090 environment:

| Execution Mode | Files Required | VRAM Overhead (Denoising) | Step Latency (Tile 512) | Speedup |
| :--- | :--- | :--- | :--- | :--- |
| **Stock CCSR (FP16 PyTorch)** | Full Checkpoint (~3.2 GB) | ~3.8 GiB | ~113 ms / step | 1.0Γ— (Baseline) |
| **CCSR TensorRT RTX** | Engine + Aux (~1.85 GB total) | ~2.2 GiB | ~**24 ms / step** | **~4.7Γ— faster** |

---

## πŸš€ Usage in ComfyUI

TensorRT engine execution for CCSR is integrated natively into the **[ComfyUI-NunchakuFluxLoraStacker](https://github.com/ussoewwin/ComfyUI-NunchakuFluxLoraStacker)** custom-node pack.

### Workflow Example

<p align="center">
  <img src="https://huggingface.co/ussoewwin/CCSR-TensorRT-Engine/resolve/main/png/ccsrtensor.png" alt="CCSR TensorRT RTX Engine Workflow in ComfyUI" width="800">
</p>

### Installation & Setup

1. **Install the Custom Node Pack**:
   ```bash
   cd ComfyUI/custom_nodes
   git clone https://github.com/ussoewwin/ComfyUI-NunchakuFluxLoraStacker.git
   ```

2. **Place Engine & Aux Files**:
   Download both `ccsr_apply_f16io.rtxplan` and `ccsr_trt_aux.safetensors` and place them directly into the engine directory:
   ```
   ComfyUI/custom_nodes/ComfyUI-NunchakuFluxLoraStacker/nodes/CCSR/trt_engines/
   β”œβ”€β”€ ccsr_apply_f16io.rtxplan
   └── ccsr_trt_aux.safetensors
   ```

3. **In ComfyUI**:
   - Add **`Load CCSR Model (TensorRT)`** (`LoadCCSRModelTensorRT`). The node automatically discovers `.rtxplan` files in `trt_engines/` and loads the companion `ccsr_trt_aux.safetensors`.
   - Connect the `ccsr_model` output to **`CCSR Upscale (TRT)`** (`CCSR_Upscale_TRT`).
   - Connect an input image to `image`.
   - Configure upscale parameters:
     - `tile_size`: 512 (fixed to match the compiled static engine shape)
     - `tile_stride`: 256 (recommended for seamless blending)
     - `color_fix_type`: `adain` (or `wavelet` / `none`)
     - `steps`: Effective diffusion step count (densified schedule guarantees exact execution of requested step count)

---

## πŸ“œ Credits & License

- **Original CCSR Implementation & Weights**: [csslc/CCSR](https://github.com/csslc/CCSR) (Apache-2.0 License)
- **Research Paper**: *"Creative Content Super-Resolution with Pre-trained Diffusion Model"* by Liang et al.
- **ComfyUI CCSR Node Foundation**: [kijai/ComfyUI-CCSR](https://github.com/kijai/ComfyUI-CCSR) & [Kijai/ccsr-safetensors](https://huggingface.co/Kijai/ccsr-safetensors) (Apache-2.0 License)
- **TensorRT Engine & ComfyUI Loader**: [ussoewwin/ComfyUI-NunchakuFluxLoraStacker](https://github.com/ussoewwin/ComfyUI-NunchakuFluxLoraStacker) (CCSR nodes located under `nodes/CCSR/`)
- **License**: **Apache-2.0**