Text-to-Image
Diffusion Single File
comfyui
convrot
int8
hadamard
ruwwww commited on
Commit
c68ee71
·
verified ·
1 Parent(s): 21fbee0

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +24 -10
README.md CHANGED
@@ -10,19 +10,35 @@ tags:
10
  - diffusion-single-file
11
  - convrot
12
  - int8
 
13
  pipeline_tag: text-to-image
14
  ---
15
 
16
- # Anima DiT Base v1.0 - ConvRot INT8 Quantized
17
 
18
- This repository provides the **ConvRot INT8** quantized checkpoint of [circlestone-labs/Anima](https://huggingface.co/circlestone-labs/Anima) (Base v1.0, Cosmos 2 DiT architecture).
19
 
20
- ## Key Advantages of ConvRot INT8 vs MXFP8 / FP8
21
 
22
- - **Orthogonal Outlier Smoothing:** Unlike standard FP8/MXFP8 quantization which suffers from severe activation outliers in DiT channels, ConvRot applies a group-wise Regular Hadamard Transform (RHT, $N_0=256$) to disperse outliers evenly across all channels.
23
- - **Superior Semantic Fidelity (~90% Winrate vs MXFP8):** Generates outputs that are visually and structurally almost identical to the BF16 ground truth while eliminating color banding, blurry textures, or anatomical distortion.
24
- - **High Generation Speed:** Runs at **~2.59 it/s (~11.5s total)** on RTX 5060 Ti 16GB at 832x1216 (30 steps, batched CFG $B=2$) with native ComfyUI compile.
 
 
25
  - **VRAM Savings:** File size reduced from 3.89 GB down to **2.41 GB (~38% reduction)** with peak active VRAM under ~5.2 GB.
 
 
 
 
 
 
 
 
 
 
 
 
 
26
 
27
  ## Preserved Layers (100% BF16 Fidelity)
28
 
@@ -38,11 +54,9 @@ Quantized with ConvRot ($N_0=256$):
38
 
39
  ## Usage in ComfyUI
40
 
41
- This model embeds native ComfyUI `comfy_quant` metadata. It requires **ZERO custom node registration**:
42
- 1. Download `anima-base-v1.0-convrot-int8.safetensors` to `ComfyUI/models/diffusion_models/`.
43
- 2. Select it directly in the standard `UNETLoader` node.
44
  3. Launch ComfyUI with `--fast --use-ck-attention`.
45
- 4. ComfyUI and `comfy-kitchen` will automatically route the layers to hardware-fused `ck.int8_linear` kernels on your GPU.
46
 
47
  For the full optimization recipe and converter tools, see:
48
  https://github.com/ruwwww/anima-fastpath-recipe
 
10
  - diffusion-single-file
11
  - convrot
12
  - int8
13
+ - hadamard
14
  pipeline_tag: text-to-image
15
  ---
16
 
17
+ # Anima DiT Collection - ConvRot INT8 Quantized Models
18
 
19
+ Official high-fidelity **ConvRot INT8** quantized checkpoints for the **Anima DiT** family (Cosmos 2 DiT architecture) created by [circlestone-labs/Anima](https://huggingface.co/circlestone-labs/Anima).
20
 
21
+ ## Why ConvRot INT8?
22
 
23
+ Standard FP8/MXFP8 quantization on Diffusion Transformers introduces activation outlier truncation errors, often leading to color banding, blurry eye/facial details, or subtle anatomical deformities.
24
+
25
+ ConvRot (**Convolution-like Regular Hadamard Rotation**, arXiv:2512.03673) solves this by pre-rotating weight and activation coordinates with an orthogonal regular Hadamard matrix ($N_0=256$, $H H^\top = I$).
26
+ - **~90% Winrate vs MXFP8:** Produces outputs nearly identical to BF16 ground truth in blind side-by-side evaluations.
27
+ - **Fast Execution:** Reaches **~2.59 it/s (~11.5s total)** on RTX 5060 Ti 16GB (832x1216, 30 steps, batched CFG $B=2$)—only ~0.3s difference from raw MXFP8.
28
  - **VRAM Savings:** File size reduced from 3.89 GB down to **2.41 GB (~38% reduction)** with peak active VRAM under ~5.2 GB.
29
+ - **Zero Custom Nodes:** Standard ComfyUI `comfy_quant` metadata allows `comfy-kitchen` to dispatch `ck.int8_linear` kernels out of the box.
30
+
31
+ ## Available Checkpoints
32
+
33
+ | Checkpoint File | Base Model | Suggested Steps | Suggested CFG | Description |
34
+ | :--- | :--- | :--- | :--- | :--- |
35
+ | `anima-base-v1.0-convrot-int8.safetensors` | Base v1.0 | 25 - 30 | 4.0 - 5.0 | Standard production foundation model |
36
+ | `anima-turbo-v1.0-convrot-int8.safetensors` | Turbo v1.0 | 8 - 12 | 1.0 - 2.0 | Distilled fast few-step generator |
37
+ | `anima-turbo-v1.1-convrot-int8.safetensors` | Turbo v1.1 | 8 - 12 | 1.0 - 2.0 | Updated turbo with improved sharpness |
38
+ | `anima-aesthetic-v1.0b-convrot-int8.safetensors` | Aesthetic v1.0b | 25 - 30 | 4.0 - 5.0 | Fine-tuned for anime aesthetic fidelity |
39
+ | `anima-preview-convrot-int8.safetensors` | Preview 1 | 25 - 30 | 4.0 - 5.0 | Early release checkpoint |
40
+ | `anima-preview2-convrot-int8.safetensors` | Preview 2 | 25 - 30 | 4.0 - 5.0 | Preview edition v2 |
41
+ | `anima-preview3-base-convrot-int8.safetensors` | Preview 3 Base | 25 - 30 | 4.0 - 5.0 | Preview edition v3 |
42
 
43
  ## Preserved Layers (100% BF16 Fidelity)
44
 
 
54
 
55
  ## Usage in ComfyUI
56
 
57
+ 1. Place any downloaded checkpoint in `ComfyUI/models/diffusion_models/`.
58
+ 2. Select it directly in the `UNETLoader` node.
 
59
  3. Launch ComfyUI with `--fast --use-ck-attention`.
 
60
 
61
  For the full optimization recipe and converter tools, see:
62
  https://github.com/ruwwww/anima-fastpath-recipe