Coderdw commited on
Commit
00a8f55
·
verified ·
1 Parent(s): d4ceb4a

Add dynamic-resolution model graphs

Browse files
README.md CHANGED
@@ -12,12 +12,9 @@ tags:
12
 
13
  # ERNIE-Image-Turbo ncnn 权重
14
 
15
- 这是 [`baidu/ERNIE-Image-Turbo`](https://huggingface.co/baidu/ERNIE-Image-Turbo)
16
- 的实验性 C++/ncnn 移植所使用的 FP32 部署文件。模型包包含配套 C++ CLI 所需的
17
- Tokenizer、Text Encoder、36 层 DiT、VAE Decoder,以及官方 seed=42 回归
18
- latent。
19
 
20
- 请将整个仓库下载到同一个本地目录,并把该目录传给运行程序:
21
 
22
  ```bash
23
  hf download Coderdw/ernie-image-ncnn-vulkan \
@@ -26,12 +23,17 @@ hf download Coderdw/ernie-image-ncnn-vulkan \
26
  ./build/ernie_image_cli \
27
  --model-dir models/ernie-image-turbo \
28
  --device vulkan \
 
29
  --rng-mode portable \
30
  --seed 123 \
31
- --prompt '一只黑白相间的中华田园犬' \
 
 
32
  --output image.png
33
  ```
34
 
 
 
35
  模型目录结构:
36
 
37
  ```text
@@ -42,10 +44,6 @@ vae/ decoder 与归一化常量
42
  rng/ 官方 CUDA/BF16 seed=42 回归 latent
43
  ```
44
 
45
- 当前正确性基线使用 FP32,因此模型包约为 43 GiB。当前固定 batch=1、分辨率
46
- 1024 x 1024、推理步数 8,并且不包含 Prompt Enhancer。请使用配套代码仓库内置
47
- 的 ncnn 源码;其 Vulkan GELU 与 SDPA mask 行为属于已验证运行契约的一部分。
48
 
49
- 这些文件转换自 `baidu/ERNIE-Image-Turbo` revision
50
- `bc68c81e2a1730a394d5fc9fae70713dee940140`。许可证与模型信息请参阅本目录中的
51
- `LICENSE` 以及上游模型页面。
 
12
 
13
  # ERNIE-Image-Turbo ncnn 权重
14
 
15
+ 这是 [`baidu/ERNIE-Image-Turbo`](https://huggingface.co/baidu/ERNIE-Image-Turbo) 的 C++/ncnn FP32 部署模型,配套代码位于 [`everythingfornothing/ernie-image-ncnn-vulkan`](https://github.com/everythingfornothing/ernie-image-ncnn-vulkan)。模型包包含 Tokenizer、Text Encoder、36 层 DiT、VAE Decoder 和官方 seed=42 回归 latent。
 
 
 
16
 
17
+ ## 下载与运行
18
 
19
  ```bash
20
  hf download Coderdw/ernie-image-ncnn-vulkan \
 
23
  ./build/ernie_image_cli \
24
  --model-dir models/ernie-image-turbo \
25
  --device vulkan \
26
+ --gpu-index 0 \
27
  --rng-mode portable \
28
  --seed 123 \
29
+ --width 1376 \
30
+ --height 768 \
31
+ --prompt '一只黑白相间的中华田园犬在草地上奔跑' \
32
  --output image.png
33
  ```
34
 
35
+ 图片宽高默认为 1024×1024,必须为正数且均为 16 的倍数;动态分辨率使用 `portable` RNG。同一套权重已完整运行 1024×1024、1376×768、768×1376 和 528×784,并支持运行时动态文本长度。`reference` RNG 仅用于官方 1024×1024、seed=42 回归。
36
+
37
  模型目录结构:
38
 
39
  ```text
 
44
  rng/ 官方 CUDA/BF16 seed=42 回归 latent
45
  ```
46
 
47
+ 当前模型包采用 FP32,约为 43 GiB,固定 batch=1 和 Turbo 8-step,不包含 Prompt Enhancer。请使用配套代码仓库内置的 ncnn 源码,其中 Vulkan GELU 与 SDPA mask 行为属于已验证运行契约。
 
 
48
 
49
+ 这些文件转换自 `baidu/ERNIE-Image-Turbo` revision `bc68c81e2a1730a394d5fc9fae70713dee940140`,遵循 Apache-2.0 许可证。
 
 
dit/frontend.ncnn.param CHANGED
@@ -4,7 +4,7 @@ Input in0 0 1 in0
4
  Input in1 0 1 in1
5
  Input in2 0 1 in2
6
  Convolution conv_1 1 1 in0 3 0=4096 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=524288
7
- Reshape reshape_7 1 1 3 4 0=4096 1=4096
8
  Permute transpose_8 1 1 4 out0 0=1
9
  Gemm gemm_0 1 1 in2 out1 10=-1 2=0 3=1 4=0 5=1 6=1 7=0 8=4096 9=3072
10
  MemoryData pnnx_fold_133 0 1 7 0=2048 1=1
 
4
  Input in1 0 1 in1
5
  Input in2 0 1 in2
6
  Convolution conv_1 1 1 in0 3 0=4096 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=524288
7
+ Reshape reshape_7 1 1 3 4 0=-1 1=4096
8
  Permute transpose_8 1 1 4 out0 0=1
9
  Gemm gemm_0 1 1 in2 out1 10=-1 2=0 3=1 4=0 5=1 6=1 7=0 8=4096 9=3072
10
  MemoryData pnnx_fold_133 0 1 7 0=2048 1=1
dit/output_head.ncnn.param CHANGED
@@ -1,5 +1,5 @@
1
  7767517
2
- 14 15
3
  Input in0 0 1 in0
4
  Input in1 0 1 in1
5
  InnerProduct linear_2 1 1 in1 2 0=8192 1=1 2=33554432
@@ -10,7 +10,4 @@ ExpandDims unsqueeze_6 1 1 4 7 -23303=1,0
10
  BinaryOp add_0 1 1 6 8 0=0 1=1 2=1.0
11
  BinaryOp mul_1 2 1 5 8 9 0=2
12
  BinaryOp add_2 2 1 9 7 10 0=0
13
- Gemm gemm_0 1 1 10 11 10=4 2=0 3=1 4=0 5=1 6=1 7=0 8=128 9=4096
14
- Crop slice_0 1 1 11 12 -23310=1,4096 -23311=1,0 -23309=1,0
15
- Permute transpose_4 1 1 12 13 0=1
16
- Reshape reshape_3 1 1 13 out0 0=64 1=64 2=128
 
1
  7767517
2
+ 11 12
3
  Input in0 0 1 in0
4
  Input in1 0 1 in1
5
  InnerProduct linear_2 1 1 in1 2 0=8192 1=1 2=33554432
 
10
  BinaryOp add_0 1 1 6 8 0=0 1=1 2=1.0
11
  BinaryOp mul_1 2 1 5 8 9 0=2
12
  BinaryOp add_2 2 1 9 7 10 0=0
13
+ Gemm gemm_0 1 1 10 out0 10=4 2=0 3=1 4=0 5=1 6=1 7=0 8=128 9=4096
 
 
 
vae/decoder.ncnn.param CHANGED
@@ -1,5 +1,5 @@
1
  7767517
2
- 134 149
3
  Input in0 0 1 in0
4
  Convolution conv_3 1 1 in0 1 0=32 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024
5
  Convolution conv_4 1 1 1 2 0=512 1=3 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=147456
@@ -11,13 +11,13 @@ GroupNorm gn_40 1 1 7 8 0=32 1=512 2=1.000000e
11
  Swish silu_71 1 1 8 9
12
  Convolution conv_6 1 1 9 10 0=512 1=3 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=2359296
13
  BinaryOp add_0 2 1 3 10 11 0=0
14
- Split splitncnn_1 1 2 11 12 13
15
- Reshape reshape_99 1 1 13 14 0=16384 1=512
16
  GroupNorm gn_41 1 1 14 15 0=32 1=512 2=1.000000e-6 3=1
17
  Permute transpose_101 1 1 15 16 0=1
18
  MultiHeadAttention attention_69 1 1 16 17 0=512 1=1 2=262144 3=512 4=512
19
  Permute transpose_102 1 1 17 18 0=1
20
- Reshape reshape_100 1 1 18 19 0=128 1=128 2=512
21
  BinaryOp add_1 2 1 19 12 20 0=0
22
  Split splitncnn_2 1 2 20 21 22
23
  GroupNorm gn_42 1 1 22 23 0=32 1=512 2=1.000000e-6 3=1
 
1
  7767517
2
+ 134 150
3
  Input in0 0 1 in0
4
  Convolution conv_3 1 1 in0 1 0=32 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024
5
  Convolution conv_4 1 1 1 2 0=512 1=3 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=147456
 
11
  Swish silu_71 1 1 8 9
12
  Convolution conv_6 1 1 9 10 0=512 1=3 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=2359296
13
  BinaryOp add_0 2 1 3 10 11 0=0
14
+ Split splitncnn_1 1 3 11 12 13 shape_ref
15
+ Reshape reshape_99 1 1 13 14 0=-1 1=512
16
  GroupNorm gn_41 1 1 14 15 0=32 1=512 2=1.000000e-6 3=1
17
  Permute transpose_101 1 1 15 16 0=1
18
  MultiHeadAttention attention_69 1 1 16 17 0=512 1=1 2=262144 3=512 4=512
19
  Permute transpose_102 1 1 17 18 0=1
20
+ Reshape reshape_100 2 1 18 shape_ref 19 6="1w,1h,512"
21
  BinaryOp add_1 2 1 19 12 20 0=0
22
  Split splitncnn_2 1 2 20 21 22
23
  GroupNorm gn_42 1 1 22 23 0=32 1=512 2=1.000000e-6 3=1