Instructions to use Cccccz/HY with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use Cccccz/HY with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("Cccccz/HY", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
File size: 10,611 Bytes
ba798d3 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 | # Data
本文只描述当前使用的 v4 数据体系。早期 v1 dense-text 数据和 v2 direct-K/V 数据已经完成其历史实验,
不再作为后续训练的数据标准。
## 1. 数据集总览
当前训练数据集名称为 `hyworldplay_predictor_prefeature_v4`,schema 为
`predictor_context_prefeature_bf16_v1`。
| 项目 | 数值 |
| --- | ---: |
| 图文 pair | 100 |
| 每个 pair 的动作轨迹 | 4 |
| 轨迹总数 | 400 |
| 每个视频帧数 | 125 |
| latent frames | 32 |
| 每条轨迹 chunks | 8 |
| 每个 chunk latent frames | 4 |
| 每个 chunk 去噪步 | 4 |
| 完整 manifest records | 3200 |
| 可训练 records | 2800 |
| 每个 record 的监督 pair | `0→1`、`1→2`、`2→3` |
| 训练 pair 总数 | 8400 |
| Context blocks | `[0,1,52,53]` |
| 保存精度 | BF16 safetensors |
| 实际 NVMe 占用 | 约 1.3 TB |
`train_eval_manifest.jsonl` 排除 chunk 0,因为 v4 需要前一 chunk 的 same-timestep hidden;
`predictor-disca` 为保持相同训练分布也沿用这 2800 条记录,但不会加载 previous-chunk hidden 或
Context pre-feature。
## 2. 数据来源和图像预处理
训练图文 pair 来自:
```text
/mnt/s3files/s3-us-west2-default/zoubin/cz/projects/VBench/
vbench2_beta_i2v/data/origin
```
选择方法:
1. 按路径名排序全部图片;
2. 使用 Python `random.Random(0).sample(...)` 随机选择 100 张;
3. caption 使用源图片文件名去掉扩展名后的文本;
4. 图像先保持宽高比,用 Lanczos 放大至覆盖目标区域,再做中心裁剪;
5. 保存为 832×480 JPEG,quality 95、无 chroma subsampling。
预处理入口:
```bash
python tools/prepare_predictor_prefeature_cases.py
```
输出为 `cases.jsonl`、`dataset_config.json` 和 `input_images/case_XXXX.jpg`。
## 3. 动作与视频配置
四种动作轨迹固定为:
| action_id | 名称 | pose |
| ---: | --- | --- |
| 0 | `w_s` | `w-15,s-16` |
| 1 | `a_d` | `a-15,d-16` |
| 2 | `up_down` | `up-15,down-16` |
| 3 | `left_right` | `left-15,right-16` |
生成配置为 seed 0、480×832、125 帧、32 latent frames、8 chunks、每 chunk 4 latent frames、
每个 chunk 4 个去噪步。构建训练 tensor 的同时保存每条轨迹对应的 125 帧 MP4,作为训练集
Full-DiT 评测参考视频。
## 4. Teacher 数据语义
Teacher 使用 HY-WorldPlay action-distilled AR checkpoint:
```text
/mnt/s3files/s3-us-west2-default/zoubin/cz/checkpoints/hy_worldplay/
huggingface/hub/models--tencent--HY-WorldPlay/snapshots/
f4c29235647707b571479a69b569e4166f9f5bf8/
ar_distilled_action_model/diffusion_pytorch_model.safetensors
```
数据构建使用精确 Full-DiT 路径:
- 四个去噪步都运行完整 54 层;
- 保留正式 WorldPlay text K/V 和 AR history vision memory;
- 关闭 denoise cache、history cache reuse、Context pooling/pruning;
- 关闭 sparse attention、SageAttention 和 FP8 GEMM;
- `prompt_rewrite=false`;
- Transformer 与落盘 feature 使用 BF16;
- offloading 只影响显存,不改变模型定义。
Context prefill 的语义是 `joint_noncausal_full_dit_prefill`:同一目标 chunk 选中的历史帧作为完整窗口
联合、非因果计算。因此深层 pre-feature 与目标窗口绑定,不能将历史 chunk 跨窗口去重。
## 5. 文件组织
```text
hyworldplay_predictor_prefeature_v4/
├── dataset_config.json
├── cases.jsonl
├── manifest.jsonl # 3200 records
├── train_eval_manifest.jsonl # 2800 records,排除 chunk 0
├── input_images/ # 100 张 832×480 图像
├── cases/ # 100 个 case tensors
├── text_kv_exact_all54_bf16/ # 100 个 Full Teacher 54 层 Text K/V
├── steps/ # 3200 个 chunk step tensors
│ └── case_XXXX/action_YY/chunk_ZZ.safetensors
├── context_prefeature/ # 12800 个 pre-feature tensors
│ ├── block_00/
│ ├── block_01/
│ ├── block_52/
│ └── block_53/
├── videos/ # 400 个 Full-DiT MP4
├── manifests/ # 8 个 worker manifests
└── logs/
```
实际空间分布:
| 目录 | 文件数 | 大小 |
| --- | ---: | ---: |
| `cases/` | 100 | 2.97 GB |
| `text_kv_exact_all54_bf16/` | 100 | 37.20 GiB |
| `steps/` | 3200 | 337.60 GB |
| `context_prefeature/` | 12800 | 1022.63 GB |
| `videos/` | 400 | 0.40 GB |
| `input_images/` | 100 | 0.02 GB |
## 6. Tensor schema
### 6.1 Case tensor
每个 case 只保存一次图像条件和四个候选 block 的 Text K/V:
```text
image_condition_latent [1,32,1,30,52] BF16
block_{00,01,52,53}_k_txt [1,16,903,128] BF16
block_{00,01,52,53}_v_txt [1,16,903,128] BF16
text_valid_mask [1,903] bool
```
Text K/V 在磁盘中补齐到 903 tokens,真实长度由 `text_valid_mask` 指示。
### 6.2 Full Teacher Text K/V
`text_kv_exact_all54_bf16/` 是离线数据集的一部分,不是 pre-feature,也不是 rollout sidecar。每个
case 一个文件:
```text
case_XXXX_text_kv_exact_all54_bf16.safetensors
block_{00..53}_k_txt [1,16,903,128] BF16
block_{00..53}_v_txt [1,16,903,128] BF16
text_valid_mask [1,903] bool
```
真实有效长度为 739–751 tokens,训练读取时按 `text_valid_mask` 去掉补零区,再交给 Teacher/Predictor。
构建会将 blocks `0、1、52、53` 与原 case tensor 逐元素比对,确保它来自相同的 frozen text prefill。
### 6.3 Step tensor
每个 chunk 保存共享动作/相机/RoPE 信息,以及四个去噪步的监督目标:
```text
action_labels [1,4] int64
target_viewmats [1,4,4,4] BF16
target_Ks [1,4,3,3] BF16
rope_temporal_size [1] int64
start_rope_start_idx [1] int64
step_{0..3}_timestep [1] FP32
step_{0..3}_noisy_sample [1,32,4,30,52] BF16
step_{0..3}_frame_condition [1,4,2048] BF16
step_{0..3}_final_hidden [1,6240,2048] BF16
step_{0..3}_velocity [1,32,4,30,52] BF16
```
65 通道 Predictor 输入不重复落盘。Dataset 使用 noisy sample、case-level image latent 和 I2V mask
重建 `[B,65,4,30,52]` 输入。
### 6.4 Context pre-feature
每个 record、每个 block 保存 K/V 线性投影之前的 `img_modulated`:
```text
img_modulated [1,S_context,2048] BF16
context_valid_mask [1,S_context] bool
selected_frame_indices [context_frames] int64
context_viewmats [1,context_frames,4,4] BF16
context_Ks [1,context_frames,3,3] BF16
rope_temporal_size [1] int64
start_rope_start_idx [1] int64
```
`S_context = context_frames × 30 × 52`。八个 chunks 的 Context frames 依次为
`0,4,8,12,16,20,20,20`。磁盘文件保持变长且不 padding;v4 collate 按 batch 最大 Context 长度
动态 padding,并用 valid mask 屏蔽补齐部分。
训练时由冻结的 Teacher `img_attn_k`、`img_attn_v`、`img_attn_k_norm` 和相同的 RoPE/ProPE 路径
重建 Context Vision K/V。直接 K/V 与重建 K/V 的验收阈值为 relative L2 ≤ 5e-3、cosine ≥ 0.9999。
## 7. Manifest 关系
主键是 `(case_id, action_id, chunk_id)`。每条记录包含 case、step、四个 Context 文件的相对路径,
以及 caption、源图、动作、seed、分辨率、Teacher checkpoint 和 Context frame 数。
合并 manifest 时:
- 检查 3200 个主键唯一;
- 验证所有引用文件存在;
- chunk 1–7 增加 `previous_chunk_id` 和 `previous_step_tensor_file`;
- `manifest.jsonl` 保留全部 3200 条;
- `train_eval_manifest.jsonl` 保留 chunk 1–7,共 2800 条。
## 8. 构建、恢复和持久化
NVMe 是构建和训练读取位置,S3 是持久化 source of truth:
```text
NVMe:
/mnt/local_nvme/zoubin/cz/hyworldplay_predictor_prefeature_v4
S3:
s3://s3-us-west2-default/zoubin/cz/projects/
HY-WorldPlay-DEV-Predictor/datasets/hyworldplay_predictor_prefeature_v4
项目挂载路径:
/mnt/s3files/s3-us-west2-default/zoubin/cz/projects/
HY-WorldPlay-DEV-Predictor/datasets/hyworldplay_predictor_prefeature_v4
```
完整构建流程:
```bash
python tools/prepare_predictor_prefeature_cases.py
bash tools/launch_predictor_prefeature_build.sh
```
第二条命令使用 8 个 GPU worker,先写 NVMe,随后合并 manifest 并同步到 S3。writer 支持按
`case/action/chunk` 跳过完整文件,任务中断后可重新执行。
全 54 层精确 Text K/V 单独构建并持久化到同一数据集:
```bash
bash tools/launch_predictor_text_kv_all54_build.sh
```
新机器训练前执行:
```bash
bash tools/ensure_predictor_prefeature_nvme.sh
```
需要全层 Text K/V 的 rollout 训练再执行:
```bash
bash tools/ensure_predictor_text_kv_all54_nvme.sh
```
如果 NVMe 不完整,该脚本使用 `aws s3 sync` 从持久前缀恢复并重新检查数量。手动将合法更新同步回
S3 使用:
```bash
bash tools/sync_predictor_prefeature_to_s3.sh
```
不要把 NVMe 当作唯一副本,也不要在未核对目标前缀时使用 `aws s3 sync --delete`。
## 9. 完整性验收
```bash
python tools/validate_predictor_prefeature_dataset.py
```
验收条件:
- 100 case tensors;
- 3200 step tensors;
- 12800 Context pre-feature tensors;
- 400 个 125 帧、832×480 视频;
- 3200 条无重复 manifest records 和 2800 条训练 records;
- 所有 tensor shape、dtype、finite、mask 和 Context frame 数正确;
- 四个 block 的 selected frame indices 一致;
- frozen final layer 能由 final hidden 与 frame condition 复现 velocity;
- 抽样重建的 65 通道输入与 Teacher 实际输入一致。
## 10. Validation/Test holdout
独立评估集为 `hyworldplay_predictor_vbench_val25_test50`:
- VBench metadata 共 355 个 pair;
- 排除训练集 100 个 pair 后剩 255 个;
- 使用 seed `20260718` 随机选择 75 个;
- validation 25 个、test 50 个,两者互斥且均与训练集互斥;
- 每个 pair 生成四种动作;
- 已保存 300 个 Full-DiT 和 300 个 Reuse 视频;
- Full 为 steps `[0,1,2,3]`,Reuse 为 Full `[0,3]`、复用 `[1,2]`。
入口:
```bash
bash tools/launch_vbench_holdout_full_reuse.sh
```
持久目录:
```text
datasets/hyworldplay_predictor_vbench_val25_test50
```
|