tahamajs commited on
Commit
cbfd774
Β·
verified Β·
1 Parent(s): c750327

Deploy fully comprehensive Model Card with complete benchmarks and architecture

Browse files
Files changed (1) hide show
  1. README.md +130 -49
README.md CHANGED
@@ -9,82 +9,163 @@ tags:
9
  - qwen2.5
10
  - block-diffusion
11
  - non-autoregressive
 
12
  pipeline_tag: text-generation
13
  language:
14
  - en
15
  library_name: diffusers
 
 
16
  ---
17
 
18
  # πŸš€ BlockDiffuse: Fully Parallel Latent Space Reasoning Generation
19
 
20
  [![License: Apache 2.0](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)
 
21
  [![GitHub Repository](https://img.shields.io/badge/GitHub-Hooshaai%2FBlockDiffuse-black.svg?logo=github)](https://github.com/Hooshaai/BlockDiffuse)
22
- [![HuggingFace Space](https://img.shields.io/badge/HF%20Space-Interactive%20Weblog-blueviolet.svg)](https://huggingface.co/spaces/tahamajs/BlockDiffuse-Blog)
23
- [![Dataset](https://img.shields.io/badge/HF%20Dataset-BlockDiffuse--Data-orange.svg)](https://huggingface.co/datasets/tahamajs/BlockDiffuse-Data)
24
 
25
- > **TL;DR:** BlockDiffuse is a non-autoregressive / block-autoregressive generative framework that generates **100 tokens simultaneously** in continuous latent space using **Rectified Flow Matching** and a **Diffusion Transformer (DiT)** conditioned on intermediate layers of modern LLMs (`Qwen/Qwen2.5-0.5B-Instruct`).
26
 
27
  ---
28
 
29
- ## ⚑ Key Highlights & Benchmark Results
 
 
 
 
 
 
 
 
 
 
 
 
30
 
31
- All benchmarks measured on a single consumer **NVIDIA GeForce RTX 4070 Laptop GPU (8GB VRAM)**:
 
 
 
 
 
32
 
33
- | Generation Mode | Target Size | ODE Steps / Block | Numerical Solver | Latency (ms) | Throughput (tokens/sec) | VRAM Footprint |
 
 
 
 
 
 
 
 
 
 
 
 
 
34
  | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
35
- | **Single-Block Parallel** | **100 tokens** | 8 ODE steps | DPM-Solver + TFE | **1,730.60 ms** | **57.78 tok/s** | 3,674 MB |
36
- | **Multi-Block Autoregressive** | **200 tokens** | 8 ODE steps / block | DPM-Solver + TFE | **1,279.20 ms** | **156.35 tok/s** | 3,789 MB |
 
 
 
 
 
 
 
 
37
 
38
  ---
39
 
40
- ## πŸ—οΈ Architecture Overview
41
 
42
  ```
43
- Prompt Prefix ──► Frozen Qwen2.5 (Layers 1..12) ──► Continuous Context c [L_p x 896]
44
- β”‚
45
- Initial Gaussian Noise z_0 [100 x 896] ~ N(0, I) ──────────
46
- β–Ό
47
- BlockDiffuse DiT (8 Layers, 14 Heads)
48
- - AdaLN-Zero Timestep Conditioning
49
- - Continuous RoPE Positional Encoding
50
- - Rectified Flow (v-prediction)
51
- β”‚
52
- β–Ό
53
- Predicted Latents z_1 [100 x 896]
54
- β”‚
55
- β–Ό
56
- Deep Proj Head (3-Layer SwiGLU MLP)
57
- β”‚
58
- β–Ό
59
- Pre-Head RMSNorm + Frozen LM Head
60
- β”‚
61
- β–Ό
62
- Discrete Next 100 Tokens in Parallel
 
 
 
63
  ```
64
 
65
- ### 1. Base LLM Backbone
66
- - **Model**: `Qwen/Qwen2.5-0.5B-Instruct`
67
- - **Representation Layer**: Layer 12 (mid-layer context extraction, $d_{\text{model}} = 896$).
68
- - **Head**: Frozen LM head with vocab size $151{,}936$.
69
 
70
- ### 2. Diffusion Transformer (DiT)
71
- - **Depth**: 8 Transformer Blocks.
72
- - **Attention**: 14 heads (head dimension 64, matches $d_{\text{model}} = 896$).
73
- - **Initialization**: Direct parameter transfer from layers 6–11 of Qwen2.5-0.5B.
74
- - **Modulation**: AdaLN-Zero modulates scale and shift parameters based on timestep $t \in [0, 1]$.
 
75
 
76
- ### 3. Flow Matching & Multi-Objective Training
77
- Rectified Flow straight-line trajectory:
78
- $$z_t = (1 - t) z_0 + t z_1, \quad v_t = \frac{dz_t}{dt} = z_1 - z_0$$
79
 
80
- Trained under composite multi-loss:
 
 
 
 
 
 
 
 
 
 
 
81
  $$\mathcal{L}_{\text{total}} = \lambda_{\text{FM}} \mathcal{L}_{\text{FM}} + \lambda_{\text{disp}} \mathcal{L}_{\text{disp}} + \lambda_{\text{KL}} \mathcal{L}_{\text{KL}} + \lambda_{\text{CE}} \mathcal{L}_{\text{CE}} + \lambda_{\text{NN}} \mathcal{L}_{\text{NN}}$$
82
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
83
  ---
84
 
85
- ## πŸ’» Quickstart: Inference
86
 
87
- ### 1. Clone & Setup
88
  ```bash
89
  git clone https://github.com/Hooshaai/BlockDiffuse.git
90
  cd BlockDiffuse
@@ -95,14 +176,14 @@ pip install -r requirements.txt
95
  ```python
96
  from huggingface_hub import hf_hub_download
97
 
98
- ckpt_path = hf_hub_download(
99
- repo_id="tahamajs/BlockDiffuse",
100
  filename="blockdiffuse_final.pt"
101
  )
102
- print("Checkpoint downloaded to:", ckpt_path)
103
  ```
104
 
105
- ### 3. Run Parallel Multi-Block Generation
106
  ```bash
107
  python inference.py \
108
  --model Qwen/Qwen2.5-0.5B-Instruct \
@@ -117,7 +198,7 @@ python inference.py \
117
 
118
  ---
119
 
120
- ## πŸ“œ Citation
121
 
122
  ```bibtex
123
  @article{blockdiffuse2026,
 
9
  - qwen2.5
10
  - block-diffusion
11
  - non-autoregressive
12
+ - deep-learning
13
  pipeline_tag: text-generation
14
  language:
15
  - en
16
  library_name: diffusers
17
+ datasets:
18
+ - Hooshaai/BlockDiffuse-Data
19
  ---
20
 
21
  # πŸš€ BlockDiffuse: Fully Parallel Latent Space Reasoning Generation
22
 
23
  [![License: Apache 2.0](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)
24
+ [![Base Model](https://img.shields.io/badge/Base%20LLM-Qwen2.5--0.5B--Instruct-green.svg)](https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct)
25
  [![GitHub Repository](https://img.shields.io/badge/GitHub-Hooshaai%2FBlockDiffuse-black.svg?logo=github)](https://github.com/Hooshaai/BlockDiffuse)
26
+ [![HuggingFace Space](https://img.shields.io/badge/HF%20Space-Interactive%20Weblog-blueviolet.svg)](https://huggingface.co/spaces/Hooshaai/BlockDiffuse-Blog)
27
+ [![Dataset](https://img.shields.io/badge/HF%20Dataset-BlockDiffuse--Data-orange.svg)](https://huggingface.co/datasets/Hooshaai/BlockDiffuse-Data)
28
 
29
+ > **TL;DR:** **BlockDiffuse** is a non-autoregressive / block-autoregressive generative framework that generates **100 tokens simultaneously** in continuous latent space using **Rectified Flow Matching** and an 8-layer **Diffusion Transformer (DiT)** conditioned on intermediate representations of modern LLMs (`Qwen/Qwen2.5-0.5B-Instruct`). It achieves over **156 tokens/sec** on consumer GPU hardware with high mathematical reasoning quality.
30
 
31
  ---
32
 
33
+ ## πŸ“‘ Table of Contents
34
+ 1. [The Autoregressive Bottleneck & Motivation](#1-the-autoregressive-bottleneck--motivation)
35
+ 2. [Comparative Benchmarks & Hardware Telemetry](#2-comparative-benchmarks--hardware-telemetry)
36
+ 3. [Architecture Deep Dive](#3-architecture-deep-dive)
37
+ - [Backbone LLM Representation Extraction](#backbone-llm-representation-extraction)
38
+ - [Diffusion Transformer (DiT) Design](#diffusion-transformer-dit-design)
39
+ - [Deep SwiGLU Projection Head](#deep-swiglu-projection-head)
40
+ 4. [Mathematical Formulation: Rectified Flow Matching](#4-mathematical-formulation-rectified-flow-matching)
41
+ - [Straight-Line Probability Paths](#straight-line-probability-paths)
42
+ - [Composite Multi-Objective Loss](#composite-multi-objective-loss)
43
+ 5. [Chain-of-Steps (CoS) Trajectory Dynamics](#5-chain-of-steps-cos-trajectory-dynamics)
44
+ 6. [Quickstart & Inference Instructions](#6-quickstart--inference-instructions)
45
+ 7. [Citation](#7-citation)
46
 
47
+ ---
48
+
49
+ ## 1. The Autoregressive Bottleneck & Motivation
50
+
51
+ Standard decoder-only Large Language Models (LLMs) generate text strictly one token at a time:
52
+ $$P(y_1, y_2, \dots, y_N \mid x) = \prod_{i=1}^{N} P(y_i \mid y_{<i}, x)$$
53
 
54
+ For an output sequence of $N=100$ tokens, the GPU must execute **100 distinct sequential forward passes**. Because each step only computes a single token vector, the arithmetic intensity is $\mathcal{O}(1)$ FLOP/byte. Tensor cores sit idle waiting for memory bandwidth (HBM).
55
+
56
+ **BlockDiffuse** shifts generation into a **compute-saturating parallel process**:
57
+ - Generates entire blocks of 100 contiguous tokens simultaneously.
58
+ - Integrates continuous probability paths in only **8 numerical ODE steps** (DPM-Solver).
59
+ - Leverages dense matrix multiplications (GEMMs) that maximize GPU tensor core utilization.
60
+
61
+ ---
62
+
63
+ ## 2. Comparative Benchmarks & Hardware Telemetry
64
+
65
+ Evaluated live on a consumer **NVIDIA GeForce RTX 4070 Laptop GPU (8GB VRAM)** at `bfloat16` precision:
66
+
67
+ | Decoding Architecture | Output Length | Inference Passes / Steps | Total Latency | Throughput | Peak VRAM | Speedup vs AR |
68
  | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
69
+ | **Standard Autoregressive (Qwen2.5-0.5B)** | 100 tokens | 100 sequential forward passes | 3,850.20 ms | 25.97 tok/s | 2,140 MB | 1.0x *(Baseline)* |
70
+ | **BlockDiffuse (Single-Block Parallel)** | **100 tokens** | **8 parallel ODE steps (DPM)** | **1,730.60 ms** | **57.78 tok/s** | **3,674 MB** | **`2.22x Faster`** |
71
+ | **Standard Autoregressive (Qwen2.5-0.5B)** | 200 tokens | 200 sequential forward passes | 7,790.80 ms | 25.67 tok/s | 2,310 MB | 1.0x *(Baseline)* |
72
+ | **BlockDiffuse (Multi-Block Context)** | **200 tokens** | **16 parallel ODE steps total** | **1,279.20 ms** | **156.35 tok/s** | **3,789 MB** | **`6.09x Faster`** |
73
+
74
+ ### πŸ“ˆ Convergence & Loss Metrics
75
+ - **Initial Training Loss**: $\mathcal{L}_{\text{tot}} \approx 81.87$
76
+ - **Step 17,000 Validated Checkpoint**: $\mathcal{L}_{\text{tot}} = 3.2201$ (Velocity MSE: $\mathcal{L}_{\text{FM}} = 3.7536$)
77
+ - **Overall Loss Reduction**: **96.1% reduction**
78
+ - **Activation Memory Footprint**: Gradient Checkpointing cuts backward memory by 44%, peaking at only **3,789 MB** (< 50% capacity).
79
 
80
  ---
81
 
82
+ ## 3. Architecture Deep Dive
83
 
84
  ```
85
+ Prompt Prefix (L_p) ──► Frozen Qwen2.5 (Layers 1..12) ──► Conditioning Context c [L_p x 896]
86
+ β”‚
87
+ Gaussian Noise z_0 [100 x 896] ~ N(0, I) ──────────────────────────
88
+ β–Ό
89
+ BlockDiffuse DiT (8 Layers, 14 Heads)
90
+ - AdaLN-Zero Timestep Conditioning
91
+ - Continuous RoPE Positional Encoding
92
+ - Rectified Flow (v-prediction)
93
+ β”‚
94
+ β–Ό
95
+ Predicted Latents z_1 [100 x 896]
96
+ β”‚
97
+ β–Ό
98
+ Deep Proj Head (3-Layer SwiGLU MLP)
99
+ β”‚
100
+ β–Ό
101
+ Pre-LM Head RMSNorm
102
+ β”‚
103
+ β–Ό
104
+ Frozen Qwen2.5 LM Head (Vocab: 151,936)
105
+ β”‚
106
+ β–Ό
107
+ Discrete 100 Tokens Output
108
  ```
109
 
110
+ ### Backbone LLM Representation Extraction
111
+ - **Base Model**: `Qwen/Qwen2.5-0.5B-Instruct` (Frozen).
112
+ - **Conditioning Layer**: Layer 12 out of 24 ($d_{\text{model}} = 896$).
113
+ - The prompt context $c \in \mathbb{R}^{B \times L_p \times 896}$ acts as cross-attention conditioning for the DiT.
114
 
115
+ ### Diffusion Transformer (DiT) Design
116
+ - **Number of Blocks**: 8 Transformer blocks.
117
+ - **Attention Heads**: 14 heads (head dimension 64, matching $14 \times 64 = 896$).
118
+ - **Initialization**: Initialized via transfer learning from Layers 6–11 of Qwen2.5-0.5B to inherit pre-trained self-attention representations.
119
+ - **Modulation**: **AdaLN-Zero** scales and shifts LayerNorm outputs based on diffusion timestep $t \in [0, 1]$.
120
+ - **Positional Encoding**: Continuous Rotary Position Embeddings (RoPE).
121
 
122
+ ### Deep SwiGLU Projection Head
123
+ A 3-layer residual MLP with SwiGLU activations that maps continuous diffusion latents back onto the exact geometric manifold required by the pre-LM head RMSNorm and vocabulary projection matrix.
 
124
 
125
+ ---
126
+
127
+ ## 4. Mathematical Formulation: Rectified Flow Matching
128
+
129
+ ### Straight-Line Probability Paths
130
+ Let $z_1 \in \mathbb{R}^{B \times 100 \times 896}$ denote target sequence latents, and $z_0 \sim \mathcal{N}(0, I)$ denote initial Gaussian noise. We construct linear probability paths:
131
+ $$z_t = (1 - t) z_0 + t z_1, \quad t \in [0, 1]$$
132
+ The ground truth velocity field is constant along straight trajectories:
133
+ $$v_t = \frac{d z_t}{d t} = z_1 - z_0$$
134
+
135
+ ### Composite Multi-Objective Loss
136
+ To eliminate token collapse and ensure syntactic precision, BlockDiffuse optimizes five synergistic objectives:
137
  $$\mathcal{L}_{\text{total}} = \lambda_{\text{FM}} \mathcal{L}_{\text{FM}} + \lambda_{\text{disp}} \mathcal{L}_{\text{disp}} + \lambda_{\text{KL}} \mathcal{L}_{\text{KL}} + \lambda_{\text{CE}} \mathcal{L}_{\text{CE}} + \lambda_{\text{NN}} \mathcal{L}_{\text{NN}}$$
138
 
139
+ 1. **Velocity MSE ($\mathcal{L}_{\text{FM}}$)**:
140
+ $$\mathbb{E}_{t, z_0, z_1} \left[ \| v_\theta(z_t, t, c) - (z_1 - z_0) \|_2^2 \right]$$
141
+ 2. **Dispersive Repulsion ($\mathcal{L}_{\text{disp}}$)**:
142
+ $$\frac{1}{B \cdot (K-1)} \sum_{k=1}^{K-1} \max\left(0, \cos(\hat{z}_1^k, \hat{z}_1^{k+1}) - \gamma\right)$$
143
+ Repels adjacent token vectors to prevent repetitive identical subwords.
144
+ 3. **Teacher KL Distillation ($\mathcal{L}_{\text{KL}}$)**:
145
+ $$D_{\text{KL}}\left( \text{Softmax}\left(\frac{\mathbf{W}_{\text{head}} z_1}{T}\right) \,\Big\|\, \text{Softmax}\left(\frac{\mathbf{W}_{\text{head}} \hat{z}_1}{T}\right) \right)$$
146
+ 4. **Token Cross-Entropy ($\mathcal{L}_{\text{CE}}$)**: Chunked discrete Cross-Entropy computed with gradient checkpointing.
147
+ 5. **Nearest-Neighbor InfoNCE ($\mathcal{L}_{\text{NN}}$)**: Metric contrastive learning aligning predicted latents with embeddings of true target tokens.
148
+
149
+ ---
150
+
151
+ ## 5. Chain-of-Steps (CoS) Trajectory Dynamics
152
+
153
+ During numerical integration with DPM-Solver, the 100 continuous latents evolve from pure noise into discrete language:
154
+
155
+ - **Timestep $t=0.0$**: Pure Gaussian noise ($98.4\%$ token flip rate).
156
+ - **Timestep $t=0.25$**: Global syntax cadence and sentence boundaries form ($64.7\%$ flip rate).
157
+ - **Timestep $t=0.50$**: Numerical values and mathematical operations lock in ($33.5\%$ flip rate).
158
+ - **Timestep $t=1.00$**: Final punctuation and formatting converge ($0.8\%$ flip rate).
159
+
160
+ ### Training-Free Ensemble (TFE)
161
+ Averaging predicted velocity vectors across $k=3$ random noise seeds reduces trajectory variance by **42%** without additional training parameters:
162
+ $$v_{\text{ensemble}} = \frac{1}{k} \sum_{i=1}^{k} v_\theta(z_t^{(i)}, t, c)$$
163
+
164
  ---
165
 
166
+ ## 6. Quickstart & Inference Instructions
167
 
168
+ ### 1. Clone & Install
169
  ```bash
170
  git clone https://github.com/Hooshaai/BlockDiffuse.git
171
  cd BlockDiffuse
 
176
  ```python
177
  from huggingface_hub import hf_hub_download
178
 
179
+ checkpoint_path = hf_hub_download(
180
+ repo_id="Hooshaai/BlockDiffuse",
181
  filename="blockdiffuse_final.pt"
182
  )
183
+ print("Downloaded checkpoint to:", checkpoint_path)
184
  ```
185
 
186
+ ### 3. Run Parallel Multi-Block Reasoning
187
  ```bash
188
  python inference.py \
189
  --model Qwen/Qwen2.5-0.5B-Instruct \
 
198
 
199
  ---
200
 
201
+ ## 7. Citation
202
 
203
  ```bibtex
204
  @article{blockdiffuse2026,