File size: 5,611 Bytes
c137565
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
---
license: apache-2.0
tags:
- diffusion
- flow-matching
- rectified-flow
- text-generation
- reasoning
- qwen2.5
- block-diffusion
- non-autoregressive
pipeline_tag: text-generation
language:
- en
library_name: diffusers
---

# πŸš€ BlockDiffuse: Fully Parallel Latent Space Reasoning Generation

[![License: Apache 2.0](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)
[![GitHub Repository](https://img.shields.io/badge/GitHub-Hooshaai%2FBlockDiffuse-black.svg?logo=github)](https://github.com/Hooshaai/BlockDiffuse)
[![HuggingFace Space](https://img.shields.io/badge/HF%20Space-Interactive%20Weblog-blueviolet.svg)](https://huggingface.co/spaces/tahamajs/BlockDiffuse-Blog)
[![Dataset](https://img.shields.io/badge/HF%20Dataset-BlockDiffuse--Data-orange.svg)](https://huggingface.co/datasets/tahamajs/BlockDiffuse-Data)

> **TL;DR:** BlockDiffuse is a non-autoregressive / block-autoregressive generative framework that generates **100 tokens simultaneously** in continuous latent space using **Rectified Flow Matching** and a **Diffusion Transformer (DiT)** conditioned on intermediate layers of modern LLMs (`Qwen/Qwen2.5-0.5B-Instruct`).

---

## ⚑ Key Highlights & Benchmark Results

All benchmarks measured on a single consumer **NVIDIA GeForce RTX 4070 Laptop GPU (8GB VRAM)**:

| Generation Mode | Target Size | ODE Steps / Block | Numerical Solver | Latency (ms) | Throughput (tokens/sec) | VRAM Footprint |
| :--- | :--- | :--- | :--- | :--- | :--- | :--- |
| **Single-Block Parallel** | **100 tokens** | 8 ODE steps | DPM-Solver + TFE | **1,730.60 ms** | **57.78 tok/s** | 3,674 MB |
| **Multi-Block Autoregressive** | **200 tokens** | 8 ODE steps / block | DPM-Solver + TFE | **1,279.20 ms** | **156.35 tok/s** | 3,789 MB |

---

## πŸ—οΈ Architecture Overview

```
Prompt Prefix ──► Frozen Qwen2.5 (Layers 1..12) ──► Continuous Context c [L_p x 896]
                                                          β”‚
Initial Gaussian Noise z_0 [100 x 896] ~ N(0, I) ──────────
                                                          β–Ό
                                             BlockDiffuse DiT (8 Layers, 14 Heads)
                                             - AdaLN-Zero Timestep Conditioning
                                             - Continuous RoPE Positional Encoding
                                             - Rectified Flow (v-prediction)
                                                          β”‚
                                                          β–Ό
                                            Predicted Latents z_1 [100 x 896]
                                                          β”‚
                                                          β–Ό
                                             Deep Proj Head (3-Layer SwiGLU MLP)
                                                          β”‚
                                                          β–Ό
                                            Pre-Head RMSNorm + Frozen LM Head
                                                          β”‚
                                                          β–Ό
                                            Discrete Next 100 Tokens in Parallel
```

### 1. Base LLM Backbone
- **Model**: `Qwen/Qwen2.5-0.5B-Instruct`
- **Representation Layer**: Layer 12 (mid-layer context extraction, $d_{\text{model}} = 896$).
- **Head**: Frozen LM head with vocab size $151{,}936$.

### 2. Diffusion Transformer (DiT)
- **Depth**: 8 Transformer Blocks.
- **Attention**: 14 heads (head dimension 64, matches $d_{\text{model}} = 896$).
- **Initialization**: Direct parameter transfer from layers 6–11 of Qwen2.5-0.5B.
- **Modulation**: AdaLN-Zero modulates scale and shift parameters based on timestep $t \in [0, 1]$.

### 3. Flow Matching & Multi-Objective Training
Rectified Flow straight-line trajectory:
$$z_t = (1 - t) z_0 + t z_1, \quad v_t = \frac{dz_t}{dt} = z_1 - z_0$$

Trained under composite multi-loss:
$$\mathcal{L}_{\text{total}} = \lambda_{\text{FM}} \mathcal{L}_{\text{FM}} + \lambda_{\text{disp}} \mathcal{L}_{\text{disp}} + \lambda_{\text{KL}} \mathcal{L}_{\text{KL}} + \lambda_{\text{CE}} \mathcal{L}_{\text{CE}} + \lambda_{\text{NN}} \mathcal{L}_{\text{NN}}$$

---

## πŸ’» Quickstart: Inference

### 1. Clone & Setup
```bash
git clone https://github.com/Hooshaai/BlockDiffuse.git
cd BlockDiffuse
pip install -r requirements.txt
```

### 2. Download Checkpoint from Hugging Face
```python
from huggingface_hub import hf_hub_download

ckpt_path = hf_hub_download(
    repo_id="tahamajs/BlockDiffuse",
    filename="blockdiffuse_final.pt"
)
print("Checkpoint downloaded to:", ckpt_path)
```

### 3. Run Parallel Multi-Block Generation
```bash
python inference.py \
    --model Qwen/Qwen2.5-0.5B-Instruct \
    --checkpoint ./checkpoints_improved/blockdiffuse_final.pt \
    --prompt "<|im_start|>system\nYou are a helpful assistant that solves problems step by step.<|im_end|>\n<|im_start|>user\nA bookstore has 140 books on Monday. On Tuesday, they sell 45 books. On Wednesday, they receive 80 books. How many remain?<|im_end|>\n<|im_start|>assistant\n" \
    --max_blocks 2 \
    --steps 8 \
    --solver dpm_solver \
    --use_tfe \
    --tfe_seeds 3
```

---

## πŸ“œ Citation

```bibtex
@article{blockdiffuse2026,
  title={BlockDiffuse: Fully Parallel Latent Space Reasoning Generation with Diffusion Transformers},
  author={Hooshaai Research},
  journal={GitHub / HuggingFace Technical Report},
  year={2026},
  url={https://github.com/Hooshaai/BlockDiffuse}
}
```