Trinity / README.md
KBlueLeaf's picture
Authors: Tzu-Sian Wang shares first authorship (equal contribution) and is also at Kohaku Lab
06efa43 verified
|
Raw History Blame Contribute Delete
9.3 kB
---
license: apache-2.0
tags:
- floorplanning
- chip-design
- eda
- diffusion
- rectified-flow
- physics-informed
---
# Trinity: One Differentiable Physics for Training, Refining and Scoring Generative Floorplanners
**Shih-Ying Yeh**<sup>♡♠</sup><sub>*†</sub>, **Tzu-Sian Wang**<sup>♡♠</sup><sub>*</sub>, **Xuehai Wang**<sup>♣</sup><sub>△</sub>, **Jia-Hua Lee**<sup>★</sup>, **Daniel Z. Kaplan**<sup>◇</sup>, **Ming-Qi Xu**<sup>♡</sup>, **Wuqian Tang**<sup>♡</sup>, **Chun-Yao Wang**<sup>♡</sup>, **Shang-Hong Lai**<sup>♡</sup>, **Chun-Yi Lee**<sup>★</sup>
<sup>♡</sup>National Tsing Hua University · <sup>♠</sup>Kohaku Lab · <sup>♣</sup>Karolinska Institutet · <sup>△</sup>Stockholm University · <sup>◇</sup>realiz.ai · <sup>★</sup>National Taiwan University
<sup>*</sup>Equal contribution · <sup>†</sup>Corresponding author: kohaku@kblueleaf.net
[Project page](https://kohaku-lab.github.io/Trinity/) · [Code](https://github.com/Kohaku-Lab/Trinity) · arXiv (coming soon)
![Trinity sampling, refining and legalizing a 60-block FloorSet chip](https://huggingface.co/KBlueLeaf/Trinity/resolve/main/assets/pipeline_n60.webp)
*Sampling (32 steps, no guidance), refinement (400 steps on the same physics) and legalization (one linear program) of a 60-block FloorSet chip. Left and right panels: the layout before and after each step.*
This repository holds the released denoisers of **Trinity**: the flagship model and eight
smaller sizes, all trained on [FloorSet](https://github.com/IntelLabs/FloorSet) with the same
recipe. Given a chip (its blocks, their areas and shape limits, the netlist, the fixed and
preplaced blocks, the cluster, MIB and boundary constraints), a model generates the block
positions and shapes. The code that loads and runs them is
[Kohaku-Lab/Trinity](https://github.com/Kohaku-Lab/Trinity).
## The idea
Existing generative floorplanners train only to reproduce reference layouts and leave the rules
of the chip to corrections added afterwards: guidance in the sampler, post-hoc loops and a
legalizer, each in its own form, with only the final layout ever scored. Trinity writes every
constraint and objective of floorplanning once, as six differentiable functions of the layout
(overlap, grouping, MIB, boundary, wirelength, area), and uses that one physics three times:
* **training**: as a term of the denoiser's loss, so the network learns the correction and
sampling needs no guidance;
* **refining**: as the energy a closed-form refiner descends on a finished sample;
* **scoring**: as a soft cost that measures a layout at every stage, before legalization.
![Prior pipelines against Trinity, and soft cost along refinement](https://huggingface.co/KBlueLeaf/Trinity/resolve/main/assets/teaser.png)
## Model details
![Architecture and pipeline](https://huggingface.co/KBlueLeaf/Trinity/resolve/main/assets/arch.png)
| | |
|---|---|
| network | plain transformer, one token per block: RMSNorm, SwiGLU MLP, QK-norm, a netlist bias in attention, head dimension 64 |
| netlist | spectral-drawing graph positional encoding, plus the info-mover residual (a graph convolution along the netlist between attention layers) |
| generative model | rectified flow; the network predicts the clean layout, the sampler is Euler with any number of steps (NFE) |
| physics term | all six constraint terms on the predicted clean layout, weight λ = 0.01 |
| training data | FloorSet-Lite training chips, 21 to 120 blocks |
| recipe | AdamW, peak learning rate 2 × 10⁻⁴ with cosine decay, effective batch 256, 200,000 steps, one GPU, wire-dropout mixture; identical for every size |
| weights | EMA weights, `model.safetensors` per size, with a `config.json` that rebuilds the model |
## Released sizes
Each subfolder is one size, named by its hidden width `d` and layer count `l`. All sizes share
the recipe above and differ only in width and depth. Raw samples (no refiner, no legalizer) on
our 12,000-chip FloorSet validation split; soft cost at 32 / 1 sampling steps, lower is better.
| subfolder | width × depth | parameters | overlap (NFE 32) | soft cost, NFE 32 / 1 |
|---|---|---|---|---|
| **`flagship`** | 768 × 12 | 92.9 M | **0.0225** | **1.514 / 1.679** |
| `d512_l16` | 512 × 16 | 56.0 M | 0.0226 | 1.530 / 1.679 |
| `d512` | 512 × 12 | 42.1 M | 0.0229 | 1.534 / 1.682 |
| `d384` | 384 × 12 | 23.3 M | 0.0277 | 1.612 / 1.759 |
| `d256` | 256 × 12 | 10.6 M | 0.0297 | 1.703 / 1.812 |
| `d192` | 192 × 12 | 5.9 M | 0.0347 | 1.827 / 1.903 |
| `l8` | 768 × 8 | 62.2 M | 0.0239 | 1.593 / 1.740 |
| `l6` | 768 × 6 | 46.8 M | 0.0309 | 1.733 / 1.865 |
| `l4` | 768 × 4 | 31.5 M | 0.0483 | 2.128 / 2.258 |
![Raw soft cost against parameters for every size and the four ported placers](https://huggingface.co/KBlueLeaf/Trinity/resolve/main/assets/sizes.png)
Width is cheaper than depth: `d512` matches the flagship within its seed spread at 45% of the
parameters, while removing layers costs more per parameter, and `l4` is the one size that
cannot carry the netlist far enough. Even `d192` at 5.9 M parameters stays below every ported
placer's raw soft cost (FlowPlace 2.46, ChipDiffusion 2.80, MacroDiff+ 3.85, DiffPlace 6.23).
## Results
Four recent diffusion placers (FlowPlace, ChipDiffusion, MacroDiff+, DiffPlace) were
re-implemented on the same data and training recipe, and every model was scored at every
stage on 12,000 held-out FloorSet chips. The soft cost extends the FloorSet contest's hard
cost to layouts that still overlap, so raw, refined and legalized layouts share one scale.
![Raw soft cost against sampling steps, and our refiner against each placer's own loop](https://huggingface.co/KBlueLeaf/Trinity/resolve/main/assets/results.png)
| | |
|---|---|
| raw soft cost, physics term on vs. off (same transformer) | **−26%** |
| steps for our refiner to match each placer's own loop | **16–660× fewer** |
| refined soft cost vs. the best existing pipeline | **−36%** |
| soft cost vs. hard cost, rank agreement across settings (Spearman) | **0.94** |
| FloorSet validation set, mean hard cost (best of 48) | **1.014 at 1.63 s per chip** |
Every method through its own pipeline at 32 sampling steps (the refined column uses each
placer's own post-hoc loop at its released setting, Trinity uses 400 steps of its refiner):
| method | raw overlap | refined soft cost | legalized hard cost, first draw / best of 4 | seconds per draw |
|---|---|---|---|---|
| **Trinity (flagship)** | **0.023** | **1.06** | **1.234 / 1.089** | **0.28** |
| FlowPlace | 0.044 | no loop | 1.567 / 1.343 | 0.36 |
| ChipDiffusion | 0.051 | 2.05 | 1.521 / 1.300 | 0.41 |
| MacroDiff+ | 0.064 | 1.66 | 1.546 / 1.308 | 0.39 |
| DiffPlace | 0.113 | 3.91 | 1.537 / 1.313 | 0.36 |
| PARSAC (simulated annealing, 10⁷ steps) | 0.000 | no loop | 1.685 / 1.348 | 1,120 |
The physics term needs no reference layout, so it keeps supervising where the data runs out.
On GSRC, with chips beyond FloorSet's 21 to 120 blocks and a 10% dead-space outline (best
legalized layout over 4 seeds × 16 draws):
| GSRC | with physics: wirelength / draws that fit | without physics: wirelength / draws that fit | PARSAC wirelength |
|---|---|---|---|
| n100 | 294.7 mm / 36% | 293.9 mm / 80% | 303.5 mm |
| n200 | **554.9 mm / 100%** | 560.1 mm / 0% (outline 440 × 498 > 440 × 440) | 576.0 mm |
| n300 | 681.3 mm / 0% | 679.8 mm / 0% | 705.2 mm |
## Usage
```bash
git clone https://github.com/Kohaku-Lab/Trinity.git
cd Trinity
pip install -e .
```
```python
from trinity.floorplan.data import load_validation_case
from trinity.floorplan.scoring import validate
from trinity.hub import load_model
model = load_model("KBlueLeaf/Trinity", scale="flagship", device="cuda") # or "d512", "d192", ...
placer = model.placer(samples=16) # sampler -> closed-form refiner -> legalizer, best of 16
instance = load_validation_case(60) # FloorSet validation case with 60 blocks
placement = placer.solve(instance)
print(validate(placement, "full").score)
```
The sampler writes the known answers into its state by default (preplaced and fixed blocks,
one shared shape per MIB group). The paper's experiments sample free:
`model.placer(projections=[])`. Training, evaluation and every baseline are in the
[code repository](https://github.com/Kohaku-Lab/Trinity).
## Limitations
* The models are trained on FloorSet chips with 21 to 120 blocks. Far beyond that range the
generator degrades: at 300 blocks (GSRC n300) no model fits the 10% dead-space outline.
* A raw sample is not a legal floorplan. Use the placer (refiner and legalizer) for legal
layouts; the raw output is meant to be refined.
* The models place blocks of a floorplan only; they do not route or place standard cells.
## Citation
```bibtex
@misc{trinity2026,
title = {Trinity: One Differentiable Physics for Training, Refining and Scoring Generative Floorplanners},
author = {Yeh, Shih-Ying and Wang, Tzu-Sian and Wang, Xuehai and Lee, Jia-Hua and Kaplan, Daniel Z. and
Xu, Ming-Qi and Tang, Wuqian and Wang, Chun-Yao and Lai, Shang-Hong and Lee, Chun-Yi},
year = {2026},
note = {arXiv identifier to appear}
}
```
## License
Apache-2.0. FloorSet and the other third-party data keep their own licenses.