Download README.md from KBlueLeaf/Trinity: direct link, hf CLI and curl.
- Browser
- Download file 9.22 kB
-
https://huggingface.co/KBlueLeaf/Trinity/resolve/main/README.md
- Command line
-
hf download hf://KBlueLeaf/Trinity/README.md
-
curl -L -o README.md https://huggingface.co/KBlueLeaf/Trinity/resolve/main/README.md
license: apache-2.0
tags:
- floorplanning
- chip-design
- eda
- diffusion
- rectified-flow
- physics-informed
Trinity: One Differentiable Physics for Training, Refining and Scoring Generative Floorplanners
Shih-Ying Yeh♡♠†, Tzu-Sian Wang♡, Xuehai Wang♣△, Jia-Hua Lee★, Daniel Z. Kaplan◇, Ming-Qi Xu♡, Wuqian Tang♡, Chun-Yao Wang♡, Shang-Hong Lai♡, Chun-Yi Lee★
♡National Tsing Hua University · ♠Kohaku Lab · ♣Karolinska Institutet · △Stockholm University · ◇realiz.ai · ★National Taiwan University †Corresponding author: kohaku@kblueleaf.net
Project page · Code · arXiv (coming soon)
Sampling (32 steps, no guidance), refinement (400 steps on the same physics) and legalization (one linear program) of a 60-block FloorSet chip. Left and right panels: the layout before and after each step.
This repository holds the released denoisers of Trinity: the flagship model and eight smaller sizes, all trained on FloorSet with the same recipe. Given a chip (its blocks, their areas and shape limits, the netlist, the fixed and preplaced blocks, the cluster, MIB and boundary constraints), a model generates the block positions and shapes. The code that loads and runs them is Kohaku-Lab/Trinity.
The idea
Existing generative floorplanners train only to reproduce reference layouts and leave the rules of the chip to corrections added afterwards: guidance in the sampler, post-hoc loops and a legalizer, each in its own form, with only the final layout ever scored. Trinity writes every constraint and objective of floorplanning once, as six differentiable functions of the layout (overlap, grouping, MIB, boundary, wirelength, area), and uses that one physics three times:
- training: as a term of the denoiser's loss, so the network learns the correction and sampling needs no guidance;
- refining: as the energy a closed-form refiner descends on a finished sample;
- scoring: as a soft cost that measures a layout at every stage, before legalization.
Model details
| network | plain transformer, one token per block: RMSNorm, SwiGLU MLP, QK-norm, a netlist bias in attention, head dimension 64 |
| netlist | spectral-drawing graph positional encoding, plus the info-mover residual (a graph convolution along the netlist between attention layers) |
| generative model | rectified flow; the network predicts the clean layout, the sampler is Euler with any number of steps (NFE) |
| physics term | all six constraint terms on the predicted clean layout, weight λ = 0.01 |
| training data | FloorSet-Lite training chips, 21 to 120 blocks |
| recipe | AdamW, peak learning rate 2 × 10⁻⁴ with cosine decay, effective batch 256, 200,000 steps, one GPU, wire-dropout mixture; identical for every size |
| weights | EMA weights, model.safetensors per size, with a config.json that rebuilds the model |
Released sizes
Each subfolder is one size, named by its hidden width d and layer count l. All sizes share
the recipe above and differ only in width and depth. Raw samples (no refiner, no legalizer) on
our 12,000-chip FloorSet validation split; soft cost at 32 / 1 sampling steps, lower is better.
| subfolder | width × depth | parameters | overlap (NFE 32) | soft cost, NFE 32 / 1 |
|---|---|---|---|---|
flagship |
768 × 12 | 92.9 M | 0.0225 | 1.514 / 1.679 |
d512_l16 |
512 × 16 | 56.0 M | 0.0226 | 1.530 / 1.679 |
d512 |
512 × 12 | 42.1 M | 0.0229 | 1.534 / 1.682 |
d384 |
384 × 12 | 23.3 M | 0.0277 | 1.612 / 1.759 |
d256 |
256 × 12 | 10.6 M | 0.0297 | 1.703 / 1.812 |
d192 |
192 × 12 | 5.9 M | 0.0347 | 1.827 / 1.903 |
l8 |
768 × 8 | 62.2 M | 0.0239 | 1.593 / 1.740 |
l6 |
768 × 6 | 46.8 M | 0.0309 | 1.733 / 1.865 |
l4 |
768 × 4 | 31.5 M | 0.0483 | 2.128 / 2.258 |
Width is cheaper than depth: d512 matches the flagship within its seed spread at 45% of the
parameters, while removing layers costs more per parameter, and l4 is the one size that
cannot carry the netlist far enough. Even d192 at 5.9 M parameters stays below every ported
placer's raw soft cost (FlowPlace 2.46, ChipDiffusion 2.80, MacroDiff+ 3.85, DiffPlace 6.23).
Results
Four recent diffusion placers (FlowPlace, ChipDiffusion, MacroDiff+, DiffPlace) were re-implemented on the same data and training recipe, and every model was scored at every stage on 12,000 held-out FloorSet chips. The soft cost extends the FloorSet contest's hard cost to layouts that still overlap, so raw, refined and legalized layouts share one scale.
| raw soft cost, physics term on vs. off (same transformer) | −26% |
| steps for our refiner to match each placer's own loop | 16–660× fewer |
| refined soft cost vs. the best existing pipeline | −36% |
| soft cost vs. hard cost, rank agreement across settings (Spearman) | 0.94 |
| FloorSet validation set, mean hard cost (best of 48) | 1.014 at 1.63 s per chip |
Every method through its own pipeline at 32 sampling steps (the refined column uses each placer's own post-hoc loop at its released setting, Trinity uses 400 steps of its refiner):
| method | raw overlap | refined soft cost | legalized hard cost, first draw / best of 4 | seconds per draw |
|---|---|---|---|---|
| Trinity (flagship) | 0.023 | 1.06 | 1.234 / 1.089 | 0.28 |
| FlowPlace | 0.044 | no loop | 1.567 / 1.343 | 0.36 |
| ChipDiffusion | 0.051 | 2.05 | 1.521 / 1.300 | 0.41 |
| MacroDiff+ | 0.064 | 1.66 | 1.546 / 1.308 | 0.39 |
| DiffPlace | 0.113 | 3.91 | 1.537 / 1.313 | 0.36 |
| PARSAC (simulated annealing, 10⁷ steps) | 0.000 | no loop | 1.685 / 1.348 | 1,120 |
The physics term needs no reference layout, so it keeps supervising where the data runs out. On GSRC, with chips beyond FloorSet's 21 to 120 blocks and a 10% dead-space outline (best legalized layout over 4 seeds × 16 draws):
| GSRC | with physics: wirelength / draws that fit | without physics: wirelength / draws that fit | PARSAC wirelength |
|---|---|---|---|
| n100 | 294.7 mm / 36% | 293.9 mm / 80% | 303.5 mm |
| n200 | 554.9 mm / 100% | 560.1 mm / 0% (outline 440 × 498 > 440 × 440) | 576.0 mm |
| n300 | 681.3 mm / 0% | 679.8 mm / 0% | 705.2 mm |
Usage
git clone https://github.com/Kohaku-Lab/Trinity.git
cd Trinity
pip install -e .
from trinity.floorplan.data import load_validation_case
from trinity.floorplan.scoring import validate
from trinity.hub import load_model
model = load_model("KBlueLeaf/Trinity", scale="flagship", device="cuda") # or "d512", "d192", ...
placer = model.placer(samples=16) # sampler -> closed-form refiner -> legalizer, best of 16
instance = load_validation_case(60) # FloorSet validation case with 60 blocks
placement = placer.solve(instance)
print(validate(placement, "full").score)
The sampler writes the known answers into its state by default (preplaced and fixed blocks,
one shared shape per MIB group). The paper's experiments sample free:
model.placer(projections=[]). Training, evaluation and every baseline are in the
code repository.
Limitations
- The models are trained on FloorSet chips with 21 to 120 blocks. Far beyond that range the generator degrades: at 300 blocks (GSRC n300) no model fits the 10% dead-space outline.
- A raw sample is not a legal floorplan. Use the placer (refiner and legalizer) for legal layouts; the raw output is meant to be refined.
- The models place blocks of a floorplan only; they do not route or place standard cells.
Citation
@misc{trinity2026,
title = {Trinity: One Differentiable Physics for Training, Refining and Scoring Generative Floorplanners},
author = {Yeh, Shih-Ying and Wang, Tzu-Sian and Wang, Xuehai and Lee, Jia-Hua and Kaplan, Daniel Z. and
Xu, Ming-Qi and Tang, Wuqian and Wang, Chun-Yao and Lai, Shang-Hong and Lee, Chun-Yi},
year = {2026},
note = {arXiv identifier to appear}
}
License
Apache-2.0. FloorSet and the other third-party data keep their own licenses.




