--- license: apache-2.0 tags: - floorplanning - chip-design - eda - diffusion - rectified-flow - physics-informed --- # Trinity: One Differentiable Physics for Training, Refining and Scoring Generative Floorplanners **Shih-Ying Yeh**♡♠*†, **Tzu-Sian Wang**♡♠*, **Xuehai Wang**♣△, **Jia-Hua Lee**★, **Daniel Z. Kaplan**◇, **Ming-Qi Xu**♡, **Wuqian Tang**♡, **Chun-Yao Wang**♡, **Shang-Hong Lai**♡, **Chun-Yi Lee**★ ♡National Tsing Hua University · ♠Kohaku Lab · ♣Karolinska Institutet · △Stockholm University · ◇realiz.ai · ★National Taiwan University *Equal contribution · †Corresponding author: kohaku@kblueleaf.net [Project page](https://kohaku-lab.github.io/Trinity/) · [Code](https://github.com/Kohaku-Lab/Trinity) · arXiv (coming soon) ![Trinity sampling, refining and legalizing a 60-block FloorSet chip](https://huggingface.co/KBlueLeaf/Trinity/resolve/main/assets/pipeline_n60.webp) *Sampling (32 steps, no guidance), refinement (400 steps on the same physics) and legalization (one linear program) of a 60-block FloorSet chip. Left and right panels: the layout before and after each step.* This repository holds the released denoisers of **Trinity**: the flagship model and eight smaller sizes, all trained on [FloorSet](https://github.com/IntelLabs/FloorSet) with the same recipe. Given a chip (its blocks, their areas and shape limits, the netlist, the fixed and preplaced blocks, the cluster, MIB and boundary constraints), a model generates the block positions and shapes. The code that loads and runs them is [Kohaku-Lab/Trinity](https://github.com/Kohaku-Lab/Trinity). ## The idea Existing generative floorplanners train only to reproduce reference layouts and leave the rules of the chip to corrections added afterwards: guidance in the sampler, post-hoc loops and a legalizer, each in its own form, with only the final layout ever scored. Trinity writes every constraint and objective of floorplanning once, as six differentiable functions of the layout (overlap, grouping, MIB, boundary, wirelength, area), and uses that one physics three times: * **training**: as a term of the denoiser's loss, so the network learns the correction and sampling needs no guidance; * **refining**: as the energy a closed-form refiner descends on a finished sample; * **scoring**: as a soft cost that measures a layout at every stage, before legalization. ![Prior pipelines against Trinity, and soft cost along refinement](https://huggingface.co/KBlueLeaf/Trinity/resolve/main/assets/teaser.png) ## Model details ![Architecture and pipeline](https://huggingface.co/KBlueLeaf/Trinity/resolve/main/assets/arch.png) | | | |---|---| | network | plain transformer, one token per block: RMSNorm, SwiGLU MLP, QK-norm, a netlist bias in attention, head dimension 64 | | netlist | spectral-drawing graph positional encoding, plus the info-mover residual (a graph convolution along the netlist between attention layers) | | generative model | rectified flow; the network predicts the clean layout, the sampler is Euler with any number of steps (NFE) | | physics term | all six constraint terms on the predicted clean layout, weight λ = 0.01 | | training data | FloorSet-Lite training chips, 21 to 120 blocks | | recipe | AdamW, peak learning rate 2 × 10⁻⁴ with cosine decay, effective batch 256, 200,000 steps, one GPU, wire-dropout mixture; identical for every size | | weights | EMA weights, `model.safetensors` per size, with a `config.json` that rebuilds the model | ## Released sizes Each subfolder is one size, named by its hidden width `d` and layer count `l`. All sizes share the recipe above and differ only in width and depth. Raw samples (no refiner, no legalizer) on our 12,000-chip FloorSet validation split; soft cost at 32 / 1 sampling steps, lower is better. | subfolder | width × depth | parameters | overlap (NFE 32) | soft cost, NFE 32 / 1 | |---|---|---|---|---| | **`flagship`** | 768 × 12 | 92.9 M | **0.0225** | **1.514 / 1.679** | | `d512_l16` | 512 × 16 | 56.0 M | 0.0226 | 1.530 / 1.679 | | `d512` | 512 × 12 | 42.1 M | 0.0229 | 1.534 / 1.682 | | `d384` | 384 × 12 | 23.3 M | 0.0277 | 1.612 / 1.759 | | `d256` | 256 × 12 | 10.6 M | 0.0297 | 1.703 / 1.812 | | `d192` | 192 × 12 | 5.9 M | 0.0347 | 1.827 / 1.903 | | `l8` | 768 × 8 | 62.2 M | 0.0239 | 1.593 / 1.740 | | `l6` | 768 × 6 | 46.8 M | 0.0309 | 1.733 / 1.865 | | `l4` | 768 × 4 | 31.5 M | 0.0483 | 2.128 / 2.258 | ![Raw soft cost against parameters for every size and the four ported placers](https://huggingface.co/KBlueLeaf/Trinity/resolve/main/assets/sizes.png) Width is cheaper than depth: `d512` matches the flagship within its seed spread at 45% of the parameters, while removing layers costs more per parameter, and `l4` is the one size that cannot carry the netlist far enough. Even `d192` at 5.9 M parameters stays below every ported placer's raw soft cost (FlowPlace 2.46, ChipDiffusion 2.80, MacroDiff+ 3.85, DiffPlace 6.23). ## Results Four recent diffusion placers (FlowPlace, ChipDiffusion, MacroDiff+, DiffPlace) were re-implemented on the same data and training recipe, and every model was scored at every stage on 12,000 held-out FloorSet chips. The soft cost extends the FloorSet contest's hard cost to layouts that still overlap, so raw, refined and legalized layouts share one scale. ![Raw soft cost against sampling steps, and our refiner against each placer's own loop](https://huggingface.co/KBlueLeaf/Trinity/resolve/main/assets/results.png) | | | |---|---| | raw soft cost, physics term on vs. off (same transformer) | **−26%** | | steps for our refiner to match each placer's own loop | **16–660× fewer** | | refined soft cost vs. the best existing pipeline | **−36%** | | soft cost vs. hard cost, rank agreement across settings (Spearman) | **0.94** | | FloorSet validation set, mean hard cost (best of 48) | **1.014 at 1.63 s per chip** | Every method through its own pipeline at 32 sampling steps (the refined column uses each placer's own post-hoc loop at its released setting, Trinity uses 400 steps of its refiner): | method | raw overlap | refined soft cost | legalized hard cost, first draw / best of 4 | seconds per draw | |---|---|---|---|---| | **Trinity (flagship)** | **0.023** | **1.06** | **1.234 / 1.089** | **0.28** | | FlowPlace | 0.044 | no loop | 1.567 / 1.343 | 0.36 | | ChipDiffusion | 0.051 | 2.05 | 1.521 / 1.300 | 0.41 | | MacroDiff+ | 0.064 | 1.66 | 1.546 / 1.308 | 0.39 | | DiffPlace | 0.113 | 3.91 | 1.537 / 1.313 | 0.36 | | PARSAC (simulated annealing, 10⁷ steps) | 0.000 | no loop | 1.685 / 1.348 | 1,120 | The physics term needs no reference layout, so it keeps supervising where the data runs out. On GSRC, with chips beyond FloorSet's 21 to 120 blocks and a 10% dead-space outline (best legalized layout over 4 seeds × 16 draws): | GSRC | with physics: wirelength / draws that fit | without physics: wirelength / draws that fit | PARSAC wirelength | |---|---|---|---| | n100 | 294.7 mm / 36% | 293.9 mm / 80% | 303.5 mm | | n200 | **554.9 mm / 100%** | 560.1 mm / 0% (outline 440 × 498 > 440 × 440) | 576.0 mm | | n300 | 681.3 mm / 0% | 679.8 mm / 0% | 705.2 mm | ## Usage ```bash git clone https://github.com/Kohaku-Lab/Trinity.git cd Trinity pip install -e . ``` ```python from trinity.floorplan.data import load_validation_case from trinity.floorplan.scoring import validate from trinity.hub import load_model model = load_model("KBlueLeaf/Trinity", scale="flagship", device="cuda") # or "d512", "d192", ... placer = model.placer(samples=16) # sampler -> closed-form refiner -> legalizer, best of 16 instance = load_validation_case(60) # FloorSet validation case with 60 blocks placement = placer.solve(instance) print(validate(placement, "full").score) ``` The sampler writes the known answers into its state by default (preplaced and fixed blocks, one shared shape per MIB group). The paper's experiments sample free: `model.placer(projections=[])`. Training, evaluation and every baseline are in the [code repository](https://github.com/Kohaku-Lab/Trinity). ## Limitations * The models are trained on FloorSet chips with 21 to 120 blocks. Far beyond that range the generator degrades: at 300 blocks (GSRC n300) no model fits the 10% dead-space outline. * A raw sample is not a legal floorplan. Use the placer (refiner and legalizer) for legal layouts; the raw output is meant to be refined. * The models place blocks of a floorplan only; they do not route or place standard cells. ## Citation ```bibtex @misc{trinity2026, title = {Trinity: One Differentiable Physics for Training, Refining and Scoring Generative Floorplanners}, author = {Yeh, Shih-Ying and Wang, Tzu-Sian and Wang, Xuehai and Lee, Jia-Hua and Kaplan, Daniel Z. and Xu, Ming-Qi and Tang, Wuqian and Wang, Chun-Yao and Lai, Shang-Hong and Lee, Chun-Yi}, year = {2026}, note = {arXiv identifier to appear} } ``` ## License Apache-2.0. FloorSet and the other third-party data keep their own licenses.