Commit History
Move deprecated sm_120 TRT engines under legacy/ 19ab22a verified
Move deprecated sm_90 TRT engines under legacy/ ceaecd4 verified
retire fp8_fast; remove the superseded sm_90 fp8 engines (miscalibrated encoder, 8192-ceiling decoders) 3963f0e verified
retire fp8_fast; remove the superseded sm_90 fp8 engines (miscalibrated encoder, 8192-ceiling decoders) f061405 verified
retire fp8_fast; remove the superseded sm_90 fp8 engines (miscalibrated encoder, 8192-ceiling decoders) 4b4cf7c verified
retire fp8_fast; remove the superseded sm_90 fp8 engines (miscalibrated encoder, 8192-ceiling decoders) 935685c verified
retire fp8_fast; remove the superseded sm_90 fp8 engines (miscalibrated encoder, 8192-ceiling decoders) 276f110 verified
retire fp8_fast; remove the superseded sm_90 fp8 engines (miscalibrated encoder, 8192-ceiling decoders) 0996c8a verified
retire fp8_fast; remove the superseded sm_90 fp8 engines (miscalibrated encoder, 8192-ceiling decoders) cf76999 verified
SAME-L/SAME-S: bake the limiter ceiling so the chunkable decoders keep the ('latent','pcm') signature; add the SAME-S fp8 chunkable pair 17258a7 verified
SAME-L/SAME-S: bake the limiter ceiling so the chunkable decoders keep the ('latent','pcm') signature; add the SAME-S fp8 chunkable pair 7ca4180 verified
SAME-L/SAME-S: bake the limiter ceiling so the chunkable decoders keep the ('latent','pcm') signature; add the SAME-S fp8 chunkable pair 775a981 verified
SAME-L/SAME-S: bake the limiter ceiling so the chunkable decoders keep the ('latent','pcm') signature; add the SAME-S fp8 chunkable pair aea1cf0 verified
SAME-L/SAME-S: bake the limiter ceiling so the chunkable decoders keep the ('latent','pcm') signature; add the SAME-S fp8 chunkable pair 9a01fee verified
SAME-S: add the chunkable bf16 pair (two profiles, limiter) baa3cd7 verified
SAME-S: add the chunkable bf16 pair (two profiles, limiter) 7c73bf8 verified
SAME-L encoders: low band 64 -> 256 (1.3-1.5x faster, cos 0.99831 -> 0.99969) 1a8702f verified
SAME-L encoders: low band 64 -> 256 (1.3-1.5x faster, cos 0.99831 -> 0.99969) d0d486b verified
SAME-L: remove miscalibrated enc_fp8.trt (activation scales 1.3-28x too small; superseded by enc_fp8_chunkable.trt on sm_90) 2c8b02a verified
SAME-L: remove miscalibrated enc_fp8.trt (activation scales 1.3-28x too small; superseded by enc_fp8_chunkable.trt on sm_90) e00128c verified
SAME-L: add enc_fp8_chunkable.trt (fp16, two-profile chunkable) 7036ca2 verified
SAME-L: add dec_fp8_chunkable_limiter.trt (fp16, two-profile chunkable, runtime limiter ceiling) 093c4d6 verified
SAME-L: add enc_fp16_chunkable.trt (fp16, two-profile chunkable) eb343c9 verified
SAME-L: add dec_fp16_chunkable_limiter.trt (fp16, two-profile chunkable, runtime limiter ceiling) 5dba282 verified
Delete old fp16mixed names + retired medium dit_bf16.trt (fp16 rename shipped) 48e34e2 verified
Rename fp16mixed->fp16 DiT/T5Gemma engines+onnx (copies; deletes follow after code merge) 3f29675 verified
Remove redundant w8_bf16 engines (int8 weights fold to bf16 at build β identical size/speed to the bf16 baseline) 63cc70d verified
Add sm_120 TRT engines (AOT SWA, wide profile) for the quantized tiers 4e8cc21 verified
Add sm_90 TRT engines (wide profile) for the quantized tiers d593f74 verified
Add fp8 tier for sm-music / sm-sfx DiTs (clean fp8-linear graft on fp16mixed) c3aa6e0 verified
Add fp8 tier for sm-music / sm-sfx DiTs (clean fp8-linear graft on fp16mixed) 1489a17 verified
sm_120 sa3-m: add the fp8 medium-DiT tier (176 fp8 GEMM + 96 bf16 fused MHA, 1.43-1.95x vs fp16mixed) (#7) f8e0a0f
sa3-m fp8: ship calibrated-bakedmin engine (176 fp8 GEMM + 96 bf16 fused MHA + baked fp32 RoPE; calibration by @ryanontheinside #47) 4697212 verified
sm_120 SAME-L: AOT (graph-capturable) engines + tensorRT/README.md (#6) a05f2d4
Upload tensorRT/sm_90/sa3-m/dit_fp8.trt with huggingface_hub feddb8b verified
Fix the bf16 medium DiT: bake RoPE's tables into the graph (clipping 3.112% -> 0.014%) (#5) 6152b1d
Rebuild medium DiT fp16-mixed engines (sm_90 + new sm_120): attention now fuses, 4.3x faster (#4) 06debc3
Add bf16 medium DiT TRT engine (sm_90) β FMHA-fused, medium default (#1) 9a38efb
Add FP32 SAME-L + SAME-S TRT decoders 773ac48
Cortexelus Claude Opus 4.7 (1M context) commited on
Decoder ONNX: clip+scale in FP32 before int32 cast a429ffa
Cortexelus Claude Opus 4.7 (1M context) commited on
DiT engines: replace BF16 with FP16-mixed (FP32 islands) + add FP32 variants 97651ef
Cortexelus Claude Opus 4.7 (1M context) commited on
DiT TRT engines: rebuild with profile min=1 (sm_90) 5f27021
Cortexelus Claude Opus 4.7 (1M context) commited on
T5Gemma: FP16-mixed (FP32 attention island) β fixes BF16 numerical bug 23f5c64
Cortexelus commited on
Decoder TRT engines: PCM-baked-in (sm_90) 106e6fd
Cortexelus commited on
Drop L-range suffix from SAME-L decoder filename 633bd26
Cortexelus commited on
Reorganize tensorRT/ into per-arch subdirs; rebuild DiT engines with conditioner baked in 08c1abe
Cortexelus commited on
Remove padding_embedding.pt β these ship with the inference code repo e8c6e3e
Cortexelus commited on
Remove pipeline_state from tensorRT/ β it ships with the inference code repo, not here f6ada03
Cortexelus commited on
Add pipeline_state + per-model padding_embedding for SA3 TRT pipeline 267e477
Cortexelus commited on