Weight-only INT4 for RTX 3090/4090-class cards, where FP8 and NVFP4 are emulated or unusable. Full recipe on every card.
Alex Αdamopoulos
aleada
·
AI & ML interests
Verified open-weight inference for consumer GPUs. W4A16 packs, reproducible benchmarks and tools that inspect what quantized models actually contain.
Recent Activity
updated a collection 20 days ago
W4A16 for consumer Ampere updated a collection 20 days ago
W4A16 for consumer Ampere updated a model 20 days ago
aleada/Nemotron-3-Nano-4B-W4A16Organizations
None yet