Thor Lin
coolthor
·
AI & ML interests
On-device & edge LLM inference on NVIDIA GB10 / DGX Spark.
Quantization (NVFP4 W4A4/W4A16, FP8), vLLM serving, speculative
decoding (EAGLE-3), and multimodal/omni models. Measure-first
benchmarking — I publish the numbers, including the ones that fail.
Recent Activity
liked a model 5 days ago
convaiinnovations/laya updated a model 11 days ago
coolthor/MiniMax-H3-pruned-NVFP4 updated a model 11 days ago
coolthor/H3-Super-Acceleration-TuringOrganizations
None yet