Thor Lin
coolthor
AI & ML interests
On-device & edge LLM inference on NVIDIA GB10 / DGX Spark.
Quantization (NVFP4 W4A4/W4A16, FP8), vLLM serving, speculative
decoding (EAGLE-3), and multimodal/omni models. Measure-first
benchmarking — I publish the numbers, including the ones that fail.
Recent Activity
updated a model 5 days ago
coolthor/Sulphur-2-distilled-Q4_K_M-GGUF updated a model 5 days ago
coolthor/Laguna-S-2.1-attnQ8-DFlash-Q4-GGUF published a model 5 days ago
coolthor/Sulphur-2-distilled-Q4_K_M-GGUFOrganizations
None yet