Local LLM benchmarking and model tuning on NVIDIA DGX Spark. Quantisation, speculative decoding, abliteration.