Local AI, LLMs, multimodal models, and generative AI. Interested in running and optimizing models locally, model quantization, inference performance, long-context LLMs, uncensored models, and GPU acceleration.
Exploring Qwen, ComfyUI, image/video generation, vLLM, llama.cpp, Ollama, SGLang, CUDA, and NVIDIA GPUs — with a strong focus on getting the most performance out of consumer hardware and self-hosted infrastructure.
Cloud, Kubernetes, DevOps/SRE, homelabs, and building practical AI infrastructure are also part of the fun.