AI & ML interests

Reference implementations of LLM inference at the metal — Gemma 4 and Llama on CUDA, built from explicit C++23 components you can read and understand. Runs on consumer hardware.

Recent Activity

toddt  updated a model 16 days ago
mila-llm/gpt2-small
toddt  updated a model 16 days ago
mila-llm/Llama-3.1-8B-Instruct-fp4
toddt  updated a model 16 days ago
mila-llm/Llama-3.2-3B-Instruct-fp4
View all activity