I work on benchmark-driven model evaluation, scientific ML, and structured reasoning challenges, including ARC-AGI, model reasoning, and protein function prediction.
My focus is on building testable pipelines, studying failure modes, and understanding where models genuinely generalize versus where they only appear to pattern-match.
Product leader and applied ML practitioner.