Running Reproduction: Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals ๐ฏ Explore code logs, traces, and workspace in an interactive logbook