# RLVR baseline checkpoint pi1, grpo, semantic step 800. Historical pi1 pure GRPO; matches per-checkpoint evaluation provenance. Pure GRPO/RLVR baseline (Teacher rejection disabled). This repository contains inference weights and tokenizer files, not optimizer state.