# RLVR baseline checkpoint dsr128, grpo, semantic step 800. Original DSR128 pure GRPO; weight hashes from historical single-answer evaluation protocol. Pure GRPO/RLVR baseline (Teacher rejection disabled). This repository contains inference weights and tokenizer files, not optimizer state.