Trained Verifier Models
Surrogate code verifiers across three model sizes trained using multiple different algorithms as described in the Aletheia paper
Text Generation • 2B • Updated • 27Note Our flagship verifiers, trained using a GRPO-style algorithm with a 16k generation limit
Aletheia-Bench/GRPO-Think-7B-16k
Text Generation • 8B • Updated • 26Note Our flagship verifiers, trained using a GRPO-style algorithm with a 16k generation limit
Aletheia-Bench/GRPO-Think-14B-16k
Text Generation • 15B • Updated • 19Note Our flagship verifiers, trained using a GRPO-style algorithm with a 16k generation limit
Aletheia-Bench/GRPO-Think-1.5B-4k
Text Generation • 2B • Updated • 18Note A variant of the flagship verifiers trained with a 4k generation limit
Aletheia-Bench/GRPO-Think-7B-4k
Text Generation • 8B • Updated • 27Note A variant of the flagship verifiers trained with a 4k generation limit
Aletheia-Bench/GRPO-Think-14B-4k
Text Generation • 15B • Updated • 24Note A variant of the flagship verifiers trained with a 4k generation limit
Aletheia-Bench/GRPO-Think-1.5B-8k
Text Generation • 2B • Updated • 44Note A variant of our flagship verifiers, trained with a 8k generation limit
Aletheia-Bench/GRPO-Think-7B-8k
Text Generation • 8B • Updated • 30Note A variant of our flagship verifiers, trained with a 8k generation limit
Aletheia-Bench/GRPO-Think-14B-8k
Text Generation • 15B • Updated • 34 • 1Note A variant of our flagship verifiers, trained with a 8k generation limit
Aletheia-Bench/GRPO-Instruct-1.5B
Text Generation • 2B • Updated • 26Note A GRPO training ablation that does not generate thinking traces
Aletheia-Bench/GRPO-Instruct-7B
Text Generation • 8B • Updated • 27Note A GRPO training ablation that does not generate thinking traces
Aletheia-Bench/GRPO-Instruct-14B
Text Generation • 15B • Updated • 23Note A GRPO training ablation that does not generate thinking traces
Aletheia-Bench/DPO-Think-1.5B
Text Generation • 2B • Updated • 25Note A verifier trained completely offline using DPO
Aletheia-Bench/DPO-Think-7B
Text Generation • 8B • Updated • 22Note A verifier trained completely offline using DPO
Aletheia-Bench/DPO-Think-14B
Text Generation • 15B • Updated • 45 • 2Note A verifier trained completely offline using DPO
Aletheia-Bench/BatchOnline-GRPO-1.5B
Text Generation • 2B • Updated • 36Note A variant of our flagship verifiers, where the generation policy is synced every 4 gradient updates
Aletheia-Bench/BatchOnline-GRPO-7B
Text Generation • 8B • Updated • 29 • 1Note A variant of our flagship verifiers, where the generation policy is synced every 4 gradient updates
Aletheia-Bench/BatchOnline-GRPO-14B
Text Generation • 15B • Updated • 26 • 1Note A variant of our flagship verifiers, where the generation policy is synced every 4 gradient updates
Aletheia-Bench/RAFT-1.5B
Text Generation • 2B • Updated • 21Note A verifier trained using on-policy rejection sampling on only positive samples
Aletheia-Bench/RAFT-7B
Text Generation • 8B • Updated • 28Note A verifier trained using on-policy rejection sampling on only positive samples
Aletheia-Bench/RAFT-14B
Text Generation • 15B • Updated • 23Note A verifier trained using on-policy rejection sampling on only positive samples
-
Aletheia: What Makes RLVR For Code Verifiers Tick?
Paper • 2601.12186 • Published • 1