CRP Process Reward Model (PRM) โ€” DeBERTa-v3-large

A sequence-pair classifier that scores whether a reasoning step is entailed by / consistent with its prior steps and the original problem. Returns VALID or INVALID. Used by crp.vr.prm.ProcessRewardVerifier as an advisory step-level judge inside the CRP Verification Relay.

Model description

  • Architecture: microsoft/deberta-v3-large sequence classification.
  • Labels: VALID, INVALID.
  • Held-out AUC: 0.793 (independent step-level harness on prm800k test steps).
  • Curated reasoning cases: 8/10 correct.
  • Training data: PRM800K, Math-Shepherd, MMLU-Pro-CoT, RLHFlow/Mistral-PRM-Data, and agentic synthetic examples.
  • Inference budget: 400 ms on CPU; degrades to UNKNOWN if exceeded.

Intended use

from transformers import pipeline
pm = pipeline('text-classification', model='AutoCyberAI/crp-prm-deberta-v1', top_k=None)
text = 'premises: Server returned 502. The load balancer health check is failing. [SEP] step: The database is the root cause.'
print(pm(text))  # [{'label': 'VALID', 'score': ...}]

Limitations

  • Trained primarily on math and multiple-choice reasoning data; transfer to arbitrary agentic steps is a work in progress.
  • Wired as an advisory scorer in CRP โ€” hard INVALID gating remains with symbolic verifiers and checkpoints.
  • The exported prm_threshold (0.675) was calibrated on a training-matched mix and does not necessarily transfer to real-world distributions; tune via CRP_PRM_THRESHOLD.

Citation

@misc{crp-prm-deberta-v1,
  title={{CRP Process Reward Model}},
  author={{AutoCyber AI}},
  year={2026},
  howpublished={\url{https://huggingface.co/AutoCyberAI/crp-prm-deberta-v1}}
}

This model is part of the Context Relay Protocol (CRP) v6 Phase A managed-model suite. Learn more at https://crprotocol.io.

Downloads last month
23
Safetensors
Model size
0.4B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for AutoCyberAI/crp-prm-deberta-v1

Finetuned
(311)
this model

Evaluation results

  • ROC AUC (held-out, independent harness) on prm800k held-out test steps
    self-reported
    0.793