Text Classification
Transformers
Safetensors
deberta-v2
process-reward-model
reasoning-verification
step-verification
crp
context-relay-protocol
Eval Results (legacy)
text-embeddings-inference
Instructions to use AutoCyberAI/crp-prm-deberta-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AutoCyberAI/crp-prm-deberta-v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="AutoCyberAI/crp-prm-deberta-v1")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("AutoCyberAI/crp-prm-deberta-v1") model = AutoModelForSequenceClassification.from_pretrained("AutoCyberAI/crp-prm-deberta-v1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
CRP Process Reward Model (PRM) โ DeBERTa-v3-large
A sequence-pair classifier that scores whether a reasoning step is entailed by / consistent with its prior steps and the original problem. Returns VALID or INVALID. Used by crp.vr.prm.ProcessRewardVerifier as an advisory step-level judge inside the CRP Verification Relay.
Model description
- Architecture:
microsoft/deberta-v3-largesequence classification. - Labels:
VALID,INVALID. - Held-out AUC: 0.793 (independent step-level harness on prm800k test steps).
- Curated reasoning cases: 8/10 correct.
- Training data: PRM800K, Math-Shepherd, MMLU-Pro-CoT, RLHFlow/Mistral-PRM-Data, and agentic synthetic examples.
- Inference budget: 400 ms on CPU; degrades to
UNKNOWNif exceeded.
Intended use
from transformers import pipeline
pm = pipeline('text-classification', model='AutoCyberAI/crp-prm-deberta-v1', top_k=None)
text = 'premises: Server returned 502. The load balancer health check is failing. [SEP] step: The database is the root cause.'
print(pm(text)) # [{'label': 'VALID', 'score': ...}]
Limitations
- Trained primarily on math and multiple-choice reasoning data; transfer to arbitrary agentic steps is a work in progress.
- Wired as an advisory scorer in CRP โ hard INVALID gating remains with symbolic verifiers and checkpoints.
- The exported
prm_threshold(0.675) was calibrated on a training-matched mix and does not necessarily transfer to real-world distributions; tune viaCRP_PRM_THRESHOLD.
Citation
@misc{crp-prm-deberta-v1,
title={{CRP Process Reward Model}},
author={{AutoCyber AI}},
year={2026},
howpublished={\url{https://huggingface.co/AutoCyberAI/crp-prm-deberta-v1}}
}
This model is part of the Context Relay Protocol (CRP) v6 Phase A managed-model suite. Learn more at https://crprotocol.io.
- Downloads last month
- 23
Model tree for AutoCyberAI/crp-prm-deberta-v1
Base model
microsoft/deberta-v3-largeEvaluation results
- ROC AUC (held-out, independent harness) on prm800k held-out test stepsself-reported0.793