CPM-jev / README.md
link921's picture
Add complete v0.1 Research Preview model card
4dee0fc verified
|
Raw History Blame Contribute Delete
8.74 kB
---
license: apache-2.0
base_model:
- openbmb/MiniCPM5-2B-Base
datasets:
- SargeDev/jev-distill-corpus-v3
pipeline_tag: text-classification
tags:
- minicpm
- minicpm5
- system-one
- decision-model
- jev-style
- probability
- calibration
- agent
- routing
- text-classification
---
# CPM-jev
**Version:** `v0.1-research-preview`
> An experimental Jev-style / System-One decision model based on MiniCPM5-2B-Base. This is an independent community project and is not affiliated with or endorsed by TypeSafe AI or OpenBMB.
CPM-jev is an independently developed **LoRA fine-tune of `openbmb/MiniCPM5-2B-Base`**, with an added decision head for scoring candidate actions. It is a research preview, **not a chat model**, **not production-ready**, and should not use `generate()` as its primary decision interface.
## Model Description
Given a state, a question, and candidate options, CPM-jev assigns one scalar score to each option and normalizes those scores into a probability distribution:
```text
state + question + candidate options
-> decision scorer
-> probability distribution
```
The probabilities are useful for ranking, routing, selective prediction, and confidence-aware escalation. They must not be interpreted as universally calibrated real-world probabilities.
Base model weights are not duplicated in this repository. Download `openbmb/MiniCPM5-2B-Base` separately.
## Architecture
- Backbone: [`openbmb/MiniCPM5-2B-Base`](https://huggingface.co/openbmb/MiniCPM5-2B-Base)
- Adaptation: LoRA, rank 16, alpha 32, dropout 0.05
- LoRA targets: `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj`
- Decision head: FP32 linear scalar head over the last non-padding token
- Candidate normalization: raw softmax across the options for one question
- Maximum sequence length: 512 tokens, left truncation
Each candidate is encoded independently with the same prompt template. The scalar scores are only comparable among the options supplied in the same call.
## Training Data
CPM-jev was trained on the `train` split of [`SargeDev/jev-distill-corpus-v3`](https://huggingface.co/datasets/SargeDev/jev-distill-corpus-v3), using exactly **655,806 samples**. Probability targets in that dataset include data distilled from a Jev teacher.
No samples from `validation`, `calibration`, `test`, `test_set_30k`, or `ood` were used for training. Those splits were reserved for model selection, calibration analysis, final evaluation, and OOD evaluation as appropriate.
## Training Procedure
- Objective: soft-target cross-entropy over candidate distributions
- Epochs: 1
- Optimizer steps: 40,988
- Effective batch size: 16 (micro-batch 4, gradient accumulation 4)
- Learning rate: `1e-4`
- Weight decay: `0.01`
- Warmup ratio: `0.03`
- Seed: 42
- Maximum length: 512
- Trainable parameters: 25,118,721 (about 1.10% of total)
The final weights are the completed Stage 3 run. No evaluation split was mixed into training.
## Evaluation
The primary final benchmark is `jev-distill-corpus-v3/test_set_30k`, containing **29,955 samples**. Reported values use raw probabilities.
| Metric | Result |
| --- | ---: |
| Accuracy | 0.878317 |
| NLL | 0.727553 |
| Brier score | 0.011975 |
| ECE | 0.212610 |
| Signed ECE | -0.212610 |
| Correct mean confidence | 0.692363 |
| Incorrect mean confidence | 0.473300 |
Compared with the Stage 2 100k rebaseline on the same benchmark:
| Metric | Absolute change |
| --- | ---: |
| Accuracy | +4.934 percentage points |
| NLL | -0.034981 |
| Brier score | -0.018019 |
| ECE | +0.034864 |
Accuracy improved, but calibration error worsened. This trade-off is important when deciding whether to abstain or escalate.
## Selective Accuracy
| Threshold | Coverage | Accuracy | Risk |
| --------- | -------- | -------- | ----- |
| >=0.70 | 41.92% | 99.68% | 0.32% |
| >=0.80 | 27.98% | 99.80% | 0.20% |
| >=0.90 | 16.03% | 99.90% | 0.10% |
These figures come from `jev-distill-corpus-v3/test_set_30k`. They do **not** imply the same accuracy on arbitrary real-world tasks, distributions, prompts, option sets, or deployment environments.
## Calibration
Temperature scaling fitted on the calibration split produced `T = 1.008740`. It did not improve overall ECE consistently across `validation`, `calibration`, `test`, and `test_set_30k`, so the release defaults to **raw probabilities** (`use_temperature: false`).
The negative signed ECE on the primary in-domain benchmark indicates that the model is substantially under-confident there. High ECE remains a central limitation.
## OOD Evaluation
On the held-out `ood` split (13,058 samples), raw probabilities produced:
| Metric | Result |
| --- | ---: |
| Accuracy | 0.878542 |
| NLL | 0.881370 |
| Brier score | 0.196659 |
| ECE | 0.093474 |
| Signed ECE | +0.093387 |
The positive signed ECE and much larger Brier score show a different, over-confident OOD failure mode. OOD calibration is materially weaker than the in-domain selective-accuracy table suggests.
## Intended Use
Research and prototyping uses include:
- agent tool routing
- model routing
- retry and recovery decisions
- next-action selection
- result judging
- confidence-aware escalation
- System-1 / System-2 routing
The model is intended to compare explicitly supplied options. Applications should define abstention and human-escalation policies and validate them on their own workload.
## Limitations
1. The model is clearly under-confident on the reported in-domain evaluation.
2. ECE remains high.
3. OOD calibration is materially weaker than in-domain calibration and exhibits over-confidence.
4. OOD metrics are accuracy 0.878542, NLL 0.881370, Brier 0.196659, ECE 0.093474, and signed ECE +0.093387.
5. The model must not be described as fully calibrated.
6. Current evidence supports selective decision and routing use more strongly than treating its probability as an absolute real-world probability.
7. Do not directly use its probabilities for automated medical, financial, legal, safety-critical, or other high-risk decisions.
8. Results are specific to the released corpus and evaluation procedure; independent external evaluation is still needed.
9. The model is not a chat model and `generate()` is not its decision interface.
## Installation
```bash
git clone https://huggingface.co/link921/CPM-jev
cd CPM-jev
python -m venv .venv
# Windows: .venv\Scripts\activate
# Linux/macOS: source .venv/bin/activate
pip install -r requirements.txt
```
The first load downloads the base model unless it is already cached. You can pass a local base-model directory with `base_model=...` or `--base-model ...`.
## Inference Example
```python
from inference import DecisionModel
model = DecisionModel(".")
result = model.decide(
state="The previous tool call failed twice.",
question="What should the agent do next?",
options=[
"retry",
"switch_tool",
"ask_user",
],
)
print(result)
```
Output schema (illustrative values):
```json
{
"options": [
"retry",
"switch_tool",
"ask_user"
],
"probabilities": [
0.08,
0.84,
0.08
],
"choice": "switch_tool",
"confidence": 0.84
}
```
Run the included example:
```bash
python example.py
```
Or use the CLI:
```bash
python inference.py --model-dir . --state "The previous tool call failed twice." --question "What should the agent do next?" --options retry switch_tool ask_user
```
Do not call `generate()` to obtain the primary decision. CPM-jev compares candidate scores and applies softmax across the supplied options.
## Citation
If this research preview is useful, cite the repository and the upstream model and dataset:
```bibtex
@software{cpm_jev_2026,
title = {CPM-jev: A MiniCPM5-2B Jev-style Decision Model},
year = {2026},
version = {v0.1-research-preview},
url = {https://huggingface.co/link921/CPM-jev}
}
```
## License
This repository is released under the Apache License 2.0. The upstream base model and dataset are currently marked Apache-2.0 on Hugging Face. Users remain responsible for reviewing upstream licenses, dataset content, and applicable laws for their intended use.
## Acknowledgements
- Thanks to [OpenBMB](https://huggingface.co/openbmb) for [`MiniCPM5-2B-Base`](https://huggingface.co/openbmb/MiniCPM5-2B-Base).
- Thanks to [SargeDev](https://huggingface.co/SargeDev) for [`jev-distill-corpus-v3`](https://huggingface.co/datasets/SargeDev/jev-distill-corpus-v3).
- The corpus contains probability targets that include data distilled from a Jev teacher. This release is an independent community project and does not claim official status or endorsement from TypeSafe AI, OpenBMB, or the dataset authors.