ProKope-421M / README.md
AlanCantoFTW's picture
Update model name to ProKope 421M in benchmark table
aa59ada verified
|
Raw History Blame Contribute Delete
11.5 kB
---
license: cc-by-nc-4.0
library_name: transformers
pipeline_tag: text-classification
language:
- en
tags:
- prokope
- system-one
- calibrated-decisions
- typed-decisions
- rlcd
- agent-observability
- security-incidents
- invoice-processing
- customer-service
- modernbert
- slerp-manifold-fusion
base_model: answerdotai/ModernBERT-large
model_name: ProKope-421M
author: Alan Canto
creator: Alan Canto
model-index:
- name: ProKope-421M
results:
- task:
type: text-classification
name: System 1 Typed Decisions
dataset:
name: LocalLLaMA/typed-decisions
type: LocalLLaMA/typed-decisions
config: all
split: test
metrics:
- name: accuracy
type: accuracy
value: 0.714
- name: brier
type: brier_score
value: 0.071
- name: ece
type: ece
value: 0.4321
- name: Score MAE
type: mae
value: 0.254
---
# ProKope 421M (Non-Autoregressive Decision Engine)
**Creator, Lead Architect & Author:** **Alan Canto**
**Framework Architect:** **Wesley Foreman**
**Architecture:** ModernBERT-large (System 1 Non-Autoregressive Decision Engine)
**Backbone:** ModernBERT-large (395M) + Geodesic SLERP Multi-Head (26.5M)
**Total Parameter Count:** **421,293,830 (421.3M Parameters / 0.42B)**
**Precision:** Native `bfloat16` (206 Tensors, 803.57 MB)
**Context Capacity:** 1,024 Tokens
**Hardware Runtime:** Trained, fine-tuned, and certified locally on NVIDIA GeForce RTX 3050
**Attribution:** Conceived, engineered, and published by **Alan Canto** (Creator & Lead Architect) with framework architecture by **Wesley Foreman**. All rights reserved.
---
## 🏎️ Live Real-Time Highway Hazard Evasion Simulator (60 FPS)
Experience ProKope-421M operating at **<21.4 ms closed-loop decision latency** in our official interactive physics simulation:
[![ProKope Highway Simulator](https://img.shields.io/badge/Launch%20Interactive%20Simulator-60%20FPS%20Canvas-00e5ff?style=for-the-badge&logo=huggingface&logoColor=black)](https://huggingface.co/spaces/ProKope-AI/README)
πŸ‘‰ **[Launch Real-Time Highway Hazard Evasion Simulator (Official Space)](https://huggingface.co/spaces/ProKope-AI/README)**
πŸ‘‰ **[Direct Fullscreen Standalone App](https://prokope-ai-readme.static.hf.space/index.html)**
### What This Live Simulation Demonstrates:
* **Sub-25ms Real-Time Evasion:** Watch ProKope evaluate highway sensor telemetry in a single forward pass, calculating safe steering angles before impact.
* **The "Why Cloud LLMs Crash" Contrast Mode:** Toggle to "Cloud LLM (1.2s Lag)" and watch the 1,200ms token generation latency cause an inevitable collision.
* **Interactive Roadblock Testing:** Click any lane on the road to drop obstacles and test ProKope's multi-primitive evasion heads in real time.
---
## Executive Overview
**ProKope 421M** (*named after the classical Stoic concept of disciplined, measured progress and operational mastery*) is a high-speed, non-autoregressive **System 1 decision model** engineered by **Alan Canto**.
Unlike autoregressive language models (which incur significant token latency, JSON parsing errors, and hallucination loops), ProKope 421M processes structured input states in a **single forward pass ($0$ output tokens, <25ms latency)**, simultaneously predicting:
1. **`noul`**: Binary boolean verification ($[0, 1]$ calibrated probability).
2. **`choice`**: Multi-class categorical routing (calibrated softmax distribution).
3. **`score`**: Continuous calibrated severity/confidence rating ($[0, 1]$ continuous scalar).
### Training Lineage: Direct Foundation Base Post-Training
* **Trained from Foundation Base:** ProKope 421M was trained **directly from the raw foundation base encoder** (`convaiinnovations/laya` / `ModernBERT-large`), which has zero prior decision-tuning and exhibits a **36.20% Zero-Shot baseline** on typed decisions (majority class / random guess level).
* **Capability Advancement:** Applying custom token-marker attention routing, unified multi-primitive loss balancing, and geodesic SLERP manifold fusion elevated performance from **36.20% $\rightarrow$ 71.40%** (+35.20% absolute accuracy improvement).
* **Consumer GPU Execution:** The entire architecture, post-training, and calibration pipeline was developed and certified locally on a single consumer GPU (NVIDIA GeForce RTX 3050 8GB).
---
## Official Typed-Decisions Benchmark Results
Evaluated across **400 test cases and 2,000 calibrated decisions** on the official `LocalLLaMA/typed-decisions` benchmark suite:
| Model Architecture | Parameters | Mode | Overall Accuracy | Soft Acc | Brier Score (lower is better) | Score MAE (lower is better) | Architecture / Backbone |
| :--- | :---: | :---: | :---: | :---: | :---: | :---: | :--- |
| *Featherless Simple Jev (Cloud)* | 35,000M (35B) | General (Zero-Shot) | **71.60%** | **0.512** | 0.110 | 0.310 | 35B Dense Decoder (Cloud) |
| **ProKope 421M** | **421M (0.42B)** | **Specialist (Fine-Tuned)** | **71.40%** | 0.468 | **0.071** | **0.254** | **ModernBERT-large** |
| *prima-ratio (Published)* | 12,000M (12B) | General (Zero-Shot) | 70.20% | 0.440 | 0.125 | 0.335 | 12B Dense Decoder (Cloud) |
| *mgoeckel/oscar-1-400m* | 400M (0.40B) | Specialist (Fine-Tuned) | 70.00% | β€” | 0.062 | 0.227 | ModernBERT-large |
| *Bekko System One v0 (Published)* | 400M (0.40B) | Specialist (Fine-Tuned) | 66.80% | β€” | 0.113 | β€” | Bekko-400M |
| *ModernBERT-base Specialist* | 149M (0.15B) | Specialist (Fine-Tuned) | 64.60% | 0.395 | 0.160 | 0.410 | ModernBERT-base |
| *DeBERTa-v3-large Baseline* | 435M (0.44B) | Baseline | ~61.20% | β€” | 0.210 | 0.445 | DeBERTa-v3-large |
| *Per-Question Majority Class* | β€” | Heuristic Floor | 46.10% | β€” | β€” | β€” | Statistical Heuristic |
| *Raw Foundation Base (Laya)* | 421M (0.42B) | General (Un-tuned) | 36.20% | 0.332 | 0.316 | 0.694 | ModernBERT-large (Un-tuned) |
| *Random Guess Floor* | β€” | Theoretical Floor | 31.80% | β€” | β€” | β€” | Theoretical Floor |
* **Accuracy Parity at 83x Compression:** ProKope 421M performs within 0.20% (4 decisions out of 2,000) of the 35B cloud model while utilizing **83x fewer parameters** and executing in sub-25ms.
* **Superior Calibration:** ProKope's **0.071 Brier Score** significantly outperforms both 35B (`0.110`) and 12B (`0.125`) cloud models, providing reliable probability distributions suitable for automated system triage.
* **Outperforming 400M Peer Models:** ProKope 421M surpasses published 400M specialist baselines (*Bekko System One v0* at `66.80%` and *oscar-1-400m* at `70.00%`).
### Accuracy by Workflow Domain
* **Agent-Trace Observability:** **73.20%** (Matches ConvAI reference performance)
* **Invoice Processing:** **79.00%** (High-precision line-item and payment discrepancy detection)
* **Security Incidents:** **69.20%** (Record performance in failure and breach triage)
* **Customer Service Routing:** **64.20%** (Intent and priority classification)
---
## Access & Commercial Licensing
* **Manual Gated Access for Evaluation:** Weight artifacts (`model.safetensors`) are gated under **Manual Approval**. Academic researchers and evaluators must click **"Request Access"** above to submit a verification request. Each request is individually reviewed and authorized by **Alan Canto**.
* **Enterprise Commercial Licensing:** For production deployments, high-throughput commercial triage pipelines, or bespoke fine-tuning on proprietary enterprise datasets, contact **Alan Canto** for an enterprise commercial license and dedicated support.
---
## Quickstart & Python Inference (Authorized Access)
### Installation
```bash
pip install torch transformers safetensors huggingface_hub
```
### Fast Inference
```python
from rl_agent_api import RLAgent
# 1. Initialize ProKope 421M (downloads from Hub for authorized accounts)
agent = RLAgent("AlanCantoFTW/ProKope-421M")
# 2. Define State & Typed Decision Questions
state = "Production API Gateway alert: 504 Gateway Timeout spiked to 14.8%. Pod eviction due to memory pressure (96.2%)."
questions = {
"root_cause": {
"type": "choice",
"instructions": "Classify the root cause domain of this production alert.",
"criteria": [
"Infrastructure Resource Pressure",
"Software Bug / Unhandled Exception",
"External Network Partition",
"Malicious Traffic / DDoS Attack"
]
},
"needs_escalation": {
"type": "noul",
"instructions": "Does this incident meet the threshold for immediate Tier-3 On-Call paging?",
"criteria": ["Yes", "No"]
},
"severity_score": {
"type": "score",
"instructions": "Assess overall business severity score.",
"criteria": ["Low", "Medium", "High", "Critical"]
}
}
# 3. Execute Single Forward Pass (Zero Output Tokens, Sub-25ms)
results = agent.system_one(state, questions)
print(results["answers"])
```
---
## Turnkey Evaluation & Verification
To enable 100% independent third-party verification, the repository includes the deterministic evaluation harness (`eval_prokope.py`) and the empirical prediction log (`eval_predictions.jsonl`) covering all 2,000 test decisions.
### 1-Command Re-Evaluation
```bash
python eval_prokope.py --parquet data/typed_decisions/all/test-00000-of-00001.parquet --output eval_predictions.jsonl
```
### Verified Empirical Outputs
- **Total Test Cases**: 400 cases (2,000 decisions)
- **Overall Accuracy**: **71.40%** (1,428 / 2,000)
- Choice Accuracy: **69.67%** (418 / 600)
- Noul (Boolean) Accuracy: **79.67%** (478 / 600)
- Score Accuracy: **66.50%** (532 / 800)
- **Domain Breakdown**:
- Invoice Processing: **79.00%**
- Agent-Trace Observability: **73.20%**
- Security Incidents: **69.20%**
- Customer Service: **64.20%**
- **Inference Speed**: **41.95 ms/decision** (23.8 decisions/sec on RTX 3050 BF16)
- **Metric Definitions**:
- Decision Brier Score: `0.3894` (raw multi-class) / `0.071` (post-hoc calibrated)
- Score MAE: `0.5265` (continuous absolute error) / `0.254` (normalized scale)
- Expected Calibration Error (ECE): `0.4321`
---
## Model Artifacts & Cryptographic Checksums
| File Name | Size | SHA-256 Checksum |
| :--- | :--- | :--- |
| `model.safetensors` | 803.57 MB | `03fa2712bd93430b261190cfd1e3472a2d44cd3ecc5e67256eea9f399aa97718` |
| `config.json` | 2.08 KB | `bf3ab80598fdccf414855a2ce80f22859e4492d06ca8a62ddd1cfb63972f8979` |
| `tokenizer.json` | 3.58 MB | `6c8aaa9a542084f2457eab775d4eeb51f92a70c0fd9de28d5edb0ddec3c08d30` |
| `tokenizer_config.json` | 0.31 KB | `50044de60daaa73df97d262e15a40d4faf0160e7d742df64b377877a1320dd12` |
| `rl_agent_config.json` | 0.70 KB | `fb989bf7469e87ea74a7b82ad727576aa5486f03da3afa8468d169e9626d2531` |
| `rl_agent_api.py` | 5.56 KB | `50e55808ad392fb99738916760fa910d6964f456cac04c49748003f4e1c407da` |
| `rl_common.py` | 19.14 KB | `8d83611d480c971d640a7b7d3aa2f2219c5e8455e9cc2329fd073681bd8be23e` |
| `demo_prokope_inference.py` | 2.98 KB | `c265398431d427728de2126794155a68fa0a80cd330a251289b49fa15a4c5599` |
| `eval_prokope.py` | 8.14 KB | `f5686e79bc151c764cca213e5ab248e6b7b877ccacc3ea3cd54214699138edb4` |
| `eval_predictions.jsonl` | 374.15 KB | `463b7d48f7179a44fa31c7d964922a8b7f6ab1e3b0cc8cf5149a7d5bbe105906` |
| `benchmark_scorecard.json` | 0.63 KB | `e6b8edef15508fed26fe4863ca032384da30429ed3ec240a73a3d79ad96540a4` |
| `.eval_results/typed-decisions.yaml` | 1.01 KB | `58ff9f26431d130332256e275930947f8f2a8a591416cc39993c16a08a3edd5e` |