Spaces:
Running
Running
File size: 1,797 Bytes
4b258eb ce58634 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 | ---
title: README
emoji: 🏦
colorFrom: gray
colorTo: blue
sdk: static
pinned: false
---
# Dissei Data
**Outcome-graded RL environments for institutional financial judgment.**
Dissei builds reinforcement-learning environments from real institutional credit
decisions — each one a documented transaction rebuilt as an agentic task, worked with
analyst tools and graded against what actually happened.
- **Point-in-time by construction.** A model sees only what was knowable at the decision
date; every exhibit is anchor-gated, so hindsight is excluded structurally — not by an
instruction it can ignore.
- **Graded, not checked.** Reward decomposes into fail-closed critical gates, an
expert-authored graded rubric, a contradiction penalty, and a retrieval modulator that
rewards grounding in the record.
- **Grading automation, disclosed.** Rubrics are authored and anchored by credit
practitioners who carried real risk; execution runs on LLM judges at temperature 0 that
apply the sealed rubric and never invent criteria. Answers are graded twice — against
best practice, and against the realized outcome.
- **Sealed by architecture.** Answer material is time-locked and anchor-gated; every asset
ships with a per-asset contamination statement and right-to-train lineage.
- **Operators, not an annotation farm.** The corpus is the by-product of an operating
lending business — incentive-bearing decisions with real consequences, not opinions
commissioned for a dataset.
Environments (RL post-training) and sealed evaluations are available to frontier labs
**under agreement**. We do not publish tasks, rubrics, graders, or leaderboards.
**Methods & findings:** https://dissei.ai/research
**Contact:** info@dissei.credit · Mutual NDA before anything sensitive.
|