File size: 1,797 Bytes
4b258eb
 
 
 
 
 
 
 
 
ce58634
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
---
title: README
emoji: 🏦
colorFrom: gray
colorTo: blue
sdk: static
pinned: false
---

# Dissei Data
**Outcome-graded RL environments for institutional financial judgment.**

Dissei builds reinforcement-learning environments from real institutional credit
decisions — each one a documented transaction rebuilt as an agentic task, worked with
analyst tools and graded against what actually happened.

- **Point-in-time by construction.** A model sees only what was knowable at the decision
  date; every exhibit is anchor-gated, so hindsight is excluded structurally — not by an
  instruction it can ignore.
- **Graded, not checked.** Reward decomposes into fail-closed critical gates, an
  expert-authored graded rubric, a contradiction penalty, and a retrieval modulator that
  rewards grounding in the record.
- **Grading automation, disclosed.** Rubrics are authored and anchored by credit
  practitioners who carried real risk; execution runs on LLM judges at temperature 0 that
  apply the sealed rubric and never invent criteria. Answers are graded twice — against
  best practice, and against the realized outcome.
- **Sealed by architecture.** Answer material is time-locked and anchor-gated; every asset
  ships with a per-asset contamination statement and right-to-train lineage.
- **Operators, not an annotation farm.** The corpus is the by-product of an operating
  lending business — incentive-bearing decisions with real consequences, not opinions
  commissioned for a dataset.

Environments (RL post-training) and sealed evaluations are available to frontier labs
**under agreement**. We do not publish tasks, rubrics, graders, or leaderboards.

**Methods & findings:** https://dissei.ai/research
**Contact:** info@dissei.credit · Mutual NDA before anything sensitive.