aayushiii18's picture
Upload README.md
0d43549 verified
|
Raw History Blame Contribute Delete
3.52 kB
---
license: apache-2.0
tags:
- legal
- bert
- legal-bert
- contract-risk
- cuad
- pytorch
pipeline_tag: text-classification
library_name: transformers
---
# RiskLoop Representative Model Checkpoints
This repository hosts the **reproduced representative model checkpoints** for [RiskLoop](https://github.com/aayushiii18/RiskLoop), an NLP system for automated contract risk detection based on Legal-BERT fine-tuning on the CUAD (Contract Understanding Atticus Dataset) benchmark.
> [!IMPORTANT]
> **Provenance & Reproduction Note:**
> These checkpoints are **reproduced representative checkpoints** created under strictly controlled, frozen training conditions (Condition A) and evaluated in RiskLoop's frozen Phase 6 test evaluation.
> They are **NOT** the original historical official checkpoints from earlier development phases, as those original historical binaries were lost. These reproduced models represent the exact frozen representative checkpoints used to generate the Phase 6 decision analysis and final test metrics.
---
## 1. Model Architecture & Training Details
- **Backbone Model:** `nlpaueb/legal-bert-base-uncased`
- **Model Type:** Single-task span classification models (QA-style start/end logit heads over transformer sequence outputs).
- **Training Strategy:** Condition A (single-task fine-tuning with fixed seeds, 4 epochs, max sequence length 512, document stride 256).
---
## 2. Model Checkpoint Registry & Integrity Hashes
The repository contains three frozen model binary files (`.pt` PyTorch state dictionaries):
| Checkpoint Filename | Target Contract Task | Experimental Condition | Random Seed | File Size (Bytes) | SHA-256 Hash |
| :--- | :--- | :--- | :--- | :--- | :--- |
| `run_03_best_model.pt` | **Cap On Liability** | Condition A | `44` | 438,001,498 | `8689ef00a0718713f4ae8bdf8b48d3441b5d9a2ccd13ae322185862551c6f674` |
| `run_04_best_model.pt` | **Anti-Assignment** | Condition A | `42` | 438,001,498 | `ea6a75254798342e0cc3925dd8adca1ba9b5bb60c6a631ddf14057748c01947c` |
| `run_07_best_model.pt` | **Termination For Convenience** | Condition A | `42` | 438,001,498 | `a8a28902ef3dbc28e706135e8599f01cfa8c0e9c4c05cc081818e3027deda119` |
---
## 3. Frozen Phase 6 Test Evaluation Metrics
The reproduced representative checkpoints achieved the following metrics on the held-out frozen CUAD test evaluation set (Phase 6):
| Task Name | Representative Checkpoint | Test F1 Score | Test ROC-AUC | Test PR-AUC |
| :--- | :--- | :--- | :--- | :--- |
| **Cap On Liability** | `run_03_best_model.pt` | `0.7727273` | `0.9960492` | `0.9147721` |
| **Anti-Assignment** | `run_04_best_model.pt` | `0.8083624` | `0.9944473` | `0.9146714` |
| **Termination For Convenience** | `run_07_best_model.pt` | `0.7483871` | `0.9939074` | `0.7926948` |
---
## 4. Inference Score Interpretation
In RiskLoop's inference engine and Streamlit demo interface:
- Output scores are **uncalibrated raw logit deltas** ($\text{logit}_1 - \text{logit}_0$).
- High positive scores indicate strong model activation for clause presence.
- Scores are **not** calibrated probabilities or confidence percentages.
---
## 5. Usage & Legal Disclaimer
- **Intended Use:** Portfolio evaluation, academic research, and interactive open-source demonstration of contract risk detection.
- **Legal Disclaimer:** This software and model predictions do **NOT** constitute legal advice, formal contract audit, or legal guarantee. Users should consult qualified legal professionals for actual contract review.