Clause risk detection — TextCNN

Multi-label classifier flagging six risk-relevant clause types in commercial contracts, fine-tuned on CUAD v1.

Source and training code: https://github.com/ManasDasri/NNDL

Labels

Index Label
0 Cap on Liability
1 Non-Compete
2 License Grant
3 Audit Rights
4 Termination for Convenience
5 Insurance

Six independent sigmoid outputs, not a softmax: contracts routinely carry several of these clause types at once.

Results

Held-out test split, split by contract so no document appears in training.

Label Precision Recall F1 Support
Cap on Liability 0.760 0.731 0.745 104
Non-Compete 0.400 0.351 0.374 57
License Grant 0.806 0.767 0.786 146
Audit Rights 0.651 0.793 0.715 87
Termination for Convenience 0.491 0.540 0.514 50
Insurance 0.877 0.814 0.844 70

Chunk-level macro F1: 0.663 Document-level macro F1: 0.840 (max-pooled over windows)

How it was trained

  • Contracts are split into overlapping windows of 300 words with 100 words of overlap, because 97% of CUAD contracts exceed a 512-token context and 70% exceed 4,096.
  • A window takes a label when it meaningfully overlaps an annotated clause span.
  • Loss is BCE weighted by each label's negative-to-positive ratio, which runs 22x to 68x. 85% of windows carry no label at all, so an unweighted loss collapses to predicting nothing.
  • Splits are assigned per contract, never per window.

Thresholds

Decision thresholds are tuned per label on the validation split rather than fixed at 0.5, because each label has its own positive rate. Using 0.5 will cost you recall.

Label Threshold
Cap on Liability 0.78
Non-Compete 0.28
License Grant 0.32
Audit Rights 0.39
Termination for Convenience 0.22
Insurance 0.83

Usage

This is a TextCNN, not a transformers architecture, so there is no from_pretrained for it. Load it with the training repository:

from legal_risk_classifier.runtime import CNNRuntime

runtime = CNNRuntime('path/to/this/download')
prediction = runtime.predict_document(contract_text)
print(prediction.predicted, prediction.document_scores)

Limitations

  • Trained on 358 contracts. The rarest label, Non-Compete, appears in around 1.5% of windows and is correspondingly the weakest.
  • Run-to-run variance is about ±0.03 macro F1 on identical settings, so small differences between checkpoints are not meaningful.
  • CUAD is US commercial contracts. Behaviour on other jurisdictions or contract families is untested.
  • This is a research artifact and not legal advice. It is a triage aid for locating clauses a lawyer should read, not a substitute for reading them.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train KenX049/cuad-clause-risk-cnn