File size: 5,537 Bytes
65a8cd3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
6aed982
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
65a8cd3
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
---
license: mit
base_model: microsoft/mdeberta-v3-base
tags:
- causal-extraction
- causality
- cause-effect
- span-extraction
- causal-news-corpus
- multilingual
language:
- en
- es
- fr
- de
- pt
- tr
- ru
- ar
- zh
- ja
---

# causal-span-pointer-v2

A **span-pointer** causal extraction model: given a sentence it predicts the
**cause**, **effect** and **signal** spans as start/end pointers, decoded under
ordering/non-overlap constraints with beam search (top-2 relations per sentence).
Fine-tuned from [`microsoft/mdeberta-v3-base`](https://huggingface.co/microsoft/mdeberta-v3-base)
on the [Causal News Corpus](https://github.com/tanfiona/CausalNewsCorpus) Subtask-2
(CC0), augmented with synthetic cause/effect data in 6 languages (en, es, fr, de, nl, tr)
and hard negatives that train the causal gate. Architecture reimplemented from the CNC
baseline (MIT).

## Benchmark (Causal News Corpus Subtask 2)

Official scorer (`evaluation/subtask2`: FairEval + best-combination alignment), V2 dev:

```
Overall   Precision 0.708   Recall 0.694   F1 0.699
Cause     F1 0.726
Effect    F1 0.702
Signal    F1 0.661

Causal gate precision (940 multilingual negatives): 0.999
Synthetic test per role, 6 languages: Cause 0.975 / Effect 0.972 / Signal 0.946
```

**This beats the organizer's 0.627 dev baseline** and the 2022 shared-task winner
(0.542, test); it trails the 2023 winner (0.728, test). Same scorer, same dev set.

### vs a few-shot LLM

A prompted general LLM does not match this fine-tune. Qwen2.5-7B-Instruct, few-shot on
the same dev set and official scorer, scores **0.24** F1 with a plain causal prompt and
**0.41** with a scheme-aware prompt (vs **0.70** here). Even on a capability-fair subset
-- causality a reader recognises without CNC's broad purpose/motive/implicit
conventions -- the LLM reaches ~0.45 vs this model's ~0.63. The residual gap is exact
span-boundary precision, which fine-tuning on the annotation provides.

## Usage

This is a custom architecture, so inference goes through the `causal_span_model`
package (not `AutoModel`):

```python
from huggingface_hub import snapshot_download
from causal_span_model.pointer.submission import load_pointer, predict_sentence

local_dir = snapshot_download("Berk/causal-span-pointer-v2")
model, tokenizer = load_pointer(local_dir)
print(predict_sentence(model, tokenizer, "Heavy rainfall caused severe flooding."))
# ['<ARG0>Heavy rainfall</ARG0> <SIG0>caused</SIG0> <ARG1>severe flooding</ARG1> .', ...]
```

`<ARG0>` = cause, `<ARG1>` = effect, `<SIG0>` = signal. The prediction is a list of
tagged relation strings (up to two per sentence).

### Multilingual

Trained on English CNC spans plus synthetic cause/effect data in 6 languages, and
multilingual at inference (mDeBERTa encoder + script-aware segmentation). Use
`predict_relations`, which returns character-exact
spans in any script:

```python
from causal_span_model.pointer.infer import predict_relations

predict_relations(model, tokenizer, "暴雨导致该地区发生严重洪灾。")
# [{'cause': '暴雨', 'effect': '该地区发生严重洪灾', 'signal': '导致'}]
predict_relations(model, tokenizer, "Las fuertes lluvias provocaron inundaciones.")
# [{'cause': 'Las fuertes lluvias', 'effect': 'inundaciones', 'signal': 'provocaron'}]
```

Verified on es/fr/de/pt/tr/ru/ar and CJK (zh/ja).

## Notes

- It is NOT compatible with a generic token-classification ONNX consumer -- it
  needs its own start/end + beam-search decoder (provided by the package).
- It has a built-in **causal gate** (a causal/non-causal head, ~0.85 accuracy on
  CNC dev): `predict_relations` returns `[]` on text it judges non-causal, so it
  is safe to run on arbitrary input. Beam duplicates are collapsed to one relation
  per distinct cause->effect.

## Companion causal gate (`token_gate/`)

The repo also ships a fine-tuned **token-aware causal gate** in `token_gate/`: a
`paraphrase-multilingual-MiniLM-L12-v2` sequence classifier (P(causal); id2label
`{0: non_causal, 1: causal}`) that decides whether a sentence expresses a causal relation
before the pointer extracts spans. Unlike a frozen-embedding gate it keys on the relation, not
the topic, so it separates a verb-causal sentence from its plain twin ("The GPU cluster
increased training throughput" -> 0.99 vs "The GPU cluster is installed in rack 4" -> 0.01).

reasongraph >= 0.7.1:

```python
from reasongraph import CausalPointerExtractor
ex = CausalPointerExtractor(
    model="Berk/causal-span-pointer-v2",
    token_gate="hf://Berk/causal-span-pointer-v2/token_gate",
    token_gate_threshold=0.10)
```

Or directly:

```python
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
tok = AutoTokenizer.from_pretrained("Berk/causal-span-pointer-v2", subfolder="token_gate")
m = AutoModelForSequenceClassification.from_pretrained("Berk/causal-span-pointer-v2", subfolder="token_gate")
enc = tok("It flooded because it rained.", return_tensors="pt")
p_causal = torch.softmax(m(**enc).logits, -1)[0, 1].item()  # ~0.98
```

At cutoff **0.10** on the 39-case reviewed set / 60 plain facts: keeps 75% of causal hop
facts, rejects **100%** of plain facts, synthetic-dev F1 0.98, ~5 ms/sentence on CPU. It is
stricter than the companion embedding gate (`embed_gate_mlp.joblib`, which keeps ~92% of hop
facts but lets ~25% of plain facts through): use the token gate when clean plain-fact
rejection matters, the embedding gate for maximum causal recall.

## License

MIT (weights and code). Training data is CC0-1.0.