File size: 5,537 Bytes
65a8cd3 6aed982 65a8cd3 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 | ---
license: mit
base_model: microsoft/mdeberta-v3-base
tags:
- causal-extraction
- causality
- cause-effect
- span-extraction
- causal-news-corpus
- multilingual
language:
- en
- es
- fr
- de
- pt
- tr
- ru
- ar
- zh
- ja
---
# causal-span-pointer-v2
A **span-pointer** causal extraction model: given a sentence it predicts the
**cause**, **effect** and **signal** spans as start/end pointers, decoded under
ordering/non-overlap constraints with beam search (top-2 relations per sentence).
Fine-tuned from [`microsoft/mdeberta-v3-base`](https://huggingface.co/microsoft/mdeberta-v3-base)
on the [Causal News Corpus](https://github.com/tanfiona/CausalNewsCorpus) Subtask-2
(CC0), augmented with synthetic cause/effect data in 6 languages (en, es, fr, de, nl, tr)
and hard negatives that train the causal gate. Architecture reimplemented from the CNC
baseline (MIT).
## Benchmark (Causal News Corpus Subtask 2)
Official scorer (`evaluation/subtask2`: FairEval + best-combination alignment), V2 dev:
```
Overall Precision 0.708 Recall 0.694 F1 0.699
Cause F1 0.726
Effect F1 0.702
Signal F1 0.661
Causal gate precision (940 multilingual negatives): 0.999
Synthetic test per role, 6 languages: Cause 0.975 / Effect 0.972 / Signal 0.946
```
**This beats the organizer's 0.627 dev baseline** and the 2022 shared-task winner
(0.542, test); it trails the 2023 winner (0.728, test). Same scorer, same dev set.
### vs a few-shot LLM
A prompted general LLM does not match this fine-tune. Qwen2.5-7B-Instruct, few-shot on
the same dev set and official scorer, scores **0.24** F1 with a plain causal prompt and
**0.41** with a scheme-aware prompt (vs **0.70** here). Even on a capability-fair subset
-- causality a reader recognises without CNC's broad purpose/motive/implicit
conventions -- the LLM reaches ~0.45 vs this model's ~0.63. The residual gap is exact
span-boundary precision, which fine-tuning on the annotation provides.
## Usage
This is a custom architecture, so inference goes through the `causal_span_model`
package (not `AutoModel`):
```python
from huggingface_hub import snapshot_download
from causal_span_model.pointer.submission import load_pointer, predict_sentence
local_dir = snapshot_download("Berk/causal-span-pointer-v2")
model, tokenizer = load_pointer(local_dir)
print(predict_sentence(model, tokenizer, "Heavy rainfall caused severe flooding."))
# ['<ARG0>Heavy rainfall</ARG0> <SIG0>caused</SIG0> <ARG1>severe flooding</ARG1> .', ...]
```
`<ARG0>` = cause, `<ARG1>` = effect, `<SIG0>` = signal. The prediction is a list of
tagged relation strings (up to two per sentence).
### Multilingual
Trained on English CNC spans plus synthetic cause/effect data in 6 languages, and
multilingual at inference (mDeBERTa encoder + script-aware segmentation). Use
`predict_relations`, which returns character-exact
spans in any script:
```python
from causal_span_model.pointer.infer import predict_relations
predict_relations(model, tokenizer, "暴雨导致该地区发生严重洪灾。")
# [{'cause': '暴雨', 'effect': '该地区发生严重洪灾', 'signal': '导致'}]
predict_relations(model, tokenizer, "Las fuertes lluvias provocaron inundaciones.")
# [{'cause': 'Las fuertes lluvias', 'effect': 'inundaciones', 'signal': 'provocaron'}]
```
Verified on es/fr/de/pt/tr/ru/ar and CJK (zh/ja).
## Notes
- It is NOT compatible with a generic token-classification ONNX consumer -- it
needs its own start/end + beam-search decoder (provided by the package).
- It has a built-in **causal gate** (a causal/non-causal head, ~0.85 accuracy on
CNC dev): `predict_relations` returns `[]` on text it judges non-causal, so it
is safe to run on arbitrary input. Beam duplicates are collapsed to one relation
per distinct cause->effect.
## Companion causal gate (`token_gate/`)
The repo also ships a fine-tuned **token-aware causal gate** in `token_gate/`: a
`paraphrase-multilingual-MiniLM-L12-v2` sequence classifier (P(causal); id2label
`{0: non_causal, 1: causal}`) that decides whether a sentence expresses a causal relation
before the pointer extracts spans. Unlike a frozen-embedding gate it keys on the relation, not
the topic, so it separates a verb-causal sentence from its plain twin ("The GPU cluster
increased training throughput" -> 0.99 vs "The GPU cluster is installed in rack 4" -> 0.01).
reasongraph >= 0.7.1:
```python
from reasongraph import CausalPointerExtractor
ex = CausalPointerExtractor(
model="Berk/causal-span-pointer-v2",
token_gate="hf://Berk/causal-span-pointer-v2/token_gate",
token_gate_threshold=0.10)
```
Or directly:
```python
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
tok = AutoTokenizer.from_pretrained("Berk/causal-span-pointer-v2", subfolder="token_gate")
m = AutoModelForSequenceClassification.from_pretrained("Berk/causal-span-pointer-v2", subfolder="token_gate")
enc = tok("It flooded because it rained.", return_tensors="pt")
p_causal = torch.softmax(m(**enc).logits, -1)[0, 1].item() # ~0.98
```
At cutoff **0.10** on the 39-case reviewed set / 60 plain facts: keeps 75% of causal hop
facts, rejects **100%** of plain facts, synthetic-dev F1 0.98, ~5 ms/sentence on CPU. It is
stricter than the companion embedding gate (`embed_gate_mlp.joblib`, which keeps ~92% of hop
facts but lets ~25% of plain facts through): use the token gate when clean plain-fact
rejection matters, the embedding gate for maximum causal recall.
## License
MIT (weights and code). Training data is CC0-1.0.
|