EventTwin — 中文同事件判别器
判断两个中文新闻事件描述是否指同一个现实发生的事件。
模型组成
本仓库包含一个三模型集成系统(AUROC 0.981):
| 模型 | 位置 | 底座 | 参数量 | 单模型 AUROC |
|---|---|---|---|---|
| v15(主力) | 根目录 | Qwen3-Reranker-4B | 4B | 0.963 |
| v8 | ensemble/v8/ |
bge-reranker-v2-m3 | 568M | 0.909 |
| v10a | ensemble/v10a/ |
bge-reranker-v2-m3 | 568M | 0.903 |
集成公式
import torch, json
from transformers import AutoModelForCausalLM, AutoTokenizer
# 加载三个模型(各目录下的 calibration.json 存有温度校准值)
# score = sigmoid(0.8×logit(v15/T15) + 0.1×logit(v8/T8) + 0.1×logit(v10a/T10a)) / 0.5
性能
| 指标 | v15 单模型 | 三模型集成 | 教师(Jev) |
|---|---|---|---|
| AUROC | 0.963 | 0.981 | 0.998 |
| gray 层 | 0.974 | 0.979 | 1.000 |
| pos 层 | 0.845 | 0.922 | 0.992 |
| ECE | 0.107 | 0.096 | 0.062 |
使用方式
v15 单模型(推荐快速部署)
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained("MaYiding/EventTwin", torch_dtype=torch.bfloat16).cuda().eval()
tok = AutoTokenizer.from_pretrained("MaYiding/EventTwin")
tok.padding_side = "left"
cal = json.load(open("calibration.json")) # 温度 T
PFX = "<|im_start|>system\nJudge whether the Document meets the requirements based on the Query. Only give me the judgment. The judgment should be yes or no.<|im_end|>\n<|im_start|>user\nQuery: "
SFX = "\nDocument: "
SFX2 = "\nJudgment: <|im_end|>\n<|im_start|>assistant\n"
yes_id = tok("yes", add_special_tokens=False)["input_ids"][0]
no_id = tok("no", add_special_tokens=False)["input_ids"][0]
def judge(event_a, event_b):
text = f"{PFX}{event_a}{SFX}{event_b}{SFX2}"
inp = tok([text], return_tensors="pt", add_special_tokens=False).to("cuda")
with torch.no_grad():
logits = model(**inp).logits[:, -1, :].float()
two = torch.stack([logits[:, no_id], logits[:, yes_id]], dim=-1)
return float(torch.softmax(two / cal["temperature"], dim=-1)[0, 1])
score = judge("小米YU7正式上市,售价25.35万元起", "小米发布YU7 SUV,起售价25.35万")
# score ≈ 0.95(同一事件)
训练数据
EventTwin-Data:1000 对分层金标 + 81K 训练对
技术报告
- 19 版本迭代完整对决报告(GitHub
ml/benchmark/学生模型对决报告.md) - 训练配方:提及式数据 + LoRA + Label Smoothing + 温度校准
- Downloads last month
- 25
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support