README / README.md
lianghsun's picture
Trim the org card; detail moves to opentwbench.ai
ef416ff verified
|
Raw History Blame Contribute Delete
2.37 kB
---
title: README
emoji: ๐Ÿ๏ธ
colorFrom: blue
colorTo: green
sdk: static
pinned: true
---
# OpenTWBench
**An open evaluation suite for Taiwan-domain knowledge in large language models.**
๐Ÿ”— **[opentwbench.ai](https://opentwbench.ai)** โ€” leaderboard, subject map, and how the data is built
---
Most Chinese-language benchmarks are built from mainland sources, in Simplified
Chinese, about mainland institutions. A model can score well on them while knowing
nothing about the Labor Standards Act, the Household Registration Act, or how a
Taiwanese pharmacist is licensed. These datasets exist to make that gap measurable.
The questions come from Taiwan's **national examinations**, published as open data by
the **Ministry of Examination (่€ƒ้ธ้ƒจ)** โ€” professionally written, officially
answer-keyed, and covering essentially every regulated profession in the country.
## The datasets
One dataset per **ๅญธ็ง‘ (academic subject)**, named `tw-<subject>-bench`, in the
**Twinkle Eval MCQ** format (`question`, `A`โ€“`D`, `answer`) with provenance columns
alongside. Subjects are the unit, not the profession sitting the exam:
็ฅž็ถ“็–พ็—…็‰ฉ็†ๆฒป็™‚ๅญธ and ้ชจ็ง‘็–พ็—…็‰ฉ็†ๆฒป็™‚ๅญธ are separate benchmarks, because a model can
know one and not the other.
```python
from datasets import load_dataset
ds = load_dataset("OpenTWBench/tw-medicine-bench", split="test")
```
## Reading the scores honestly
- These are **public past papers**, very likely in pretraining corpora. Use the suite
for *relative* comparison between models, not to claim a model "passes" an exam.
- **Check the parse rate before believing a low score** โ€” a model that knows the
answer but writes it in an unexpected shape is a formatting failure.
- **Shuffle the options.** The source papers carry an answer-position skew.
Taiwan **law** is covered separately and in depth by
[`lianghsun/tw-legal-benchmark-v2`](https://huggingface.co/datasets/lianghsun/tw-legal-benchmark-v2);
questions already published there are excluded from this org, item by item.
## Contributing
A wrong answer key is worse than a missing question. Open a discussion on the dataset
it affects, quoting the `paper_id` and `q_no`, and it can be traced to the original PDF.
*Questions and answer keys are official publications of ่€ƒ้ธ้ƒจ, released as open data.
Packaging and metadata: Apache-2.0.*