Spaces:
Running
Running
|
Download README.md from OpenTWBench/README: direct link, hf CLI and curl.
- Browser
- Download file 2.37 kB
-
https://huggingface.co/spaces/OpenTWBench/README/resolve/main/README.md
- Command line
-
hf download hf://spaces/OpenTWBench/README/README.md
-
curl -L -o README.md https://huggingface.co/spaces/OpenTWBench/README/resolve/main/README.md
2.37 kB
| title: README | |
| emoji: ๐๏ธ | |
| colorFrom: blue | |
| colorTo: green | |
| sdk: static | |
| pinned: true | |
| # OpenTWBench | |
| **An open evaluation suite for Taiwan-domain knowledge in large language models.** | |
| ๐ **[opentwbench.ai](https://opentwbench.ai)** โ leaderboard, subject map, and how the data is built | |
| --- | |
| Most Chinese-language benchmarks are built from mainland sources, in Simplified | |
| Chinese, about mainland institutions. A model can score well on them while knowing | |
| nothing about the Labor Standards Act, the Household Registration Act, or how a | |
| Taiwanese pharmacist is licensed. These datasets exist to make that gap measurable. | |
| The questions come from Taiwan's **national examinations**, published as open data by | |
| the **Ministry of Examination (่้ธ้จ)** โ professionally written, officially | |
| answer-keyed, and covering essentially every regulated profession in the country. | |
| ## The datasets | |
| One dataset per **ๅญธ็ง (academic subject)**, named `tw-<subject>-bench`, in the | |
| **Twinkle Eval MCQ** format (`question`, `A`โ`D`, `answer`) with provenance columns | |
| alongside. Subjects are the unit, not the profession sitting the exam: | |
| ็ฅ็ถ็พ็ ็ฉ็ๆฒป็ๅญธ and ้ชจ็ง็พ็ ็ฉ็ๆฒป็ๅญธ are separate benchmarks, because a model can | |
| know one and not the other. | |
| ```python | |
| from datasets import load_dataset | |
| ds = load_dataset("OpenTWBench/tw-medicine-bench", split="test") | |
| ``` | |
| ## Reading the scores honestly | |
| - These are **public past papers**, very likely in pretraining corpora. Use the suite | |
| for *relative* comparison between models, not to claim a model "passes" an exam. | |
| - **Check the parse rate before believing a low score** โ a model that knows the | |
| answer but writes it in an unexpected shape is a formatting failure. | |
| - **Shuffle the options.** The source papers carry an answer-position skew. | |
| Taiwan **law** is covered separately and in depth by | |
| [`lianghsun/tw-legal-benchmark-v2`](https://huggingface.co/datasets/lianghsun/tw-legal-benchmark-v2); | |
| questions already published there are excluded from this org, item by item. | |
| ## Contributing | |
| A wrong answer key is worse than a missing question. Open a discussion on the dataset | |
| it affects, quoting the `paper_id` and `q_no`, and it can be traced to the original PDF. | |
| *Questions and answer keys are official publications of ่้ธ้จ, released as open data. | |
| Packaging and metadata: Apache-2.0.* | |