Genshin Impact Character Knowledge Benchmark
A multiple-choice benchmark for evaluationg an LLM's knowledge about Genshin Impact characters (both playable and non-playable). It contains 1570 questions in total, with 10 questions per character. Data was taken directly from the Genshin Impact Fandom wiki, so scoring is unambiguous. Benchmark date cutoff is 13th September 2026, no offline LLM will be able to achieve 100% yet.
Possible use cases
- Evaluating the knowledge of an LLM about Genshin Impact characters
- Deciding on an LLM for chatting about or interacting with Genshin Impact
- Improving an LLM's knowledge about Genshin Impact characters by injecting into context
Methodology
Each question is asked 3 times, with four possible answer options each, re-shuffled everytime to prevent position bias or lucky guesses. Without shuffling, smaller LLMs tend to pick A in 70% of cases, regardless of content. A question only counts as correct if the model picks the right answer 3 out of 3 times.
Leaderboard
| # | Model | Params | Score (3/3) |
|---|---|---|---|
| 1 | Qwen3.5-9B-Q8_0 | 9B | 29.81% |
| 2 | Ornith-1.5-9B-Q8_0 | 9B | 25.80% |
| 3 | Mistral-v0.3-7B-Q6_K | 7B | 23.63% |
| 4 | MiniCPM5-2B-Q8_0 | 2B | 13.76% |
| 5 | LFM2.5-2.6B-Q8_0 | 2.6B | 13.57% |
| 6 | LFM2.5-1.2B-Instruct-Q8_0 | 1.2B | 13.38% |
| 7 | LFM2-700M-Q8_0 | 700M | 11.59% |
| 8 | LFM-8B-A1B-Q6_K | 8.3B MoE, ~1.5B active | 11.21% |
| 9 | LFM2.5-350M-Q8_0 | 350M | 7.39% |
| 10 | LFM2.5-230M-Q8_0 | 230M | 6.43% |
All models are tested at the standard koboldcpp API temperature value of 0.7
Community contributions of larger LLM evaluations are welcome!
Usage
pip install requests
python Benchmark.py --file Benchmark-Data.json --host YOUR-LOCAL-LLM-API-ADDRESS
Results are written into results.json and printed in the terminal after a successful run