Game Review Reader (small)
Classifies one claim from a Steam review: which of 26 subjects it is
about, and whether it is praise, a complaint or neutral. Fine-tuned from
intfloat/multilingual-e5-small, run e5small-pool-person-rules-s1. Part of
SteamGauge.
Use
model.onnxtakesinput_idsandattention_maskfor a pair tokenised withtokenizer.json: the claim, then the text of its review around it, at most 128 tokens. For a game played only in a VR headset, the review text starts with "Played in a VR headset."- It returns
subject_logits(26, in the order ofsubjectsinreader.json),polarity_logits(praise, complaint, neutral) andpooled. reader.jsonholds the confidence each subject (thresholds) and each language (language_thresholds) needs before the reader answers. A claim gets an answer when its top subject's probability clears both lines. A language with no line, or missing from the list, gets no answers.
Subjects: accessibility, atmosphere, audio, bugs, community, compatibility, content, controls, difficulty, gameplay, genre, graphics, language, licensing, mods, monetisation, multiplayer, offtopic, performance, policy, price, story, tutorial, updates, verdict, vr.
Results
On 5,080 claims from games held out of training:
- Subject accuracy 0.708, macro F1 0.618.
- Polarity macro F1 0.795.
- Calibration error 0.080; area under the risk-coverage curve 0.116.
- With one confidence line at 0.59: answers 78% of claims, accuracy 0.796 (95% interval 0.783 to 0.808, 3,963 claims).
- Lowest per-subject F1:
accessibility0.00,vr0.34,community0.38,policy0.45,licensing0.49. - No answers in: bulgarian, danish, dutch, finnish, greek, indonesian, norwegian, portuguese, romanian, vietnamese.
Sizes
| Reader | Parameters | Download | Accuracy at 80% coverage | Accuracy at 90% coverage | GPU memory | One large game on the GPU | One small game on the CPU |
|---|---|---|---|---|---|---|---|
| small | 118M | 244 MB | 77.9% | 73.8% | 1.2 GB | 62 s | 36 s |
| standard | 559M | 1.1 GB | 84.7% | 82.0% | 2.6 GB | 125 s | 285 s |
| 4B (not published) | 4.0B | 8.8 GB | 83.6% | 79.8% | 12.5 GB | 23 min | 49 min |
- Accuracy at a coverage: 458 claims from ten games held out of training, labelled under the current category sheet. Each reader answers the given share of claims it is most confident about, scored on DirectML (
frontier.py reader). - GPU memory and one large game on the GPU: one read of game 920210 (117,664 claims) through the desktop app on DirectML, on an otherwise idle NVIDIA GeForce RTX 4090. Memory is the peak during the read minus what the GPU held before. The 4B's GPU and CPU figures are from a 4B of the same architecture.
- One small game on the CPU: 1,416 claims of game 1888930 on the CPU only, loading included, on an AMD Ryzen 9 5950X with other work running.
Training
36,017 labelled claims, with 4,310 more for validation, from Aureliolo/game-review-claims. SteamGauge commit 6d55fc474a28.
Limits
The training labels were written by Claude models from a category sheet, and the results above measure agreement with those labels.
Licence
Apache-2.0. Base model intfloat/multilingual-e5-small: MIT.
Model tree for Aureliolo/game-review-reader-small
Base model
intfloat/multilingual-e5-small