Game Review Reader (small)

Classifies one claim from a Steam review: which of 26 subjects it is about, and whether it is praise, a complaint or neutral. Fine-tuned from intfloat/multilingual-e5-small, run e5small-pool-person-rules-s1. Part of SteamGauge.

Use

  • model.onnx takes input_ids and attention_mask for a pair tokenised with tokenizer.json: the claim, then the text of its review around it, at most 128 tokens. For a game played only in a VR headset, the review text starts with "Played in a VR headset."
  • It returns subject_logits (26, in the order of subjects in reader.json), polarity_logits (praise, complaint, neutral) and pooled.
  • reader.json holds the confidence each subject (thresholds) and each language (language_thresholds) needs before the reader answers. A claim gets an answer when its top subject's probability clears both lines. A language with no line, or missing from the list, gets no answers.

Subjects: accessibility, atmosphere, audio, bugs, community, compatibility, content, controls, difficulty, gameplay, genre, graphics, language, licensing, mods, monetisation, multiplayer, offtopic, performance, policy, price, story, tutorial, updates, verdict, vr.

Results

On 5,080 claims from games held out of training:

  • Subject accuracy 0.708, macro F1 0.618.
  • Polarity macro F1 0.795.
  • Calibration error 0.080; area under the risk-coverage curve 0.116.
  • With one confidence line at 0.59: answers 78% of claims, accuracy 0.796 (95% interval 0.783 to 0.808, 3,963 claims).
  • Lowest per-subject F1: accessibility 0.00, vr 0.34, community 0.38, policy 0.45, licensing 0.49.
  • No answers in: bulgarian, danish, dutch, finnish, greek, indonesian, norwegian, portuguese, romanian, vietnamese.

Sizes

Reader Parameters Download Accuracy at 80% coverage Accuracy at 90% coverage GPU memory One large game on the GPU One small game on the CPU
small 118M 244 MB 77.9% 73.8% 1.2 GB 62 s 36 s
standard 559M 1.1 GB 84.7% 82.0% 2.6 GB 125 s 285 s
4B (not published) 4.0B 8.8 GB 83.6% 79.8% 12.5 GB 23 min 49 min
  • Accuracy at a coverage: 458 claims from ten games held out of training, labelled under the current category sheet. Each reader answers the given share of claims it is most confident about, scored on DirectML (frontier.py reader).
  • GPU memory and one large game on the GPU: one read of game 920210 (117,664 claims) through the desktop app on DirectML, on an otherwise idle NVIDIA GeForce RTX 4090. Memory is the peak during the read minus what the GPU held before. The 4B's GPU and CPU figures are from a 4B of the same architecture.
  • One small game on the CPU: 1,416 claims of game 1888930 on the CPU only, loading included, on an AMD Ryzen 9 5950X with other work running.

Training

36,017 labelled claims, with 4,310 more for validation, from Aureliolo/game-review-claims. SteamGauge commit 6d55fc474a28.

Limits

The training labels were written by Claude models from a category sheet, and the results above measure agreement with those labels.

Licence

Apache-2.0. Base model intfloat/multilingual-e5-small: MIT.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Aureliolo/game-review-reader-small

Quantized
(284)
this model

Dataset used to train Aureliolo/game-review-reader-small