Add Typed Decisions benchmark results

#2
by codelion - opened
Files changed (1) hide show
  1. .eval_results/typed-decisions.yaml +36 -0
.eval_results/typed-decisions.yaml ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ - dataset:
2
+ id: LocalLLaMA/typed-decisions
3
+ task_id: accuracy
4
+ value: 0.701
5
+ source:
6
+ url: https://huggingface.co/datasets/LocalLLaMA/typed-decisions
7
+ name: Typed Decisions leaderboard (README)
8
+ user: LocalLLaMA
9
+ notes: "General model (any question schema at request time), not fitted by us on this benchmark's train split. Scored on the full 400-case test split (2,000 decisions): one request per case with all five questions, text only, the benchmark's own metrics, MLX on an Apple M3 Max."
10
+ - dataset:
11
+ id: LocalLLaMA/typed-decisions
12
+ task_id: kl_from_gold
13
+ value: 0.2126
14
+ source:
15
+ url: https://huggingface.co/datasets/LocalLLaMA/typed-decisions
16
+ name: Typed Decisions leaderboard (README)
17
+ user: LocalLLaMA
18
+ notes: "General model (any question schema at request time), not fitted by us on this benchmark's train split. Scored on the full 400-case test split (2,000 decisions): one request per case with all five questions, text only, the benchmark's own metrics, MLX on an Apple M3 Max."
19
+ - dataset:
20
+ id: LocalLLaMA/typed-decisions
21
+ task_id: brier
22
+ value: 0.1104
23
+ source:
24
+ url: https://huggingface.co/datasets/LocalLLaMA/typed-decisions
25
+ name: Typed Decisions leaderboard (README)
26
+ user: LocalLLaMA
27
+ notes: "General model (any question schema at request time), not fitted by us on this benchmark's train split. Scored on the full 400-case test split (2,000 decisions): one request per case with all five questions, text only, the benchmark's own metrics, MLX on an Apple M3 Max."
28
+ - dataset:
29
+ id: LocalLLaMA/typed-decisions
30
+ task_id: ece
31
+ value: 0.1204
32
+ source:
33
+ url: https://huggingface.co/datasets/LocalLLaMA/typed-decisions
34
+ name: Typed Decisions leaderboard (README)
35
+ user: LocalLLaMA
36
+ notes: "General model (any question schema at request time), not fitted by us on this benchmark's train split. Scored on the full 400-case test split (2,000 decisions): one request per case with all five questions, text only, the benchmark's own metrics, MLX on an Apple M3 Max."