codelion commited on
Commit
35dddc5
·
verified ·
1 Parent(s): 5c646a4

Add Typed Decisions benchmark results

Browse files

Adds this model's scores on [Typed Decisions](https://huggingface.co/datasets/LocalLLaMA/typed-decisions) as `.eval_results/typed-decisions.yaml`, so they are listed on the benchmark's Hub leaderboard.

| Accuracy | KL from gold (lower is better) | Brier (lower is better) | ECE (lower is better) |
|---|---|---|---|
| 0.7185 | 0.1987 | 0.106 | 0.1185 |

General model (any question schema at request time), not fitted by us on this benchmark's train split. Scored on the full 400-case test split (2,000 decisions): one request per case with all five questions, text only, the benchmark's own metrics, MLX on an Apple M3 Max.

Once merged, the model appears on the leaderboard. Quantizations and other derived models are hidden by default; turn the base-model switch off to see them. While this PR is open the Hub marks the scores as community-provided.

If you would rather not list the model, or a number looks wrong, close this PR or tell us and we will correct it. Scored by codelion (OptiQ).

Files changed (1) hide show
  1. .eval_results/typed-decisions.yaml +36 -0
.eval_results/typed-decisions.yaml ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ - dataset:
2
+ id: LocalLLaMA/typed-decisions
3
+ task_id: accuracy
4
+ value: 0.7185
5
+ source:
6
+ url: https://huggingface.co/datasets/LocalLLaMA/typed-decisions
7
+ name: Typed Decisions leaderboard (README)
8
+ user: LocalLLaMA
9
+ notes: "General model (any question schema at request time), not fitted by us on this benchmark's train split. Scored on the full 400-case test split (2,000 decisions): one request per case with all five questions, text only, the benchmark's own metrics, MLX on an Apple M3 Max."
10
+ - dataset:
11
+ id: LocalLLaMA/typed-decisions
12
+ task_id: kl_from_gold
13
+ value: 0.1987
14
+ source:
15
+ url: https://huggingface.co/datasets/LocalLLaMA/typed-decisions
16
+ name: Typed Decisions leaderboard (README)
17
+ user: LocalLLaMA
18
+ notes: "General model (any question schema at request time), not fitted by us on this benchmark's train split. Scored on the full 400-case test split (2,000 decisions): one request per case with all five questions, text only, the benchmark's own metrics, MLX on an Apple M3 Max."
19
+ - dataset:
20
+ id: LocalLLaMA/typed-decisions
21
+ task_id: brier
22
+ value: 0.106
23
+ source:
24
+ url: https://huggingface.co/datasets/LocalLLaMA/typed-decisions
25
+ name: Typed Decisions leaderboard (README)
26
+ user: LocalLLaMA
27
+ notes: "General model (any question schema at request time), not fitted by us on this benchmark's train split. Scored on the full 400-case test split (2,000 decisions): one request per case with all five questions, text only, the benchmark's own metrics, MLX on an Apple M3 Max."
28
+ - dataset:
29
+ id: LocalLLaMA/typed-decisions
30
+ task_id: ece
31
+ value: 0.1185
32
+ source:
33
+ url: https://huggingface.co/datasets/LocalLLaMA/typed-decisions
34
+ name: Typed Decisions leaderboard (README)
35
+ user: LocalLLaMA
36
+ notes: "General model (any question schema at request time), not fitted by us on this benchmark's train split. Scored on the full 400-case test split (2,000 decisions): one request per case with all five questions, text only, the benchmark's own metrics, MLX on an Apple M3 Max."