sileod commited on
Commit
6fa4fce
路
verified 路
1 Parent(s): afa49af

v3: same recipe as v2 trained for 24,000 steps (Decision Index 0.3 public 29.05); v2 kept at tag v2

Browse files
Files changed (3) hide show
  1. README.md +13 -11
  2. decider_config.json +2 -2
  3. model.safetensors +1 -1
README.md CHANGED
@@ -13,8 +13,10 @@ tags:
13
 
14
  # tasksource-decider-nano
15
 
16
- This repository was `tasksource/tasksource-jev-nano-v1`. Its `main` (tag `v2`) is now a new model; the previous
17
- multi-vector model (v1) stays at tag [`v1`](https://huggingface.co/tasksource/tasksource-decider-nano/tree/v1).
 
 
18
 
19
  A 150M-parameter typed-decision model for the [Decision Index](https://huggingface.co/spaces/multimodalart/jev-decision-index).
20
  It takes a `state` and typed questions (`choice`, `noul`) and returns a probability for every option. It does not
@@ -22,9 +24,9 @@ generate text.
22
 
23
  | | |
24
  |---|---|
25
- | Decision Index 0.3, public index | **26.77** |
26
  | Requests | 140,171 `ok`, 7 `unsupported` (pairs above 8,192 tokens) |
27
- | Latency, 1脳 NVIDIA A10 23 GB, one request at a time | median 45 ms 路 mean 71 ms 路 p80 71 ms |
28
  | Complete results | [runs/tasksource-decider-nano](https://huggingface.co/datasets/tasksource/decision-index-results/tree/main/runs/tasksource-decider-nano) |
29
 
30
  ## Architecture
@@ -64,7 +66,7 @@ python -m decision_index pipeline --engine decision_index_engine:DeciderEngine \
64
 
65
  ## Training
66
 
67
- 12,000 steps, batches of 8 requests, learning rate 2e-5, inputs up to 2,048 tokens, seed 7. The loss is soft
68
  cross-entropy over each question's options. The training mix has 240,000 requests:
69
 
70
  | Share | Source |
@@ -84,10 +86,10 @@ request.
84
 
85
  | Area | Skill | Raw |
86
  |---|---:|---:|
87
- | Knowledge & Reasoning | 6.0 | 29.5 |
88
- | Language Understanding | 28.2 | 47.9 |
89
- | Retrieval & Classification | 42.0 | 58.2 |
90
- | Tools & Automation | 44.6 | 49.8 |
91
- | Arts & Human Taste | 13.9 | 43.5 |
92
 
93
- Raw index 45.1; kit `9eb2dbe`, edition 0.3, 140,171 requests scored.
 
13
 
14
  # tasksource-decider-nano
15
 
16
+ This repository was `tasksource/tasksource-jev-nano-v1`. Its `main` (tag `v3`) is the joint cross-encoder trained for
17
+ 24,000 steps. The same model trained for 12,000 steps (public index 26.77) stays at tag
18
+ [`v2`](https://huggingface.co/tasksource/tasksource-decider-nano/tree/v2), and the previous multi-vector model at tag
19
+ [`v1`](https://huggingface.co/tasksource/tasksource-decider-nano/tree/v1).
20
 
21
  A 150M-parameter typed-decision model for the [Decision Index](https://huggingface.co/spaces/multimodalart/jev-decision-index).
22
  It takes a `state` and typed questions (`choice`, `noul`) and returns a probability for every option. It does not
 
24
 
25
  | | |
26
  |---|---|
27
+ | Decision Index 0.3, public index | **29.05** |
28
  | Requests | 140,171 `ok`, 7 `unsupported` (pairs above 8,192 tokens) |
29
+ | Latency, 1脳 NVIDIA A10 23 GB, one request at a time | median 73 ms 路 mean 96 ms 路 p80 95 ms |
30
  | Complete results | [runs/tasksource-decider-nano](https://huggingface.co/datasets/tasksource/decision-index-results/tree/main/runs/tasksource-decider-nano) |
31
 
32
  ## Architecture
 
66
 
67
  ## Training
68
 
69
+ 24,000 steps, batches of 8 requests, learning rate 2e-5, inputs up to 2,048 tokens, seed 7. The loss is soft
70
  cross-entropy over each question's options. The training mix has 240,000 requests:
71
 
72
  | Share | Source |
 
86
 
87
  | Area | Skill | Raw |
88
  |---|---:|---:|
89
+ | Knowledge & Reasoning | 6.6 | 29.8 |
90
+ | Language Understanding | 33.2 | 51.5 |
91
+ | Retrieval & Classification | 47.1 | 61.3 |
92
+ | Tools & Automation | 43.3 | 48.2 |
93
+ | Arts & Human Taste | 14.1 | 43.2 |
94
 
95
+ Raw index 46.41; kit `9eb2dbe`, edition 0.3, 140,171 requests scored.
decider_config.json CHANGED
@@ -4,7 +4,7 @@
4
  "base": "cross-encoder/ettin-reranker-150m-v1",
5
  "train_max_length": 2048,
6
  "inference_max_length": 8192,
7
- "step": 12000,
8
  "train": {
9
  "arch": "joint",
10
  "opt_pool": "marker",
@@ -16,7 +16,7 @@
16
  "batch_requests": 8,
17
  "max_k": 32,
18
  "seed": 7,
19
- "total_steps": 12000,
20
  "max_length": 2048,
21
  "pack": 0,
22
  "init": null,
 
4
  "base": "cross-encoder/ettin-reranker-150m-v1",
5
  "train_max_length": 2048,
6
  "inference_max_length": 8192,
7
+ "step": 24000,
8
  "train": {
9
  "arch": "joint",
10
  "opt_pool": "marker",
 
16
  "batch_requests": 8,
17
  "max_k": 32,
18
  "seed": 7,
19
+ "total_steps": 24000,
20
  "max_length": 2048,
21
  "pack": 0,
22
  "init": null,
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:a96dcfdf2d4e7da62429ba9f836535adf3a328f30cb51366095ad2c38fb13659
3
  size 596074556
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ca302526bc031a820677508630f72806c7f757fe25191d5eb1e154f658191717
3
  size 596074556