Feature Extraction
Transformers
Safetensors
modernbert
typed-decisions
decision-index
decision-model
cross-encoder
text-embeddings-inference
Instructions to use tasksource/tasksource-decider-nano with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use tasksource/tasksource-decider-nano with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="tasksource/tasksource-decider-nano")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("tasksource/tasksource-decider-nano") model = AutoModel.from_pretrained("tasksource/tasksource-decider-nano", device_map="auto") - Notebooks
- Google Colab
- Kaggle
v3: same recipe as v2 trained for 24,000 steps (Decision Index 0.3 public 29.05); v2 kept at tag v2
Browse files- README.md +13 -11
- decider_config.json +2 -2
- model.safetensors +1 -1
README.md
CHANGED
|
@@ -13,8 +13,10 @@ tags:
|
|
| 13 |
|
| 14 |
# tasksource-decider-nano
|
| 15 |
|
| 16 |
-
This repository was `tasksource/tasksource-jev-nano-v1`. Its `main` (tag `
|
| 17 |
-
|
|
|
|
|
|
|
| 18 |
|
| 19 |
A 150M-parameter typed-decision model for the [Decision Index](https://huggingface.co/spaces/multimodalart/jev-decision-index).
|
| 20 |
It takes a `state` and typed questions (`choice`, `noul`) and returns a probability for every option. It does not
|
|
@@ -22,9 +24,9 @@ generate text.
|
|
| 22 |
|
| 23 |
| | |
|
| 24 |
|---|---|
|
| 25 |
-
| Decision Index 0.3, public index | **
|
| 26 |
| Requests | 140,171 `ok`, 7 `unsupported` (pairs above 8,192 tokens) |
|
| 27 |
-
| Latency, 1脳 NVIDIA A10 23 GB, one request at a time | median
|
| 28 |
| Complete results | [runs/tasksource-decider-nano](https://huggingface.co/datasets/tasksource/decision-index-results/tree/main/runs/tasksource-decider-nano) |
|
| 29 |
|
| 30 |
## Architecture
|
|
@@ -64,7 +66,7 @@ python -m decision_index pipeline --engine decision_index_engine:DeciderEngine \
|
|
| 64 |
|
| 65 |
## Training
|
| 66 |
|
| 67 |
-
|
| 68 |
cross-entropy over each question's options. The training mix has 240,000 requests:
|
| 69 |
|
| 70 |
| Share | Source |
|
|
@@ -84,10 +86,10 @@ request.
|
|
| 84 |
|
| 85 |
| Area | Skill | Raw |
|
| 86 |
|---|---:|---:|
|
| 87 |
-
| Knowledge & Reasoning | 6.
|
| 88 |
-
| Language Understanding |
|
| 89 |
-
| Retrieval & Classification |
|
| 90 |
-
| Tools & Automation |
|
| 91 |
-
| Arts & Human Taste |
|
| 92 |
|
| 93 |
-
Raw index
|
|
|
|
| 13 |
|
| 14 |
# tasksource-decider-nano
|
| 15 |
|
| 16 |
+
This repository was `tasksource/tasksource-jev-nano-v1`. Its `main` (tag `v3`) is the joint cross-encoder trained for
|
| 17 |
+
24,000 steps. The same model trained for 12,000 steps (public index 26.77) stays at tag
|
| 18 |
+
[`v2`](https://huggingface.co/tasksource/tasksource-decider-nano/tree/v2), and the previous multi-vector model at tag
|
| 19 |
+
[`v1`](https://huggingface.co/tasksource/tasksource-decider-nano/tree/v1).
|
| 20 |
|
| 21 |
A 150M-parameter typed-decision model for the [Decision Index](https://huggingface.co/spaces/multimodalart/jev-decision-index).
|
| 22 |
It takes a `state` and typed questions (`choice`, `noul`) and returns a probability for every option. It does not
|
|
|
|
| 24 |
|
| 25 |
| | |
|
| 26 |
|---|---|
|
| 27 |
+
| Decision Index 0.3, public index | **29.05** |
|
| 28 |
| Requests | 140,171 `ok`, 7 `unsupported` (pairs above 8,192 tokens) |
|
| 29 |
+
| Latency, 1脳 NVIDIA A10 23 GB, one request at a time | median 73 ms 路 mean 96 ms 路 p80 95 ms |
|
| 30 |
| Complete results | [runs/tasksource-decider-nano](https://huggingface.co/datasets/tasksource/decision-index-results/tree/main/runs/tasksource-decider-nano) |
|
| 31 |
|
| 32 |
## Architecture
|
|
|
|
| 66 |
|
| 67 |
## Training
|
| 68 |
|
| 69 |
+
24,000 steps, batches of 8 requests, learning rate 2e-5, inputs up to 2,048 tokens, seed 7. The loss is soft
|
| 70 |
cross-entropy over each question's options. The training mix has 240,000 requests:
|
| 71 |
|
| 72 |
| Share | Source |
|
|
|
|
| 86 |
|
| 87 |
| Area | Skill | Raw |
|
| 88 |
|---|---:|---:|
|
| 89 |
+
| Knowledge & Reasoning | 6.6 | 29.8 |
|
| 90 |
+
| Language Understanding | 33.2 | 51.5 |
|
| 91 |
+
| Retrieval & Classification | 47.1 | 61.3 |
|
| 92 |
+
| Tools & Automation | 43.3 | 48.2 |
|
| 93 |
+
| Arts & Human Taste | 14.1 | 43.2 |
|
| 94 |
|
| 95 |
+
Raw index 46.41; kit `9eb2dbe`, edition 0.3, 140,171 requests scored.
|
decider_config.json
CHANGED
|
@@ -4,7 +4,7 @@
|
|
| 4 |
"base": "cross-encoder/ettin-reranker-150m-v1",
|
| 5 |
"train_max_length": 2048,
|
| 6 |
"inference_max_length": 8192,
|
| 7 |
-
"step":
|
| 8 |
"train": {
|
| 9 |
"arch": "joint",
|
| 10 |
"opt_pool": "marker",
|
|
@@ -16,7 +16,7 @@
|
|
| 16 |
"batch_requests": 8,
|
| 17 |
"max_k": 32,
|
| 18 |
"seed": 7,
|
| 19 |
-
"total_steps":
|
| 20 |
"max_length": 2048,
|
| 21 |
"pack": 0,
|
| 22 |
"init": null,
|
|
|
|
| 4 |
"base": "cross-encoder/ettin-reranker-150m-v1",
|
| 5 |
"train_max_length": 2048,
|
| 6 |
"inference_max_length": 8192,
|
| 7 |
+
"step": 24000,
|
| 8 |
"train": {
|
| 9 |
"arch": "joint",
|
| 10 |
"opt_pool": "marker",
|
|
|
|
| 16 |
"batch_requests": 8,
|
| 17 |
"max_k": 32,
|
| 18 |
"seed": 7,
|
| 19 |
+
"total_steps": 24000,
|
| 20 |
"max_length": 2048,
|
| 21 |
"pack": 0,
|
| 22 |
"init": null,
|
model.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 596074556
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:ca302526bc031a820677508630f72806c7f757fe25191d5eb1e154f658191717
|
| 3 |
size 596074556
|