assemsabry commited on
Commit
fc7091c
·
verified ·
1 Parent(s): c40a32c

docs: link complete GitHub source repository

Browse files
Files changed (1) hide show
  1. README.md +104 -54
README.md CHANGED
@@ -1,83 +1,133 @@
1
  ---
2
  license: other
3
- license_name: tokenai-neo-dataset-license
4
- license_link: https://huggingface.co/datasets/tokenaii/Neo-dataset/blob/main/legal/DATASET_LICENSE.md
5
- language:
6
- - en
7
- task_categories:
8
- - text-classification
9
  tags:
10
  - system-one
11
  - decision-model
12
  - rlcd
13
- - synthetic
14
  - tool-routing
 
15
  ---
16
 
17
- # NEO - Decision Model Dataset
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
18
 
19
- <p align="center"><img src="assets/neo-cover.png" alt="NEO - Decision Model Dataset" width="360"></p>
 
 
 
 
20
 
21
- TokenAI is a non-profit startup founded in 2025 by Assem Sabry, based in Alexandria, Egypt. The dataset,
22
- model, source code, documentation, and related materials are owned by TokenAI.
23
- Contact: info@tokenai.llc · https://tokenai.llc
24
 
25
- Canonical reference: `tokenaii/Neo` — https://huggingface.co/tokenaii/Neo
 
 
26
 
27
- Complete source repository: [github.com/tokenaii/Neo](https://github.com/tokenaii/Neo).
28
- It contains the complete project source code, training code, synthetic-data
29
- generation code, evaluation and benchmark scripts, configurations, examples,
30
- documentation, licenses, and reproducibility materials. This repository
31
- contains the dataset files, manifest, provenance information, and dataset legal
32
- notices.
33
 
34
- Synthetic, schema-driven decision records for training and evaluating Neo, a System One decision model.
 
 
 
 
35
 
36
- ## Contents
37
 
38
- - 1,000,000 individual decisions.
39
- - 400,000 state records.
40
- - `choice`, `score`, and `noul` questions.
41
- - English-only state text, questions, labels, and tool descriptions.
42
- - Tool routing, urgency, human-review, ambiguity, distractor tools, option permutations, and unknowable cases.
43
 
44
- ## Provenance
 
 
 
 
 
45
 
46
- The records are generated by the `neo-synthetic` generator in the Neo repository. No rows from public datasets are copied into the generated corpus. Public datasets are used only for schema design, sanity checks, or separately documented evaluation.
47
 
48
- Generator: `neo-english-synthetic`.
49
- Generator seed: `20261003`.
 
 
 
50
 
51
- The complete source file used by training has SHA-256
52
- `3eff0bcf1ced0e748698fefefcf8069d2cc997888d7568fddb5ccd6a1a651c50`.
53
- The four files under `data/` are a byte-preserving concatenation of that
54
- training source.
55
 
56
- ## Splits
 
 
 
 
 
57
 
58
- - `train`: 799,882 decisions
59
- - `calibration`: 100,138 decisions
60
- - `test`: 99,980 decisions
61
 
62
- The calibration split is reserved for probability calibration. The test split must not be used to select checkpoints, thresholds, or temperature.
 
 
 
 
 
63
 
64
- ## Intended use and limitations
65
 
66
- This dataset is a research corpus for typed decision models. It is not a substitute for domain-specific validation, human review, or safety testing. Synthetic distributions can contain template artifacts and must not be treated as real user behavior.
 
 
 
67
 
68
- ## Identity and license
 
69
 
70
- This dataset must remain identified as the **TokenAI Neo Decision Model Dataset** and must
71
- retain the canonical reference `tokenaii/Neo`. Rebranding, renaming,
72
- white-labeling, redistributing, rehosting, or publishing it under another
73
- identity is prohibited. The dataset must not be used without preserving its
74
- reference and provenance. See [the dataset license](legal/DATASET_LICENSE.md)
75
- and [the ownership notice](legal/NOTICE.md).
76
 
77
- Commercial, monetized, paid, sponsored, advertising-supported, production,
78
- client, resale, or financially beneficial use is forbidden. Any model or
79
- experiment trained with this dataset must clearly state that it used
80
- `tokenaii/Neo-dataset`.
81
 
82
- Third-party datasets referenced by the project are not included in this release
83
- and retain their own terms.
 
 
 
 
1
  ---
2
  license: other
3
+ license_name: tokenai-neo-model-license-v3.0
4
+ license_link: https://huggingface.co/tokenaii/Neo/blob/main/licenses/MODEL_LICENSE.md
5
+ library_name: pytorch
6
+ pipeline_tag: text-classification
 
 
7
  tags:
8
  - system-one
9
  - decision-model
10
  - rlcd
 
11
  - tool-routing
12
+ - en
13
  ---
14
 
15
+ # NEO - Decision Model
16
+
17
+ <p align="center"><img src="assets/neo-cover.png" alt="NEO - Decision Model" width="360"></p>
18
+
19
+ ## Ownership and contact
20
+
21
+ TokenAI is a non-profit startup founded in 2025 by Assem Sabry, based in
22
+ Alexandria, Egypt. The Neo model, source code, training data, documentation,
23
+ and related materials are owned by TokenAI.
24
+
25
+ Contact: **info@tokenai.llc** · https://tokenai.llc
26
+
27
+ ## Summary
28
+
29
+ Neo is a compact TokenAI System One decision model for selecting actions from a
30
+ declared candidate set. It reads English application context and a typed
31
+ decision schema, then returns calibrated outputs for `choice`, `score`, and
32
+ `noul` in one encoder forward pass. It is not a conversational or free-form
33
+ text-generation model.
34
+
35
+ Canonical model: `tokenaii/Neo` — https://huggingface.co/tokenaii/Neo
36
+
37
+ Training dataset: `tokenaii/Neo-dataset` — https://huggingface.co/datasets/tokenaii/Neo-dataset
38
+
39
+ ## Model architecture
40
+
41
+ - Bidirectional Transformer encoder, initialized and trained from scratch.
42
+ - Approximately 110 million parameters.
43
+ - 8 Transformer layers; hidden size 512; 8 attention heads.
44
+ - Intermediate size 2048; maximum context 1024 tokens.
45
+ - Typed heads: `choice`, `score`, and `noul`.
46
+ - Choice head: up to 32 candidate options per question.
47
+ - Score head: up to 5 ordered levels.
48
+ - Noul head: binary probability for a yes/no or escalation decision.
49
+ - A request can contain up to 8 independent decision questions.
50
+
51
+ ## Input and output contract
52
 
53
+ The input is English text plus a predefined decision schema and candidate
54
+ options. The output contains typed probabilities and the selected index or
55
+ level. Applications should apply their own confidence thresholds, abstention
56
+ rules, validation, logging, and human review. Neo does not execute tools and
57
+ does not replace authorization or policy enforcement.
58
 
59
+ ## Intended uses
 
 
60
 
61
+ Non-commercial research and education for tool routing, workflow selection,
62
+ request classification, department routing, escalation detection, validation
63
+ gates, and selecting the next action in an agent pipeline.
64
 
65
+ ## Out-of-scope uses
 
 
 
 
 
66
 
67
+ The license prohibits commercial, monetized, paid, sponsored, client-facing,
68
+ production-business, or financially beneficial use. Do not use Neo for
69
+ unsupervised medical, legal, financial, employment, housing, admissions,
70
+ insurance, credit, safety-critical, government-benefit, law-enforcement, or
71
+ irreversible decisions.
72
 
73
+ ## Training data and procedure
74
 
75
+ The English-only corpus contains 400,000 synthetic records and 1,000,000
76
+ independent decisions: 799,882 training, 100,138 calibration, and 99,980 test
77
+ decisions. The catalog contains 20 tools; records commonly present one target
78
+ and five distractors.
 
79
 
80
+ The training run uses RLCD-inspired weighted soft-target training from scratch,
81
+ 20 epochs, approximately 15,640 optimizer steps, batch size 1024, bfloat16,
82
+ learning rate `0.0002`, choice loss weight `8.0`, and score/noul weights `1.0`.
83
+ The generator is identified as `neo-english-synthetic-v1` in the project
84
+ manifests. The final benchmark report is produced after training and is not
85
+ claimed by this card until verified.
86
 
87
+ ## Tokenizer
88
 
89
+ Neo uses the standard English `bert-base-uncased` tokenizer and vocabulary. A
90
+ custom tokenizer was not built for this release, so no custom tokenizer is
91
+ claimed or published under `tokenizer/`. The Transformer parameters are Neo's
92
+ own randomly initialized and trained weights; the tokenizer vocabulary is not
93
+ the model architecture.
94
 
95
+ ## Release status
 
 
 
96
 
97
+ The previous model weights were removed from this repository before the final
98
+ English-only training release. The current repository contains documentation,
99
+ configuration, tokenizer metadata, and licenses while the English training and
100
+ automatic evaluation pipeline complete. Do not infer benchmark accuracy from
101
+ the architecture or data counts. Verified results will be published only after
102
+ the local evaluation artifacts are reviewed.
103
 
104
+ ## Repository layout
 
 
105
 
106
+ - `licenses/MODEL_LICENSE.md` — full model license.
107
+ - `licenses/NOTICE.md` — attribution and ownership notice.
108
+ - `assets/neo-cover.png` — repository artwork.
109
+ - `docs/MODEL_SPECIFICATION.md` — technical specification.
110
+ - `tokenizer/` — reserved for a custom tokenizer only if one is actually built.
111
+ - `training_config.json` — reproducibility metadata when a release includes it.
112
 
113
+ ## License and attribution
114
 
115
+ Use is governed by [the TokenAI Neo Model License v3.0](licenses/MODEL_LICENSE.md).
116
+ Redistribution, renaming, rebranding, white-labeling, Derivative Models,
117
+ Commercial Use, and Financial Benefit are prohibited without written permission
118
+ from TokenAI. Every permitted downstream report or model must state:
119
 
120
+ > This work was trained using the TokenAI Neo Decision Model Dataset:
121
+ > `tokenaii/Neo-dataset`.
122
 
123
+ Permission requests must be sent to **info@tokenai.llc**. The complete legal
124
+ terms, trademark policy, notice-and-takedown process, patent reservation,
125
+ contribution/CLA rules, and dependency obligations are in the linked license.
 
 
 
126
 
127
+ ## Limitations
 
 
 
128
 
129
+ Neo is trained on synthetic English decision records and may fail on unseen
130
+ schemas, ambiguous contexts, distribution shifts, adversarial options, or
131
+ languages other than English. Calibration and accuracy are task-dependent.
132
+ Always validate with held-out, task-specific data and add human oversight for
133
+ high-impact workflows.