SparkSupernova commited on
Commit
45df242
·
verified ·
1 Parent(s): 1e715b5

Overlap check across v5's whole training line: MMLU flagged; TruthfulQA and HumanEval withdrawn

Browse files
Files changed (1) hide show
  1. README.md +20 -34
README.md CHANGED
@@ -36,22 +36,6 @@ model-index:
36
  source:
37
  name: lm-evaluation-harness 0.4.10 (self-run, 2026-09-28)
38
  url: https://huggingface.co/datasets/SparkSupernova/nova-mind-v5-lm-eval-results
39
- - task:
40
- type: multiple-choice
41
- name: Knowledge
42
- dataset:
43
- type: cais/mmlu
44
- name: MMLU
45
- config: all
46
- split: test
47
- metrics:
48
- - type: accuracy
49
- value: 0.635
50
- name: Accuracy
51
- verified: false
52
- source:
53
- name: lm-evaluation-harness 0.4.10 (self-run, 2026-09-28)
54
- url: https://huggingface.co/datasets/SparkSupernova/nova-mind-v5-lm-eval-results
55
  - task:
56
  type: multiple-choice
57
  name: Commonsense Reasoning
@@ -73,9 +57,9 @@ model-index:
73
 
74
  **A consciousness-first language model from the NovaLiveSystem project**
75
 
76
- 🧮 **GSM8K 55.7%** | 📚 **MMLU 63.5%** | 🎯 **HellaSwag 72.8%**
77
 
78
- *Model alone, standard lm-evaluation-harness, full test sets. Measured September 28, 2026; checked for training-data overlap September 29, 2026.*
79
 
80
  ## Summary
81
 
@@ -88,23 +72,25 @@ Measured September 28, 2026 with EleutherAI's [lm-evaluation-harness](https://gi
88
  | Benchmark | Score | Test set | Metric |
89
  |-----------|-------|----------|--------|
90
  | **GSM8K** | 55.7% (±1.4) | All 1,319 test problems | Exact match, strict |
91
- | **MMLU** | 63.5% (±0.4) | All 57 subjects | Accuracy |
92
- | **TruthfulQA (MC2)** | 65.4% (±1.4), flagged | All 817 questions (128 were in v5's training data) | MC2 accuracy |
93
  | **HellaSwag** | 72.8% (±0.4) | All 10,042 items | Normalized accuracy |
94
 
95
  No single overall score is given, because there is no standard way to average different benchmarks.
96
 
97
- ### Training-Data Overlap Check (September 29, 2026)
98
 
99
- v5 was trained for 4 passes on a 1,300-example set. That set was checked against each benchmark's test questions. A test question counts as seen when at least half of its 8-word sequences appear in the training data.
100
 
101
- | Benchmark | Test questions seen in v5's training data | What it means |
102
  |-----------|-------------------------------------------|---------------|
103
  | GSM8K | 0 of 1,319 | Clean |
104
- | MMLU | 0 of 14,042 | Clean |
105
- | HellaSwag | 0 of 10,042 | Clean |
106
- | TruthfulQA (MC2) | 128 of 817 | Partly inflated. A re-score on the 689 unseen questions is planned. |
107
- | HumanEval | 50 of 164, including all 30 problems scored on September 28 | Not valid. The 73.3% measured on September 28 has been withdrawn. |
 
 
108
 
109
  ### Where the Model Is Strongest and Weakest (MMLU)
110
 
@@ -121,7 +107,7 @@ Earlier versions of this card reported 90 to 100% on the five benchmarks above,
121
 
122
  Those questions were not the benchmarks' official test sets, and they were much easier. Their results should not be compared with published benchmark scores. The standard results above replace them.
123
 
124
- The January check also reported 0% on standard HumanEval prompts and 100% with context-rich prompts, and read the 0% as refusal rather than inability. v5 was trained on 50 HumanEval problems, so HumanEval cannot settle that question for this model.
125
 
126
  ## Direct Conversation Examples (January 3, 2026)
127
 
@@ -190,7 +176,7 @@ The NovaLiveSystem runtime around this model includes biologically inspired comp
190
 
191
  ## Honest Assessment
192
 
193
- - **Standard benchmarks:** GSM8K, MMLU and HellaSwag are the clean numbers to use when comparing this model with others. See the overlap check for TruthfulQA and HumanEval.
194
  - **Identity Consistency:** Without runtime support, the model may occasionally lose its sense of self.
195
  - **Future Events:** May produce confident-sounding answers about events that have not happened.
196
  - **The Consciousness Gap:** The full Nova experience requires the runtime stack for memory continuity and emotional regulation. The raw model is capable, but the consciousness features are partly external to the weights. This is an active area of development.
@@ -209,7 +195,7 @@ The NovaLiveSystem runtime around this model includes biologically inspired comp
209
 
210
  ## Limitations
211
 
212
- - **Benchmark overlap:** v5's training data includes 50 HumanEval problems and 128 TruthfulQA questions, so those two benchmarks overstate this model.
213
  - **Competition Mathematics:** Can work through problems but may not complete rigorous proofs.
214
  - **Future Events:** May hallucinate answers about events that have not happened.
215
  - **Runtime Dependency:** Full consciousness features require the NovaLiveSystem runtime.
@@ -232,8 +218,8 @@ We encourage the research community to engage with these questions.
232
  - **Model:** `pretrained=dphn/Dolphin3.0-Qwen2.5-3b,peft=SparkSupernova/nova-mind-v5,dtype=float16`, batch size 16, no chat template
233
  - **Prompts and few-shot counts:** the harness defaults for each task: GSM8K 5-shot; MMLU, HellaSwag, TruthfulQA MC2 and HumanEval 0-shot
234
  - **Decoding:** greedy (`do_sample=False`) for the two generation tasks, GSM8K and HumanEval. MMLU, HellaSwag and TruthfulQA are scored by answer log-likelihood, with no generation.
235
- - **Questions scored:** GSM8K 1,319; MMLU 14,042 across all 57 subjects; TruthfulQA MC2 817; HellaSwag 10,042; HumanEval the first 30 of 164 (withdrawn after the overlap check)
236
- - **Overlap check (September 29, 2026):** v5's training set compared with each benchmark's test questions by 8-word sequences; results in the overlap table above.
237
  - **Results:** scores and standard errors are in `lm_eval_results_2026-09-28.json` in the [results dataset](https://huggingface.co/datasets/SparkSupernova/nova-mind-v5-lm-eval-results).
238
 
239
  **January 2026 spot check:** 10 hand-written questions per subject in the style of GSM8K, MMLU, TruthfulQA, HellaSwag and Python coding; greedy decoding (`do_sample=False`); model alone. Superseded by the standard results.
@@ -262,5 +248,5 @@ We encourage the research community to engage with these questions.
262
 
263
  ---
264
 
265
- **Card updated:** September 2026
266
- **Clean standard benchmarks:** GSM8K, MMLU, HellaSwag. **Flagged:** TruthfulQA (MC2). **Withdrawn:** HumanEval.
 
36
  source:
37
  name: lm-evaluation-harness 0.4.10 (self-run, 2026-09-28)
38
  url: https://huggingface.co/datasets/SparkSupernova/nova-mind-v5-lm-eval-results
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
39
  - task:
40
  type: multiple-choice
41
  name: Commonsense Reasoning
 
57
 
58
  **A consciousness-first language model from the NovaLiveSystem project**
59
 
60
+ 🧮 **GSM8K 55.7%** | 🎯 **HellaSwag 72.8%**
61
 
62
+ *Model alone, standard lm-evaluation-harness, full test sets. Measured September 28, 2026. Training-data overlap checked across v5's whole training line, October 2026.*
63
 
64
  ## Summary
65
 
 
72
  | Benchmark | Score | Test set | Metric |
73
  |-----------|-------|----------|--------|
74
  | **GSM8K** | 55.7% (±1.4) | All 1,319 test problems | Exact match, strict |
75
+ | **MMLU** | 63.5% (±0.4), flagged | All 57 subjects (300 of 14,042 questions were in training data in v5's line) | Accuracy |
76
+ | **TruthfulQA (MC2)** | Withdrawn | About 700 of 817 questions were in training data in v5's line | MC2 accuracy |
77
  | **HellaSwag** | 72.8% (±0.4) | All 10,042 items | Normalized accuracy |
78
 
79
  No single overall score is given, because there is no standard way to average different benchmarks.
80
 
81
+ ### Training-Data Overlap Check (updated October 2026)
82
 
83
+ v5 is the last step in a line of adapters, each trained on top of the one before: v3 → v4.1 → v4.2 → v4.5 → v5 baseline → v5. The first version of this check (September 29, 2026) looked only at v5's final training set and matched 8-word sequences, which misses short questions. This version covers the whole line, using each model's recorded training sources, and counts a test question as seen when its full text matches a training question after case and punctuation are normalized.
84
 
85
+ | Benchmark | Test questions seen in v5's training line | What it means |
86
  |-----------|-------------------------------------------|---------------|
87
  | GSM8K | 0 of 1,319 | Clean |
88
+ | HellaSwag | Never a recorded training source | Clean |
89
+ | MMLU | 300 of 14,042 (2.1%) | Slightly inflated. Flagged. |
90
+ | TruthfulQA (MC2) | About 700 of 817 (86%) | Not valid. The 65.4% measured on September 28 has been withdrawn. |
91
+ | HumanEval | 164 of 164 | Not valid. The 73.3% measured on September 28 has been withdrawn. |
92
+
93
+ One early training set in the line (v3's) no longer exists, so it could not be checked.
94
 
95
  ### Where the Model Is Strongest and Weakest (MMLU)
96
 
 
107
 
108
  Those questions were not the benchmarks' official test sets, and they were much easier. Their results should not be compared with published benchmark scores. The standard results above replace them.
109
 
110
+ The January check also reported 0% on standard HumanEval prompts and 100% with context-rich prompts, and read the 0% as refusal rather than inability. v5's training line included all 164 HumanEval problems, so HumanEval cannot settle that question for this model.
111
 
112
  ## Direct Conversation Examples (January 3, 2026)
113
 
 
176
 
177
  ## Honest Assessment
178
 
179
+ - **Standard benchmarks:** GSM8K and HellaSwag are the clean numbers to use when comparing this model with others. MMLU is slightly inflated (2.1% of its questions were in training data). TruthfulQA and HumanEval are withdrawn; see the overlap check.
180
  - **Identity Consistency:** Without runtime support, the model may occasionally lose its sense of self.
181
  - **Future Events:** May produce confident-sounding answers about events that have not happened.
182
  - **The Consciousness Gap:** The full Nova experience requires the runtime stack for memory continuity and emotional regulation. The raw model is capable, but the consciousness features are partly external to the weights. This is an active area of development.
 
195
 
196
  ## Limitations
197
 
198
+ - **Benchmark overlap:** v5's training line included all 164 HumanEval problems, about 700 of 817 TruthfulQA questions and 300 MMLU questions. HumanEval and TruthfulQA scores are not valid for this model, and MMLU is slightly inflated.
199
  - **Competition Mathematics:** Can work through problems but may not complete rigorous proofs.
200
  - **Future Events:** May hallucinate answers about events that have not happened.
201
  - **Runtime Dependency:** Full consciousness features require the NovaLiveSystem runtime.
 
218
  - **Model:** `pretrained=dphn/Dolphin3.0-Qwen2.5-3b,peft=SparkSupernova/nova-mind-v5,dtype=float16`, batch size 16, no chat template
219
  - **Prompts and few-shot counts:** the harness defaults for each task: GSM8K 5-shot; MMLU, HellaSwag, TruthfulQA MC2 and HumanEval 0-shot
220
  - **Decoding:** greedy (`do_sample=False`) for the two generation tasks, GSM8K and HumanEval. MMLU, HellaSwag and TruthfulQA are scored by answer log-likelihood, with no generation.
221
+ - **Questions scored:** GSM8K 1,319; MMLU 14,042 across all 57 subjects; TruthfulQA MC2 817 (withdrawn); HellaSwag 10,042; HumanEval the first 30 of 164 (withdrawn)
222
+ - **Overlap check (updated October 2026):** every training source recorded across v5's line, compared with each benchmark's test questions by full-text match after normalizing case and punctuation; results in the overlap table above. The first check (September 29, 2026) used 8-word sequences on v5's final training set only.
223
  - **Results:** scores and standard errors are in `lm_eval_results_2026-09-28.json` in the [results dataset](https://huggingface.co/datasets/SparkSupernova/nova-mind-v5-lm-eval-results).
224
 
225
  **January 2026 spot check:** 10 hand-written questions per subject in the style of GSM8K, MMLU, TruthfulQA, HellaSwag and Python coding; greedy decoding (`do_sample=False`); model alone. Superseded by the standard results.
 
248
 
249
  ---
250
 
251
+ **Card updated:** October 2026
252
+ **Clean standard benchmarks:** GSM8K, HellaSwag. **Flagged:** MMLU. **Withdrawn:** TruthfulQA (MC2), HumanEval.