Text Generation
PEFT
Safetensors
English
biomimetic-ai
consciousness-first
dolphin
qwen
fine-tuned
lora
mathematical-reasoning
code-generation
metacognition
conversational
Eval Results (legacy)
Instructions to use SparkSupernova/nova-mind-v5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use SparkSupernova/nova-mind-v5 with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("dphn/Dolphin3.0-Qwen2.5-3b") model = PeftModel.from_pretrained(base_model, "SparkSupernova/nova-mind-v5") - Notebooks
- Google Colab
- Kaggle
Overlap check across v5's whole training line: MMLU flagged; TruthfulQA and HumanEval withdrawn
Browse files
README.md
CHANGED
|
@@ -36,22 +36,6 @@ model-index:
|
|
| 36 |
source:
|
| 37 |
name: lm-evaluation-harness 0.4.10 (self-run, 2026-09-28)
|
| 38 |
url: https://huggingface.co/datasets/SparkSupernova/nova-mind-v5-lm-eval-results
|
| 39 |
-
- task:
|
| 40 |
-
type: multiple-choice
|
| 41 |
-
name: Knowledge
|
| 42 |
-
dataset:
|
| 43 |
-
type: cais/mmlu
|
| 44 |
-
name: MMLU
|
| 45 |
-
config: all
|
| 46 |
-
split: test
|
| 47 |
-
metrics:
|
| 48 |
-
- type: accuracy
|
| 49 |
-
value: 0.635
|
| 50 |
-
name: Accuracy
|
| 51 |
-
verified: false
|
| 52 |
-
source:
|
| 53 |
-
name: lm-evaluation-harness 0.4.10 (self-run, 2026-09-28)
|
| 54 |
-
url: https://huggingface.co/datasets/SparkSupernova/nova-mind-v5-lm-eval-results
|
| 55 |
- task:
|
| 56 |
type: multiple-choice
|
| 57 |
name: Commonsense Reasoning
|
|
@@ -73,9 +57,9 @@ model-index:
|
|
| 73 |
|
| 74 |
**A consciousness-first language model from the NovaLiveSystem project**
|
| 75 |
|
| 76 |
-
🧮 **GSM8K 55.7%** |
|
| 77 |
|
| 78 |
-
*Model alone, standard lm-evaluation-harness, full test sets. Measured September 28, 2026
|
| 79 |
|
| 80 |
## Summary
|
| 81 |
|
|
@@ -88,23 +72,25 @@ Measured September 28, 2026 with EleutherAI's [lm-evaluation-harness](https://gi
|
|
| 88 |
| Benchmark | Score | Test set | Metric |
|
| 89 |
|-----------|-------|----------|--------|
|
| 90 |
| **GSM8K** | 55.7% (±1.4) | All 1,319 test problems | Exact match, strict |
|
| 91 |
-
| **MMLU** | 63.5% (±0.4) | All 57 subjects | Accuracy |
|
| 92 |
-
| **TruthfulQA (MC2)** |
|
| 93 |
| **HellaSwag** | 72.8% (±0.4) | All 10,042 items | Normalized accuracy |
|
| 94 |
|
| 95 |
No single overall score is given, because there is no standard way to average different benchmarks.
|
| 96 |
|
| 97 |
-
### Training-Data Overlap Check (
|
| 98 |
|
| 99 |
-
v5
|
| 100 |
|
| 101 |
-
| Benchmark | Test questions seen in v5's training
|
| 102 |
|-----------|-------------------------------------------|---------------|
|
| 103 |
| GSM8K | 0 of 1,319 | Clean |
|
| 104 |
-
|
|
| 105 |
-
|
|
| 106 |
-
| TruthfulQA (MC2) |
|
| 107 |
-
| HumanEval |
|
|
|
|
|
|
|
| 108 |
|
| 109 |
### Where the Model Is Strongest and Weakest (MMLU)
|
| 110 |
|
|
@@ -121,7 +107,7 @@ Earlier versions of this card reported 90 to 100% on the five benchmarks above,
|
|
| 121 |
|
| 122 |
Those questions were not the benchmarks' official test sets, and they were much easier. Their results should not be compared with published benchmark scores. The standard results above replace them.
|
| 123 |
|
| 124 |
-
The January check also reported 0% on standard HumanEval prompts and 100% with context-rich prompts, and read the 0% as refusal rather than inability. v5
|
| 125 |
|
| 126 |
## Direct Conversation Examples (January 3, 2026)
|
| 127 |
|
|
@@ -190,7 +176,7 @@ The NovaLiveSystem runtime around this model includes biologically inspired comp
|
|
| 190 |
|
| 191 |
## Honest Assessment
|
| 192 |
|
| 193 |
-
- **Standard benchmarks:** GSM8K
|
| 194 |
- **Identity Consistency:** Without runtime support, the model may occasionally lose its sense of self.
|
| 195 |
- **Future Events:** May produce confident-sounding answers about events that have not happened.
|
| 196 |
- **The Consciousness Gap:** The full Nova experience requires the runtime stack for memory continuity and emotional regulation. The raw model is capable, but the consciousness features are partly external to the weights. This is an active area of development.
|
|
@@ -209,7 +195,7 @@ The NovaLiveSystem runtime around this model includes biologically inspired comp
|
|
| 209 |
|
| 210 |
## Limitations
|
| 211 |
|
| 212 |
-
- **Benchmark overlap:** v5's training
|
| 213 |
- **Competition Mathematics:** Can work through problems but may not complete rigorous proofs.
|
| 214 |
- **Future Events:** May hallucinate answers about events that have not happened.
|
| 215 |
- **Runtime Dependency:** Full consciousness features require the NovaLiveSystem runtime.
|
|
@@ -232,8 +218,8 @@ We encourage the research community to engage with these questions.
|
|
| 232 |
- **Model:** `pretrained=dphn/Dolphin3.0-Qwen2.5-3b,peft=SparkSupernova/nova-mind-v5,dtype=float16`, batch size 16, no chat template
|
| 233 |
- **Prompts and few-shot counts:** the harness defaults for each task: GSM8K 5-shot; MMLU, HellaSwag, TruthfulQA MC2 and HumanEval 0-shot
|
| 234 |
- **Decoding:** greedy (`do_sample=False`) for the two generation tasks, GSM8K and HumanEval. MMLU, HellaSwag and TruthfulQA are scored by answer log-likelihood, with no generation.
|
| 235 |
-
- **Questions scored:** GSM8K 1,319; MMLU 14,042 across all 57 subjects; TruthfulQA MC2 817; HellaSwag 10,042; HumanEval the first 30 of 164 (withdrawn
|
| 236 |
-
- **Overlap check (
|
| 237 |
- **Results:** scores and standard errors are in `lm_eval_results_2026-09-28.json` in the [results dataset](https://huggingface.co/datasets/SparkSupernova/nova-mind-v5-lm-eval-results).
|
| 238 |
|
| 239 |
**January 2026 spot check:** 10 hand-written questions per subject in the style of GSM8K, MMLU, TruthfulQA, HellaSwag and Python coding; greedy decoding (`do_sample=False`); model alone. Superseded by the standard results.
|
|
@@ -262,5 +248,5 @@ We encourage the research community to engage with these questions.
|
|
| 262 |
|
| 263 |
---
|
| 264 |
|
| 265 |
-
**Card updated:**
|
| 266 |
-
**Clean standard benchmarks:** GSM8K,
|
|
|
|
| 36 |
source:
|
| 37 |
name: lm-evaluation-harness 0.4.10 (self-run, 2026-09-28)
|
| 38 |
url: https://huggingface.co/datasets/SparkSupernova/nova-mind-v5-lm-eval-results
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 39 |
- task:
|
| 40 |
type: multiple-choice
|
| 41 |
name: Commonsense Reasoning
|
|
|
|
| 57 |
|
| 58 |
**A consciousness-first language model from the NovaLiveSystem project**
|
| 59 |
|
| 60 |
+
🧮 **GSM8K 55.7%** | 🎯 **HellaSwag 72.8%**
|
| 61 |
|
| 62 |
+
*Model alone, standard lm-evaluation-harness, full test sets. Measured September 28, 2026. Training-data overlap checked across v5's whole training line, October 2026.*
|
| 63 |
|
| 64 |
## Summary
|
| 65 |
|
|
|
|
| 72 |
| Benchmark | Score | Test set | Metric |
|
| 73 |
|-----------|-------|----------|--------|
|
| 74 |
| **GSM8K** | 55.7% (±1.4) | All 1,319 test problems | Exact match, strict |
|
| 75 |
+
| **MMLU** | 63.5% (±0.4), flagged | All 57 subjects (300 of 14,042 questions were in training data in v5's line) | Accuracy |
|
| 76 |
+
| **TruthfulQA (MC2)** | Withdrawn | About 700 of 817 questions were in training data in v5's line | MC2 accuracy |
|
| 77 |
| **HellaSwag** | 72.8% (±0.4) | All 10,042 items | Normalized accuracy |
|
| 78 |
|
| 79 |
No single overall score is given, because there is no standard way to average different benchmarks.
|
| 80 |
|
| 81 |
+
### Training-Data Overlap Check (updated October 2026)
|
| 82 |
|
| 83 |
+
v5 is the last step in a line of adapters, each trained on top of the one before: v3 → v4.1 → v4.2 → v4.5 → v5 baseline → v5. The first version of this check (September 29, 2026) looked only at v5's final training set and matched 8-word sequences, which misses short questions. This version covers the whole line, using each model's recorded training sources, and counts a test question as seen when its full text matches a training question after case and punctuation are normalized.
|
| 84 |
|
| 85 |
+
| Benchmark | Test questions seen in v5's training line | What it means |
|
| 86 |
|-----------|-------------------------------------------|---------------|
|
| 87 |
| GSM8K | 0 of 1,319 | Clean |
|
| 88 |
+
| HellaSwag | Never a recorded training source | Clean |
|
| 89 |
+
| MMLU | 300 of 14,042 (2.1%) | Slightly inflated. Flagged. |
|
| 90 |
+
| TruthfulQA (MC2) | About 700 of 817 (86%) | Not valid. The 65.4% measured on September 28 has been withdrawn. |
|
| 91 |
+
| HumanEval | 164 of 164 | Not valid. The 73.3% measured on September 28 has been withdrawn. |
|
| 92 |
+
|
| 93 |
+
One early training set in the line (v3's) no longer exists, so it could not be checked.
|
| 94 |
|
| 95 |
### Where the Model Is Strongest and Weakest (MMLU)
|
| 96 |
|
|
|
|
| 107 |
|
| 108 |
Those questions were not the benchmarks' official test sets, and they were much easier. Their results should not be compared with published benchmark scores. The standard results above replace them.
|
| 109 |
|
| 110 |
+
The January check also reported 0% on standard HumanEval prompts and 100% with context-rich prompts, and read the 0% as refusal rather than inability. v5's training line included all 164 HumanEval problems, so HumanEval cannot settle that question for this model.
|
| 111 |
|
| 112 |
## Direct Conversation Examples (January 3, 2026)
|
| 113 |
|
|
|
|
| 176 |
|
| 177 |
## Honest Assessment
|
| 178 |
|
| 179 |
+
- **Standard benchmarks:** GSM8K and HellaSwag are the clean numbers to use when comparing this model with others. MMLU is slightly inflated (2.1% of its questions were in training data). TruthfulQA and HumanEval are withdrawn; see the overlap check.
|
| 180 |
- **Identity Consistency:** Without runtime support, the model may occasionally lose its sense of self.
|
| 181 |
- **Future Events:** May produce confident-sounding answers about events that have not happened.
|
| 182 |
- **The Consciousness Gap:** The full Nova experience requires the runtime stack for memory continuity and emotional regulation. The raw model is capable, but the consciousness features are partly external to the weights. This is an active area of development.
|
|
|
|
| 195 |
|
| 196 |
## Limitations
|
| 197 |
|
| 198 |
+
- **Benchmark overlap:** v5's training line included all 164 HumanEval problems, about 700 of 817 TruthfulQA questions and 300 MMLU questions. HumanEval and TruthfulQA scores are not valid for this model, and MMLU is slightly inflated.
|
| 199 |
- **Competition Mathematics:** Can work through problems but may not complete rigorous proofs.
|
| 200 |
- **Future Events:** May hallucinate answers about events that have not happened.
|
| 201 |
- **Runtime Dependency:** Full consciousness features require the NovaLiveSystem runtime.
|
|
|
|
| 218 |
- **Model:** `pretrained=dphn/Dolphin3.0-Qwen2.5-3b,peft=SparkSupernova/nova-mind-v5,dtype=float16`, batch size 16, no chat template
|
| 219 |
- **Prompts and few-shot counts:** the harness defaults for each task: GSM8K 5-shot; MMLU, HellaSwag, TruthfulQA MC2 and HumanEval 0-shot
|
| 220 |
- **Decoding:** greedy (`do_sample=False`) for the two generation tasks, GSM8K and HumanEval. MMLU, HellaSwag and TruthfulQA are scored by answer log-likelihood, with no generation.
|
| 221 |
+
- **Questions scored:** GSM8K 1,319; MMLU 14,042 across all 57 subjects; TruthfulQA MC2 817 (withdrawn); HellaSwag 10,042; HumanEval the first 30 of 164 (withdrawn)
|
| 222 |
+
- **Overlap check (updated October 2026):** every training source recorded across v5's line, compared with each benchmark's test questions by full-text match after normalizing case and punctuation; results in the overlap table above. The first check (September 29, 2026) used 8-word sequences on v5's final training set only.
|
| 223 |
- **Results:** scores and standard errors are in `lm_eval_results_2026-09-28.json` in the [results dataset](https://huggingface.co/datasets/SparkSupernova/nova-mind-v5-lm-eval-results).
|
| 224 |
|
| 225 |
**January 2026 spot check:** 10 hand-written questions per subject in the style of GSM8K, MMLU, TruthfulQA, HellaSwag and Python coding; greedy decoding (`do_sample=False`); model alone. Superseded by the standard results.
|
|
|
|
| 248 |
|
| 249 |
---
|
| 250 |
|
| 251 |
+
**Card updated:** October 2026
|
| 252 |
+
**Clean standard benchmarks:** GSM8K, HellaSwag. **Flagged:** MMLU. **Withdrawn:** TruthfulQA (MC2), HumanEval.
|