jsbai-aaron commited on
Commit
12e8eb1
·
verified ·
1 Parent(s): 7971753

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +3 -1
README.md CHANGED
@@ -37,7 +37,7 @@ We reserved 121 real software bugs that the model never saw during training. Bef
37
 
38
  | Benchmark | Qwen3.5-4B (base) | **JSBAI-Coder-4B** | JSBAI-Coder-4B-NVFP4 |
39
  |---|---|---|---|
40
- | **Generalization test** (121 unseen bugs, tests run to verify) | 10.1% | **82.9%** | coming soon |
41
  | **Live-60** (60 real-world engineering tasks, solved end-to-end in containers) | 15.0% | **21.7%** | 15.0% |
42
  | **Instruction-following** (IFEval) | 84.66 | **87.21** | 86.37 |
43
  | MMLU-Pro | 64.0% | **70.0%** | 66.85% |
@@ -45,6 +45,8 @@ We reserved 121 real software bugs that the model never saw during training. Bef
45
 
46
  The instruction-following score *improved* over the base model. The coding gains cost nothing on general quality. Gains of this kind usually trade one for the other.
47
 
 
 
48
  Our decontamination protocol is published with the model: none of these benchmark problems overlap the training data.
49
 
50
  ## What it does
 
37
 
38
  | Benchmark | Qwen3.5-4B (base) | **JSBAI-Coder-4B** | JSBAI-Coder-4B-NVFP4 |
39
  |---|---|---|---|
40
+ | **Generalization test** (121 unseen bugs, tests run to verify) | 10.1% | **82.9%** | 38.0%* |
41
  | **Live-60** (60 real-world engineering tasks, solved end-to-end in containers) | 15.0% | **21.7%** | 15.0% |
42
  | **Instruction-following** (IFEval) | 84.66 | **87.21** | 86.37 |
43
  | MMLU-Pro | 64.0% | **70.0%** | 66.85% |
 
45
 
46
  The instruction-following score *improved* over the base model. The coding gains cost nothing on general quality. Gains of this kind usually trade one for the other.
47
 
48
+ *NVFP4 generalization: 12/32 on a 32-instance subset (the same slice our comparisons use). The quantization costs roughly half the generalization capability.
49
+
50
  Our decontamination protocol is published with the model: none of these benchmark problems overlap the training data.
51
 
52
  ## What it does