Text Generation
GGUF
code
imatrix
conversational
Nerdsking commited on
Commit
1a8c757
·
verified ·
1 Parent(s): 75114f1

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +6 -6
README.md CHANGED
@@ -94,9 +94,9 @@ We did not considered it for our score, but "if" considered those extra 5 questi
94
  <td>Evaluated score (zero-shot, strict HumanEval pass@1, using unquantized weigths bf16)</td>
95
  </tr>
96
  <tr>
97
- <td>StarCoder2-3B</td>
98
- <td>~33.6</td>
99
- <td>Reported in third-party performance overview; may differ by protocol</td>
100
  </tr>
101
  <tr>
102
  <td>Qwen/Qwen3-4B-GGUF (<u>bigger</u> size model)</td>
@@ -104,9 +104,9 @@ We did not considered it for our score, but "if" considered those extra 5 questi
104
  <td>Indicative proxy from published code-task performance breakdowns (not a strict HumanEval pass@1)</td>
105
  </tr>
106
  <tr>
107
- <td>Wizard Coder 3B*</td>
108
- <td>~31.6 (estimate)</td>
109
- <td>Indicative proxy from published code-task performance breakdowns (not a strict HumanEval pass@1)</td>
110
  </tr>
111
  <tr>
112
  <td>CodeLlama 7B‑Python (<u>much bigger</u> size model)</td>
 
94
  <td>Evaluated score (zero-shot, strict HumanEval pass@1, using unquantized weigths bf16)</td>
95
  </tr>
96
  <tr>
97
+ <td>Gemma 3 27B</td>
98
+ <td>76</td>
99
+ <td>Our model surpasses a model 9x bigger by 12%</td>
100
  </tr>
101
  <tr>
102
  <td>Qwen/Qwen3-4B-GGUF (<u>bigger</u> size model)</td>
 
104
  <td>Indicative proxy from published code-task performance breakdowns (not a strict HumanEval pass@1)</td>
105
  </tr>
106
  <tr>
107
+ <td>Llama 3.3 70B</td>
108
+ <td>~83% </td>
109
+ <td>Our model surpasses a model 23x bigger</td>
110
  </tr>
111
  <tr>
112
  <td>CodeLlama 7B‑Python (<u>much bigger</u> size model)</td>