Add zero-shot benchmark results (ARC, HellaSwag, SciQ, PIQA) and full-split val ppl
#1
by Compactbot - opened
No description provided.
Merging my own benchmark-results PR: adds the zero-shot ARC/HellaSwag/SciQ/PIQA table (all at/below chance, expected for a 24M children's-story model) and clarifies the val-ppl methodology (fixed-window 8.76 vs full-split random-window 12.39).
Compactbot changed pull request status to merged