VISTA-24M / EVALUATION.md
AwakeningOS's picture
Release VISTA-24M: model, architecture diagrams, training recipe and evaluation evidence
9287d39 verified
|
Raw History Blame Contribute Delete
3.04 kB

Evaluation files and interpretation

Official evaluator: https://github.com/babylm-org/babylm-eval

Recorded revision: 6f825c291e2c4c78ad33b1935fd64d45f52642dc.

Files

File or folder Scope
evaluation/final_scores.json Selected 80M raw model, downstream scores, AoA trajectory, aggregate formulas' results and snapshot-comparison record
evaluation/checkpoints/010M.json … 100M.json Full zero-shot results; checkpoint/freeze identities and sample counts
evaluation/learning_curve.csv Same checkpoint results in a tabular form
evaluation/predictions/zero_shot/ Official-format predictions for the selected 80M model, including Reading
evaluation/predictions/finetune/ Official-format predictions from seven task-specific fine-tuning runs
evaluation/aoa_surprisal_no_context.json All 152,095 numerical records (19 checkpoints × 8,005 contexts); context text omitted
evaluation/aoa_score.json Raw correlation from the official AoA scoring implementation

NLP = mean(BLiMP, Supplement, EWoK, Entity, COMPS, mean(PIQA_parallel, PIQA_nonparallel), SuperGLUE).

human_like = mean(Reading, 100 * AoA_correlation).

overall = (7 * NLP + 2 * human_like) / 9.

The learning-curve chart omits SuperGLUE, Reading and AoA and averages the six zero-shot categories only. Global PIQA contributes once. Its parallel and nonparallel splits receive equal weight. Reading uses the mean of the evaluator's eye-tracking and self-paced reading scores. AoA is a correlation, not MSE or accuracy.

The score summaries retain the evaluator's reported precision. Public-table comparison uses the saved 2026-09-15 snapshot and does not constitute an official placement. Final scores and all release claims refer to the raw 80M model; task fine-tuning is a separate operation.

Re-running

Use the pinned official repository and its data download instructions. The model backend is causal; the track is strict-small. Load AwakeningOS/VISTA-24M, with trust_remote_code=True, and use the supplied tokenizer. The standard intermediate names are available as Hub branches. Use the official launchers for full zero-shot scoring, Reading, AoA and fine-tuning. The original local AoA run modified only checkpoint resolution to read local folders; the released branches allow standard Hub resolution.

Original score-producing inference used BF16 on CUDA with the FlashAttention backend. To select that backend, load AutoConfig, set config.dense_config['backend']='flash', then pass that config to from_pretrained; install a compatible FlashAttention build. The SDPA default is convenient for CPU and general use. Kernel/dtype changes can affect borderline preferences, so report those conditions with any rerun.

Fast evaluations for the early 1–9M checkpoints and a submission-ready collated file are outside this package. The public checkpoint series makes those further evaluations possible. No benchmark is started by importing the model or opening this repository.