Slayer149 Balanced ARC-E fine-tune v2

This is a separate v2 release; the original v1 repository remains unchanged. This checkpoint fine-tunes the 149,333,081-parameter Slayer149 Balanced base on the clean ARC-Easy training split with general-text replay. It is a research base model, not an instruction-tuned assistant.

Evaluation

Benchmark Result Count
ARC-Easy raw accuracy 64.90% 1,542 / 2,376
WikiText-2 byte perplexity 2.2719 full test split
WikiText-2 token perplexity 29.277 full test split
BLiMP raw accuracy 78.89% 52,859 / 67,000

Evaluation used the pinned GLINT-derived harness: 256-token context and no added BOS. ARC uses raw summed answer log-likelihood; the board metric is raw accuracy. The checkpoint was selected by ARC-Easy validation acc_norm before test evaluation. Test results did not select the checkpoint. Full result files and per-example benchmark details are included in this repository.

Fine-tuning provenance

  • Base: SlayerLab/Slayer149-balanced, complete step 59,368 export; base weights SHA-256 b1353f99662b209dca70a50eb06fdb99bba4671486a1695236b2752514cfce63.
  • Fine-tune: epoch 2, update 1,118; 149,333,081 unique parameters.
  • Objective: ARC-Easy answer-choice ranking and correct-answer cross-entropy, mixed with clean general-text replay.
  • Training data: pinned ARC-Easy train split only; validation/test examples were excluded.
  • Selected checkpoint SHA-256: 79e4d06b1b95408c4b47e6d9fab34ed2a26609997b80d45698e6ec23db8724e8.
  • Safetensors SHA-256: 4663361d3acdae2ce766f980c549649f9db21a77a26213034902370f775f69eb.
  • Tokenizer SHA-256: d1ffe47f6fd8cc19ba7c3e5d2537699d6a99d09b584d3319944f4e40fe6c44ea.
  • Dataset revision: 210d026faf9955653af8916fad021475a3f00453.

The base model is Apache-2.0. ARC is credited to AI2; its dataset card lists CC-BY-SA-4.0. The dataset is not redistributed here.

Load locally

from load_model import load_model
import torch
model, tokenizer = load_model(".", device="cpu")
ids = tokenizer.encode("What is the result?", add_special_tokens=False).ids
with torch.inference_mode():
    logits, _ = model(torch.tensor([ids]))

This package uses the included native PyTorch architecture and tokenizer; it does not require Transformers remote code.

Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SlayerLab/Slayer149-Balanced-ARC-E-ft-v2

Finetuned
(2)
this model