You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Gemma4 Salem-Spencer adapters

Rule-SFT and RL LoRA adapters for budgeted Salem-Spencer set construction. This repository contains adapters and tokenizer files; base weights and optimizer states are not included.

Exact base model

This is the pre-trained, non-IT base. The original model is distributed under its upstream Apache-2.0 terms; see its model card and license link.

Layout and dependency chain

sft/                    rule-SFT adapter, completed update 32
rl/checkpoint-200/       RL adapter at completed step 200
rl/checkpoint-449/       RL adapter at completed step 449
rl/checkpoint-910/       RL adapter at completed step 910

The required order is:

Pinned raw base -> apply and merge sft/ -> apply the selected rl/ adapter.

The RL adapters were trained against the SFT-merged reference. Applying an RL adapter directly to the raw Google base skips the SFT update and produces a different model. Each RL adapter intentionally leaves automatic base selection unset; supply the reconstructed SFT-merged model explicitly as below.

The SFT stage used 32 updates and 256 unique single-action rule examples. These teach legal ADD/REMOVE/DONE JSON actions; they are not complete-set solution targets. The RL adapters use rank 32, alpha 32, FP32 trainable adapters and the budget-best-v1 search protocol.

Load a checkpoint

from pathlib import Path
import torch
from huggingface_hub import snapshot_download
from peft import PeftModel
from transformers import AutoTokenizer, Gemma4ForConditionalGeneration

BASE = "google/gemma-4-26B-A4B"
BASE_REVISION = "24548b62aa021d562695c04aaf7758a1ea47990b"
ADAPTERS = "seanmamasde/gemma4_ssset"
step = 910  # 200, 449 or 910
path = Path(snapshot_download(ADAPTERS, allow_patterns=["sft/*", f"rl/checkpoint-{step}/*"]))

tokenizer = AutoTokenizer.from_pretrained(path / "sft")
base = Gemma4ForConditionalGeneration.from_pretrained(
    BASE, revision=BASE_REVISION, dtype=torch.bfloat16,
    attn_implementation="sdpa", device_map="auto",
)
sft = PeftModel.from_pretrained(
    base, path / "sft", is_trainable=False, autocast_adapter_dtype=False,
)
sft_base = sft.merge_and_unload(safe_merge=True)
model = PeftModel.from_pretrained(
    sft_base, path / f"rl/checkpoint-{step}", is_trainable=False,
    autocast_adapter_dtype=True,
).eval()

Reference software versions: Transformers 5.14.1, PEFT 0.20.0, PyTorch 2.11.0+cu129, and vLLM 0.26.0+cu129. Match precision and merge behavior when reproducing a run. The saved SFT adapter is BF16; keep it BF16 during the SFT merge rather than automatically upcasting it to FP32. The saved RL adapters are FP32. Re-merging can still differ in low-order bits across numerical backends; the originally merged SFT weight files are not part of this adapter-only release.

Prompt and evaluation protocol

These adapters consume structured state-table prompts with full-menu ADD/REMOVE IDs. They are action policies, not general chat instruction models. The llmcosolver.ssset compatibility profile and Gemma4 example provide that protocol.

Core search settings were H192/L80/K8/E960/R2, 24 raw proposals, max backup, UCB 0.4, temperature 1, top_p 1 and top_k=-1. Reward uses actual cardinality gain with a 0.25 efficiency bonus on positive gains. RL used two policy updates followed by one separate expert imitation update.

Published checkpoints 200/449/910 are the actual saved adapters used in the major inference runs. The last checkpoint is 910. Evaluation time windows and prompt visibility changed across runs: 200 used 30-minute windows, 449 used 60-minute windows, and 910 used 180-minute windows (with an accepted partial final target). The 449/910 evaluations exposed full REMOVE menus. Those results should not be interpreted as a time-budget-controlled comparison of weights alone.

Data

The primary n=1..115 witness corpus is available at https://huggingface.co/datasets/seanmamasde/extremal .

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for seanmamasde/gemma4_ssset

Adapter
(2)
this model