Instructions to use seanmamasde/gemma4_ssset with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use seanmamasde/gemma4_ssset with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Gemma4 Salem-Spencer adapters
Rule-SFT and RL LoRA adapters for budgeted Salem-Spencer set construction. This repository contains adapters and tokenizer files; base weights and optimizer states are not included.
Exact base model
- Repository:
google/gemma-4-26B-A4B - Revision:
24548b62aa021d562695c04aaf7758a1ea47990b - Pinned tree: https://huggingface.co/google/gemma-4-26B-A4B/tree/24548b62aa021d562695c04aaf7758a1ea47990b
This is the pre-trained, non-IT base. The original model is distributed under its upstream Apache-2.0 terms; see its model card and license link.
Layout and dependency chain
sft/ rule-SFT adapter, completed update 32
rl/checkpoint-200/ RL adapter at completed step 200
rl/checkpoint-449/ RL adapter at completed step 449
rl/checkpoint-910/ RL adapter at completed step 910
The required order is:
Pinned raw base -> apply and merge sft/ -> apply the selected rl/ adapter.
The RL adapters were trained against the SFT-merged reference. Applying an RL adapter directly to the raw Google base skips the SFT update and produces a different model. Each RL adapter intentionally leaves automatic base selection unset; supply the reconstructed SFT-merged model explicitly as below.
The SFT stage used 32 updates and 256 unique single-action rule examples. These
teach legal ADD/REMOVE/DONE JSON actions; they are not complete-set solution targets.
The RL adapters use rank 32, alpha 32, FP32 trainable adapters and the
budget-best-v1 search protocol.
Load a checkpoint
from pathlib import Path
import torch
from huggingface_hub import snapshot_download
from peft import PeftModel
from transformers import AutoTokenizer, Gemma4ForConditionalGeneration
BASE = "google/gemma-4-26B-A4B"
BASE_REVISION = "24548b62aa021d562695c04aaf7758a1ea47990b"
ADAPTERS = "seanmamasde/gemma4_ssset"
step = 910 # 200, 449 or 910
path = Path(snapshot_download(ADAPTERS, allow_patterns=["sft/*", f"rl/checkpoint-{step}/*"]))
tokenizer = AutoTokenizer.from_pretrained(path / "sft")
base = Gemma4ForConditionalGeneration.from_pretrained(
BASE, revision=BASE_REVISION, dtype=torch.bfloat16,
attn_implementation="sdpa", device_map="auto",
)
sft = PeftModel.from_pretrained(
base, path / "sft", is_trainable=False, autocast_adapter_dtype=False,
)
sft_base = sft.merge_and_unload(safe_merge=True)
model = PeftModel.from_pretrained(
sft_base, path / f"rl/checkpoint-{step}", is_trainable=False,
autocast_adapter_dtype=True,
).eval()
Reference software versions: Transformers 5.14.1, PEFT 0.20.0, PyTorch 2.11.0+cu129, and vLLM 0.26.0+cu129. Match precision and merge behavior when reproducing a run. The saved SFT adapter is BF16; keep it BF16 during the SFT merge rather than automatically upcasting it to FP32. The saved RL adapters are FP32. Re-merging can still differ in low-order bits across numerical backends; the originally merged SFT weight files are not part of this adapter-only release.
Prompt and evaluation protocol
These adapters consume structured state-table prompts with full-menu ADD/REMOVE
IDs. They are action policies, not general chat instruction models. The
llmcosolver.ssset compatibility profile and Gemma4 example provide that protocol.
Core search settings were H192/L80/K8/E960/R2, 24 raw proposals, max backup, UCB 0.4, temperature 1, top_p 1 and top_k=-1. Reward uses actual cardinality gain with a 0.25 efficiency bonus on positive gains. RL used two policy updates followed by one separate expert imitation update.
Published checkpoints 200/449/910 are the actual saved adapters used in the major inference runs. The last checkpoint is 910. Evaluation time windows and prompt visibility changed across runs: 200 used 30-minute windows, 449 used 60-minute windows, and 910 used 180-minute windows (with an accepted partial final target). The 449/910 evaluations exposed full REMOVE menus. Those results should not be interpreted as a time-budget-controlled comparison of weights alone.
Data
The primary n=1..115 witness corpus is available at https://huggingface.co/datasets/seanmamasde/extremal .
- Downloads last month
- -
Model tree for seanmamasde/gemma4_ssset
Base model
google/gemma-4-26B-A4B