ZeroHero Crafter 5M โ€” SFT epoch 12

A 5,033,600-parameter, randomly initialized Qwen3-architecture agent trained with supervised imitation of Crafter expert reasoning and actions. These are original supervised pretraining checkpoints, before PPO. They do not contain pretrained Qwen weights and use their own tokenizer.

Versions

Revision SFT epoch Optimizer updates Validation response-token CE Validation response-token accuracy
epoch-02 2 1,310 2.137118 0.506485
epoch-12 12 7,860 1.684756 0.580853

main points to epoch 12 after publication. This revision contains epoch 12. These are teacher-forced validation diagnostics, not environment success rates.

Architecture and training

  • 6 layers; hidden size 256; intermediate size 608; 8 attention heads and 4 KV heads.
  • Context: 4,096 tokens; shared byte-level BPE vocabulary: 4,096 tokens.
  • Input: BALROG Crafter instructions, current public observation/inventory, and four previous public observations and executed actions.
  • Output: free-form reasoning followed by ACTION: <action>.
  • Random initialization; training seed 7101; response-only cross entropy with prompt labels masked.
  • AdamW; learning rate 2e-4 with 100-update warmup; effective batch 32; microbatch 8.
  • Train/validation split: 96/10 expert worlds; 20,949/2,028 tokenized examples after filtering.
  • Tokenizer fitted on the training partition only. Saved weights are FP32.

Loading

The saved config records Transformers 4.51.1. Use a Transformers version with Qwen3 support.

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "RLMM/zerohero-crafter-5m"
revision = "epoch-12"  # choose "epoch-02" or "epoch-12"
tokenizer = AutoTokenizer.from_pretrained(repo, revision=revision)
model = AutoModelForCausalLM.from_pretrained(repo, revision=revision)
model.eval()

Always load the tokenizer and model from the same revision. No custom remote code is needed. This is a task-specific completion model; use the repository's public_prompt serializer instead of applying a generic chat template. With generation, use at most 1,024 new tokens and keep the total prompt plus response within the 4,096-token context. The reasoning targets are learned text and should not be treated as verified explanations.

Code and PPO initialization

These two checkpoints were used as actor initializations for separate Crafter PPO runs. The full PPO experiment needs the environment/dependencies and input data files described in the code repository; this model repository contains inference/SFT checkpoint files, not the dataset, environment, PPO critic or optimizer recovery state.

Integrity

model.safetensors SHA-256: 7de94e4fc028ab5c4b755983f9943522f4c6d38bb5a34d0bec2a3c4abf42b515.

provenance.json records checksums for all seven original checkpoint files. checkpoint_complete.json retains the original SFT completion metadata.

Downloads last month
21
Safetensors
Model size
5.03M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support