Explore a small model learning to choose from game states
Jev-Omni decisions, SAM masks, and reviewable evidence