chess_phase_smolvla_baseline

SmolVLA trained to move a chess piece with an SO-101 arm in one fixed simulated scene (MuJoCo). It is the simple baseline of the chess robot's "hand": first get reliable closed-loop success in the easiest setting, then add variety back one change at a time. Trained fresh from lerobot/smolvla_base, not from the earlier model trained on the varied data.

Status: not yet reliable. In a closed-loop test it completed 36.7% of 60 moves (details below). It has not been run on a real arm.

What it does

  • Input: two 640x480 camera images and the arm's joint positions, at 30 Hz.
    • observation.images.overhead: a webcam 65 cm straight above the board, 44 degrees top to bottom, the arm at the top edge of the image. Fed to the model as camera1.
    • observation.images.wrist: the official SO-101 wrist camera. Fed to the model as camera2.
    • The piece to move has a red square with an opaque outline drawn over its square; the target square has a blue one. Both are drawn on the images before they reach the policy.
    • observation.state: 5 arm joints in degrees and the gripper 0-100 (LeRobot's SO-101 follower units with use_degrees=True).
    • The instruction is always "move the piece on the red square to the blue square".
  • Output: the next joint targets in the same units, in chunks of 50 steps.
  • The training video was stored at high quality (H.264, CRF 18, full 4:4:4 colour), so the markers keep their colour. For the closest match, pass live frames through the same encode and decode (training_look(image, 18, "yuv444p") in the project's sim/camera_effects.py).

The scene (all fixed)

  • Board with 25 mm squares, square to the robot and flush against its 6 cm deck; the robot plays black. One Staunton set, black and white pieces, centred in their squares. White arm.
  • Tray on the right, one table, the same two lamps, no clutter, a clean camera image.
  • Ten moves from the start position: e7e5, d7d5, g8f6, b8c6, c7c5 (the robot's side) and e2e4, d2d4, g1f3, b1c3, c2c4 (the far side). The start is the same every time and only the red and blue squares change, so doing different moves correctly shows it reads the markers.

Test in simulation

60 closed-loop episodes, each a random one of the ten moves from seeds not used in training. An episode counts as a success if the piece ends within 6 mm of the target square, upright, and no other piece moved more than 2 mm. 20 s limit.

Success: 36.7%

  • Gripper within half a square of the marked piece: 95.0% of episodes (median closest approach 3.3 mm).
  • Marked piece lifted: 68.3%.
move success episodes
b1-c3 16.7% 6
b8-c6 14.3% 7
c2-c4 0.0% 5
c7-c5 66.7% 6
d2-d4 33.3% 6
d7-d5 60.0% 5
e2-e4 60.0% 10
e7-e5 33.3% 6
g1-f3 33.3% 3
g8-f6 33.3% 6

More tests

Two further 60-episode tests on held-out seeds gave 51.7% and 41.7%, so about 43% over 180 episodes (a set of 60 swings by about 12 points by chance). On the final set the median sideways error when the jaws start closing is 4.5 mm (2.5 in successes, 6.5 in failures); the main failure is knocking the piece over while closing.

What goes wrong

  • It reads the markers. The gripper reaches within half a square of the marked piece in 95% of episodes, for every one of the ten moves.
  • It fails at the grasp. A grasp inspection over 20 episodes compared the pinch point, at the moment the policy starts closing the jaws, with the grasp the scripted expert would make on the same piece. Grasps that held were 0.8-0.9 mm off; failures were 4-6 mm off. With the jaws off-centre it pushes the piece over (7 of 20) or closes beside it (6 of 20).
  • Executing 10 steps of each 50-step chunk before predicting again, instead of all 50, did not help (4 of 15).
  • ACT trained on the same data reached 14%.
  • A first fine-tune on correction demonstrations of its own mistakes reached 20%. In those corrections the teacher rose above the piece and paused before grasping again, and the model learned to hover over pawns. That model was not kept.
  • A second attempt, where the teacher lines up from where the arm is and goes straight down, reached 30% against this model's 51.7% on the same episodes: the model copied the low sideways alignment imprecisely and knocked more pieces over. It was not kept either.

Training data

Machanize/playful, folder chess-sim/datasets/baseline_v1 (private): 300 successful episodes (89008 frames) of a scripted inverse-kinematics expert at 30 fps, about 30 per move. The expert's start pose and speed vary slightly between episodes.

Training

From lerobot/smolvla_base, 10000 steps at batch 64 on one RTX 4090 in 1.9 h, LeRobot 0.4.4, default SmolVLA settings (vision encoder frozen, action expert trained). Overhead and wrist cameras are renamed to camera1 and camera2.

Use

from lerobot.policies.smolvla.modeling_smolvla import SmolVLAPolicy
policy = SmolVLAPolicy.from_pretrained("Machanize/chess_phase_smolvla_baseline")

Load the pre- and post-processors saved with the model using make_pre_post_processors. They apply the training normalisation and rename the camera keys.

Limitations

  • Simulation only, one scene, ten moves. Other moves, board positions, camera views, lighting or piece sets are untested and expected to fail until the variety is added back.
  • The overlay is required: without the red and blue squares it does not know what to move.
Downloads last month
35
Safetensors
Model size
0.5B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Model tree for Machanize/chess_phase_smolvla_baseline

Finetuned
(8015)
this model