MolmoBot-DROID Directional Pick/Place SimpleDesk — Step 10,000

This repository contains the 10,000-step intermediate checkpoint from the lr_custom_16k MolmoBot finetuning run. It was finetuned from allenai/MolmoBot-DROID for Franka/FR3 tabletop pick-and-place commands that require placing one object to the left or right of another object.

The root config.yaml and model.pt form the unsharded MolmoBot checkpoint expected by the MolmoBot server. This is not the final 16,000-step checkpoint; the step is deliberately included in both the repository name and this model card.

Checkpoint identity

Field Value
Checkpoint step 10,000
Training run lr_custom_16k
Checkpoint format MolmoBot unsharded (config.yaml + model.pt)
Base model allenai/MolmoBot-DROID
Base model revision cbe6ec358958d07ddfb20d3aa54e560e9e1b18c9
Training W&B run ccwatson/molmobot_finetune/ouuzdjlg
Training Slurm job 7107997 (lr-custom-ie-fe-16k-b16-32)
Checkpoint created 2026-07-17 12:11 UTC
model.pt size 19,992,166,548 bytes
model.pt SHA-256 027daf40cbf87cfebc4dcee0aa16b0d6c0a87c2bfb30229ed77f9c1545d600da
config.yaml SHA-256 7eeb51134b93fc17955f3bd850240074d6b9faf3dcf527177861940c7c098400

The complete resolved model/trainer configuration is in config.yaml. The submission-time environment is preserved as metadata/training_launch.env, and machine-readable provenance is in metadata/provenance.json.

Training data

Training used ccwatson/directional_pick_place_simpledesk_diverse_targets_tested40_200demos_20260716_molmobot at dataset revision 5446191605f7eb3020983d83f707153d9719ecfb.

The dataset contains 200 successful scripted trajectories derived from 40 valid benchmark episodes, with five attempts per episode:

Split Source episodes Trajectories HDF5 files Videos
Train 36 180 180 360
Validation 4 20 20 40

The two cameras are droid_shoulder_light_randomization and wrist_camera_zed_mini. Prompts were not randomized during training. Normalization was computed from the train split using action quantiles and state min/max statistics for the eight-dimensional Franka action and state vectors. The exact statistics are included as normalization_stats.yaml (SHA-256 972a88e10f02976ce5a4d1c90a2070a62e4daa08a5fe1cecb4ca00e8216eb15d).

Main training settings

Setting Value
Intended run length 16,000 steps
This checkpoint 10,000 steps
Device / global batch size 16 / 32
Sequence length 928
Observation frames 2
Observation step delta 8
Maximum images 4
Maximum frames 1
Single-frame preprocessing false
Action preset franka_joint
Camera preset franka_droid
Action horizon in data 16 × 8
Weighted sampling enabled
Prompt randomization disabled
Image augmentation enabled
LLM finetuning enabled
Embedding finetuning mode ae
Flow timestamps 8
Connector learning rate 5e-6
Vision encoder learning rate 5e-6
LLM learning rate 1e-5
Action expert learning rate 1e-4
Validation interval / max examples 200 / 200
Save and external-eval interval 2,000 steps
Trainer and optimizer state reset from base model
Hardware 1× NVIDIA B200, 28 CPUs, 224 GiB host memory

Closed-loop evaluation

This checkpoint was evaluated on all 40 episodes of FrankaPickAndPlaceRelativeNextToSimpleDeskDiverseTargetsTested40Bench (20 left and 20 right tasks) through the VLA evaluation harness and MolmoSpaces. The evaluation used natural-language prompts, Filament rendering, 10 parallel shards, a 500-step task horizon, and both DROID cameras.

Server action settings were:

  • execution horizon: 8
  • state mode: cross_attn
  • maximum joint delta: 0.2
Metric Result
Strict success 25 / 40 (62.5%)
Lax success 26 / 40 (65.0%)
Strict oracle-done rate 25 / 40 (62.5%)
Lax oracle-done rate 27 / 40 (67.5%)
Average episode length 274.675 steps
Average successful episode length 140.08 steps
Min / max episode length 32 / 499 steps
Failed shards 0
Missing / duplicate / unexpected episodes 0 / 0 / 0

The eval completed successfully (eval_returncode=0) in 2,016.92 seconds. See the evaluation W&B run, metadata/external_eval_request.json, and metadata/external_eval_result.json for the complete recorded configuration and result.

Repository files

  • model.pt: unsharded MolmoBot weights.
  • config.yaml: complete resolved model and trainer configuration.
  • normalization_stats.yaml: exact state/action normalization statistics.
  • metadata/training_launch.env: exact experiment launch configuration.
  • metadata/external_eval_request.json: checkpoint identity and requested eval settings.
  • metadata/external_eval_result.json: complete eval result and coverage accounting.
  • metadata/provenance.json: source revisions, hashes, run identity, and artifact inventory.

Intended use and limitations

This checkpoint is intended for research evaluation of the matching Franka/FR3 DROID-camera setup and directional SimpleDesk tasks. It should be served with the MolmoBot inference stack and the included normalization stats. It has only been measured on the 40-episode simulated benchmark described above. The 62.5% result is a single benchmark evaluation, not a claim of broad robotic capability or real-world safety. Validate action scaling, cameras, workspace limits, collision handling, and emergency-stop behavior before any physical-robot use.

License

This derivative checkpoint follows the Apache-2.0 license declared by the base allenai/MolmoBot-DROID model. Users remain responsible for the terms of the base model, training dataset, simulator assets, and other dependencies.

Downloads last month
14
Video Preview
loading

Model tree for ccwatson/MolmoBot-DROID-Directional-PickPlace-SimpleDesk-Step-10000

Finetuned
(4)
this model

Dataset used to train ccwatson/MolmoBot-DROID-Directional-PickPlace-SimpleDesk-Step-10000

Evaluation results

  • Strict success rate on MolmoSpaces SimpleDesk Diverse Targets Tested40
    self-reported
    0.625
  • Lax success rate on MolmoSpaces SimpleDesk Diverse Targets Tested40
    self-reported
    0.650