Instructions to use helen9975/molmoact2_telearms_v4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use helen9975/molmoact2_telearms_v4 with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForImageTextToText model = AutoModelForImageTextToText.from_pretrained("helen9975/molmoact2_telearms_v4", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
molmoact2_telearms_v4 (final checkpoint, step 100,000)
MolmoAct2 full fine-tune of allenai/MolmoAct2-BimanualYAM on mliu48/telearms-yam-cube-v4
(53,832 simulated bimanual YAM demonstrations, 48 tasks, 30 Hz, 3 × 224×224 cameras).
- Trainer:
allenai/molmoact2experiments/launch_scripts/train_lerobot.py(upstream6070080), converted witholmo.hf_model.convert_molmoact2_to_hf. - Full fine-tune (LLM, ViT, connector, LM head, action expert), unpacked + dynamic sequence length.
- 16 × H100 (2 nodes × 8), batch 16/GPU (global 256), 100k-step schedule; LRs: LLM 2e-5, ViT 1e-5, connector 1e-5, action expert 1e-4 (upstream full-FT recipe × sqrt(256/64)).
- Mixture tag
telearms_yam_sim: absolute joint-pose control, 14-D state/action (per arm 6 joints + gripper), cameras[camera_front, camera_left, camera_right](same order as BimanualYAM's[top, left, right]), 30-step action chunks. norm_stats.jsonholds this dataset's absolute action/state statistics (tagtelearms_yam_sim).- Token embedding / LM head have 155,648 rows (base checkpoint: 154,624): the trainer adds special tokens at fine-tuning start; the bundled tokenizer matches.
- Not yet evaluated in simulation.
- Downloads last month
- 3
Model tree for helen9975/molmoact2_telearms_v4
Base model
allenai/MolmoAct2-BimanualYAM