VTAM Cucumber Peel: no-force-output ablation (action expert, step 30000)

Stage-II action expert of VTAM (Video-Tactile-Action Model) for the real-robot Cucumber Peel task, trained without force in the action output. It is the force-regression ablation of the full VTAM model for this task.

What differs from the full model

  • The action target is sliced to the first 7 dimensions (valid_act_dim: 7: xyz, rpy, gripper); the force dimensions of the 10-D absolute action are dropped. The expert's input/output width is 23 (7 action + 16 state).
  • There is no force loss. A test suite (tests/test_force_ablation.py, 30 tests) checked that only 7 action dimensions are predicted and supervised before launch.
  • Same dataset, normalization statistics and frozen Stage-I video/world model as the full-force model. Training was capped at 50k steps and stopped at 30k to match the other ablation.

Training

30,000 steps, 4x A100, global batch 64, learning rate 5e-5, seed 42.

Files

  • diffusion_pytorch_model.safetensors, config.json: action-expert weights (step 30000).
  • action_model_task_cucumber_peel_force_in_action_action_full_gelsight_noforce.yaml: training config.
  • task_cucumber_peel_force_in_action_stats.json: action/state normalization statistics.

Usage

Load with the VTAM codebase and set action_dim=7 in MVActor; the policy outputs 7-D actions (no force).

Downloads last month
5
Safetensors
Model size
2B params
Tensor type
BF16
·
Video Preview
loading