VTAM β€” tactile world-model ablation ("late fuse") stage-B3 action experts

Stage-B3 (train_mode: action_full) action-expert checkpoints for the VTAM tactile world-model ablation, on three contact-rich tasks: chip, cucumber_peel, hard_wipe.

These checkpoints have not been evaluated. No success rates, no open-loop metrics, no comparison numbers exist for them yet. Nothing in this card should be read as a performance claim.

What the ablation is

The question is whether predictive tactile world modeling is worth anything, as distinct from merely giving the policy access to touch. Two arms, with the policy's access to tactile held constant:

arm where tactile enters world model sees touch?
full/ (control) GelSight is a DiT view: it passes through all 28 blocks and mixes with vision in cross-view attention, so it is part of the modeled world state the action expert reads. yes
latefuse/ (ablation) GelSight is VAE-encoded and handed straight to the action expert's cross-attention KV, skipping every DiT block. The world model is vision-only and never sees a tactile token. no

full βˆ’ latefuse therefore isolates the value of predictive tactile world modeling.

The ablation arm's only extra weight is a zero-initialized tactile marker embedding (1 Γ— 2048 = 2048 params), so the two arms stay parameter-comparable. This is deliberately not a dual-cross-attention late fuse (which would add a second ~73.5M-param cross-attention tower and confound "where touch enters" with "how much capacity the arm has"). The 4200-byte difference in safetensors size between the two arms is that marker embedding plus header.

Tactile covers the memory frames only in the ablation arm. Under noisy_video the full model's future positions are pure noise, so feeding clean future tactile latents to the ablation would hand it an observation the control never gets β€” and one that does not exist at deploy time.

For both arms valid_cam is identical (realsense_0, realsense_1, gelsight); the trainer peels the trailing tactile view off the DiT view stack in the ablation arm so the DiT runs V=2 while tactile latents travel separately.

Step counts differ on purpose

Each late-fuse arm's train_steps was set to its own task's FULL baseline last saved checkpoint, so each pair is compared at equal optimizer steps. The step counts are therefore not uniform across tasks, and that is intentional, not an inconsistency:

task latefuse step FULL baseline step matched?
chip 30000 30000 yes
hard_wipe 30000 30000 yes
cucumber_peel 50000 50000 (May lineage) yes

World-model warm-start time is also held fixed: every arm warm-starts from a stage-B2 world model at step_50000, the ablation from the vision-only B2 and the control from the GelSight B2 of the same lineage.

All three control arms are published

full/ contains chip, cucumber_peel and hard_wipe, so each ablation arm can be read against its own matched control.

One note on cucumber_peel, because two different FULL runs exist and only one is the right control:

  • The control published here is the May lineage run task_cucumber_peel_force_in_action_action_full_gelsight/2026_05_20_12_18_18 at step_50000, trained on the same 120-episode dataset the ablation arm uses and warm-started from the May GelSight world model (task_cucumber_peel_force_in_action_video_gelsight/2026_05_15_04_14_06, step_50000). This is the run the ablation config names as its baseline.
  • A March run (2026_03_03_06_46_25, to step_30000) also exists and is byte-intact, but it was trained on the older 61-episode copy of the dataset and stops at a different step. It is not a valid control for this ablation arm and is deliberately not uploaded, so that nobody picks it up by mistake.

Architecture

LTX-Video multi-view DiT (LTXVideoTransformer3DModel, VTAM models/ltx_models/transformer_ltx_multiview.py) with an action expert head:

  • 28 layers, 32 attention heads, head dim 64, cross_attention_dim 2048
  • in_channels / out_channels 128, qk_norm: rms_norm_across_heads
  • action expert: action_in_channels / action_out_channels 26, 16 heads, head dim 32
  • action(10) = xyz(3) + rpy(3) + gripper(1) + force(3); state(16); concatenated on the channel axis β†’ 26 in / 26 out
  • ablation arm only: tactile_late_fuse: true, tactile_late_fuse_max_view: 1

Training: bf16, DeepSpeed ZeRO-2, AdamW lr 5e-5, constant_with_warmup (1000 warmup), batch size 16, gradient_accumulation_steps 1, noisy_video: true, pixel_wise_timestep: true, seed 42.

Checkpoint provenance

Exact source run directory and step for every file published here. All were verified byte-complete before upload (safetensors header's declared end equals actual file size).

latefuse/ β€” ablation arm

task source run directory step bytes
chip /projects/behf/haorany7/VTAM/outputs/task_chip_force_in_action_action_full_latefuse/2026_09_17_08_52_31 30000 4224918756
cucumber_peel /projects/behf/haorany7/VTAM/outputs/task_cucumber_peel_force_in_action_action_full_latefuse/2026_09_18_17_39_08 50000 4224918756
hard_wipe /projects/behf/haorany7/VTAM/outputs/task_hard_wipe_force_in_action_action_full_latefuse/2026_09_17_08_55_50 30000 4224918756

full/ β€” control arm

task source run directory step bytes
chip /work/hdd/bekg/vtam/outputs/task_chip_force_in_action_action_full_gelsight/2026_02_28_03_54_57 30000 4224914556
cucumber_peel /work/hdd/behf/haorany7/VTAM/outputs/task_cucumber_peel_force_in_action_action_full_gelsight/2026_05_20_12_18_18 50000 4224914556
hard_wipe /work/hdd/behe/WORLD-MODEL-TOUCH/outputs/task_hard_wipe_force_in_action_action_full_gelsight/task_hard_wipe_force_in_action_action_full_gelsight/2026_03_03_06_48_59 30000 4224914556

Stage-B2 warm starts

arm / task B2 world model step
latefuse chip chip_task_video/2026_02_19_05_27_36 50000
latefuse cucumber_peel task_cucumber_peel_force_in_action_video/2026_05_15_16_21_43 50000
latefuse hard_wipe task_hard_wipe_force_in_action_video/2026_03_02_06_38_22 50000
full chip chip_task_video_gelsight/2026_02_19_03_40_15 50000
full cucumber_peel task_cucumber_peel_force_in_action_video_gelsight/2026_05_15_04_14_06 50000
full hard_wipe task_hard_wipe_force_in_action_video_gelsight/2026_03_01_11_45_56 50000

Layout

latefuse/<task>/step_<N>/diffusion_pytorch_model.safetensors
latefuse/<task>/step_<N>/config.json      # needed to load the shard
latefuse/<task>/run_config.json           # resolved config snapshot from the run
latefuse/<task>/train_config.yaml         # launch config from the VTAM repo
full/<task>/...                           # same layout

train_config.yaml is the repo's launch config; run_config.json is the resolved snapshot the run itself archived. Both are included so the weights are reproducible from what is published.

Note: full/chip/train_config.yaml and full/hard_wipe/train_config.yaml are the repo's *_action_full_gelsight.yaml files as they stand today. Their model_path and some data_roots point at /workspace or other paths that are dead now; the run's own run_config.json records what was actually used.

full/cucumber_peel/ deliberately has no train_config.yaml. The only checked-in candidate, action_model_task_cucumber_peel_force_in_action_action_full_gelsight.yaml, describes the March 61-episode lineage (it warm-starts from /workspace/.../2026_03_01_11_40_38 and reads the older data root), so it does not describe the May run published here and shipping it under that name would be misleading. Use full/cucumber_peel/run_config.json, which is the resolved config the run itself archived and is authoritative for these weights.

Downloads last month
-
Video Preview
loading