Instructions to use JensenYuan/VTAM_tactile_latefuse_ablation with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use JensenYuan/VTAM_tactile_latefuse_ablation with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("JensenYuan/VTAM_tactile_latefuse_ablation", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
VTAM β tactile world-model ablation ("late fuse") stage-B3 action experts
Stage-B3 (train_mode: action_full) action-expert checkpoints for the VTAM
tactile world-model ablation, on three contact-rich tasks: chip,
cucumber_peel, hard_wipe.
These checkpoints have not been evaluated. No success rates, no open-loop metrics, no comparison numbers exist for them yet. Nothing in this card should be read as a performance claim.
What the ablation is
The question is whether predictive tactile world modeling is worth anything, as distinct from merely giving the policy access to touch. Two arms, with the policy's access to tactile held constant:
| arm | where tactile enters | world model sees touch? |
|---|---|---|
full/ (control) |
GelSight is a DiT view: it passes through all 28 blocks and mixes with vision in cross-view attention, so it is part of the modeled world state the action expert reads. | yes |
latefuse/ (ablation) |
GelSight is VAE-encoded and handed straight to the action expert's cross-attention KV, skipping every DiT block. The world model is vision-only and never sees a tactile token. | no |
full β latefuse therefore isolates the value of predictive tactile world
modeling.
The ablation arm's only extra weight is a zero-initialized tactile marker
embedding (1 Γ 2048 = 2048 params), so the two arms stay parameter-comparable.
This is deliberately not a dual-cross-attention late fuse (which would add a
second ~73.5M-param cross-attention tower and confound "where touch enters" with
"how much capacity the arm has"). The 4200-byte difference in safetensors size
between the two arms is that marker embedding plus header.
Tactile covers the memory frames only in the ablation arm. Under
noisy_video the full model's future positions are pure noise, so feeding clean
future tactile latents to the ablation would hand it an observation the control
never gets β and one that does not exist at deploy time.
For both arms valid_cam is identical (realsense_0, realsense_1,
gelsight); the trainer peels the trailing tactile view off the DiT view stack
in the ablation arm so the DiT runs V=2 while tactile latents travel separately.
Step counts differ on purpose
Each late-fuse arm's train_steps was set to its own task's FULL baseline
last saved checkpoint, so each pair is compared at equal optimizer steps. The
step counts are therefore not uniform across tasks, and that is intentional, not
an inconsistency:
| task | latefuse step | FULL baseline step | matched? |
|---|---|---|---|
chip |
30000 | 30000 | yes |
hard_wipe |
30000 | 30000 | yes |
cucumber_peel |
50000 | 50000 (May lineage) | yes |
World-model warm-start time is also held fixed: every arm warm-starts from a
stage-B2 world model at step_50000, the ablation from the vision-only B2 and
the control from the GelSight B2 of the same lineage.
All three control arms are published
full/ contains chip, cucumber_peel and hard_wipe, so each ablation arm
can be read against its own matched control.
One note on cucumber_peel, because two different FULL runs exist and only one
is the right control:
- The control published here is the May lineage run
task_cucumber_peel_force_in_action_action_full_gelsight/2026_05_20_12_18_18atstep_50000, trained on the same 120-episode dataset the ablation arm uses and warm-started from the May GelSight world model (task_cucumber_peel_force_in_action_video_gelsight/2026_05_15_04_14_06,step_50000). This is the run the ablation config names as its baseline. - A March run (
2026_03_03_06_46_25, tostep_30000) also exists and is byte-intact, but it was trained on the older 61-episode copy of the dataset and stops at a different step. It is not a valid control for this ablation arm and is deliberately not uploaded, so that nobody picks it up by mistake.
Architecture
LTX-Video multi-view DiT (LTXVideoTransformer3DModel, VTAM
models/ltx_models/transformer_ltx_multiview.py) with an action expert head:
- 28 layers, 32 attention heads, head dim 64,
cross_attention_dim2048 in_channels/out_channels128,qk_norm: rms_norm_across_heads- action expert:
action_in_channels/action_out_channels26, 16 heads, head dim 32 - action(10) = xyz(3) + rpy(3) + gripper(1) + force(3); state(16); concatenated on the channel axis β 26 in / 26 out
- ablation arm only:
tactile_late_fuse: true,tactile_late_fuse_max_view: 1
Training: bf16, DeepSpeed ZeRO-2, AdamW lr 5e-5, constant_with_warmup (1000
warmup), batch size 16, gradient_accumulation_steps 1, noisy_video: true,
pixel_wise_timestep: true, seed 42.
Checkpoint provenance
Exact source run directory and step for every file published here. All were verified byte-complete before upload (safetensors header's declared end equals actual file size).
latefuse/ β ablation arm
| task | source run directory | step | bytes |
|---|---|---|---|
chip |
/projects/behf/haorany7/VTAM/outputs/task_chip_force_in_action_action_full_latefuse/2026_09_17_08_52_31 |
30000 | 4224918756 |
cucumber_peel |
/projects/behf/haorany7/VTAM/outputs/task_cucumber_peel_force_in_action_action_full_latefuse/2026_09_18_17_39_08 |
50000 | 4224918756 |
hard_wipe |
/projects/behf/haorany7/VTAM/outputs/task_hard_wipe_force_in_action_action_full_latefuse/2026_09_17_08_55_50 |
30000 | 4224918756 |
full/ β control arm
| task | source run directory | step | bytes |
|---|---|---|---|
chip |
/work/hdd/bekg/vtam/outputs/task_chip_force_in_action_action_full_gelsight/2026_02_28_03_54_57 |
30000 | 4224914556 |
cucumber_peel |
/work/hdd/behf/haorany7/VTAM/outputs/task_cucumber_peel_force_in_action_action_full_gelsight/2026_05_20_12_18_18 |
50000 | 4224914556 |
hard_wipe |
/work/hdd/behe/WORLD-MODEL-TOUCH/outputs/task_hard_wipe_force_in_action_action_full_gelsight/task_hard_wipe_force_in_action_action_full_gelsight/2026_03_03_06_48_59 |
30000 | 4224914556 |
Stage-B2 warm starts
| arm / task | B2 world model | step |
|---|---|---|
latefuse chip |
chip_task_video/2026_02_19_05_27_36 |
50000 |
latefuse cucumber_peel |
task_cucumber_peel_force_in_action_video/2026_05_15_16_21_43 |
50000 |
latefuse hard_wipe |
task_hard_wipe_force_in_action_video/2026_03_02_06_38_22 |
50000 |
full chip |
chip_task_video_gelsight/2026_02_19_03_40_15 |
50000 |
full cucumber_peel |
task_cucumber_peel_force_in_action_video_gelsight/2026_05_15_04_14_06 |
50000 |
full hard_wipe |
task_hard_wipe_force_in_action_video_gelsight/2026_03_01_11_45_56 |
50000 |
Layout
latefuse/<task>/step_<N>/diffusion_pytorch_model.safetensors
latefuse/<task>/step_<N>/config.json # needed to load the shard
latefuse/<task>/run_config.json # resolved config snapshot from the run
latefuse/<task>/train_config.yaml # launch config from the VTAM repo
full/<task>/... # same layout
train_config.yaml is the repo's launch config; run_config.json is the
resolved snapshot the run itself archived. Both are included so the weights are
reproducible from what is published.
Note: full/chip/train_config.yaml and full/hard_wipe/train_config.yaml are
the repo's *_action_full_gelsight.yaml files as they stand today. Their
model_path and some data_roots point at /workspace or other paths that are
dead now; the run's own run_config.json records what was actually used.
full/cucumber_peel/ deliberately has no train_config.yaml. The only
checked-in candidate,
action_model_task_cucumber_peel_force_in_action_action_full_gelsight.yaml,
describes the March 61-episode lineage (it warm-starts from
/workspace/.../2026_03_01_11_40_38 and reads the older data root), so it does
not describe the May run published here and shipping it under that name would be
misleading. Use full/cucumber_peel/run_config.json, which is the resolved
config the run itself archived and is authoritative for these weights.
- Downloads last month
- -