3D-DLP-VC trained on GenericShapes-RGB β€” reproduction checkpoint

A 3D-DLP-VC (RGB-voxel Deep Latent Particles) model trained from scratch for an independent reproduction of ICML 2026 paper #10351, 3D-DLP: Self-supervised 3D Object-centric Scene Representation Learning (OpenReview vIotI25gJz).

Trained with the authors' code completely unmodified β€” Eubooks3003/3d-dlp at commit c1e90ab β€” via python train_dlp_voxel.py -d repro_gs.

The authors release no checkpoints, so this is, as far as we can tell, the first publicly available trained 3D-DLP model.

What it is

64Β³ RGB voxel grids β†’ 24 foreground latent particles + 1 background particle. Each foreground particle carries a 3D keypoint position, a per-axis scale, a transparency, and appearance features. Training is pure self-supervised reconstruction β€” no instance labels or masks.

Files

File Purpose
best.pt best-validation checkpoint ({"model": state_dict, "epoch", "best_metric", …})
hparams.json the exact config the run used β€” pass as --dlp-cfg
train_gs.log full training log (per-epoch loss / KL / PSNR)
repro_gs.json the training config, for train_dlp_voxel.py -d

Training setup, and how it differs from the paper

This checkpoint Paper
Data GenericShapes-RGB, 40,000 scenes, re-created GenericShapes, ~40,000 scenes, unreleased
Particles 24 + background, 64 K-means proposals, 5 iterations same
Optimiser Adam, lr 8 Γ— 10⁻⁡, Ξ»_chroma = 500 same
Batch size 16 32
Hardware 1 Γ— RTX PRO 5000 Blackwell (48 GB) 1 Γ— GH200
Schedule far shorter than the paper's ~48 h ~48 h

Batch 32 at 64Β³ with 24 particles exceeds 48 GB, so batch 16 was used with the learning rate unchanged. Both deviations make this checkpoint a lower bound on the paper's quality β€” treat the metrics below as a floor, not as a replication of the paper's Table 1 (which is on a different, unreleased corpus anyway).

Measured behaviour (held-out test split)

Ground-truth per-voxel instance labels exist for this synthetic corpus and were never used in training β€” only for evaluation.

Metric Value Control
Foreground ARI (object separation) 0.485 Β± 0.120 model's own untrained K-means prior 0.362; random 0.000
FG/BG balanced accuracy, unsupervised 0.799 Β± 0.107 chance 0.500
FG/BG IoU 0.627 β€”
mBO (per-object best overlap) 0.151 β€”
Masked PSNR (repo's eval/eval_vox.py) 14.80 dB β€”
Occupancy IoU (same) 0.305 β€”
Active particles vs. true objects 15.6 vs 4.5 β€”
Position-edit slope (ideal 1.0) 1.0024 (RΒ² 0.9959, n=2972) randomly-init model 0.003
Scale-edit log–log slope (ideal 1.0) 1.0040 (RΒ² 0.973) β€”
Drift of other particles under an edit 0.000 β€”

Metrics are over 200 (Claim 1) / 50 (Claim 2) held-out test scenes.

Important caveat. On this corpus a plain K-means on colour alone scores FG-ARI 0.932 β€” better than the model β€” because the generator gives each object a distinct hue. The checkpoint beats geometry-only clustering (0.402), joint colour+geometry clustering (0.482) and its own untrained prior, but this dataset cannot demonstrate that the segmentation is good in absolute terms.

The live numbers, controls, adversarial checks and figures are in the reproduction logbook: https://huggingface.co/spaces/rvt832/repro-3d-dlp-self-supervised-3d-object-centric-scene-representation-learning

Usage

git clone https://github.com/Eubooks3003/3d-dlp.git && cd 3d-dlp
git checkout c1e90ab41ba9fb4bd674eaad0b93c09d1fa5e0cd
hf download rvt832/3d-dlp-repro-genericshapes-vc --local-dir ckpt
import torch, json
from voxel_models import DLP                     # from the cloned repo

cfg = json.load(open("ckpt/hparams.json"))
model = DLP(cdim=3, image_size=64, n_kp_enc=cfg["n_kp_enc"], ...)   # see repo for the full arg list
model.load_state_dict(torch.load("ckpt/best.pt", map_location="cpu")["model"], strict=False)
model.eval()

out = model(voxels, deterministic=True)          # voxels: [B, 3, 64, 64, 64] in [0,1]
out["alpha_masks"]   # [B, 24, 1, 64, 64, 64] β€” per-particle segmentation, obj_on-gated
out["bg_mask"]       # [B, 1, 64, 64, 64]     β€” background mask
out["z"], out["z_scale"], out["obj_on"]         # keypoints, scales, transparencies
out["rec"]           # reconstruction

repro/eval_claim1.py and repro/eval_claim2.py in the reproduction workspace show the exact model-construction call and the decomposition read-out.

Limitations

  • Trained only on synthetic single-table scenes with 3–6 convex primitives and saturated per-object colours. It has not been evaluated on MimicGen, RLBench, ShapeNet scenes or real RGB-D.
  • Over-segments: roughly 15 active particles for ~4.6 objects, so per-object mean-best-overlap is low even where the ARI is good.
  • The voxel grid used here is anisotropic (3.8 mm vertically vs 13 mm laterally), which makes learned particles full-height columns; vertical position/scale edits saturate against the grid.

Citation

Cite the original paper:

@inproceedings{zhang2026threeddlp,
  title     = {3D-DLP: Self-supervised 3D Object-centric Scene Representation Learning},
  author    = {Zhang, Ellina and Iyengar, Madhavan and Zadeh, Amir and Li, Chuan
               and Held, David and Pathak, Deepak and Daniel, Tal},
  booktitle = {International Conference on Machine Learning (ICML)},
  year      = {2026},
  url       = {https://arxiv.org/abs/2606.19451}
}

This checkpoint is an independent reproduction artifact and is not affiliated with, or endorsed by, the paper's authors.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train rvt832/3d-dlp-repro-genericshapes-vc

Paper for rvt832/3d-dlp-repro-genericshapes-vc