3D-DLP-VC trained on GenericShapes-RGB β reproduction checkpoint
A 3D-DLP-VC (RGB-voxel Deep Latent Particles) model trained from scratch for an independent
reproduction of ICML 2026 paper #10351,
3D-DLP: Self-supervised 3D Object-centric Scene Representation Learning
(OpenReview vIotI25gJz).
Trained with the authors' code completely unmodified β
Eubooks3003/3d-dlp at commit
c1e90ab β
via python train_dlp_voxel.py -d repro_gs.
The authors release no checkpoints, so this is, as far as we can tell, the first publicly available trained 3D-DLP model.
What it is
64Β³ RGB voxel grids β 24 foreground latent particles + 1 background particle. Each foreground particle carries a 3D keypoint position, a per-axis scale, a transparency, and appearance features. Training is pure self-supervised reconstruction β no instance labels or masks.
Files
| File | Purpose |
|---|---|
best.pt |
best-validation checkpoint ({"model": state_dict, "epoch", "best_metric", β¦}) |
hparams.json |
the exact config the run used β pass as --dlp-cfg |
train_gs.log |
full training log (per-epoch loss / KL / PSNR) |
repro_gs.json |
the training config, for train_dlp_voxel.py -d |
Training setup, and how it differs from the paper
| This checkpoint | Paper | |
|---|---|---|
| Data | GenericShapes-RGB, 40,000 scenes, re-created | GenericShapes, ~40,000 scenes, unreleased |
| Particles | 24 + background, 64 K-means proposals, 5 iterations | same |
| Optimiser | Adam, lr 8 Γ 10β»β΅, Ξ»_chroma = 500 | same |
| Batch size | 16 | 32 |
| Hardware | 1 Γ RTX PRO 5000 Blackwell (48 GB) | 1 Γ GH200 |
| Schedule | far shorter than the paper's ~48 h | ~48 h |
Batch 32 at 64Β³ with 24 particles exceeds 48 GB, so batch 16 was used with the learning rate unchanged. Both deviations make this checkpoint a lower bound on the paper's quality β treat the metrics below as a floor, not as a replication of the paper's Table 1 (which is on a different, unreleased corpus anyway).
Measured behaviour (held-out test split)
Ground-truth per-voxel instance labels exist for this synthetic corpus and were never used in training β only for evaluation.
| Metric | Value | Control |
|---|---|---|
| Foreground ARI (object separation) | 0.485 Β± 0.120 | model's own untrained K-means prior 0.362; random 0.000 |
| FG/BG balanced accuracy, unsupervised | 0.799 Β± 0.107 | chance 0.500 |
| FG/BG IoU | 0.627 | β |
| mBO (per-object best overlap) | 0.151 | β |
Masked PSNR (repo's eval/eval_vox.py) |
14.80 dB | β |
| Occupancy IoU (same) | 0.305 | β |
| Active particles vs. true objects | 15.6 vs 4.5 | β |
| Position-edit slope (ideal 1.0) | 1.0024 (RΒ² 0.9959, n=2972) | randomly-init model 0.003 |
| Scale-edit logβlog slope (ideal 1.0) | 1.0040 (RΒ² 0.973) | β |
| Drift of other particles under an edit | 0.000 | β |
Metrics are over 200 (Claim 1) / 50 (Claim 2) held-out test scenes.
Important caveat. On this corpus a plain K-means on colour alone scores FG-ARI 0.932 β better than the model β because the generator gives each object a distinct hue. The checkpoint beats geometry-only clustering (0.402), joint colour+geometry clustering (0.482) and its own untrained prior, but this dataset cannot demonstrate that the segmentation is good in absolute terms.
The live numbers, controls, adversarial checks and figures are in the reproduction logbook: https://huggingface.co/spaces/rvt832/repro-3d-dlp-self-supervised-3d-object-centric-scene-representation-learning
Usage
git clone https://github.com/Eubooks3003/3d-dlp.git && cd 3d-dlp
git checkout c1e90ab41ba9fb4bd674eaad0b93c09d1fa5e0cd
hf download rvt832/3d-dlp-repro-genericshapes-vc --local-dir ckpt
import torch, json
from voxel_models import DLP # from the cloned repo
cfg = json.load(open("ckpt/hparams.json"))
model = DLP(cdim=3, image_size=64, n_kp_enc=cfg["n_kp_enc"], ...) # see repo for the full arg list
model.load_state_dict(torch.load("ckpt/best.pt", map_location="cpu")["model"], strict=False)
model.eval()
out = model(voxels, deterministic=True) # voxels: [B, 3, 64, 64, 64] in [0,1]
out["alpha_masks"] # [B, 24, 1, 64, 64, 64] β per-particle segmentation, obj_on-gated
out["bg_mask"] # [B, 1, 64, 64, 64] β background mask
out["z"], out["z_scale"], out["obj_on"] # keypoints, scales, transparencies
out["rec"] # reconstruction
repro/eval_claim1.py and repro/eval_claim2.py in the reproduction workspace show the exact
model-construction call and the decomposition read-out.
Limitations
- Trained only on synthetic single-table scenes with 3β6 convex primitives and saturated per-object colours. It has not been evaluated on MimicGen, RLBench, ShapeNet scenes or real RGB-D.
- Over-segments: roughly 15 active particles for ~4.6 objects, so per-object mean-best-overlap is low even where the ARI is good.
- The voxel grid used here is anisotropic (3.8 mm vertically vs 13 mm laterally), which makes learned particles full-height columns; vertical position/scale edits saturate against the grid.
Citation
Cite the original paper:
@inproceedings{zhang2026threeddlp,
title = {3D-DLP: Self-supervised 3D Object-centric Scene Representation Learning},
author = {Zhang, Ellina and Iyengar, Madhavan and Zadeh, Amir and Li, Chuan
and Held, David and Pathak, Deepak and Daniel, Tal},
booktitle = {International Conference on Machine Learning (ICML)},
year = {2026},
url = {https://arxiv.org/abs/2606.19451}
}
This checkpoint is an independent reproduction artifact and is not affiliated with, or endorsed by, the paper's authors.