- TripoSplat → Qwen-Image 2.1 — 3D Angle Workflow V2
- 1. Stage One — Create the 3D Model
- 2. Stage One — Optional Orbit Preview
- 3. Stage Two — Select a New Camera View
- 4. Qwen-Image 2.1 — Generate the New View
- 5. Qwen-Image 2.1 Configuration
- 6. AnyAngle LoRA
- 7. Qwen Generation Settings
- 8. LCS Tone Adjustment
- 9. Latent Routing
- 10. Final Image
- 11. Comparison Output
- 🚀 Quick Start
- 🧠 The Core Idea
- 📁 Required Models
- 🔌 Custom Nodes / Components
- ⚠️ Important Notes
- 📚 References
- 📜 License
TripoSplat → Qwen-Image 2.1 — 3D Angle Workflow V2
This workflow has 2 stages
1. Create the 3D model using TripoSplat
A single reference image is converted into a 3D Gaussian Splat and exported as an SPZ 3D model.2. Make a new view with Qwen-Image 2.1 using the 3D model
The SPZ model is loaded into ComfyUI's Load3D viewer. You choose the desired camera angle interactively, and that rendered 3D view is sent to Qwen-Image 2.1 together with the original reference image. Qwen then generates the new view.┌──────────────────────────────────────────────────────────────┐ │ STAGE 1 │ │ CREATE THE 3D MODEL │ └──────────────────────────────────────────────────────────────┘ Reference Image │ ▼ ┌─────────────┐ │ TripoSplat │ └──────┬──────┘ │ ▼ 3D Gaussian Splat │ ▼ SplatToFile3D │ ▼ SPZ ┌──────────────────────────────────────────────────────────────┐ │ STAGE 2 │ │ CREATE A NEW CAMERA VIEW │ └──────────────────────────────────────────────────────────────┘ SPZ 3D Model │ ▼ ┌────────┐ │ Load3D │ ◄── interactively choose camera angle └────┬───┘ │ ▼ 3D View IMAGE │ ├──────────────────────┐ │ │ Original Reference Selected 3D View │ │ └──────────┬───────────┘ ▼ ┌──────────────────┐ │ Qwen-Image 2.1 │ │ + AnyAngle LoRA │ └────────┬─────────┘ │ ▼ New Camera View
✨ What this workflow does
The workflow separates 3D spatial control from 2D image generation.
SPATIAL CONTROL
│
▼
Reference ──► TripoSplat ──► SPZ ──► Load3D
│
│ choose angle
▼
3D View Image
│
▼
┌────────────────────┴───────────────────┐
│ │
Original Image 3D View
│ │
└────────────────┬───────────────────────┘
▼
Qwen-Image 2.1
│
▼
FINAL NEW VIEW
The important idea is that Load3D provides the camera/viewpoint reference, while Qwen-Image 2.1 creates the final image.
1. Stage One — Create the 3D Model
TripoSplat
The first stage starts with a single reference image:
LoadImage
│
▼
Image to Gaussian Splat (TripoSplat)
│
▼
3D Gaussian Splat
The current workflow uses:
| Setting | Value |
|---|---|
| Remove background | False |
| Number of Gaussians | 262144 |
| Seed | 81 |
| TripoSplat model | triposplat_fp16.safetensors |
| CLIP Vision | dino_v3_vit_h.safetensors |
| Splat VAE | triposplat_vae_decoder_fp16.safetensors |
| Flux VAE | flux2-vae.safetensors |
| Background removal model | birefnet.safetensors |
| Preview | Enabled |
The workflow note explains that 262144 is the final Gaussian count / octree density. Increasing the count oversamples existing points and can increase VRAM/time usage; it does not automatically create additional source detail.
SPZ Export
The active 3D export path is:
TripoSplat
│
▼
SplatToFile3D
│
│ format = spz
▼
.SPZ
The workflow uses SPZ as the primary 3D model format for Stage 2.
The saved model prefix is:
3d/ComfyUI-TripoSplat-SPZ-model
So the resulting file will be saved as an .spz model.
Why SPZ?
The workflow keeps the Gaussian Splat representation instead of converting it to a conventional polygon mesh before the camera-selection stage.
Gaussian Splat
│
└──► SPZ
│
▼
Load3D
The GLB mesh branch remains available in the workflow as an experimental/bypassed alternative, but it is not required for the main TripoSplat → Load3D → Qwen workflow.
2. Stage One — Optional Orbit Preview
The workflow also creates an orbit preview directly from the Gaussian Splat.
TripoSplat
│
▼
CreateCameraInfo
│
▼
RenderSplat
│
▼
CreateVideo
│
▼
SaveVideo
Current camera settings:
| Parameter | Value |
|---|---|
| Mode | orbit |
| Yaw | 35° |
| Pitch | 30° |
| Distance | 1.5 |
| Target X | 0 |
| Target Y | 0 |
| Target Z | 0 |
| Roll | 0° |
| FOV | 50° |
| Zoom | 1 |
| Camera | perspective |
Current RenderSplat settings:
| Parameter | Value |
|---|---|
| Width | 1024 |
| Height | 1024 |
| Frames | 75 |
| Splat scale | 1 |
| Sharpen | 2 |
| Headlight shading | 0 |
| Opacity threshold | 0 |
| Render style | color |
| Background | #848484 |
The video is saved with the prefix:
video/ComfyUI_TripoSplat
This orbit branch is primarily a preview / visualization of the generated Gaussian Splat.
3. Stage Two — Select a New Camera View
The generated SPZ model is loaded into:
Load3D
The current workflow loads:
3d/ComfyUI_TripoSplat_00015_.spz
When using a new reference image, replace this with the SPZ generated by Stage 1.
The important connection is:
SPZ
│
▼
Load3D
│
└── IMAGE ───────────────► Qwen image_2
Load3D also exposes:
imagemaskmesh_pathnormalcamera_inforecording_videomodel_3dmodel_3d_info
For this workflow, the important output is the IMAGE produced by the selected 3D view.
🎥 The Load3D Camera
Use the interactive Load3D viewport to choose the desired camera:
┌───────────────┐
│ LOAD3D │
│ │
│ 3D MODEL │
│ ◉ │
│ / │
│ / camera │
│ / │
└───────────────┘
│
▼
VIEW IMAGE
You can change:
- Rotation
- Camera angle
- Elevation
- Distance
- Framing
- Perspective
- Subject placement
This is the creative camera-control stage of the workflow.
4. Qwen-Image 2.1 — Generate the New View
The selected Load3D view is passed into:
TextEncodeQwenImage21
The workflow uses two image references:
image_1 = original reference image
image_2 = selected Load3D 3D view
The current prompt is:
Change the camera angle from <image2> to <image1>.
Conceptually:
IMAGE 1 IMAGE 2
Original Reference Selected 3D View
│ │
│ │
└──────────────┬────────────────┘
▼
TextEncodeQwenImage21
│
▼
Qwen conditioning
│
▼
KSampler
│
▼
VAE Decode
│
▼
New View
The workflow therefore uses the 3D model to establish where the camera should be, while Qwen handles the final image reconstruction.
5. Qwen-Image 2.1 Configuration
Diffusion Model
qwen_image_2.1_int8_convrot.safetensors
Loaded with:
UNETLoader
weight_dtype = default
Text Encoder
qwen3vl_8b_int8_convrot.safetensors
Loaded as:
type = qwen_image
device = default
VAE
qwen_image_2.1_vae_bf16.safetensors
6. AnyAngle LoRA
The workflow applies:
Qwen_i21/QI2.1_AnyAngle.safetensors
through:
LoraLoaderModelOnly
Current strength:
1.0
Pipeline:
Qwen-Image 2.1
│
▼
QI2.1 AnyAngle LoRA
│
▼
QwenImage21Cache
│
▼
KSampler
7. Qwen Generation Settings
Current ResolutionSelector:
| Setting | Value |
|---|---|
| Aspect ratio | 1:1 (Square) |
| Megapixels | 1 |
| Multiple | 32 |
Current output resolution:
1024 × 1024
Current KSampler:
| Setting | Value |
|---|---|
| Steps | 25 |
| CFG | 1 |
| Sampler | euler |
| Scheduler | simple |
| Denoise | 1 |
| Seed | Randomized after generation |
8. LCS Tone Adjustment
The workflow uses:
LCSToneAdjust
with:
| Parameter | Value |
|---|---|
| Preset | Custom |
| Contrast | 0.95 |
| Brightness | 0.03 |
| Saturation | 0.9 |
| Color temperature | 0.02 |
| Start step | 30 |
| End step | 33 |
The adjusted model is then passed through:
QwenImage21Cache
before sampling.
9. Latent Routing
The workflow contains a latent switch:
TextEncodeQwenImage21
│
│ latent
▼
ComfySwitchNode ◄──── EmptyLatentImage
│
▼
KSampler
The current switch is:
False
The EmptyLatentImage branch uses the selected output resolution.
Current empty latent:
1024 × 1024
batch size = 1
10. Final Image
The generation path is:
KSampler
│
▼
VAEDecode
│
▼
SaveImageAdvanced
Current output:
Format: PNG
Bit depth: 8-bit
Color space: sRGB
Filename prefix:
Qwen_image_2.1/Qwen_image_2.1-3D-Angle
11. Comparison Output
The workflow includes an image comparison and presentation section.
┌─────────────────────┐
│ Original Image │
└──────────┬──────────┘
│
▼
┌───────────┐
│ ImageReel │
└─────┬─────┘
│
├── Original
├── 3D View
└── Qwen Image 2.1
│
▼
ImageReelComposit
│
▼
Comparison
The reel labels are:
Image1
image 2
Qwen Image 2.1
The final comparison image is saved with:
Qwen_image_2.1/Qwen_image_2.1-3D-Angle-Compare
An ImageCompare node is also included for direct comparison between the source and generated result.
🚀 Quick Start
Stage 1 — Create the 3D Model
1. Load your reference image
Replace the image in:
LoadImage
The current workflow uses:
Kr-Queen_0992_.jpg
The same source image is also prepared at:
1024 × 1024
for the Qwen stage.
2. Run TripoSplat
Queue the workflow to generate the Gaussian Splat.
The active outputs are:
TripoSplat
├──► Orbit Preview
└──► SPZ 3D Model
3. Use the generated SPZ
The workflow's current Load3D reference is:
3d/ComfyUI_TripoSplat_00015_.spz
Replace this with the SPZ generated from your current reference.
Stage 2 — Create the New View
4. Open the Load3D viewer
Load the SPZ into:
Load3D
5. Choose the desired camera angle
Rotate and position the model until the desired viewpoint is established.
6. Run Qwen-Image 2.1
Qwen receives:
Original Image
+
Selected 3D View
↓
Qwen-Image 2.1
↓
New Camera View
🧠 The Core Idea
This workflow is intentionally not simply:
Image → Qwen → Different Image
Instead, it introduces an explicit 3D spatial stage:
Image
│
▼
┌──────────────┐
│ TripoSplat │
└──────┬───────┘
│
▼
SPZ 3D
│
▼
┌──────────────┐
│ Load3D │
│ │
│ Camera Angle │
└──────┬───────┘
│
▼
3D View Image
│
▼
┌──────────────┐
│ Qwen-Image │
│ 2.1 │
└──────┬───────┘
│
▼
New View
This provides a practical separation between:
Where the camera is
→ controlled in 3D
and
How the final image looks
→ generated by Qwen-Image 2.1
📁 Required Models
TripoSplat
ComfyUI/
└── models/
├── background_removal/
│ └── birefnet.safetensors
│
├── diffusion_models/
│ └── triposplat_fp16.safetensors
│
├── clip_vision/
│ └── dino_v3_vit_h.safetensors
│
└── vae/
├── triposplat_vae_decoder_fp16.safetensors
└── flux2-vae.safetensors
Qwen-Image 2.1
ComfyUI/
└── models/
├── diffusion_models/
│ └── qwen_image_2.1_int8_convrot.safetensors
│
├── text_encoders/
│ └── qwen3vl_8b_int8_convrot.safetensors
│
├── vae/
│ └── qwen_image_2.1_vae_bf16.safetensors
│
└── loras/
└── Qwen_i21/
└── QI2.1_AnyAngle.safetensors
🔌 Custom Nodes / Components
The workflow uses:
TripoSplat
Image to Gaussian Splat (TripoSplat)SplatToFile3DRenderSplatCreateCameraInfo
ComfyUI 3D
Load3DSaveGLBSplatToMesh
Qwen
UNETLoaderCLIPLoaderVAELoaderTextEncodeQwenImage21QwenImage21CacheKSamplerVAEDecodeLoraLoaderModelOnly
Additional workflow nodes
LCSToneAdjustLCSLoadDataResolutionSelectorLayerUtility: ImageReelLayerUtility: ImageReelCompositImageCompare
⚠️ Important Notes
SPZ is the primary 3D path
The main workflow uses:
TripoSplat → SPZ → Load3D
The GLB mesh branch is an optional experimental export path and is bypassed in the TripoSplat note.
Unseen geometry
TripoSplat reconstructs a 3D representation from a single image. Areas not visible in the source may therefore be incomplete or less reliable.
Extreme camera rotations can expose these limitations.
Qwen is the final reconstruction
The final Qwen image is a generated interpretation of the selected 3D viewpoint. It is not intended to be a pixel-perfect render of the SPZ.
📚 References
TripoSplat
- GitHub: https://github.com/VAST-AI-Research/TripoSplat
- Paper: https://arxiv.org/abs/2605.16355
- Model: https://huggingface.co/VAST-AI/TripoSplat
Qwen-Image 2.1
The workflow uses the Comfy-Org Qwen-Image 2.1 model files:
qwen_image_2.1_int8_convrot.safetensors
qwen3vl_8b_int8_convrot.safetensors
qwen_image_2.1_vae_bf16.safetensors
📜 License
This README describes a ComfyUI workflow configuration.
The workflow itself may be distributed under the license selected by the workflow author. The licenses and terms of the underlying models and third-party components remain applicable.
Check the respective terms for:
- TripoSplat
- Qwen-Image 2.1
- QI2.1 AnyAngle LoRA
- BiRefNet
- ComfyUI
- Third-party custom nodes
Workflow
TripoSplat → Gaussian Splat → SPZ → Load3D Camera → Qwen-Image 2.1 AnyAngle
Stage 1: Create the 3D model.
Stage 2: Use the 3D model to define a new camera view and generate the final image with Qwen-Image 2.1.
Model tree for zuanfilm/3D_Model_Multiview_Qwen_i21
Base model
Qwen/Qwen-Image-2.1