GlenCarpenter/SwarmUI-SharpSplat

JavaScript

1

59 commits

updated Sep 17, 2026

See the code

README

SwarmUI-SharpSplat

A SwarmUI extension that turns images into 3D Gaussian Splats directly inside the browser.

https://github.com/user-attachments/assets/c71d7912-4fa1-4b15-a6fe-c7ea75f13da8

Nine reconstruction models are supported:

  • ml-sharp (default) — Apple's monocular 3DGS model. Takes a single image and produces a Gaussian Splat in seconds.
  • TripoSplat — VAST-AI's TripoSplat. Takes a single image and produces a full 3D Gaussian Splat using a latent diffusion pipeline with spherical harmonics; often higher fidelity than ml-sharp, especially for object-centric subjects.
  • VGGT — Facebook's Visual Geometry Grounded Transformer (CVPR 2025 Best Paper). Works with a single image or multiple images of the same scene from different angles; more views produce a denser, more accurate point cloud.
  • InstantSplat — NVIDIA's InstantSplat. Takes multiple images and uses MASt3R geometry initialisation to produce a coloured point cloud.
  • Pixal3D — TencentARC's native ComfyUI image-to-3D pipeline. Takes a single image and produces a PBR-textured GLB mesh with camera-aware conditioning.
  • TRELLIS.2 — Microsoft's native ComfyUI image-to-3D pipeline. Takes a single image and produces a PBR-textured GLB mesh.
  • MoGe-1 - Microsoft's single-image geometry estimation baseline, without metric scale.
  • MoGe-2 - Adds metric scale and predicted normals.
  • MoGe-3 - Adds sparse volumetric refinement for finer geometry; uses the ViT-L checkpoint.

Note: VGGT and InstantSplat output geometry-initialised point clouds represented as Gaussians with fixed scale and opacity — they are not the result of a full 3DGS training optimisation loop. Results are usable for previewing and exporting but will not match the quality of a dedicated 3DGS training pipeline.

Gaussian results are saved as .ply (default) or .splat; Pixal3D and TRELLIS.2 results are saved as .glb. MoGe supports textured GLB, untextured GLB, Gaussian PLY, and Gaussian SPLAT. All formats are rendered interactively in the dedicated Splat Viewer tab.


How It Works

Single image           Single image        1+ images (VGGT)     2+ images (InstantSplat)
(ml-sharp)             (TripoSplat)               │                        │
     │                      │                     ▼                        ▼
     ▼                      ▼          [Splat Viewer → drop images & select model]
[Generate 3D Splat button]  │                     │                        │
     │                      │                     │                        │
     └──────────┬───────────┴─────────────────────┴────────────────────────┘
                ▼  (base64 → server)
       SharpSplat API (C#)
                │
    ┌───────────┼─────────────┬───────────┐
    ▼           ▼             ▼           ▼
 ml-sharp   TripoSplat      VGGT     InstantSplat
(single img)(single img)  (1+ imgs)  (2+ imgs, MASt3R)
    │           │             │           │
    └───────────┴─────────────┴───────────┘
                ▼
           .ply output
                │
          [if format = splat]
                ▼
           ply2splat → .splat
           saved to Output/{user}/splats/
                │
                ▼
     Splat Viewer tab (WebGL)

Requirements

  • SwarmUI with a working ComfyUI backend (provides the Python environment).
  • An NVIDIA GPU is strongly recommended for reconstruction.
  • Internet access on first use to download model weights and install Python dependencies.

GPU and memory disclosure: Pixal3D and TRELLIS.2 are substantially more resource-intensive than the splat viewer and the lighter reconstruction paths. They keep the GPU busy for an extended period and can require significant VRAM, system RAM, temporary working memory, and disk space while generating geometry, remeshing, and baking PBR textures. Exact requirements depend on the GPU, backend configuration, and input, but low-VRAM systems may run slowly due to model offloading or fail with an out-of-memory error. Close other GPU-heavy applications and avoid running multiple native 3D jobs concurrently.

Python dependencies are installed automatically on first use:

PackagePurpose
ml-sharpMonocular 3DGS reconstruction (single image)
TripoSplatSingle-image 3DGS reconstruction with SH (model weights ~3 GB, downloaded from HuggingFace on first use)
VGGTMulti-view 3D reconstruction
huggingface_hubDownloads VGGT / TripoSplat model weights (first run only)
InstantSplatMASt3R-based multi-view reconstruction (cloned from GitHub on first use, ~1.2 GB checkpoint downloaded automatically)
ply2splatPLY → .splat conversion (only needed when output format is .splat)

Pixal3D and TRELLIS.2 use ComfyUI's native nodes. Their Comfy-Org model files are downloaded automatically on first use and verified by SHA-256. A fresh TRELLIS.2 installation downloads approximately 8.9 GB; Pixal3D requires a similarly large download plus its MoGe camera-estimation model. These downloads are separate from the temporary memory and output storage used during generation.

MoGe also uses native ComfyUI inference, including MoGe-3 support present in current ComfyUI source even though the ComfyUI tutorial only describes MoGe-1/2. Update ComfyUI and its dependencies, then restart SwarmUI and the backend after installing this extension update. Gaussian output additionally uses the extension's SharpSplatMoGeToSplat adapter and ComfyUI's SplatToFile3D / SaveGaussianSplat nodes. No separate Microsoft MoGe Python installation is required.

Only the selected MoGe checkpoint is downloaded on first use, into SwarmUI's geometry_estimation model folder forwarded to ComfyUI. Downloads use the existing lock, temporary-file cleanup, and SHA-256 verification. Existing files are reused; MoGe-2 also reuses the checkpoint installed for Pixal3D. Sources are the official Comfy-Org/MoGe repository:

ModelCheckpointApproximate Download
MoGe-1moge_1_vitl_fp16.safetensors628 MB
MoGe-2moge_2_vitl_normal_fp16.safetensors662 MB
MoGe-3moge_3_vitl_fp16.safetensors741 MB

Compact SPLAT export automatically installs ply2splat when missing using the existing NumPy-constrained installer. It does not install ml-sharp or VGGT.

MoGe Regression Checks

From the extension directory:

dotnet run --project tests/WorkflowChecks.csproj -p:StaticWebAssetsEnabled=false
python -m unittest discover -s tests -p test_moge.py -v
node --check Assets/sharp_splat.js

Use the ComfyUI Python environment (or another environment with PyTorch) for the Python tests. These checks cover workflow selection, API validation, and geometry-to-Gaussian conversion without downloading model weights or running inference.


Installation

  1. In SwarmUI, go to Server → Extensions.
  2. Click Install Extension and provide this repository's URL.
  3. Restart SwarmUI.

Or clone manually into src/Extensions/:

cd src/Extensions
git clone https://github.com/GlenCarpenter/SwarmUI-SharpSplat

Usage

Generating a splat from a single image (ml-sharp / TripoSplat)

  1. Generate any image in the Generate tab.
  2. Click Generate 3D Splat in the image button bar.
  3. Wait for inference (30–120 seconds depending on GPU).
  4. The Splat Viewer tab opens automatically with the result loaded.

Note: The Generate 3D Splat button in the image viewer uses the selected single-image model, including Pixal3D, TRELLIS.2, and all MoGe versions. For MoGe it also uses the selected MoGe output format. Use the Splat Viewer tab directly for VGGT and InstantSplat.

Generating MoGe geometry or Gaussian splats

  1. Select MoGe-1, MoGe-2, or MoGe-3 in Splat Viewer → Settings → Reconstruction model.
  2. Choose MoGe output: Textured GLB, Untextured GLB, Gaussian PLY, or Gaussian SPLAT.
  3. Set Resolution level (0-9, default 9). MoGe-3 also exposes Refinement steps (0-8, default 3; 0 disables refinement).
  4. Supply one image in the Input Image panel and generate, or use the image viewer's generation button.

GLB outputs are meshes, with the source image embedded as a texture when enabled. Gaussian PLY and SPLAT contain geometry-derived Gaussians, with RGB-derived spherical-harmonic colors, adaptive isotropic scales, and fixed opacity. They are not optimized 3DGS reconstructions, and PLY here means Gaussian PLY, not a triangle mesh or plain point cloud. MoGe estimates visible surfaces only; hidden surfaces and object backsides remain missing. See the MoGe project and MoGe-3 project page.

All outputs are saved in the user's splats directory, listed in the viewer, and downloadable. The output choice and inference settings persist between sessions. The <sharpsplat> automatic prompt tag does not select MoGe.

You can also drop or browse to an image in the Splat Viewer sidebar directly.

Choosing between ml-sharp and TripoSplat:

  • ml-sharp is faster and works well for a wide range of subjects.
  • TripoSplat uses a latent diffusion pipeline that tends to produce higher-quality geometry and colour on object-centric subjects, at the cost of longer inference time (~2–5 minutes on first run while models download).

Generating a splat from a single image (TripoSplat — Splat Viewer)

TripoSplat can also be used from the Splat Viewer tab, which is useful for generating from an image that wasn't produced by SwarmUI.

  1. Open the Splat Viewer tab.
  2. In Settings, change Reconstruction model to TripoSplat.
  3. Drop a single image onto the dropzone, or click Browse to select one.
  4. Click Generate Splat.
  5. On first use, model weights (~3 GB total) are downloaded from HuggingFace automatically before inference begins.

Model files downloaded on first use:

FileDescription
diffusion_models/triposplat_fp16.safetensorsMain denoising UNet
clip_vision/dino_v3_vit_h.safetensorsDINOv3 ViT-H image encoder
vae/triposplat_vae_decoder_fp16.safetensorsSplat VAE decoder
vae/flux2-vae.safetensorsFlux VAE (image encoding)
background_removal/birefnet.safetensorsOptional background removal

Generating a splat from multiple images (VGGT)

VGGT can work from a single image but produces significantly better results with multiple photos of the same subject taken from different angles (like a photogrammetry capture).

  1. Open the Splat Viewer tab.
  2. In Settings, change Reconstruction model to VGGT.
  3. The dropzone in the Input Image section will now accept multiple files.
  4. Drop all your images onto the dropzone, or click Browse to select them.
    Each added image appears as a thumbnail — click × on a thumbnail to remove it.
  5. Click Generate Splat.
  6. VGGT inference runs through the ComfyUI backend (VRAM-managed like normal generations), falling back to a direct subprocess if no ComfyUI backend is available.

Tips for best results:

  • Use 5–20 overlapping photos that cover the subject from many angles.
  • Keep consistent lighting across shots.
  • Avoid motion blur and reflective surfaces.
  • Images are resized to 518 × 518 before inference. If your images are not square, enable Pad images to square in Settings (see below) to preserve the full frame.

Generating a splat from multiple images (InstantSplat)

InstantSplat uses NVIDIA's MASt3R geometry initialisation pipeline and requires at least 2 images. On first use it clones the InstantSplat repository and downloads a ~1.2 GB MASt3R checkpoint automatically.

  1. Open the Splat Viewer tab.
  2. In Settings, change Reconstruction model to InstantSplat.
  3. The dropzone in the Input Image section will accept multiple files.
  4. Drop at least 2 images onto the dropzone, or click Browse to select them.
    Each added image appears as a thumbnail — click × on a thumbnail to remove it.
  5. Click Generate Splat.
  6. Inference runs through the ComfyUI backend, falling back to a direct subprocess if no ComfyUI backend is available.

Tips for best results:

  • Provide at least 2–3 overlapping images; more views improve geometry.
  • Keep consistent lighting and avoid motion blur.
  • Enable Pad images to square in Settings if your images are not square.

Settings

Open Settings in the Splat Viewer sidebar to configure:

SettingDescription
Open in viewer after generationAutomatically navigate to the Splat Viewer tab when a splat finishes.
Reconstruction modelml-sharp, TripoSplat, VGGT, InstantSplat, Pixal3D, TRELLIS.2, MoGe-1, MoGe-2, or MoGe-3.
MoGe outputTextured GLB (default), untextured GLB, Gaussian PLY, or compact Gaussian SPLAT. Separate from the other models' output preference.
Resolution level(MoGe) 0-9, default 9. Higher levels capture more detail at higher inference cost.
Refinement steps(MoGe-3) 0-8, default 3. Sparse refinement passes; 0 disables refinement.
Pad images to square(VGGT / InstantSplat) Resize each input image to fit within a square and pad with neutral grey rather than centre-cropping. Useful when your source images are landscape or portrait. Low-confidence grey border splats are filtered out automatically.
Output formatPLY (default, no conversion) or SPLAT (compact binary, requires ply2splat).
Generate Repair Prompt buttonShows the Generate Repair Prompt button in the Export Canvas section. Intended for use with the ml-sharp repair LoRA — see below. Off by default.

All settings are remembered between sessions.

Automatic generation with the <sharpsplat> prompt tag

Add <sharpsplat> anywhere in your prompt to automatically generate a splat from every image produced by that generation, without clicking the button manually.

a photo of a red apple on a wooden table <sharpsplat>
  • The tag is stripped before it reaches the model — it has no effect on image content.
  • Splat generation runs as a node inside the same ComfyUI job as the image.
  • Batch generations produce one splat per image.
  • The tag uses whichever single-image model is currently selected in Settings → Reconstruction model (ml-sharp or TripoSplat). VGGT and InstantSplat are not available via the prompt tag.

The tag is available in the prompt autocomplete — type <sharpsplat to see it suggested.

Viewing previous splats

  1. Click the Splat Viewer top-level tab.
  2. Previously generated .ply and .splat files are listed in the left sidebar, newest first.
  3. Click any entry to load it into the viewer.
  4. Use ↺ Refresh to update the list after generating new splats.

Exporting the canvas

The Export Canvas section in the Splat Viewer sidebar lets you capture the current rendered frame as a PNG.

  1. Load a splat and position the camera as desired.
  2. Open Export Canvas in the sidebar and click Export Canvas.
  3. Choose a crop ratio from the Resolution dropdown:
    • None (Full) — captures the entire canvas at its current resolution.
    • Aspect ratio presets (1:1, 4:3, 16:9, etc.) — the largest centered crop of that ratio.
    • Custom — enter your own width and height; the largest centered crop matching that ratio is used.
  4. A blue overlay on the canvas shows the region that will be captured.
  5. Click Save to Outputs to save the PNG to Output/local/splats_export/ with a filename of splatname_timestamp.png, or Download to download it directly to your browser's download folder.
  6. Click Cancel to dismiss without exporting.

Generating a repair prompt (ml-sharp)

The Generate Repair Prompt button produces a prompt pre-filled with the current camera movement delta, designed for use with the flux2-klein9b-lora-mlsharp-3d-repair LoRA. The workflow is:

  1. Generate a splat from a single image using ml-sharp.
  2. Export the initial view as a PNG — this becomes image 1 (the reference).
  3. Orbit to the angle you want repaired, then export again — this becomes image 2.
  4. Click Generate Repair Prompt. The prompt is copied to your clipboard with the camera movement encoded as JSON.
  5. Use the copied prompt together with the two exported images and the repair LoRA to inpaint/repair the missing or distorted areas of the novel view.

The prompt takes the form:

Referring to the scene in image 1, restore the perspective of the scene in image 2. Repair the perspective and missing areas. The camera has moved by: {"x":0,"y":0,"z":0,"pitch":0,"yaw":0,"roll":0}

Position values (x, y, z) are world-space translation deltas relative to the initial camera position at scene load. Rotation values (pitch, yaw, roll) are in degrees.

Note: This feature is designed for ml-sharp splats. The repair LoRA was trained on ml-sharp output — results with TripoSplat splats may vary. VGGT and InstantSplat produce multi-view reconstructions with different geometry characteristics that the repair LoRA was not trained for.

This button is hidden by default. Enable it in Settings → Generate Repair Prompt button.

Viewer controls

ActionControl
OrbitLeft-click + drag
ZoomScroll wheel
PanRight-click + drag

Mesh lighting

When a .glb mesh is selected, the sidebar shows a Lighting section. The controls update the rendered view immediately:

ControlEffect
ExposureAdjusts tone-mapping exposure for the complete rendered image.
FillControls soft hemisphere illumination, including light from above and darker fill from below.
KeyControls the intensity of the main directional light.
AzimuthRotates the key light horizontally around the model.
ElevationMoves the key light above or below the model.
Reset LightingRestores the default exposure, intensities, and key-light direction.

Lighting settings are remembered between browser sessions. They affect the interactive viewer and images captured with Export Canvas, but they do not modify the GLB file or its embedded KHR_lights_punctual lights. Lighting controls are hidden for Gaussian splats because splat colors are rendered without the mesh lighting setup.

The built-in mesh viewer is intended for previewing generated assets and producing quick canvas captures, not for replacing a full 3D content-creation application. Downloaded GLB files can be imported into Blender through File → Import → glTF 2.0 (.glb/.gltf), or opened in other software with glTF 2.0 support. From there, continue the workflow with persistent scene lights, cameras, materials, animation, compositing, and renderer-specific effects. The appearance may vary between applications because environment lighting, color management, shadows, and some renderer settings are not stored portably in the GLB.


Roadmap

  • Export PLY as SPLAT or KSPLAT
  • Export SPLAT as KSPLAT

GlenCarpenter/SwarmUI-SharpSplat

JavaScript

1

59 commits

updated Sep 17, 2026

See the code

README

SwarmUI-SharpSplat

A SwarmUI extension that turns images into 3D Gaussian Splats directly inside the browser.

https://github.com/user-attachments/assets/c71d7912-4fa1-4b15-a6fe-c7ea75f13da8

Nine reconstruction models are supported:

  • ml-sharp (default) — Apple's monocular 3DGS model. Takes a single image and produces a Gaussian Splat in seconds.
  • TripoSplat — VAST-AI's TripoSplat. Takes a single image and produces a full 3D Gaussian Splat using a latent diffusion pipeline with spherical harmonics; often higher fidelity than ml-sharp, especially for object-centric subjects.
  • VGGT — Facebook's Visual Geometry Grounded Transformer (CVPR 2025 Best Paper). Works with a single image or multiple images of the same scene from different angles; more views produce a denser, more accurate point cloud.
  • InstantSplat — NVIDIA's InstantSplat. Takes multiple images and uses MASt3R geometry initialisation to produce a coloured point cloud.
  • Pixal3D — TencentARC's native ComfyUI image-to-3D pipeline. Takes a single image and produces a PBR-textured GLB mesh with camera-aware conditioning.
  • TRELLIS.2 — Microsoft's native ComfyUI image-to-3D pipeline. Takes a single image and produces a PBR-textured GLB mesh.
  • MoGe-1 - Microsoft's single-image geometry estimation baseline, without metric scale.
  • MoGe-2 - Adds metric scale and predicted normals.
  • MoGe-3 - Adds sparse volumetric refinement for finer geometry; uses the ViT-L checkpoint.

Note: VGGT and InstantSplat output geometry-initialised point clouds represented as Gaussians with fixed scale and opacity — they are not the result of a full 3DGS training optimisation loop. Results are usable for previewing and exporting but will not match the quality of a dedicated 3DGS training pipeline.

Gaussian results are saved as .ply (default) or .splat; Pixal3D and TRELLIS.2 results are saved as .glb. MoGe supports textured GLB, untextured GLB, Gaussian PLY, and Gaussian SPLAT. All formats are rendered interactively in the dedicated Splat Viewer tab.


How It Works

Single image           Single image        1+ images (VGGT)     2+ images (InstantSplat)
(ml-sharp)             (TripoSplat)               │                        │
     │                      │                     ▼                        ▼
     ▼                      ▼          [Splat Viewer → drop images & select model]
[Generate 3D Splat button]  │                     │                        │
     │                      │                     │                        │
     └──────────┬───────────┴─────────────────────┴────────────────────────┘
                ▼  (base64 → server)
       SharpSplat API (C#)
                │
    ┌───────────┼─────────────┬───────────┐
    ▼           ▼             ▼           ▼
 ml-sharp   TripoSplat      VGGT     InstantSplat
(single img)(single img)  (1+ imgs)  (2+ imgs, MASt3R)
    │           │             │           │
    └───────────┴─────────────┴───────────┘
                ▼
           .ply output
                │
          [if format = splat]
                ▼
           ply2splat → .splat
           saved to Output/{user}/splats/
                │
                ▼
     Splat Viewer tab (WebGL)

Requirements

  • SwarmUI with a working ComfyUI backend (provides the Python environment).
  • An NVIDIA GPU is strongly recommended for reconstruction.
  • Internet access on first use to download model weights and install Python dependencies.

GPU and memory disclosure: Pixal3D and TRELLIS.2 are substantially more resource-intensive than the splat viewer and the lighter reconstruction paths. They keep the GPU busy for an extended period and can require significant VRAM, system RAM, temporary working memory, and disk space while generating geometry, remeshing, and baking PBR textures. Exact requirements depend on the GPU, backend configuration, and input, but low-VRAM systems may run slowly due to model offloading or fail with an out-of-memory error. Close other GPU-heavy applications and avoid running multiple native 3D jobs concurrently.

Python dependencies are installed automatically on first use:

PackagePurpose
ml-sharpMonocular 3DGS reconstruction (single image)
TripoSplatSingle-image 3DGS reconstruction with SH (model weights ~3 GB, downloaded from HuggingFace on first use)
VGGTMulti-view 3D reconstruction
huggingface_hubDownloads VGGT / TripoSplat model weights (first run only)
InstantSplatMASt3R-based multi-view reconstruction (cloned from GitHub on first use, ~1.2 GB checkpoint downloaded automatically)
ply2splatPLY → .splat conversion (only needed when output format is .splat)

Pixal3D and TRELLIS.2 use ComfyUI's native nodes. Their Comfy-Org model files are downloaded automatically on first use and verified by SHA-256. A fresh TRELLIS.2 installation downloads approximately 8.9 GB; Pixal3D requires a similarly large download plus its MoGe camera-estimation model. These downloads are separate from the temporary memory and output storage used during generation.

MoGe also uses native ComfyUI inference, including MoGe-3 support present in current ComfyUI source even though the ComfyUI tutorial only describes MoGe-1/2. Update ComfyUI and its dependencies, then restart SwarmUI and the backend after installing this extension update. Gaussian output additionally uses the extension's SharpSplatMoGeToSplat adapter and ComfyUI's SplatToFile3D / SaveGaussianSplat nodes. No separate Microsoft MoGe Python installation is required.

Only the selected MoGe checkpoint is downloaded on first use, into SwarmUI's geometry_estimation model folder forwarded to ComfyUI. Downloads use the existing lock, temporary-file cleanup, and SHA-256 verification. Existing files are reused; MoGe-2 also reuses the checkpoint installed for Pixal3D. Sources are the official Comfy-Org/MoGe repository:

ModelCheckpointApproximate Download
MoGe-1moge_1_vitl_fp16.safetensors628 MB
MoGe-2moge_2_vitl_normal_fp16.safetensors662 MB
MoGe-3moge_3_vitl_fp16.safetensors741 MB

Compact SPLAT export automatically installs ply2splat when missing using the existing NumPy-constrained installer. It does not install ml-sharp or VGGT.

MoGe Regression Checks

From the extension directory:

dotnet run --project tests/WorkflowChecks.csproj -p:StaticWebAssetsEnabled=false
python -m unittest discover -s tests -p test_moge.py -v
node --check Assets/sharp_splat.js

Use the ComfyUI Python environment (or another environment with PyTorch) for the Python tests. These checks cover workflow selection, API validation, and geometry-to-Gaussian conversion without downloading model weights or running inference.


Installation

  1. In SwarmUI, go to Server → Extensions.
  2. Click Install Extension and provide this repository's URL.
  3. Restart SwarmUI.

Or clone manually into src/Extensions/:

cd src/Extensions
git clone https://github.com/GlenCarpenter/SwarmUI-SharpSplat

Usage

Generating a splat from a single image (ml-sharp / TripoSplat)

  1. Generate any image in the Generate tab.
  2. Click Generate 3D Splat in the image button bar.
  3. Wait for inference (30–120 seconds depending on GPU).
  4. The Splat Viewer tab opens automatically with the result loaded.

Note: The Generate 3D Splat button in the image viewer uses the selected single-image model, including Pixal3D, TRELLIS.2, and all MoGe versions. For MoGe it also uses the selected MoGe output format. Use the Splat Viewer tab directly for VGGT and InstantSplat.

Generating MoGe geometry or Gaussian splats

  1. Select MoGe-1, MoGe-2, or MoGe-3 in Splat Viewer → Settings → Reconstruction model.
  2. Choose MoGe output: Textured GLB, Untextured GLB, Gaussian PLY, or Gaussian SPLAT.
  3. Set Resolution level (0-9, default 9). MoGe-3 also exposes Refinement steps (0-8, default 3; 0 disables refinement).
  4. Supply one image in the Input Image panel and generate, or use the image viewer's generation button.

GLB outputs are meshes, with the source image embedded as a texture when enabled. Gaussian PLY and SPLAT contain geometry-derived Gaussians, with RGB-derived spherical-harmonic colors, adaptive isotropic scales, and fixed opacity. They are not optimized 3DGS reconstructions, and PLY here means Gaussian PLY, not a triangle mesh or plain point cloud. MoGe estimates visible surfaces only; hidden surfaces and object backsides remain missing. See the MoGe project and MoGe-3 project page.

All outputs are saved in the user's splats directory, listed in the viewer, and downloadable. The output choice and inference settings persist between sessions. The <sharpsplat> automatic prompt tag does not select MoGe.

You can also drop or browse to an image in the Splat Viewer sidebar directly.

Choosing between ml-sharp and TripoSplat:

  • ml-sharp is faster and works well for a wide range of subjects.
  • TripoSplat uses a latent diffusion pipeline that tends to produce higher-quality geometry and colour on object-centric subjects, at the cost of longer inference time (~2–5 minutes on first run while models download).

Generating a splat from a single image (TripoSplat — Splat Viewer)

TripoSplat can also be used from the Splat Viewer tab, which is useful for generating from an image that wasn't produced by SwarmUI.

  1. Open the Splat Viewer tab.
  2. In Settings, change Reconstruction model to TripoSplat.
  3. Drop a single image onto the dropzone, or click Browse to select one.
  4. Click Generate Splat.
  5. On first use, model weights (~3 GB total) are downloaded from HuggingFace automatically before inference begins.

Model files downloaded on first use:

FileDescription
diffusion_models/triposplat_fp16.safetensorsMain denoising UNet
clip_vision/dino_v3_vit_h.safetensorsDINOv3 ViT-H image encoder
vae/triposplat_vae_decoder_fp16.safetensorsSplat VAE decoder
vae/flux2-vae.safetensorsFlux VAE (image encoding)
background_removal/birefnet.safetensorsOptional background removal

Generating a splat from multiple images (VGGT)

VGGT can work from a single image but produces significantly better results with multiple photos of the same subject taken from different angles (like a photogrammetry capture).

  1. Open the Splat Viewer tab.
  2. In Settings, change Reconstruction model to VGGT.
  3. The dropzone in the Input Image section will now accept multiple files.
  4. Drop all your images onto the dropzone, or click Browse to select them.
    Each added image appears as a thumbnail — click × on a thumbnail to remove it.
  5. Click Generate Splat.
  6. VGGT inference runs through the ComfyUI backend (VRAM-managed like normal generations), falling back to a direct subprocess if no ComfyUI backend is available.

Tips for best results:

  • Use 5–20 overlapping photos that cover the subject from many angles.
  • Keep consistent lighting across shots.
  • Avoid motion blur and reflective surfaces.
  • Images are resized to 518 × 518 before inference. If your images are not square, enable Pad images to square in Settings (see below) to preserve the full frame.

Generating a splat from multiple images (InstantSplat)

InstantSplat uses NVIDIA's MASt3R geometry initialisation pipeline and requires at least 2 images. On first use it clones the InstantSplat repository and downloads a ~1.2 GB MASt3R checkpoint automatically.

  1. Open the Splat Viewer tab.
  2. In Settings, change Reconstruction model to InstantSplat.
  3. The dropzone in the Input Image section will accept multiple files.
  4. Drop at least 2 images onto the dropzone, or click Browse to select them.
    Each added image appears as a thumbnail — click × on a thumbnail to remove it.
  5. Click Generate Splat.
  6. Inference runs through the ComfyUI backend, falling back to a direct subprocess if no ComfyUI backend is available.

Tips for best results:

  • Provide at least 2–3 overlapping images; more views improve geometry.
  • Keep consistent lighting and avoid motion blur.
  • Enable Pad images to square in Settings if your images are not square.

Settings

Open Settings in the Splat Viewer sidebar to configure:

SettingDescription
Open in viewer after generationAutomatically navigate to the Splat Viewer tab when a splat finishes.
Reconstruction modelml-sharp, TripoSplat, VGGT, InstantSplat, Pixal3D, TRELLIS.2, MoGe-1, MoGe-2, or MoGe-3.
MoGe outputTextured GLB (default), untextured GLB, Gaussian PLY, or compact Gaussian SPLAT. Separate from the other models' output preference.
Resolution level(MoGe) 0-9, default 9. Higher levels capture more detail at higher inference cost.
Refinement steps(MoGe-3) 0-8, default 3. Sparse refinement passes; 0 disables refinement.
Pad images to square(VGGT / InstantSplat) Resize each input image to fit within a square and pad with neutral grey rather than centre-cropping. Useful when your source images are landscape or portrait. Low-confidence grey border splats are filtered out automatically.
Output formatPLY (default, no conversion) or SPLAT (compact binary, requires ply2splat).
Generate Repair Prompt buttonShows the Generate Repair Prompt button in the Export Canvas section. Intended for use with the ml-sharp repair LoRA — see below. Off by default.

All settings are remembered between sessions.

Automatic generation with the <sharpsplat> prompt tag

Add <sharpsplat> anywhere in your prompt to automatically generate a splat from every image produced by that generation, without clicking the button manually.

a photo of a red apple on a wooden table <sharpsplat>
  • The tag is stripped before it reaches the model — it has no effect on image content.
  • Splat generation runs as a node inside the same ComfyUI job as the image.
  • Batch generations produce one splat per image.
  • The tag uses whichever single-image model is currently selected in Settings → Reconstruction model (ml-sharp or TripoSplat). VGGT and InstantSplat are not available via the prompt tag.

The tag is available in the prompt autocomplete — type <sharpsplat to see it suggested.

Viewing previous splats

  1. Click the Splat Viewer top-level tab.
  2. Previously generated .ply and .splat files are listed in the left sidebar, newest first.
  3. Click any entry to load it into the viewer.
  4. Use ↺ Refresh to update the list after generating new splats.

Exporting the canvas

The Export Canvas section in the Splat Viewer sidebar lets you capture the current rendered frame as a PNG.

  1. Load a splat and position the camera as desired.
  2. Open Export Canvas in the sidebar and click Export Canvas.
  3. Choose a crop ratio from the Resolution dropdown:
    • None (Full) — captures the entire canvas at its current resolution.
    • Aspect ratio presets (1:1, 4:3, 16:9, etc.) — the largest centered crop of that ratio.
    • Custom — enter your own width and height; the largest centered crop matching that ratio is used.
  4. A blue overlay on the canvas shows the region that will be captured.
  5. Click Save to Outputs to save the PNG to Output/local/splats_export/ with a filename of splatname_timestamp.png, or Download to download it directly to your browser's download folder.
  6. Click Cancel to dismiss without exporting.

Generating a repair prompt (ml-sharp)

The Generate Repair Prompt button produces a prompt pre-filled with the current camera movement delta, designed for use with the flux2-klein9b-lora-mlsharp-3d-repair LoRA. The workflow is:

  1. Generate a splat from a single image using ml-sharp.
  2. Export the initial view as a PNG — this becomes image 1 (the reference).
  3. Orbit to the angle you want repaired, then export again — this becomes image 2.
  4. Click Generate Repair Prompt. The prompt is copied to your clipboard with the camera movement encoded as JSON.
  5. Use the copied prompt together with the two exported images and the repair LoRA to inpaint/repair the missing or distorted areas of the novel view.

The prompt takes the form:

Referring to the scene in image 1, restore the perspective of the scene in image 2. Repair the perspective and missing areas. The camera has moved by: {"x":0,"y":0,"z":0,"pitch":0,"yaw":0,"roll":0}

Position values (x, y, z) are world-space translation deltas relative to the initial camera position at scene load. Rotation values (pitch, yaw, roll) are in degrees.

Note: This feature is designed for ml-sharp splats. The repair LoRA was trained on ml-sharp output — results with TripoSplat splats may vary. VGGT and InstantSplat produce multi-view reconstructions with different geometry characteristics that the repair LoRA was not trained for.

This button is hidden by default. Enable it in Settings → Generate Repair Prompt button.

Viewer controls

ActionControl
OrbitLeft-click + drag
ZoomScroll wheel
PanRight-click + drag

Mesh lighting

When a .glb mesh is selected, the sidebar shows a Lighting section. The controls update the rendered view immediately:

ControlEffect
ExposureAdjusts tone-mapping exposure for the complete rendered image.
FillControls soft hemisphere illumination, including light from above and darker fill from below.
KeyControls the intensity of the main directional light.
AzimuthRotates the key light horizontally around the model.
ElevationMoves the key light above or below the model.
Reset LightingRestores the default exposure, intensities, and key-light direction.

Lighting settings are remembered between browser sessions. They affect the interactive viewer and images captured with Export Canvas, but they do not modify the GLB file or its embedded KHR_lights_punctual lights. Lighting controls are hidden for Gaussian splats because splat colors are rendered without the mesh lighting setup.

The built-in mesh viewer is intended for previewing generated assets and producing quick canvas captures, not for replacing a full 3D content-creation application. Downloaded GLB files can be imported into Blender through File → Import → glTF 2.0 (.glb/.gltf), or opened in other software with glTF 2.0 support. From there, continue the workflow with persistent scene lights, cameras, materials, animation, compositing, and renderer-specific effects. The appearance may vary between applications because environment lighting, color management, shadows, and some renderer settings are not stored portably in the GLB.


Roadmap

  • Export PLY as SPLAT or KSPLAT
  • Export SPLAT as KSPLAT