This is a collective repository for all 3D and 4D Object Generation papers
20
15 commits
updated May 22, 2026
A curated collection of cutting-edge research papers on 3D and 4D Content Generation.
If you find this repo useful, please consider giving it a ⭐!
| List | Description | |
|---|---|---|
| ✏️ | Awesome 3D/4D Editing | Papers on 3D and 4D scene/object editing |
| 🔄 | Awesome FeedForward 3D/4D Reconstruction | Papers on feed-forward 3D/4D reconstruction |
Score Distillation Sampling approaches for text/image-to-3D generation.
| # | Paper | Link |
|---|---|---|
| 1 | DreamFusion: Text-to-3D Using 2D Diffusion | 📄 |
| 2 | Latent-NeRF for Shape-Guided Generation of 3D Shapes and Textures | 📄 |
| 3 | Magic3D: High-Resolution Text-to-3D Content Creation | 📄 |
| 4 | NerfDiff: Single-Image View Synthesis with NeRF-Guided Distillation from 3D-Aware Diffusion | 📄 |
| 5 | DreamBooth3D: Subject-Driven Text-to-3D Generation | 📄 |
| 6 | Fantasia3D: Disentangling Geometry and Appearance for High-Quality Text-to-3D Content Creation | 📄 |
| 7 | Make-It-3D: High-Fidelity 3D Creation from a Single Image with Diffusion Prior | 📄 |
| 8 | TextMesh: Generation of Realistic 3D Meshes From Text Prompts | 📄 |
| 9 | IT3D: Improved Text-to-3D Generation with Explicit View Synthesis | 📄 |
| 10 | Text-to-3D Using Gaussian Splatting | 📄 |
| 11 | DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content Creation | 📄 |
| 12 | GaussianDreamer: Fast Generation from Text to 3D Gaussians by Bridging 2D and 3D Diffusion Models | 📄 |
| 13 | DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion Prior | 📄 |
| 14 | Learn to Optimize Denoising Scores for 3D Generation | 📄 |
| 15 | REPARO: Compositional 3D Assets Generation with Differentiable 3D Layout Alignment | 📄 |
| 16 | Recon3D: High Quality 3D Reconstruction from a Single Image Using Generated Back-View Explicit Priors | 📄 |
| 17 | COMOGen: A Controllable Text-to-3D Multi-Object Generation Framework | 📄 |
| 18 | Enhancing Single Image to 3D Generation Using Gaussian Splatting and Hybrid Diffusion Priors | 📄 |
| 19 | ModeDreamer: Mode Guiding Score Distillation for Text-to-3D Generation | 📄 |
| 20 | Enhanced 3D Generation by 2D Editing | 📄 |
| 21 | Diverse Score Distillation | 📄 |
| 22 | Chirpy3D: Continuous Part Latents for Creative 3D Bird Generation | 📄 |
| 23 | Probability-Flow Distillation: Exact Wasserstein Gradient Flow for High-Fidelity 3D Generation | 📄 |
SDS-based dynamic 4D content generation.
| # | Paper | Link |
|---|---|---|
| 1 | Text-to-4D Dynamic Scene Generation | 📄 |
| 2 | Consistent4D: Consistent 360° Dynamic Object Generation from Monocular Video | 📄 |
| 3 | Animate124: Animating One Image to 4D Dynamic Scene | 📄 |
| 4 | 4D-fy: Text-to-4D Generation Using Hybrid Score Distillation Sampling | 📄 |
| 5 | DreamGaussian4D: Generative 4D Gaussian Splatting | 📄 |
| 6 | SC4D: Sparse-Controlled Video-to-4D Generation and Motion Transfer | 📄 |
Multi-view diffusion models for 3D-consistent image generation and reconstruction.
| # | Paper | Link |
|---|---|---|
| 1 | Zero-1-to-3: Zero-Shot One Image to 3D Object | 📄 |
| 2 | One-2-3-45: Any Single Image to 3D Mesh in 45 Seconds Without Per-Shape Optimization | 📄 |
| 3 | MVDiffusion: Enabling Holistic Multi-View Image Generation with Correspondence-Aware Diffusion | 📄 |
| 4 | MVDream: Multi-View Diffusion for 3D Generation | 📄 |
| 5 | SyncDreamer: Generating Multiview-Consistent Images from a Single-View Image | 📄 |
| 6 | Wonder3D: Single Image to 3D Using Cross-Domain Diffusion | 📄 |
| 7 | Zero123++: A Single Image to Consistent Multi-View Diffusion Base Model | 📄 |
| 8 | DMV3D: Denoising Multi-View Diffusion Using 3D Large Reconstruction Model | 📄 |
| 9 | ViVid-1-to-3: Novel View Synthesis with Video Diffusion Models | 📄 |
| 10 | ImageDream: Image-Prompt Multi-View Diffusion for 3D Generation | 📄 |
| 11 | 4DGen: Grounded 4D Content Generation with Spatial-Temporal Consistency | 📄 |
| 12 | EscherNet: A Generative Model for Scalable View Synthesis | 📄 |
| 13 | LGM: Large Multi-View Gaussian Model for High-Resolution 3D Content Creation | 📄 |
| 14 | IM-3D: Iterative Multiview Diffusion and Reconstruction for High-Quality 3D Generation | 📄 |
| 15 | MVDiffusion++: A Dense High-Resolution Multi-View Diffusion Model for 3D Object Reconstruction | 📄 |
| 16 | V3D: Video Diffusion Models Are Effective 3D Generators | 📄 |
| 17 | Make-Your-3D: Fast and Consistent Subject-Driven 3D Content Generation | 📄 |
| 18 | SV3D: Novel Multi-View Synthesis and 3D Generation from a Single Image Using Latent Video Diffusion | 📄 |
| 19 | VFusion3D: Learning Scalable 3D Generative Models from Video Diffusion Models | 📄 |
| 20 | CAT3D: Create Anything in 3D with Multi-View Diffusion Models | 📄 |
| 21 | Era3D: High-Resolution Multiview Diffusion Using Efficient Row-Wise Attention | 📄 |
| 22 | Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models | 📄 |
| 23 | Enhancing Single Image to 3D Generation Using Gaussian Splatting and Hybrid Diffusion Priors | 📄 |
| 24 | DreamCraft3D++: Efficient Hierarchical 3D Generation with Multi-Plane Reconstruction Model | 📄 |
| 25 | MVPaint: Synchronized Multi-View Diffusion for Painting Anything 3D | 📄 |
| 26 | Edify 3D: Scalable High-Quality 3D Asset Generation | 📄 |
| 27 | Direct and Explicit 3D Generation from a Single Image | 📄 |
| 28 | Fancy123: One Image to High-Quality 3D Mesh Generation via Plug-and-Play Deformation | 📄 |
| 29 | ModeDreamer: Mode Guiding Score Distillation for Text-to-3D Generation | 📄 |
| 30 | RIGI: Rectifying Image-to-3D Generation Inconsistency via Uncertainty-Aware Learning | 📄 |
| 31 | Turbo3D: Ultra-Fast Text-to-3D Generation | 📄 |
| 32 | Gen-3Diffusion: Realistic Image-to-3D Generation via 2D & 3D Diffusion Synergy | 📄 |
| 33 | PartGen: Part-Level 3D Generation and Reconstruction with Multi-View Diffusion Models | 📄 |
Multi-view generation approaches for dynamic 4D content.
| # | Paper | Link |
|---|---|---|
| 1 | 4DGen: Grounded 4D Content Generation with Spatial-Temporal Consistency | 📄 |
| 2 | STAG4D: Spatial-Temporal Anchored Generative 4D Gaussians | 📄 |
| 3 | Diffusion4D: Fast Spatial-Temporal Consistent 4D Generation via Video Diffusion Models | 📄 |
| 4 | L4GM: Large 4D Gaussian Reconstruction Model | 📄 |
| 5 | SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency | 📄 |
Large Reconstruction Model (LRM) based feed-forward 3D generation.
| # | Paper | Link |
|---|---|---|
| 1 | LRM: Large Reconstruction Model for Single Image to 3D | 📄 |
| 2 | Instant3D: Fast Text-to-3D with Sparse-View Generation and Large Reconstruction Model | 📄 |
| 3 | DMV3D: Denoising Multi-View Diffusion Using 3D Large Reconstruction Model | 📄 |
| 4 | PF-LRM: Pose-Free Large Reconstruction Model for Joint Pose and Shape Prediction | 📄 |
| 5 | TripoSR: Fast 3D Object Reconstruction from a Single Image | 📄 |
| 6 | GRM: Large Gaussian Reconstruction Model for Efficient 3D Reconstruction and Generation | 📄 |
| 7 | InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-View Large Reconstruction Models | 📄 |
| 8 | M-LRM: Multi-View Large Reconstruction Model | 📄 |
| 9 | LRM-Zero: Training Large Reconstruction Models with Synthesized Data | 📄 |
| 10 | ControLRM: Fast and Controllable 3D Generation via Large Reconstruction Model | 📄 |
| 11 | Enhancing Single Image to 3D Generation Using Gaussian Splatting and Hybrid Diffusion Priors | 📄 |
| 12 | 3D-Adapter: Geometry-Consistent Multi-View Diffusion for High-Quality 3D Generation | 📄 |
| 13 | Fancy123: One Image to High-Quality 3D Mesh Generation via Plug-and-Play Deformation | 📄 |
Naïve 3D native generation methods (diffusion in 3D space, autoregressive, etc.).
| # | Paper | Link |
|---|---|---|
| 1 | DiffRF: Rendering-Guided 3D Radiance Field Diffusion | 📄 |
| 2 | Point-E: A System for Generating 3D Point Clouds from Complex Prompts | 📄 |
| 3 | Shap-E: Generating Conditional 3D Implicit Functions | 📄 |
| 4 | GSD: View-Guided Gaussian Splatting Diffusion for 3D Reconstruction | 📄 |
| 5 | 3DTopia-XL: Scaling High-Quality 3D Asset Generation via Primitive Diffusion | 📄 |
| 6 | SeMv-3D: Towards Semantic and Multi-View Consistency for General Text-to-3D Generation with Triplane Priors | 📄 |
| 7 | L3DG: Latent 3D Gaussian Diffusion | 📄 |
| 8 | LucidFusion: Generating 3D Gaussians with Arbitrary Unposed Images | 📄 |
| 9 | GaussianAnything: Interactive Point Cloud Latent Diffusion for 3D Generation | 📄 |
| 10 | SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-Scale 3D VQVAE | 📄 |
| 11 | A Lesson in Splats: Teacher-Guided Diffusion for 3D Gaussian Splats Generation with 2D Supervision | 📄 |
| 12 | TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction | 📄 |
| 13 | RELATE3D: REfocusing Latent Adapter for Targeted Local Enhancement and Editing in 3D Generation | 📄 |
| 14 | AssetFormer: Modular 3D Assets Generation with Autoregressive Transformer | 📄 |
| 15 | VAR-3D: View-Aware Auto-Regressive Model for Text-to-3D Generation via a 3D Tokenizer | 📄 |
| 16 | Fuse3D: Generating 3D Assets Controlled by Multi-Image Fusion | 📄 |
| 17 | GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation | 📄 |
| 18 | Omni123: Exploring 3D Native Foundation Models with Limited 3D Data | 📄 |
| 19 | Hitem3D 2.0: Multi-View Guided Native 3D Texture Generation | 📄 |
Implicit latent space representations for 3D shape generation.
| # | Paper | Link |
|---|---|---|
| 1 | 3DShape2VecSet: A 3D Shape Representation for Neural Fields and Generative Diffusion Models | 📄 |
| 2 | Michelangelo: Conditional 3D Shape Generation Based on Shape-Image-Text Aligned Latent Representation | 📄 |
| 3 | CraftsMan3D: High-Fidelity Mesh Generation with 3D Native Generation and Interactive Geometry Refiner | 📄 |
| 4 | CLAY: A Controllable Large-Scale Generative Model for Creating High-Quality 3D Assets | 📄 |
| 5 | Dora: Sampling and Benchmarking for 3D Shape Variational Auto-Encoders | 📄 |
| 6 | TripoSG: High-Fidelity 3D Shape Synthesis Using Large-Scale Rectified Flow Models | 📄 |
| 7 | Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets | 📄 |
| 8 | Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material | 📄 |
| 9 | Hunyuan3D 2.5: Towards High-Fidelity 3D Assets Generation with Ultimate Details | 📄 |
| 10 | Seed3D 1.0: From images to high-fidelity simulation-ready 3D assets | 📄 |
| 11 | LATTICE: Democratize high-fidelity 3D generation at scale | 📄 |
| 12 | UltraShape 1.0: High-Fidelity 3D Shape Generation via Scalable Geometric Refinement | 📄 |
| 13 | Seed3D 2.0: Advancing High-Fidelity Simulation-Ready 3D Content Generation | 📄 |
| 14 | Pose-Aware Diffusion for 3D Generation | 📄 |
| 15 | ROAR-3D: Routing arbitrary views for high-fidelity 3D generation | 📄 |
| # | Paper | Link |
|---|---|---|
| 1 | Sculpt4D: Generating 4D Shapes via Sparse-Attention Diffusion Transformers | 📄 |
Explicit latent space representations (Gaussians, structured latents, etc.) for generation.
| # | Paper | Link |
|---|---|---|
| 1 | GaussianCube: A Structured and Explicit Radiance Representation for 3D Generative Modeling | 📄 |
| 2 | Structured 3D Latents for Scalable and Versatile 3D Generation | 📄 |
| 3 | SynCity: Training-Free Generation of 3D Worlds | 📄 |
| 4 | Hi3DGen: High-Fidelity 3D Geometry Generation from Images via Normal Bridging | 📄 |
| 5 | DSO: Aligning 3D Generators with Simulation Feedback for Physical Soundness | 📄 |
| 6 | Sparc3D: Sparse Representation and Construction for High-Resolution 3D Shapes Modeling | 📄 |
| 7 | Ultra3D: Efficient and High-Fidelity 3D Generation with Part Attention | 📄 |
| 8 | Few-Step Flow for 3D Generation via Marginal-Data Transport Distillation | 📄 |
| 9 | ReconViaGen: Towards Accurate Multi-View 3D Object Reconstruction via Generation | 📄 |
| 10 | SAM 3D: 3Dfy Anything in Images | 📄 |
| 11 | Wukong's 72 transformations: High-fidelity textured 3D morphing via flow models | 📄 |
| 12 | Native and Compact Structured Latents for 3D Generation | 📄 |
| 13 | MorphAny3D: Unleashing the power of structured latent in 3D morphing | 📄 |
| 14 | Muses: Designing, Composing, Generating Nonexistent Fantasy 3D Creatures without Training | 📄 |
| 15 | Interp3D: Correspondence-aware Interpolation for Generative Textured 3D Morphing | 📄 |
| 16 | RelaxFlow: Text-Driven Amodal 3D Generation | 📄 |
| 17 | MV-SAM3D: Adaptive multi-view fusion for layout-aware 3D generation | 📄 |
| 18 | Points-to-3D: Structure-Aware 3D Generation with Point Cloud Priors | 📄 |
| 19 | Pixal3D: Pixel-Aligned 3D Generation from Images | 📄 |
| # | Paper | Link |
|---|---|---|
| 1 | SS4D: Native 4D Generative Model via Structured Spacetime Latents | 📄 |
Triplane-based 3D representations for generation.
| # | Paper | Link |
|---|---|---|
| 1 | Efficient Geometry-Aware 3D Generative Adversarial Networks | 📄 |
| 2 | GET3D: A Generative Model of High Quality 3D Textured Shapes Learned from Images | 📄 |
| 3 | Rodin: A Generative Model for Sculpting 3D Digital Avatars Using Diffusion | 📄 |
| 4 | 3DGen: Triplane Latent Diffusion for Textured Mesh Generation | 📄 |
| 5 | Triplane Meets Gaussian Splatting: Fast and Generalizable Single-View 3D Reconstruction with Transformers | 📄 |
| 6 | AGG: Amortized Generative 3D Gaussians for Single Image to 3D | 📄 |
| 7 | LN3Diff: Scalable Latent Neural Fields Diffusion for Speedy 3D Generation | 📄 |
| 8 | Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer | 📄 |
| 9 | DiffGS: Functional Gaussian Splatting Diffusion | 📄 |
Direct mesh generation via autoregressive or diffusion-based approaches.
| # | Paper | Link |
|---|---|---|
| 1 | MeshGPT: Generating Triangle Meshes with Decoder-Only Transformers | 📄 |
| 2 | MeshAnything: Artist-Created Mesh Generation with Autoregressive Transformers | 📄 |
| 3 | MeshAnything V2: Artist-Created Mesh Generation with Adjacent Mesh Tokenization | 📄 |
| 4 | EdgeRunner: Auto-Regressive Auto-Encoder for Artistic Mesh Generation | 📄 |
| 5 | Scaling Mesh Generation via Compressive Tokenization | 📄 |
| 6 | LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models | 📄 |
| 7 | TreeMeshGPT: Artistic Mesh Generation with Autoregressive Tree Sequencing | 📄 |
| 8 | DeepMesh: Auto-Regressive Artist-Mesh Creation with Reinforcement Learning | 📄 |
| 9 | MeshCraft: Exploring Efficient and Controllable Mesh Generation with Flow-Based DiTs | 📄 |
| 10 | Mesh-RFT: Enhancing Mesh Generation via Fine-Grained Reinforcement Fine-Tuning | 📄 |
| 11 | Topology-Preserved Auto-Regressive Mesh Generation in the Manner of Weaving Silk | 📄 |
| 12 | XSpecMesh: Quality-Preserving Auto-Regressive Mesh Generation Acceleration | 📄 |
| 13 | VertexRegen: Mesh Generation with Continuous Level of Detail | 📄 |
| 14 | FastMesh: Efficient Artistic Mesh Generation via Component Decoupling | 📄 |
| 15 | MeshMosaic: Scaling Artist Mesh Generation via Local-to-Global Assembly | 📄 |
| 16 | ARMesh: Autoregressive Mesh Generation via Next-Level-of-Detail Prediction | 📄 |
| 17 | FlashMesh: Faster and Better Autoregressive Mesh Synthesis via Structured Speculation | 📄 |
| 18 | PartDiffuser: Part-Wise 3D Mesh Generation via Discrete Diffusion | 📄 |
| 19 | TopGen: Learning structural layouts and cross-fields for quadrilateral mesh generation | 📄 |
| 20 | FACE: A face-based autoregressive representation for high-fidelity and efficient mesh generation | 📄 |
| 21 | Strips as Tokens: Artist Mesh Generation with Native UV Segmentation | 📄 |
Part-aware and compositional 3D generation.
| # | Paper | Link |
|---|---|---|
| 1 | Part123: Part-Aware 3D Reconstruction from a Single-View Image | 📄 |
| 2 | MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation | 📄 |
| 3 | PartGen: Part-Level 3D Generation and Reconstruction with Multi-View Diffusion Models | 📄 |
| 4 | PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers | 📄 |
| 5 | Efficient Part-Level 3D Object Generation via Dual Volume Packing | 📄 |
| 6 | Assembler: Scalable 3D Part Assembly via Anchor Point Diffusion | 📄 |
| 7 | OmniPart: Part-Aware 3D Generation with Semantic Decoupling and Structural Cohesion | 📄 |
| 8 | From One to More: Contextual Part Latents for 3D Generation | 📄 |
| 9 | AutoPartGen: Autoregressive 3D Part Generation and Discovery | 📄 |
| 10 | X-Part: High Fidelity and Structure Coherent Shape Decomposition | 📄 |
| 11 | FullPart: Generating Each 3D Part at Full Resolution | 📄 |
| 12 | UniPart: Part-Level 3D Generation with Unified 3D Geom-Seg Latents | 📄 |
| 13 | EI-Part: Explode for Completion and Implode for Refinement | 📄 |
| 14 | DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising | 📄 |
Articulated object generation, rigging, and animation.
| # | Paper | Link |
|---|---|---|
| 1 | URDFormer: A Pipeline for Constructing Articulated Simulation Environments from Real-World Images | 📄 |
| 2 | Puppet-Master: Scaling Interactive Video Generation as a Motion Prior for Part-Level Dynamics | 📄 |
| 3 | Articulate-Anything: Automatic Modeling of Articulated Objects via a Vision-Language Foundation Model | 📄 |
| 4 | SINGAPO: Single Image Controlled Generation of Articulated Parts in Objects | 📄 |
| 5 | ArtFormer: Controllable Generation of Diverse 3D Articulated Objects | 📄 |
| 6 | MeshArt: Generating Articulated Meshes with Structure-Guided Transformers | 📄 |
| 7 | RigAnything: Template-Free Autoregressive Rigging for Diverse 3D Assets | 📄 |
| 8 | MagicArticulate: Make Your 3D Models Articulation-Ready | 📄 |
| 9 | One Model to Rig Them All: Diverse Skeleton Rigging with UniRig | 📄 |
| 10 | Anymate: A Dataset and Baselines for Learning 3D Object Rigging | 📄 |
| 11 | DreamArt: Generating Interactable Articulated Objects from a Single Image | 📄 |
| 12 | Puppeteer: Rig and Animate Your 3D Models | 📄 |
| 13 | Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation | 📄 |
| 14 | FreeArt3D: Training-Free Articulated Object Generation Using 3D Diffusion | 📄 |
| 15 | PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image | 📄 |
| 16 | Particulate: Feed-Forward 3D Object Articulation | 📄 |
| 17 | ART: Articulated Reconstruction Transformer | 📄 |
| 18 | Choreographing a World of Dynamic Objects | 📄 |
| 19 | RigMo: Unifying Rig and Motion Learning for Generative Animation | 📄 |
| 20 | PALUM: Part-Based Attention Learning for Unified Motion Retargeting | 📄 |
| 21 | Motion 3-to-4: 3D Motion Reconstruction for 4D Synthesis | 📄 |
| 22 | ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusion | 📄 |
| 23 | Skin Tokens: A Learned Compact Representation for Unified Autoregressive Rigging | 📄 |
| 24 | PAct: Part-Decomposed Single-View Articulated Object Generation | 📄 |
| 25 | ArtLLM: Generating Articulated Assets via 3D LLM | 📄 |
| 26 | AniGen: Unified S3 Fields for Animatable 3D Asset Generation | 📄 |
| 27 | AnimateAnyMesh++: A Flexible 4D Foundation Model for High-Fidelity Text-Driven Mesh Animation | 📄 |
| 28 | PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual World | 📄 |
| 29 | Rigel3D: Rig-aware Latents for Animation-Ready 3D Asset Generation | 📄 |
| 30 | R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow | 📄 |
| 31 | Articraft: An Agentic System for Scalable Articulated 3D Asset Generation | 📄 |
If you find this repository useful, please consider giving it a ⭐
Made with ❤️ for the 3D Vision Community
15 commits
This is a collective repository for all 3D and 4D Object Generation papers
20
15 commits
updated May 22, 2026
A curated collection of cutting-edge research papers on 3D and 4D Content Generation.
If you find this repo useful, please consider giving it a ⭐!
| List | Description | |
|---|---|---|
| ✏️ | Awesome 3D/4D Editing | Papers on 3D and 4D scene/object editing |
| 🔄 | Awesome FeedForward 3D/4D Reconstruction | Papers on feed-forward 3D/4D reconstruction |
Score Distillation Sampling approaches for text/image-to-3D generation.
| # | Paper | Link |
|---|---|---|
| 1 | DreamFusion: Text-to-3D Using 2D Diffusion | 📄 |
| 2 | Latent-NeRF for Shape-Guided Generation of 3D Shapes and Textures | 📄 |
| 3 | Magic3D: High-Resolution Text-to-3D Content Creation | 📄 |
| 4 | NerfDiff: Single-Image View Synthesis with NeRF-Guided Distillation from 3D-Aware Diffusion | 📄 |
| 5 | DreamBooth3D: Subject-Driven Text-to-3D Generation | 📄 |
| 6 | Fantasia3D: Disentangling Geometry and Appearance for High-Quality Text-to-3D Content Creation | 📄 |
| 7 | Make-It-3D: High-Fidelity 3D Creation from a Single Image with Diffusion Prior | 📄 |
| 8 | TextMesh: Generation of Realistic 3D Meshes From Text Prompts | 📄 |
| 9 | IT3D: Improved Text-to-3D Generation with Explicit View Synthesis | 📄 |
| 10 | Text-to-3D Using Gaussian Splatting | 📄 |
| 11 | DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content Creation | 📄 |
| 12 | GaussianDreamer: Fast Generation from Text to 3D Gaussians by Bridging 2D and 3D Diffusion Models | 📄 |
| 13 | DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion Prior | 📄 |
| 14 | Learn to Optimize Denoising Scores for 3D Generation | 📄 |
| 15 | REPARO: Compositional 3D Assets Generation with Differentiable 3D Layout Alignment | 📄 |
| 16 | Recon3D: High Quality 3D Reconstruction from a Single Image Using Generated Back-View Explicit Priors | 📄 |
| 17 | COMOGen: A Controllable Text-to-3D Multi-Object Generation Framework | 📄 |
| 18 | Enhancing Single Image to 3D Generation Using Gaussian Splatting and Hybrid Diffusion Priors | 📄 |
| 19 | ModeDreamer: Mode Guiding Score Distillation for Text-to-3D Generation | 📄 |
| 20 | Enhanced 3D Generation by 2D Editing | 📄 |
| 21 | Diverse Score Distillation | 📄 |
| 22 | Chirpy3D: Continuous Part Latents for Creative 3D Bird Generation | 📄 |
| 23 | Probability-Flow Distillation: Exact Wasserstein Gradient Flow for High-Fidelity 3D Generation | 📄 |
SDS-based dynamic 4D content generation.
| # | Paper | Link |
|---|---|---|
| 1 | Text-to-4D Dynamic Scene Generation | 📄 |
| 2 | Consistent4D: Consistent 360° Dynamic Object Generation from Monocular Video | 📄 |
| 3 | Animate124: Animating One Image to 4D Dynamic Scene | 📄 |
| 4 | 4D-fy: Text-to-4D Generation Using Hybrid Score Distillation Sampling | 📄 |
| 5 | DreamGaussian4D: Generative 4D Gaussian Splatting | 📄 |
| 6 | SC4D: Sparse-Controlled Video-to-4D Generation and Motion Transfer | 📄 |
Multi-view diffusion models for 3D-consistent image generation and reconstruction.
| # | Paper | Link |
|---|---|---|
| 1 | Zero-1-to-3: Zero-Shot One Image to 3D Object | 📄 |
| 2 | One-2-3-45: Any Single Image to 3D Mesh in 45 Seconds Without Per-Shape Optimization | 📄 |
| 3 | MVDiffusion: Enabling Holistic Multi-View Image Generation with Correspondence-Aware Diffusion | 📄 |
| 4 | MVDream: Multi-View Diffusion for 3D Generation | 📄 |
| 5 | SyncDreamer: Generating Multiview-Consistent Images from a Single-View Image | 📄 |
| 6 | Wonder3D: Single Image to 3D Using Cross-Domain Diffusion | 📄 |
| 7 | Zero123++: A Single Image to Consistent Multi-View Diffusion Base Model | 📄 |
| 8 | DMV3D: Denoising Multi-View Diffusion Using 3D Large Reconstruction Model | 📄 |
| 9 | ViVid-1-to-3: Novel View Synthesis with Video Diffusion Models | 📄 |
| 10 | ImageDream: Image-Prompt Multi-View Diffusion for 3D Generation | 📄 |
| 11 | 4DGen: Grounded 4D Content Generation with Spatial-Temporal Consistency | 📄 |
| 12 | EscherNet: A Generative Model for Scalable View Synthesis | 📄 |
| 13 | LGM: Large Multi-View Gaussian Model for High-Resolution 3D Content Creation | 📄 |
| 14 | IM-3D: Iterative Multiview Diffusion and Reconstruction for High-Quality 3D Generation | 📄 |
| 15 | MVDiffusion++: A Dense High-Resolution Multi-View Diffusion Model for 3D Object Reconstruction | 📄 |
| 16 | V3D: Video Diffusion Models Are Effective 3D Generators | 📄 |
| 17 | Make-Your-3D: Fast and Consistent Subject-Driven 3D Content Generation | 📄 |
| 18 | SV3D: Novel Multi-View Synthesis and 3D Generation from a Single Image Using Latent Video Diffusion | 📄 |
| 19 | VFusion3D: Learning Scalable 3D Generative Models from Video Diffusion Models | 📄 |
| 20 | CAT3D: Create Anything in 3D with Multi-View Diffusion Models | 📄 |
| 21 | Era3D: High-Resolution Multiview Diffusion Using Efficient Row-Wise Attention | 📄 |
| 22 | Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models | 📄 |
| 23 | Enhancing Single Image to 3D Generation Using Gaussian Splatting and Hybrid Diffusion Priors | 📄 |
| 24 | DreamCraft3D++: Efficient Hierarchical 3D Generation with Multi-Plane Reconstruction Model | 📄 |
| 25 | MVPaint: Synchronized Multi-View Diffusion for Painting Anything 3D | 📄 |
| 26 | Edify 3D: Scalable High-Quality 3D Asset Generation | 📄 |
| 27 | Direct and Explicit 3D Generation from a Single Image | 📄 |
| 28 | Fancy123: One Image to High-Quality 3D Mesh Generation via Plug-and-Play Deformation | 📄 |
| 29 | ModeDreamer: Mode Guiding Score Distillation for Text-to-3D Generation | 📄 |
| 30 | RIGI: Rectifying Image-to-3D Generation Inconsistency via Uncertainty-Aware Learning | 📄 |
| 31 | Turbo3D: Ultra-Fast Text-to-3D Generation | 📄 |
| 32 | Gen-3Diffusion: Realistic Image-to-3D Generation via 2D & 3D Diffusion Synergy | 📄 |
| 33 | PartGen: Part-Level 3D Generation and Reconstruction with Multi-View Diffusion Models | 📄 |
Multi-view generation approaches for dynamic 4D content.
| # | Paper | Link |
|---|---|---|
| 1 | 4DGen: Grounded 4D Content Generation with Spatial-Temporal Consistency | 📄 |
| 2 | STAG4D: Spatial-Temporal Anchored Generative 4D Gaussians | 📄 |
| 3 | Diffusion4D: Fast Spatial-Temporal Consistent 4D Generation via Video Diffusion Models | 📄 |
| 4 | L4GM: Large 4D Gaussian Reconstruction Model | 📄 |
| 5 | SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency | 📄 |
Large Reconstruction Model (LRM) based feed-forward 3D generation.
| # | Paper | Link |
|---|---|---|
| 1 | LRM: Large Reconstruction Model for Single Image to 3D | 📄 |
| 2 | Instant3D: Fast Text-to-3D with Sparse-View Generation and Large Reconstruction Model | 📄 |
| 3 | DMV3D: Denoising Multi-View Diffusion Using 3D Large Reconstruction Model | 📄 |
| 4 | PF-LRM: Pose-Free Large Reconstruction Model for Joint Pose and Shape Prediction | 📄 |
| 5 | TripoSR: Fast 3D Object Reconstruction from a Single Image | 📄 |
| 6 | GRM: Large Gaussian Reconstruction Model for Efficient 3D Reconstruction and Generation | 📄 |
| 7 | InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-View Large Reconstruction Models | 📄 |
| 8 | M-LRM: Multi-View Large Reconstruction Model | 📄 |
| 9 | LRM-Zero: Training Large Reconstruction Models with Synthesized Data | 📄 |
| 10 | ControLRM: Fast and Controllable 3D Generation via Large Reconstruction Model | 📄 |
| 11 | Enhancing Single Image to 3D Generation Using Gaussian Splatting and Hybrid Diffusion Priors | 📄 |
| 12 | 3D-Adapter: Geometry-Consistent Multi-View Diffusion for High-Quality 3D Generation | 📄 |
| 13 | Fancy123: One Image to High-Quality 3D Mesh Generation via Plug-and-Play Deformation | 📄 |
Naïve 3D native generation methods (diffusion in 3D space, autoregressive, etc.).
| # | Paper | Link |
|---|---|---|
| 1 | DiffRF: Rendering-Guided 3D Radiance Field Diffusion | 📄 |
| 2 | Point-E: A System for Generating 3D Point Clouds from Complex Prompts | 📄 |
| 3 | Shap-E: Generating Conditional 3D Implicit Functions | 📄 |
| 4 | GSD: View-Guided Gaussian Splatting Diffusion for 3D Reconstruction | 📄 |
| 5 | 3DTopia-XL: Scaling High-Quality 3D Asset Generation via Primitive Diffusion | 📄 |
| 6 | SeMv-3D: Towards Semantic and Multi-View Consistency for General Text-to-3D Generation with Triplane Priors | 📄 |
| 7 | L3DG: Latent 3D Gaussian Diffusion | 📄 |
| 8 | LucidFusion: Generating 3D Gaussians with Arbitrary Unposed Images | 📄 |
| 9 | GaussianAnything: Interactive Point Cloud Latent Diffusion for 3D Generation | 📄 |
| 10 | SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-Scale 3D VQVAE | 📄 |
| 11 | A Lesson in Splats: Teacher-Guided Diffusion for 3D Gaussian Splats Generation with 2D Supervision | 📄 |
| 12 | TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction | 📄 |
| 13 | RELATE3D: REfocusing Latent Adapter for Targeted Local Enhancement and Editing in 3D Generation | 📄 |
| 14 | AssetFormer: Modular 3D Assets Generation with Autoregressive Transformer | 📄 |
| 15 | VAR-3D: View-Aware Auto-Regressive Model for Text-to-3D Generation via a 3D Tokenizer | 📄 |
| 16 | Fuse3D: Generating 3D Assets Controlled by Multi-Image Fusion | 📄 |
| 17 | GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation | 📄 |
| 18 | Omni123: Exploring 3D Native Foundation Models with Limited 3D Data | 📄 |
| 19 | Hitem3D 2.0: Multi-View Guided Native 3D Texture Generation | 📄 |
Implicit latent space representations for 3D shape generation.
| # | Paper | Link |
|---|---|---|
| 1 | 3DShape2VecSet: A 3D Shape Representation for Neural Fields and Generative Diffusion Models | 📄 |
| 2 | Michelangelo: Conditional 3D Shape Generation Based on Shape-Image-Text Aligned Latent Representation | 📄 |
| 3 | CraftsMan3D: High-Fidelity Mesh Generation with 3D Native Generation and Interactive Geometry Refiner | 📄 |
| 4 | CLAY: A Controllable Large-Scale Generative Model for Creating High-Quality 3D Assets | 📄 |
| 5 | Dora: Sampling and Benchmarking for 3D Shape Variational Auto-Encoders | 📄 |
| 6 | TripoSG: High-Fidelity 3D Shape Synthesis Using Large-Scale Rectified Flow Models | 📄 |
| 7 | Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets | 📄 |
| 8 | Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material | 📄 |
| 9 | Hunyuan3D 2.5: Towards High-Fidelity 3D Assets Generation with Ultimate Details | 📄 |
| 10 | Seed3D 1.0: From images to high-fidelity simulation-ready 3D assets | 📄 |
| 11 | LATTICE: Democratize high-fidelity 3D generation at scale | 📄 |
| 12 | UltraShape 1.0: High-Fidelity 3D Shape Generation via Scalable Geometric Refinement | 📄 |
| 13 | Seed3D 2.0: Advancing High-Fidelity Simulation-Ready 3D Content Generation | 📄 |
| 14 | Pose-Aware Diffusion for 3D Generation | 📄 |
| 15 | ROAR-3D: Routing arbitrary views for high-fidelity 3D generation | 📄 |
| # | Paper | Link |
|---|---|---|
| 1 | Sculpt4D: Generating 4D Shapes via Sparse-Attention Diffusion Transformers | 📄 |
Explicit latent space representations (Gaussians, structured latents, etc.) for generation.
| # | Paper | Link |
|---|---|---|
| 1 | GaussianCube: A Structured and Explicit Radiance Representation for 3D Generative Modeling | 📄 |
| 2 | Structured 3D Latents for Scalable and Versatile 3D Generation | 📄 |
| 3 | SynCity: Training-Free Generation of 3D Worlds | 📄 |
| 4 | Hi3DGen: High-Fidelity 3D Geometry Generation from Images via Normal Bridging | 📄 |
| 5 | DSO: Aligning 3D Generators with Simulation Feedback for Physical Soundness | 📄 |
| 6 | Sparc3D: Sparse Representation and Construction for High-Resolution 3D Shapes Modeling | 📄 |
| 7 | Ultra3D: Efficient and High-Fidelity 3D Generation with Part Attention | 📄 |
| 8 | Few-Step Flow for 3D Generation via Marginal-Data Transport Distillation | 📄 |
| 9 | ReconViaGen: Towards Accurate Multi-View 3D Object Reconstruction via Generation | 📄 |
| 10 | SAM 3D: 3Dfy Anything in Images | 📄 |
| 11 | Wukong's 72 transformations: High-fidelity textured 3D morphing via flow models | 📄 |
| 12 | Native and Compact Structured Latents for 3D Generation | 📄 |
| 13 | MorphAny3D: Unleashing the power of structured latent in 3D morphing | 📄 |
| 14 | Muses: Designing, Composing, Generating Nonexistent Fantasy 3D Creatures without Training | 📄 |
| 15 | Interp3D: Correspondence-aware Interpolation for Generative Textured 3D Morphing | 📄 |
| 16 | RelaxFlow: Text-Driven Amodal 3D Generation | 📄 |
| 17 | MV-SAM3D: Adaptive multi-view fusion for layout-aware 3D generation | 📄 |
| 18 | Points-to-3D: Structure-Aware 3D Generation with Point Cloud Priors | 📄 |
| 19 | Pixal3D: Pixel-Aligned 3D Generation from Images | 📄 |
| # | Paper | Link |
|---|---|---|
| 1 | SS4D: Native 4D Generative Model via Structured Spacetime Latents | 📄 |
Triplane-based 3D representations for generation.
| # | Paper | Link |
|---|---|---|
| 1 | Efficient Geometry-Aware 3D Generative Adversarial Networks | 📄 |
| 2 | GET3D: A Generative Model of High Quality 3D Textured Shapes Learned from Images | 📄 |
| 3 | Rodin: A Generative Model for Sculpting 3D Digital Avatars Using Diffusion | 📄 |
| 4 | 3DGen: Triplane Latent Diffusion for Textured Mesh Generation | 📄 |
| 5 | Triplane Meets Gaussian Splatting: Fast and Generalizable Single-View 3D Reconstruction with Transformers | 📄 |
| 6 | AGG: Amortized Generative 3D Gaussians for Single Image to 3D | 📄 |
| 7 | LN3Diff: Scalable Latent Neural Fields Diffusion for Speedy 3D Generation | 📄 |
| 8 | Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer | 📄 |
| 9 | DiffGS: Functional Gaussian Splatting Diffusion | 📄 |
Direct mesh generation via autoregressive or diffusion-based approaches.
| # | Paper | Link |
|---|---|---|
| 1 | MeshGPT: Generating Triangle Meshes with Decoder-Only Transformers | 📄 |
| 2 | MeshAnything: Artist-Created Mesh Generation with Autoregressive Transformers | 📄 |
| 3 | MeshAnything V2: Artist-Created Mesh Generation with Adjacent Mesh Tokenization | 📄 |
| 4 | EdgeRunner: Auto-Regressive Auto-Encoder for Artistic Mesh Generation | 📄 |
| 5 | Scaling Mesh Generation via Compressive Tokenization | 📄 |
| 6 | LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models | 📄 |
| 7 | TreeMeshGPT: Artistic Mesh Generation with Autoregressive Tree Sequencing | 📄 |
| 8 | DeepMesh: Auto-Regressive Artist-Mesh Creation with Reinforcement Learning | 📄 |
| 9 | MeshCraft: Exploring Efficient and Controllable Mesh Generation with Flow-Based DiTs | 📄 |
| 10 | Mesh-RFT: Enhancing Mesh Generation via Fine-Grained Reinforcement Fine-Tuning | 📄 |
| 11 | Topology-Preserved Auto-Regressive Mesh Generation in the Manner of Weaving Silk | 📄 |
| 12 | XSpecMesh: Quality-Preserving Auto-Regressive Mesh Generation Acceleration | 📄 |
| 13 | VertexRegen: Mesh Generation with Continuous Level of Detail | 📄 |
| 14 | FastMesh: Efficient Artistic Mesh Generation via Component Decoupling | 📄 |
| 15 | MeshMosaic: Scaling Artist Mesh Generation via Local-to-Global Assembly | 📄 |
| 16 | ARMesh: Autoregressive Mesh Generation via Next-Level-of-Detail Prediction | 📄 |
| 17 | FlashMesh: Faster and Better Autoregressive Mesh Synthesis via Structured Speculation | 📄 |
| 18 | PartDiffuser: Part-Wise 3D Mesh Generation via Discrete Diffusion | 📄 |
| 19 | TopGen: Learning structural layouts and cross-fields for quadrilateral mesh generation | 📄 |
| 20 | FACE: A face-based autoregressive representation for high-fidelity and efficient mesh generation | 📄 |
| 21 | Strips as Tokens: Artist Mesh Generation with Native UV Segmentation | 📄 |
Part-aware and compositional 3D generation.
| # | Paper | Link |
|---|---|---|
| 1 | Part123: Part-Aware 3D Reconstruction from a Single-View Image | 📄 |
| 2 | MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation | 📄 |
| 3 | PartGen: Part-Level 3D Generation and Reconstruction with Multi-View Diffusion Models | 📄 |
| 4 | PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers | 📄 |
| 5 | Efficient Part-Level 3D Object Generation via Dual Volume Packing | 📄 |
| 6 | Assembler: Scalable 3D Part Assembly via Anchor Point Diffusion | 📄 |
| 7 | OmniPart: Part-Aware 3D Generation with Semantic Decoupling and Structural Cohesion | 📄 |
| 8 | From One to More: Contextual Part Latents for 3D Generation | 📄 |
| 9 | AutoPartGen: Autoregressive 3D Part Generation and Discovery | 📄 |
| 10 | X-Part: High Fidelity and Structure Coherent Shape Decomposition | 📄 |
| 11 | FullPart: Generating Each 3D Part at Full Resolution | 📄 |
| 12 | UniPart: Part-Level 3D Generation with Unified 3D Geom-Seg Latents | 📄 |
| 13 | EI-Part: Explode for Completion and Implode for Refinement | 📄 |
| 14 | DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising | 📄 |
Articulated object generation, rigging, and animation.
| # | Paper | Link |
|---|---|---|
| 1 | URDFormer: A Pipeline for Constructing Articulated Simulation Environments from Real-World Images | 📄 |
| 2 | Puppet-Master: Scaling Interactive Video Generation as a Motion Prior for Part-Level Dynamics | 📄 |
| 3 | Articulate-Anything: Automatic Modeling of Articulated Objects via a Vision-Language Foundation Model | 📄 |
| 4 | SINGAPO: Single Image Controlled Generation of Articulated Parts in Objects | 📄 |
| 5 | ArtFormer: Controllable Generation of Diverse 3D Articulated Objects | 📄 |
| 6 | MeshArt: Generating Articulated Meshes with Structure-Guided Transformers | 📄 |
| 7 | RigAnything: Template-Free Autoregressive Rigging for Diverse 3D Assets | 📄 |
| 8 | MagicArticulate: Make Your 3D Models Articulation-Ready | 📄 |
| 9 | One Model to Rig Them All: Diverse Skeleton Rigging with UniRig | 📄 |
| 10 | Anymate: A Dataset and Baselines for Learning 3D Object Rigging | 📄 |
| 11 | DreamArt: Generating Interactable Articulated Objects from a Single Image | 📄 |
| 12 | Puppeteer: Rig and Animate Your 3D Models | 📄 |
| 13 | Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation | 📄 |
| 14 | FreeArt3D: Training-Free Articulated Object Generation Using 3D Diffusion | 📄 |
| 15 | PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image | 📄 |
| 16 | Particulate: Feed-Forward 3D Object Articulation | 📄 |
| 17 | ART: Articulated Reconstruction Transformer | 📄 |
| 18 | Choreographing a World of Dynamic Objects | 📄 |
| 19 | RigMo: Unifying Rig and Motion Learning for Generative Animation | 📄 |
| 20 | PALUM: Part-Based Attention Learning for Unified Motion Retargeting | 📄 |
| 21 | Motion 3-to-4: 3D Motion Reconstruction for 4D Synthesis | 📄 |
| 22 | ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusion | 📄 |
| 23 | Skin Tokens: A Learned Compact Representation for Unified Autoregressive Rigging | 📄 |
| 24 | PAct: Part-Decomposed Single-View Articulated Object Generation | 📄 |
| 25 | ArtLLM: Generating Articulated Assets via 3D LLM | 📄 |
| 26 | AniGen: Unified S3 Fields for Animatable 3D Asset Generation | 📄 |
| 27 | AnimateAnyMesh++: A Flexible 4D Foundation Model for High-Fidelity Text-Driven Mesh Animation | 📄 |
| 28 | PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual World | 📄 |
| 29 | Rigel3D: Rig-aware Latents for Animation-Ready 3D Asset Generation | 📄 |
| 30 | R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow | 📄 |
| 31 | Articraft: An Agentic System for Scalable Articulated 3D Asset Generation | 📄 |
If you find this repository useful, please consider giving it a ⭐
Made with ❤️ for the 3D Vision Community
15 commits