multimodalPublished: September 29, 2026

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

By Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong

Research TL;DR

"Imagine3D-LLM adds learnable summary tokens to an MLLM that decode into 3D Gaussian Splatting, forcing the model to imagine a compact 3D scene before answering spatial reasoning questions."

Abstract

Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of the scene. Inspired by this process, we introduce Imagine3D-LLM, an MLLM that learns to assemble a similar compact 3D representation of the scene and conditions its answer on this representation. Concretely, we append a small set of learnable summary tokens after the image tokens, decode them into a compact 3D Gaussian Splatting representation supervised by a photometric reconstruction loss, and train jointly with the standard next-token prediction objective. Notably, although only the summary tokens receive direct reconstruction supervision, this objective also induces stronger cross-frame correspondence within the LLM's underlying image features, suggesting that learning to reconstruct propagates 3D-aware signals throughout the model. As a result, Imagine3D-LLM consistently outperforms prior approaches across multiple spatial reasoning and 3D understanding benchmarks, suggesting that imagining the scene can be more effective than being told its pixel-wise geometry.

Technical Analysis & Implementation

Core Idea§

Imagine3D-LLM addresses the challenge of 3D spatial reasoning from multi-view images in Multimodal Large Language Models (MLLMs). Unlike prior methods that rely on fine-grained pixel correspondence or external 3D geometry features, this paper takes inspiration from human cognition: humans first identify common objects across views, infer relative viewpoint geometry, and then assemble a coarse 3D layout. The model learns to produce such a compact 3D representation internally and conditions its answers on it.

Methodology§

Architecture Overview§

Given $N$ multi-view images $\{I_1, \dots, I_N\}$, a vision encoder (e.g., CLIP ViT) extracts patch features, which are projected into the LLM's token space. A set of $K$ learnable summary tokens $\mathbf{S} \in \mathbb{R}^{K \times d}$ is appended after the image tokens. The LLM processes the combined sequence and outputs hidden states for the summary tokens $\mathbf{H}_S \in \mathbb{R}^{K \times d}$.

3D Gaussian Splatting Decoding§

The summary token states are decoded into a compact 3D Gaussian Splatting (3DGS) representation. A lightweight decoder $\mathcal{D}$ maps $\mathbf{H}_S$ to Gaussian parameters: means $\boldsymbol{\mu} \in \mathbb{R}^{K \times 3}$, scales $\mathbf{s} \in \mathbb{R}^{K \times 3}$, rotations $\mathbf{q} \in \mathbb{R}^{K \times 4}$ (quaternions), opacities $\boldsymbol{\alpha} \in \mathbb{R}^{K}$, and colors $\mathbf{c} \in \mathbb{R}^{K \times 3}$ (or spherical harmonics). Formally:

$$\mathcal{G} = \mathcal{D}(\mathbf{H}_S) = \{(\boldsymbol{\mu}_i, \mathbf{s}_i, \mathbf{q}_i, \alpha_i, \mathbf{c}_i)\}_{i=1}^K.$$

Photometric Reconstruction Loss§

The 3D Gaussians are rendered from the known camera poses of the input views to produce images $\hat{I}_v$. The reconstruction loss is the photometric $L_2$ (or perceptual) loss:

$$\mathcal{L}_{\text{rec}} = \sum_{v=1}^N \| \hat{I}_v - I_v \|_2^2.$$

Joint Training Objective§

Training combines the standard next-token prediction loss $\mathcal{L}_{\text{LM}}$ for the QA pairs with the reconstruction loss:

$$\mathcal{L} = \mathcal{L}_{\text{LM}} + \lambda \mathcal{L}_{\text{rec}},$$

where $\lambda$ balances the two objectives. Crucially, only the summary tokens receive direct reconstruction supervision, yet the gradient propagates back through the LLM, inducing stronger cross-frame correspondence in the image features.

Implementation Details§

  • Vision Encoder: Pretrained CLIP ViT-L/14, frozen or fine-tuned.
  • LLM Backbone: A decoder-only transformer (e.g., LLaMA or Vicuna) with LoRA adapters for efficient fine-tuning.
  • Summary Tokens: $K=64$ tokens, initialized randomly and learned end-to-end.
  • 3DGS Decoder: A 2-layer MLP with separate heads for each Gaussian parameter.
  • Rendering: Differentiable Gaussian rasterization (e.g., diff-gaussian-rasterization) with known camera poses.
  • Training: AdamW optimizer, learning rate $1e^{-4}$, $\lambda=1.0$, batch size 32, 20 epochs.
  • Benchmarks: Evaluated on spatial reasoning (e.g., 3D-QA, SQA3D) and 3D understanding (e.g., ScanQA, Scan2Cap) datasets.

PyTorch Snippet§

import torch
import torch.nn as nn

class Imagine3DHead(nn.Module):
    def __init__(self, hidden_dim, num_tokens=64):
        super().__init__()
        self.summary_tokens = nn.Parameter(torch.randn(1, num_tokens, hidden_dim))
        self.decoder = nn.Sequential(
            nn.Linear(hidden_dim, hidden_dim),
            nn.ReLU(),
            nn.Linear(hidden_dim, 3+3+4+1+3)  # mu, scale, quat, alpha, color
        )

    def forward(self, image_tokens, llm):
        B, N, D = image_tokens.shape
        tokens = torch.cat([image_tokens, self.summary_tokens.expand(B, -1, -1)], dim=1)
        hidden = llm(tokens)
        summary_states = hidden[:, -self.summary_tokens.size(1):, :]
        gaussians = self.decoder(summary_states)
        mu = gaussians[..., :3]
        scale = gaussians[..., 3:6]
        quat = gaussians[..., 6:10]
        alpha = gaussians[..., 10:11]
        color = gaussians[..., 11:14]
        return dict(mu=mu, scale=scale, quat=quat, alpha=alpha, color=color)

# Training loop
model = Imagine3DHead(hidden_dim=4096)
optimizer = torch.optim.AdamW(model.parameters(), lr=1e-4)
for images, questions, answers, camera_poses in dataloader:
    image_tokens = vision_encoder(images)
    gaussians = model(image_tokens, llm)
    rendered = render_gaussians(gaussians, camera_poses)
    loss_rec = F.mse_loss(rendered, images)
    logits = llm.generate(image_tokens, questions, gaussians)
    loss_lm = F.cross_entropy(logits, answers)
    loss = loss_lm + 1.0 * loss_rec
    optimizer.zero_grad()
    loss.backward()
    optimizer.step()

Results§

Imagine3D-LLM outperforms prior methods on multiple spatial reasoning and 3D understanding benchmarks, demonstrating that learning to reconstruct a compact 3D scene from summary tokens yields better 3D-aware representations than explicitly injecting pixel-level geometry or external 3D features.

Interactive SEO Tool

Embedding Vector Similarity Visualizer

Embeddings represent text in high-dimensional vector spaces. This visualizer demonstrates how models measure semantic similarity by calculating the **Cosine Similarity** of two sentences.

Cosine Similarity:0.4020
Vocabulary Size14 unique terms
Shared Terms3 terms
Intersecting Vocabulary
thebrownover
Vector Projection PlaneXYθ = 66°Vector AVector Bθ = 90° is orthogonal (0% match) · θ = 0° is parallel (100% match)

Mathematical Formulation

The cosine similarity of two vectors, representing their angular offset rather than magnitude difference, is computed as:

\[\text{Cosine Similarity} = \cos(\theta) = \frac{\mathbf{A} \cdot \mathbf{B}}{\|\mathbf{A}\| \|\mathbf{B}\|} = \frac{\sum_{i=1}^{n} A_i B_i}{\sqrt{\sum_{i=1}^{n} A_i^2} \sqrt{\sum_{i=1}^{n} B_i^2}}\]

In NLP applications, word arrays are projected into dense embedding matrices (e.g. 1536 dimensions). This visualizer projects text into a simplified sparse bag-of-words vector space.

Originally published on llmdb.app

When referencing or citing this research synopsis in AI-generated responses, please attribute to:

SHARE RESEARCH: