visionPublished: October 2, 2026

Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis

By Keerthi Kaashyap, Dennis Anthony, Akshay Krishnan, Nhi Ngoc Nguyen, Jeremy Collins, James Hays, Shreyas Kousik, Animesh Garg

Research TL;DR

"SNAP improves geometric representation learning by using a pose-conditioned local decoder and latent-space reconstruction, preventing decoder expressivity from suppressing encoder features."

Abstract

This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene structure, thereby enabling transferable multi-view geometric representations. Yet, existing encoder-based NVS methods yield poor representations. This is not because of a lack of supervisory signal, but rather due to inconspicuous architectural choices: \textit{spatially expressive decoders} that dilute representational capabilities of the scene encoder, and \textit{low-level pixel-space targets} that hinder feature learning. We present SNAP, a self-supervised encoder-decoder transformer that addresses both through a pose-conditioned local decoder and a latent-space reconstruction objective. SNAP is task agnostic, and we show that it is competitive with special-purpose geometry-supervised methods. SNAP also performs competitively against self-supervised representations across five tasks: visual localization, pose estimation, point correspondence, depth estimation, and robot manipulation. Remarkably, SNAP's patch features exhibit emergent viewpoint invariance that approaches heavily supervised models despite lower compute and data budgets. Under camera shifts where standard 2D representations collapse, SNAP degrades more gracefully, revealing that restricting decoder expressivity actively prevents the suppression of transferable geometric structure. https://snap-nvs.github.io

Technical Analysis & Implementation

Core Problem & Motivation§

Novel View Synthesis (NVS) should learn transferable 3D geometric representations via multi-view supervision. However, existing encoder-based NVS methods produce poor representations. The authors identify two culprits: spatially expressive decoders that let the decoder "cheat" by reconstructing fine details without forcing the encoder to capture geometry, and low-level pixel-space targets that waste capacity on photometric details rather than geometry. They propose SNAP, a self-supervised encoder-decoder transformer that fixes both issues.

Key Contributions§

  1. Pose-conditioned local decoder: The decoder processes only a local patch around the query pixel, conditioned on the target camera pose. This restricts decoder expressivity, forcing the encoder to store global geometric structure.
  2. Latent-space reconstruction objective: Instead of predicting RGB pixels, SNAP reconstructs features from a frozen pre-trained encoder (e.g., DINOv2). This focuses learning on semantics and geometry, not low-level appearance.
  3. Emergent viewpoint invariance: Despite lower compute and data budgets, SNAP's patch features approach the viewpoint invariance of heavily supervised models, and degrade more gracefully under camera shifts.

Methodology§

Given a source image $I_s$ and its camera pose $P_s$, an encoder $E_\theta$ extracts a feature map $F = E_\theta(I_s) \in \mathbb{R}^{H'\times W'\times C}$. For a target view with pose $P_t$, a query pixel $p_t$ is projected into the source view via a pose-conditioned transformation. A local decoder $D_\phi$ takes the projected feature patch and the relative pose $\Delta P = P_t P_s^{-1}$, and outputs a latent feature $\hat{z}$ for that pixel. The training loss is a latent-space reconstruction loss:

$$\mathcal{L} = \left\| \hat{z} - \text{sg}(E_{\text{teacher}}(I_t)[p_t]) \right\|_2^2$$

where $E_{\text{teacher}}$ is a frozen pre-trained encoder (stop-gradient sg prevents collapse), and $I_t$ is the target view image. The decoder is deliberately local: it only sees a small patch ($k\times k$ features), preventing it from aggregating global information and thus forcing the encoder to be geometrically expressive.

Implementation Details§

  • Encoder: A Vision Transformer (ViT) with patch size 16, yielding patch features.
  • Decoder: A lightweight transformer that attends to a local neighborhood of features (e.g., $3\times3$ patch) and is conditioned on the relative camera pose via AdaLN or cross-attention.
  • Teacher: Frozen DINOv2 (or similar) provides target features; no labels are used.
  • Training: Uses standard NVS datasets (e.g., CO3D, RealEstate10K) with random source-target view pairs.
  • Tasks: Evaluated on visual localization, pose estimation, point correspondence, depth estimation, and robot manipulation.

Code Snippet (PyTorch)§

import torch
import torch.nn as nn
import torch.nn.functional as F

class LocalDecoder(nn.Module):
    def __init__(self, dim, patch_size=3):
        super().__init__()
        self.patch_size = patch_size
        self.pose_mlp = nn.Sequential(
            nn.Linear(12, dim), nn.ReLU(), nn.Linear(dim, dim)
        )
        self.attn = nn.MultiheadAttention(dim, 4, batch_first=True)
        self.norm = nn.LayerNorm(dim)

    def forward(self, local_feats, rel_pose):
        # local_feats: (B, k*k, C), rel_pose: (B, 12)
        B, N, C = local_feats.shape
        pose_emb = self.pose_mlp(rel_pose).unsqueeze(1)  # (B,1,C)
        tokens = local_feats + pose_emb
        attn_out, _ = self.attn(tokens, tokens, tokens)
        out = self.norm(attn_out + tokens)
        return out.mean(dim=1)  # (B, C)

class SNAP(nn.Module):
    def __init__(self, encoder, teacher, dim=768):
        super().__init__()
        self.encoder = encoder
        self.teacher = teacher  # frozen
        self.decoder = LocalDecoder(dim)

    def forward(self, src_img, tgt_img, rel_pose):
        # Encode source
        feats_s = self.encoder(src_img)  # (B, H', W', C)
        # Extract local patch around projected pixel (simplified)
        # ... use grid_sample with projected coordinates
        local_patch = F.unfold(feats_s, self.decoder.patch_size)
        # Decode
        pred = self.decoder(local_patch, rel_pose)
        # Target from frozen teacher
        with torch.no_grad():
            feats_t = self.teacher(tgt_img)
            # sample target feature at query pixel
            target = feats_t  # (placeholder)
        loss = F.mse_loss(pred, target)
        return loss

Results & Implications§

SNAP matches or exceeds special-purpose geometry-supervised methods and self-supervised representations across five tasks. Its features show viewpoint invariance approaching heavily supervised models, with lower compute. The paper argues that restricting decoder expressivity is a general principle for learning transferable representations from NVS, challenging the trend of ever-more-powerful decoders.

SHARE RESEARCH: