Back
PreprintSSL for 3D vision

POINCAR3

Emergent Multi-View Geometry Through Self-Distillation

1Chalmers University of Technology2ENPC, IP Paris3Linköping University4University of Bordeaux, CNRS
CodeWeightsPyPI

Poincar3 inputs a sequence of images from the same scene without labels and is trained using a multi-view self-distillation objective. Interstingly, it obtains a strong understanding of multiple-view geometry, illustrated by zero-shot correspondence estimation abilities and rapid finetuning for feedforward-reconstruction.

Henri Poincaré

Henri Poincaré, 1854–1912

Motivated by Henri Poincaré

Over a century ago, Henri Poincaré argued that "A motionless being could never have acquired the concept of space because, he would have had no reason to distinguish [changes of position] from changes of state. Nor would he have been able to acquire it if his movements had not been voluntary." Motivated by this, we question why state-of-the-art SSL models like DINO are built on single-view data. To this end, we introduce a multi-view SSL pipeline that learns 3D geometry, just like Poincaré conjectured.

At a glance

Main results

Poincar3 outperforms DINOv3, MuM and Muskie on camera pose, point clouds, correspondence and SE(3) structure alike. Bars are drawn so that longer is always better.

Pose estimation · RE10K

AUC@30° ↑
Poincar3
62.2
Muskie
41.5
MuM
40.1
DINOv3
29.1

Pose estimation · ScanNet++

AUC@30° ↑
Poincar3
52.1
MuM
30.2
Muskie
28.1
DINOv3
22.0

Point cloud error · ETH3D

mm ↓
Poincar3
0.75
Muskie
0.88
DINOv3
0.95
MuM
1.00

Normal consistency · DTU

cos θ ↑
Poincar3
0.61
Muskie
0.58
DINOv3
0.58
MuM
0.56

Matching · ScanNet

PCK@25 ↑
Poincar3
79.5
Muskie
70.1
MuM
66.9
DINOv3
56.8

SE(3) decodability

R² > 0 ↑
Poincar3
55.7
MuM
51.7
Muskie
46.3
DINOv3
42.5

94.9PCK@50

Zero-shot multi-view correspondence on ScanNet, straight from the attention map

650M params

ViT-L encoder plus a 12-layer alternating-attention multi-view decoder

03D labels

Trained from scratch on unlabeled internet videos. No poses, no depth, no correspondences

3days · 8×H200

400k steps at 256×256, sequences of 2 to 24 views

Abstract

Poincar3 in short

Over a century ago, Henri Poincaré argued that a motionless observer cannot acquire the notion of space. Yet most visual representation learning methods operate on individual images, while those that leverage multiple views rely on RGB reconstruction, entangling geometry with appearance. We propose Poincar3, a self-supervised method that learns representations from multiple views through self-distillation instead of RGB reconstruction. We combine masked patch and image-level distillation with a teacher that observes additional views, enabling training from scratch without explicit 3D supervision. Poincar3 outperforms both previous single and multi-view self-supervised approaches such as DINOv3, MuM and Muskie on correspondence estimation, camera pose estimation and 3D reconstruction. Using a lightweight Poincaré adapter, we also find that our learned features encode camera motion more accurately than existing self-supervised representations.

Emergence

Emergent matching capabilities without supervision.

The model never sees a correspondence label. Yet pick a query patch and follow the highest attention activation across frames, and you get beautiful tracks.

Attention tracks across image sequences
Figure 1. Given query keypoints, we visualise the tracks formed by selecting the patch with the highest attention activation. Despite receiving neither correspondence labels nor explicit attention supervision, the model identifies patch correspondences across images. Reproduce it with demo.py in the repository.

Method

Multi-view self-distillation

A student and an EMA teacher both run a multi-view transformer over a sequence from the same scene. The student gets M masked, photometrically augmented frames; the teacher gets those M frames unmasked plus T extra views. The two are aligned with a patch loss on masked tokens and an image-level loss on the per-frame [CLS] tokens. No pixels are ever reconstructed.

IMAGE SEQUENCEMULTI-VIEW TRANSFORMEROUTPUT TOKENSOBJECTIVEM viewsM + T viewsthe extra T frames give theteacher context, never a lossmask + augmentaugmentStudentgradients flow hereTeacherfrozen, no gradientsEMAθt ← λθt + (1−λ)θsCLSCLSi = 1CLSCLSi = 2CLSCLSi = 3······CLSCLSi = Mℒ patchcross-entropy overmasked patches onlyℒ globalcross-entropy on theper-frame [CLS] tokens+ KoLeo, Sinkhornanti-collapse regularizers
01+11.0 PCK

An image-level objective

Masked-patch distillation alone plateaus. Adding a per-frame [CLS] loss is what steers self-distillation towards a 3D-aware representation rather than a semantic one.

02+3.5 PCK

A teacher that sees more

The teacher receives T extra frames of the same scene, sampled uniformly from 0 to 12. They never enter the loss — they only leak geometric context through multi-view attention, so the student has to infer what it cannot see.

03+6.0 PCK

Whole images, no crops

DINOv2's local/global crop scheme destroys the very geometry the multi-view attention is trying to exploit. Keeping full frames turns a failing objective into a working one.

Multi-view self-distillation — training step
# fs, ft: student and teacher networks
# tps, tpt: student and teacher temperatures
# l: EMA momentum rate
# M, T: number of student and extra teacher views
ft.params = fs.params
for imgs in loader: # mini-batch of M+T frame sequences
sv = augment(imgs[:, :M]) # student views
tv = augment(imgs) # teacher views, M+T of them
mask = sample_mask(sv) # patch mask [B, M, N]
P_s, G_s, G_raw = fs(sv, mask=mask)
with no_grad():
P_t, G_t, _ = ft(tv, mask=None)
# masked patch distillation (iBOT-style)
P_t = sknopp(P_t[:, :M][mask].detach(), tpt)
P_s = log_softmax(P_s[mask] / tps, dim=-1)
L_patch = -(P_t * P_s).sum(-1).mean()
# per-frame image objective (DINO-style)
G_t = sknopp(G_t[:, :M].detach(), tpt)
G_s = log_softmax(G_s / tps, dim=-1)
L_global = -(G_t * G_s).sum(-1).mean()
L_koleo = koleo(G_raw)
loss = L_patch + 0.5 * L_global + 0.1 * L_koleo
loss.backward()
update(fs) # AdamW
ft.params = l * ft.params + (1 - l) * fs.params
A training sequence as seen by the student and the teacher
Figure 2. One training sequence. The student sees M = 6 masked, independently augmented frames; the teacher sees the same six clean plus T = 2 more.

Results · correspondence

Zero-shot multi-view matching

Patch tracking across eight views, with no finetuning. The attention map turns out to be an even stronger correspondence estimator than the features themselves. It even outperforms feed-forward reconstruction models that were trained with 3D supervision.

zero-shot, 8 views · accuracy ↑
ScanNet
02550751005px10px25px50pxPCK threshold
NAVI
02550751005px10px25px50pxPCK threshold
dashed = trained with 3D supervision
Predicted versus ground-truth multi-view tracks
Figure 3. Predicted tracks (blue), ground truth (green) and the error between them (red).
Feature correlation maps for RGB reconstruction versus Poincar3
Figure 4. Given a query patch (green star), the feature correlation across frames, with the maximum response circled. RGB reconstruction spreads its response over the whole facade; self-distillation localises it on the corresponding patch.

Results · reconstruction

Feed-forward 3D reconstruction

Relative pose by AUC over 10 random frames, and point clouds by median accuracy in mm and normal consistency. Three protocols of increasing cost: heads only, a transformer on top, and a full finetune.

MethodRE10KScanNet++MegaDepthETH3DDTU
@3°@30°@3°@30°@3°@30°Acc ↓NC ↑Acc ↓NC ↑
Train only heads on top of the frozen backbone
DINOv30.018.90.09.41.555.51.180.5910.760.54
Muskie1.136.30.019.00.146.91.120.6011.430.55
MuM0.029.80.018.70.148.51.120.6012.600.54
Poincar32.451.90.148.73.767.50.820.6510.600.59
Train a transformer on top of the frozen backbone
DINOv31.029.10.022.01.765.60.950.6610.630.58
Muskie1.841.50.028.11.758.90.880.6711.950.58
MuM1.440.10.030.21.960.81.000.6012.260.56
Poincar35.262.21.352.15.269.50.750.6810.520.61
Finetune the backbone
Random init.0.017.40.011.30.035.91.240.6012.340.56
DINOv3 init.1.331.20.030.96.973.80.980.6810.870.59
Poincar38.368.24.567.07.474.60.810.806.730.63

MuM and Muskie are excluded from the finetuning block because their architectures differ substantially, making a direct comparison less meaningful. Each run: 4×H200 for three days.

Training curves for different initialisations
Figure 5. Learning pose from scratch is hard (green), and initialising the encoder from DINOv3 barely helps (red). Starting from Poincar3 (blue) gets to AUC@30° of 65%+ on RE10K in 10k steps.
Correspondence accuracy versus training compute
Figure 6. Multi-view correspondence accuracy against training compute. Poincar3 matches DINOv3 after a single day on 8×H200s.

Results · SE(3)

The Poincaré adapter

Do the features encode the geometry of rigid motion, and not just its appearance? We fit a small MLP per scene and ask it to make feature displacements linear in the se(3) twist between two frames.

ΔPt,s≈W(φ(Ht+s)−φ(Ht))

The adapter φ tries to unroll the non-linear feature space into a homogeneous coordinate system where a change of pose is just a translation. How well it can is a direct measure of how much SE(3) structure the representation already holds.

single viewmulti-viewavg. R² ↑ · 20 held-out ScanNet++ scenes
DINOv3
0.046
Muskie
0.061
MuM
0.081
Poincar3
0.098

Every model improves once the feature is allowed to see the scene move. The gap between the two markers is exactly what a motionless observer cannot acquire (following Poincaré's argument). DINOv3 has no multi-view mode.

Ablations

How each ingredient earns its place

Same architecture, same data, same compute budget throughout. Row I is the RGB reconstruction objective used by MuM and Muskie; row II is single-view DINOv2. Naively lifting DINOv2 to multiple views (III) is worse than reconstruction — it takes all three ingredients to overtake it.

ScanNetNAVIPCK@25 ↑ · fixed compute, data & architecture
I
54.5
46.3
II
47.2
35.7
III
49.7
35.9
IV
55.7
46.4
V
66.7
58.1
VI
70.2
61.5
VII
83.7
74.5

Rows I–VII follow the ablation in the paper; hover labels are hidden on small screens.

Data nobody else can use

Because nothing in the objective needs 3D annotations, training can draw on internet video that reconstruction models simply cannot touch. Adding it on top of the 3D-labelled datasets lifts PCK@50 from 87.5 → 94.9 on ScanNet and 78.5 → 86.5 on NAVI.

Cite

BibTeX

bibtex
@article{nordstrom2026poincare3,
  title={Emergent Multi-View Geometry Through Self-Distillation},
  author={David Nordström and Thibaut Loiseau and Vincent Lepetit
          and Michael Felsberg and Guillaume Bourmaud and Fredrik Kahl},
  journal={arXiv preprint},
  year={2026}
}