POINCAR3
Emergent Multi-View Geometry
Through Self-Distillation
Poincar3 inputs a sequence of images from the same scene without labels and is trained using a multi-view self-distillation objective. Interstingly, it obtains a strong understanding of multiple-view geometry, illustrated by zero-shot correspondence estimation abilities and rapid finetuning for feedforward-reconstruction.

Henri Poincaré, 1854–1912
Motivated by Henri Poincaré
At a glance
Main results
Poincar3 outperforms DINOv3, MuM and Muskie on camera pose, point clouds, correspondence and SE(3) structure alike. Bars are drawn so that longer is always better.
Pose estimation · RE10K
AUC@30° ↑Pose estimation · ScanNet++
AUC@30° ↑Point cloud error · ETH3D
mm ↓Normal consistency · DTU
cos θ ↑Matching · ScanNet
PCK@25 ↑SE(3) decodability
R² > 0 ↑94.9PCK@50
Zero-shot multi-view correspondence on ScanNet, straight from the attention map
650M params
ViT-L encoder plus a 12-layer alternating-attention multi-view decoder
03D labels
Trained from scratch on unlabeled internet videos. No poses, no depth, no correspondences
3days · 8×H200
400k steps at 256×256, sequences of 2 to 24 views
Abstract
Poincar3 in short
Over a century ago, Henri Poincaré argued that a motionless observer cannot acquire the notion of space. Yet most visual representation learning methods operate on individual images, while those that leverage multiple views rely on RGB reconstruction, entangling geometry with appearance. We propose Poincar3, a self-supervised method that learns representations from multiple views through self-distillation instead of RGB reconstruction. We combine masked patch and image-level distillation with a teacher that observes additional views, enabling training from scratch without explicit 3D supervision. Poincar3 outperforms both previous single and multi-view self-supervised approaches such as DINOv3, MuM and Muskie on correspondence estimation, camera pose estimation and 3D reconstruction. Using a lightweight Poincaré adapter, we also find that our learned features encode camera motion more accurately than existing self-supervised representations.
Emergence
Emergent matching capabilities without supervision.
The model never sees a correspondence label. Yet pick a query patch and follow the highest attention activation across frames, and you get beautiful tracks.

demo.py in the repository.Method
Multi-view self-distillation
A student and an EMA teacher both run a multi-view transformer over a sequence from the same scene. The student gets M masked, photometrically augmented frames; the teacher gets those M frames unmasked plus T extra views. The two are aligned with a patch loss on masked tokens and an image-level loss on the per-frame [CLS] tokens. No pixels are ever reconstructed.
An image-level objective
Masked-patch distillation alone plateaus. Adding a per-frame [CLS] loss is what steers self-distillation towards a 3D-aware representation rather than a semantic one.
A teacher that sees more
The teacher receives T extra frames of the same scene, sampled uniformly from 0 to 12. They never enter the loss — they only leak geometric context through multi-view attention, so the student has to infer what it cannot see.
Whole images, no crops
DINOv2's local/global crop scheme destroys the very geometry the multi-view attention is trying to exploit. Keeping full frames turns a failing objective into a working one.

Results · correspondence
Zero-shot multi-view matching
Patch tracking across eight views, with no finetuning. The attention map turns out to be an even stronger correspondence estimator than the features themselves. It even outperforms feed-forward reconstruction models that were trained with 3D supervision.


Results · reconstruction
Feed-forward 3D reconstruction
Relative pose by AUC over 10 random frames, and point clouds by median accuracy in mm and normal consistency. Three protocols of increasing cost: heads only, a transformer on top, and a full finetune.
| Method | RE10K | ScanNet++ | MegaDepth | ETH3D | DTU | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| @3° | @30° | @3° | @30° | @3° | @30° | Acc ↓ | NC ↑ | Acc ↓ | NC ↑ | |
| Train only heads on top of the frozen backbone | ||||||||||
| DINOv3 | 0.0 | 18.9 | 0.0 | 9.4 | 1.5 | 55.5 | 1.18 | 0.59 | 10.76 | 0.54 |
| Muskie | 1.1 | 36.3 | 0.0 | 19.0 | 0.1 | 46.9 | 1.12 | 0.60 | 11.43 | 0.55 |
| MuM | 0.0 | 29.8 | 0.0 | 18.7 | 0.1 | 48.5 | 1.12 | 0.60 | 12.60 | 0.54 |
| Poincar3 | 2.4 | 51.9 | 0.1 | 48.7 | 3.7 | 67.5 | 0.82 | 0.65 | 10.60 | 0.59 |
| Train a transformer on top of the frozen backbone | ||||||||||
| DINOv3 | 1.0 | 29.1 | 0.0 | 22.0 | 1.7 | 65.6 | 0.95 | 0.66 | 10.63 | 0.58 |
| Muskie | 1.8 | 41.5 | 0.0 | 28.1 | 1.7 | 58.9 | 0.88 | 0.67 | 11.95 | 0.58 |
| MuM | 1.4 | 40.1 | 0.0 | 30.2 | 1.9 | 60.8 | 1.00 | 0.60 | 12.26 | 0.56 |
| Poincar3 | 5.2 | 62.2 | 1.3 | 52.1 | 5.2 | 69.5 | 0.75 | 0.68 | 10.52 | 0.61 |
| Finetune the backbone | ||||||||||
| Random init. | 0.0 | 17.4 | 0.0 | 11.3 | 0.0 | 35.9 | 1.24 | 0.60 | 12.34 | 0.56 |
| DINOv3 init. | 1.3 | 31.2 | 0.0 | 30.9 | 6.9 | 73.8 | 0.98 | 0.68 | 10.87 | 0.59 |
| Poincar3 | 8.3 | 68.2 | 4.5 | 67.0 | 7.4 | 74.6 | 0.81 | 0.80 | 6.73 | 0.63 |
MuM and Muskie are excluded from the finetuning block because their architectures differ substantially, making a direct comparison less meaningful. Each run: 4×H200 for three days.


Results · SE(3)
The Poincaré adapter
Do the features encode the geometry of rigid motion, and not just its appearance? We fit a small MLP per scene and ask it to make feature displacements linear in the se(3) twist between two frames.
ΔPt,s≈W(φ(Ht+s)−φ(Ht))
The adapter φ tries to unroll the non-linear feature space into a homogeneous coordinate system where a change of pose is just a translation. How well it can is a direct measure of how much SE(3) structure the representation already holds.
Every model improves once the feature is allowed to see the scene move. The gap between the two markers is exactly what a motionless observer cannot acquire (following Poincaré's argument). DINOv3 has no multi-view mode.
Ablations
How each ingredient earns its place
Same architecture, same data, same compute budget throughout. Row I is the RGB reconstruction objective used by MuM and Muskie; row II is single-view DINOv2. Naively lifting DINOv2 to multiple views (III) is worse than reconstruction — it takes all three ingredients to overtake it.
Rows I–VII follow the ablation in the paper; hover labels are hidden on small screens.
Data nobody else can use
Because nothing in the objective needs 3D annotations, training can draw on internet video that reconstruction models simply cannot touch. Adding it on top of the 3D-labelled datasets lifts PCK@50 from 87.5 → 94.9 on ScanNet and 78.5 → 86.5 on NAVI.
Cite
BibTeX
@article{nordstrom2026poincare3,
title={Emergent Multi-View Geometry Through Self-Distillation},
author={David Nordström and Thibaut Loiseau and Vincent Lepetit
and Michael Felsberg and Guillaume Bourmaud and Fredrik Kahl},
journal={arXiv preprint},
year={2026}
}