Abstract
We introduce OMa (Object Matching), an object-centric reconstruction pipeline based on dense matching and correspondence-based Structure-from-Motion. OMa fine-tunes a dense matcher on synthetic render pairs whose background pixels are zeroed, so the mask localizes the object for the matcher, and reconstructs using only correspondences within the propagated masks. Its pose accuracy degrades only moderately after background removal. We benchmark our method against the state-of-the-art feed-forward static-scene reconstructors VGGT, MASt3R, MapAnything, and π3. In the foreground-only scenario, the strongest competing feed-forward method collapses on HANDAL (37.86° / 15.93 cm) and loses accuracy on HOPEv2 (6.68° / 1.89 cm), and remains competitive only on NAVI at 2.82° / 1.50 cm, while being competitive on the background-retained sequences. Because OMa uses foreground correspondences alone, the foreground-only regime extends to objects that move independently of the scene. There, OMa attains the highest F-score on HANDAL and, with the background visible, on HOPEv2.
How it works
OMa takes an RGB video of a single rigid object with known camera intrinsics and a segmentation mask of the object in the first frame. It outputs the object's 3D point cloud and the camera poses relative to the object (up to a similarity transform), without depth, a 3D template, category priors, or object-specific training.
- Mask propagation. The first-frame mask is propagated through the video with the off-the-shelf tracker SAM2; foreground refers to the pixels inside these masks.
- Foreground correspondence. A dense matcher (UFM) estimates dense flow and per-pixel certainty between frames. We keep only confident foreground-to-foreground matches. The matcher is fine-tuned on masked synthetic render pairs (background zeroed), so it stays accurate on texture-poor object interiors without background cues.
- Causal keyframe selection. A matchability criterion selects keyframes and relocalizes frames online, during capture, without future observations or exhaustive all-pairs matching, and builds the viewgraph for reconstruction.
- Reconstruction. The surviving foreground correspondences are chained into multi-view tracks and passed to robust, incremental correspondence-based Structure-from-Motion, which recovers the object point cloud and object-relative camera poses.
Contributions
- Template-free object-centric reconstruction. Recover a previously unseen rigid object's point cloud and object-relative camera poses from RGB video and a first-frame mask, with no depth, 3D template, category priors, or object-specific training.
- Foreground-only reconstruction. Dense-matching-based SfM stays accurate where the evaluated feed-forward reconstructors deteriorate after background removal; because it relies on foreground correspondences alone, it extends to objects that move independently of the scene (for example, a hand-held object).
- Causal frame selection. A matchability-based criterion selects keyframes and relocalizes online, an alternative to fixed every-k subsampling and to offline all-pairs selection.
Results: static capture
One representative object per dataset. Columns compare OMa against feed-forward reconstructors, with the textured ground-truth mesh at the right. The baselines reconstruct from every 8th video frame, OMa from its selected keyframes, all with the background masked to black. Point clouds are pose-aligned to the ground-truth mesh.
| OMa (ours) | VGGT | MASt3R | MapAnything | π3 | GT mesh | |
|---|---|---|---|---|---|---|
| HANDAL static |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
| F@5 0.27 | 0.03 | 0.01 | 0.02 | 0.09 | ||
| HOPEv2 static |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
| F@5 0.79 | 0.30 | 0.28 | 0.27 | 0.50 | ||
| NAVI static |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
| F@5 0.70 | 0.48 | 0.25 | 0.30 | 0.70 |
F@5 is the mean fraction of reconstructed points within 5 mm of the ground-truth mesh, aggregated per dataset over the masked-input arm. OMa recovers the object shape where the feed-forward baselines fragment or collapse.
Results: dynamic, hand-held capture
Here the object moves independently of the camera and the background, and a hand occludes it. For each object, the top strip shows frames of the capture and the grid below shows the reconstructions. OMa uses its adaptive keyframes, the baselines every k-th frame, giving 20 to 25 keyframes.
HANDAL obj_000022 (drill)
![]() |
![]() |
![]() |
![]() |
![]() |
| Input video frames (uniformly sampled) | ||||
![]() |
![]() |
![]() |
| OMa (ours) | VGGT | MASt3R |
![]() |
![]() |
![]() |
| MapAnything | π3 | GT mesh |
HOPEv2 obj_000006 (cookies box)
![]() |
![]() |
![]() |
![]() |
![]() |
| Input video frames (uniformly sampled) | ||||
![]() |
![]() |
![]() |
| OMa (ours) | VGGT | MASt3R |
![]() |
![]() |
![]() |
| MapAnything | π3 | GT mesh |
The dynamic setting remains hard: reconstruction succeeds for some objects, but robustness is limited for all methods, including ours.
BibTeX
@misc{jelinek2026oma,
title = {OMa: Dense Object Matching for Dense Reconstruction},
author = {Jel\'inek, Tom\'a\v{s} and Mishkin, Dmytro and Matas, Ji\v{r}\'i},
year = {2026},
note = {Preprint},
url = {https://tjelinek.github.io/oma/}
}







































