OMa: Dense Object Matching for Dense Reconstruction

Visual Recognition Group, Czech Technical University in Prague
OMa teaser: selected keyframes on the left, reconstructed object point cloud with object-relative camera poses on the right.

Reconstructing an object from a video capture. Left: keyframes chosen online by a matchability criterion. Right: the reconstructed object point cloud with camera poses relative to the object (up to scale). The background is replaced with solid black at every stage of the pipeline; it is kept here for illustration only. Example: HOPEv2 obj_000001.

Abstract

We introduce OMa (Object Matching), an object-centric reconstruction pipeline based on dense matching and correspondence-based Structure-from-Motion. OMa fine-tunes a dense matcher on synthetic render pairs whose background pixels are zeroed, so the mask localizes the object for the matcher, and reconstructs using only correspondences within the propagated masks. Its pose accuracy degrades only moderately after background removal. We benchmark our method against the state-of-the-art feed-forward static-scene reconstructors VGGT, MASt3R, MapAnything, and π3. In the foreground-only scenario, the strongest competing feed-forward method collapses on HANDAL (37.86° / 15.93 cm) and loses accuracy on HOPEv2 (6.68° / 1.89 cm), and remains competitive only on NAVI at 2.82° / 1.50 cm, while being competitive on the background-retained sequences. Because OMa uses foreground correspondences alone, the foreground-only regime extends to objects that move independently of the scene. There, OMa attains the highest F-score on HANDAL and, with the background visible, on HOPEv2.

How it works

OMa takes an RGB video of a single rigid object with known camera intrinsics and a segmentation mask of the object in the first frame. It outputs the object's 3D point cloud and the camera poses relative to the object (up to a similarity transform), without depth, a 3D template, category priors, or object-specific training.

  1. Mask propagation. The first-frame mask is propagated through the video with the off-the-shelf tracker SAM2; foreground refers to the pixels inside these masks.
  2. Foreground correspondence. A dense matcher (UFM) estimates dense flow and per-pixel certainty between frames. We keep only confident foreground-to-foreground matches. The matcher is fine-tuned on masked synthetic render pairs (background zeroed), so it stays accurate on texture-poor object interiors without background cues.
  3. Causal keyframe selection. A matchability criterion selects keyframes and relocalizes frames online, during capture, without future observations or exhaustive all-pairs matching, and builds the viewgraph for reconstruction.
  4. Reconstruction. The surviving foreground correspondences are chained into multi-view tracks and passed to robust, incremental correspondence-based Structure-from-Motion, which recovers the object point cloud and object-relative camera poses.

Contributions

  • Template-free object-centric reconstruction. Recover a previously unseen rigid object's point cloud and object-relative camera poses from RGB video and a first-frame mask, with no depth, 3D template, category priors, or object-specific training.
  • Foreground-only reconstruction. Dense-matching-based SfM stays accurate where the evaluated feed-forward reconstructors deteriorate after background removal; because it relies on foreground correspondences alone, it extends to objects that move independently of the scene (for example, a hand-held object).
  • Causal frame selection. A matchability-based criterion selects keyframes and relocalizes online, an alternative to fixed every-k subsampling and to offline all-pairs selection.

Results: static capture

One representative object per dataset. Columns compare OMa against feed-forward reconstructors, with the textured ground-truth mesh at the right. The baselines reconstruct from every 8th video frame, OMa from its selected keyframes, all with the background masked to black. Point clouds are pose-aligned to the ground-truth mesh.

OMa (ours)VGGTMASt3R MapAnythingπ3GT mesh
HANDAL
static
OMa reconstruction, HANDAL static VGGT reconstruction, HANDAL static MASt3R reconstruction, HANDAL static MapAnything reconstruction, HANDAL static Pi3 reconstruction, HANDAL static Ground-truth mesh, HANDAL static
F@5 0.270.030.01 0.020.09
HOPEv2
static
OMa reconstruction, HOPEv2 static VGGT reconstruction, HOPEv2 static MASt3R reconstruction, HOPEv2 static MapAnything reconstruction, HOPEv2 static Pi3 reconstruction, HOPEv2 static Ground-truth mesh, HOPEv2 static
F@5 0.790.300.28 0.270.50
NAVI
static
OMa reconstruction, NAVI static VGGT reconstruction, NAVI static MASt3R reconstruction, NAVI static MapAnything reconstruction, NAVI static Pi3 reconstruction, NAVI static Ground-truth mesh, NAVI static
F@5 0.700.480.25 0.300.70

F@5 is the mean fraction of reconstructed points within 5 mm of the ground-truth mesh, aggregated per dataset over the masked-input arm. OMa recovers the object shape where the feed-forward baselines fragment or collapse.

Results: dynamic, hand-held capture

Here the object moves independently of the camera and the background, and a hand occludes it. For each object, the top strip shows frames of the capture and the grid below shows the reconstructions. OMa uses its adaptive keyframes, the baselines every k-th frame, giving 20 to 25 keyframes.

HANDAL obj_000022 (drill)

Input frame 1, HANDAL dynamic Input frame 2, HANDAL dynamic Input frame 3, HANDAL dynamic Input frame 4, HANDAL dynamic Input frame 5, HANDAL dynamic
Input video frames (uniformly sampled)
OMa reconstruction, HANDAL dynamic VGGT reconstruction, HANDAL dynamic MASt3R reconstruction, HANDAL dynamic
OMa (ours)VGGTMASt3R
MapAnything reconstruction, HANDAL dynamic Pi3 reconstruction, HANDAL dynamic Ground-truth mesh, HANDAL dynamic
MapAnythingπ3GT mesh

HOPEv2 obj_000006 (cookies box)

Input frame 1, HOPEv2 dynamic Input frame 2, HOPEv2 dynamic Input frame 3, HOPEv2 dynamic Input frame 4, HOPEv2 dynamic Input frame 5, HOPEv2 dynamic
Input video frames (uniformly sampled)
OMa reconstruction, HOPEv2 dynamic VGGT reconstruction, HOPEv2 dynamic MASt3R reconstruction, HOPEv2 dynamic
OMa (ours)VGGTMASt3R
MapAnything reconstruction, HOPEv2 dynamic Pi3 reconstruction, HOPEv2 dynamic Ground-truth mesh, HOPEv2 dynamic
MapAnythingπ3GT mesh

The dynamic setting remains hard: reconstruction succeeds for some objects, but robustness is limited for all methods, including ours.

BibTeX

@misc{jelinek2026oma,
  title  = {OMa: Dense Object Matching for Dense Reconstruction},
  author = {Jel\'inek, Tom\'a\v{s} and Mishkin, Dmytro and Matas, Ji\v{r}\'i},
  year   = {2026},
  note   = {Preprint},
  url    = {https://tjelinek.github.io/oma/}
}