TAPVid-MV: Tracking Any Point in 3D Across Multiple Views

The first benchmark for long-term 3D point tracking across several synchronized, moving cameras.

Skanda Koppula*, Frano Rajič*, Abdullah Faiz Ur Rahman, Yi Yang, Ignacio Rocco, Jeet Thakwani, Rishabh Kabra, Andrew Zisserman, Joao Carreira, Siyu Tang, Carl Doersch, Gabriel Brostow

Overview

The seven TAPVid-MV subsets — DROID, EgoExo4D, Harmony4D, PACE, Hi4D, Waymo, and Perpetua — each shown as three camera views of one sequence with its camera-motion glyph and size.
The seven subsets. Each column shows three views of one sequence with its ground-truth tracks, and below them the camera trajectories of that sequence (white dots are static cameras, blue paths moving ones), how many of its cameras move, how far they travel, and the size of the subset.

Inspect any sequence in 3D

Every one of the 284 sequences is released as a Rerun recording — multi-view RGB, colored point cloud, ground-truth 3D tracks, and camera frusta. Pick one below, or browse the full index.

Choose a sequence and press Load the viewer streams the recording from the release bucket

BibTeX

@article{tapvidmv2026,
  title   = {TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views},
  author  = {Skanda Koppula and Frano Raji{\v{c}} and Abdullah Faiz Ur Rahman and Yi Yang and Ignacio Rocco and Jeet Thakwani and Rishabh Kabra and Andrew Zisserman and Joao Carreira and Siyu Tang and Carl Doersch and Gabriel Brostow},
  year    = {2026}
}

Citing the source datasets

TAPVid-MV is built on six existing datasets, whose authors made this benchmark possible. If you use a subset, please cite the dataset it comes from alongside this work: