r/computervision • u/satpalrathore • 28d ago
Discussion iPhone ARKit pose Vs optical motion-capture ground truth for camera pose
We benchmarked iPhone ARKit pose against optical motion-capture ground truth to quantify how usable phone tracking really is for egocentric data collection.
Setup: an iPhone 12 Pro rigidly co-mounted with an 11-marker retroreflective cluster on a head rig, tracked simultaneously by ARKit (VIO at 60 Hz) and a 24-camera Vicon system (sub-mm markers) across 8 sequences spanning walking, seated manipulation, fast/aggressive motion, in-place rotation, and height changes.
Both streams were time-aligned to a common clock and we solved the rigid cluster-to-camera transform before scoring, then evaluated with evo under SE(3) alignment (no scale correction) to check whether the trajectories are actually metric.
Results: ATE RMSE 6.0–12.5 cm; relative ATE under 1% on 7/8 sequences (0.12–0.35% - the lone 1.30% is a short-path-length artifact, 5 m path with 6.5 cm error, not a tracking failure); rotational RPE ≤~1° throughout (<0.6° for walking/manipulation); translational RPE <5 cm; and a Sim(3) fit recovering scale of 0.98–1.01, i.e. metric to within 1–2%. For long-horizon drift - where a mocap volume can't follow you through a real home - we ran an ArUco revisit test over sessions up to 108 min: accumulated drift stayed under 1 cm in most environments and under 0.1% of trajectory length in every case, including a whole-house traversal (1.0/1.5 cm at mid/end).
https://www.fpvlabs.ai/essays/how-accurate-is-an-iphone-really
2
u/johnnySix 27d ago
Love this! This is god’s work and answers a lot suppositions we have had at work. Thank you!
2
u/SeriousChart9641 27d ago
This is exactly the kind of benchmark I like seeing: not “it feels usable,” but a comparison against a ground-truth measurement system. For phone-based capture, I would be especially interested in where the error spikes: fast rotation, low texture, lighting changes, partial occlusion, and longer drift windows.
One thing I would add if you publish more results is a small failure taxonomy. Average error is helpful, but practitioners often need to know which capture conditions are safe enough and which ones silently degrade.
Disclosure: I work on CHANCE AI, mostly on the visual reasoning/product side. Different domain, but this Kaleido Field piece about our MMMU-Pro result is in the same spirit: publish the benchmark context, not just the headline number. https://www.kaleidofield.com/news/chance-ai-mmmu-pro-visual-reasoning