Benchmarking Human and DNN Biases in Monocular Depth Estimation

Talk Presentation 41.23: Monday, May 18, 2026, 8:15 – 9:45 am, Talk Room 2
Session: 3D Shape and Space Perception

Yuki Kubota1, Taiki Fukiage1; 1Communication Science Laboratories, NTT, Inc.

Human depth perception from a single image is systematically biased, yet the characteristics of these distortions remain insufficiently understood. Meanwhile, modern monocular depth estimation (MDE) models achieve high physical accuracy, raising a central question: to what extent do such models reproduce—or diverge from—human perceptual biases? To address this, we constructed two human-annotated depth datasets using established benchmarks: NYU (indoor scenes) and KITTI (outdoor scenes). These datasets enabled direct comparisons between human observers and 69 deep neural networks (DNNs), spanning diverse architectures, training strategies, datasets, and output formats. Human data were obtained by asking participants to report the absolute distances to four simultaneously marked target points in each photograph. Model accuracy was quantified using scale-invariant RMSE. Human–model similarity was defined as a partial correlation between model and human error patterns, obtained by repeatedly correlating fixed model errors with split-half averages of observer errors while controlling for ground-truth depth. We further applied an affine decomposition that isolates per-image affine distortions (scale, shift, horizontal shear, and vertical shear) from residual error. Across both datasets, humans showed robust and systematic deviations from physical ground truth, as indicated by high split-half human–human partial correlations of error patterns (NYU: 0.808; KITTI: 0.671). Examining the relationship between accuracy and human similarity revealed a clear pattern: similarity increased with accuracy up to approximately human-level performance, but declined for models surpassing that range—indicating a distinct accuracy–similarity trade-off. Notably, this trade-off was substantially more pronounced in the KITTI dataset. Overall, our findings demonstrate that human-like behavior in MDE does not emerge simply by improving metric accuracy. Instead, the divergence suggests that the strategies used by state-of-the-art DNNs, which potentially rely on dataset-specific cues to maximize precision, fundamentally differ from the generalized perceptual heuristics employed by human observers.