Sogomonyan, A. (2026). Multi-View Robustness and 3D Aggregation for Object Detection Models [Diploma Thesis, Technische Universität Wien]. reposiTUm. https://doi.org/10.34726/hss.2026.130312
E193 - Institut für Visual Computing and Human-Centered Technology
-
Date (published):
2026
-
Number of Pages:
75
-
Keywords:
Object Detection; Multi-View Detection; Computer Vision; Indoor Scene Understanding; Deep Learning; YOLO; Faster R-CNN; DETR; ScanNet++; Multi-View Consistency
en
Abstract:
Modern object detection models such as YOLO, Faster R-CNN, and DETR achieve strong results on 2D images, but it remains an open question how reliably they detect the same physical object when it is observed from different viewpoints. We hypothesize that off-the-shelf 2D detectors are weak at multi-view object detection in indoor scenes.This thesis tests that hypothesis by benchmarking four COCO-trained detector families (Faster R-CNN, YOLOv8, DETR, RF-DETR) across 10 ScanNet++ indoor scenes comprising 7,694 frames and 279 COCO-valid physical object instances. Alongside classical mean Average Precision (mAP), we introduce the Multi-View Consistency Score (MVCS), defined as the fraction of visible frames in which the same physical instance is detected at IoU >= tau, as a per-instance measure of cross-view reliability.Using strict COCO-valid filtering, the benchmark shows that even the best detector (Faster R-CNN) achieves only 41% MVCS@0.50, meaning it detects objects in fewer than half of their visible frames. RF-DETR leads at stricter localization thresholds (MVCS@0.75 = 28.1%), indicating tighter bounding-box agreement when detections occur, and also achieves the highest full-scope mAP (0.4533 at IoU 0.50).Per-instance correlation between AP and MVCS is strong (Pearson 0.77-0.87) but not perfect, confirming that the two metrics capture different aspects of detection quality. Per-scene and per-class analyses reveal that scene composition, in particular the share of large and visually distinctive objects, is the strongest predictor of consistency. A supplementary 3D projection workflow is used to inspect spatial detection patterns and failure modes.
en
Additional information:
Arbeit an der Bibliothek noch nicht eingelangt - Daten nicht geprüft