<div class="csl-bib-body">
<div class="csl-entry">Sogomonyan, A. (2026). <i>Multi-View Robustness and 3D Aggregation for Object Detection Models</i> [Diploma Thesis, Technische Universität Wien]. reposiTUm. https://doi.org/10.34726/hss.2026.130312</div>
</div>
-
dc.identifier.uri
https://doi.org/10.34726/hss.2026.130312
-
dc.identifier.uri
http://hdl.handle.net/20.500.12708/229105
-
dc.description
Arbeit an der Bibliothek noch nicht eingelangt - Daten nicht geprüft
-
dc.description.abstract
Modern object detection models such as YOLO, Faster R-CNN, and DETR achieve strong results on 2D images, but it remains an open question how reliably they detect the same physical object when it is observed from different viewpoints. We hypothesize that off-the-shelf 2D detectors are weak at multi-view object detection in indoor scenes.This thesis tests that hypothesis by benchmarking four COCO-trained detector families (Faster R-CNN, YOLOv8, DETR, RF-DETR) across 10 ScanNet++ indoor scenes comprising 7,694 frames and 279 COCO-valid physical object instances. Alongside classical mean Average Precision (mAP), we introduce the Multi-View Consistency Score (MVCS), defined as the fraction of visible frames in which the same physical instance is detected at IoU >= tau, as a per-instance measure of cross-view reliability.Using strict COCO-valid filtering, the benchmark shows that even the best detector (Faster R-CNN) achieves only 41% MVCS@0.50, meaning it detects objects in fewer than half of their visible frames. RF-DETR leads at stricter localization thresholds (MVCS@0.75 = 28.1%), indicating tighter bounding-box agreement when detections occur, and also achieves the highest full-scope mAP (0.4533 at IoU 0.50).Per-instance correlation between AP and MVCS is strong (Pearson 0.77-0.87) but not perfect, confirming that the two metrics capture different aspects of detection quality. Per-scene and per-class analyses reveal that scene composition, in particular the share of large and visually distinctive objects, is the strongest predictor of consistency. A supplementary 3D projection workflow is used to inspect spatial detection patterns and failure modes.
en
dc.language
English
-
dc.language.iso
en
-
dc.rights.uri
http://rightsstatements.org/vocab/InC/1.0/
-
dc.subject
Object Detection
en
dc.subject
Multi-View Detection
en
dc.subject
Computer Vision
en
dc.subject
Indoor Scene Understanding
en
dc.subject
Deep Learning
en
dc.subject
YOLO
en
dc.subject
Faster R-CNN
en
dc.subject
DETR
en
dc.subject
ScanNet++
en
dc.subject
Multi-View Consistency
en
dc.title
Multi-View Robustness and 3D Aggregation for Object Detection Models
en
dc.type
Thesis
en
dc.type
Hochschulschrift
de
dc.rights.license
In Copyright
en
dc.rights.license
Urheberrechtsschutz
de
dc.identifier.doi
10.34726/hss.2026.130312
-
dc.contributor.affiliation
TU Wien, Österreich
-
dc.rights.holder
Artur Sogomonyan
-
dc.publisher.place
Wien
-
tuw.version
vor
-
tuw.thesisinformation
Technische Universität Wien
-
tuw.publication.orgunit
E193 - Institut für Visual Computing and Human-Centered Technology
-
dc.type.qualificationlevel
Diploma
-
dc.identifier.libraryid
AC17910588
-
dc.description.numberOfPages
75
-
dc.thesistype
Diplomarbeit
de
dc.thesistype
Diploma Thesis
en
dc.rights.identifier
In Copyright
en
dc.rights.identifier
Urheberrechtsschutz
de
tuw.advisor.staffStatus
exstaff
-
item.mimetype
application/pdf
-
item.cerifentitytype
Publications
-
item.grantfulltext
open
-
item.fulltext
with Fulltext
-
item.openairetype
master thesis
-
item.openairecristype
http://purl.org/coar/resource_type/c_bdcc
-
item.languageiso639-1
en
-
item.openaccessfulltext
Open Access
-
crisitem.author.dept
E307-04 - Forschungsbereich Maschinenbauinformatik und Virtuelle Produktentwicklung
-
crisitem.author.parentorg
E307 - Institut für Konstruktionswissenschaften und Produktentwicklung