Burges, M., Zambanini, S., & Sablatnig, R. (2025). Interactive Object Detection for Tiny Objects in Large Remotely Sensed Images. In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) (pp. 4704–4713). IEEE. https://doi.org/10.1109/WACV61041.2025.00461
2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV 2025)
en
Event date:
26-Feb-2025 - 6-Mar-2025
-
Event place:
Tucson, AZ, United States of America (the)
-
Number of Pages:
10
-
Publisher:
IEEE
-
Peer reviewed:
Yes
-
Keywords:
interactive; object detection; remote sensing
en
Abstract:
This paper highlights the potential of a Human-In-the-Loop (HIL) in interactive object detection methods. Although automation in computer vision is advancing rapidly, certain critical tasks, such as detecting UneXploded Ordnance (UXO), space/marine debris, or the generation of new datasets, require 100% recall and near-perfect precision. These tasks are often performed manually since automatic methods do not achieve the necessary accuracy. However, interactive object detection frameworks can potentially enhance annotation speed while maintaining the recall and accuracy of manual annotation. We propose IRTDETR, an interactive and real-time object detection method for very large imagery to address this. Using either point or bounding box annotations provided by a HIL, it globally relates the full image with the annotator inputs via a cross-attention-like mechanism, employs an attention loss to maximize the classification score based on similarity, and reuses portions of the network outputs during iterative refinements to conserve resources. We conduct experiments on five different datasets (Tiny-DOTA, CHAI, AITOD, SarDET, and COCO) to verify the efficacy of our approach. Our method surpasses existing interactive annotation approaches, achieving a higher mean Average Precision (mAP) with the same number of clicks. Additionally, we validate the annotation efficiency of our method in a user study, demonstrating it is 2.46× quicker and asks for only 72% of the task load (NASA-TLX) compared to fully manual annotation.
en
Project title:
Domain-adaptive Remote sensing Image Analysis with Human-in-the-loop: 880883 (FFG - Österr. Forschungsförderungs- gesellschaft mbH)
-
Research Areas:
Visual Computing and Human-Centered Technology: 100%