<div class="csl-bib-body">
<div class="csl-entry">Tekaya, N., Waldner, M., & Zeppelzauer, M. (2025). A Matter of Time: Revealing the Structure of Time in Vision-Language Models. In L. Rossetto, D.-T. Dang-Nguyen, P. Chen, S. Rudinac, W.-H. Cheng, & J. Benois-Pineau (Eds.), <i>MM ’25 : Proceedings of the 33rd ACM International Conference on Multimedia</i> (pp. 12371–12380). Association for Computing Machinery. https://doi.org/10.1145/3746027.3758163</div>
</div>
-
dc.identifier.uri
http://hdl.handle.net/20.500.12708/221607
-
dc.description.abstract
Large-scale vision-language models (VLMs) such as CLIP have gained popularity for their generalizable and expressive multimodal representations. By leveraging large-scale training data with diverse textual metadata, VLMs acquire open-vocabulary capabilities, solving tasks beyond their training scope. This paper investigates the temporal awareness of VLMs, assessing their ability to position visual content in time. We introduce TIME10k, a benchmark dataset of over 10,000 images with temporal ground truth, and evaluate the time-awareness of 37 VLMs by a novel methodology. Our investigation reveals that temporal information is structured along a low-dimensional, non-linear manifold in the VLM embedding space. Based on this insight, we propose methods to derive an explicit ''timeline'' representation from the embedding space. These representations model time and its chronological progression and thereby facilitate temporal reasoning tasks. Our timeline approaches achieve competitive to superior accuracy compared to a prompt-based baseline while being computationally efficient. All code and data are available at https://tekayanidham.github.io/timeline-page/.
en
dc.description.sponsorship
FWF - Österr. Wissenschaftsfonds
-
dc.language.iso
en
-
dc.relation.ispartofseries
ACM Conference Proceedings
-
dc.rights.uri
http://creativecommons.org/licenses/by/4.0/
-
dc.subject
Multimodal representations
en
dc.subject
Vision-language models
en
dc.subject
Time modeling
en
dc.subject
Time reasoning
en
dc.subject
Time estimation
en
dc.subject
Benchmark dataset
en
dc.title
A Matter of Time: Revealing the Structure of Time in Vision-Language Models
en
dc.type
Inproceedings
en
dc.type
Konferenzbeitrag
de
dc.rights.license
Creative Commons Namensnennung 4.0 International
de
dc.rights.license
Creative Commons Attribution 4.0 International
en
dc.contributor.affiliation
University of Applied Sciences St Pölten, Austria
-
dc.contributor.affiliation
University of Applied Sciences St Pölten, Austria
-
dc.contributor.editoraffiliation
Dublin City University, Ireland
-
dc.contributor.editoraffiliation
University of Bergen, Norway
-
dc.contributor.editoraffiliation
La Trobe University, Australia
-
dc.contributor.editoraffiliation
University of Amsterdam, Netherlands (the)
-
dc.contributor.editoraffiliation
National Taiwan University, Taiwan (Province of China)