<div class="csl-bib-body">
<div class="csl-entry">Shashaani, S., & Knees, P. (2025). Towards Playlist Continuation Through Large-Scale Context and Audio-Based Music Representations. In A. Ferraro, L. Porcaro, & C. Bauer (Eds.), <i>Proceedings of the 3rd Music Recommender Systems Workshop (MuRS 2025) co-located with the 19th ACM Conference on Recommender Systems (RecSys 2025)</i>. http://hdl.handle.net/20.500.12708/223057</div>
</div>
-
dc.identifier.uri
http://hdl.handle.net/20.500.12708/223057
-
dc.description.abstract
Music recommendation research faces several challenges when modeling the complex relationships between users, items, and the circumstances under which they interact. In spite of access to commercial catalogs and large customer bases, academic research builds its findings on publicly shared datasets. However, these often only contain selected data modalities, limited catalogs, or temporally restricted snapshots of interaction data. Moreover, they might eventually vanish due to licensing issues. A strategy to overcome some of these limitations could consist in learning multimodal representation learning for playlist continuation. For instance, while metadata and interaction data can be used to learn item representations, content-based data can be used to predict representations for tracks where audio is available but interaction data is lacking.
To address this specific case, in this work, as a first pointer into the overall direction, we explore the integration of deep audio features extracted directly from MP3 files for music playlist completion. We first generate track embeddings using a Convolutional Neural Network (CNN) trained on a subset of MP3 files with the Spotify Million Playlist Dataset (MPSD), using pre-learned Word2Vec embeddings as labels. These embeddings serve as item representations in sequential recommender models such as Bidirectional Encoder Representations from Transformers for Sequential Recommendation (BERT4Rec) and Self-Attentive Sequential Recommendation (SASRec). We evaluate four approaches: (1) training the entire recommender model from scratch, (2) incorporating Word2Vec embeddings as item vector in recommenders, (3) incorporating CNN-predicted embeddings only for the last tracks in a playlist while using Word2Vec embeddings for others, and train the remaining model’s parameter, and (4) replacing the CNN with a dilated CNN in the third approach. Our experiments show that audio-based features can enhance playlist continuation, especially in cold-start scenarios, while offering potential for improved explainability over traditional metadata-based methods.
en
dc.description.sponsorship
WWTF Wiener Wissenschafts-, Forschu und Technologiefonds
-
dc.language.iso
en
-
dc.relation.ispartofseries
CEUR Workshop Proceedings
-
dc.rights.uri
http://creativecommons.org/licenses/by/4.0/
-
dc.subject
Music Recommendation
en
dc.subject
Representation Learning
en
dc.subject
Item Embeddings
en
dc.subject
Playlist Completion
en
dc.title
Towards Playlist Continuation Through Large-Scale Context and Audio-Based Music Representations