| October 1, 2019
Project Overview
Digital music services need to build personalized playlists whose successive tracks remain coherent to the listener. This project addressed abrupt transitions in playlists for Qobuz, whose catalogue contained more than 40 million tracks. The original brief proposed an unsupervised model using audio spectrograms to extract latent representations and measure distances between tracks. These distances should capture characteristics such as tempo without preventing playlists from mixing musical styles. The approach could also support recommendations based on overall musical similarity, with computational cost assessed early to ensure that it could scale to the catalogue.
The students developed a proof of concept for classifying pairs of tracks as suitable or unsuitable for a transition. They extracted acoustic descriptors including tempo, chroma, spectral features and MFCCs, and evaluated distance measures and feature weighting on 300 manually labelled pairs. They also compared neural networks and decision trees, and explored visual features extracted from spectrograms with a convolutional neural network. Their work highlighted the cost of feature extraction, the subjectivity of transition labels and the need for a larger annotated dataset. A web interface was proposed to collect additional judgements through crowdsourcing.
References and resources
- Martin F. McKinney and Jeroen Breebaart (2003), Features for Audio and Music Classification.
- Geoffroy Peeters (2004), A large set of audio features for sound description (similarity and classification) in the CUIDADO project.
- Joe George and Lior Shamir (2014), Computer analysis of similarities between albums in popular music.
- Music feature extraction in Python.
- Librosa feature extraction documentation.
📎 Presentation