P2-11: PeakNetFP: Peak-based Neural Audio Fingerprinting Robust to Extreme Time Stretching
Guillem Cortès-Sebastià, Benjamin Martin, Emilio Molina, Xavier Serra, Romain Hennequin
Subjects: Indexing and querying ; Open Review ; Reproducibility ; Similarity metrics ; Fingerprinting ; Evaluation, datasets, and reproducibility ; MIR tasks ; Pattern matching and detection
Presented In-person
4-minute short-format presentation
This work introduces PeakNetFP, the first neural audio fingerprinting (AFP) system designed specifically around spectral peaks. This novel system is designed to leverage the sparse spectral coordinates typically computed by traditional peak-based AFP methods. PeakNetFP performs hierarchical point feature extraction techniques similar to the computer vision model PointNet++, and is trained using contrastive learning like in the state-of-the-art deep learning AFP, NeuralFP. This combination allows PeakNetFP to outperform conventional AFP systems and achieves comparable performance to NeuralFP when handling challenging time-stretched audio data. In extensive evaluation, PeakNetFP maintains a Top-1 hit rate of over 90% for stretching factors ranging from 50% to 200%. Moreover, PeakNetFP offers significant efficiency advantages: compared to NeuralFP, it has 100 times fewer parameters and uses 11 times smaller input data. These features make PeakNetFP a lightweight and efficient solution for AFP tasks where time stretching is involved. Overall, this system represents a promising direction for future AFP technologies, as it successfully merges the lightweight nature of peak-based AFP with the adaptability and pattern recognition capabilities of neural network-based approaches, paving the way for more scalable and efficient solutions in the field.
Q2 ( I am an expert on the topic of the paper.)
Strongly agree
Q3 ( The title and abstract reflect the content of the paper.)
Strongly agree
Q4 (The paper discusses, cites and compares with all relevant related work.)
Strongly agree
Q6 (Readability and paper organization: The writing and language are clear and structured in a logical manner.)
Agree
Q7 (The paper adheres to ISMIR 2025 submission guidelines (uses the ISMIR 2025 template, has at most 6 pages of technical content followed by “n” pages of references or ethical considerations, references are well formatted). If you selected “No”, please explain the issue in your comments.)
Yes
Q8 (Relevance of the topic to ISMIR: The topic of the paper is relevant to the ISMIR community. Note that submissions of novel music-related topics, tasks, and applications are highly encouraged. If you think that the paper has merit but does not exactly match the topics of ISMIR, please do not simply reject the paper but instead communicate this to the Program Committee Chairs. Please do not penalize the paper when the proposed method can also be applied to non-music domains if it is shown to be useful in music domains.)
Strongly agree
Q9 (Scholarly/scientific quality: The content is scientifically correct.)
Strongly agree
Q11 (Novelty of the paper: The paper provides novel methods, applications, findings or results. Please do not narrowly view "novelty" as only new methods or theories. Papers proposing novel musical applications of existing methods from other research fields are considered novel at ISMIR conferences.)
Agree
Q12 (The paper provides all the necessary details or material to reproduce the results described in the paper. Keep in mind that ISMIR respects the diversity of academic disciplines, backgrounds, and approaches. Although ISMIR has a tradition of publishing open datasets and open-source projects to enhance the scientific reproducibility, ISMIR accepts submissions using proprietary datasets and implementations that are not sharable. Please do not simply reject the paper when proprietary datasets or implementations are used.)
Agree
Q13 (Pioneering proposals: This paper proposes a novel topic, task or application. Since this is intended to encourage brave new ideas and challenges, papers rated “Strongly Agree” and “Agree” can be highlighted, but please do not penalize papers rated “Disagree” or “Strongly Disagree”. Keep in mind that it is often difficult to provide baseline comparisons for novel topics, tasks, or applications. If you think that the novelty is high but the evaluation is weak, please do not simply reject the paper but carefully assess the value of the paper for the community.)
Disagree (Standard topic, task, or application)
Q14 (Reusable insights: The paper provides reusable insights (i.e. the capacity to gain an accurate and deep understanding). Such insights may go beyond the scope of the paper, domain or application, in order to build up consistent knowledge across the MIR community.)
Agree
Q15 (Please explain your assessment of reusable insights in the paper.)
The idea of point cloud processing can be used for other applications in music and audio processing
Q16 ( Write ONE line (in your own words) with the main take-home message from the paper.)
Audio fingerprinting method using only the spectral peaks and PointNet++, leading to efficient deployment
Q17 (This paper is of award-winning quality.)
No
Q19 (Potential to generate discourse: The paper will generate discourse at the ISMIR conference or have a large influence/impact on the future of the ISMIR community.)
Agree
Q20 (Overall evaluation (to be completed before the discussion phase): Please first evaluate before the discussion phase. Keep in mind that minor flaws can be corrected, and should not be a reason to reject a paper. Please familiarize yourself with the reviewer guidelines at https://ismir.net/reviewer-guidelines.)
Weak accept
Q21 (Main review and comments for the authors (to be completed before the discussion phase). Please summarize strengths and weaknesses of the paper. It is essential that you justify the reason for the overall evaluation score in detail. Keep in mind that belittling or sarcastic comments are not appropriate.)
Contribution: - The paper proposes to use PointNet++ for the Audio Fingerprinting (AFP) task - This leads to a reduction in model size and input data size - Results are on par with the NeuralFP system.
Limitations: - The experimental evaluation is very limited. - The pointNet++ algorithm, as used for AFP, is not described. - What is the feature vector at every peak? (Line 347 talks about feature aggregation) - Is the magnitude of the peak considered? - Is the displacement (distance + direction) between peaks considered for deriving the aggregated features? What is the mathematical formula used?
Q22 (Final recommendation (to be completed after the discussion phase) Please give a final recommendation after the discussion phase. In the final recommendation, please do not simply average the scores of the reviewers. Note that the number of recommendation options for reviewers is different from the number of options here. We encourage you to take a stand, and preferably avoid “weak accepts” or “weak rejects” if possible.)
Strong accept
Q23 (Meta-review and final comments for authors (to be completed after the discussion phase))
All reviewers unanimously agree to accept this paper. The authors could do minor edits to include the reviewer suggestions in the final version of the paper.
Q2 ( I am an expert on the topic of the paper.)
Strongly agree
Q3 (The title and abstract reflect the content of the paper.)
Strongly agree
Q4 (The paper discusses, cites and compares with all relevant related work)
Strongly agree
Q6 (Readability and paper organization: The writing and language are clear and structured in a logical manner.)
Strongly agree
Q7 (The paper adheres to ISMIR 2025 submission guidelines (uses the ISMIR 2025 template, has at most 6 pages of technical content followed by “n” pages of references or ethical considerations, references are well formatted). If you selected “No”, please explain the issue in your comments.)
Yes
Q8 (Relevance of the topic to ISMIR: The topic of the paper is relevant to the ISMIR community. Note that submissions of novel music-related topics, tasks, and applications are highly encouraged. If you think that the paper has merit but does not exactly match the topics of ISMIR, please do not simply reject the paper but instead communicate this to the Program Committee Chairs. Please do not penalize the paper when the proposed method can also be applied to non-music domains if it is shown to be useful in music domains.)
Strongly agree
Q9 (Scholarly/scientific quality: The content is scientifically correct.)
Strongly agree
Q11 (Novelty of the paper: The paper provides novel methods, applications, findings or results. Please do not narrowly view "novelty" as only new methods or theories. Papers proposing novel musical applications of existing methods from other research fields are considered novel at ISMIR conferences.)
Strongly agree
Q12 (The paper provides all the necessary details or material to reproduce the results described in the paper. Keep in mind that ISMIR respects the diversity of academic disciplines, backgrounds, and approaches. Although ISMIR has a tradition of publishing open datasets and open-source projects to enhance the scientific reproducibility, ISMIR accepts submissions using proprietary datasets and implementations that are not sharable. Please do not simply reject the paper when proprietary datasets or implementations are used.)
Agree
Q13 (Pioneering proposals: This paper proposes a novel topic, task or application. Since this is intended to encourage brave new ideas and challenges, papers rated "Strongly Agree" and "Agree" can be highlighted, but please do not penalize papers rated "Disagree" or "Strongly Disagree". Keep in mind that it is often difficult to provide baseline comparisons for novel topics, tasks, or applications. If you think that the novelty is high but the evaluation is weak, please do not simply reject the paper but carefully assess the value of the paper for the community.)
Agree (Novel topic, task, or application)
Q14 (Reusable insights: The paper provides reusable insights (i.e. the capacity to gain an accurate and deep understanding). Such insights may go beyond the scope of the paper, domain or application, in order to build up consistent knowledge across the MIR community.)
Agree
Q15 (Please explain your assessment of reusable insights in the paper.)
The authors actually tried to keep part of their model similar to an existing model (with published code) to make the performance comparable. They used an existing and public dataset (which they augmented). They also published their own model (anonymized).
Q16 (Write ONE line (in your own words) with the main take-home message from the paper.)
The authors proposed to combine the benefits of peak-based and NN-based audio fingerprinting approaches, by generating fingerprints from peaks rather than spectrograms directly. They showed that their model was able to perform better than a peak-based method and comparable to a full NN-based method, while reducing the size of the final inputs and the model itself.
Q17 (Would you recommend this paper for an award?)
No
Q19 (Potential to generate discourse: The paper will generate discourse at the ISMIR conference or have a large influence/impact on the future of the ISMIR community.)
Agree
Q20 (Overall evaluation: Keep in mind that minor flaws can be corrected, and should not be a reason to reject a paper. Please familiarize yourself with the reviewer guidelines at https://ismir.net/reviewer-guidelines)
Strong accept
Q21 (Main review and comments for the authors. Please summarize strengths and weaknesses of the paper. It is essential that you justify the reason for the overall evaluation score in detail. Keep in mind that belittling or sarcastic comments are not appropriate.)
General comments:
- This is a well-written paper, with a practical contribution and solid references. I wish the description of the model will come with better/more explanations and clearer/larger plots. The evaluation could also be a bit improved, to better showcase the performance of the proposed model. See the details comments below.
Detailed comments:
1. - Reference [18] does not seem to be about audio fingerprinting but codec identification. Perhaps you meant to cite the following patent from this company? R. Coover and Z. Rafii, “Methods and Apparatus to Fingerprint an Audio Signal via Normalization,” 16/453,654, Mar. 2020.
-
Figure 1 is too small.
-
When you say "raw local maxima," do you mean that you also use the amplitude of the peaks? And how do you define a local maxima? Do you use a specific window size or maximum peak distance(s)? How do you deal with silences or segments with sparse energy? I would have liked to see some information about the spectrogram itself.
-
I would also make Figures 2 and 3 bigger (including the fonts). It would really help the reader to understand the system.
-
I am unclear what SA and MSG are in Figures 2 and 3. What kind of features gets into the MLP then? Section 3.2 is a very important section as it is meant to explain the core of the algorithm. I would (try to) make it clearer. I didn't fully understand the steps, how the features would look like at every step, and what is actually being learned.
-
"NT-Xent"
-
What is a "faiss" index?
-
Could you write a bit more about the dataset, the type of audio in it? Does it come with noise too?
-
It would make things easier for the reader if Table 1 had more information in the caption. What are j and A, B, C, for example?
-
Why not (also) use the more classic TP/FP/FN, and/or recall and precision as the metrics?
-
The fonts in Figure 4 are too small; I cannot really see the values in the plot.
-
Some of the results showed in section 4.2 could be turned into (perhaps more convincing) tables (or perhaps you were too short on space).
6. - Please, fix your bibliography: - Be consistent in the way you list every entry - Avoid repetition (some entries have the name of the conference and the location listed twice) - Use capital letters when needed (e.g., proper nouns and acronyms)
Q2 ( I am an expert on the topic of the paper.)
Agree
Q3 (The title and abstract reflect the content of the paper.)
Strongly agree
Q4 (The paper discusses, cites and compares with all relevant related work)
Strongly agree
Q6 (Readability and paper organization: The writing and language are clear and structured in a logical manner.)
Strongly agree
Q7 (The paper adheres to ISMIR 2025 submission guidelines (uses the ISMIR 2025 template, has at most 6 pages of technical content followed by “n” pages of references or ethical considerations, references are well formatted). If you selected “No”, please explain the issue in your comments.)
Yes
Q8 (Relevance of the topic to ISMIR: The topic of the paper is relevant to the ISMIR community. Note that submissions of novel music-related topics, tasks, and applications are highly encouraged. If you think that the paper has merit but does not exactly match the topics of ISMIR, please do not simply reject the paper but instead communicate this to the Program Committee Chairs. Please do not penalize the paper when the proposed method can also be applied to non-music domains if it is shown to be useful in music domains.)
Strongly agree
Q9 (Scholarly/scientific quality: The content is scientifically correct.)
Strongly agree
Q11 (Novelty of the paper: The paper provides novel methods, applications, findings or results. Please do not narrowly view "novelty" as only new methods or theories. Papers proposing novel musical applications of existing methods from other research fields are considered novel at ISMIR conferences.)
Strongly agree
Q12 (The paper provides all the necessary details or material to reproduce the results described in the paper. Keep in mind that ISMIR respects the diversity of academic disciplines, backgrounds, and approaches. Although ISMIR has a tradition of publishing open datasets and open-source projects to enhance the scientific reproducibility, ISMIR accepts submissions using proprietary datasets and implementations that are not sharable. Please do not simply reject the paper when proprietary datasets or implementations are used.)
Strongly agree
Q13 (Pioneering proposals: This paper proposes a novel topic, task or application. Since this is intended to encourage brave new ideas and challenges, papers rated "Strongly Agree" and "Agree" can be highlighted, but please do not penalize papers rated "Disagree" or "Strongly Disagree". Keep in mind that it is often difficult to provide baseline comparisons for novel topics, tasks, or applications. If you think that the novelty is high but the evaluation is weak, please do not simply reject the paper but carefully assess the value of the paper for the community.)
Strongly Agree (Very novel topic, task, or application)
Q14 (Reusable insights: The paper provides reusable insights (i.e. the capacity to gain an accurate and deep understanding). Such insights may go beyond the scope of the paper, domain or application, in order to build up consistent knowledge across the MIR community.)
Strongly agree
Q15 (Please explain your assessment of reusable insights in the paper.)
The paper provides an innovative use of a hierarchical grouping algorithm for the task of Neural Audio Fingerprinting, resulting in a considerably smaller model size. While the authors present its usage for the task of time-stretching, it can be used for any neural fingerprinting approach.
Q16 (Write ONE line (in your own words) with the main take-home message from the paper.)
This paper presents an efficient method for neural audio fingerprinting - utilizing a hierarchical grouping algorithm, the authors show that their approach performs almost on par with a spectral based approach, but requires only roughly 1% of model size.
Q17 (Would you recommend this paper for an award?)
Yes
Q18 ( If yes, please explain why it should be awarded.)
This paper is very well written and clearly outlines the motivation. With neural audio fingerprinting, the authors tackle an important topic in MIR, which on top of artist discovery can aid digital rights management and fair artist compensation. In the era of generative music AI, such fingerprinting systems become even more crucial. The authors present a smart, signal processing based feature engineering approach for a more efficient implementation that can help save resources which is a further important factor for the paper to be outstanding. To me, the combination of these points reasonably justify the nomination for a best paper award.
Q19 (Potential to generate discourse: The paper will generate discourse at the ISMIR conference or have a large influence/impact on the future of the ISMIR community.)
Agree
Q20 (Overall evaluation: Keep in mind that minor flaws can be corrected, and should not be a reason to reject a paper. Please familiarize yourself with the reviewer guidelines at https://ismir.net/reviewer-guidelines)
Strong accept
Q21 (Main review and comments for the authors. Please summarize strengths and weaknesses of the paper. It is essential that you justify the reason for the overall evaluation score in detail. Keep in mind that belittling or sarcastic comments are not appropriate.)
The authors of this paper present a novel approach to neural audio fingerprinting, combining the neural self-supervised approach such as in NeuralFP with peak-based algorithms. With substantially less parameters, their proposed model performs comparably well with an implementation of NeuralFP.
The paper is very well written and follows a clear and logical structure. The motivation is outlined well and a good overview is given over related previous works. Section 3.2 "Hierarchical peak set feature extraction", being a crucial part of this work could be presented in a little more detail as PointNet++ will be less familiar to the MIR community.
Some minor notes on the paper:
Line 100 ff: "This method is commonly used by DJs to synchronize the tempo of different songs within a mix or to create remixes that are either slowed down or sped up [20]". -> Yes, while traditionally DJs apply pitch-shifting which incorporates both time and freqeuncy changes of the signal. In fact, it would have been interesting to see the apporach for manipulation in both of these domains. To the best of my knowledge, time stretching and pitch shifting are often used by video creators to circumvent licensing costs for music they use in their productions. This could be added to the motivation (there might be internet sources at least).
Line 389: The "candidate pruning" could be described in a little more detail. Also, this term does not appear in the referenced NeuralFP paper - I suggest you are referring to what they denote as "Sequence search"?
Line 400: "and it comes" -> "that comes"
Line 405: It is unclear to me where the 278 seconds average length come from.
Line 408: "in" -> "as in" ?
Line 433: What motivates the somehow uncommon batch size of 240?
Line 502: With "query size", do you refer to the segments' size in seconds? It would be favorable if you would stick to a clear unified terminology here (you usually use "query length").
I strongly suggest that this paper will be accepted for ISMIR 2025. I believe that this paper would also deserve a nomination for a best paper award. As noted at point 18:
This paper is very well written and clearly outlines the motivation. With neural audio fingerprinting, the authors tackle an important topic in MIR, which on top of artist discovery can aid digital rights management and fair artist compensation. In the era of generative music AI, such fingerprinting systems become even more crucial. The authors present a smart, signal processing based feature engineering approach for a more efficient implementation that can help save resources which is a further important factor for the paper to be outstanding. To me, the combination of these points reasonably justify the nomination for a best paper award.
Q2 ( I am an expert on the topic of the paper.)
Strongly agree
Q3 (The title and abstract reflect the content of the paper.)
Strongly agree
Q4 (The paper discusses, cites and compares with all relevant related work)
Agree
Q6 (Readability and paper organization: The writing and language are clear and structured in a logical manner.)
Strongly agree
Q7 (The paper adheres to ISMIR 2025 submission guidelines (uses the ISMIR 2025 template, has at most 6 pages of technical content followed by “n” pages of references or ethical considerations, references are well formatted). If you selected “No”, please explain the issue in your comments.)
Yes
Q8 (Relevance of the topic to ISMIR: The topic of the paper is relevant to the ISMIR community. Note that submissions of novel music-related topics, tasks, and applications are highly encouraged. If you think that the paper has merit but does not exactly match the topics of ISMIR, please do not simply reject the paper but instead communicate this to the Program Committee Chairs. Please do not penalize the paper when the proposed method can also be applied to non-music domains if it is shown to be useful in music domains.)
Strongly agree
Q9 (Scholarly/scientific quality: The content is scientifically correct.)
Agree
Q11 (Novelty of the paper: The paper provides novel methods, applications, findings or results. Please do not narrowly view "novelty" as only new methods or theories. Papers proposing novel musical applications of existing methods from other research fields are considered novel at ISMIR conferences.)
Agree
Q12 (The paper provides all the necessary details or material to reproduce the results described in the paper. Keep in mind that ISMIR respects the diversity of academic disciplines, backgrounds, and approaches. Although ISMIR has a tradition of publishing open datasets and open-source projects to enhance the scientific reproducibility, ISMIR accepts submissions using proprietary datasets and implementations that are not sharable. Please do not simply reject the paper when proprietary datasets or implementations are used.)
Strongly agree
Q13 (Pioneering proposals: This paper proposes a novel topic, task or application. Since this is intended to encourage brave new ideas and challenges, papers rated "Strongly Agree" and "Agree" can be highlighted, but please do not penalize papers rated "Disagree" or "Strongly Disagree". Keep in mind that it is often difficult to provide baseline comparisons for novel topics, tasks, or applications. If you think that the novelty is high but the evaluation is weak, please do not simply reject the paper but carefully assess the value of the paper for the community.)
Agree (Novel topic, task, or application)
Q14 (Reusable insights: The paper provides reusable insights (i.e. the capacity to gain an accurate and deep understanding). Such insights may go beyond the scope of the paper, domain or application, in order to build up consistent knowledge across the MIR community.)
Strongly agree
Q15 (Please explain your assessment of reusable insights in the paper.)
The paper presents a clean, reusable framework for combining traditional sparse audio fingerprinting features with neural embeddings. While the architecture and training pipeline are adapted from prior work (e.g., NeuralFP, PointNet++), the insight of applying them to spectral peaks under strong time-stretching conditions opens opportunities for further development in lightweight industrial-scale systems.
Q16 (Write ONE line (in your own words) with the main take-home message from the paper.)
PeakNetFP is a neural fingerprinting system built on sparse spectral peaks that offers robustness to time stretching with significantly reduced model size and input dimensionality.
Q17 (Would you recommend this paper for an award?)
No
Q19 (Potential to generate discourse: The paper will generate discourse at the ISMIR conference or have a large influence/impact on the future of the ISMIR community.)
Agree
Q20 (Overall evaluation: Keep in mind that minor flaws can be corrected, and should not be a reason to reject a paper. Please familiarize yourself with the reviewer guidelines at https://ismir.net/reviewer-guidelines)
Strong accept
Q21 (Main review and comments for the authors. Please summarize strengths and weaknesses of the paper. It is essential that you justify the reason for the overall evaluation score in detail. Keep in mind that belittling or sarcastic comments are not appropriate.)
Dear authors, thank you for your paper. It is clearly written, focused, and well-executed around a central idea.
The motivation is strong: the paper addresses a real-world gap by combining sparse peak-based input with neural representations. This hybrid strategy is well-motivated for low-resource or privacy-constrained environments. Your framing of the practical trade-offs (e.g., computation on client devices, data volume, dense vs. sparse representations) is excellent and grounds the paper in relevant applications.
The technical implementation is clearly explained, and the evaluation is thorough. The comparison with QuadFP and NeuralFP is carefully done, and results are clearly presented. I appreciated the validation of your own QuadFP implementation and the transparency around dataset differences.
There are a few areas for improvement: * While the paper focuses on robustness to time stretching, omitting pitch shifting is a limitation. Since pitch and tempo changes often co-occur in real-world transformations, I encourage you to include this in future work (which you mention at the end). * The “11× smaller input data” claim in the abstract and conclusion is interesting, but the exact basis or formula behind this figure is not explained. Please consider clarifying how it’s computed. * The categorization of related work into “local descriptor-based” approaches could benefit from a clearer definition and concrete examples. * The use of “faiss” should be capitalized to “Faiss”.
One of the strongest aspects of the paper is that you commit to sharing both code and dataset. This makes it easier to reproduce your results and build on top of the work, and I expect it will encourage follow-up research from both your team and others.
While this is not a “revolutionary” paper in terms of algorithmic novelty, it is a solid contribution with strong execution and real practical relevance. I strongly support its acceptance.