This repository contains the implementation for the paper:
- "Nikhil Raghav, Avisek Gupta, Md Sahidullah, and Swagatam Das, Self-Tuning Spectral Clustering for Speaker Diarization, to appear in Proc. of ICASSP 2025.
The details of the technique can be found here.
The poster can be accessed here.
🎥 YouTube Presentation: Watch the presentation here
Our implementation is based on a modified version of the AMI recipe provided in the SpeechBrain toolkit.
- Follow the installation guidelines for the SpeechBrain toolkit provided here
The following three files were modified from the exisitng AMI recipe, and were adapted for the experiments on the DIHARD-III dataset. It contains, the scripts for the proposed SC-pNA technique:
- experiment.py located at /speechbrain/recipes/AMI/Diarization/experiment.py
- ecapa_tdnn.yaml located at /speechbrain/recipes/AMI/Diarization/ecapa_tdnn.yaml
- diarization.py located at /speechbrain/speechbrain/processing/diarization.py
If you find our approach useful in your research, please consider citing:
@INPROCEEDINGS{10890194,
author={Raghav, Nikhil and Gupta, Avisek and Sahidullah, Md and Das, Swagatam},
booktitle={ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
title={Self-Tuning Spectral Clustering for Speaker Diarization},
year={2025},
volume={},
number={},
pages={1-5},
keywords={Laplace equations;Costs;Clustering algorithms;Signal processing;Gaussian distribution;Acoustics;Computational efficiency;Sparse matrices;Speech processing;Tuning;speaker diarization;spectral clustering;matrix sparsification;eigengap;DIHARD-III},
doi={10.1109/ICASSP49660.2025.10890194}}
This project is licensed under the MIT License. The full terms of the MIT License can be found in the LICENSE.md file at the root of this project.