Katalog Plus
Bibliothek der Frankfurt UAS
Bald neuer Katalog: sichern Sie sich schon vorab Ihre persönlichen Merklisten im Nutzerkonto: Anleitung.
Dieses Ergebnis aus BASE kann Gästen nicht angezeigt werden.  Login für vollen Zugriff.

Joint optimization of diffusion probabilistic-based multichannel speech enhancement with far-field speaker verification

Title: Joint optimization of diffusion probabilistic-based multichannel speech enhancement with far-field speaker verification
Authors: Dowerah, Sandipana; Serizel, Romain; Jouvet, Denis; Mohammadamini, M; Matrouf, Driss
Contributors: Speech Modeling for Facilitating Oral-Based Communication (MULTISPEECH); Centre Inria de l'Université de Lorraine; Institut National de Recherche en Informatique et en Automatique (Inria)-Institut National de Recherche en Informatique et en Automatique (Inria)-Department of Natural Language Processing & Knowledge Discovery (LORIA - NLPKD); Laboratoire Lorrain de Recherche en Informatique et ses Applications (LORIA); Institut National de Recherche en Informatique et en Automatique (Inria)-CentraleSupélec-Université de Lorraine (UL)-Centre National de la Recherche Scientifique (CNRS)-Institut National de Recherche en Informatique et en Automatique (Inria)-CentraleSupélec-Université de Lorraine (UL)-Centre National de la Recherche Scientifique (CNRS)-Laboratoire Lorrain de Recherche en Informatique et ses Applications (LORIA); Institut National de Recherche en Informatique et en Automatique (Inria)-CentraleSupélec-Université de Lorraine (UL)-Centre National de la Recherche Scientifique (CNRS)-CentraleSupélec-Université de Lorraine (UL)-Centre National de la Recherche Scientifique (CNRS); Laboratoire Informatique d'Avignon (LIA); Avignon Université (AU)-Centre d'Enseignement et de Recherche en Informatique - CERI; ANR-18-CE33-0014,ROBOVOX,ROBOVOX - Identification vocale robuste pour les robots de sécurité mobiles(2018)
Source: IEEE SLT 2022 ; https://hal.science/hal-03671583 ; IEEE SLT 2022, Jan 2023, Doha, Qatar
Publisher Information: CCSD
Publication Year: 2023
Collection: Université de Lorraine: HAL
Subject Terms: multichannel speech enhancement; diffusion model; far-field speaker verification; [INFO.INFO-HC]Computer Science [cs]/Human-Computer Interaction [cs.HC]
Subject Geographic: Doha; Qatar
Description: International audience ; Today's smart devices using speaker verification are getting equipped with multiple microphones resulting in improving spatial ambiguity and directivity. However, unlike any other speech-based applications, the performance of speaker verification degrades in far-field scenarios due to the adverse effects of a noisy environment and room reverberation. This paper presents a novel multichannel speech enhancement module based on the diffusion probabilistic model. It is used as the front-end of the ECAPA-TDNN speaker verification system in far-field scenarios under a noisy-reverberant environment. The proposed system incorporates a two-stage training approach. In the first stage, both speech enhancement and speaker verification modules are trained individually. In the second stage, both the modules are combined to jointly trained them. We use similaritypreserving knowledge distillation loss that guides the network to produce similar activation for enhanced signals to that of clean speech signals. Using joint optimization with knowledge distillation loss achieved the best performance on both the evaluation composed of synthetic clips similar to those used at training and on unseen recorded clips from the VOiCES dataset.
Document Type: conference object
Language: English
Availability: https://hal.science/hal-03671583; https://hal.science/hal-03671583v2/document; https://hal.science/hal-03671583v2/file/SLT_2022.pdf
Rights: info:eu-repo/semantics/OpenAccess
Accession Number: edsbas.149D1DD2
Database: BASE