论文标题

现代化开放式语音语言标识

Modernizing Open-Set Speech Language Identification

论文作者

Eyceoz, Mustafa, Lee, Justin, Beigi, Homayoon

论文摘要

尽管大多数现代语音语言识别方法都是封闭式的,但我们希望查看是否可以修改并适应开放式问题。切换到开放集问题时,解决方案将获得拒绝音频输入的能力,而当它无法匹配我们的任何已知语言选项时。我们通过调整两种现代的最先进的方法来解决开放设定的任务:第一次使用CRNN引起注意,第二款使用TDNN。除了使用MFCC,对数频谱功能和音调增强输入功能嵌入外,我们还将尝试两种方法来检测台外语言检测:一个使用阈值,另一个实质上是执行验证任务。我们将比较TDNN和CRNN的性能以及我们的检测方法。

While most modern speech Language Identification methods are closed-set, we want to see if they can be modified and adapted for the open-set problem. When switching to the open-set problem, the solution gains the ability to reject an audio input when it fails to match any of our known language options. We tackle the open-set task by adapting two modern-day state-of-the-art approaches to closed-set language identification: the first using a CRNN with attention and the second using a TDNN. In addition to enhancing our input feature embeddings using MFCCs, log spectral features, and pitch, we will be attempting two approaches to out-of-set language detection: one using thresholds, and the other essentially performing a verification task. We will compare both the performance of the TDNN and the CRNN, as well as our detection approaches.

扫码加入交流群

加入微信交流群

微信交流群二维码

扫码加入学术交流群,获取更多资源