Subband fusion of complex spectrogram for fake speech detection
文献类型:期刊论文
作者 | Fan, Cunhang3; Xue, Jun3; Dong, Shunbo3; Ding, Mingming3; Yi, Jiangyan2; Li, Jinpeng1; Lv, Zhao3 |
刊名 | SPEECH COMMUNICATION |
出版日期 | 2023-11-01 |
卷号 | 155页码:8 |
ISSN号 | 0167-6393 |
关键词 | Automatic speaker verification Complex spectrogram Fake speech detection Phase information Subband |
DOI | 10.1016/j.specom.2023.102988 |
通讯作者 | Lv, Zhao(kjlz@ahu.edu.cn) |
英文摘要 | The phase information was shown useful in fake speech detection. However, the most common reason why phase-based features are not widely used is phase wrapping. This makes the original phase hard to model directly. Therefore, it remains a challenge how to utilize the phase information effectively. To address this issue, this paper proposes a novel subband fusion of the complex spectrogram method for fake speech detection. The complex spectrogram is used as the input feature, containing both amplitude and phase spectrogram. In addition, subbands of the complex spectrogram are modeled separately. The idea is motivated by the fact that each frequency band has a different effect on the fake speech detection task. Finally, to make full use of the subbands, the subband results are fused. Experimental results on the ASVspoof 2019 LA dataset show that our proposed system achieves an equal error rate (EER) of 0.68% and a minimum tandem detection cost function (min t-DCF) of 0.0224. |
WOS关键词 | SPEAKER VERIFICATION ; PHASE ; COUNTERMEASURES ; FEATURES |
资助项目 | National Key Research and Development Program of China[2021ZD0201502] ; National Natural Science Foundation of China (NSFC)[61972437] ; National Natural Science Foundation of China (NSFC)[62201002] ; Excellent Youth Foundation of Anhui Scientific Committee, China[2208085J05] ; Special Fund for Key Program of Science and Technology of Anhui Province, China[202203a07020008] ; Open Research Projects of Zhejiang Lab, China[2021KH0AB06] ; Open Projects Program of National Laboratory of Pattern Recognition, China[202200014] |
WOS研究方向 | Acoustics ; Computer Science |
语种 | 英语 |
出版者 | ELSEVIER |
WOS记录号 | WOS:001149053400001 |
资助机构 | National Key Research and Development Program of China ; National Natural Science Foundation of China (NSFC) ; Excellent Youth Foundation of Anhui Scientific Committee, China ; Special Fund for Key Program of Science and Technology of Anhui Province, China ; Open Research Projects of Zhejiang Lab, China ; Open Projects Program of National Laboratory of Pattern Recognition, China |
源URL | [http://ir.ia.ac.cn/handle/173211/55428] |
专题 | 多模态人工智能系统全国重点实验室 |
通讯作者 | Lv, Zhao |
作者单位 | 1.Univ Chinese Acad Sci, Ningbo Inst Life & Hlth Ind, Ningbo, Peoples R China 2.Chinese Acad Sci, Inst Automat, Beijing 100190, Peoples R China 3.Anhui Univ, Sch Comp Sci & Technol, Anhui Prov Key Lab Multimodal Cognit Computat, Hefei 230601, Peoples R China |
推荐引用方式 GB/T 7714 | Fan, Cunhang,Xue, Jun,Dong, Shunbo,et al. Subband fusion of complex spectrogram for fake speech detection[J]. SPEECH COMMUNICATION,2023,155:8. |
APA | Fan, Cunhang.,Xue, Jun.,Dong, Shunbo.,Ding, Mingming.,Yi, Jiangyan.,...&Lv, Zhao.(2023).Subband fusion of complex spectrogram for fake speech detection.SPEECH COMMUNICATION,155,8. |
MLA | Fan, Cunhang,et al."Subband fusion of complex spectrogram for fake speech detection".SPEECH COMMUNICATION 155(2023):8. |
入库方式: OAI收割
来源:自动化研究所
浏览0
下载0
收藏0
其他版本
除非特别说明,本系统中所有内容都受版权保护,并保留所有权利。