中国科学院机构知识库网格
Chinese Academy of Sciences Institutional Repositories Grid
Subband fusion of complex spectrogram for fake speech detection

文献类型:期刊论文

作者Fan, Cunhang3; Xue, Jun3; Dong, Shunbo3; Ding, Mingming3; Yi, Jiangyan2; Li, Jinpeng1; Lv, Zhao3
刊名SPEECH COMMUNICATION
出版日期2023-11-01
卷号155页码:8
ISSN号0167-6393
关键词Automatic speaker verification Complex spectrogram Fake speech detection Phase information Subband
DOI10.1016/j.specom.2023.102988
通讯作者Lv, Zhao(kjlz@ahu.edu.cn)
英文摘要The phase information was shown useful in fake speech detection. However, the most common reason why phase-based features are not widely used is phase wrapping. This makes the original phase hard to model directly. Therefore, it remains a challenge how to utilize the phase information effectively. To address this issue, this paper proposes a novel subband fusion of the complex spectrogram method for fake speech detection. The complex spectrogram is used as the input feature, containing both amplitude and phase spectrogram. In addition, subbands of the complex spectrogram are modeled separately. The idea is motivated by the fact that each frequency band has a different effect on the fake speech detection task. Finally, to make full use of the subbands, the subband results are fused. Experimental results on the ASVspoof 2019 LA dataset show that our proposed system achieves an equal error rate (EER) of 0.68% and a minimum tandem detection cost function (min t-DCF) of 0.0224.
WOS关键词SPEAKER VERIFICATION ; PHASE ; COUNTERMEASURES ; FEATURES
资助项目National Key Research and Development Program of China[2021ZD0201502] ; National Natural Science Foundation of China (NSFC)[61972437] ; National Natural Science Foundation of China (NSFC)[62201002] ; Excellent Youth Foundation of Anhui Scientific Committee, China[2208085J05] ; Special Fund for Key Program of Science and Technology of Anhui Province, China[202203a07020008] ; Open Research Projects of Zhejiang Lab, China[2021KH0AB06] ; Open Projects Program of National Laboratory of Pattern Recognition, China[202200014]
WOS研究方向Acoustics ; Computer Science
语种英语
出版者ELSEVIER
WOS记录号WOS:001149053400001
资助机构National Key Research and Development Program of China ; National Natural Science Foundation of China (NSFC) ; Excellent Youth Foundation of Anhui Scientific Committee, China ; Special Fund for Key Program of Science and Technology of Anhui Province, China ; Open Research Projects of Zhejiang Lab, China ; Open Projects Program of National Laboratory of Pattern Recognition, China
源URL[http://ir.ia.ac.cn/handle/173211/55428]  
专题多模态人工智能系统全国重点实验室
通讯作者Lv, Zhao
作者单位1.Univ Chinese Acad Sci, Ningbo Inst Life & Hlth Ind, Ningbo, Peoples R China
2.Chinese Acad Sci, Inst Automat, Beijing 100190, Peoples R China
3.Anhui Univ, Sch Comp Sci & Technol, Anhui Prov Key Lab Multimodal Cognit Computat, Hefei 230601, Peoples R China
推荐引用方式
GB/T 7714
Fan, Cunhang,Xue, Jun,Dong, Shunbo,et al. Subband fusion of complex spectrogram for fake speech detection[J]. SPEECH COMMUNICATION,2023,155:8.
APA Fan, Cunhang.,Xue, Jun.,Dong, Shunbo.,Ding, Mingming.,Yi, Jiangyan.,...&Lv, Zhao.(2023).Subband fusion of complex spectrogram for fake speech detection.SPEECH COMMUNICATION,155,8.
MLA Fan, Cunhang,et al."Subband fusion of complex spectrogram for fake speech detection".SPEECH COMMUNICATION 155(2023):8.

入库方式: OAI收割

来源:自动化研究所

浏览0
下载0
收藏0
其他版本

除非特别说明,本系统中所有内容都受版权保护,并保留所有权利。