中国科学院机构知识库网格系统: Subband fusion of complex spectrogram for fake speech detection

Subband fusion of complex spectrogram for fake speech detection

文献类型：期刊论文


作者	Fan, Cunhang3 ; Xue, Jun 3; Dong, Shunbo 3; Ding, Mingming 3; Yi, Jiangyan2 ; Li, Jinpeng1 ; Lv, Zhao 3
刊名	SPEECH COMMUNICATION
出版日期	2023-11-01
卷号	155 页码:8
关键词	Automatic speaker verification Complex spectrogram Fake speech detection Phase information Subband
ISSN号	0167-6393
DOI	10.1016/j.specom.2023.102988
通讯作者	Lv, Zhao(kjlz@ahu.edu.cn)
英文摘要	The phase information was shown useful in fake speech detection. However, the most common reason why phase-based features are not widely used is phase wrapping. This makes the original phase hard to model directly. Therefore, it remains a challenge how to utilize the phase information effectively. To address this issue, this paper proposes a novel subband fusion of the complex spectrogram method for fake speech detection. The complex spectrogram is used as the input feature, containing both amplitude and phase spectrogram. In addition, subbands of the complex spectrogram are modeled separately. The idea is motivated by the fact that each frequency band has a different effect on the fake speech detection task. Finally, to make full use of the subbands, the subband results are fused. Experimental results on the ASVspoof 2019 LA dataset show that our proposed system achieves an equal error rate (EER) of 0.68% and a minimum tandem detection cost function (min t-DCF) of 0.0224.
WOS关键词	SPEAKER VERIFICATION ; PHASE ; COUNTERMEASURES ; FEATURES
资助项目	National Key Research and Development Program of China[2021ZD0201502] ; National Natural Science Foundation of China (NSFC)[61972437] ; National Natural Science Foundation of China (NSFC)[62201002] ; Excellent Youth Foundation of Anhui Scientific Committee, China[2208085J05] ; Special Fund for Key Program of Science and Technology of Anhui Province, China[202203a07020008] ; Open Research Projects of Zhejiang Lab, China[2021KH0AB06] ; Open Projects Program of National Laboratory of Pattern Recognition, China[202200014]
WOS研究方向	Acoustics ; Computer Science
语种	英语
WOS记录号	WOS:001149053400001
出版者	ELSEVIER
资助机构	National Key Research and Development Program of China ; National Natural Science Foundation of China (NSFC) ; Excellent Youth Foundation of Anhui Scientific Committee, China ; Special Fund for Key Program of Science and Technology of Anhui Province, China ; Open Research Projects of Zhejiang Lab, China ; Open Projects Program of National Laboratory of Pattern Recognition, China
源URL	[http://ir.ia.ac.cn/handle/173211/55428]
专题	多模态人工智能系统全国重点实验室
通讯作者	Lv, Zhao
作者单位	1.Univ Chinese Acad Sci, Ningbo Inst Life & Hlth Ind, Ningbo, Peoples R China 2.Chinese Acad Sci, Inst Automat, Beijing 100190, Peoples R China 3.Anhui Univ, Sch Comp Sci & Technol, Anhui Prov Key Lab Multimodal Cognit Computat, Hefei 230601, Peoples R China
推荐引用方式 GB/T 7714	Fan, Cunhang,Xue, Jun,Dong, Shunbo,et al. Subband fusion of complex spectrogram for fake speech detection[J]. SPEECH COMMUNICATION,2023,155:8.
APA	Fan, Cunhang.,Xue, Jun.,Dong, Shunbo.,Ding, Mingming.,Yi, Jiangyan.,...&Lv, Zhao.(2023).Subband fusion of complex spectrogram for fake speech detection.SPEECH COMMUNICATION,155,8.
MLA	Fan, Cunhang,et al."Subband fusion of complex spectrogram for fake speech detection".SPEECH COMMUNICATION 155(2023):8.

入库方式： OAI收割

来源：自动化研究所

下载0

Subband fusion of complex spectrogram for fake speech detection

其他版本