中国科学院机构知识库网格
Chinese Academy of Sciences Institutional Repositories Grid
Spelling correction of non-word errors in Uyghur-Chinese machine translation

文献类型:期刊论文

作者Dong, Rui1,2,3; Yang, Yating1,2; Jiang, Tonghai1
刊名Information
出版日期2019
卷号10期号:6页码:1-9
关键词spelling correction natural language processing machine translation language model Uyghur
ISSN号20782489
英文摘要

This research was conducted to solve the out-of-vocabulary problem caused by Uyghur spelling errors in Uyghur-Chinese machine translation, so as to improve the quality of Uyghur-Chinese machine translation. This paper assesses three spelling correction methods based on machine translation: 1. Using a Bilingual Evaluation Understudy (BLEU) score; 2. Using a Chinese language model; 3. Using a bilingual language model. The best results were achieved in both the spelling correction task and the machine translation task by using the BLEU score for spelling correction. A maximum F1 score of 0.72 was reached for spelling correction, and the translation result increased the BLEU score by 1.97 points, relative to the baseline system. However, the method of using a BLEU score for spelling correction requires the support of a bilingual parallel corpus, which is a supervised method that can be used in corpus pre-processing. Unsupervised spelling correction can be performed by using either a Chinese language model or a bilingual language model. These two methods can be easily extended to other languages, such as Arabic.

源URL[http://ir.xjipc.cas.cn/handle/365002/7798]  
专题新疆理化技术研究所_多语种信息技术研究室
通讯作者Dong, Rui1,2,3; Jiang, Tonghai1
作者单位1.100049, China
2.University of the Chinese Academy of Sciences, Beijing
3.Xinjiang Laboratory of Minority Speech and Language Information Processing, Urumqi
4.830011, China
5.Xinjiang Technical Institute of Physics and Chemistry Chinese Academy of Science, Urumqi
推荐引用方式
GB/T 7714
Dong, Rui1,2,3,Yang, Yating1,2,Jiang, Tonghai1. Spelling correction of non-word errors in Uyghur-Chinese machine translation[J]. Information,2019,10(6):1-9.
APA Dong, Rui1,2,3,Yang, Yating1,2,&Jiang, Tonghai1.(2019).Spelling correction of non-word errors in Uyghur-Chinese machine translation.Information,10(6),1-9.
MLA Dong, Rui1,2,3,et al."Spelling correction of non-word errors in Uyghur-Chinese machine translation".Information 10.6(2019):1-9.

入库方式: OAI收割

来源:新疆理化技术研究所

浏览0
下载0
收藏0
其他版本

除非特别说明,本系统中所有内容都受版权保护,并保留所有权利。