船舶与海洋工程

AIS拼音船名到汉字的智能翻译技术研究

  • 潘明阳 ,
  • 李琦 ,
  • 盛尊阔 ,
  • 韩斌 ,
  • 李超 ,
  • 李邵喜
展开
  • (1.大连海事大学 航海学院, 辽宁 大连 116026; 2.亿海蓝(北京)数据技术股份公司,北京 100089)
潘明阳(1975 — ),男,博士,副教授,E-mail:panmingyang@dlmu.edu.cn.

收稿日期: 2019-10-22

  修回日期: 2019-12-04

  网络出版日期: 2019-12-04

基金资助

国家自然科学基金资助项目(51579025);中央高校基本科研业务费专项资金资助项目(3132019400).

Intelligent translation technology of AIS ship name from Pinyin to Chinese

  • PAN Ming-yang ,
  • LI Qi ,
  • SHENG Zun-kuo ,
  • HAN Bin ,
  • LI Chao ,
  • LI Shao-xi
Expand
  • (1. Navigation College, Dalian Maritime University, Dalian 116026, China;2.Elane Inc, Beijing 100089, China)

Received date: 2019-10-22

  Revised date: 2019-12-04

  Online published: 2019-12-04

Supported by

 

摘要

针对航海领域中船舶自动识别系统(AIS)无法利用中文给我国使用者带来的识别障碍问题,研究了AIS拼音信息到汉字的智能翻译技术. 在建立标准化汉字和拼音船名语料库的基础上,分别搭建了基于Seq2Seq和Transformer框架的智能船名翻译的深度学习模型. 通过在同一数据集上的性能对比分析,Transformer模型具有更好的效果.为弥补Transformer模型受语料库规模限制而带来的翻译损失,进一步研究了其与隐马尔科夫链(HMM)的联合翻译模型,最终,在测试集上达到了98.92 %的准确率,实现了对AIS拼音船名的精准匹配和合理翻译. 该模型同样适用于AIS中目的港等拼音信息到汉字的翻译,对于提升AIS信息使用者的体验具有实际应用价值.

本文引用格式

潘明阳 , 李琦 , 盛尊阔 , 韩斌 , 李超 , 李邵喜 . AIS拼音船名到汉字的智能翻译技术研究[J]. 大连海事大学学报, 2020 , 46(2) : 41 -48 . DOI: 10.16411/j.cnki.issn1006-7736.2020.02.006

Abstract

In order to solve the problem that Automatic Identification System (AIS) can't make use of Chinese in the field of navigation,which has brought identification problems to Chinese users, this paper studies the intelligent translation technology from AIS Pinyin information to Chinese characters. Based on the establishment of standardized Chinese characters and Pinyin ship name corpus, deep learning models of intelligent ship names translation based on Seq2Seq and Transformer frameworks were built respectively. Through the performance comparison analysis on the same dataset, it is found that the Transformer model has better effects. In order to compensate for the translation loss caused by the limitation of the corpus size of the Transformer model, the combined translation model of the Transformer model and the Hidden Markov Chain (HMM) was further studied. Finally, the accuracy of the test set is 98.92%, and the accurate matching and reasonable translation of the AIS Pinyin ship names are realized. The model is also applicable to the translation of other Pinyin information into Chinese characters in AIS, such as destination port. It has application value for improving the experience of AIS information users.

参考文献

[1] 张颖.船舶自动识别系统(AIS)接口数据的研究与应用[D]. 大连海事大学, 2009.
[2]于光峰.船载AIS信息采集与解码技术研究[J][J].电子技术与软件工程, 2013, (21):91-92
[3]Peter F Brown, Stephen A Della Pietra, Vincent J Della Pietra, and Robert L Mercer.The Mathematics of Statistical Machine Translation: Parameter Estimation[J].Computational Linguistics, 1993, 19(02):263-311
[4]Philipp Koehn, Franz J Och, and Daniel Marcu.Statistical Phrase-Based Translation[C]. Proceedings of the 2003 Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technology, 2003:48-54.
[5]Kyunghyun Cho, ?Bart van Merrienboer, ?Caglar Gulcehre, ?Dzmitry Bahdanau, ?Fethi Bougares, ?Holger Schwenk, ?Yoshua Bengio.Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation[C]. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2014:1724–1734.
[6]Ilya Sutskever, Oriol Vinyals, and Quoc V Le.Sequence to Sequence Learning with Neural Networks[C].NIPS' 14 Proceedings of the 27th International Conference on Neural Information Processing Systems, 2014, Volume 2: 3104-3112 .
[7]Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural Machine Translation by Jointly Learning to Align and Translate.[J].rXiv:1409.0473, 2014 ., [J]., :-
[8]Vaswani A, Shazeer N, Parmar N, et al.Attention Is All You Need[C]. 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, 2017.
[9]魏瑾.基于统计的汉英机器翻译技术的研究[D]. 国防科学技术大学, 2006.
[10]赵静.基于统计的汉英机器翻译技术的研究[J].电子设计工程, 2016, 24(21):69-71
[11]张胜刚.基于神经网络的维汉机器翻译研究[D]. 新疆大学, 2018.
[12]刘婉婉, 苏依拉, 乌尼尔, 仁庆道尔吉.基于的蒙汉机器翻译的研究[J].计算机工程与科学, 2018, 40(10):178-184
[13]巨同升.机器学习在汉字智能拼音输入中的应用[J][J].山东理工大学学报(自然科学版), 2005, (03):86-88
[14]Chung Euisok, Park, Jeon Gue.Sentence-Chain Based Seq2seq Model for Corpus Expansion[J]. [J].ETRI JOURNAL, 2017, (39):455-466
[15]Chorowski Jan, Jaitly Navdeep.Towards better decoding and language model integration in sequence to sequence models[C]. 18th Annual Conference of the International-Speech-Communication-Association (INTERSPEECH 2017). Stockholm, SWEDEN, AUG 20-24, 2017:523-527.
[16]Kou Tanaka, Hirokazu Kameoka, Takuhiro Kaneko, Nobukatsu Hojo.Sequence-to-Sequence Voice Conversion with Attention and Context Preservation Mechanisms[J]. [J].AttS2S-VC, 2018, :-
[17]Wang Zuoying, Sun Jian.The Inhomogeneous HMM with General Topological Structure and Its Application in Language Identification between Mandarin and English[J]. [J].Journal of Electronics & Information Technology, 2007, :867-869
[18]Eddy SR.Profile hidden Markov models[J]. [J].BIOINFORMATICS, , 1998, :755-763
文章导航

/