Ar: Auto-repair the synthetic data for neural machine translation

S Cheng, S Kuang, R Weng, H Yu, C Zhu… - arXiv preprint arXiv …, 2020 - arxiv.org
arXiv preprint arXiv:2004.02196, 2020arxiv.org
Compared with only using limited authentic parallel data as training corpus, many studies
have proved that incorporating synthetic parallel data, which generated by back translation
(BT) or forward translation (FT, or selftraining), into the NMT training process can
significantly improve translation quality. However, as a well-known shortcoming, synthetic
parallel data is noisy because they are generated by an imperfect NMT system. As a result,
the improvements in translation quality bring by the synthetic parallel data are greatly …
Compared with only using limited authentic parallel data as training corpus, many studies have proved that incorporating synthetic parallel data, which generated by back translation (BT) or forward translation (FT, or selftraining), into the NMT training process can significantly improve translation quality. However, as a well-known shortcoming, synthetic parallel data is noisy because they are generated by an imperfect NMT system. As a result, the improvements in translation quality bring by the synthetic parallel data are greatly diminished. In this paper, we propose a novel Auto- Repair (AR) framework to improve the quality of synthetic data. Our proposed AR model can learn the transformation from low quality (noisy) input sentence to high quality sentence based on large scale monolingual data with BT and FT techniques. The noise in synthetic parallel data will be sufficiently eliminated by the proposed AR model and then the repaired synthetic parallel data can help the NMT models to achieve larger improvements. Experimental results show that our approach can effective improve the quality of synthetic parallel data and the NMT model with the repaired synthetic data achieves consistent improvements on both WMT14 EN!DE and IWSLT14 DE!EN translation tasks.
arxiv.org
以上显示的是最相近的搜索结果。 查看全部搜索结果

Google学术搜索按钮

example.edu/paper.pdf
搜索
获取 PDF 文件
引用
References