查看文章

arxiv.org 中的 [PDF]

Improving multilingual translation by representation and gradient regularization

作者

Yilin Yang, Akiko Eriguchi, Alexandre Muzio, Prasad Tadepalli, Stefan Lee, Hany Hassan

发表日期

2021/9/10

期刊

arXiv preprint arXiv:2109.04778

简介

Multilingual Neural Machine Translation (NMT) enables one model to serve all translation directions, including ones that are unseen during training, i.e. zero-shot translation. Despite being theoretically attractive, current models often produce low quality translations -- commonly failing to even produce outputs in the right target language. In this work, we observe that off-target translation is dominant even in strong multilingual systems, trained on massive multilingual corpora. To address this issue, we propose a joint approach to regularize NMT models at both representation-level and gradient-level. At the representation level, we leverage an auxiliary target language prediction task to regularize decoder outputs to retain information about the target language. At the gradient level, we leverage a small amount of direct data (in thousands of sentence pairs) to regularize model gradients. Our results demonstrate that our approach is highly effective in both reducing off-target translation occurrences and improving zero-shot translation performance by +5.59 and +10.38 BLEU on WMT and OPUS datasets respectively. Moreover, experiments show that our method also works well when the small amount of direct data is not available.

引用总数

被引用次数：29

20222023202412 11 6

学术搜索中的文章

Improving multilingual translation by representation and gradient regularization

Y Yang, A Eriguchi, A Muzio, P Tadepalli, S Lee… - arXiv preprint arXiv:2109.04778, 2021

被引用次数：29 相关文章所有 4 个版本