Handling named entities and compound verbs in phrase-based statistical machine translation

Authors :: Pal, Santanu
Kumar Naskar, Sudip
Pecina, Pavel
Bandyopadhyay, Sivaji
Way, Andy
Pal, Santanu
Kumar Naskar, Sudip
Pecina, Pavel
Bandyopadhyay, Sivaji
Way, Andy
Publication Year :: 2010
Abstract: Data preprocessing plays a crucial role in phrase-based statistical machine translation (PB-SMT). In this paper, we show how single-tokenization of two types of multi-word expressions (MWE), namely named entities (NE) and compound verbs, as well as their prior alignment can boost the performance of PB-SMT. Single-tokenization of compound verbs and named entities (NE) provides significant gains over the baseline PB-SMT system. Automatic alignment of NEs substantially improves the overall MT performance, and thereby the word alignment quality indirectly. For establishing NE alignments, we transliterate source NEs into the target language and then compare them with the target NEs. Target language NEs are first converted into a canonical form before the comparison takes place. Our best system achieves statistically significant improvements (4.59 BLEU points absolute, 52.5% relative improvement) on an English—Bangla translation task.

Tools