Author: "Chen, Hsin-Hsi" / Region: jizah (egypt) - Searchworks@Jio Institute Digital Library Search Results

1. Learning English–Chinese bilingual word representations from sentence-aligned parallel corpus.

Author: Yen, An-Zi, Huang, Hen-Hsen, and Chen, Hsin-Hsi
Subjects: *INTERNATIONAL alliances, *VOCABULARY
Abstract: Highlights • We investigate the approaches of learning bilingual word representations with and without word alignment comprehensively. • In the approaches without word alignment, we propose four types of methods to formulate the contexts for Skip-gram modelling. • In the approaches with word alignment, we induce word alignment links by word alignment tools to learn bilingual semantics. • We explore the impact of different word alignment tools and alignment directions in learning bilingual word representations. Abstract Representation of words in different languages is fundamental for various cross-lingual applications. In the past researches, there was an argument in using or not using word alignment in learning bilingual word representations. This paper presents a comprehensive empirical study on the uses of parallel corpus to learn the word representations in the embedding space. Various non-alignment and alignment approaches are explored to formulate the contexts for Skip-gram modeling. In the approaches without word alignment, concatenating A and B, concatenating B and A, interleaving A with B, shuffling A and B, and using A and B separately are considered, where A and B denote parallel sentences in two languages. In the approaches with word alignment, three word alignment tools, including GIZA++, TsinghuaAligner, and fast_align, are employed to align words in sentences A and B. The effects of alignment direction from A to B or from B to A are also discussed. To deal with the unaligned words in the word alignment approach, two alternatives, using the words aligned with their immediate neighbors and using the words in the interleaving approach, are explored. We evaluate the performance of the adopted approaches in four tasks, including bilingual dictionary induction, cross-lingual information retrieval, cross-lingual analogy reasoning, and cross-lingual word semantic relatedness. These tasks cover the issues of translation, reasoning, and information access. Experimental results show the word alignment approach with conditional interleaving achieves the best performance in most of the tasks. [ABSTRACT FROM AUTHOR]
Published: 2019
Full Text: View/download PDF

Searchworks

Select search scope, currently: Articles

Catalog

books, media & more in Jio Institute collections

Articles

journal articles & other e-resources

Refine your results

1 results on '"Chen, Hsin-Hsi"'

1. Learning English–Chinese bilingual word representations from sentence-aligned parallel corpus.

Catalog

Searchworks

Select search scope, currently: Articles Catalog books, media & more in Jio Institute collections Articles journal articles & other e-resources

Search

Search Constraints

Refine your results

Search Limiters

Publication Year Range

Language

Publication Type

Database

1 results on '"Chen, Hsin-Hsi"'

Search Results

Catalog

Select search scope, currently: Articles

Catalog

books, media & more in Jio Institute collections

Articles

journal articles & other e-resources