Reusing monolingual pre-trained models by cross-connecting seq2seq models for machine translation

Oh, Jiun; Choi, Yong-Suk

doi:10.3390/app11188737

Detailed Information

Cited 1 time in webofscience

Cited 1 time in scopus

Metadata Downloads

Reusing monolingual pre-trained models by cross-connecting seq2seq models for machine translationopen access

Authors: Oh, Jiun; Choi, Yong-Suk

Issue Date: Sep-2021

Publisher: MDPI

Keywords: natural language processing; transfer learning; neural machine translation

Citation: Applied Sciences-basel, v.11, no.18, pp 1 - 13

Pages: 13

Indexed: SCIE
SCOPUS

Journal Title: Applied Sciences-basel

Volume: 11

Number: 18

Start Page: 1

End Page: 13

URI: https://scholarworks.bwise.kr/hanyang/handle/2021.sw.hanyang/141039

DOI: 10.3390/app11188737

ISSN: 2076-3417

Abstract: This work uses sequence-to-sequence (seq2seq) models pre-trained on monolingual corpora for machine translation. We pre-train two seq2seq models with monolingual corpora for the source and target languages, then combine the encoder of the source language model and the decoder of the target language model, i.e., the cross-connection. We add an intermediate layer between the pre-trained encoder and the decoder to help the mapping of each other since the modules are pre-trained completely independently. These monolingual pre-trained models can work as a multilingual pre-trained model because one model can be cross-connected with another model pre-trained on any other language, while their capacity is not affected by the number of languages. We will demonstrate that our method improves the translation performance significantly over the random baseline. Moreover, we will analyze the appropriate choice of the intermediate layer, the importance of each part of a pre-trained model, and the performance change along with the size of the bitext.

Files in This Item

applsci-11-08737-v2.pdf 525.78 kB

Appears in Collections: 서울 공과대학 > 서울 컴퓨터소프트웨어학부 > 1. Journal Articles

Show full item record

qrcode

Related Researcher

Researcher Choi, Yong Suk photo

Choi, Yong Suk: COLLEGE OF ENGINEERING (SCHOOL OF COMPUTER SCIENCE)

Read more

Altmetrics

Total Views & Downloads

RSS_1.0 RSS_2.0 ATOM_1.0

222, Wangsimni-ro, Seongdong-gu, Seoul, 04763, Korea+82-2-2220-1366

Certain data included herein are derived from the © Web of Science of Clarivate Analytics. All rights reserved.
You may not copy or re-distribute this material in whole or in part without the prior written consent of Clarivate Analytics.

Detailed Information

Related Researcher

Altmetrics

Total Views & Downloads

BROWSE