Charles Explorer logo
🇬🇧

Genre Transfer in NMT: Creating Synthetic Spoken Parallel Sentences using Written Parallel Data

Publication at Faculty of Mathematics and Physics |
2023

Abstract

Text style transfer (TST) aims to control attributes in a given text without changing the content. The matter gets complicated when the boundary separating two styles gets blurred.

We can notice similar difficulties in the case of parallel datasets in spoken and written genres. Genuine spoken features like filler words and repetitions in the existing spoken genre parallel datasets are often cleaned during transcription and translation, making the texts closer to written datasets.

This poses several problems for spoken genre-specific tasks like simultaneous speech translation. This paper seeks to address the challenge of improving spoken language translations.

We start by creating a genre classifier for individual sentences and then try two approaches for data augmentation using written examples: (1) a novel method that involves assembling and disassembling spoken and written neural machine translation (NMT) models, and (2) a rule-based method to inject spoken features. Though the observed results for (1) are not promising, we get some interesting insights into the solution.

The model proposed in (1) fine-tuned on the synthesized data from (2) produces naturally looking spoken translations for writtenRIGHTWARDS ARROWspoken genre transfer in En-Hi translation systems. We use this system to produce a second-stage En-Hi synthetic corpus, which however lacks appropriate alignments of explicit spoken features across the languages.

For the final evaluation, we fine-tune Hi-En spoken translation systems on the synthesized parallel corpora. We observe that the parallel corpus synthesized using our rule-based method produces the best results.