LREC 2014 Proceedings

Summary of the paper

Title	Dual Subtitles as Parallel Corpora
Authors	Shikun Zhang, Wang Ling and Chris Dyer
Abstract	In this paper, we leverage the existence of dual subtitles as a source of parallel data. Dual subtitles present viewers with two languages simultaneously, and are generally aligned in the segment level, which removes the need to automatically perform this alignment. This is desirable as extracted parallel data does not contain alignment errors present in previous work that aligns different subtitle files for the same movie. We present a simple heuristic to detect and extract dual subtitles and show that more than 20 million sentence pairs can be extracted for the Mandarin-English language pair. We also show that extracting data from this source can be a viable solution for improving Machine Translation systems in the domain of subtitles.
Topics	Machine Translation, SpeechToSpeech Translation, Text Mining
Full paper	Dual Subtitles as Parallel Corpora
Bibtex	@InProceedings{ZHANG14.1199, author = {Shikun Zhang and Wang Ling and Chris Dyer}, title = {Dual Subtitles as Parallel Corpora}, booktitle = {Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)}, year = {2014}, month = {may}, date = {26-31}, address = {Reykjavik, Iceland}, editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Thierry Declerck and Hrafn Loftsson and Bente Maegaard and Joseph Mariani and Asuncion Moreno and Jan Odijk and Stelios Piperidis}, publisher = {European Language Resources Association (ELRA)}, isbn = {978-2-9517408-8-4}, language = {english} }