LREC 2016 Proceedings

Summary of the paper

Title	The United Nations Parallel Corpus v1.0
Authors	Michał Ziemski, Marcin Junczys-Dowmunt and Bruno Pouliquen
Abstract	This paper describes the creation process and statistics of the official United Nations Parallel Corpus, the first parallel corpus composed from United Nations documents published by the original data creator. The parallel corpus presented consists of manually translated UN documents from the last 25 years (1990 to 2014) for the six official UN languages, Arabic, Chinese, English, French, Russian, and Spanish. The corpus is freely available for download under a liberal license. Apart from the pairwise aligned documents, a fully aligned subcorpus for the six official UN languages is distributed. We provide baseline BLEU scores of our Moses-based SMT systems trained with the full data of language pairs involving English and for all possible translation directions of the six-way subcorpus.
Topics	Machine Translation, SpeechToSpeech Translation, Multilinguality, Metadata
Full paper	The United Nations Parallel Corpus v1.0
Bibtex	@InProceedings{ZIEMSKI16.1195, author = {Michał Ziemski and Marcin Junczys-Dowmunt and Bruno Pouliquen}, title = {The United Nations Parallel Corpus v1.0}, booktitle = {Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016)}, year = {2016}, month = {may}, date = {23-28}, location = {Portorož, Slovenia}, editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Thierry Declerck and Sara Goggi and Marko Grobelnik and Bente Maegaard and Joseph Mariani and Helene Mazo and Asuncion Moreno and Jan Odijk and Stelios Piperidis}, publisher = {European Language Resources Association (ELRA)}, address = {Paris, France}, isbn = {978-2-9517408-9-1}, language = {english} }