Summary of the paper

Title Automatic Corpus Extension for Data-driven Natural Language Generation
Authors Elena Manishina, Bassam Jabaian, Stéphane Huet and Fabrice Lefevre
Abstract As data-driven approaches started to make their way into the Natural Language Generation (NLG) domain, the need for automation of corpus building and extension became apparent. Corpus creation and extension in data-driven NLG domain traditionally involved manual paraphrasing performed by either a group of experts or with resort to crowd-sourcing. Building the training corpora manually is a costly enterprise which requires a lot of time and human resources. We propose to automate the process of corpus extension by integrating automatically obtained synonyms and paraphrases. Our methodology allowed us to significantly increase the size of the training corpus and its level of variability (the number of distinct tokens and specific syntactic structures). Our extension solutions are fully automatic and require only some initial validation. The human evaluation results confirm that in many cases native speakers favor the outputs of the model built on the extended corpus.
Topics Corpus (Creation, Annotation, etc.), Natural Language Generation, Usability, User Satisfaction
Full paper Automatic Corpus Extension for Data-driven Natural Language Generation
Bibtex @InProceedings{MANISHINA16.571,
  author = {Elena Manishina and Bassam Jabaian and Stéphane Huet and Fabrice Lefevre},
  title = {Automatic Corpus Extension for Data-driven Natural Language Generation},
  booktitle = {Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016)},
  year = {2016},
  month = {may},
  date = {23-28},
  location = {Portoro┼ż, Slovenia},
  editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Thierry Declerck and Sara Goggi and Marko Grobelnik and Bente Maegaard and Joseph Mariani and Helene Mazo and Asuncion Moreno and Jan Odijk and Stelios Piperidis},
  publisher = {European Language Resources Association (ELRA)},
  address = {Paris, France},
  isbn = {978-2-9517408-9-1},
  language = {english}
Powered by ELDA © 2016 ELDA/ELRA