LREC 2014 Proceedings

Summary of the paper

Title	Lexical Substitution Dataset for German
Authors	Kostadin Cholakov, Chris Biemann, Judith Eckle-Kohler and Iryna Gurevych
Abstract	This article describes a lexical substitution dataset for German. The whole dataset contains 2,040 sentences from the German Wikipedia, with one target word in each sentence. There are 51 target nouns, 51 adjectives, and 51 verbs randomly selected from 3 frequency groups based on the lemma frequency list of the German WaCKy corpus. 200 sentences have been annotated by 4 professional annotators and the remaining sentences by 1 professional annotator and 5 additional annotators who have been recruited via crowdsourcing. The resulting dataset can be used to evaluate not only lexical substitution systems, but also different sense inventories and word sense disambiguation systems.
Topics	Corpus (Creation, Annotation, etc.), Semantics
Full paper	Lexical Substitution Dataset for German
Bibtex	@InProceedings{CHOLAKOV14.545, author = {Kostadin Cholakov and Chris Biemann and Judith Eckle-Kohler and Iryna Gurevych}, title = {Lexical Substitution Dataset for German}, booktitle = {Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)}, year = {2014}, month = {may}, date = {26-31}, address = {Reykjavik, Iceland}, editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Thierry Declerck and Hrafn Loftsson and Bente Maegaard and Joseph Mariani and Asuncion Moreno and Jan Odijk and Stelios Piperidis}, publisher = {European Language Resources Association (ELRA)}, isbn = {978-2-9517408-8-4}, language = {english} }