Summary of the paper

Title Corpus and Evaluation Measures for Automatic Plagiarism Detection
Authors Alberto Barrón-Cedeño, Martin Potthast, Paolo Rosso, Benno Stein and Andreas Eiselt
Abstract The simple access to texts on digital libraries and the World Wide Web has led to an increased number of plagiarism cases in recent years, which renders manual plagiarism detection infeasible at large. Various methods for automatic plagiarism detection have been developed whose objective is to assist human experts in the analysis of documents for plagiarism. The methods can be divided into two main approaches: intrinsic and external. Unlike other tasks in natural language processing and information retrieval, it is not possible to publish a collection of real plagiarism cases for evaluation purposes since they cannot be properly anonymized. Therefore, current evaluations found in the literature are incomparable and, very often not even reproducible. Our contribution in this respect is a newly developed large-scale corpus of artificial plagiarism useful for the evaluation of intrinsic as well as external plagiarism detection. Additionally, new detection performance measures tailored to the evaluation of plagiarism detection algorithms are proposed.
Topics Information Extraction, Information Retrieval, Corpus (creation, annotation, etc.), Evaluation methodologies
Full paper Corpus and Evaluation Measures for Automatic Plagiarism Detection
Slides Corpus and Evaluation Measures for Automatic Plagiarism Detection
Bibtex @InProceedings{BARRNCEDEO10.35,
  author = {Alberto Barrón-Cedeño and Martin Potthast and Paolo Rosso and Benno Stein and Andreas Eiselt},
  title = {Corpus and Evaluation Measures for Automatic Plagiarism Detection},
  booktitle = {Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)},
  year = {2010},
  month = {may},
  date = {19-21},
  address = {Valletta, Malta},
  editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Bente Maegaard and Joseph Mariani and Jan Odijk and Stelios Piperidis and Mike Rosner and Daniel Tapias},
  publisher = {European Language Resources Association (ELRA)},
  isbn = {2-9517408-6-7},
  language = {english}
Powered by ELDA © 2010 ELDA/ELRA