Summary of the paper

Title A Corpus for Evaluating Semantic Multilingual Web Retrieval Systems: The Sense Folder Corpus
Authors Ernesto William De Luca
Abstract In this paper, we present the multilingual Sense Folder Corpus. After the analysis of different corpora, we describe the requirements that have to be satisfied for evaluating semantic multilingual retrieval approaches. Justified by the unfulfilled requirements explained, we start creating a small bilingual hand-tagged corpus of 502 documents retrieved from Web searches. The documents contained in this collection have been created using Google queries. A single ambiguous word has been searched and related documents (approx. the first 60 documents for every keyword) have been retrieved. The document collection has been extended at the query word level, using single ambiguous words for English (argument, bank, chair, network and rule) and for Italian (argomento, lingua, regola, rete and stampa). The search and annotation process has been done both in a monolingual way for the English and the Italian language. 252 English and 250 Italian documents have been retrieved from Google and saved in their original rank. The performance of semantic multilingual retrieval systems has been evaluated using such a corpus with three baselines (“Random”, “First Sense” and “Most Frequent Sense”) that are formally presented and discussed. The fine-grained evaluation of the Sense Folder approach is discussed in details.
Topics Corpus (creation, annotation, etc.), Document Classification, Text categorisation, Information Extraction, Information Retrieval
Full paper A Corpus for Evaluating Semantic Multilingual Web Retrieval Systems: The Sense Folder Corpus
Slides -
Bibtex @InProceedings{DELUCA10.816,
  author = {Ernesto William De Luca},
  title = {A Corpus for Evaluating Semantic Multilingual Web Retrieval Systems: The Sense Folder Corpus},
  booktitle = {Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)},
  year = {2010},
  month = {may},
  date = {19-21},
  address = {Valletta, Malta},
  editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Bente Maegaard and Joseph Mariani and Jan Odijk and Stelios Piperidis and Mike Rosner and Daniel Tapias},
  publisher = {European Language Resources Association (ELRA)},
  isbn = {2-9517408-6-7},
  language = {english}
Powered by ELDA © 2010 ELDA/ELRA