Title The Italian NESPOLE! Corpus: A Multilingual Database with Interlingua Annotation in Tourism and Medical Domains
Author(s) Nadia Mana (1), Roldano Cattoni (1), Emanuele Pianta (1), Franca Rossi (1), Fabio Pianesi (1), Susanne Burger (2)

(1) ITC-irst, Centro per la Ricerca Scientifica e Tecnologica, Trento, Italy, {mana,cattoni,pianta,frarossi,pianesi}@itc.it; (2) Interactive Systems Laboratories, Carnegie Mellon University, Pittsburgh, USA, sburger@cs.cmu.edu

Abstract This paper presents the Italian NESPOLE! Database. The database consists of three parts: The first two, called DB-1 and DB-2 concern the tourism domain, while the third part, DB-3, concentrates on the medical domain. The database includes audio files, transcriptions, Interlingua annotations in IF (Interchange Format) and translations into English, French and German. We describe how the database was built (data collection set-up, scenarios, recording procedure, data transcription and annotation) and statistically illustrates the corpus by providing a data analysis focused on language and spontaneous phenomena.
Keyword(s) Spoken dialogues, multilingual, multimodal, interlingua annotations, tourism and medical domains
Language(s) Italian plus translations in English, French and German
