Summary of the paper

Title New Bilingual Speech Databases for Audio Diarization
Authors David Tavarez, Eva Navas, Daniel Erro, Ibon Saratxaga and Inma Hernaez
Abstract This paper describes the process of collecting and recording two new bilingual speech databases in Spanish and Basque. They are designed primarily for speaker diarization in two different application domains: broadcast news audio and recorded meetings. First, both databases have been manually segmented. Next, several diarization experiments have been carried out in order to evaluate them. Our baseline speaker diarization system has been applied to both databases with around 30% of DER for broadcast news audio and 40% of DER for recorded meetings. Also, the behavior of the system when different languages are used by the same speaker has been tested.
Topics Corpus (Creation, Annotation, etc.), Other
Full paper New Bilingual Speech Databases for Audio Diarization
Bibtex @InProceedings{TAVAREZ14.799,
  author = {David Tavarez and Eva Navas and Daniel Erro and Ibon Saratxaga and Inma Hernaez},
  title = {New Bilingual Speech Databases for Audio Diarization},
  booktitle = {Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)},
  year = {2014},
  month = {may},
  date = {26-31},
  address = {Reykjavik, Iceland},
  editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Thierry Declerck and Hrafn Loftsson and Bente Maegaard and Joseph Mariani and Asuncion Moreno and Jan Odijk and Stelios Piperidis},
  publisher = {European Language Resources Association (ELRA)},
  isbn = {978-2-9517408-8-4},
  language = {english}
Powered by ELDA © 2014 ELDA/ELRA