Summary of the paper

Title Phonetically Balanced Code-Mixed Speech Corpus for Hindi-English Automatic Speech Recognition
Authors Ayushi Pandey, Brij Mohan Lal Srivastava, Rohit Kumar, Bhanu Teja Nellore, Kasi Sai Teja and Suryakanth V Gangashetty
Abstract The paper presents the development of a phonetically balanced read speech corpus of code-mixed Hindi-English. Phonetic balance in the corpus has been created by selecting sentences that contained triphones lower in frequency than a predefined threshold. The assumption with a compulsory inclusion of such rare units was that the high frequency triphones will inevitably be included. Using this metric, the Pearson's correlation coefficient of the phonetically balanced corpus with a large code-mixed reference corpus was recorded to be 0.996. The data for corpus creation has been extracted from selected sections of Hindi newspapers.These sections contain frequent English insertions in a matrix of Hindi sentence. Statistics on the phone and triphone distribution have been presented, to graphically display the phonetic likeness between the reference corpus and the corpus sampled through our method.
Topics Speech Resource/Database, Multilinguality, Speech Recognition/Understanding
Full paper Phonetically Balanced Code-Mixed Speech Corpus for Hindi-English Automatic Speech Recognition
Bibtex @InProceedings{PANDEY18.940,
  author = {Ayushi Pandey and Brij Mohan Lal Srivastava and Rohit Kumar and Bhanu Teja Nellore and Kasi Sai Teja and Suryakanth V Gangashetty},
  title = "{Phonetically Balanced Code-Mixed Speech Corpus for Hindi-English Automatic Speech Recognition}",
  booktitle = {Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018)},
  year = {2018},
  month = {May 7-12, 2018},
  address = {Miyazaki, Japan},
  editor = {Nicoletta Calzolari (Conference chair) and Khalid Choukri and Christopher Cieri and Thierry Declerck and Sara Goggi and Koiti Hasida and Hitoshi Isahara and Bente Maegaard and Joseph Mariani and Hélène Mazo and Asuncion Moreno and Jan Odijk and Stelios Piperidis and Takenobu Tokunaga},
  publisher = {European Language Resources Association (ELRA)},
  isbn = {979-10-95546-00-9},
  language = {english}
  }
Powered by ELDA © 2018 ELDA/ELRA