Title Manual vs Assisted Transcription of Prepared and Spontaneous Speech
Authors Thierry Bazillon, Yannick Estève and Daniel Luzzati
Abstract Our paper focuses on the gain which can be achieved on human transcription of spontaneous and prepared speech, by using the assistance of an ASR system. This experiment has shown interesting results, first about the duration of the transcription task itself: even with the combination of prepared speech + ASR, an experimented annotator needs approximately 4 hours to transcribe 1 hours of audio data. Then, using an ASR system is mostly time-saving, although this gain is much more significant on prepared speech: assisted transcriptions are up to 4 times faster than manual ones. This ratio falls to 2 with spontaneous speech, because of ASR limits for these data. Detailed results reveal interesting correlations between the transcription task and phenomena such as Word Error Rate, telephonic or non-native speech turns, the number of fillers or propers nouns. The latter make spelling correction very time-consuming with prepared speech because of their frequency. As a consequence, watching for low averages of proper nouns may be a way to detect spontaneous speech.
Language Single language
Topics Corpus (creation, annotation, etc.), Speech resource/database, Speech recognition and understanding
Full paper Manual vs Assisted Transcription of Prepared and Spontaneous Speech
