Summary of the paper

Title Work on Spoken (Multimodal) Language Corpora in South Africa
Authors Jens Allwood, Harald Hammarström, Andries Hendrikse, Mtholeni N. Ngcobo, Nozibele Nomdebevana, Laurette Pretorius and Mac van der Merwe
Abstract This paper describes past, ongoing and planned work on the collection and transcription of spoken language samples for all the South African official languages and as part of this the training of researchers in corpus linguistic research skills. More specifically the work has involved (and still involves) establishing an international corpus linguistic network linked to a network hub at a UNISA website and the development of research tools, a corpus research guide and workbook for multimodal communication and spoken language corpus research. As an example of the work we are doing and hope to do more of in the future, we present a small pilot study of the influence of English and Afrikaans on the 100 most frequent words in spoken Xhosa as this is evidenced in the corpus of spoken interaction we have gathered so far. Other planned work, besides work on spoken language phenomena, involves comparison of spoken and written language and work on communicative body movements (gestures) and their relation to speech.
Topics Corpus (creation, annotation, etc.), Discourse annotation, representation and processing, Multilinguality
Full paper Work on Spoken (Multimodal) Language Corpora in South Africa
Slides -
Bibtex @InProceedings{ALLWOOD10.438,
  author = {Jens Allwood and Harald Hammarström and Andries Hendrikse and Mtholeni N. Ngcobo and Nozibele Nomdebevana and Laurette Pretorius and Mac van der Merwe},
  title = {Work on Spoken (Multimodal) Language Corpora in South Africa},
  booktitle = {Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)},
  year = {2010},
  month = {may},
  date = {19-21},
  address = {Valletta, Malta},
  editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Bente Maegaard and Joseph Mariani and Jan Odijk and Stelios Piperidis and Mike Rosner and Daniel Tapias},
  publisher = {European Language Resources Association (ELRA)},
  isbn = {2-9517408-6-7},
  language = {english}
Powered by ELDA © 2010 ELDA/ELRA