Building Natural Language Processing Tools for Runyakitara
Abstract
This paper describes an endeavour to build natural language processing (NLP) tools for Runyakitara, a group of four closely related Bantu languages spoken in western Uganda. In contrast with major world languages such as English, for which corpora are comparatively abundant and NLP tools are well developed, computational linguistic resources for Runyakitara are in short supply. First therefore, we need to collect corpora for these languages, before we can proceed to the design of a spell-checker, grammar-checker and applications for computer-assisted language learning (CALL). We explain how we are collecting primary data for a new Runya Corpus of speech and writing, we outline the design of a morphological analyser, and discuss how we can use these new resources to build NLP tools. We are initially working with Runyankore-Rukiga, a closely-related pair of Runyakitara languages, and we frame our project in the context of NLP for low-resource languages, as well as CALL for the preservation of endangered languages. We put our project forward as a test case for the revitalization of endangered languages through education and technology. © 2020 Walter de Gruyter GmbH, Berlin/Boston.
Author(s)
Katushemererwe, F., Caines, A., Buttery, P.
Year
2020
Countries
Uganda
Language
English
Research Method
Quantitative
Article Type
Peer-Reviewed Articles
Keywords
Primary education | Language of instruction | Excel Import
Full Citation Style
Research Team
Katushemererwe, F.
Caines, A.
Buttery, P.
© 2026 AERD Database System. All Rights Reserved. Developed by Strategic Technologies.