Topic modelling to support English text selection for translation into South Africa's other official languages
DOI:
https://doi.org/10.55492/dhasa.v4i01.4447Keywords:
DHASA, topic modelling, text selection, translation, under-resourced languagesAbstract
Appropriate training data is a prerequisite for the development of natural language processing (NLP) techniques. Vast amounts of language data are typically required to develop NLP tools that perform at state-of-the-art level. Such abundant resources are currently only available in a few languages. The remaining languages have to find alternative ways to become ``NLP-enabled''. The aim of the study reported on here is to make more language data available to support NLP development in the official languages of South Africa. In this paper we present the idea of generating text data by means of translation. We also propose the use of topic modelling to identify text in a highly resourced source language that will yield meaningful translations in under-resourced target languages. More specifically, the paper describes how topic modelling was used to identify English Wikipedia articles that should be suitable for translation into South Africa's 10 other official languages.
Downloads
Published
Issue
Section
License
Copyright (c) 2023 Jocelyn Mazarura, Febe de Wet
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.