On the Usage of Semantics, Syntax, and Morphology for Noun Classification in IsiZulu

Sayed, Imaan; Mahlaza, Zola; van der Leek, Alexander; Mopp, Jonathan; Keet, C. Maria

On the Usage of Semantics, Syntax, and Morphology for Noun Classification in IsiZulu

Failid

2025_resourceful_1_23.pdf (212.19 KB)

Kuupäev

2025-03

Autorid

Sayed, Imaan

Mahlaza, Zola

van der Leek, Alexander

Mopp, Jonathan

Keet, C. Maria

Kirjastaja

University of Tartu Library

Abstrakt

There is limited work aimed at solving the core task of noun classification for Nguni languages. The task focuses on identifying the semantic categorisation of each noun and plays a crucial role in the ability to form semantically and morphologically valid sentences. The work by Byamugisha (2022) was the first to tackle the problem for a related, but non-Nguni, language. While there have been efforts to replicate it for a Nguni language, there has been no effort focused on comparing the technique used in the original work vs. contemporary neural methods or a number of traditional machine learning classification techniques that do not rely on human-guided knowledge to the same extent. We reproduce Byamugisha (2022)’s work with different configurations to account for differences in access to datasets and resources, compare the approach with a pre-trained transformer-based model, and traditional machine learning models that relyon less human-guided knowledge. The newly created data-driven models outperform the knowledge-infused models, with the best performing models achieving an F1 score of 0.97.

URI

https://aclanthology.org/2025.resourceful-1.0/
https://hdl.handle.net/10062/107121

Kollektsioonid

Proceedings of the Third Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2025)

Kirje täielik lehekülg

On the Usage of Semantics, Syntax, and Morphology for Noun Classification in IsiZulu

Failid

Kuupäev

Autorid

Ajakirja pealkiri

Ajakirja ISSN

Köite pealkiri

Kirjastaja

Abstrakt

Kirjeldus

Märksõnad

Viide

URI

Kollektsioonid