"I Need More Context and an English Translation": Analysing How LLMs Identify Personal Information in Komi, Polish, and English

Ilinykh, Nikolai; Szawerna, Maria Irena

"I Need More Context and an English Translation": Analysing How LLMs Identify Personal Information in Komi, Polish, and English

Failid

2025_resourceful_1_32.pdf (380.43 KB)

Kuupäev

2025-03

Autorid

Ilinykh, Nikolai

Szawerna, Maria Irena

Kirjastaja

University of Tartu Library

Abstrakt

Automatic identification of personal information (PI) is particularly difficult for languages with limited linguistic resources. Recently, large language models (LLMs) have been applied to various tasks involving low-resourced languages, but their capability to process PI in such contexts remains under-explored. In this paper we provide a qualitative analysis of the outputs from three LLMs prompted to identify PI in texts written in Komi (Permyak and Zyrian), Polish, and English. Our analysis highlights challenges in using pre-trained LLMs for PI identification in both low- and medium-resourced languages. It also motivates the need to develop LLMs that understand the differences in how PI is expressed across languages with varying levels of availability of linguistic resources.

URI

https://hdl.handle.net/10062/107129

Kollektsioonid

Proceedings of the Third Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2025)

Kirje täielik lehekülg

"I Need More Context and an English Translation": Analysing How LLMs Identify Personal Information in Komi, Polish, and English

Failid

Kuupäev

Autorid

Ajakirja pealkiri

Ajakirja ISSN

Köite pealkiri

Kirjastaja

Abstrakt

Kirjeldus

Märksõnad

Viide

URI

Kollektsioonid