Comparing Human and Machine Translations of Generative Language Model Evaluation Datasets

de Vroe, Sander Bijl; Stampoulidis, George; Hakala, Kai; Rouhe, Aku; van Heeswijk, Mark; Karlgren, Jussi

Comparing Human and Machine Translations of Generative Language Model Evaluation Datasets

Failid

2025_nodalida_1_9.pdf (128.31 KB)

Kuupäev

2025-03

Autorid

Kirjastaja

University of Tartu Library

Abstrakt

The evaluation of Large Language Models (LLMs) is one of the crucial current challenges in the field of Natural Language Processing (NLP) and becomes even more challenging in the multilingual setting. Since the majority of the community's benchmarks exist only in English, test sets are now being machine translated at scale into dozens of languages. This work explores the feasibility of that approach, comparing a Finnish machine translation (MT) of ARC-Challenge with a new human translated version. Our findings suggest that since absolute scores are fairly close and model size rankings are preserved, machine translation is adequate in this case. Surprisingly, however, the datasets reverse the order of base models compared to their chat-finetuned counterparts.

URI

https://hdl.handle.net/10062/107200

Kollektsioonid

Proceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies (NoDaLiDa/Baltic-HLT 2025)

Kirje täielik lehekülg

Comparing Human and Machine Translations of Generative Language Model Evaluation Datasets

Failid

Kuupäev

Autorid

Ajakirja pealkiri

Ajakirja ISSN

Köite pealkiri

Kirjastaja

Abstrakt

Kirjeldus

Märksõnad

Viide

URI

Kollektsioonid