WikiQA-IS: Assisted Benchmark Generation and Automated Evaluation of Icelandic Cultural Knowledge in LLMs

Arnardóttir, ÞórunnEinarsson, Elías BjarturIngvarsson Juto, GarðarHelgason, Þorvaldur PállEinarsson, HafsteinnTudor, Crina MadalinaDebess, Iben NyholmBruton, MicaellaScalvini, BarbaraIlinykh, NikolaiHoldt, Špela Arhar2025-02-142025-02-142025-03https://aclanthology.org/2025.resourceful-1.0/https://hdl.handle.net/10062/107117This paper presents WikiQA-IS, a novel question-answering dataset focusing on Icelandic culture and history, along with an automated pipeline for dataset generation and evaluation. Leveraging GPT-4 to create questions and answers based on Icelandic Wikipedia articles and news sources, we produced a high-quality corpus of 2,000 question-answer pairs. We introduce an automatic evaluation method using GPT-4o as a judge, which shows strong agreement with human evaluations. Our benchmark reveals varying performances across different language models, with closed-source models generally outperforming open-weights alternatives. This work contributes a resource for evaluating language models' knowledge of Icelandic culture and offers a replicable framework for creating similar datasets in other cultural contexts.enAttribution-NonCommercial-NoDerivatives 4.0 Internationalhttps://creativecommons.org/licenses/by-nc-nd/4.0/WikiQA-IS: Assisted Benchmark Generation and Automated Evaluation of Icelandic Cultural Knowledge in LLMsArticle