Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–6 of 6 results for author: Aldabe, I

Searching in archive cs. Search in all archives.
.
  1. High-Order Question Generation in a Multilingual Educational Context

    Authors: Suna-Şeyma Uçar, Itziar Aldabe, Nora Aranberri, Orphée De Clercq

    Abstract: Critical thinking is a fundamental skill that helps learners move beyond simple memorization. One way to develop this skill is through high-order questioning. However, crafting such questions remains a challenge for educators, and classroom practices tend to rely on low-order questions. Large Language Models have demonstrated strong capabilities in generating high-order questions, especially when… ▽ More

    Submitted 15 July, 2026; originally announced July 2026.

    Comments: This paper was accepted at the 15th edition of the Language Resources and Evaluation Conference (LREC 2026)

    Journal ref: Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026), pp. 760-769

  2. arXiv:2506.07597  [pdf, ps, other

    cs.CL

    Instructing Large Language Models for Low-Resource Languages: A Systematic Study for Basque

    Authors: Oscar Sainz, Naiara Perez, Julen Etxaniz, Joseba Fernandez de Landa, Itziar Aldabe, Iker García-Ferrero, Aimar Zabala, Ekhi Azurmendi, German Rigau, Eneko Agirre, Mikel Artetxe, Aitor Soroa

    Abstract: Instructing language models with user intent requires large instruction datasets, which are only available for a limited set of languages. In this paper, we explore alternatives to conventional instruction adaptation pipelines in low-resource scenarios. We assume a realistic scenario for low-resource languages, where only the following are available: corpora in the target language, existing open-w… ▽ More

    Submitted 13 March, 2026; v1 submitted 9 June, 2025; originally announced June 2025.

    Comments: Accepted at EMNLP 2025 Main Conference

  3. arXiv:2403.20266  [pdf, other

    cs.CL cs.AI cs.LG

    Latxa: An Open Language Model and Evaluation Suite for Basque

    Authors: Julen Etxaniz, Oscar Sainz, Naiara Perez, Itziar Aldabe, German Rigau, Eneko Agirre, Aitor Ormazabal, Mikel Artetxe, Aitor Soroa

    Abstract: We introduce Latxa, a family of large language models for Basque ranging from 7 to 70 billion parameters. Latxa is based on Llama 2, which we continue pretraining on a new Basque corpus comprising 4.3M documents and 4.2B tokens. Addressing the scarcity of high-quality benchmarks for Basque, we further introduce 4 multiple choice evaluation datasets: EusProficiency, comprising 5,169 questions from… ▽ More

    Submitted 20 September, 2024; v1 submitted 29 March, 2024; originally announced March 2024.

    Comments: ACL 2024

    Journal ref: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14952--14972. 2024

  4. arXiv:2203.08111  [pdf, other

    cs.CL cs.AI cs.LG

    Does Corpus Quality Really Matter for Low-Resource Languages?

    Authors: Mikel Artetxe, Itziar Aldabe, Rodrigo Agerri, Olatz Perez-de-Viñaspre, Aitor Soroa

    Abstract: The vast majority of non-English corpora are derived from automatically filtered versions of CommonCrawl. While prior work has identified major issues on the quality of these datasets (Kreutzer et al., 2021), it is not clear how this impacts downstream performance. Taking representation learning in Basque as a case study, we explore tailored crawling (manually identifying and scraping websites wit… ▽ More

    Submitted 26 October, 2022; v1 submitted 15 March, 2022; originally announced March 2022.

    Comments: EMNLP 2022

  5. arXiv:1702.00700  [pdf, ps, other

    cs.CL cs.AI

    Multilingual and Cross-lingual Timeline Extraction

    Authors: Egoitz Laparra, Rodrigo Agerri, Itziar Aldabe, German Rigau

    Abstract: In this paper we present an approach to extract ordered timelines of events, their participants, locations and times from a set of multilingual and cross-lingual data sources. Based on the assumption that event-related information can be recovered from different documents written in different languages, we extend the Cross-document Event Ordering task presented at SemEval 2015 by specifying two ne… ▽ More

    Submitted 2 February, 2017; originally announced February 2017.

    Comments: 20 pages, 7 tables, 7 figures; submitted to Knowledge Based Systems (Elsevier), January, 2017

  6. arXiv:1507.03462  [pdf, other

    cs.CL

    Supervised Hierarchical Classification for Student Answer Scoring

    Authors: Itziar Aldabe, Oier Lopez de Lacalle, Iñigo Lopez-Gazpio, Montse Maritxalar

    Abstract: This paper describes a hierarchical system that predicts one label at a time for automated student response analysis. For the task, we build a classification binary tree that delays more easily confused labels to later stages using hierarchical processes. In particular, the paper describes how the hierarchical classifier has been built and how the classification task has been broken down into bina… ▽ More

    Submitted 13 July, 2015; originally announced July 2015.

    Comments: 5 pages with references