Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–8 of 8 results for author: Frenda, S

Searching in archive cs. Search in all archives.
.
  1. arXiv:2606.02911   

    cs.CL

    The Ghost Annotator: a Framework to Explore Human Label Variation in Content Moderation through Conformal Prediction

    Authors: Mirko Lai, Alessandra Urbinati, Simona Frenda, Fabiana Vernero, Marco Antonio Stranisci

    Abstract: Current research primarily focuses on model performance, while comparatively less attention has been devoted to uncertainty estimation, particularly in settings where LLMs are increasingly used to generate annotated data. We introduce a framework combining conformal prediction with Collaborative Filtering-style annotators' representation to model LLM behavior in relation to human annotators and to… ▽ More

    Submitted 15 July, 2026; v1 submitted 1 June, 2026; originally announced June 2026.

    Comments: The publishing of this preprint is contextual with the ACL ARR cycle system. After an encouraging review in January we revised and submit the paper on Arxiv. However, a new batch of reviewers raised additional issues that will lead to significant revisions of the experimental setting. Therefore, we decide to withdraw the manuscript

  2. arXiv:2512.04759  [pdf, ps, other

    cs.CL

    Challenging the Abilities of Large Language Models in Italian: a Community Initiative

    Authors: Malvina Nissim, Danilo Croce, Viviana Patti, Pierpaolo Basile, Giuseppe Attanasio, Elio Musacchio, Matteo Rinaldi, Federico Borazio, Maria Francis, Jacopo Gili, Daniel Scalena, Begoña Altuna, Ekhi Azurmendi, Valerio Basile, Luisa Bentivogli, Arianna Bisazza, Marianna Bolognesi, Dominique Brunato, Tommaso Caselli, Silvia Casola, Maria Cassese, Mauro Cettolo, Claudia Collacciani, Leonardo De Cosmo, Maria Pia Di Buono , et al. (56 additional authors not shown)

    Abstract: The rapid progress of Large Language Models (LLMs) has transformed natural language processing and broadened its impact across research and society. Yet, systematic evaluation of these models, especially for languages beyond English, remains limited. "Challenging the Abilities of LAnguage Models in ITAlian" (CALAMITA) is a large-scale collaborative benchmarking initiative for Italian, coordinated… ▽ More

    Submitted 4 December, 2025; originally announced December 2025.

  3. Are you sure? Measuring models bias in content moderation through uncertainty

    Authors: Alessandra Urbinati, Mirko Lai, Simona Frenda, Marco Antonio Stranisci

    Abstract: Automatic content moderation is crucial to ensuring safety in social media. Language Model-based classifiers are being increasingly adopted for this task, but it has been shown that they perpetuate racial and social biases. Even if several resources and benchmark corpora have been developed to challenge this issue, measuring the fairness of models in content moderation remains an open issue. In th… ▽ More

    Submitted 28 October, 2025; v1 submitted 21 September, 2025; originally announced September 2025.

    Comments: accepted at Findings of ACL: EMNLP 2025

  4. arXiv:2508.04638  [pdf, ps, other

    cs.CL

    Can NLP Tackle Hate Speech in the Real World? Stakeholder-Informed Feedback and Survey on Counterspeech

    Authors: Tanvi Dinkar, Aiqi Jiang, Simona Frenda, Poppy Gerrard-Abbott, Nancie Gunson, Gavin Abercrombie, Ioannis Konstas

    Abstract: Counterspeech, i.e. the practice of responding to online hate speech, has gained traction in NLP as a promising intervention. While early work emphasised collaboration with non-governmental organisation stakeholders, recent research trends have shifted toward automated pipelines that reuse a small set of legacy datasets, often without input from affected communities. This paper presents a systemat… ▽ More

    Submitted 6 August, 2025; originally announced August 2025.

  5. arXiv:2505.20624  [pdf, ps, other

    cs.CL

    POLAR: A Benchmark for Multilingual, Multicultural, and Multi-Event Online Polarization

    Authors: Usman Naseem, Robert Geislinger, Juan Ren, Sarah Kohail, Rudy Garrido Veliz, P Sam Sahil, Yiran Zhang, Marco Antonio Stranisci, Idris Abdulmumin, Özge Alacam, Cengiz Acartürk, Aisha Jabr, Saba Anwar, Abinew Ali Ayele, Simona Frenda, Alessandra Teresa Cignarella, Elena Tutubalina, Oleg Rogov, Aung Kyaw Htet, Xintong Wang, Surendrabikram Thapa, Kritesh Rauniyar, Tanmoy Chakraborty, Arfeen Zeeshan, Dheeraj Kodati , et al. (18 additional authors not shown)

    Abstract: Online polarization poses a growing challenge for democratic discourse, yet most computational social science research remains monolingual, culturally narrow, or event-specific. We introduce POLAR, a multilingual, multicultural, and multi-event dataset with over 110K instances in 22 languages drawn from diverse online platforms and real-world events. Polarization is annotated along three axes, nam… ▽ More

    Submitted 5 February, 2026; v1 submitted 26 May, 2025; originally announced May 2025.

    Comments: Preprint

  6. arXiv:2412.19168  [pdf, other

    cs.CL

    GFG -- Gender-Fair Generation: A CALAMITA Challenge

    Authors: Simona Frenda, Andrea Piergentili, Beatrice Savoldi, Marco Madeddu, Martina Rosola, Silvia Casola, Chiara Ferrando, Viviana Patti, Matteo Negri, Luisa Bentivogli

    Abstract: Gender-fair language aims at promoting gender equality by using terms and expressions that include all identities and avoid reinforcing gender stereotypes. Implementing gender-fair strategies is particularly challenging in heavily gender-marked languages, such as Italian. To address this, the Gender-Fair Generation challenge intends to help shift toward gender-fair language in written communicatio… ▽ More

    Submitted 30 December, 2024; v1 submitted 26 December, 2024; originally announced December 2024.

    Comments: To refer to this paper please cite the CEUR-ws publication available at https://ceur-ws.org/Vol-3878/

  7. arXiv:2207.10652  [pdf, other

    cs.CL

    O-Dang! The Ontology of Dangerous Speech Messages

    Authors: Marco A. Stranisci, Simona Frenda, Mirko Lai, Oscar Araque, Alessandra T. Cignarella, Valerio Basile, Viviana Patti, Cristina Bosco

    Abstract: Inside the NLP community there is a considerable amount of language resources created, annotated and released every day with the aim of studying specific linguistic phenomena. Despite a variety of attempts in order to organize such resources has been carried on, a lack of systematic methods and of possible interoperability between resources are still present. Furthermore, when storing linguistic i… ▽ More

    Submitted 13 July, 2022; originally announced July 2022.

  8. arXiv:2205.15627  [pdf, other

    cs.CL

    APPReddit: a Corpus of Reddit Posts Annotated for Appraisal

    Authors: Marco Antonio Stranisci, Simona Frenda, Eleonora Ceccaldi, Valerio Basile, Rossana Damiano, Viviana Patti

    Abstract: Despite the large number of computational resources for emotion recognition, there is a lack of data sets relying on appraisal models. According to Appraisal theories, emotions are the outcome of a multi-dimensional evaluation of events. In this paper, we present APPReddit, the first corpus of non-experimental data annotated according to this theory. After describing its development, we compare ou… ▽ More

    Submitted 31 May, 2022; originally announced May 2022.