Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–13 of 13 results for author: Assenmacher, D

Searching in archive cs. Search in all archives.
.
  1. arXiv:2510.18582  [pdf, ps, other

    cs.CL

    Beyond the Explicit: A Bilingual Dataset for Dehumanization Detection in Social Media

    Authors: Dennis Assenmacher, Paloma Piot, Katarina Laken, David Jurgens, Claudia Wagner

    Abstract: Digital dehumanization, although a critical issue, remains largely overlooked within the field of computational linguistics and Natural Language Processing. The prevailing approach in current research concentrating primarily on a single aspect of dehumanization that identifies overtly negative statements as its core marker. This focus, while crucial for understanding harmful online communications,… ▽ More

    Submitted 21 October, 2025; originally announced October 2025.

  2. arXiv:2509.25063  [pdf, ps, other

    cs.CY cs.CL

    Learning from Convenience Samples: A Case Study on Fine-Tuning LLMs for Survey Non-response in the German Longitudinal Election Study

    Authors: Tobias Holtdirk, Dennis Assenmacher, Arnim Bleier, Claudia Wagner

    Abstract: Survey researchers face two key challenges: the rising costs of probability samples and missing data (e.g., non-response or attrition), which can undermine inference and increase the use of convenience samples. Recent work explores using large language models (LLMs) to simulate respondents via persona-based prompts, often without labeled data. We study a more practical setting where partial survey… ▽ More

    Submitted 29 September, 2025; originally announced September 2025.

  3. arXiv:2506.18576  [pdf, ps, other

    cs.CL cs.CY

    A Modular Taxonomy for Hate Speech Definitions and Its Impact on Zero-Shot LLM Classification Performance

    Authors: Matteo Melis, Gabriella Lapesa, Dennis Assenmacher

    Abstract: Detecting harmful content is a crucial task in the landscape of NLP applications for Social Good, with hate speech being one of its most dangerous forms. But what do we mean by hate speech, how can we define it, and how does prompting different definitions of hate speech affect model performance? The contribution of this work is twofold. At the theoretical level, we address the ambiguity surroundi… ▽ More

    Submitted 23 June, 2025; originally announced June 2025.

  4. arXiv:2410.11745  [pdf, other

    cs.CL cs.HC

    Personas with Attitudes: Controlling LLMs for Diverse Data Annotation

    Authors: Leon Fröhling, Gianluca Demartini, Dennis Assenmacher

    Abstract: We present a novel approach for enhancing diversity and control in data annotation tasks by personalizing large language models (LLMs). We investigate the impact of injecting diverse persona descriptions into LLM prompts across two studies, exploring whether personas increase annotation diversity and whether the impacts of individual personas on the resulting annotations are consistent and control… ▽ More

    Submitted 15 October, 2024; originally announced October 2024.

    Comments: 21 pages, 13 figures

  5. arXiv:2406.04892  [pdf, other

    cs.CL

    Sexism Detection on a Data Diet

    Authors: Rabiraj Bandyopadhyay, Dennis Assenmacher, Jose M. Alonso Moral, Claudia Wagner

    Abstract: There is an increase in the proliferation of online hate commensurate with the rise in the usage of social media. In response, there is also a significant advancement in the creation of automated tools aimed at identifying harmful text content using approaches grounded in Natural Language Processing and Deep Learning. Although it is known that training Deep Learning models require a substantial am… ▽ More

    Submitted 7 June, 2024; originally announced June 2024.

    Comments: Accepted at ACM WebSci 2024 Workshop in DHOW: Diffusion of Harmful Content on Online Web Workshop

  6. The Unseen Targets of Hate -- A Systematic Review of Hateful Communication Datasets

    Authors: Zehui Yu, Indira Sen, Dennis Assenmacher, Mattia Samory, Leon Fröhling, Christina Dahn, Debora Nozza, Claudia Wagner

    Abstract: Machine learning (ML)-based content moderation tools are essential to keep online spaces free from hateful communication. Yet, ML tools can only be as capable as the quality of the data they are trained on allows them. While there is increasing evidence that they underperform in detecting hateful communications directed towards specific identities and may discriminate against them, we know surpris… ▽ More

    Submitted 14 May, 2024; originally announced May 2024.

    Comments: 20 pages, 14 figures

  7. arXiv:2404.14244  [pdf, other

    cs.CR cs.AI cs.CY cs.LG cs.SI

    AI-Generated Faces in the Real World: A Large-Scale Case Study of Twitter Profile Images

    Authors: Jonas Ricker, Dennis Assenmacher, Thorsten Holz, Asja Fischer, Erwin Quiring

    Abstract: Recent advances in the field of generative artificial intelligence (AI) have blurred the lines between authentic and machine-generated content, making it almost impossible for humans to distinguish between such media. One notable consequence is the use of AI-generated images for fake profiles on social media. While several types of disinformation campaigns and similar incidents have been reported… ▽ More

    Submitted 3 October, 2024; v1 submitted 22 April, 2024; originally announced April 2024.

    Comments: International Symposium on Research in Attacks, Intrusions and Defenses (RAID), 2024

  8. arXiv:2311.01270  [pdf, other

    cs.CL cs.CY

    People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection

    Authors: Indira Sen, Dennis Assenmacher, Mattia Samory, Isabelle Augenstein, Wil van der Aalst, Claudia Wagner

    Abstract: NLP models are used in a variety of critical social computing tasks, such as detecting sexist, racist, or otherwise hateful content. Therefore, it is imperative that these models are robust to spurious features. Past work has attempted to tackle such spurious features using training data augmentation, including Counterfactually Augmented Data (CADs). CADs introduce minimal changes to existing trai… ▽ More

    Submitted 25 February, 2024; v1 submitted 2 November, 2023; originally announced November 2023.

    Comments: Preprint of EMNLP'23 paper

  9. arXiv:2302.00546  [pdf, other

    cs.SI cs.CY

    You are a Bot! -- Studying the Development of Bot Accusations on Twitter

    Authors: Dennis Assenmacher, Leon Fröhling, Claudia Wagner

    Abstract: The characterization and detection of bots with their presumed ability to manipulate society on social media platforms have been subject to many research endeavors over the last decade. In the absence of ground truth data (i.e., accounts that are labeled as bots by experts or self-declare their automated nature), researchers interested in the characterization and detection of bots may want to tap… ▽ More

    Submitted 31 March, 2024; v1 submitted 1 February, 2023; originally announced February 2023.

    Comments: 11 pages, 7 figures

  10. arXiv:2301.11429  [pdf, other

    cs.SI cs.CY

    Just Another Day on Twitter: A Complete 24 Hours of Twitter Data

    Authors: Juergen Pfeffer, Daniel Matter, Kokil Jaidka, Onur Varol, Afra Mashhadi, Jana Lasser, Dennis Assenmacher, Siqi Wu, Diyi Yang, Cornelia Brantner, Daniel M. Romero, Jahna Otterbacher, Carsten Schwemmer, Kenneth Joseph, David Garcia, Fred Morstatter

    Abstract: At the end of October 2022, Elon Musk concluded his acquisition of Twitter. In the weeks and months before that, several questions were publicly discussed that were not only of interest to the platform's future buyers, but also of high relevance to the Computational Social Science research community. For example, how many active users does the platform have? What percentage of accounts on the site… ▽ More

    Submitted 11 April, 2023; v1 submitted 26 January, 2023; originally announced January 2023.

  11. arXiv:2003.07595  [pdf, other

    cs.CY cs.HC

    FakeYou! -- A Gamified Approach for Building and Evaluating Resilience Against Fake News

    Authors: Lena Clever, Dennis Assenmacher, Kilian Müller, Moritz Vinzent Seiler, Dennis M. Riehle, Mike Preuss, Christian Grimme

    Abstract: Nowadays fake news are heavily discussed in public and political debates. Even though the phenomenon of intended false information is rather old, misinformation reaches a new level with the rise of the internet and participatory platforms. Due to Facebook and Co., purposeful false information - often called fake news - can be easily spread by everyone. Because of a high data volatility and variety… ▽ More

    Submitted 17 March, 2020; originally announced March 2020.

    Comments: accepted for Disinformation in Open Online Media - 2nd Multidisciplinary International Symposium, MISDOOM 2020

  12. Computational Methods in Professional Communication

    Authors: André Calero Valdez, Lena Adam, Dennis Assenmacher, Laura Burbach, Malte Bonart, Lena Frischlich, Philipp Schaer

    Abstract: The digitization of the world has also led to a digitization of communication processes. Traditional research methods fall short in understanding communication in digital worlds as the scope has become too large in volume, variety, and velocity to be studied using traditional approaches. In this paper, we present computational methods and their use in public and mass communication research and how… ▽ More

    Submitted 2 January, 2020; originally announced January 2020.

    Journal ref: 2019 IEEE International Professional Communication Conference (ProComm), Aachen, Germany, 2019, pp. 275-285

  13. arXiv:1902.06691  [pdf, other

    cs.CY cs.LG

    Openbots

    Authors: Dennis Assenmacher, Lena Adam, Lena Frischlich, Heike Trautmann, Christian Grimme

    Abstract: Social bots have recently gained attention in the context of public opinion manipulation on social media platforms. While a lot of research effort has been put into the classification and detection of such (semi-)automated programs, it is still unclear how sophisticated those bots actually are, which platforms they target, and where they originate from. To answer these questions, we gathered repos… ▽ More

    Submitted 19 February, 2019; v1 submitted 14 February, 2019; originally announced February 2019.

    Comments: Fixed typos