Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 78 results for author: El-Assady, M

Searching in archive cs. Search in all archives.
.
  1. arXiv:2607.19386  [pdf, ps, other

    cs.LG cs.CL

    Building Fast, Evaluating Slow: Pipeline Choices Dominate Autointerpretability Score Variance

    Authors: Sinie van der Ben, Neele Roch, Anna Hedström, Mennatallah El-Assady

    Abstract: Cross-paper comparison of sparse autoencoder (SAE) interpretability often relies on autointerpretability scores. In this evaluation pipeline, a language model (LM) explains each feature, and another LM scores the explanation. For these comparisons to be meaningful, scores must reflect stable properties of the features rather than confounding aspects of the evaluation pipeline. Through systematic e… ▽ More

    Submitted 1 July, 2026; originally announced July 2026.

  2. arXiv:2607.07521  [pdf, ps, other

    cs.HC cs.AI

    Creativity from Friction: Human-AI Interaction for Exploratory Structural Design

    Authors: Ricardo Maia Avelino, Rita Sevastjanova, Tom Van Mele, Philippe Block, Mennatallah El-Assady

    Abstract: AI agents that generate final answers based on user input often do not meet the needs of creative fields. Fields such as structural design and architecture need interactive systems that help users externalise and develop ideas, explore alternatives, and refine partial solutions. The final product of such designs needs to comply with many constraints concerning, e.g., spatial configuration, mechani… ▽ More

    Submitted 8 July, 2026; originally announced July 2026.

    Comments: Accepted at ICML 2026, Workshop on Human-AI Co-Creativity

  3. arXiv:2607.05655  [pdf, ps, other

    cs.HC

    GeoXplain: On-the-Fly Visual Explanations for Weather Foundation Models

    Authors: Clemens Walter Koprolin, Leonardo Trentini, Benedikt Soja, Mennatallah El-Assady, Christina Humer

    Abstract: Weather and climate foundation models produce high-dimensional forecasts whose learned relationships are difficult to inspect with static plots alone. GeoXplain is an interactive Python-based visualization toolkit for exploring geospatial attribution maps across climate variables, atmospheric pressure levels, and forecast time. The toolkit accepts attribution bundles containing attribution grids t… ▽ More

    Submitted 6 July, 2026; originally announced July 2026.

    Comments: 9 pages, 6 figures. Submitted to VISxClimate at IEEE VIS 2026

  4. arXiv:2607.04501  [pdf, ps, other

    cs.HC cs.LG

    From Interaction to Intent: Inferring User Objectives from Provenance Logs

    Authors: Steffen Holter, Tobias Stähle, Arpit Narechania, Mennatallah El-Assady

    Abstract: The ability to automatically infer analytic intent from user interaction histories could enable interactive AI systems to proactively assist users during exploratory data analysis. In this paper, we examine whether provenance logs -- detailed records capturing sequences and timing of user interactions -- can be used to classify user intentions in visual exploration tasks. To investigate this, we r… ▽ More

    Submitted 5 July, 2026; originally announced July 2026.

    Comments: 13 pages, 11 figures

  5. arXiv:2606.26987  [pdf, ps, other

    cs.CL cs.AI

    Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs

    Authors: Sinie van der Ben, Raphaël Baur, Yannick Metz, Mennatallah El-Assady

    Abstract: Recent work identified emotion vectors in Claude Sonnet 4.5, which are internal representations that encode emotion concepts, causally influence behavior, and exhibit geometry mirroring human psychological structure. We test the generality of these findings in two open-weight models, Apertus-8B-Instruct-2509 and Gemma-4-E4B-it, extracting emotion contrast vectors across all layers, using two model… ▽ More

    Submitted 25 June, 2026; originally announced June 2026.

  6. arXiv:2606.19344  [pdf, ps, other

    cs.CL cs.AI

    Exposing the Unsaid: Visualizing Hidden LLM Bias through Stochastic Path Aggregation

    Authors: Matteo Pelossi, Rita Sevastjanova, Thilo Spinner, Mennatallah El-Assady

    Abstract: Large Language Models (LLMs) exhibit representational and syntactic biases that are difficult to evaluate due to the stochastic nature of text generation. Standard auditing methods rely on a single output inspection or static automated metrics. These approaches obscure the underlying probability distributions and fail to capture biases hidden in lower-probability generation branches. This paper in… ▽ More

    Submitted 24 April, 2026; originally announced June 2026.

    Comments: 14 pages

  7. Do we have the knowledge we need? Rethinking human-AI decision-making in corporations

    Authors: Anne S. R. Marx, Ricardo M. Avelino, Torbjørn Netland, Mennatallah El-Assady

    Abstract: Organizational knowledge is fragmented across a variety of software systems, tacit expertise, and manual documents that have traditionally been designed for human consumption. As AI systems are increasingly deployed and granted decision-making roles, they require access to this knowledge. This raises two questions: how should organizations store and maintain knowledge so that it remains accessible… ▽ More

    Submitted 2 April, 2026; originally announced June 2026.

    Comments: Proceedings of AutomationXP26 Workshop of the 2026 CHI Conference on Human Factors in Computing Systems, April 14, 2026, Barcelona, Spain. ACM, New York, NY, USA, 8 pages

  8. arXiv:2606.07545  [pdf, ps, other

    cs.CY

    Reshaping Undergraduate Computer Science Education in the Generative AI Era

    Authors: Yi-Chieh Lee, Nattapat Boonprakong, Yugin Tan, Harold Soh, Alex Potanin, Viraj Kumar, Anoop K. Sinha, Chen Qian, Paul Denny, Mennatallah El-Assady, Ian Oakley, Jake Renzella, Amy Zhang, Jat Singh, Wee Sun Lee, Hsuan-Tien Lin, Jane L. E, Anthony Tang, Margaret M. Burnett, Sowmya Somanath, Renwen Zhang, Vicky Charisi, Alexandra I. Cristea

    Abstract: Generative AI represents a turning point for Computer Science (CS) education. In recent decades, post-secondary CS education has largely focused on what has been seen as practical software engineering skills: implementation-level programming, debugging, testing, and software design, analysis, and documentation. However, this framing is becoming less tenable as generative AI automates many of these… ▽ More

    Submitted 11 June, 2026; v1 submitted 2 May, 2026; originally announced June 2026.

    Comments: Workshop report

  9. arXiv:2605.25541  [pdf, ps, other

    cs.CG cs.AI cs.HC cs.LG

    TopoAlign: Topology-Aware Visual Representation Alignment

    Authors: Xinyuan Yan, Rita Sevastjanova, Mennatallah El-Assady, Bei Wang

    Abstract: Neural networks encode inputs as high-dimensional vectors, known as representations, that capture how models process data by encoding task-relevant structure and semantics. Representation alignment refers to the degree to which different models, layers, or training conditions produce similar representations for the same inputs, with important implications for model interpretation, selection, and r… ▽ More

    Submitted 25 May, 2026; originally announced May 2026.

  10. arXiv:2605.15932  [pdf, ps, other

    cs.HC

    GEMS -- Guided Evolutionary Molecule Design for Sustainable Chemicals

    Authors: Coelina Robinson, Franziska Weissbach, Kjell Jorner, Mennatallah El-Assady, Christina Humer

    Abstract: Designing safe and sustainable chemicals is critical to combat chemical pollution in our environment. Computational and AI-assisted methods have been developed to aid de novo molecule design. However, data on the environmental impacts of chemical compounds are sparse, resulting in low-fidelity machine learning (ML) oracles and unreliable candidate proposals. Furthermore, many automated molecular d… ▽ More

    Submitted 4 July, 2026; v1 submitted 15 May, 2026; originally announced May 2026.

  11. arXiv:2604.19247  [pdf, ps, other

    cs.HC cs.MA cs.SE

    BONSAI: A Mixed-Initiative Workspace for Human-AI Co-Development of Visual Analytics Applications

    Authors: Thilo Spinner, Matthias Miller, Fabian Sperrle-Roth, Mennatallah El-Assady

    Abstract: Developing Visual Analytics (VA) applications requires integrating complex machine learning models with expressive interactive interfaces. Developers face a stark trade-off: building tightly-coupled monoliths plagued by fragile interdependencies, or relying on restrictive, simplistic frameworks. Meanwhile, unconstrained, single-shot AI code generation promises speed but yields unstructured, unaudi… ▽ More

    Submitted 21 April, 2026; originally announced April 2026.

    Comments: 9 pages paper, 2 pages references, 10 figures

    ACM Class: H.5.2; I.2.11; D.2.6; D.2.11; H.5.3

  12. arXiv:2604.12545  [pdf, ps, other

    cs.AI cs.CY

    Cross-Cultural Simulation of Citizen Emotional Responses to Bureaucratic Red Tape Using LLM Agents

    Authors: Wanchun Ni, Jiugeng Sun, Yixian Liu, Mennatallah El-Assady

    Abstract: Improving policymaking is a central concern in public administration. Prior human subject studies reveal substantial cross-cultural differences in citizens' emotional responses to red tape during policy implementation. While LLM agents offer opportunities to simulate human-like responses and reduce experimental costs, their ability to generate culturally appropriate emotional responses to red tape… ▽ More

    Submitted 14 April, 2026; originally announced April 2026.

    Comments: To appear in the CHI 2026 Workshop on PoliSim

  13. arXiv:2603.29322  [pdf, ps, other

    cs.HC

    VACP: Visual Analytics Context Protocol

    Authors: Tobias Stähle, Péter Ferenc Gyarmati, Thilo Spinner, Rita Sevastjanova, Dominik Moritz, Mennatallah El-Assady

    Abstract: The rise of AI agents introduces a fundamental shift in Visual Analytics (VA), in which agents act as a new user group. Current agentic approaches - based on computer vision and raw DOM access - fail to perform VA tasks accurately and efficiently. This paper introduces the Visual Analytics Context Protocol (VACP), a framework designed to make VA applications "agent-ready" that extends generic prot… ▽ More

    Submitted 31 March, 2026; originally announced March 2026.

  14. arXiv:2603.17815  [pdf, ps, other

    cs.CL

    Process Supervision for Chain-of-Thought Reasoning via Monte Carlo Net Information Gain

    Authors: Corentin Royer, Debarun Bhattacharjya, Gaetano Rossiello, Andrea Giovannini, Mennatallah El-Assady

    Abstract: Multi-step reasoning improves the capabilities of large language models (LLMs) but increases the risk of errors propagating through intermediate steps. Process reward models (PRMs) mitigate this by scoring each step individually, enabling fine-grained supervision and improved reliability. Existing methods for training PRMs rely on costly human annotations or computationally intensive automatic lab… ▽ More

    Submitted 18 March, 2026; originally announced March 2026.

  15. arXiv:2603.01795  [pdf, ps, other

    cs.HC cs.AI cs.CL

    PleaSQLarify: Visual Pragmatic Repair for Natural Language Database Querying

    Authors: Robin Shing Moon Chan, Rita Sevastjanova, Mennatallah El-Assady

    Abstract: Natural language database interfaces broaden data access, yet they remain brittle under input ambiguity. Standard approaches often collapse uncertainty into a single query, offering little support for mismatches between user intent and system interpretation. We reframe this challenge through pragmatic inference: while users economize expressions, systems operate on priors over the action space tha… ▽ More

    Submitted 2 March, 2026; originally announced March 2026.

    Comments: Accepted at CHI'26, main track

  16. arXiv:2602.15206  [pdf, ps, other

    cs.LG cs.AI

    MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference

    Authors: Raphaël Baur, Yannick Metz, Maria Gkoulta, Mennatallah El-Assady, Giorgia Ramponi, Thomas Kleine Buening

    Abstract: Reward learning typically relies on a single feedback type or combines multiple feedback types using manually weighted loss terms. Currently, it remains unclear how to jointly learn reward functions from heterogeneous feedback types such as demonstrations, comparisons, ratings, and stops that provide qualitatively different signals. We address this challenge by formulating reward learning from mul… ▽ More

    Submitted 19 June, 2026; v1 submitted 16 February, 2026; originally announced February 2026.

    Comments: Accepted at ICML 2026

  17. arXiv:2512.23372  [pdf, ps, other

    cs.HC

    A Design Space for Intelligent Agents in Mixed-Initiative Visual Analytics

    Authors: Tobias Stähle, Matthijs Jansen op de Haar, Sophia Boyer, Rita Sevastjanova, Arpit Narechania, Mennatallah El-Assady

    Abstract: Mixed-initiative visual analytics (VA) systems, where human and artificial intelligence (AI) agents collaborate as equal partners during analysis, represented a paradigm shift in human-computer interaction. With recent advances in AI, these systems have seen an increase in sophisticated software agents that have improved task planning, reasoning, and completion capabilities. However, while existin… ▽ More

    Submitted 29 December, 2025; originally announced December 2025.

  18. arXiv:2512.07483  [pdf, ps, other

    cs.HC

    SemanticTours: A Conceptual Framework for Non-Linear, Knowledge Graph-Driven Data Tours

    Authors: Daniel Fürst, Matthijs Jansen op de Haar, Mennatallah El-Assady, Daniel A Keim, Maximilian T. Fischer

    Abstract: Interactive tours help users explore datasets and provide onboarding. They rely on a linear sequence of views, showing a curated set of relevant data selections and introduce user interfaces. Existing frameworks of tours, however, often do not allow for branching and refining hypotheses outside of a rigid sequence, which is important in knowledge-centric domains such as law. For example, lawyers p… ▽ More

    Submitted 8 December, 2025; originally announced December 2025.

    Comments: 14 pages, 9 figures, 2 tables

    ACM Class: H.5.2

  19. arXiv:2509.26210  [pdf, ps, other

    cs.HC

    Dia-Lingle: A Gamified Interface for Dialectal Data Collection

    Authors: Jiugeng Sun, Rita Sevastjanova, Sina Ahmadi, Rico Sennrich, Mennatallah El-Assady

    Abstract: Dialects suffer from the scarcity of computational textual resources as they exist predominantly in spoken rather than written form and exhibit remarkable geographical diversity. Collecting dialect data and subsequently integrating it into current language technologies present significant obstacles. Gamification has been proven to facilitate remote data collection processes with great ease and on… ▽ More

    Submitted 30 September, 2025; originally announced September 2025.

  20. arXiv:2509.20901  [pdf, ps, other

    cs.HC

    CafGa: Customizing Feature Attributions to Explain Language Models

    Authors: Alan Boyle, Furui Cheng, Vilém Zouhar, Mennatallah El-Assady

    Abstract: Feature attribution methods, such as SHAP and LIME, explain machine learning model predictions by quantifying the influence of each input component. When applying feature attributions to explain language models, a basic question is defining the interpretable components. Traditional feature attribution methods, commonly treat individual words as atomic units. This is highly computationally ineffici… ▽ More

    Submitted 25 September, 2025; originally announced September 2025.

    Comments: 6 Pages (excl. bibliography and appendix), 5 Figures and 2 Tables. Will be presented at EMNLP 2025 (Demo Track)

  21. Beyond Quantification: Navigating Uncertainty in Professional AI Systems

    Authors: Sylvie Delacroix, Diana Robinson, Umang Bhatt, Jacopo Domenicucci, Jessica Montgomery, Gael Varoquaux, Carl Henrik Ek, Vincent Fortuin, Yulan He, Tom Diethe, Neill Campbell, Mennatallah El-Assady, Soren Hauberg, Ivana Dusparic, Neil Lawrence

    Abstract: The growing integration of large language models across professional domains transforms how experts make critical decisions in healthcare, education, and law. While significant research effort focuses on getting these systems to communicate their outputs with probabilistic measures of reliability, many consequential forms of uncertainty in professional contexts resist such quantification. A physic… ▽ More

    Submitted 3 September, 2025; originally announced September 2025.

    Journal ref: RSS Data Science (2025)

  22. DxHF: Providing High-Quality Human Feedback for LLM Alignment via Interactive Decomposition

    Authors: Danqing Shi, Furui Cheng, Tino Weinkauf, Antti Oulasvirta, Mennatallah El-Assady

    Abstract: Human preferences are widely used to align large language models (LLMs) through methods such as reinforcement learning from human feedback (RLHF). However, the current user interfaces require annotators to compare text paragraphs, which is cognitively challenging when the texts are long or unfamiliar. This paper contributes by studying the decomposition principle as an approach to improving the qu… ▽ More

    Submitted 24 July, 2025; originally announced July 2025.

  23. arXiv:2507.18607  [pdf, ps, other

    cs.CG cs.LG

    Explainable Mapper: Charting LLM Embedding Spaces Using Perturbation-Based Explanation and Verification Agents

    Authors: Xinyuan Yan, Rita Sevastjanova, Sinie van der Ben, Mennatallah El-Assady, Bei Wang

    Abstract: Large language models (LLMs) produce high-dimensional embeddings that capture rich semantic and syntactic relationships between words, sentences, and concepts. Investigating the topological structures of LLM embedding spaces via mapper graphs enables us to understand their underlying structures. Specifically, a mapper graph summarizes the topological structure of the embedding space, where each no… ▽ More

    Submitted 24 July, 2025; originally announced July 2025.

  24. arXiv:2507.05010  [pdf, ps, other

    cs.CL

    Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification

    Authors: Chenfei Xiong, Jingwei Ni, Yu Fan, Vilém Zouhar, Donya Rooein, Lorena Calvo-Bartolomé, Alexander Hoyle, Zhijing Jin, Mrinmaya Sachan, Markus Leippold, Dirk Hovy, Mennatallah El-Assady, Elliott Ash

    Abstract: We introduce Co-DETECT (Collaborative Discovery of Edge cases in TExt ClassificaTion), a novel mixed-initiative annotation framework that integrates human expertise with automatic annotation guided by large language models (LLMs). Co-DETECT starts with an initial, sketch-level codebook and dataset provided by a domain expert, then leverages the LLM to annotate the data and identify edge cases that… ▽ More

    Submitted 7 July, 2025; originally announced July 2025.

  25. arXiv:2505.07610  [pdf, ps, other

    cs.CL cs.AI

    Concept-Level Explainability for Auditing & Steering LLM Responses

    Authors: Kenza Amara, Rita Sevastjanova, Mennatallah El-Assady

    Abstract: As large language models (LLMs) become widely deployed, concerns about their safety and alignment grow. An approach to steer LLM behavior, such as mitigating biases or defending against jailbreaks, is to identify which parts of a prompt influence specific aspects of the model's output. Token-level attribution methods offer a promising solution, but still struggle in text generation, explaining the… ▽ More

    Submitted 19 May, 2025; v1 submitted 12 May, 2025; originally announced May 2025.

    Comments: 9 pages, 7 figures, Submission to Neurips 2025

  26. arXiv:2504.10504  [pdf, other

    cs.CL cs.GR

    LayerFlow: Layer-wise Exploration of LLM Embeddings using Uncertainty-aware Interlinked Projections

    Authors: Rita Sevastjanova, Robin Gerling, Thilo Spinner, Mennatallah El-Assady

    Abstract: Large language models (LLMs) represent words through contextual word embeddings encoding different language properties like semantics and syntax. Understanding these properties is crucial, especially for researchers investigating language model capabilities, employing embeddings for tasks related to text similarity, or evaluating the reasons behind token importance as measured through attribution… ▽ More

    Submitted 9 April, 2025; originally announced April 2025.

  27. arXiv:2502.21038  [pdf, other

    cs.LG cs.AI

    Reward Learning from Multiple Feedback Types

    Authors: Yannick Metz, András Geiszl, Raphaël Baur, Mennatallah El-Assady

    Abstract: Learning rewards from preference feedback has become an important tool in the alignment of agentic models. Preference-based feedback, often implemented as a binary comparison between multiple completions, is an established method to acquire large-scale human feedback. However, human feedback in other contexts is often much more diverse. Such diverse feedback can better support the goals of a human… ▽ More

    Submitted 28 February, 2025; originally announced February 2025.

    Comments: Published as a conference paper at ICLR 2025

  28. Challenges and Opportunities for Visual Analytics in Jurisprudence

    Authors: Daniel Fürst, Mennatallah El-Assady, Daniel A. Keim, Maximilian T. Fischer

    Abstract: Legal exploration, analysis, and interpretation remain complex and demanding tasks, even for experienced legal scholars, due to the domain-specific language, tacit legal concepts, and intentional ambiguities embedded in legal texts. In related, text-based domains, Visual Analytics (VA) has become an indispensable tool for navigating documents, representing knowledge, and supporting analytical reas… ▽ More

    Submitted 19 November, 2025; v1 submitted 9 December, 2024; originally announced December 2024.

    Comments: 34 pages, 3 figures, 1 table

    ACM Class: H.5.2

    Journal ref: Artificial Intelligence and Law 2025

  29. arXiv:2411.11761  [pdf, other

    cs.LG cs.HC

    Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework

    Authors: Yannick Metz, David Lindner, Raphaël Baur, Mennatallah El-Assady

    Abstract: Reinforcement Learning from Human feedback (RLHF) has become a powerful tool to fine-tune or train agentic machine learning models. Similar to how humans interact in social contexts, we can use many types of feedback to communicate our preferences, intentions, and knowledge to an RL agent. However, applications of human feedback in RL are often limited in scope and disregard human factors. In this… ▽ More

    Submitted 20 February, 2025; v1 submitted 18 November, 2024; originally announced November 2024.

  30. arXiv:2410.01690  [pdf, other

    cs.AI

    Why context matters in VQA and Reasoning: Semantic interventions for VLM input modalities

    Authors: Kenza Amara, Lukas Klein, Carsten Lüth, Paul Jäger, Hendrik Strobelt, Mennatallah El-Assady

    Abstract: The various limitations of Generative AI, such as hallucinations and model failures, have made it crucial to understand the role of different modalities in Visual Language Model (VLM) predictions. Our work investigates how the integration of information from image and text modalities influences the performance and behavior of VLMs in visual question answering (VQA) and reasoning tasks. We measure… ▽ More

    Submitted 2 October, 2024; originally announced October 2024.

  31. arXiv:2409.16756  [pdf, other

    cs.CV

    Navigating the Maze of Explainable AI: A Systematic Approach to Evaluating Methods and Metrics

    Authors: Lukas Klein, Carsten T. Lüth, Udo Schlegel, Till J. Bungert, Mennatallah El-Assady, Paul F. Jäger

    Abstract: Explainable AI (XAI) is a rapidly growing domain with a myriad of proposed methods as well as metrics aiming to evaluate their efficacy. However, current studies are often of limited scope, examining only a handful of XAI methods and ignoring underlying design parameters for performance, such as the model architecture or the nature of input data. Moreover, they often rely on one or a few metrics a… ▽ More

    Submitted 2 January, 2025; v1 submitted 25 September, 2024; originally announced September 2024.

    Comments: Accepted at NeurIPS 2024

  32. arXiv:2409.00413  [pdf, other

    cs.HC

    iToT: An Interactive System for Customized Tree-of-Thought Generation

    Authors: Alan Boyle, Isha Gupta, Sebastian Hönig, Lukas Mautner, Kenza Amara, Furui Cheng, Mennatallah El-Assady

    Abstract: As language models have become increasingly successful at a wide array of tasks, different prompt engineering methods have been developed alongside them in order to adapt these models to new tasks. One of them is Tree-of-Thoughts (ToT), a prompting strategy and framework for language model inference and problem-solving. It allows the model to explore multiple solution paths and select the best cou… ▽ More

    Submitted 31 August, 2024; originally announced September 2024.

    Comments: 6 pages excl. figures and comments; 8 figures. Will appear in IEEE 2024 NLVIZ Workshop

  33. arXiv:2407.17998  [pdf, other

    cs.HC cs.LG

    iNNspector: Visual, Interactive Deep Model Debugging

    Authors: Thilo Spinner, Daniel Fürst, Mennatallah El-Assady

    Abstract: Deep learning model design, development, and debugging is a process driven by best practices, guidelines, trial-and-error, and the personal experiences of model developers. At multiple stages of this process, performance and internal model data can be logged and made available. However, due to the sheer complexity and scale of this data and process, model developers often resort to evaluating thei… ▽ More

    Submitted 25 July, 2024; originally announced July 2024.

    Comments: 41 pages paper, 4 pages references, 3 pages appendix, 19 figures, 2 tables

  34. arXiv:2407.17431  [pdf, other

    cs.HC

    ProvenanceWidgets: A Library of UI Control Elements to Track and Dynamically Overlay Analytic Provenance

    Authors: Arpit Narechania, Kaustubh Odak, Mennatallah El-Assady, Alex Endert

    Abstract: We present ProvenanceWidgets, a Javascript library of UI control elements such as radio buttons, checkboxes, and dropdowns to track and dynamically overlay a user's analytic provenance. These in situ overlays not only save screen space but also minimize the amount of time and effort needed to access the same information from elsewhere in the UI. In this paper, we discuss how we design modular UI c… ▽ More

    Submitted 24 July, 2024; originally announced July 2024.

    Comments: 11 pages, 8 figures. To appear in IEEE VIS 2024

  35. arXiv:2407.05427  [pdf, other

    cs.HC cs.IR

    MelodyVis: Visual Analytics for Melodic Patterns in Sheet Music

    Authors: Matthias Miller, Daniel Fürst, Maximilian T. Fischer, Hanna Hauptmann, Daniel Keim, Mennatallah El-Assady

    Abstract: Manual melody detection is a tedious task requiring high expertise level, while automatic detection is often not expressive or powerful enough. Thus, we present MelodyVis, a visual application designed in collaboration with musicology experts to explore melodic patterns in digital sheet music. MelodyVis features five connected views, including a Melody Operator Graph and a Voicing Timeline. The sy… ▽ More

    Submitted 7 July, 2024; originally announced July 2024.

    Comments: 9+2 pages, 9 figures, preprint, originally submitted to IEEE VIS 23, revision

    ACM Class: I.5.4; H.3.3; J.5.7

  36. arXiv:2406.02329  [pdf, other

    cs.CL cs.LG

    On Affine Homotopy between Language Encoders

    Authors: Robin SM Chan, Reda Boumasmoud, Anej Svete, Yuxin Ren, Qipeng Guo, Zhijing Jin, Shauli Ravfogel, Mrinmaya Sachan, Bernhard Schölkopf, Mennatallah El-Assady, Ryan Cotterell

    Abstract: Pre-trained language encoders -- functions that represent text as vectors -- are an integral component of many NLP tasks. We tackle a natural question in language encoder analysis: What does it mean for two encoders to be similar? We contend that a faithful measure of similarity needs to be \emph{intrinsic}, that is, task-independent, yet still be informative of \emph{extrinsic} similarity -- the… ▽ More

    Submitted 18 December, 2024; v1 submitted 4 June, 2024; originally announced June 2024.

    Comments: 10 pages, Accepted at NeurIPS 2024 (Main)

  37. arXiv:2405.08468  [pdf, other

    cs.CL cs.AI

    Challenges and Opportunities in Text Generation Explainability

    Authors: Kenza Amara, Rita Sevastjanova, Mennatallah El-Assady

    Abstract: The necessity for interpretability in natural language processing (NLP) has risen alongside the growing prominence of large language models. Among the myriad tasks within NLP, text generation stands out as a primary objective of autoregressive models. The NLP community has begun to take a keen interest in gaining a deeper understanding of text generation, leading to the development of model-agnost… ▽ More

    Submitted 14 May, 2024; originally announced May 2024.

    Comments: 17 pages, 5 figures, xAI-2024 Conference, Main track

  38. arXiv:2405.00708  [pdf, ps, other

    cs.CL cs.AI cs.HC cs.LG

    Understanding Large Language Model Behaviors through Interactive Counterfactual Generation and Analysis

    Authors: Furui Cheng, Vilém Zouhar, Robin Shing Moon Chan, Daniel Fürst, Hendrik Strobelt, Mennatallah El-Assady

    Abstract: Understanding the behavior of large language models (LLMs) is crucial for ensuring their safe and reliable use. However, existing explainable AI (XAI) methods for LLMs primarily rely on word-level explanations, which are often computationally inefficient and misaligned with human reasoning processes. Moreover, these methods often treat explanation as a one-time output, overlooking its inherently i… ▽ More

    Submitted 7 August, 2025; v1 submitted 23 April, 2024; originally announced May 2024.

    ACM Class: I.2.7; H.5.2

  39. arXiv:2404.12056  [pdf, other

    cs.HC cs.AI

    Deconstructing Human-AI Collaboration: Agency, Interaction, and Adaptation

    Authors: Steffen Holter, Mennatallah El-Assady

    Abstract: As full AI-based automation remains out of reach in most real-world applications, the focus has instead shifted to leveraging the strengths of both human and AI agents, creating effective collaborative systems. The rapid advances in this area have yielded increasingly more complex systems and frameworks, while the nuance of their characterization has gotten more vague. Similarly, the existing conc… ▽ More

    Submitted 18 April, 2024; originally announced April 2024.

    Comments: 10 pages, 4 figures. Accepted to Proceedings of EuroVis 2024

  40. arXiv:2403.07627  [pdf, other

    cs.HC cs.LG

    generAItor: Tree-in-the-Loop Text Generation for Language Model Explainability and Adaptation

    Authors: Thilo Spinner, Rebecca Kehlbeck, Rita Sevastjanova, Tobias Stähle, Daniel A. Keim, Oliver Deussen, Mennatallah El-Assady

    Abstract: Large language models (LLMs) are widely deployed in various downstream tasks, e.g., auto-completion, aided writing, or chat-based text generation. However, the considered output candidates of the underlying search algorithm are under-explored and under-explained. We tackle this shortcoming by proposing a tree-in-the-loop approach, where a visual representation of the beam search tree is the centra… ▽ More

    Submitted 12 March, 2024; originally announced March 2024.

    Comments: 24 pages paper, 4 pages references, 3 pages appendix, 8 figures

    ACM Class: I.2.7; H.5.2

  41. arXiv:2402.09259  [pdf, other

    cs.CL cs.AI

    SyntaxShap: Syntax-aware Explainability Method for Text Generation

    Authors: Kenza Amara, Rita Sevastjanova, Mennatallah El-Assady

    Abstract: To harness the power of large language models in safety-critical domains, we need to ensure the explainability of their predictions. However, despite the significant attention to model interpretability, there remains an unexplored domain in explaining sequence-to-sequence tasks using methods tailored for textual data. This paper introduces SyntaxShap, a local, model-agnostic explainability method… ▽ More

    Submitted 3 June, 2024; v1 submitted 14 February, 2024; originally announced February 2024.

    Comments: Accepted to ACL 2024

  42. arXiv:2402.02827  [pdf, other

    cs.LG eess.SY

    PowerGraph: A power grid benchmark dataset for graph neural networks

    Authors: Anna Varbella, Kenza Amara, Blazhe Gjorgiev, Mennatallah El-Assady, Giovanni Sansavini

    Abstract: Power grids are critical infrastructures of paramount importance to modern society and, therefore, engineered to operate under diverse conditions and failures. The ongoing energy transition poses new challenges for the decision-makers and system operators. Therefore, developing grid analysis algorithms is important for supporting reliable operations. These key tools include power flow analysis and… ▽ More

    Submitted 31 October, 2024; v1 submitted 5 February, 2024; originally announced February 2024.

    Comments: 21 pages, 8 figures, conference paper

  43. RELIC: Investigating Large Language Model Responses using Self-Consistency

    Authors: Furui Cheng, Vilém Zouhar, Simran Arora, Mrinmaya Sachan, Hendrik Strobelt, Mennatallah El-Assady

    Abstract: Large Language Models (LLMs) are notorious for blending fact with fiction and generating non-factual content, known as hallucinations. To address this challenge, we propose an interactive system that helps users gain insight into the reliability of the generated text. Our approach is based on the idea that the self-consistency of multiple samples generated by the same LLM relates to its confidence… ▽ More

    Submitted 4 April, 2024; v1 submitted 28 November, 2023; originally announced November 2023.

  44. arXiv:2310.13544  [pdf, other

    cs.CL cs.HC

    A Diachronic Perspective on User Trust in AI under Uncertainty

    Authors: Shehzaad Dhuliawala, Vilém Zouhar, Mennatallah El-Assady, Mrinmaya Sachan

    Abstract: In a human-AI collaboration, users build a mental model of the AI system based on its reliability and how it presents its decision, e.g. its presentation of system confidence and an explanation of the output. Modern NLP systems are often uncalibrated, resulting in confidently incorrect predictions that undermine user trust. In order to build trustworthy AI, we must understand how user trust is dev… ▽ More

    Submitted 20 October, 2023; originally announced October 2023.

    Comments: EMNLP 2023, 14 pages (8+6)

  45. arXiv:2310.11252  [pdf, other

    cs.CL cs.AI cs.HC

    Revealing the Unwritten: Visual Investigation of Beam Search Trees to Address Language Model Prompting Challenges

    Authors: Thilo Spinner, Rebecca Kehlbeck, Rita Sevastjanova, Tobias Stähle, Daniel A. Keim, Oliver Deussen, Andreas Spitz, Mennatallah El-Assady

    Abstract: The growing popularity of generative language models has amplified interest in interactive methods to guide model outputs. Prompt refinement is considered one of the most effective means to influence output among these methods. We identify several challenges associated with prompting large language models, categorized into data- and model-specific, linguistic, and socio-linguistic challenges. A co… ▽ More

    Submitted 17 October, 2023; originally announced October 2023.

    Comments: 9 pages paper, 2 pages references, 7 figures

    ACM Class: H.5.2; I.2.7

  46. arXiv:2309.16223  [pdf, other

    cs.AI cs.LG

    GInX-Eval: Towards In-Distribution Evaluation of Graph Neural Network Explanations

    Authors: Kenza Amara, Mennatallah El-Assady, Rex Ying

    Abstract: Diverse explainability methods of graph neural networks (GNN) have recently been developed to highlight the edges and nodes in the graph that contribute the most to the model predictions. However, it is not clear yet how to evaluate the correctness of those explanations, whether it is from a human or a model perspective. One unaddressed bottleneck in the current evaluation procedure is the problem… ▽ More

    Submitted 5 November, 2023; v1 submitted 28 September, 2023; originally announced September 2023.

    Comments: Preprint, Submitted to ICLR2024

  47. A Heuristic Approach for Dual Expert/End-User Evaluation of Guidance in Visual Analytics

    Authors: Davide Ceneda, Christopher Collins, Mennatallah El-Assady, Silvia Miksch, Christian Tominski, Alessio Arleo

    Abstract: Guidance can support users during the exploration and analysis of complex data. Previous research focused on characterizing the theoretical aspects of guidance in visual analytics and implementing guidance in different scenarios. However, the evaluation of guidance-enhanced visual analytics solutions remains an open research question. We tackle this question by introducing and validating a practic… ▽ More

    Submitted 24 August, 2023; originally announced August 2023.

    Comments: Accepted to IEEE VIS 2023

  48. arXiv:2308.04332  [pdf, other

    cs.LG cs.HC

    RLHF-Blender: A Configurable Interactive Interface for Learning from Diverse Human Feedback

    Authors: Yannick Metz, David Lindner, Raphaël Baur, Daniel Keim, Mennatallah El-Assady

    Abstract: To use reinforcement learning from human feedback (RLHF) in practical applications, it is crucial to learn reward models from diverse sources of human feedback and to consider human factors involved in providing feedback of different types. However, the systematic study of learning from diverse types of feedback is held back by limited standardized tooling available to researchers. To bridge this… ▽ More

    Submitted 8 August, 2023; originally announced August 2023.

    Comments: 14 pages, 3 figures

    Journal ref: ICML2023 Interactive Learning from Implicit Human Feedback Workshop

  49. arXiv:2307.08494  [pdf, other

    cs.HC cs.LG

    Visual Explanations with Attributions and Counterfactuals on Time Series Classification

    Authors: Udo Schlegel, Daniela Oelke, Daniel A. Keim, Mennatallah El-Assady

    Abstract: With the rising necessity of explainable artificial intelligence (XAI), we see an increase in task-dependent XAI methods on varying abstraction levels. XAI techniques on a global level explain model behavior and on a local level explain sample predictions. We propose a visual analytics workflow to support seamless transitions between global and local explanations, focusing on attributions and coun… ▽ More

    Submitted 14 July, 2023; originally announced July 2023.

    Comments: 14 pages, 6 figures, submitted to IEEE TVCG in December

  50. arXiv:2306.12146  [pdf, other

    cs.CL cs.HC

    Which Spurious Correlations Impact Reasoning in NLI Models? A Visual Interactive Diagnosis through Data-Constrained Counterfactuals

    Authors: Robin Chan, Afra Amini, Mennatallah El-Assady

    Abstract: We present a human-in-the-loop dashboard tailored to diagnosing potential spurious features that NLI models rely on for predictions. The dashboard enables users to generate diverse and challenging examples by drawing inspiration from GPT-3 suggestions. Additionally, users can receive feedback from a trained NLI model on how challenging the newly created example is and make refinements based on the… ▽ More

    Submitted 21 June, 2023; originally announced June 2023.

    Comments: 7 pages, Accepted at ACL 2023: System Demonstrations