Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–19 of 19 results for author: Tewolde, E

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.12125  [pdf, ps, other

    cs.GT cs.AI cs.CL cs.MA

    Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation

    Authors: Akash Kundu, Emanuel Tewolde, Ratip Emin Berker, Samuel F. Brown, Vincent Conitzer

    Abstract: As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes. Prior literature has argued that cooperation problems such as the Prisoner's Dilemma are resolvable in settings where agents know they follow very similar decision making patterns, as for example… ▽ More

    Submitted 12 August, 2026; originally announced August 2026.

    Comments: 41 pages, 18 Figures, 4 Tables, 16 Listings

    MSC Class: 68T05; 68T37; 68T42; 91A05; 91A06; 91A10; 91A35 ACM Class: I.2; J.4; K.4

  2. arXiv:2605.08060  [pdf, ps, other

    cs.CL cs.AI cs.GT cs.MA

    The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agents

    Authors: Jiayuan Liu, Tianqin Li, Shiyi Du, Xin Luo, Haoxuan Zeng, Emanuel Tewolde, Tai Sing Lee, Tonghan Wang, Carl Kingsford, Vincent Conitzer

    Abstract: Context window expansion is often treated as a straightforward capability upgrade for LLMs, but we find it systematically fails in multi-agent social dilemmas. Across 7 LLMs and 4 games over 500 rounds, expanding accessible history degrades cooperation in 18 of 28 model--game settings, a pattern we term the memory curse. We isolate the underlying mechanism through three analyses. First, lexical an… ▽ More

    Submitted 8 May, 2026; originally announced May 2026.

  3. arXiv:2604.15267  [pdf, ps, other

    cs.GT cs.AI cs.CL cs.CY cs.MA

    CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas

    Authors: Emanuel Tewolde, Xiao Zhang, David Guzman Piedrahita, Vincent Conitzer, Zhijing Jin

    Abstract: It is increasingly important that LLM agents interact effectively and safely with other goal-pursuing agents, yet, recent works report the opposite trend: LLMs with stronger reasoning capabilities behave _less_ cooperatively in mixed-motive games such as the prisoner's dilemma and public goods settings. Indeed, our experiments show that recent models -- with or without reasoning enabled -- consist… ▽ More

    Submitted 4 July, 2026; v1 submitted 16 April, 2026; originally announced April 2026.

    Comments: Published paper at the International Conference on Machine Learning (ICML) 2026. 65 pages, 38 Figures, 8 Tables, 17 Listings

    MSC Class: 68T05; 68T42; 91A05; 91A06; 91A10; 91A20; ACM Class: I.2; J.4; K.4

  4. arXiv:2602.15252  [pdf, ps, other

    cs.GT cs.AI cs.LG

    Decision Making under Imperfect Recall: Algorithms and Benchmarks

    Authors: Emanuel Tewolde, Brian Hu Zhang, Ioannis Anagnostides, Tuomas Sandholm, Vincent Conitzer

    Abstract: In game theory, imperfect-recall decision problems model situations in which an agent forgets information it held before. They encompass games such as the ``absentminded driver'' and team games with limited communication. In this paper, we introduce the first benchmark suite for imperfect-recall decision problems. Our benchmarks capture a variety of problem types, including ones concerning privacy… ▽ More

    Submitted 16 February, 2026; originally announced February 2026.

    Comments: 39 pages, 71 figures, 4 table

    MSC Class: 68Q25; 68T01; 68T05; 68T42; 90C23; 90C26; 90C30; 91A14; 91A18; 91A26; 91A27; 91A35; 91A68 ACM Class: I.2; J.4; F.2

  5. arXiv:2602.06855  [pdf, ps, other

    cs.AI

    AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents

    Authors: Alisia Lupidi, Bhavul Gauri, Thomas Simon Foster, Bassel Al Omari, Despoina Magka, Alberto Pepe, Alexis Audran-Reiss, Muna Aghamelu, Nicolas Baldwin, Lucia Cipolina-Kun, Jean-Christophe Gagnon-Audet, Chee Hau Leow, Sandra Lefdal, Hossam Mossalam, Abhinav Moudgil, Saba Nazir, Emanuel Tewolde, Isabel Urrego, Jordi Armengol Estape, Amar Budhiraja, Gaurav Chaurasia, Abhishek Charnalia, Derek Dunfield, Karen Hambardzumyan, Daniel Izcovich , et al. (12 additional authors not shown)

    Abstract: LLM agents hold significant promise for advancing scientific research. To accelerate this progress, we introduce AIRS-Bench (the AI Research Science Benchmark), a suite of 20 tasks sourced from state-of-the-art machine learning papers. These tasks span diverse domains, including language modeling, mathematics, bioinformatics, and time series forecasting. AIRS-Bench tasks assess agentic capabilitie… ▽ More

    Submitted 16 February, 2026; v1 submitted 6 February, 2026; originally announced February 2026.

    Comments: 49 pages, 14 figures, 10 tables

  6. arXiv:2512.16895  [pdf, ps, other

    cs.GT

    On the Edge of Core (Non-)Emptiness: An Automated Reasoning Approach to Approval-Based Multi-Winner Voting

    Authors: Ratip Emin Berker, Emanuel Tewolde, Vincent Conitzer, Mingyu Guo, Marijn Heule, Lirong Xia

    Abstract: Core stability is a natural and well-studied notion for group fairness in multi-winner voting, where the task is to select a committee from a pool of candidates. We study the setting where voters either approve or disapprove of each candidate; here, it remains a major open problem whether a core-stable committee always exists. In this work, we develop an approach based on mixed-integer linear prog… ▽ More

    Submitted 18 December, 2025; originally announced December 2025.

    Comments: 29 pages, 1 figure, main body to be published in Proceedings of the Fortieth AAAI Conference on Artificial Intelligence (AAAI-26), Singapore, 2026

    MSC Class: 91B12; 91B14; 68Q25; 68T01; 68V05; 90C11 ACM Class: F.2; I.2; J.4

  7. arXiv:2512.03318  [pdf, ps, other

    cs.AI

    Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia

    Authors: Chandler Smith, Marwa Abdulhai, Manfred Diaz, Marko Tesic, Rakshit S. Trivedi, Alexander Sasha Vezhnevets, Lewis Hammond, Jesse Clifton, Minsuk Chang, Edgar A. Duéñez-Guzmán, John P. Agapiou, Jayd Matyas, Danny Karmon, Akash Kundu, Aliaksei Korshuk, Ananya Ananya, Arrasy Rahman, Avinaash Anand Kulandaivel, Bain McHale, Beining Zhang, Buyantuev Alexander, Carlos Saith Rodriguez Rojas, Caroline Wang, Chetan Talele, Chenao Liu , et al. (61 additional authors not shown)

    Abstract: Large Language Model (LLM) agents have demonstrated impressive capabilities for social interaction and are increasingly being deployed in situations where they might engage with both human and artificial agents. These interactions represent a critical frontier for LLM-based agents, yet existing evaluation methods fail to measure how well these capabilities generalize to novel social situations. In… ▽ More

    Submitted 2 December, 2025; originally announced December 2025.

    Comments: Published at NeurIPS Datasets and Benchmarks 2025, 10 pages

    MSC Class: 68T42 ACM Class: I.2.6

  8. arXiv:2511.15593  [pdf, ps, other

    cs.AI

    What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity

    Authors: Alexis Audran-Reiss, Jordi Armengol-Estapé, Karen Hambardzumyan, Amar Budhiraja, Martin Josifoski, Edan Toledo, Rishi Hazra, Despoina Magka, Michael Shvartsman, Parth Pathak, Justine T Kao, Lucia Cipolina-Kun, Bhavul Gauri, Jean-Christophe Gagnon-Audet, Emanuel Tewolde, Jenny Zhang, Taco Cohen, Yossi Adi, Tatiana Shavrina, Yoram Bachrach

    Abstract: AI research agents offer the promise to accelerate scientific progress by automating the design, implementation, and training of machine learning models. However, the field is still in its infancy, and the key factors driving the success or failure of agent trajectories are not fully understood. We examine the role that ideation diversity plays in agent performance. First, we analyse agent traject… ▽ More

    Submitted 9 December, 2025; v1 submitted 19 November, 2025; originally announced November 2025.

  9. arXiv:2511.03968  [pdf, ps, other

    cs.GT

    The Complexity of Equilibrium Refinements in Potential Games

    Authors: Ioannis Anagnostides, Maria-Florina Balcan, Kiriaki Fragkia, Tuomas Sandholm, Emanuel Tewolde, Brian Hu Zhang

    Abstract: The complexity of computing equilibrium refinements has been at the forefront of algorithmic game theory research, but it has remained open in the seminal class of potential games; we close this fundamental gap in this paper. We first show that computing a pure(-strategy) perfect or proper equilibrium is $\mathsf{PLS}$-complete in concise potential games in normal form. For pure perfect equilibr… ▽ More

    Submitted 10 February, 2026; v1 submitted 5 November, 2025; originally announced November 2025.

    Comments: The abstract has been abridged due to arXiv length constraints. The previous version of this preprint contained results concerning normal-form proper equilibria; these results have now been extended and moved to a separate paper

  10. arXiv:2510.17067  [pdf, ps, other

    cs.GT cs.LG math.OC

    Convergence of Regret Matching in Potential Games and Constrained Optimization

    Authors: Ioannis Anagnostides, Emanuel Tewolde, Brian Hu Zhang, Ioannis Panageas, Vincent Conitzer, Tuomas Sandholm

    Abstract: Regret matching (RM) -- and its modern variants -- is a foundational online algorithm that has been at the heart of many AI breakthrough results in solving benchmark zero-sum games, such as poker. Yet, surprisingly little is known so far in theory about its convergence beyond two-player zero-sum games. For example, whether regret matching converges to Nash equilibria in potential games has been an… ▽ More

    Submitted 17 November, 2025; v1 submitted 19 October, 2025; originally announced October 2025.

    Comments: V2 extends the convergence bounds to simultaneous RM+

  11. arXiv:2505.19212  [pdf, ps, other

    cs.CL cs.AI cs.CY

    When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas

    Authors: Steffen Backmann, David Guzman Piedrahita, Terry Jingchen Zhang, Emanuel Tewolde, Rada Mihalcea, Bernhard Schölkopf, Zhijing Jin

    Abstract: Recent advances in LLMs have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical alignment a critical concern. While prior work has examined LLMs' moral judgment and strategic behavior separately, there is limited understanding of how they act when moral imperatives directly conflict with profit incentives. We introduce \msimfull (\msim… ▽ More

    Submitted 24 July, 2026; v1 submitted 25 May, 2025; originally announced May 2025.

  12. arXiv:2502.18605  [pdf, other

    cs.GT cs.LG math.OC

    Expected Variational Inequalities

    Authors: Brian Hu Zhang, Ioannis Anagnostides, Emanuel Tewolde, Ratip Emin Berker, Gabriele Farina, Vincent Conitzer, Tuomas Sandholm

    Abstract: Variational inequalities (VIs) encompass many fundamental problems in diverse areas ranging from engineering to economics and machine learning. However, their considerable expressivity comes at the cost of computational intractability. In this paper, we introduce and analyze a natural relaxation -- which we refer to as expected variational inequalities (EVIs) -- where the goal is to find a distrib… ▽ More

    Submitted 27 February, 2025; v1 submitted 25 February, 2025; originally announced February 2025.

    Comments: V2 expands on the related work

  13. arXiv:2502.18582  [pdf, ps, other

    stat.ML cs.GT cs.LG

    Learning and Computation of $Φ$-Equilibria at the Frontier of Tractability

    Authors: Brian Hu Zhang, Ioannis Anagnostides, Emanuel Tewolde, Ratip Emin Berker, Gabriele Farina, Vincent Conitzer, Tuomas Sandholm

    Abstract: $Φ$-equilibria -- and the associated notion of $Φ$-regret -- are a powerful and flexible framework at the heart of online learning and game theory, whereby enriching the set of deviations $Φ$ begets stronger notions of rationality. Recently, Daskalakis, Farina, Fishelson, Pipis, and Schneider (STOC '24) -- abbreviated as DFFPS -- settled the existence of efficient algorithms when $Φ… ▽ More

    Submitted 12 December, 2025; v1 submitted 25 February, 2025; originally announced February 2025.

    Comments: V2 makes minor corrections

  14. arXiv:2501.08905  [pdf, other

    cs.GT cs.AI cs.CC cs.MA

    Computing Game Symmetries and Equilibria That Respect Them

    Authors: Emanuel Tewolde, Brian Hu Zhang, Caspar Oesterheld, Tuomas Sandholm, Vincent Conitzer

    Abstract: Strategic interactions can be represented more concisely, and analyzed and solved more efficiently, if we are aware of the symmetries within the multiagent system. Symmetries also have conceptual implications, for example for equilibrium selection. We study the computational complexity of identifying and using symmetries. Using the classical framework of normal-form games, we consider game symmetr… ▽ More

    Submitted 27 February, 2025; v1 submitted 15 January, 2025; originally announced January 2025.

    Comments: Long and updated version to the published paper in the Proceedings of the 39th Annual AAAI Conference on Artificial Intelligence (AAAI 2025). 24 pages, 2 figures, 1 table

    MSC Class: 91A05; 91A06; 91A10; 91A26; 91A35; 91A68; 68Q17; 68Q25; 68T01 ACM Class: I.2; J.4; F.2

  15. arXiv:2412.19659  [pdf, ps, other

    cs.GT

    The Value of Recall in Extensive-Form Games

    Authors: Ratip Emin Berker, Emanuel Tewolde, Ioannis Anagnostides, Tuomas Sandholm, Vincent Conitzer

    Abstract: Imperfect-recall games, in which players may forget previously acquired information, have found many practical applications, ranging from game abstractions to team games and testing AI agents. In this paper, we quantify the utility gain by endowing a player with perfect recall, which we call the value of recall (VoR). While VoR can be unbounded in general, we parameterize it in terms of various ga… ▽ More

    Submitted 27 December, 2024; originally announced December 2024.

    Comments: 36 pages, 6 figures, main body to be published in Proceedings of the Thirty-Ninth AAAI Conference on Artificial Intelligence (AAAI-25), Philadelphia, Pennsylvania, USA, 2025

    MSC Class: 91A05; 91A06; 91A10; 91A11; 91A18; 91A35; 91A68; 68T37; 68Q17; 68Q25 ACM Class: F.2; I.2; J.4

  16. arXiv:2406.15970  [pdf, ps, other

    cs.GT cs.AI cs.CC

    Imperfect-Recall Games: Equilibrium Concepts and Their Complexity

    Authors: Emanuel Tewolde, Brian Hu Zhang, Caspar Oesterheld, Manolis Zampetakis, Tuomas Sandholm, Paul W. Goldberg, Vincent Conitzer

    Abstract: We investigate optimal decision making under imperfect recall, that is, when an agent forgets information it once held before. An example is the absentminded driver game, as well as team games in which the members have limited communication capabilities. In the framework of extensive-form games with imperfect recall, we analyze the computational complexities of finding equilibria in multiplayer se… ▽ More

    Submitted 22 June, 2024; originally announced June 2024.

    Comments: Long version of the paper that got accepted to the Thirty-Third International Joint Conference on Artificial Intelligence (IJCAI 2024). 35 pages, 10 figures, 1 table

    MSC Class: 91A05; 91A06; 91A10; 91A11; 91A18; 91A35; 91A68; 68T37; 68Q17; 68Q25 ACM Class: I.2; J.4; F.2

  17. arXiv:2404.10271  [pdf, other

    cs.LG cs.AI cs.CL cs.CY cs.GT

    Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback

    Authors: Vincent Conitzer, Rachel Freedman, Jobst Heitzig, Wesley H. Holliday, Bob M. Jacobs, Nathan Lambert, Milan Mossé, Eric Pacuit, Stuart Russell, Hailey Schoelkopf, Emanuel Tewolde, William S. Zwicker

    Abstract: Foundation models such as GPT-4 are fine-tuned to avoid unsafe or otherwise problematic behavior, such as helping to commit crimes or producing racist text. One approach to fine-tuning, called reinforcement learning from human feedback, learns from humans' expressed preferences over multiple outputs. Another approach is constitutional AI, in which the input from humans is a list of high-level prin… ▽ More

    Submitted 4 June, 2024; v1 submitted 15 April, 2024; originally announced April 2024.

    Comments: 15 pages, 4 figures

    MSC Class: 68T01; 68T50; 91B14; 91B12 ACM Class: I.2.0; I.2.7; K.4.2; I.2.m; J.4

  18. arXiv:2305.17805  [pdf, other

    cs.GT cs.AI cs.CC

    The Computational Complexity of Single-Player Imperfect-Recall Games

    Authors: Emanuel Tewolde, Caspar Oesterheld, Vincent Conitzer, Paul W. Goldberg

    Abstract: We study single-player extensive-form games with imperfect recall, such as the Sleeping Beauty problem or the Absentminded Driver game. For such games, two natural equilibrium concepts have been proposed as alternative solution concepts to ex-ante optimality. One equilibrium concept uses generalized double halving (GDH) as a belief system and evidential decision theory (EDT), and another one uses… ▽ More

    Submitted 28 May, 2023; originally announced May 2023.

    Comments: Long version of the paper that got accepted to the Thirty-Second International Joint Conference on Artificial Intelligence (IJCAI-23). 10 pages and 2 figures in the main body. 17 pages and 4 figures in the appendix

    MSC Class: 91A18; 68T37; 68Q17; 91A35 ACM Class: I.2; J.4; F.2

  19. arXiv:2111.00076  [pdf, other

    cs.GT econ.TH

    Game Transformations That Preserve Nash Equilibria or Best-Response Sets

    Authors: Emanuel Tewolde, Vincent Conitzer

    Abstract: In this paper, we investigate under which conditions normal-form games are (guaranteed to be) strategically equivalent. First, we show for N-player games (N >= 3) that (A) it is NP-hard to decide whether a given strategy is a best response to some strategy profile of the opponents, and that (B) it is co-NP-hard to decide whether two games have the same best-response sets. Combining that with… ▽ More

    Submitted 22 June, 2024; v1 submitted 29 October, 2021; originally announced November 2021.

    Comments: Long version of the paper that got accepted to the Thirty-Third International Joint Conference on Artificial Intelligence (IJCAI 2024). 28 pages, 1 figures

    MSC Class: 91A10 (Primary); 91A06 (Secondary) ACM Class: J.4