Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–24 of 24 results for author: Haupt, A

Searching in archive cs. Search in all archives.
.
  1. arXiv:2607.28617  [pdf, ps, other

    cs.AI cs.CL cs.CY cs.HC

    AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

    Authors: Xiangning Lin, Shenzhe Zhu, Shu Yang, Zhenyu Zhang, Haoqian Zhang, Yipeng Zhao, Chengxuan Qian, Tianwei Wang, Ziheng Zhang, Zhenlong Yuan, Dingcheng Wang, Juncheng Wu, Yuan Si, Jiaxin Liu, Baolong Bi, Robert Mahari, Tobin South, Dazza Greenwood, Zexue He, Rishi Bommasani, Sophia Kazinnik, Andreas Haupt, Samuele Marro, Erik Brynjolfsson, Alex Pentland , et al. (1 additional authors not shown)

    Abstract: System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are rarely disclosed to the public or regulators, creating a serious trust and accountability gap in the wide deployment of AI systems. In this paper, we introduce Artificial Intelligence System Prompt Assurance (AISPA), a us… ▽ More

    Submitted 6 August, 2026; v1 submitted 30 July, 2026; originally announced July 2026.

  2. arXiv:2605.30916  [pdf, ps, other

    cs.LG cs.GT econ.TH

    Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation

    Authors: Andreas Haupt, Justin Hartenstein, Anka Reuel, Mykel Kochenderfer, Sanmi Koyejo

    Abstract: AI benchmarks have well-documented limitations, with prior work examining contamination, saturation, and construct underspecification. Aggregation has received far less attention: benchmarks are typically summarized by uniformly averaging item-level scores, implicitly treating every test item as equally valuable. We model benchmarking as a multitask principal-agent game and show that the welfare l… ▽ More

    Submitted 29 May, 2026; originally announced May 2026.

  3. arXiv:2605.27996  [pdf, ps, other

    cs.AI

    Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure

    Authors: Max Lamparth, Daniel Fein, Andreas Haupt, Marcel Hussing, Mykel J. Kochenderfer

    Abstract: Single-axis mitigations of reward-model biases (e.g., reducing proxy reliance on length, sycophancy, or style) can rotate optimization pressure onto correlated proxies rather than eliminate it, a failure mode we call reward bias substitution. The failure is enabled by a measurement-versus-optimization gap between audit and policy-induced distributions during mitigation evaluation and policy traini… ▽ More

    Submitted 28 May, 2026; v1 submitted 27 May, 2026; originally announced May 2026.

    Comments: Improved readability (mostly appendix D)

  4. arXiv:2605.18721  [pdf, ps, other

    cs.LG cs.CL

    General Preference Reinforcement Learning

    Authors: Muhammad Umer, Muhammad Ahmed Mohsin, Ahsan Bilal, Arslan Chaudhry, Andreas Haupt, Sanmi Koyejo, Emily Fox, John M. Cioffi

    Abstract: Post-training has split large language model (LLM) alignment into two largely disconnected tracks. Online reinforcement learning (RL) with verifiable rewards drives emergent reasoning on math and code but depends on a programmatic verifier that cannot reach open-ended tasks, while preference optimization handles open-ended generation yet forgoes the continuous exploration that powers online RL. Cl… ▽ More

    Submitted 21 May, 2026; v1 submitted 18 May, 2026; originally announced May 2026.

  5. arXiv:2604.23575  [pdf, ps, other

    cs.CY cs.CL cs.LG

    The Collapse of Heterogeneity in Silicon Philosophers

    Authors: Yuanming Shi, Andreas Haupt

    Abstract: Silicon samples are increasingly used as a low-cost substitute for human panels and have been shown to reproduce aggregate human opinion with high fidelity. We show that, in the alignment-relevant domain of philosophy, silicon samples systematically collapse heterogeneity. Using data from $N = {277}$ professional philosophers drawn from PhilPeople profiles, we evaluate seven proprietary and open-s… ▽ More

    Submitted 29 April, 2026; v1 submitted 26 April, 2026; originally announced April 2026.

  6. arXiv:2601.22083  [pdf, ps, other

    cs.LG cs.AI

    Latent Adversarial Regularization for Offline Preference Optimization

    Authors: Enyi Jiang, Yibo Jacky Zhang, Yinglun Xu, Andreas Haupt, Nancy Amato, Sanmi Koyejo

    Abstract: Learning from human feedback typically relies on preference optimization that constrains policy updates through token-level regularization. However, preference optimization for language models is particularly challenging because token-space similarity does not imply semantic or behavioral similarity. To address this challenge, we leverage latent-space regularization for language model preference o… ▽ More

    Submitted 2 February, 2026; v1 submitted 29 January, 2026; originally announced January 2026.

  7. arXiv:2512.03399  [pdf, ps, other

    cs.LG

    Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of Value

    Authors: Joe Edelman, Tan Zhi-Xuan, Ryan Lowe, Oliver Klingefjord, Vincent Wang-Mascianica, Matija Franklin, Ryan Othniel Kearns, Ellie Hain, Atrisha Sarkar, Michiel Bakker, Fazl Barez, David Duvenaud, Jakob Foerster, Iason Gabriel, Joseph Gubbels, Bryce Goodman, Andreas Haupt, Jobst Heitzig, Julian Jara-Ettinger, Atoosa Kasirzadeh, James Ravi Kirkpatrick, Andrew Koh, W. Bradley Knox, Philipp Koralus, Joel Lehman , et al. (8 additional authors not shown)

    Abstract: Beneficial societal outcomes cannot be guaranteed by aligning individual AI systems with the intentions of their operators or users. Even an AI system that is perfectly aligned to the intentions of its operating organization can lead to bad outcomes if the goals of that organization are misaligned with those of other institutions and individuals. For this reason, we need full-stack alignment, the… ▽ More

    Submitted 2 December, 2025; originally announced December 2025.

  8. arXiv:2510.16972  [pdf, ps, other

    econ.TH cs.CY

    Preference Measurement Error, Concentration in Recommendation Systems, and Persuasion

    Authors: Andreas Haupt

    Abstract: Algorithmic recommendation based on noisy preference measurement is prevalent in recommendation systems. This paper discusses the consequences of such recommendation on market concentration and inequality. Binary types denoting a statistical majority and minority are noisily revealed through a statistical experiment. The achievable utilities and recommendation shares for the two groups can be anal… ▽ More

    Submitted 19 October, 2025; originally announced October 2025.

    Comments: 12 pages, 3 figures

  9. arXiv:2510.11834  [pdf, ps, other

    cs.LG cs.CL

    Don't Walk the Line: Boundary Guidance for Filtered Generation

    Authors: Sarah Ball, Andreas Haupt

    Abstract: Generative models are increasingly paired with safety classifiers that filter harmful or undesirable outputs. A common strategy is to fine-tune the generator to reduce the probability of being filtered, but this can be suboptimal: it often pushes the model toward producing samples near the classifier's decision boundary, increasing both false positives and false negatives. We propose Boundary Guid… ▽ More

    Submitted 4 August, 2026; v1 submitted 13 October, 2025; originally announced October 2025.

    Comments: Accepted at ICML 2026

  10. Scaling Human Judgment in Community Notes with LLMs

    Authors: Haiwen Li, Soham De, Manon Revel, Andreas Haupt, Brad Miller, Keith Coleman, Jay Baxter, Martin Saveski, Michiel A. Bakker

    Abstract: This paper argues for a new paradigm for Community Notes in the LLM era: an open ecosystem where both humans and LLMs can write notes, and the decision of which notes are helpful enough to show remains in the hands of humans. This approach can accelerate the delivery of notes, while maintaining trust and legitimacy through Community Notes' foundational principle: A community of diverse human rater… ▽ More

    Submitted 30 June, 2025; originally announced June 2025.

  11. arXiv:2506.19882  [pdf, ps, other

    cs.LG cs.AI cs.CL cs.CY

    Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track

    Authors: Rylan Schaeffer, Joshua Kazdan, Yegor Denisov-Blanch, Brando Miranda, Matthias Gerstgrasser, Susan Zhang, Andreas Haupt, Isha Gupta, Elyas Obbad, Jesse Dodge, Jessica Zosa Forde, Francesco Orabona, Sanmi Koyejo, David Donoho

    Abstract: Science progresses by iteratively advancing and correcting humanity's understanding of the world. In machine learning (ML) research, rapid advancements have led to an explosion of publications, but have also led to misleading, incorrect, flawed or perhaps even fraudulent studies being accepted and sometimes highlighted at ML conferences due to the fallibility of peer review. While such mistakes ar… ▽ More

    Submitted 6 July, 2025; v1 submitted 23 June, 2025; originally announced June 2025.

  12. arXiv:2410.23326  [pdf, other

    q-bio.QM cs.LG

    MassSpecGym: A benchmark for the discovery and identification of molecules

    Authors: Roman Bushuiev, Anton Bushuiev, Niek F. de Jonge, Adamo Young, Fleming Kretschmer, Raman Samusevich, Janne Heirman, Fei Wang, Luke Zhang, Kai Dührkop, Marcus Ludwig, Nils A. Haupt, Apurva Kalia, Corinna Brungs, Robin Schmid, Russell Greiner, Bo Wang, David S. Wishart, Li-Ping Liu, Juho Rousu, Wout Bittremieux, Hannes Rost, Tytus D. Mak, Soha Hassoun, Florian Huber , et al. (5 additional authors not shown)

    Abstract: The discovery and identification of molecules in biological and environmental samples is crucial for advancing biomedical and chemical sciences. Tandem mass spectrometry (MS/MS) is the leading technique for high-throughput elucidation of molecular structures. However, decoding a molecular structure from its mass spectrum is exceptionally challenging, even when performed by human experts. As a resu… ▽ More

    Submitted 14 February, 2025; v1 submitted 30 October, 2024; originally announced October 2024.

  13. arXiv:2410.16600  [pdf, ps, other

    cs.GT cs.AI cs.MA

    Convex Markov Games: A New Frontier for Multi-Agent Reinforcement Learning

    Authors: Ian Gemp, Andreas Haupt, Luke Marris, Siqi Liu, Georgios Piliouras

    Abstract: Behavioral diversity, expert imitation, fairness, safety goals and others give rise to preferences in sequential decision making domains that do not decompose additively across time. We introduce the class of convex Markov games that allow general convex preferences over occupancy measures. Despite infinite time horizon and strictly higher generality than Markov games, pure strategy Nash equilibri… ▽ More

    Submitted 16 June, 2025; v1 submitted 21 October, 2024; originally announced October 2024.

    Comments: Published at ICML 2025

  14. arXiv:2401.14446  [pdf, other

    cs.CY cs.AI cs.CR

    Black-Box Access is Insufficient for Rigorous AI Audits

    Authors: Stephen Casper, Carson Ezell, Charlotte Siegmann, Noam Kolt, Taylor Lynn Curtis, Benjamin Bucknall, Andreas Haupt, Kevin Wei, Jérémy Scheurer, Marius Hobbhahn, Lee Sharkey, Satyapriya Krishna, Marvin Von Hagen, Silas Alberti, Alan Chan, Qinyi Sun, Michael Gerovitch, David Bau, Max Tegmark, David Krueger, Dylan Hadfield-Menell

    Abstract: External audits of AI systems are increasingly recognized as a key mechanism for AI governance. The effectiveness of an audit, however, depends on the degree of access granted to auditors. Recent audits of state-of-the-art AI systems have primarily relied on black-box access, in which auditors can only query the system and observe its outputs. However, white-box access to the system's inner workin… ▽ More

    Submitted 29 May, 2024; v1 submitted 25 January, 2024; originally announced January 2024.

    Comments: FAccT 2024

    Journal ref: The 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT '24), June 3-6, 2024, Rio de Janeiro, Brazil

  15. arXiv:2306.05221  [pdf, ps, other

    cs.GT

    Steering No-Regret Learners to a Desired Equilibrium

    Authors: Brian Hu Zhang, Gabriele Farina, Ioannis Anagnostides, Federico Cacciamani, Stephen Marcus McAleer, Andreas Alexander Haupt, Andrea Celli, Nicola Gatti, Vincent Conitzer, Tuomas Sandholm

    Abstract: A mediator observes no-regret learners playing an extensive-form game repeatedly across $T$ rounds. The mediator attempts to steer players toward some desirable predetermined equilibrium by giving (nonnegative) payments to players. We call this the steering problem. The steering problem captures problems several problems of interest, among them equilibrium selection and information design (persuas… ▽ More

    Submitted 17 March, 2026; v1 submitted 8 June, 2023; originally announced June 2023.

  16. arXiv:2306.05216  [pdf, ps, other

    cs.GT

    Computing Optimal Equilibria and Mechanisms via Learning in Zero-Sum Extensive-Form Games

    Authors: Brian Hu Zhang, Gabriele Farina, Ioannis Anagnostides, Federico Cacciamani, Stephen Marcus McAleer, Andreas Alexander Haupt, Andrea Celli, Nicola Gatti, Vincent Conitzer, Tuomas Sandholm

    Abstract: We introduce a new approach for computing optimal equilibria via learning in games. It applies to extensive-form settings with any number of players, including mechanism design, information design, and solution concepts such as correlated, communication, and certification equilibria. We observe that optimal equilibria are minimax equilibrium strategies of a player in an extensive-form zero-sum gam… ▽ More

    Submitted 23 May, 2024; v1 submitted 8 June, 2023; originally announced June 2023.

  17. arXiv:2302.06559  [pdf, other

    cs.CY cs.GT cs.IR econ.TH

    Recommending to Strategic Users

    Authors: Andreas Haupt, Dylan Hadfield-Menell, Chara Podimata

    Abstract: Recommendation systems are pervasive in the digital economy. An important assumption in many deployed systems is that user consumption reflects user preferences in a static sense: users consume the content they like with no other considerations in mind. However, as we document in a large-scale online survey, users do choose content strategically to influence the types of content they get recommend… ▽ More

    Submitted 13 February, 2023; originally announced February 2023.

    Comments: 35 pages

  18. arXiv:2301.13449  [pdf, other

    cs.GT econ.TH

    Certification Design for a Competitive Market

    Authors: Andreas A. Haupt, Nicole Immorlica, Brendan Lucier

    Abstract: Motivated by applications such as voluntary carbon markets and educational testing, we consider a market for goods with varying but hidden levels of quality in the presence of a third-party certifier. The certifier can provide informative signals about the quality of products, and can charge for this service. Sellers choose both the quality of the product they produce and a certification. Prices a… ▽ More

    Submitted 31 January, 2023; originally announced January 2023.

    Comments: 22 pages, 1 figure

  19. arXiv:2208.10469  [pdf, other

    cs.AI cs.GT cs.MA econ.TH

    Formal Contracts Mitigate Social Dilemmas in Multi-Agent RL

    Authors: Andreas A. Haupt, Phillip J. K. Christoffersen, Mehul Damani, Dylan Hadfield-Menell

    Abstract: Multi-agent Reinforcement Learning (MARL) is a powerful tool for training autonomous agents acting independently in a common environment. However, it can lead to sub-optimal behavior when individual incentives and group incentives diverge. Humans are remarkably capable at solving these social dilemmas. It is an open problem in MARL to replicate such cooperative behaviors in selfish agents. In this… ▽ More

    Submitted 29 January, 2024; v1 submitted 22 August, 2022; originally announced August 2022.

  20. arXiv:2208.01534  [pdf, other

    cs.IR cs.AI cs.HC

    Towards Psychologically-Grounded Dynamic Preference Models

    Authors: Mihaela Curmei, Andreas Haupt, Dylan Hadfield-Menell, Benjamin Recht

    Abstract: Designing recommendation systems that serve content aligned with time varying preferences requires proper accounting of the feedback effects of recommendations on human behavior and psychological condition. We argue that modeling the influence of recommendations on people's preferences must be grounded in psychologically plausible models. We contribute a methodology for developing grounded dynamic… ▽ More

    Submitted 6 August, 2022; v1 submitted 1 August, 2022; originally announced August 2022.

    Comments: In Sixteenth ACM Conference on Recommender Systems, September 18-23, 2022, Seattle, WA, USA, 14 pages

  21. arXiv:2205.04619  [pdf, other

    cs.LG cs.AI econ.TH

    Risk Preferences of Learning Algorithms

    Authors: Andreas Haupt, Aroon Narayanan

    Abstract: Agents' learning from feedback shapes economic outcomes, and many economic decision-makers today employ learning algorithms to make consequential choices. This note shows that a widely used learning algorithm, $\varepsilon$-Greedy, exhibits emergent risk aversion: it prefers actions with lower variance. When presented with actions of the same expectation, under a wide range of conditions,… ▽ More

    Submitted 12 December, 2023; v1 submitted 9 May, 2022; originally announced May 2022.

    Comments: 11 pages, 6 figures

  22. arXiv:2107.10323  [pdf, other

    cs.GT econ.TH

    The Optimality of Upgrade Pricing

    Authors: Dirk Bergemann, Alessandro Bonatti, Andreas Haupt, Alex Smolin

    Abstract: We consider a multiproduct monopoly pricing model. We provide sufficient conditions under which the optimal mechanism can be implemented via upgrade pricing -- a menu of product bundles that are nested in the strong set order. Our approach exploits duality methods to identify conditions on the distribution of consumer types under which (a) each product is purchased by the same set of buyers as und… ▽ More

    Submitted 2 December, 2021; v1 submitted 21 July, 2021; originally announced July 2021.

    Comments: 22 pages, 4 figures

    Report number: Web and Internet Economics 2021

  23. arXiv:2103.14375  [pdf, other

    cs.LG cs.GT

    Prior-Independent Auctions for the Demand Side of Federated Learning

    Authors: Andreas Haupt, Vaikkunth Mugunthan

    Abstract: Federated learning (FL) is a paradigm that allows distributed clients to learn a shared machine learning model without sharing their sensitive training data. While largely decentralized, FL requires resources to fund a central orchestrator or to reimburse contributors of datasets to incentivize participation. Inspired by insights from prior-independent auction design, we propose a mechanism, FIPIA… ▽ More

    Submitted 13 April, 2021; v1 submitted 26 March, 2021; originally announced March 2021.

  24. arXiv:1710.08878  [pdf, ps, other

    cs.LG stat.ML

    Classification on Large Networks: A Quantitative Bound via Motifs and Graphons

    Authors: Andreas Haupt, Mohammad Khatami, Thomas Schultz, Ngoc Mai Tran

    Abstract: When each data point is a large graph, graph statistics such as densities of certain subgraphs (motifs) can be used as feature vectors for machine learning. While intuitive, motif counts are expensive to compute and difficult to work with theoretically. Via graphon theory, we give an explicit quantitative bound for the ability of motif homomorphisms to distinguish large networks under both generat… ▽ More

    Submitted 24 October, 2017; originally announced October 2017.

    Comments: 17 pages, 2 figures, 1 table

    MSC Class: 68T05; 05C80; 62G99