Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–12 of 12 results for author: Jo, N

Searching in archive cs. Search in all archives.
.
  1. arXiv:2604.21827  [pdf, ps, other

    cs.AI cs.HC

    Alignment has a Fantasia Problem

    Authors: Nathanael Jo, Zoe De Simone, Mitchell Gordon, Ashia Wilson

    Abstract: In accomplishing complex tasks, human cognition typically progresses from abstract to concrete (e.g., from brainstorming ideas to writing an essay). With the advent of highly capable AI assistants, people now offload various parts of their task to the system. However, instruction-tuned AI systems, even when designed to infer implicit intent, lack an understanding of humans' cognitive processes. Wh… ▽ More

    Submitted 6 August, 2026; v1 submitted 23 April, 2026; originally announced April 2026.

    Comments: 10 pages, 2 figures

    Journal ref: ICLR 2026 Workshop HCAIR

  2. arXiv:2604.03529  [pdf, ps, other

    cs.HC cs.AI cs.CY

    Incentives shape how humans co-create with generative AI

    Authors: Nathanael Jo, Manish Raghavan

    Abstract: Generative AI is quickly becoming an integral part of people's everyday workflows. Early evidence has shown that while generative AI can increase individual-level productivity, it does so at the cost of collective diversity, potentially narrowing the set of ideas and perspectives produced. Our research stands in contrast to this concern: through a pre-registered randomized control trial, we show t… ▽ More

    Submitted 3 April, 2026; originally announced April 2026.

  3. arXiv:2602.24086  [pdf, ps, other

    cs.CY cs.LG

    The Subjectivity of Monoculture

    Authors: Nathanael Jo, Nikhil Garg, Manish Raghavan

    Abstract: Machine learning models -- including large language models (LLMs) -- are often said to exhibit monoculture, where outputs agree strikingly often. But what does it actually mean for models to agree too much? We argue that this question is inherently subjective, relying on two key decisions. First, the analyst must specify a baseline null model for what "independence" should look like. This choice… ▽ More

    Submitted 27 February, 2026; originally announced February 2026.

  4. arXiv:2509.19590  [pdf, ps, other

    cs.AI cs.CY cs.LG

    Position: AI Evaluations Should be Grounded on a Theory of Capability

    Authors: Nathanael Jo, Ashia Wilson

    Abstract: Evaluations of generative models are now ubiquitous, and their outcomes critically shape public and scientific expectations of AI's capabilities. Yet skepticism about their reliability continues to grow. How can we know that a reported accuracy genuinely reflects a model's underlying performance? Although benchmark results are often presented as direct measurements of capability, in practice they… ▽ More

    Submitted 15 May, 2026; v1 submitted 23 September, 2025; originally announced September 2025.

    Comments: ICML 2026 Position Paper Track

  5. arXiv:2503.15634  [pdf, other

    cs.GT cs.CY

    Homogeneous Algorithms Can Reduce Competition in Personalized Pricing

    Authors: Nathanael Jo, Kathleen Creel, Ashia Wilson, Manish Raghavan

    Abstract: Firms' algorithm development practices are often homogeneous. Whether firms train algorithms on similar data, aim at similar benchmarks, or rely on similar pre-trained models, the result is correlated predictions. We model the impact of correlated algorithms on competition in the context of personalized pricing. Our analysis reveals that (1) higher correlation diminishes consumer welfare and (2) a… ▽ More

    Submitted 19 March, 2025; originally announced March 2025.

  6. arXiv:2501.14328  [pdf, other

    cs.CR cs.AR

    Securing DRAM at Scale: ARFM-Driven Row Hammer Defense with Unveiling the Threat of Short tRC Patterns

    Authors: Nogeun Joo, Donghyuk Kim, Hyunjun Cho, Junseok Noh, Dongha Jung, Joo-Young Kim

    Abstract: To address the issue of powerful row hammer (RH) attacks, our study involved an extensive analysis of the prevalent attack patterns in the field. We discovered a strong correlation between the timing and density of the active-to-active command period, ${tRC}$, and the likelihood of RH attacks. In this paper, we introduce MARC, an innovative ARFM-driven RH mitigation IP that significantly reinforce… ▽ More

    Submitted 24 January, 2025; originally announced January 2025.

    Comments: 12 pages, 19 figures

  7. arXiv:2310.01679  [pdf, other

    cs.LG cs.CY stat.ML

    Estimating and Implementing Conventional Fairness Metrics With Probabilistic Protected Features

    Authors: Hadi Elzayn, Emily Black, Patrick Vossler, Nathanael Jo, Jacob Goldin, Daniel E. Ho

    Abstract: The vast majority of techniques to train fair models require access to the protected attribute (e.g., race, gender), either at train time or in production. However, in many important applications this protected attribute is largely unavailable. In this paper, we develop methods for measuring and reducing fairness violations in a setting with limited access to protected attribute labels. Specifical… ▽ More

    Submitted 2 October, 2023; originally announced October 2023.

  8. arXiv:2307.15691  [pdf, ps, other

    stat.ML cs.LG math.OC

    ODTlearn: A Package for Learning Optimal Decision Trees for Prediction and Prescription

    Authors: Patrick Vossler, Nathan Justin, Sina Aghaei, Nathanael Jo, Andrés Gómez, Phebe Vayanos

    Abstract: ODTlearn is an open source Python package that provides methods for learning optimal decision trees for high-stakes predictive and prescriptive tasks based on the state-of-the-art mixed-integer optimization (MIO) framework proposed in Aghaei et al. (2025). The current version of the package provides implementations for learning optimal classification trees, optimal fair classification trees, optim… ▽ More

    Submitted 14 August, 2026; v1 submitted 28 July, 2023; originally announced July 2023.

    Comments: 9 pages, 2 figures

  9. arXiv:2212.01725  [pdf, ps, other

    cs.CY cs.LG

    Fairness in Contextual Resource Allocation Systems: Metrics and Incompatibility Results

    Authors: Nathanael Jo, Bill Tang, Kathryn Dullerud, Sina Aghaei, Eric Rice, Phebe Vayanos

    Abstract: We study critical systems that allocate scarce resources to satisfy basic needs, such as homeless services that provide housing. These systems often support communities disproportionately affected by systemic racial, gender, or other injustices, so it is crucial to design these systems with fairness considerations in mind. To address this problem, we propose a framework for evaluating fairness in… ▽ More

    Submitted 3 December, 2022; originally announced December 2022.

    Comments: To be published in 37th AAAI Conference on Artificial Intelligence

  10. arXiv:2201.09932  [pdf, other

    cs.LG cs.AI math.OC

    Learning Optimal Fair Classification Trees: Trade-offs Between Interpretability, Fairness, and Accuracy

    Authors: Nathanael Jo, Sina Aghaei, Andrés Gómez, Phebe Vayanos

    Abstract: The increasing use of machine learning in high-stakes domains -- where people's livelihoods are impacted -- creates an urgent need for interpretable, fair, and highly accurate algorithms. With these needs in mind, we propose a mixed integer optimization (MIO) framework for learning optimal classification trees -- one of the most interpretable models -- that can be augmented with arbitrary fairness… ▽ More

    Submitted 25 July, 2023; v1 submitted 24 January, 2022; originally announced January 2022.

  11. arXiv:2108.13628  [pdf, other

    cs.LG cs.CY stat.ML

    Learning Optimal Prescriptive Trees from Observational Data

    Authors: Nathanael Jo, Sina Aghaei, Andrés Gómez, Phebe Vayanos

    Abstract: We consider the problem of learning an optimal prescriptive tree (i.e., an interpretable treatment assignment policy in the form of a binary tree) of moderate depth, from observational data. This problem arises in numerous socially important domains such as public health and personalized medicine, where interpretable and data-driven interventions are sought based on data gathered in deployment --… ▽ More

    Submitted 24 July, 2023; v1 submitted 31 August, 2021; originally announced August 2021.

  12. arXiv:2102.00201  [pdf, other

    cs.SD cs.IR cs.LG cs.MM eess.AS

    Melon Playlist Dataset: a public dataset for audio-based playlist generation and music tagging

    Authors: Andres Ferraro, Yuntae Kim, Soohyeon Lee, Biho Kim, Namjun Jo, Semi Lim, Suyon Lim, Jungtaek Jang, Sehwan Kim, Xavier Serra, Dmitry Bogdanov

    Abstract: One of the main limitations in the field of audio signal processing is the lack of large public datasets with audio representations and high-quality annotations due to restrictions of copyrighted commercial music. We present Melon Playlist Dataset, a public dataset of mel-spectrograms for 649,091tracks and 148,826 associated playlists annotated by 30,652 different tags. All the data is gathered fr… ▽ More

    Submitted 30 January, 2021; originally announced February 2021.

    Comments: 2021 IEEE International Conference on Acoustics, Speech and Signal Processing