Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 442 results for author: Jordan, M

Searching in archive cs. Search in all archives.
.
  1. arXiv:2607.26358  [pdf, ps, other

    cs.LG cs.AI cs.GT

    Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning

    Authors: Keegan Harris, Brian W. Lee, Ian Waudby-Smith, Philip Amortila, Nika Haghtalab, Michael I. Jordan

    Abstract: Reinforcement learning (RL) fine-tuning is widely used in language model training to improve model performance on a target task while limiting drift from a reference policy. A standard way to balance this trade-off is via a KL-regularized RL objective, although this formulation does not by itself provide a principled way to set the regularization coefficient. In practice, the coefficient is typica… ▽ More

    Submitted 28 July, 2026; originally announced July 2026.

  2. arXiv:2607.21190  [pdf

    cs.CV

    Physics-Informed Deep Learning Model for Cross-Modality Super-Resolution in Fluorescence Microscopy

    Authors: Mohammad Soltaninezhad, Elena Corbetta, Francisco Paez Larios, Paul M. Jordan, Oliver Werz, Christian Eggeling, Thomas Bocklitz

    Abstract: Cross-modality image translation offers a route to super-resolution fluorescence microscopy from low-resolution images while reducing phototoxicity and instrumentation demands. However, purely data-driven models can produce visually plausible outputs that are inconsistent with optical image formation. Here, we propose a physics-informed generative adversarial network for confocal-to-STED image tra… ▽ More

    Submitted 23 July, 2026; originally announced July 2026.

  3. arXiv:2607.08012  [pdf, ps, other

    cs.LG cs.AI cs.GT

    Provably Optimal Learning Algorithms for Assistance Games

    Authors: Nivasini Ananthakrishnan, Mark Bedaywi, Michael I. Jordan, Stuart Russell, Nika Haghtalab

    Abstract: This paper studies an online variant of the assistance games framework, where an informed agent and an uninformed agent repeatedly interact over $T$ timesteps to optimize a common reward function. While the informed agent (the human) observes a latent state of the world, the uninformed agent (the assistant) observes only the human's actions. We provide the first provably efficient learning algorit… ▽ More

    Submitted 8 July, 2026; originally announced July 2026.

  4. arXiv:2606.27315  [pdf, ps, other

    cs.LG

    Blackwell Approachability and Gradient Equilibrium are Equivalent

    Authors: Brian W. Lee, Nika Haghtalab, Michael I. Jordan, Ryan J. Tibshirani

    Abstract: Gradient equilibrium (GEQ) is a recently introduced online optimization framework that generalizes first-order stationarity from offline optimization and abstracts problems like online conformal prediction. While GEQ has curious similarities with known online learning frameworks, namely regret minimization, prior work has shown that GEQ error and regret are incomparable objectives, leaving open a… ▽ More

    Submitted 25 June, 2026; originally announced June 2026.

    Comments: 30 pages, 1 figure, accepted for presentation at COLT 2026

  5. arXiv:2606.19526  [pdf, ps, other

    cs.AR

    SPINE: A Fault Injection Profiler for Quantized Neural Networks under Accumulated Faults

    Authors: Nathan Guimarães, Ian Kersz, Leonardo R. Gobatto, Fabio Benevenuti, Michael G. Jordan, Antonio Carlos S. Beck, Fernanda L. Kastensmidt, Jose Rodrigo Azambuja

    Abstract: Deploying deep neural networks at the edge demands efficient inference under strict cost and power constraints. Quantized neural networks address these demands by replacing floating-point parameters with low-precision integers, yet their weights remain continuously exposed to radiation-induced bit-flips during inference. Fault Injection can be used to simulate those environments, but existing stud… ▽ More

    Submitted 17 June, 2026; originally announced June 2026.

    Comments: ACM/IEEE/SBC/SBMICRO Symposium on Integrated Circuits and Systems Design 2026

    ACM Class: B.8.1

  6. arXiv:2606.18867  [pdf, ps, other

    cs.LG cs.CY stat.ML

    Strategic Feature Selection

    Authors: Jivat Neet Kaur, Pratik Patil, Divya Shanmugam, Emma Pierson, Michael I. Jordan, Nika Haghtalab, Meena Jagadeesan, Ahmed Alaa, Serena Wang

    Abstract: When algorithmic predictors inform resource allocation in high-stakes domains such as healthcare, these predictors must account for strategic manipulation of input features. The typical solution is to redesign the predictor itself to explicitly account for strategic interactions. In practice, however, decision makers are often constrained to adjusting coarser levers within existing prediction pipe… ▽ More

    Submitted 17 June, 2026; originally announced June 2026.

  7. arXiv:2606.04665  [pdf, ps, other

    cs.LG

    Towards Accurate Model Selection in Deep Unsupervised Domain Adaptation

    Authors: Kaichao You, Ximei Wang, Mingsheng Long, Michael I. Jordan

    Abstract: Deep unsupervised domain adaptation (Deep UDA) methods successfully leverage rich labeled data in a source domain to boost the performance on related but unlabeled data in a target domain. However, algorithm comparison is cumbersome in Deep UDA due to the absence of accurate and standardized model selection method, posing an obstacle to further advances in the field. Existing model selection metho… ▽ More

    Submitted 3 June, 2026; originally announced June 2026.

    Comments: upload to arxiv for record

  8. arXiv:2605.30853  [pdf, ps, other

    math.OC cs.CC cs.DM math.CO

    Diffusion-Robust Optimization over Graphs

    Authors: Liviu Aolaritei, Ricky Huang, Michael I. Jordan, Paul Grigas

    Abstract: We introduce a diffusion-based uncertainty model for robust optimization on directed graphs, in which perturbations of edge weights propagate along adjacent edges and satisfy conservation constraints at nodes. This topology-aware structure is natural in networked systems where uncertainty is induced by flows and local interactions, including transportation, logistics, communication, and energy net… ▽ More

    Submitted 19 June, 2026; v1 submitted 29 May, 2026; originally announced May 2026.

    Comments: 46 pages, 6 figures

  9. arXiv:2605.30188  [pdf, ps, other

    cs.LG cs.AI stat.ML

    CalArena: A Large-Scale Post-Hoc Calibration Benchmark

    Authors: Eugène Berta, David Holzmüller, Francis Bach, Michael I. Jordan

    Abstract: Reliable probability estimates are critical in many machine learning applications, yet modern classifiers are often poorly calibrated. Post-hoc calibration provides a simple and widely used solution, but the large number of proposed methods, combined with small-scale and inconsistent evaluations, makes it difficult to determine which approaches are truly effective in practice. We introduce a large… ▽ More

    Submitted 1 June, 2026; v1 submitted 28 May, 2026; originally announced May 2026.

    Comments: 30 pages, 9 figures

  10. arXiv:2605.06987  [pdf, ps, other

    cs.LG cs.GT econ.TH stat.ML

    Response Time Enhances Alignment with Heterogeneous Preferences

    Authors: Federico Echenique, Alireza Fallah, Baihe Huang, Michael I. Jordan

    Abstract: Aligning large language models (LLMs) to human preferences typically relies on aggregating pooled feedback into a single reward model. However, this standard approach assumes that all labelers share the same underlying preferences, ignoring the fact that real-world labelers are highly heterogeneous and usually anonymous. Consequently, relying solely on binary choice data fundamentally distorts the… ▽ More

    Submitted 7 May, 2026; originally announced May 2026.

  11. arXiv:2605.06210  [pdf, ps, other

    stat.ML cs.AI cs.LG stat.AP stat.ME

    Super-Level-Set Regression: Conditional Quantiles via Volume Minimization

    Authors: Sacha Braun, Michael I. Jordan, Francis Bach

    Abstract: Constructing minimum-volume prediction regions that satisfy conditional coverage is a fundamental challenge in multivariate regression. Standard approaches rely on explicitly estimating the full conditional density and subsequently thresholding it. This two-step plug-in process is notoriously difficult, sensitive to estimation errors, and computationally expensive. One would like to instead optimi… ▽ More

    Submitted 7 May, 2026; originally announced May 2026.

  12. arXiv:2605.04266  [pdf, ps, other

    cs.LG stat.ML

    Explaining and Preventing Alignment Collapse in Iterative RLHF

    Authors: Etienne Gauthier, Francis Bach, Michael I. Jordan

    Abstract: Reinforcement learning from human feedback (RLHF) typically assumes a static or non-strategic reward model (RM). In iterative deployment, however, the policy generates the data on which the RM is retrained, creating a feedback loop. Building on the Stackelberg game formulation of this interaction, we derive an analytical decomposition of the policy's true optimization gradient into a standard poli… ▽ More

    Submitted 5 May, 2026; originally announced May 2026.

    Comments: Code at: https://github.com/GauthierE/fpo

  13. arXiv:2603.22167  [pdf, ps, other

    cs.LG cs.AI cs.GT econ.TH

    Calibeating Made Simple

    Authors: Yurong Chen, Zhiyi Huang, Michael I. Jordan, Haipeng Luo

    Abstract: We study calibeating, the problem of post-processing external forecasts online to minimize cumulative losses and match an informativeness-based benchmark. Unlike prior work, which analyzed calibeating for specific losses with specific arguments, we reduce calibeating to existing online learning techniques and obtain results for general proper losses. More concretely, we first show that calibeating… ▽ More

    Submitted 23 March, 2026; originally announced March 2026.

  14. arXiv:2603.17925  [pdf, ps, other

    stat.ME cs.LG math.ST

    Multi-Armed Sequential Hypothesis Testing by Betting

    Authors: Ricardo J. Sandoval, Ian Waudby-Smith, Michael I. Jordan

    Abstract: We consider a variant of sequential testing by betting where, at each time step, the statistician is presented with multiple data sources (arms) and obtains data by choosing one of the arms. We consider the composite global null hypothesis $\mathscr{P}$ that all arms are null in a certain sense (e.g. all dosages of a treatment are ineffective) and we are interested in rejecting $\mathscr{P}$ in fa… ▽ More

    Submitted 4 June, 2026; v1 submitted 18 March, 2026; originally announced March 2026.

  15. arXiv:2603.05619  [pdf, ps, other

    stat.AP cs.GT

    Test-then-Punish: A Statistical Approach to Repeated Games

    Authors: Aymeric Capitaine, Antoine Scheid, Etienne Boursier, Alain Durmus, Michael I. Jordan

    Abstract: We study discounted infinitely repeated games in which players agree on a cooperative mixed action profile but, at each step, observe only the realized pure actions. This form of imperfect monitoring breaks classical trigger strategies, since deviations cannot be identified with certainty. To address this problem, we study how hypothesis testing can be used to sustain cooperation. First, we develo… ▽ More

    Submitted 5 March, 2026; originally announced March 2026.

  16. arXiv:2602.24230  [pdf, ps, other

    stat.ML cs.LG

    A Variational Estimator for $L_p$ Calibration Errors

    Authors: Eugène Berta, Sacha Braun, David Holzmüller, Francis Bach, Michael I. Jordan

    Abstract: Calibration$\unicode{x2014}$the problem of ensuring that predicted probabilities align with observed class frequencies$\unicode{x2014}$is a basic desideratum for reliable prediction with machine learning systems. Calibration error is traditionally assessed via a divergence function, using the expected divergence between predictions and empirical frequencies. Accurately estimating this quantity is… ▽ More

    Submitted 27 February, 2026; originally announced February 2026.

  17. arXiv:2602.17608  [pdf, ps, other

    cs.LG cs.AI stat.ML

    Towards Anytime-Valid Statistical Watermarking

    Authors: Baihe Huang, Eric Xu, Kannan Ramchandran, Jiantao Jiao, Michael I. Jordan

    Abstract: The proliferation of Large Language Models (LLMs) necessitates efficient mechanisms to distinguish machine-generated content from human text. While statistical watermarking has emerged as a promising solution, existing methods suffer from two critical limitations: the lack of a principled approach for selecting sampling distributions and the reliance on fixed-horizon hypothesis testing, which prec… ▽ More

    Submitted 19 February, 2026; originally announced February 2026.

  18. arXiv:2602.16919  [pdf, ps, other

    cs.GT

    Signaling in Data Markets via Free Samples

    Authors: Nivasini Ananthakrishnan, Alireza Fallah, Michael I. Jordan

    Abstract: We study a setting in which a data buyer seeks to estimate an unknown parameter by purchasing samples from one of K data sellers. Each seller has privately known data quality (e.g., high vs. low variance) and a private per-sample cost. We consider a multi-stage game in which the first stage is a free-trial stage in which the sellers have the option of signaling data quality by offering a few sampl… ▽ More

    Submitted 18 February, 2026; originally announced February 2026.

  19. arXiv:2602.14952  [pdf, ps, other

    cs.LG math.OC stat.ME stat.ML

    Locally Adaptive Multi-Objective Learning

    Authors: Jivat Neet Kaur, Isaac Gibbs, Michael I. Jordan

    Abstract: We consider the general problem of learning a predictor that satisfies multiple objectives of interest simultaneously, a broad framework that captures a range of specific learning goals including calibration, regret, and multiaccuracy. We work in an online setting where the data distribution can change arbitrarily over time. Existing approaches to this problem aim to minimize the set of objectives… ▽ More

    Submitted 16 February, 2026; originally announced February 2026.

    Comments: Code is available at https://github.com/jivatneet/adaptive-multiobjective

  20. arXiv:2602.12237  [pdf, ps, other

    cs.LG cs.AI cs.CL

    Olmix: A Framework for Data Mixing Throughout LM Development

    Authors: Mayee F. Chen, Tyler Murray, David Heineman, Matt Jordan, Hannaneh Hajishirzi, Christopher Ré, Luca Soldaini, Kyle Lo

    Abstract: Data mixing -- determining the ratios of data from different domains -- is a first-order concern for training language models (LMs). While existing mixing methods show promise, they fall short when applied during real-world LM development. We present Olmix, a framework that addresses two such challenges. First, the configuration space for developing a mixing method is not well understood -- design… ▽ More

    Submitted 12 February, 2026; originally announced February 2026.

  21. arXiv:2602.12180  [pdf, ps, other

    cs.LG cs.GT

    How Sampling Shapes LLM Alignment: From One-Shot Optima to Iterative Dynamics

    Authors: Yurong Chen, Yu He, Michael I. Jordan, Fan Yao

    Abstract: Standard methods for aligning large language models with human preferences learn from pairwise comparisons among sampled candidate responses and regularize toward a reference policy. Despite their effectiveness, the effects of sampling and reference choices are poorly understood theoretically. We investigate these effects through Identity Preference Optimization, a widely used preference alignment… ▽ More

    Submitted 12 February, 2026; originally announced February 2026.

  22. arXiv:2602.03381  [pdf, ps, other

    cs.GT

    Dynamic Programming for Epistemic Uncertainty in Markov Decision Processes

    Authors: Axel Benyamine, Julien Grand-Clément, Marek Petrik, Michael I. Jordan, Alain Durmus

    Abstract: In this paper, we propose a general theory of ambiguity-averse MDPs, which treats the uncertain transition probabilities as random variables and evaluates a policy via a risk measure applied to its random return. This ambiguity-averse MDP framework unifies several models of MDPs with epistemic uncertainty for specific choices of risk measures. We extend the concepts of value functions and Bellman… ▽ More

    Submitted 7 August, 2026; v1 submitted 3 February, 2026; originally announced February 2026.

  23. arXiv:2601.05427  [pdf, ps, other

    cs.GT

    Anytime Detection of Strategic Deviations in Multi-Agent Systems

    Authors: Etienne Gauthier, Francis Bach, Michael I. Jordan

    Abstract: In many multi-agent systems, agents interact repeatedly and are expected to settle into stable, rational behavior over time. Yet in practice, behavior often drifts, and detecting such deviations in real time remains an open challenge. We introduce a sequential testing framework that monitors whether observed play is consistent with a benchmark of strategic behavior, without assuming a fixed sample… ▽ More

    Submitted 22 May, 2026; v1 submitted 8 January, 2026; originally announced January 2026.

    Comments: Code at: https://github.com/GauthierE/anytime-detection-deviation

  24. arXiv:2512.13961  [pdf, ps, other

    cs.CL cs.LG

    Olmo 3

    Authors: Team Olmo, :, Allyson Ettinger, Amanda Bertsch, Bailey Kuehl, David Graham, David Heineman, Dirk Groeneveld, Faeze Brahman, Finbarr Timbers, Hamish Ivison, Jacob Morrison, Jake Poznanski, Kyle Lo, Luca Soldaini, Matt Jordan, Mayee Chen, Michael Noukhovitch, Nathan Lambert, Pete Walsh, Pradeep Dasigi, Robert Berry, Saumya Malik, Saurabh Shah, Scott Geng , et al. (44 additional authors not shown)

    Abstract: We introduce Olmo 3, a family of state-of-the-art, fully-open language models at the 7B and 32B parameter scales. Olmo 3 model construction targets long-context reasoning, function calling, coding, instruction following, general chat, and knowledge recall. This release includes the entire model flow, i.e., the full lifecycle of the family of models, including every stage, checkpoint, data point, a… ▽ More

    Submitted 14 April, 2026; v1 submitted 15 December, 2025; originally announced December 2025.

    Comments: minor edit updates

  25. arXiv:2512.13123  [pdf, ps, other

    math.OC cs.LG math.ST stat.ML

    Stopping Rules for Stochastic Gradient Descent via Anytime-Valid Confidence Sequences

    Authors: Liviu Aolaritei, Michael I. Jordan

    Abstract: The problem of stopping stochastic gradient descent (SGD) in an online manner, based solely on the observed trajectory, is a challenging theoretical problem with significant consequences for applications. While SGD is routinely monitored as it runs, the classical theory of SGD provides guarantees only at pre-specified iteration horizons and offers no valid way to decide, based on the observed traj… ▽ More

    Submitted 20 February, 2026; v1 submitted 15 December, 2025; originally announced December 2025.

  26. arXiv:2512.11779  [pdf, ps, other

    stat.ML cs.AI cs.LG

    Conditional Coverage Diagnostics for Conformal Prediction

    Authors: Sacha Braun, David Holzmüller, Michael I. Jordan, Francis Bach

    Abstract: Evaluating conditional coverage remains one of the most persistent challenges in assessing the reliability of predictive systems. Although conformal methods can give guarantees on marginal coverage, no method can guarantee to produce sets with correct conditional coverage, leaving practitioners without a clear way to interpret local deviations. To overcome sample-inefficiency and overfitting issue… ▽ More

    Submitted 29 May, 2026; v1 submitted 12 December, 2025; originally announced December 2025.

  27. arXiv:2511.03685  [pdf, ps, other

    cs.LG cs.AI

    Structured Matrix Scaling for Multi-Class Calibration

    Authors: Eugène Berta, David Holzmüller, Michael I. Jordan, Francis Bach

    Abstract: Post-hoc recalibration methods are widely used to ensure that classifiers provide faithful probability estimates. We argue that parametric recalibration functions based on logistic regression can be motivated from a simple theoretical setting for both binary and multiclass classification. This insight motivates the use of more expressive calibration methods beyond standard temperature scaling. For… ▽ More

    Submitted 10 March, 2026; v1 submitted 5 November, 2025; originally announced November 2025.

  28. arXiv:2510.25458  [pdf, ps, other

    cs.LG cs.AI stat.ML

    Scalable Utility-Aware Multiclass Calibration

    Authors: Mahmoud Hegazy, Michael I. Jordan, Aymeric Dieuleveut

    Abstract: Ensuring that classifiers are well-calibrated, i.e., their predictions align with observed frequencies, is a minimal and fundamental requirement for classifiers to be viewed as trustworthy. Existing methods for assessing multiclass calibration often focus on specific aspects associated with prediction (e.g., top-class confidence, class-wise calibration) or utilize computationally challenging varia… ▽ More

    Submitted 29 October, 2025; originally announced October 2025.

  29. arXiv:2510.04318  [pdf, ps, other

    stat.ML cs.LG

    Adaptive Coverage Policies in Conformal Prediction

    Authors: Etienne Gauthier, Francis Bach, Michael I. Jordan

    Abstract: Traditional conformal prediction methods construct prediction sets such that the true label falls within the set with a user-specified coverage level. However, poorly chosen coverage levels can result in uninformative predictions, either producing overly conservative sets when the coverage level is too high, or empty sets when it is too low. Moreover, the fixed coverage level cannot adapt to the s… ▽ More

    Submitted 2 April, 2026; v1 submitted 5 October, 2025; originally announced October 2025.

    Comments: Code at: https://github.com/GauthierE/adaptive-coverage-policies

  30. arXiv:2509.14158  [pdf, ps, other

    cs.LG math.OC

    A Compositional Kernel Model for Feature Learning

    Authors: Feng Ruan, Keli Liu, Michael Jordan

    Abstract: We study a compositional variant of kernel ridge regression in which the predictor is applied to a coordinate-wise reweighting of the inputs. Formulated as a variational problem, this model provides a simple testbed for feature learning in compositional architectures. From the perspective of variable selection, we show how relevant variables are recovered while noise variables are eliminated. We e… ▽ More

    Submitted 3 November, 2025; v1 submitted 17 September, 2025; originally announced September 2025.

    Comments: Fix Typos

  31. arXiv:2508.20869  [pdf, ps, other

    cs.SD cs.CL cs.LG eess.AS

    OLMoASR: Open Models and Data for Training Robust Speech Recognition Models

    Authors: Huong Ngo, Matt Deitke, Martijn Bartelds, Sarah Pratt, Josh Gardner, Matt Jordan, Ludwig Schmidt

    Abstract: Improvements in training data scale and quality have led to significant advances, yet its influence in speech recognition remains underexplored. In this paper, we present a large-scale dataset, OLMoASR-Pool, and series of models, OLMoASR, to study and develop robust zero-shot speech recognition models. Beginning from OLMoASR-Pool, a collection of 3M hours of English audio and 17M transcripts, we d… ▽ More

    Submitted 28 August, 2025; originally announced August 2025.

    Comments: 17 pages, 7 figures

  32. arXiv:2508.17622  [pdf, ps, other

    stat.ML cs.LG econ.TH math.OC

    The Statistical Fairness-Accuracy Frontier

    Authors: Alireza Fallah, Michael I. Jordan, Annie Ulichney

    Abstract: We study fairness-accuracy tradeoffs when a single predictive model must serve multiple demographic groups. A useful tool for understanding this tradeoff is the fairness-accuracy (FA) Pareto frontier, which characterizes the set of models that cannot be improved in either fairness or accuracy without worsening the other. While characterizing the FA frontier requires full knowledge of the data dist… ▽ More

    Submitted 16 February, 2026; v1 submitted 24 August, 2025; originally announced August 2025.

  33. arXiv:2507.20941  [pdf, ps, other

    stat.ML cs.AI cs.LG stat.ME stat.OT

    Multivariate Standardized Residuals for Conformal Prediction

    Authors: Sacha Braun, Eugène Berta, Michael I. Jordan, Francis Bach

    Abstract: While split conformal prediction guarantees marginal coverage, approaching the stronger property of conditional coverage is essential for reliable uncertainty quantification. Naive conformal scores, however, suffer from poor conditional coverage in heteroskedastic settings. In univariate regression, this is commonly addressed by normalizing non-conformity scores using an estimated local score vari… ▽ More

    Submitted 7 May, 2026; v1 submitted 28 July, 2025; originally announced July 2025.

  34. arXiv:2507.20403  [pdf, ps, other

    econ.TH cs.LG

    A General Framework for Estimating Preferences Using Response Time Data

    Authors: Federico Echenique, Alireza Fallah, Michael I. Jordan

    Abstract: We propose a general methodology for recovering preference parameters from data on choices and response times. Our methods yield estimates with fast ($1/n$ for $n$ data points) convergence rates when specialized to the popular Drift Diffusion Model (DDM), but are broadly applicable to generalizations of the DDM as well as to alternative models of decision making that make use of response time data… ▽ More

    Submitted 31 July, 2025; v1 submitted 27 July, 2025; originally announced July 2025.

  35. arXiv:2507.06268  [pdf, ps, other

    cs.CY cs.AI stat.ML

    A Collectivist, Economic Perspective on AI

    Authors: Michael I. Jordan

    Abstract: Information technology is in the midst of a revolution in which omnipresent data collection and machine learning are impacting the human world as never before. The word ``intelligence'' is being used as a North Star for the development of this technology, with human cognition viewed as a baseline. This view neglects the fact that humans are social animals and that much of our intelligence is socia… ▽ More

    Submitted 15 December, 2025; v1 submitted 7 July, 2025; originally announced July 2025.

  36. arXiv:2507.06261  [pdf, ps, other

    cs.CL cs.AI

    Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

    Authors: Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, Luke Marris, Sam Petulla, Colin Gaffney, Asaf Aharoni, Nathan Lintz, Tiago Cardal Pais, Henrik Jacobsson, Idan Szpektor, Nan-Jiang Jiang, Krishna Haridasan, Ahmed Omran, Nikunj Saunshi, Dara Bahri, Gaurav Mishra, Eric Chu , et al. (3410 additional authors not shown)

    Abstract: In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our most capable model yet, achieving SoTA performance on frontier coding and reasoning benchmarks. In addition to its incredible coding and reasoning skills, Gemini 2.5 Pro is a thinking model that excels at multimodal unde… ▽ More

    Submitted 19 December, 2025; v1 submitted 7 July, 2025; originally announced July 2025.

    Comments: 72 pages, 17 figures

  37. arXiv:2506.20173  [pdf, ps, other

    stat.ML cs.AI cs.LG stat.ME stat.OT

    Valid Selection among Conformal Sets

    Authors: Mahmoud Hegazy, Liviu Aolaritei, Michael I. Jordan, Aymeric Dieuleveut

    Abstract: Conformal prediction offers a distribution-free framework for constructing prediction sets with coverage guarantees. In practice, multiple valid conformal prediction sets may be available, arising from different models or methodologies. However, selecting the most desirable set, such as the smallest, can invalidate the coverage guarantees. To address this challenge, we propose a stability-based ap… ▽ More

    Submitted 25 June, 2025; originally announced June 2025.

  38. arXiv:2506.13488  [pdf, ps, other

    cs.LG physics.optics quant-ph

    Imaging at the quantum limit with convolutional neural networks

    Authors: Andrew H. Proppe, Aaron Z. Goldberg, Guillaume Thekkadath, Noah Lupu-Gladstein, Kyle M. Jordan, Philip J. Bustard, Frédéric Bouchard, Duncan England, Khabat Heshami, Jeff S. Lundeen, Benjamin J. Sussman

    Abstract: Deep neural networks have been shown to achieve exceptional performance for computer vision tasks like image recognition, segmentation, and reconstruction or denoising. Here, we evaluate the ultimate performance limits of deep convolutional neural network models for image reconstruction, by comparing them against the standard quantum limit set by shot-noise and the Heisenberg limit on precision. W… ▽ More

    Submitted 16 June, 2025; originally announced June 2025.

  39. arXiv:2506.10887  [pdf, ps, other

    cs.CL cs.LG

    Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers

    Authors: Yixiao Huang, Hanlin Zhu, Tianyu Guo, Jiantao Jiao, Somayeh Sojoudi, Michael I. Jordan, Stuart Russell, Song Mei

    Abstract: Large language models (LLMs) can acquire new knowledge through fine-tuning, but this process exhibits a puzzling duality: models can generalize remarkably from new facts, yet are also prone to hallucinating incorrect information. However, the reasons for this phenomenon remain poorly understood. In this work, we argue that both behaviors stem from a single mechanism known as out-of-context reasoni… ▽ More

    Submitted 2 February, 2026; v1 submitted 12 June, 2025; originally announced June 2025.

    Comments: NeurIPS 2025, first three authors contributed equally

  40. arXiv:2506.10354  [pdf, ps, other

    math.ST cs.IT

    Revisiting mean estimation over $\ell_p$ balls: Is the MLE optimal?

    Authors: Liviu Aolaritei, Michael I. Jordan, Reese Pathak, Annie Ulichney

    Abstract: We revisit the problem of mean estimation in the Gaussian sequence model with $\ell_p$ constraints for $p \in [0, \infty]$. We demonstrate two phenomena for the behavior of the maximum likelihood estimator (MLE), which depend on the noise level, the radius of the (quasi)norm constraint, the dimension, and the norm index $p$. First, if $p$ lies between $0$ and $1 + Θ(\tfrac{1}{\log d})$, inclusive,… ▽ More

    Submitted 1 July, 2025; v1 submitted 12 June, 2025; originally announced June 2025.

    Comments: 43 pages, 3 figures

  41. arXiv:2506.05295  [pdf, ps, other

    cs.LG cs.AI stat.ML

    Sample Complexity and Representation Ability of Test-time Scaling Paradigms

    Authors: Baihe Huang, Shanda Li, Tianhao Wu, Yiming Yang, Ameet Talwalkar, Kannan Ramchandran, Michael I. Jordan, Jiantao Jiao

    Abstract: Test-time scaling paradigms have significantly advanced the capabilities of large language models (LLMs) on complex tasks. Despite their empirical success, theoretical understanding of the sample efficiency of various test-time strategies -- such as self-consistency, best-of-$n$, and self-correction -- remains limited. In this work, we first establish a separation result between two repeated sampl… ▽ More

    Submitted 12 June, 2025; v1 submitted 5 June, 2025; originally announced June 2025.

  42. arXiv:2505.18223  [pdf, ps, other

    cs.CL cs.AI

    IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis

    Authors: Hanyu Li, Haoyu Liu, Tingyu Zhu, Tianyu Guo, Zeyu Zheng, Xiaotie Deng, Michael I. Jordan

    Abstract: Large Language Models (LLMs) show promise as data analysis agents, but existing benchmarks overlook the iterative nature of the field, where experts' decisions evolve with deeper insights of the dataset. To address this, we introduce IDA-Bench, a novel benchmark evaluating LLM agents in multi-round interactive scenarios. Derived from complex Kaggle notebooks, tasks are presented as sequential natu… ▽ More

    Submitted 6 June, 2025; v1 submitted 23 May, 2025; originally announced May 2025.

  43. arXiv:2505.13732  [pdf, ps, other

    stat.ML cs.LG

    Backward Conformal Prediction

    Authors: Etienne Gauthier, Francis Bach, Michael I. Jordan

    Abstract: We introduce $\textit{Backward Conformal Prediction}$, a method that guarantees conformal coverage while providing flexible control over the size of prediction sets. Unlike standard conformal prediction, which fixes the coverage level and allows the conformal set size to vary, our approach defines a rule that constrains how prediction set sizes behave based on the observed data, and adapts the cov… ▽ More

    Submitted 12 February, 2026; v1 submitted 19 May, 2025; originally announced May 2025.

    Comments: Code available at: https://github.com/GauthierE/backward-cp

  44. arXiv:2505.13564  [pdf, ps, other

    cs.LG stat.ML

    Online Decision-Focused Learning

    Authors: Aymeric Capitaine, Maxime Haddouche, Eric Moulines, Michael I. Jordan, Etienne Boursier, Alain Durmus

    Abstract: Decision-focused learning (DFL) is an increasingly popular paradigm for training predictive models whose outputs are used in decision-making tasks. Instead of merely optimizing for predictive accuracy, DFL trains models to directly minimize the loss associated with downstream decisions. However, existing studies focus solely on scenarios where a fixed batch of data is available and the objective f… ▽ More

    Submitted 7 March, 2026; v1 submitted 19 May, 2025; originally announced May 2025.

  45. arXiv:2505.05145  [pdf, ps, other

    cs.LG cs.AI cs.CL

    Understanding In-context Learning of Addition via Activation Subspaces

    Authors: Xinyan Hu, Kayo Yin, Michael I. Jordan, Jacob Steinhardt, Lijie Chen

    Abstract: To perform few-shot learning, language models extract signals from a few input-label pairs, aggregate these into a learned prediction rule, and apply this rule to new inputs. How is this implemented in the forward pass of modern transformer models? To explore this question, we study a structured family of few-shot learning tasks for which the true prediction rule is to add an integer $k$ to the in… ▽ More

    Submitted 9 October, 2025; v1 submitted 8 May, 2025; originally announced May 2025.

  46. arXiv:2504.03560  [pdf, ps, other

    math.OC cs.LG math.ST stat.ML

    Stochastic Optimization with Optimal Importance Sampling

    Authors: Liviu Aolaritei, Bart P. G. Van Parys, Henry Lam, Michael I. Jordan

    Abstract: Importance Sampling (IS) is a widely used variance reduction technique for enhancing the efficiency of Monte Carlo methods, particularly in rare-event simulation and related applications. Despite its effectiveness, the performance of IS is highly sensitive to the choice of the proposal distribution and often requires stochastic calibration. While the design and analysis of IS have been extensively… ▽ More

    Submitted 10 February, 2026; v1 submitted 4 April, 2025; originally announced April 2025.

  47. arXiv:2503.19068  [pdf, ps, other

    stat.ML cs.AI cs.LG stat.ME stat.OT

    Minimum Volume Conformal Sets for Multivariate Regression

    Authors: Sacha Braun, Liviu Aolaritei, Michael I. Jordan, Francis Bach

    Abstract: Conformal prediction provides a principled framework for constructing predictive sets with finite-sample validity. While much of the focus has been on univariate response variables, existing multivariate methods either impose rigid geometric assumptions or rely on flexible but computationally expensive approaches that do not explicitly optimize prediction set volume. We propose an optimization-dri… ▽ More

    Submitted 18 March, 2026; v1 submitted 24 March, 2025; originally announced March 2025.

  48. arXiv:2503.13050  [pdf, other

    stat.ML cs.LG

    E-Values Expand the Scope of Conformal Prediction

    Authors: Etienne Gauthier, Francis Bach, Michael I. Jordan

    Abstract: Conformal prediction is a powerful framework for distribution-free uncertainty quantification. The standard approach to conformal prediction relies on comparing the ranks of prediction scores: under exchangeability, the rank of a future test point cannot be too extreme relative to a calibration set. This rank-based method can be reformulated in terms of p-values. In this paper, we explore an alter… ▽ More

    Submitted 6 May, 2025; v1 submitted 17 March, 2025; originally announced March 2025.

    Comments: Code available at: https://github.com/GauthierE/evalues-expand-cp

  49. arXiv:2503.07879  [pdf, ps, other

    cs.CL cs.LG

    Datasets, Documents, and Repetitions: The Practicalities of Unequal Data Quality

    Authors: Alex Fang, Hadi Pouransari, Matt Jordan, Alexander Toshev, Vaishaal Shankar, Ludwig Schmidt, Tom Gunter

    Abstract: Data filtering has become a powerful tool for improving model performance while reducing computational cost. However, as large language model compute budgets continue to grow, the limited data volume provided by heavily filtered and deduplicated datasets will become a practical constraint. In efforts to better understand how to proceed, we study model performance at various compute budgets and acr… ▽ More

    Submitted 6 November, 2025; v1 submitted 10 March, 2025; originally announced March 2025.

  50. arXiv:2503.06582  [pdf, ps, other

    econ.TH cs.GT

    Marketplace Operators Can Induce Competitive Pricing

    Authors: Tiffany Ding, Dominique Perrault-Joncas, Orit Ronen, Michael I. Jordan, Dirk Bergemann, Dean Foster, Omer Gottesman

    Abstract: As e-commerce marketplaces continue to grow in popularity, it has become increasingly important to understand the role and impact of marketplace operators on competition and social welfare. We model a marketplace operator as an entity that not only facilitates third-party sales but can also choose to directly participate in the market as a competing seller. We formalize this market structure as a… ▽ More

    Submitted 22 October, 2025; v1 submitted 9 March, 2025; originally announced March 2025.