Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 109 results for author: Rish, I

.
  1. arXiv:2608.18078  [pdf, ps, other

    cs.AI

    Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions

    Authors: Matthew Riemer, Tommaso Tosato, Amin Memarian, Maximilian Puelma Touzel, Glen Berseth, Irina Rish, Guillaume Dumas

    Abstract: This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and should be required to obtain behavioral certification before making decisions that affect economic markets. This is because integrating these agents into society could collapse the legal evidentiary distinction between competition and collusion among independent… ▽ More

    Submitted 28 May, 2026; originally announced August 2026.

    Comments: ICML 2026

  2. arXiv:2607.15449  [pdf, ps, other

    cs.LG

    Relevant and Irrelevant: A Renormalization Group Analysis of Transformer Attention

    Authors: Parviz Haggi-Mani, Irina Rish

    Abstract: Using the language of Wilsonian renormalization group theory (RG), we treat the Transformer's attention mechanism as a perturbation of the trained MLP residual-stack fixed point and ask whether it constitutes a relevant, marginal, or irrelevant operator. We derive a fixed-point shift formula and obtain four testable predictions for the fixed-point geometry, effective rank profile, layer specificit… ▽ More

    Submitted 16 July, 2026; originally announced July 2026.

  3. arXiv:2607.03580  [pdf, ps, other

    cs.LG cs.CV

    When Geometry Aligns: Dihedral Hidden-State Transformations in UNet, ViT, and DiT Architectures

    Authors: Mojtaba Faramarzi, Alex Lamb, Irina Rish

    Abstract: Diffusion architectures now encompass convolutional UNets as well as transformer-based designs such as Diffusion Transformers (DiTs), inspired by Vision Transformers (ViTs), yet the effects of structured geometric perturbations within these architectures remain poorly understood. We study this question through a unified framework that applies reflection-based elements of the dihedral group to inte… ▽ More

    Submitted 3 July, 2026; originally announced July 2026.

    Comments: Accepted at the Conference on Lifelong Learning Agents (CoLLAs), 2026

  4. arXiv:2606.12481  [pdf, ps, other

    cs.LG cs.AI

    Representing Time Series as Structured Programs for LLM Reasoning

    Authors: Jaeho Kim, Changhun Oh, Seokhyun Lee, Irina Rish, Changhee Lee

    Abstract: Large language models (LLMs) have demonstrated strong reasoning and instruction-following capabilities, making them potentially powerful tools for time-series analysis. However, time series lie outside their native textual modality, raising a fundamental question: how should time series be represented so that LLMs can reason about them effectively? Existing work typically serializes raw numerical… ▽ More

    Submitted 10 June, 2026; originally announced June 2026.

    Comments: Preprint

  5. arXiv:2606.10324  [pdf, ps, other

    cs.LG cond-mat.stat-mech stat.ML

    Rank Collapse, Fixed Points, and the Renormalization Group Structure of MLP Residual Networks

    Authors: Parviz Haggi-Mani, Irina Rish

    Abstract: The analogy between deep neural network forward passes and renormalization group (RG) flows has been repeatedly noted in the literature, but existing treatments remain qualitative: depth is described as a coarse-graining scale, attention is likened to a partition function, and representations are said to flow toward fixed points. No existing work has defined a measurable RG order parameter, tested… ▽ More

    Submitted 8 June, 2026; originally announced June 2026.

    Comments: 16 pages, 9 figures

  6. arXiv:2606.05145  [pdf, ps, other

    cs.LG cs.AI cs.CL

    Failed Reasoning Traces Tell You What Is Fixable (But Not by Reading Them)

    Authors: Nizar Islah, Istabrak Abbes, Irina Rish, Sarath Chandar, Eilif B. Muller

    Abstract: When post-trained language models fail on reasoning problems, the common test-time-scaling response is to spend more compute on additional attempts, and the failed traces play no further role. We argue this discards a crucial signal; some failures come from unlucky sampling, where more rollouts help, while others are structural and resist resampling regardless of budget. We propose that failed tra… ▽ More

    Submitted 3 June, 2026; originally announced June 2026.

  7. arXiv:2605.26248  [pdf, ps, other

    cs.LG cs.AI cs.NE

    Unified Neural Scaling Laws

    Authors: Ethan Caballero, Priyank Jaini, David Krueger, Irina Rish

    Abstract: We present a functional form (that we refer to as a Unified Neural Scaling Law (UNSL)) that accurately models and extrapolates the scaling behaviors of deep neural networks as multiple dimensions all vary simultaneously (i.e. how the evaluation metric of interest varies as one simultaneously varies the number of model parameters, training dataset size, number of training steps, number of inference… ▽ More

    Submitted 25 May, 2026; originally announced May 2026.

  8. arXiv:2605.20449  [pdf, ps, other

    cs.LG cs.AI

    LLM Pretraining Shapes a Generalizable Manifold: Insights into Cross-Modal Transfer to Time Series

    Authors: Alexis Roger, Prateek Humane, Zhenghan Tai, Gwen Legate, Andrei Mircea, Vasilii Feofanov, Irina Rish

    Abstract: Can language-pretrained transformers become effective time-series forecasters, and why? In this paper, we show that cross-modal transfer arises because language pretraining preconditions time series training with a reusable manifold. A linear probe on frozen LLM states decodes realistic time-series trajectories without paired supervision, and retrieval in this projected space yields competitive fo… ▽ More

    Submitted 19 May, 2026; originally announced May 2026.

  9. arXiv:2602.18252  [pdf, ps, other

    cs.CV cs.AI

    On the Adversarial Robustness of Discrete Image Tokenizers

    Authors: Rishika Bhagwatkar, Irina Rish, Nicolas Flammarion, Francesco Croce

    Abstract: Discrete image tokenizers encode visual inputs as sequences of tokens from a finite vocabulary and are gaining popularity in multimodal systems, including encoder-only, encoder-decoder, and decoder-only models. However, unlike CLIP encoders, their vulnerability to adversarial attacks has not been explored. Ours being the first work studying this topic, we first formulate attacks that aim to pertur… ▽ More

    Submitted 20 February, 2026; originally announced February 2026.

  10. arXiv:2602.03001  [pdf, ps, other

    cs.LG cs.AI

    Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent

    Authors: Hiroki Naganuma, Shagun Gupta, Youssef Briki, Ioannis Mitliagkas, Irina Rish, Parameswaran Raman, Hao-Jun Michael Shi

    Abstract: To maximize hardware utilization, modern machine learning systems typically employ large constant or manually tuned batch size schedules, relying on heuristics that are brittle and costly to tune. Existing adaptive strategies based on gradient noise scale (GNS) offer a principled alternative. However, their assumption of SGD's Euclidean geometry creates a fundamental mismatch with popular optimize… ▽ More

    Submitted 2 July, 2026; v1 submitted 2 February, 2026; originally announced February 2026.

    Comments: 8 pages, 2 figures, 4 tables

  11. arXiv:2512.11167  [pdf, ps, other

    cs.CV cs.AI

    Image Tiling for High-Resolution Reasoning: Balancing Local Detail with Global Context

    Authors: Anatole Jacquin de Margerie, Alexis Roger, Irina Rish

    Abstract: Reproducibility remains a cornerstone of scientific progress, yet complex multimodal models often lack transparent implementation details and accessible training infrastructure. In this work, we present a detailed reproduction and critical analysis of the Monkey Vision-Language Model (VLM) (Li et al. 2023b) published in CVPR24, a recent approach to high-resolution image understanding via image til… ▽ More

    Submitted 11 December, 2025; originally announced December 2025.

    Comments: Accepted in AAAI 2025 Workshop on Reproducible AI

  12. arXiv:2511.11622  [pdf, ps, other

    cs.LG cs.AI cs.CL

    Small Vocabularies, Big Gains: Pretraining and Tokenization in Time Series Models

    Authors: Alexis Roger, Gwen Legate, Kashif Rasul, Yuriy Nevmyvaka, Irina Rish

    Abstract: Tokenization and transfer learning are two critical components in building state of the art time series foundation models for forecasting. In this work, we systematically study the effect of tokenizer design, specifically scaling and quantization strategies, on model performance, alongside the impact of pretraining versus random initialization. We show that tokenizer configuration primarily govern… ▽ More

    Submitted 6 November, 2025; originally announced November 2025.

  13. arXiv:2510.26006  [pdf, ps, other

    cs.CV cs.CL

    CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments

    Authors: Rishika Bhagwatkar, Syrielle Montariol, Angelika Romanou, Beatriz Borges, Irina Rish, Antoine Bosselut

    Abstract: Humans can naturally identify, reason about, and explain anomalies in their environment. In computer vision, this long-standing challenge remains limited to industrial defects or unrealistic, synthetically generated anomalies, failing to capture the richness and unpredictability of real-world anomalies. In this work, we introduce CAVE, the first benchmark of real-world visual anomalies. CAVE suppo… ▽ More

    Submitted 29 October, 2025; originally announced October 2025.

    Journal ref: 2025 Conference on Empirical Methods in Natural Language Processing

  14. arXiv:2510.06108  [pdf, ps, other

    cs.LG cs.CL

    Influence Functions for Efficient Data Selection in Reasoning

    Authors: Prateek Humane, Paolo Cudrano, Daniel Z. Kaplan, Matteo Matteucci, Supriyo Chakraborty, Irina Rish

    Abstract: Fine-tuning large language models (LLMs) on chain-of-thought (CoT) data shows that a small amount of high-quality data can outperform massive datasets. Yet, what constitutes "quality" remains ill-defined. Existing reasoning methods rely on indirect heuristics such as problem difficulty or trace length, while instruction-tuning has explored a broader range of automated selection strategies, but rar… ▽ More

    Submitted 1 December, 2025; v1 submitted 7 October, 2025; originally announced October 2025.

    Comments: 4 pages, 2 figures; added link to codebase

  15. arXiv:2510.05244  [pdf, ps, other

    cs.CR

    Indirect Prompt Injections: Are Firewalls All You Need, or Stronger Benchmarks?

    Authors: Rishika Bhagwatkar, Kevin Kasa, Abhay Puri, Gabriel Huang, Irina Rish, Graham W. Taylor, Krishnamurthy Dj Dvijotham, Alexandre Lacoste

    Abstract: AI agents are vulnerable to indirect prompt injection attacks, where malicious instructions embedded in external content or tool outputs cause unintended or harmful behavior. Inspired by the well-established concept of firewalls, we show that a simple, modular, and model-agnostic defense operating at the agent--tool interface achieves perfect security with high utility across all four public bench… ▽ More

    Submitted 23 March, 2026; v1 submitted 6 October, 2025; originally announced October 2025.

  16. arXiv:2509.03503  [pdf, ps, other

    cs.LG cs.AI

    Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients

    Authors: Gwen Legate, Irina Rish, Eugene Belilovsky

    Abstract: Federated learning enables collaborative model training across numerous edge devices without requiring participants to share data; however, memory and communication constraints on these edge devices may preclude their participation in training. We consider a setting in which a subset of edge devices are below a critical memory or communication threshold required to conduct model updates. Under typ… ▽ More

    Submitted 3 September, 2025; originally announced September 2025.

  17. arXiv:2508.14079  [pdf, ps, other

    cs.LG

    A Guide to Robust Generalization: The Impact of Architecture, Pre-training, and Optimization Strategy

    Authors: Maxime Heuillet, Rishika Bhagwatkar, Jonas Ngnawé, Yann Pequignot, Alexandre Larouche, Christian Gagné, Irina Rish, Ola Ahmad, Audrey Durand

    Abstract: Deep learning models operating in the image domain are vulnerable to small input perturbations. For years, robustness to such perturbations was pursued by training models from scratch (i.e., with random initializations) using specialized loss objectives. Recently, robust fine-tuning has emerged as a more efficient alternative: instead of training from scratch, pretrained models are adapted to maxi… ▽ More

    Submitted 12 August, 2025; originally announced August 2025.

  18. arXiv:2508.09904  [pdf, ps, other

    cs.LG cs.AI

    Beyond Naïve Prompting: Strategies for Improved Context-aided Forecasting with LLMs

    Authors: Arjun Ashok, Andrew Robert Williams, Vincent Zhihao Zheng, Irina Rish, Nicolas Chapados, Étienne Marcotte, Valentina Zantedeschi, Alexandre Drouin

    Abstract: Real-world forecasting requires models to integrate not only historical data but also relevant contextual information provided in textual form. While large language models (LLMs) show promise for context-aided forecasting, critical challenges remain: we lack diagnostic tools to understand failure modes, performance remains far below their potential, and high computational costs limit practical dep… ▽ More

    Submitted 11 July, 2026; v1 submitted 13 August, 2025; originally announced August 2025.

    Comments: Published at TMLR - OpenReview link: https://openreview.net/forum?id=dkjHHFJkVI

  19. arXiv:2508.04826  [pdf, ps, other

    cs.CL cs.AI

    Persistent Instability in LLM's Personality Measurements: Effects of Scale, Reasoning, and Conversation History

    Authors: Tommaso Tosato, Saskia Helbling, Yorguin-Jose Mantilla-Ramos, Mahmood Hegazy, Alberto Tosato, David John Lemay, Irina Rish, Guillaume Dumas

    Abstract: Large language models require consistent behavioral patterns for safe deployment, yet there are indications of large variability that may lead to an instable expression of personality traits in these models. We present PERSIST (PERsonality Stability in Synthetic Text), a comprehensive evaluation framework testing 25 open-source models (1B-685B parameters) across 2 million+ responses. Using traditi… ▽ More

    Submitted 23 December, 2025; v1 submitted 6 August, 2025; originally announced August 2025.

    Comments: Accepted at AAAI 2026, Track on AI Alignment

  20. arXiv:2508.01908  [pdf, ps, other

    cs.LG cs.AI cs.CL

    Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models

    Authors: Istabrak Abbes, Gopeshh Subbaraj, Matthew Riemer, Nizar Islah, Benjamin Therien, Tsuguchika Tabaru, Hiroaki Kingetsu, Sarath Chandar, Irina Rish

    Abstract: Training large language models (LLMs) typically involves pre-training on massive corpora, only to restart the process entirely when new data becomes available. A more efficient and resource-conserving approach would be continual pre-training, where models are updated with new data rather than retraining from scratch. However, the introduction of new data often causes distribution shifts, leading t… ▽ More

    Submitted 3 August, 2025; originally announced August 2025.

  21. arXiv:2507.12367  [pdf, ps, other

    cs.SE cs.AI cs.PL

    GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities

    Authors: Diganta Misra, Nizar Islah, Victor May, Brice Rauby, Zihan Wang, Justine Gehring, Antonio Orvieto, Muawiz Chaudhary, Eilif B. Muller, Irina Rish, Samira Ebrahimi Kahou, Massimo Caccia

    Abstract: The rapid evolution of software libraries poses a considerable hurdle for code generation, necessitating continuous adaptation to frequent version updates while preserving backward compatibility. While existing code evolution benchmarks provide valuable insights, they typically lack execution-based evaluation for generating code compliant with specific library versions. To address this, we introdu… ▽ More

    Submitted 21 July, 2025; v1 submitted 16 July, 2025; originally announced July 2025.

    Comments: Version 2 of the dataset from: arXiv:2411.05830

  22. arXiv:2506.23025  [pdf, ps, other

    cs.LG cs.AI

    Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models

    Authors: Tejas Vaidhya, Ayush Kaushal, Vineet Jain, Francis Couture Harpin, Prashant Shishodia, Majid Behbahani, Yuriy Nevmyvaka, Irina Rish

    Abstract: Large language models (LLMs) are increasingly used across research and industry applications, yet their inference efficiency remains a significant challenge. As the computational power of modern GPU architectures continuously improves, their memory bandwidth and capacity have not scaled proportionally, creating a critical bottleneck during inference. To address this, we investigate ternary languag… ▽ More

    Submitted 28 June, 2025; originally announced June 2025.

  23. arXiv:2506.21570  [pdf, ps, other

    cs.CL cs.AI cs.LG

    Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting

    Authors: Roland Riachi, Kashif Rasul, Arjun Ashok, Prateek Humane, Alexis Roger, Andrew R. Williams, Yuriy Nevmyvaka, Irina Rish

    Abstract: Recent works have demonstrated the effectiveness of adapting pre-trained language models (LMs) for forecasting time series in the low-data regime. We build upon these findings by analyzing the effective transfer from language models to time series forecasting under various design choices including upstream post-training, time series tokenizer and language backbone size. In the low-data regime, the… ▽ More

    Submitted 12 June, 2025; originally announced June 2025.

  24. arXiv:2506.05447  [pdf, ps, other

    cs.LG cs.AI

    Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning

    Authors: Andrei Mircea, Supriyo Chakraborty, Nima Chitsazan, Milind Naphade, Sambit Sahu, Irina Rish, Ekaterina Lobacheva

    Abstract: This work aims to understand how scaling improves language models, specifically in terms of training dynamics. We find that language models undergo loss deceleration early in training; an abrupt slowdown in the rate of loss improvement, resulting in piecewise linear behaviour of the loss curve in log-log space. Scaling up the model mitigates this transition by (1) decreasing the loss at which dece… ▽ More

    Submitted 14 July, 2025; v1 submitted 5 June, 2025; originally announced June 2025.

    Comments: Published as a conference paper at ACL 2025

    ACM Class: I.2.7

  25. arXiv:2505.23725  [pdf, ps, other

    cs.LG

    MuLoCo: Muon is a practical inner optimizer for DiLoCo

    Authors: Benjamin Thérien, Xiaolong Huang, Aaron Defazio, Irina Rish, Eugene Belilovsky

    Abstract: DiLoCo is a powerful framework for training large language models (LLMs), enabling larger optimal batch sizes and increased accelerator utilization under networking constraints. However, DiLoCo's performance has been shown to degrade as the number of workers (K) increases (Charles et al., 2025). In this work, we posit that a related but often overlooked factor in DiLoCo's behavior is the choice of… ▽ More

    Submitted 1 June, 2026; v1 submitted 29 May, 2025; originally announced May 2025.

  26. arXiv:2503.23478  [pdf, other

    cs.LG cs.AI cs.RO

    Handling Delay in Real-Time Reinforcement Learning

    Authors: Ivan Anokhin, Rishav Rishav, Matthew Riemer, Stephen Chung, Irina Rish, Samira Ebrahimi Kahou

    Abstract: Real-time reinforcement learning (RL) introduces several challenges. First, policies are constrained to a fixed number of actions per second due to hardware limitations. Second, the environment may change while the network is still computing an action, leading to observational delay. The first issue can partly be addressed with pipelining, leading to higher throughput and potentially better polici… ▽ More

    Submitted 30 March, 2025; originally announced March 2025.

    Comments: Accepted at ICLR 2025. Code available at https://github.com/avecplezir/realtime-agent

  27. arXiv:2503.05029  [pdf, ps, other

    cs.LG cs.AI cs.CL

    Continual Pre-training of MoEs: How robust is your router?

    Authors: Benjamin Thérien, Charles-Étienne Joseph, Zain Sarwar, Ashwinee Panda, Anirban Das, Shi-Xiong Zhang, Stephen Rawls, Sambit Sahu, Eugene Belilovsky, Irina Rish

    Abstract: Sparsely-activated Mixture of Experts (MoE) transformers are promising architectures for foundation models. Compared to dense transformers that require the same amount of floating-point operations (FLOPs) per forward pass, MoEs benefit from improved sample efficiency at training time and achieve much stronger performance. Many closed-source and open-source frontier language models have thus adopte… ▽ More

    Submitted 10 November, 2025; v1 submitted 6 March, 2025; originally announced March 2025.

  28. arXiv:2503.02844  [pdf, ps, other

    cs.LG

    Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training

    Authors: Vaibhav Singh, Paul Janson, Paria Mehrbod, Adam Ibrahim, Irina Rish, Eugene Belilovsky, Benjamin Thérien

    Abstract: The ever-growing availability of unlabeled data presents both opportunities and challenges for training artificial intelligence systems. While self-supervised learning (SSL) has emerged as a powerful paradigm for extracting meaningful representations from vast amounts of unlabeled data, existing methods still struggle to adapt to the non-stationary, non-IID nature of real-world data streams withou… ▽ More

    Submitted 9 September, 2025; v1 submitted 4 March, 2025; originally announced March 2025.

    Comments: Accepted (oral) at Fourth Conference on Lifelong Learning Agents - CoLLAs 2025

  29. arXiv:2501.11566  [pdf, other

    q-bio.NC

    Artificial Neural Networks for Magnetoencephalography: A review of an emerging field

    Authors: Arthur Dehgan, Hamza Abdelhedi, Vanessa Hadid, Irina Rish, Karim Jerbi

    Abstract: Magnetoencephalography (MEG) is a cutting-edge neuroimaging technique that measures the intricate brain dynamics underlying cognitive processes with an unparalleled combination of high temporal and spatial precision. MEG data analytics has always relied on advanced signal processing and mathematical and statistical tools for various tasks ranging from data cleaning to probing the signals' rich dyn… ▽ More

    Submitted 18 May, 2025; v1 submitted 20 January, 2025; originally announced January 2025.

    Comments: Full table containing information of papers reviewed can be found in this google sheet: https://tinyurl.com/ub3s5mr

  30. arXiv:2501.09672  [pdf, ps, other

    cs.CV cs.AI

    CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models

    Authors: Alexis Roger, Prateek Humane, Daniel Z. Kaplan, Kshitij Gupta, Qi Sun, George Adamopoulos, Jonathan Siu Chi Lim, Quentin Anthony, Edwin Fennell, Irina Rish

    Abstract: The proliferation of Vision-Language Models (VLMs) in the past several years calls for rigorous and comprehensive evaluation methods and benchmarks. This work analyzes existing VLM evaluation techniques, including automated metrics, AI-based assessments, and human evaluations across diverse tasks. We first introduce Robin - a novel suite of VLMs that we built by combining Large Language Models (LL… ▽ More

    Submitted 5 August, 2025; v1 submitted 16 January, 2025; originally announced January 2025.

  31. arXiv:2412.14355  [pdf, other

    cs.LG cs.AI

    Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference

    Authors: Matthew Riemer, Gopeshh Subbaraj, Glen Berseth, Irina Rish

    Abstract: Realtime environments change even as agents perform action inference and learning, thus requiring high interaction frequencies to effectively minimize regret. However, recent advances in machine learning involve larger neural networks with longer inference times, raising questions about their applicability in realtime systems where reaction time is crucial. We present an analysis of lower bounds o… ▽ More

    Submitted 18 December, 2024; originally announced December 2024.

  32. arXiv:2411.12372  [pdf, other

    cs.CL cs.LG

    RedPajama: an Open Dataset for Training Large Language Models

    Authors: Maurice Weber, Daniel Fu, Quentin Anthony, Yonatan Oren, Shane Adams, Anton Alexandrov, Xiaozhong Lyu, Huu Nguyen, Xiaozhe Yao, Virginia Adams, Ben Athiwaratkun, Rahul Chalamala, Kezhen Chen, Max Ryabinin, Tri Dao, Percy Liang, Christopher Ré, Irina Rish, Ce Zhang

    Abstract: Large language models are increasingly becoming a cornerstone technology in artificial intelligence, the sciences, and society as a whole, yet the optimal strategies for dataset composition and filtering remain largely elusive. Many of the top-performing models lack transparency in their dataset curation and model development processes, posing an obstacle to the development of fully open language… ▽ More

    Submitted 19 November, 2024; originally announced November 2024.

    Comments: 38th Conference on Neural Information Processing Systems (NeurIPS 2024) Track on Datasets and Benchmarks

  33. arXiv:2411.07007  [pdf, other

    cs.LG cs.AI

    Non-Adversarial Inverse Reinforcement Learning via Successor Feature Matching

    Authors: Arnav Kumar Jain, Harley Wiltzer, Jesse Farebrother, Irina Rish, Glen Berseth, Sanjiban Choudhury

    Abstract: In inverse reinforcement learning (IRL), an agent seeks to replicate expert demonstrations through interactions with the environment. Traditionally, IRL is treated as an adversarial game, where an adversary searches over reward models, and a learner optimizes the reward through repeated RL procedures. This game-solving approach is both computationally expensive and difficult to stabilize. In this… ▽ More

    Submitted 22 April, 2025; v1 submitted 11 November, 2024; originally announced November 2024.

    Comments: Accepted to ICLR 2025

  34. arXiv:2411.05830  [pdf, other

    cs.SE cs.LG

    GitChameleon: Unmasking the Version-Switching Capabilities of Code Generation Models

    Authors: Nizar Islah, Justine Gehring, Diganta Misra, Eilif Muller, Irina Rish, Terry Yue Zhuo, Massimo Caccia

    Abstract: The rapid evolution of software libraries presents a significant challenge for code generation models, which must adapt to frequent version updates while maintaining compatibility with previous versions. Existing code completion benchmarks often overlook this dynamic aspect, and the one that does consider it relies on static code prediction tasks without execution-based evaluation, offering a limi… ▽ More

    Submitted 5 November, 2024; originally announced November 2024.

  35. arXiv:2411.02344  [pdf, other

    cs.LG cs.CL

    Seq-VCR: Preventing Collapse in Intermediate Transformer Representations for Enhanced Reasoning

    Authors: Md Rifat Arefin, Gopeshh Subbaraj, Nicolas Gontier, Yann LeCun, Irina Rish, Ravid Shwartz-Ziv, Christopher Pal

    Abstract: Decoder-only Transformers often struggle with complex reasoning tasks, particularly arithmetic reasoning requiring multiple sequential operations. In this work, we identify representation collapse in the model's intermediate layers as a key factor limiting their reasoning capabilities. To address this, we propose Sequential Variance-Covariance Regularization (Seq-VCR), which enhances the entropy o… ▽ More

    Submitted 20 March, 2025; v1 submitted 4 November, 2024; originally announced November 2024.

  36. arXiv:2410.18959  [pdf, ps, other

    cs.LG cs.AI stat.ML

    Context is Key: A Benchmark for Forecasting with Essential Textual Information

    Authors: Andrew Robert Williams, Arjun Ashok, Étienne Marcotte, Valentina Zantedeschi, Jithendaraa Subramanian, Roland Riachi, James Requeima, Alexandre Lacoste, Irina Rish, Nicolas Chapados, Alexandre Drouin

    Abstract: Forecasting is a critical task in decision-making across numerous domains. While historical numerical data provide a start, they fail to convey the complete context for reliable and accurate predictions. Human forecasters frequently rely on additional information, such as background knowledge and constraints, which can efficiently be communicated through natural language. However, in spite of rece… ▽ More

    Submitted 5 June, 2025; v1 submitted 24 October, 2024; originally announced October 2024.

    Comments: ICML 2025. First two authors contributed equally

  37. arXiv:2409.05817  [pdf, other

    cs.CV cs.HC

    VFA: Vision Frequency Analysis of Foundation Models and Human

    Authors: Mohammad-Javad Darvishi-Bayazi, Md Rifat Arefin, Jocelyn Faubert, Irina Rish

    Abstract: Machine learning models often struggle with distribution shifts in real-world scenarios, whereas humans exhibit robust adaptation. Models that better align with human perception may achieve higher out-of-distribution generalization. In this study, we investigate how various characteristics of large-scale computer vision models influence their alignment with human capabilities and robustness. Our f… ▽ More

    Submitted 9 September, 2024; originally announced September 2024.

  38. arXiv:2407.12327  [pdf, other

    cs.LG cs.AI cs.CL

    Spectra: Surprising Effectiveness of Pretraining Ternary Language Models at Scale

    Authors: Ayush Kaushal, Tejas Vaidhya, Arnab Kumar Mondal, Tejas Pandey, Aaryan Bhagat, Irina Rish

    Abstract: Rapid advancements in GPU computational power has outpaced memory capacity and bandwidth growth, creating bottlenecks in Large Language Model (LLM) inference. Post-training quantization is the leading method for addressing memory-related bottlenecks in LLM inference, but it suffers from significant performance degradation below 4-bit precision. This paper addresses these challenges by investigatin… ▽ More

    Submitted 11 October, 2024; v1 submitted 17 July, 2024; originally announced July 2024.

    Comments: 42 pages, 21 figures, and 13 tables

    MSC Class: 68T30 ACM Class: I.2.6; I.2.7

  39. arXiv:2407.12161  [pdf, other

    cs.AI

    Interpretability in Action: Exploratory Analysis of VPT, a Minecraft Agent

    Authors: Karolis Jucys, George Adamopoulos, Mehrab Hamidi, Stephanie Milani, Mohammad Reza Samsami, Artem Zholus, Sonia Joseph, Blake Richards, Irina Rish, Özgür Şimşek

    Abstract: Understanding the mechanisms behind decisions taken by large foundation models in sequential decision making tasks is critical to ensuring that such systems operate transparently and safely. In this work, we perform exploratory analysis on the Video PreTraining (VPT) Minecraft playing agent, one of the largest open-source vision-based agents. We aim to illuminate its reasoning mechanisms by applyi… ▽ More

    Submitted 16 July, 2024; originally announced July 2024.

    Comments: Mechanistic Interpretability Workshop at ICML 2024

  40. arXiv:2407.11121  [pdf, other

    cs.CV cs.AI cs.LG

    Towards Adversarially Robust Vision-Language Models: Insights from Design Choices and Prompt Formatting Techniques

    Authors: Rishika Bhagwatkar, Shravan Nayak, Reza Bayat, Alexis Roger, Daniel Z Kaplan, Pouya Bashivan, Irina Rish

    Abstract: Vision-Language Models (VLMs) have witnessed a surge in both research and real-world applications. However, as they are becoming increasingly prevalent, ensuring their robustness against adversarial attacks is paramount. This work systematically investigates the impact of model design choices on the adversarial robustness of VLMs against image-based attacks. Additionally, we introduce novel, cost-… ▽ More

    Submitted 15 July, 2024; originally announced July 2024.

  41. arXiv:2407.04680  [pdf, ps, other

    q-bio.NC cs.AI cs.CL

    Lost in Translation: The Algorithmic Gap Between LMs and the Brain

    Authors: Tommaso Tosato, Pascal Jr Tikeng Notsawo, Saskia Helbling, Irina Rish, Guillaume Dumas

    Abstract: Language Models (LMs) have achieved impressive performance on various linguistic tasks, but their relationship to human language processing in the brain remains unclear. This paper examines the gaps and overlaps between LMs and the brain at different levels of analysis, emphasizing the importance of looking beyond input-output behavior to examine and compare the internal processes of these systems… ▽ More

    Submitted 5 July, 2024; originally announced July 2024.

  42. arXiv:2406.00153  [pdf, ps, other

    cs.LG

    $μ$LO: Compute-Efficient Meta-Generalization of Learned Optimizers

    Authors: Benjamin Thérien, Charles-Étienne Joseph, Boris Knyazev, Edouard Oyallon, Irina Rish, Eugene Belilovsky

    Abstract: Learned optimizers (LOs) have the potential to significantly reduce the wall-clock training time of neural networks. However, they can struggle to optimize unseen tasks (meta-generalize), especially when training networks wider than those seen during meta-training. To address this, we derive the Maximal Update Parametrization ($μ$P) for two state-of-the-art learned optimizer architectures and prop… ▽ More

    Submitted 18 March, 2026; v1 submitted 31 May, 2024; originally announced June 2024.

  43. arXiv:2404.07377  [pdf, other

    cs.LG cs.AI cs.CL cs.CV cs.IT

    Deep Generative Sampling in the Dual Divergence Space: A Data-efficient & Interpretative Approach for Generative AI

    Authors: Sahil Garg, Anderson Schneider, Anant Raj, Kashif Rasul, Yuriy Nevmyvaka, Sneihil Gopal, Amit Dhurandhar, Guillermo Cecchi, Irina Rish

    Abstract: Building on the remarkable achievements in generative sampling of natural images, we propose an innovative challenge, potentially overly ambitious, which involves generating samples of entire multivariate time series that resemble images. However, the statistical challenge lies in the small sample size, sometimes consisting of a few hundred subjects. This issue is especially problematic for deep g… ▽ More

    Submitted 10 April, 2024; originally announced April 2024.

  44. arXiv:2403.08763  [pdf, other

    cs.LG cs.AI cs.CL

    Simple and Scalable Strategies to Continually Pre-train Large Language Models

    Authors: Adam Ibrahim, Benjamin Thérien, Kshitij Gupta, Mats L. Richter, Quentin Anthony, Timothée Lesort, Eugene Belilovsky, Irina Rish

    Abstract: Large language models (LLMs) are routinely pre-trained on billions of tokens, only to start the process over again once new data becomes available. A much more efficient solution is to continually pre-train these models, saving significant compute compared to re-training. However, the distribution shift induced by new data typically results in degraded performance on previous data or poor adaptati… ▽ More

    Submitted 4 September, 2024; v1 submitted 13 March, 2024; originally announced March 2024.

  45. arXiv:2402.13368  [pdf, other

    cs.LG cs.CV

    Unsupervised Concept Discovery Mitigates Spurious Correlations

    Authors: Md Rifat Arefin, Yan Zhang, Aristide Baratin, Francesco Locatello, Irina Rish, Dianbo Liu, Kenji Kawaguchi

    Abstract: Models prone to spurious correlations in training data often produce brittle predictions and introduce unintended biases. Addressing this challenge typically involves methods relying on prior knowledge and group annotation to remove spurious correlations, which may not be readily available in many applications. In this paper, we establish a novel connection between unsupervised object-centric lear… ▽ More

    Submitted 16 July, 2024; v1 submitted 20 February, 2024; originally announced February 2024.

    Journal ref: ICML 2024

  46. arXiv:2312.12868  [pdf, other

    cs.AI q-bio.NC

    Towards Machines that Trust: AI Agents Learn to Trust in the Trust Game

    Authors: Ardavan S. Nobandegani, Irina Rish, Thomas R. Shultz

    Abstract: Widely considered a cornerstone of human morality, trust shapes many aspects of human social interactions. In this work, we present a theoretical analysis of the $\textit{trust game}$, the canonical task for studying trust in behavioral and brain sciences, along with simulation results supporting our analysis. Specifically, leveraging reinforcement learning (RL) to train our AI agents, we systemat… ▽ More

    Submitted 20 December, 2023; originally announced December 2023.

  47. arXiv:2310.08278  [pdf, other

    cs.LG cs.AI

    Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting

    Authors: Kashif Rasul, Arjun Ashok, Andrew Robert Williams, Hena Ghonia, Rishika Bhagwatkar, Arian Khorasani, Mohammad Javad Darvishi Bayazi, George Adamopoulos, Roland Riachi, Nadhir Hassen, Marin Biloš, Sahil Garg, Anderson Schneider, Nicolas Chapados, Alexandre Drouin, Valentina Zantedeschi, Yuriy Nevmyvaka, Irina Rish

    Abstract: Over the past years, foundation models have caused a paradigm shift in machine learning due to their unprecedented capabilities for zero-shot and few-shot generalization. However, despite the success of foundation models in modalities such as natural language processing and computer vision, the development of foundation models for time series forecasting has lagged behind. We present Lag-Llama, a… ▽ More

    Submitted 8 February, 2024; v1 submitted 12 October, 2023; originally announced October 2023.

    Comments: First two authors contributed equally. All data, models and code used are open-source. GitHub: https://github.com/time-series-foundation-models/lag-llama

  48. arXiv:2309.14021  [pdf, other

    cs.CL cs.AI

    LORD: Low Rank Decomposition Of Monolingual Code LLMs For One-Shot Compression

    Authors: Ayush Kaushal, Tejas Vaidhya, Irina Rish

    Abstract: Low Rank Decomposition of matrix - splitting a large matrix into a product of two smaller matrix offers a means for compression that reduces the parameters of a model without sparsification, and hence delivering more speedup on modern hardware. Moreover, unlike quantization, the compressed linear layers remain fully differentiable and all the parameters trainable, while being able to leverage the… ▽ More

    Submitted 25 September, 2023; originally announced September 2023.

    Comments: 9 pages

  49. Amplifying Pathological Detection in EEG Signaling Pathways through Cross-Dataset Transfer Learning

    Authors: Mohammad-Javad Darvishi-Bayazi, Mohammad Sajjad Ghaemi, Timothee Lesort, Md Rifat Arefin, Jocelyn Faubert, Irina Rish

    Abstract: Pathology diagnosis based on EEG signals and decoding brain activity holds immense importance in understanding neurological disorders. With the advancement of artificial intelligence methods and machine learning techniques, the potential for accurate data-driven diagnoses and effective treatments has grown significantly. However, applying machine learning algorithms to real-world datasets presents… ▽ More

    Submitted 19 September, 2023; originally announced September 2023.

  50. arXiv:2308.04014  [pdf, other

    cs.CL cs.LG

    Continual Pre-Training of Large Language Models: How to (re)warm your model?

    Authors: Kshitij Gupta, Benjamin Thérien, Adam Ibrahim, Mats L. Richter, Quentin Anthony, Eugene Belilovsky, Irina Rish, Timothée Lesort

    Abstract: Large language models (LLMs) are routinely pre-trained on billions of tokens, only to restart the process over again once new data becomes available. A much cheaper and more efficient solution would be to enable the continual pre-training of these models, i.e. updating pre-trained models with new data instead of re-training them from scratch. However, the distribution shift induced by novel data t… ▽ More

    Submitted 6 September, 2023; v1 submitted 7 August, 2023; originally announced August 2023.