Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–9 of 9 results for author: Rafiee, B

Searching in archive cs. Search in all archives.
.
  1. arXiv:2605.24238  [pdf, ps, other

    cs.AI

    Toward Enactive Artificial Intelligence

    Authors: Banafsheh Rafiee, Richard Sutton

    Abstract: In this paper, we advocate for incorporating enactive approaches to perception and cognition into artificial intelligence (AI). Enactive approaches view perception as an active, skillful engagement with the world, where agents perceive by acting and by understanding how their actions shape their experience. This contrasts with classical views that treat perception as a passive internal process in… ▽ More

    Submitted 22 May, 2026; originally announced May 2026.

  2. arXiv:2605.07808  [pdf, ps, other

    cs.LG

    The Minimax Rate of Second-Order Calibration

    Authors: Kamil Ciosek, Banafsheh Rafiee, Sina Ghiassian, Nicolò Felicioni

    Abstract: We characterize the minimax rate of estimating the second-order calibration error for binary classification, which quantifies whether a higher-order predictor's epistemic-uncertainty estimate matches the conditional variance of the label probability on its level sets. Our key observation is that the sech perturbation kernel, previously used only to enforce smoothness of calibration functions, in f… ▽ More

    Submitted 8 May, 2026; originally announced May 2026.

  3. arXiv:2504.05185  [pdf, ps, other

    cs.CL

    Concise Reasoning via Reinforcement Learning

    Authors: Mehdi Fatemi, Banafsheh Rafiee, Mingjie Tang, Kartik Talamadupula

    Abstract: A major drawback of reasoning models is their excessive token usage, inflating computational cost, resource demand, and latency. We show this verbosity stems not from deeper reasoning but from reinforcement learning loss minimization when models produce incorrect answers. With unsolvable problems dominating training, this effect compounds into a systematic tendency toward longer outputs. Through t… ▽ More

    Submitted 21 November, 2025; v1 submitted 7 April, 2025; originally announced April 2025.

  4. arXiv:2210.14361  [pdf, other

    cs.LG cs.AI

    Auxiliary task discovery through generate-and-test

    Authors: Banafsheh Rafiee, Sina Ghiassian, Jun Jin, Richard Sutton, Jun Luo, Adam White

    Abstract: In this paper, we explore an approach to auxiliary task discovery in reinforcement learning based on ideas from representation learning. Auxiliary tasks tend to improve data efficiency by forcing the agent to learn auxiliary prediction and control objectives in addition to the main task of maximizing reward, and thus producing better representations. Typically these tasks are designed by people. M… ▽ More

    Submitted 20 July, 2024; v1 submitted 25 October, 2022; originally announced October 2022.

  5. arXiv:2204.00565  [pdf, other

    cs.AI cs.LG

    What makes useful auxiliary tasks in reinforcement learning: investigating the effect of the target policy

    Authors: Banafsheh Rafiee, Jun Jin, Jun Luo, Adam White

    Abstract: Auxiliary tasks have been argued to be useful for representation learning in reinforcement learning. Although many auxiliary tasks have been empirically shown to be effective for accelerating learning on the main task, it is not yet clear what makes useful auxiliary tasks. Some of the most promising results are on the pixel control, reward prediction, and the next state prediction auxiliary tasks;… ▽ More

    Submitted 1 April, 2022; originally announced April 2022.

  6. arXiv:2011.04590  [pdf, other

    cs.AI

    From Eye-blinks to State Construction: Diagnostic Benchmarks for Online Representation Learning

    Authors: Banafsheh Rafiee, Zaheer Abbas, Sina Ghiassian, Raksha Kumaraswamy, Richard Sutton, Elliot Ludvig, Adam White

    Abstract: We present three new diagnostic prediction problems inspired by classical-conditioning experiments to facilitate research in online prediction learning. Experiments in classical conditioning show that animals such as rabbits, pigeons, and dogs can make long temporal associations that enable multi-step prediction. To replicate this remarkable ability, an agent must construct an internal state repre… ▽ More

    Submitted 10 October, 2022; v1 submitted 9 November, 2020; originally announced November 2020.

  7. arXiv:2003.07417  [pdf, other

    cs.LG cs.AI cs.NE

    Improving Performance in Reinforcement Learning by Breaking Generalization in Neural Networks

    Authors: Sina Ghiassian, Banafsheh Rafiee, Yat Long Lo, Adam White

    Abstract: Reinforcement learning systems require good representations to work well. For decades practical success in reinforcement learning was limited to small domains. Deep reinforcement learning systems, on the other hand, are scalable, not dependent on domain specific prior knowledge and have been successfully used to play Atari, in 3D navigation from pixels, and to control high degree of freedom robots… ▽ More

    Submitted 16 March, 2020; originally announced March 2020.

    Comments: 10 pages; Accepted to AAMAS 2020

  8. arXiv:1805.07476  [pdf, other

    cs.LG cs.AI stat.ML

    Two geometric input transformation methods for fast online reinforcement learning with neural nets

    Authors: Sina Ghiassian, Huizhen Yu, Banafsheh Rafiee, Richard S. Sutton

    Abstract: We apply neural nets with ReLU gates in online reinforcement learning. Our goal is to train these networks in an incremental manner, without the computationally expensive experience replay. By studying how individual neural nodes behave in online training, we recognize that the global nature of ReLU gates can cause undesirable learning interference in each node's learning behavior. We propose redu… ▽ More

    Submitted 6 September, 2018; v1 submitted 18 May, 2018; originally announced May 2018.

    Comments: 16 pages

  9. arXiv:1705.04185  [pdf, other

    cs.AI cs.LG

    A First Empirical Study of Emphatic Temporal Difference Learning

    Authors: Sina Ghiassian, Banafsheh Rafiee, Richard S. Sutton

    Abstract: In this paper we present the first empirical study of the emphatic temporal-difference learning algorithm (ETD), comparing it with conventional temporal-difference learning, in particular, with linear TD(0), on on-policy and off-policy variations of the Mountain Car problem. The initial motivation for developing ETD was that it has good convergence properties under off-policy training (Sutton, Mah… ▽ More

    Submitted 12 May, 2017; v1 submitted 11 May, 2017; originally announced May 2017.

    Comments: 5 pages, Accepted to NIPS Continual Learning and Deep Networks workshop, 2016