Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–6 of 6 results for author: Ghasemlou, S

.
  1. arXiv:2607.19313  [pdf, ps, other

    cs.LG cs.AI

    Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

    Authors: Priyank Agrawal, Ankur Samanta, Shervin Ghasemlou, Jalaj Bhandari, Kavosh Asadi, Daniel Jiang, Aditya Modi

    Abstract: Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot generate any correct solutions, it receives \textit{zero} learning signal. Providing privileged guidance during training, such as solution prefixes, can help overcome this learning cliff by steering the model towards {correc… ▽ More

    Submitted 21 July, 2026; originally announced July 2026.

    Comments: 24 Pages

  2. arXiv:2602.11726  [pdf, ps, other

    cs.LG

    Dopamine: Brain Modes, Not Brains

    Authors: Shervin Ghasemlou

    Abstract: Parameter-efficient fine-tuning (PEFT) methods such as \lora{} adapt large pretrained models by adding small weight-space updates. While effective, weight deltas are hard to interpret mechanistically, and they do not directly expose \emph{which} internal computations are reused versus bypassed for a new task. We explore an alternative view inspired by neuromodulation: adaptation as a change in \em… ▽ More

    Submitted 12 February, 2026; originally announced February 2026.

  3. arXiv:2602.03713  [pdf, ps, other

    cs.IR

    Multimodal Generative Recommendation for Fusing Semantic and Collaborative Signals

    Authors: Moritz Vandenhirtz, Kaveh Hassani, Shervin Ghasemlou, Shuai Shao, Hamid Eghbalzadeh, Fuchun Peng, Jun Liu, Michael Louis Iuzzolino

    Abstract: Sequential recommender systems rank relevant items by modeling a user's interaction history and computing the inner product between the resulting user representation and stored item embeddings. To avoid the significant memory overhead of storing large item sets, the generative recommendation paradigm instead models each item as a series of discrete semantic codes. Here, the next item is predicted… ▽ More

    Submitted 3 February, 2026; originally announced February 2026.

  4. arXiv:2510.26160  [pdf, ps, other

    cs.CV

    CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark

    Authors: Jiaqi Wang, Xiao Yang, Kai Sun, Parth Suresh, Sanat Sharma, Adam Czyzewski, Derek Andersen, Surya Appini, Arkav Banerjee, Sajal Choudhary, Shervin Ghasemlou, Ziqiang Guan, Akil Iyer, Haidar Khan, Lingkun Kong, Roy Luo, Tiffany Ma, Zhen Qiao, David Tran, Wenfang Xu, Skyler Yeatman, Chen Zhou, Gunveer Gujral, Yinglong Xia, Shane Moon , et al. (16 additional authors not shown)

    Abstract: Wearable devices such as smart glasses are transforming the way people interact with their surroundings, enabling users to seek information regarding entities in their view. Multi-Modal Retrieval-Augmented Generation (MM-RAG) plays a key role in supporting such questions, yet there is still no comprehensive benchmark for this task, especially regarding wearables scenarios. To fill this gap, we pre… ▽ More

    Submitted 30 October, 2025; originally announced October 2025.

  5. arXiv:2409.06107  [pdf, other

    cs.CL cs.AI

    Doppelgänger's Watch: A Split Objective Approach to Large Language Models

    Authors: Shervin Ghasemlou, Ashish Katiyar, Aparajita Saraf, Seungwhan Moon, Mangesh Pujari, Pinar Donmez, Babak Damavandi, Anuj Kumar

    Abstract: In this paper, we investigate the problem of "generation supervision" in large language models, and present a novel bicameral architecture to separate supervision signals from their core capability, helpfulness. Doppelgänger, a new module parallel to the underlying language model, supervises the generation of each token, and learns to concurrently predict the supervision score(s) of the sequences… ▽ More

    Submitted 9 September, 2024; originally announced September 2024.

  6. arXiv:1807.08856  [pdf, ps, other

    cs.AI cs.RO

    Toward a language-theoretic foundation for planning and filtering

    Authors: Fatemeh Zahra Saberifar, Shervin Ghasemlou, Dylan A. Shell, Jason M. O'Kane

    Abstract: We address problems underlying the algorithmic question of automating the co-design of robot hardware in tandem with its apposite software. Specifically, we consider the impact that degradations of a robot's sensor and actuation suites may have on the ability of that robot to complete its tasks. We introduce a new formal structure that generalizes and consolidates a variety of well-known structure… ▽ More

    Submitted 23 July, 2018; originally announced July 2018.

    Comments: Accepted to appear in IJRR Special Issue on WAFR'16. Keywords: planning; combinatorial filter; design automation