Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 84 results for author: Hosseini, A

Searching in archive cs. Search in all archives.
.
  1. arXiv:2607.28640  [pdf, ps, other

    cs.CL cs.CV

    TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

    Authors: Andong Hua, Colton Bishop, Igor Mordatch, Arian Hosseini, Jindong Gu, Aleksandra Faust, Rebecca Roelofs, Yao Qin

    Abstract: Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observe a systematic discrepancy in model predictions under such cross-modal variations. Specifically, we define the modality gap as the difference in model performance under semantically equivalent textual and multimodal inputs. We introduce TokenSwap,… ▽ More

    Submitted 19 May, 2026; originally announced July 2026.

  2. arXiv:2607.02770  [pdf, ps, other

    cs.CL cs.AI

    Gemma 4 Technical Report

    Authors: Gemma Team, Sherif El Abd, Vaibhav Aggarwal, Robin Algayres, Alek Andreev, Olivier Bachem, Ian Ballantyne, Cormac Brick, Victor Cărbune, Michelle Casbon, Mayank Chaturvedi, Aditya Chawla, Victor Cotruta, Alice Coucke, Phil Culliton, Robert Dadashi, Lucas Dixon, Mohamed Elhawaty, Utku Evci, Clément Farabet, Johan Ferret, Filippo Galgani, Sertan Girgin, Jean-Bastien Grill, Maarten Grootendorst , et al. (298 additional authors not shown)

    Abstract: We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside improved vision and audio encoders for all model sizes, we propose a unified, encoder-free architecture… ▽ More

    Submitted 24 July, 2026; v1 submitted 2 July, 2026; originally announced July 2026.

    Comments: 17 pages, 2 figures, technical report, updated

  3. arXiv:2606.27515  [pdf, ps, other

    cs.LG physics.geo-ph

    Boundary condition fidelity for bottom-hole pressure and CO2 plume prediction in geological carbon storage

    Authors: Romal Ramadhan, Seyyed A. Hosseini, Larry W. Lake

    Abstract: Accurate prediction of bottom-hole pressure (BHP) and CO2 plume migration is essential for safe geological carbon storage, yet practical simulations often rely on truncated domains where artificial boundaries distort pressure diffusion and CO2 saturation footprints. In this study, we evaluate how boundary-condition fidelity affects BHP and CO2 plume prediction by comparing ten reduced-domain bound… ▽ More

    Submitted 25 June, 2026; originally announced June 2026.

  4. arXiv:2604.23985  [pdf, ps, other

    cs.AI cs.CL cs.LG

    Representational Curvature Modulates Behavioral Uncertainty in Large Language Models

    Authors: Jack King, Evelina Fedorenko, Eghbal A. Hosseini

    Abstract: In autoregressive large language models (LLMs), temporal straightening offers an account of how the next-token prediction objective shapes representations. Models learn to progressively straighten the representational trajectory of input sequences across layers, potentially facilitating next-token prediction via linear extrapolation. However, a direct link between this trajectory and token-level b… ▽ More

    Submitted 26 April, 2026; originally announced April 2026.

  5. arXiv:2604.22546  [pdf, ps, other

    cs.CV

    ReLIC-SGG: Relation Lattice Completion for Open-Vocabulary Scene Graph Generation

    Authors: Amir Hosseini, Sara Farahani, Xinyi Li, Suiyang Guang

    Abstract: Open-vocabulary scene graph generation (SGG) aims to describe visual scenes with flexible relation phrases beyond a fixed predicate set. Existing methods usually treat annotated triplets as positives and all unannotated object-pair relations as negatives. However, scene graph annotations are inherently incomplete: many valid relations are missing, and the same interaction can be described at diffe… ▽ More

    Submitted 25 May, 2026; v1 submitted 24 April, 2026; originally announced April 2026.

    Comments: Some errors in the experimental sections

  6. arXiv:2604.21836  [pdf, ps, other

    q-bio.NC cs.AI

    Modulating Cross-Modal Convergence with Single-Stimulus, Intra-Modal Dispersion

    Authors: Eghbal A. Hosseini, Brian Cheung, Evelina Fedorenko, Alex H. Williams

    Abstract: Neural networks exhibit a remarkable degree of representational convergence across diverse architectures, training objectives, and even data modalities. This convergence is predictive of alignment with brain representation. A recent hypothesis suggests this arises from learning the underlying structure in the environment in similar ways. However, it is unclear how individual stimuli elicit converg… ▽ More

    Submitted 23 April, 2026; originally announced April 2026.

    Journal ref: ICLR 2026 Workshop on Representational Alignment (Re-Align)

  7. arXiv:2601.22364  [pdf, ps, other

    cs.CL cs.AI

    Context Structure Reshapes the Representational Geometry of Language Models

    Authors: Eghbal A. Hosseini, Yuxuan Li, Yasaman Bahri, Declan Campbell, Andrew Kyle Lampinen

    Abstract: Large Language Models (LLMs) have been shown to organize the representations of input sequences into straighter neural trajectories in their deep layers, which has been hypothesized to facilitate next-token prediction via linear extrapolation. Language models can also adapt to diverse tasks and learn new structure in context, and recent work has shown that this in-context learning (ICL) can be ref… ▽ More

    Submitted 29 January, 2026; originally announced January 2026.

  8. arXiv:2601.18751  [pdf, ps, other

    cs.LG cs.AI

    Trust, Don't Trust, or Flip: Robust Preference-Based Reinforcement Learning with Multi-Expert Feedback

    Authors: Seyed Amir Hosseini, Maryam Abdolali, Amirhosein Tavakkoli, Fardin Ayar, Ehsan Javanmardi, Manabu Tsukada, Mahdi Javanmardi

    Abstract: Preference-based reinforcement learning (PBRL) offers a promising alternative to explicit reward engineering by learning from pairwise trajectory comparisons. However, real-world preference data often comes from heterogeneous annotators with varying reliability; some accurate, some noisy, and some systematically adversarial. Existing PBRL methods either treat all feedback equally or attempt to fil… ▽ More

    Submitted 26 January, 2026; originally announced January 2026.

    Comments: Equal contribution: Seyed Amir Hosseini and Maryam Abdolali. Corresponding author: Maryam Abdolali (maryam.abdolali@kntu.ac.ir)

  9. arXiv:2512.22255  [pdf, ps, other

    cs.AI cs.LG

    Shape of Thought: When Distribution Matters More than Correctness in Reasoning Tasks

    Authors: Abhranil Chandra, Ayush Agrawal, Arian Hosseini, Sebastian Fischmeister, Rishabh Agarwal, Navin Goyal, Aaron Courville

    Abstract: We present the surprising finding that a language model's reasoning capabilities can be improved by training on synthetic datasets of chain-of-thought (CoT) traces from more capable models, even when all of those traces lead to an incorrect final answer. Our experiments show this approach can yield better performance on reasoning tasks than training on human-annotated datasets. We hypothesize that… ▽ More

    Submitted 22 January, 2026; v1 submitted 24 December, 2025; originally announced December 2025.

  10. arXiv:2510.22243  [pdf, ps, other

    cs.CV cs.AI

    Real-Time Semantic Segmentation on FPGA for Autonomous Vehicles Using LMIINet with the CGRA4ML Framework

    Authors: Amir Mohammad Khadem Hosseini, Sattar Mirzakuchaki

    Abstract: Semantic segmentation has emerged as a fundamental problem in computer vision, gaining particular importance in real-time applications such as autonomous driving. The main challenge is achieving high accuracy while operating under computational and hardware constraints. In this research, we present an FPGA-based implementation of real-time semantic segmentation leveraging the lightweight LMIINet a… ▽ More

    Submitted 25 October, 2025; originally announced October 2025.

  11. arXiv:2510.19035  [pdf, ps, other

    cs.SE eess.SY

    Extending Resource Constrained Project Scheduling to Mega-Projects with Model-Based Systems Engineering & Hetero-functional Graph Theory

    Authors: Amirreza Hosseini, Amro M. Farid

    Abstract: Within the project management context, project scheduling serves as an indispensable component, functioning as a fundamental tool for planning, monitoring, controlling, and managing projects more broadly. Although the resource-constrained project scheduling problem (RCPSP) lies at the core of project management activities, it remains largely disconnected from the broader literature on model-based… ▽ More

    Submitted 21 October, 2025; originally announced October 2025.

  12. arXiv:2508.12712  [pdf

    cs.LG cs.CV

    Argos: A Decentralized Federated System for Detection of Traffic Signs in CAVs

    Authors: Seyed Mahdi Haji Seyed Hossein, Alireza Hosseini, Soheil Hajian Manesh, Amirali Shahriary

    Abstract: Connected and automated vehicles generate vast amounts of sensor data daily, raising significant privacy and communication challenges for centralized machine learning approaches in perception tasks. This study presents a decentralized, federated learning framework tailored for traffic sign detection in vehicular networks to enable collaborative model training without sharing raw data. The framewor… ▽ More

    Submitted 18 August, 2025; originally announced August 2025.

    Comments: 7 pages, 10 figures

    ACM Class: I.2.6; I.4.8

  13. Optimal CO2 storage management considering safety constraints in multi-stakeholder multi-site CCS projects: a Markov game perspective

    Authors: Jungang Chen, Seyyed A. Hosseini

    Abstract: Carbon capture and storage (CCS) projects typically involve a diverse array of stakeholders or players from public, private, and regulatory sectors, each with different objectives and responsibilities. Given the complexity, scale, and long-term nature of CCS operations, determining whether individual stakeholders can independently maximize their interests or whether collaborative coalition agreeme… ▽ More

    Submitted 22 January, 2026; v1 submitted 15 August, 2025; originally announced August 2025.

    Comments: 58 pages

    Journal ref: Int. J. Greenh. Gas Control 149 (2026) 104683

  14. arXiv:2508.10142  [pdf, ps, other

    cs.CL

    Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs

    Authors: Kartikeya Badola, Jonathan Simon, Arian Hosseini, Sara Marie Mc Carthy, Tsendsuren Munkhdalai, Abhimanyu Goyal, Tomáš Kočiský, Shyam Upadhyay, Bahare Fatemi, Mehran Kazemi

    Abstract: Large language models (LLMs) excel at solving problems with clear and complete statements, but often struggle with nuanced environments or interactive tasks which are common in most real-world scenarios. This highlights the critical need for developing LLMs that can effectively engage in logically consistent multi-turn dialogue, seek information and reason with incomplete data. To this end, we int… ▽ More

    Submitted 24 August, 2025; v1 submitted 13 August, 2025; originally announced August 2025.

  15. arXiv:2507.06261  [pdf, ps, other

    cs.CL cs.AI

    Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

    Authors: Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, Luke Marris, Sam Petulla, Colin Gaffney, Asaf Aharoni, Nathan Lintz, Tiago Cardal Pais, Henrik Jacobsson, Idan Szpektor, Nan-Jiang Jiang, Krishna Haridasan, Ahmed Omran, Nikunj Saunshi, Dara Bahri, Gaurav Mishra, Eric Chu , et al. (3410 additional authors not shown)

    Abstract: In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our most capable model yet, achieving SoTA performance on frontier coding and reasoning benchmarks. In addition to its incredible coding and reasoning skills, Gemini 2.5 Pro is a thinking model that excels at multimodal unde… ▽ More

    Submitted 19 December, 2025; v1 submitted 7 July, 2025; originally announced July 2025.

    Comments: 72 pages, 17 figures

  16. arXiv:2505.10585  [pdf

    cs.CV cs.CR

    Efficient Malicious UAV Detection Using Autoencoder-TSMamba Integration

    Authors: Azim Akhtarshenas, Ramin Toosi, David López-Pérez, Tohid Alizadeh, Alireza Hosseini

    Abstract: Malicious Unmanned Aerial Vehicles (UAVs) present a significant threat to next-generation networks (NGNs), posing risks such as unauthorized surveillance, data theft, and the delivery of hazardous materials. This paper proposes an integrated (AE)-classifier system to detect malicious UAVs. The proposed AE, based on a 4-layer Tri-orientated Spatial Mamba (TSMamba) architecture, effectively captures… ▽ More

    Submitted 14 May, 2025; originally announced May 2025.

    Comments: 12 pages, 6 figures and 3 tables, accepted in IbPRIA 2025, https://www.ibpria.org/2025/?page=dates

  17. arXiv:2505.04842  [pdf, ps, other

    cs.LG cs.AI

    Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers

    Authors: Kusha Sareen, Morgane M Moss, Alessandro Sordoni, Rishabh Agarwal, Arian Hosseini

    Abstract: Prevalent reinforcement learning~(RL) methods for fine-tuning LLM reasoners, such as GRPO or Leave-one-out PPO, abandon the learned value function in favor of empirically estimated returns. This hinders test-time compute scaling that relies on using the value-function for verification. Yet if parallel test-time compute is already part of the deployment plan, training should be designed to support… ▽ More

    Submitted 12 April, 2026; v1 submitted 7 May, 2025; originally announced May 2025.

  18. arXiv:2504.10070  [pdf, other

    cs.CV

    DTFSal: Audio-Visual Dynamic Token Fusion for Video Saliency Prediction

    Authors: Kiana Hooshanfar, Alireza Hosseini, Ahmad Kalhor, Babak Nadjar Araabi

    Abstract: Audio-visual saliency prediction aims to mimic human visual attention by identifying salient regions in videos through the integration of both visual and auditory information. Although visual-only approaches have significantly advanced, effectively incorporating auditory cues remains challenging due to complex spatio-temporal interactions and high computational demands. To address these challenges… ▽ More

    Submitted 16 April, 2025; v1 submitted 14 April, 2025; originally announced April 2025.

  19. arXiv:2504.03171  [pdf

    cs.CV cs.AI cs.RO

    Real-Time Roadway Obstacle Detection for Electric Scooters Using Deep Learning and Multi-Sensor Fusion

    Authors: Zeyang Zheng, Arman Hosseini, Dong Chen, Omid Shoghli, Arsalan Heydarian

    Abstract: The increasing adoption of electric scooters (e-scooters) in urban areas has coincided with a rise in traffic accidents and injuries, largely due to their small wheels, lack of suspension, and sensitivity to uneven surfaces. While deep learning-based object detection has been widely used to improve automobile safety, its application for e-scooter obstacle detection remains unexplored. This study i… ▽ More

    Submitted 4 April, 2025; originally announced April 2025.

    Comments: Accepted at ASCE International Conference on Computing in Civil Engineering (i3ce)

  20. arXiv:2504.01005  [pdf, ps, other

    cs.CL cs.AI cs.LG

    When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning

    Authors: Nishad Singhi, Hritik Bansal, Arian Hosseini, Aditya Grover, Kai-Wei Chang, Marcus Rohrbach, Anna Rohrbach

    Abstract: Scaling test-time compute has emerged as a key strategy for enhancing the reasoning capabilities of large language models (LLMs), particularly in tasks like mathematical problem-solving. A traditional approach, Self-Consistency (SC), generates multiple solutions to a problem and selects the most common answer via majority voting. Another common method involves scoring each solution with a reward m… ▽ More

    Submitted 19 October, 2025; v1 submitted 1 April, 2025; originally announced April 2025.

    Comments: COLM 2025

  21. arXiv:2502.05117  [pdf, other

    cs.HC

    Adoption of AI-Assisted E-Scooters: The Role of Perceived Trust, Safety, and Demographic Drivers

    Authors: Amit Kumar, Arman Hosseini, Arghavan Azarbayjani, Arsalan Heydarian, Omidreza Shoghli

    Abstract: E-scooters have become a more dominant mode of transport in recent years. However, the rise in their usage has been accompanied by an increase in injuries, affecting the trust and perceived safety of both users and non-users. Artificial intelligence (AI), as a cutting-edge and widely applied technology, has demonstrated potential to enhance transportation safety, particularly in driver assistance… ▽ More

    Submitted 7 February, 2025; originally announced February 2025.

  22. arXiv:2411.00728  [pdf, other

    cs.MA cs.AI cs.LG cs.RO

    Multi-Agent Deep Q-Network with Layer-based Communication Channel for Autonomous Internal Logistics Vehicle Scheduling in Smart Manufacturing

    Authors: Mohammad Feizabadi, Arman Hosseini, Zakaria Yahouni

    Abstract: In smart manufacturing, scheduling autonomous internal logistic vehicles is crucial for optimizing operational efficiency. This paper proposes a multi-agent deep Q-network (MADQN) with a layer-based communication channel (LBCC) to address this challenge. The main goals are to minimize total job tardiness, reduce the number of tardy jobs, and lower vehicle energy consumption. The method is evaluate… ▽ More

    Submitted 1 November, 2024; originally announced November 2024.

    Comments: Accepted for the 5th IFAC/INSTICC INTERNATIONAL CONFERENCE ON INNOVATIVE INTELLIGENT INDUSTRIAL PRODUCTION AND LOGISTICS

  23. arXiv:2410.18252  [pdf, other

    cs.LG cs.AI cs.CL

    Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models

    Authors: Michael Noukhovitch, Shengyi Huang, Sophie Xhonneux, Arian Hosseini, Rishabh Agarwal, Aaron Courville

    Abstract: The dominant paradigm for RLHF is online and on-policy RL: synchronously generating from the large language model (LLM) policy, labelling with a reward model, and learning using feedback on the LLM's own outputs. While performant, this paradigm is computationally inefficient. Inspired by classical deep RL literature, we propose separating generation and learning in RLHF. This enables asynchronous… ▽ More

    Submitted 26 April, 2025; v1 submitted 23 October, 2024; originally announced October 2024.

    Comments: accepted at ICLR 2025, code at https://github.com/mnoukhov/async_rlhf, integrated into the open-instruct library https://github.com/allenai/open-instruct

  24. arXiv:2410.01748  [pdf, other

    cs.LG

    Not All LLM Reasoners Are Created Equal

    Authors: Arian Hosseini, Alessandro Sordoni, Daniel Toyama, Aaron Courville, Rishabh Agarwal

    Abstract: We study the depth of grade-school math (GSM) problem-solving capabilities of LLMs. To this end, we evaluate their performance on pairs of existing math word problems together so that the answer to the second problem depends on correctly answering the first problem. Our findings reveal a significant reasoning gap in most LLMs, that is performance difference between solving the compositional pairs… ▽ More

    Submitted 2 October, 2024; originally announced October 2024.

  25. arXiv:2409.06676  [pdf, other

    eess.IV cs.CV eess.SP

    Constructing an Interpretable Deep Denoiser by Unrolling Graph Laplacian Regularizer

    Authors: Seyed Alireza Hosseini, Tam Thuc Do, Gene Cheung, Yuichi Tanaka

    Abstract: An image denoiser can be used for a wide range of restoration problems via the Plug-and-Play (PnP) architecture. In this paper, we propose a general framework to build an interpretable graph-based deep denoiser (GDD) by unrolling a solution to a maximum a posteriori (MAP) problem equipped with a graph Laplacian regularizer (GLR) as signal prior. Leveraging a recent theorem showing that any (pseudo… ▽ More

    Submitted 10 September, 2024; originally announced September 2024.

  26. arXiv:2408.16737  [pdf, other

    cs.CL cs.AI

    Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling

    Authors: Hritik Bansal, Arian Hosseini, Rishabh Agarwal, Vinh Q. Tran, Mehran Kazemi

    Abstract: Training on high-quality synthetic data from strong language models (LMs) is a common strategy to improve the reasoning performance of LMs. In this work, we revisit whether this strategy is compute-optimal under a fixed inference budget (e.g., FLOPs). To do so, we investigate the trade-offs between generating synthetic data using a stronger but more expensive (SE) model versus a weaker but cheaper… ▽ More

    Submitted 7 October, 2024; v1 submitted 29 August, 2024; originally announced August 2024.

  27. arXiv:2408.15240  [pdf, other

    cs.LG

    Generative Verifiers: Reward Modeling as Next-Token Prediction

    Authors: Lunjun Zhang, Arian Hosseini, Hritik Bansal, Mehran Kazemi, Aviral Kumar, Rishabh Agarwal

    Abstract: Verifiers or reward models are often used to enhance the reasoning performance of large language models (LLMs). A common approach is the Best-of-N method, where N candidate solutions generated by the LLM are ranked by a verifier, and the best one is selected. While LLM-based verifiers are typically trained as discriminative classifiers to score solutions, they do not utilize the text generation ca… ▽ More

    Submitted 22 February, 2025; v1 submitted 27 August, 2024; originally announced August 2024.

    Comments: ICLR 2025

  28. arXiv:2408.00802  [pdf, other

    cs.IR cs.AI cs.CL cs.LG

    Leveraging LLM Reasoning Enhances Personalized Recommender Systems

    Authors: Alicia Y. Tsai, Adam Kraft, Long Jin, Chenwei Cai, Anahita Hosseini, Taibai Xu, Zemin Zhang, Lichan Hong, Ed H. Chi, Xinyang Yi

    Abstract: Recent advancements have showcased the potential of Large Language Models (LLMs) in executing reasoning tasks, particularly facilitated by Chain-of-Thought (CoT) prompting. While tasks like arithmetic reasoning involve clear, definitive answers and logical chains of thought, the application of LLM reasoning in recommendation systems (RecSys) presents a distinct challenge. RecSys tasks revolve arou… ▽ More

    Submitted 22 July, 2024; originally announced August 2024.

    Comments: To be published at ACL 2024

  29. arXiv:2407.10310  [pdf

    cs.CY eess.SY

    Impact of Road Infrastructure and Traffic Scenarios on E-scooterists' Riding and Gaze Behavior

    Authors: Dong Chen, Arman Hosseini, Arik Smith, Zeyang Zheng, David Xiang, Arsalan Heydarian, Omid Shoghli, Bradford Campbell

    Abstract: The growing adoption of e-scooters has raised significant safety concerns, particularly due to a surge in injuries and fatalities. This study explores the relationship between road infrastructure, traffic scenarios, and e-scooterists' riding and gaze behaviors to improve road safety and user experience. A naturalistic study was conducted using instrumented e-scooters, capturing gaze patterns, fixa… ▽ More

    Submitted 16 March, 2025; v1 submitted 5 May, 2024; originally announced July 2024.

    Comments: 12 pages, 10 figures

    Journal ref: International Conference on Transportation & Development (ICTD 2025)

  30. arXiv:2406.17815  [pdf, other

    cs.CV cs.AI

    SUM: Saliency Unification through Mamba for Visual Attention Modeling

    Authors: Alireza Hosseini, Amirhossein Kazerouni, Saeed Akhavan, Michael Brudno, Babak Taati

    Abstract: Visual attention modeling, important for interpreting and prioritizing visual stimuli, plays a significant role in applications such as marketing, multimedia, and robotics. Traditional saliency prediction models, especially those based on Convolutional Neural Networks (CNNs) or Transformers, achieve notable success by leveraging large-scale annotated datasets. However, the current state-of-the-art… ▽ More

    Submitted 9 September, 2024; v1 submitted 25 June, 2024; originally announced June 2024.

    Comments: Accepted at IEEE/CVF WACV 2025

  31. arXiv:2406.04090  [pdf, other

    cs.LG cs.CV eess.IV eess.SP

    Interpretable Lightweight Transformer via Unrolling of Learned Graph Smoothness Priors

    Authors: Tam Thuc Do, Parham Eftekhar, Seyed Alireza Hosseini, Gene Cheung, Philip Chou

    Abstract: We build interpretable and lightweight transformer-like neural networks by unrolling iterative optimization algorithms that minimize graph smoothness priors -- the quadratic graph Laplacian regularizer (GLR) and the $\ell_1$-norm graph total variation (GTV) -- subject to an interpolation constraint. The crucial insight is that a normalized signal-dependent graph learning module amounts to a varian… ▽ More

    Submitted 5 November, 2024; v1 submitted 6 June, 2024; originally announced June 2024.

  32. arXiv:2405.03039  [pdf

    cs.CV eess.SY

    Performance Evaluation of Real-Time Object Detection for Electric Scooters

    Authors: Dong Chen, Arman Hosseini, Arik Smith, Amir Farzin Nikkhah, Arsalan Heydarian, Omid Shoghli, Bradford Campbell

    Abstract: Electric scooters (e-scooters) have rapidly emerged as a popular mode of transportation in urban areas, yet they pose significant safety challenges. In the United States, the rise of e-scooters has been marked by a concerning increase in related injuries and fatalities. Recently, while deep-learning object detection holds paramount significance in autonomous vehicles to avoid potential collisions,… ▽ More

    Submitted 5 May, 2024; originally announced May 2024.

    Comments: 10 pages, 3 figures

  33. arXiv:2403.17031  [pdf, other

    cs.LG

    The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization

    Authors: Shengyi Huang, Michael Noukhovitch, Arian Hosseini, Kashif Rasul, Weixun Wang, Lewis Tunstall

    Abstract: This work is the first to openly reproduce the Reinforcement Learning from Human Feedback (RLHF) scaling behaviors reported in OpenAI's seminal TL;DR summarization work. We create an RLHF pipeline from scratch, enumerate over 20 key implementation details, and share key insights during the reproduction. Our RLHF-trained Pythia models demonstrate significant gains in response quality that scale wit… ▽ More

    Submitted 23 March, 2024; originally announced March 2024.

  34. arXiv:2403.02336  [pdf, other

    cs.CV cs.AI

    Brand Visibility in Packaging: A Deep Learning Approach for Logo Detection, Saliency-Map Prediction, and Logo Placement Analysis

    Authors: Alireza Hosseini, Kiana Hooshanfar, Pouria Omrani, Reza Toosi, Ramin Toosi, Zahra Ebrahimian, Mohammad Ali Akhaee

    Abstract: In the highly competitive area of product marketing, the visibility of brand logos on packaging plays a crucial role in shaping consumer perception, directly influencing the success of the product. This paper introduces a comprehensive framework to measure the brand logo's attention on a packaging design. The proposed method consists of three steps. The first step leverages YOLOv8 for precise logo… ▽ More

    Submitted 4 March, 2024; originally announced March 2024.

  35. arXiv:2402.06457  [pdf, other

    cs.LG cs.AI cs.CL

    V-STaR: Training Verifiers for Self-Taught Reasoners

    Authors: Arian Hosseini, Xingdi Yuan, Nikolay Malkin, Aaron Courville, Alessandro Sordoni, Rishabh Agarwal

    Abstract: Common self-improvement approaches for large language models (LLMs), such as STaR, iteratively fine-tune LLMs on self-generated solutions to improve their problem-solving ability. However, these approaches discard the large amounts of incorrect solutions generated during this process, potentially neglecting valuable information in such solutions. To address this shortcoming, we propose V-STaR that… ▽ More

    Submitted 13 August, 2024; v1 submitted 9 February, 2024; originally announced February 2024.

  36. arXiv:2311.04930  [pdf, other

    cs.CL cs.AI

    Large language models implicitly learn to straighten neural sentence trajectories to construct a predictive representation of natural language

    Authors: Eghbal A. Hosseini, Evelina Fedorenko

    Abstract: Predicting upcoming events is critical to our ability to interact with our environment. Transformer models, trained on next-word prediction, appear to construct representations of linguistic input that can support diverse downstream tasks. But how does a predictive objective shape such representations? Inspired by recent work in vision (Henaff et al., 2019), we test a hypothesis about predictive r… ▽ More

    Submitted 5 November, 2023; originally announced November 2023.

    Comments: 37th Conference on Neural Information Processing Systems (NeurIPS 2023). 20 pages, 5 main figures, 7 supplementary figures

  37. arXiv:2310.18846  [pdf, other

    cs.CV

    INCODE: Implicit Neural Conditioning with Prior Knowledge Embeddings

    Authors: Amirhossein Kazerouni, Reza Azad, Alireza Hosseini, Dorit Merhof, Ulas Bagci

    Abstract: Implicit Neural Representations (INRs) have revolutionized signal representation by leveraging neural networks to provide continuous and smooth representations of complex data. However, existing INRs face limitations in capturing fine-grained details, handling noise, and adapting to diverse signal types. To address these challenges, we introduce INCODE, a novel approach that enhances the control o… ▽ More

    Submitted 28 October, 2023; originally announced October 2023.

    Comments: Accepted at WACV 2024 conference

  38. arXiv:2310.04855  [pdf, other

    cs.LG

    Epsilon non-Greedy: A Bandit Approach for Unbiased Recommendation via Uniform Data

    Authors: S. M. F. Sani, Seyed Abbas Hosseini, Hamid R. Rabiee

    Abstract: Often, recommendation systems employ continuous training, leading to a self-feedback loop bias in which the system becomes biased toward its previous recommendations. Recent studies have attempted to mitigate this bias by collecting small amounts of unbiased data. While these studies have successfully developed less biased models, they ignore the crucial fact that the recommendations generated by… ▽ More

    Submitted 7 October, 2023; originally announced October 2023.

  39. arXiv:2307.00433  [pdf

    cs.NI cs.DC

    Intelligent Traffic Control with Smart Speed Bumps

    Authors: Melvin Mokhtari, Amirreza Hosseini, Alireza Habibi, Adel Karshenas, Ali Amoomahdi

    Abstract: Traffic congestion and safety continue to pose significant challenges in urban environments. In this paper, we introduce the Smart Speed Bump (SSBump), a novel traffic calming solution that leverages the Internet of Things (IoT) and innovative non-Newtonian fluid materials to enhance road safety, optimize emergency response times, and improve the overall driving experience. The SSBump uses IoT sen… ▽ More

    Submitted 29 September, 2023; v1 submitted 1 July, 2023; originally announced July 2023.

    Comments: 7 pages, 5 figures

  40. arXiv:2306.12509  [pdf, other

    cs.CL cs.LG

    Joint Prompt Optimization of Stacked LLMs using Variational Inference

    Authors: Alessandro Sordoni, Xingdi Yuan, Marc-Alexandre Côté, Matheus Pereira, Adam Trischler, Ziang Xiao, Arian Hosseini, Friederike Niedtner, Nicolas Le Roux

    Abstract: Large language models (LLMs) can be seen as atomic units of computation mapping sequences to a distribution over sequences. Thus, they can be seen as stochastic language layers in a language network, where the learnable parameters are the natural language prompts at each layer. By stacking two such layers and feeding the output of one layer to the next, we obtain a Deep Language Network (DLN). We… ▽ More

    Submitted 4 December, 2023; v1 submitted 21 June, 2023; originally announced June 2023.

    Comments: NeurIPS 2023

  41. Emotional Framing in the Spreading of False and True Claims

    Authors: Akram Sadat Hosseini, Steffen Staab

    Abstract: The explosive growth of online misinformation, such as false claims, has affected the social behavior of online users. In order to be persuasive and mislead the audience, false claims are made to trigger emotions in their audience. This paper contributes to understanding how misinformation in social media is shaped by investigating the emotional framing that authors of the claims try to create for… ▽ More

    Submitted 29 March, 2023; originally announced March 2023.

  42. arXiv:2211.08473  [pdf, other

    cs.CL cs.LG

    On the Compositional Generalization Gap of In-Context Learning

    Authors: Arian Hosseini, Ankit Vani, Dzmitry Bahdanau, Alessandro Sordoni, Aaron Courville

    Abstract: Pretrained large generative language models have shown great performance on many tasks, but exhibit low compositional generalization abilities. Scaling such models has been shown to improve their performance on various NLP tasks even just by conditioning them on a few examples to solve the task without any fine-tuning (also known as in-context learning). In this work, we look at the gap between th… ▽ More

    Submitted 15 November, 2022; originally announced November 2022.

  43. arXiv:2107.00247  [pdf, other

    cs.LG cs.AI

    The Interplay between Distribution Parameters and the Accuracy-Robustness Tradeoff in Classification

    Authors: Alireza Mousavi Hosseini, Amir Mohammad Abouei, Mohammad Hossein Rohban

    Abstract: Adversarial training tends to result in models that are less accurate on natural (unperturbed) examples compared to standard models. This can be attributed to either an algorithmic shortcoming or a fundamental property of the training data distribution, which admits different solutions for optimal standard and adversarial classifiers. In this work, we focus on the latter case under a binary Gaussi… ▽ More

    Submitted 1 July, 2021; originally announced July 2021.

    Comments: Accepted for presentation in AML ICML workshop 2021

  44. arXiv:2105.03519  [pdf, other

    cs.CL

    Understanding by Understanding Not: Modeling Negation in Language Models

    Authors: Arian Hosseini, Siva Reddy, Dzmitry Bahdanau, R Devon Hjelm, Alessandro Sordoni, Aaron Courville

    Abstract: Negation is a core construction in natural language. Despite being very successful on many tasks, state-of-the-art pre-trained language models often handle negation incorrectly. To improve language models in this regard, we propose to augment the language modeling objective with an unlikelihood objective that is based on negated generic sentences from a raw text corpus. By training BERT with the r… ▽ More

    Submitted 7 May, 2021; originally announced May 2021.

  45. arXiv:2102.07737  [pdf, other

    eess.IV cs.CV cs.LG eess.SP physics.med-ph

    Zero-Shot Self-Supervised Learning for MRI Reconstruction

    Authors: Burhaneddin Yaman, Seyed Amir Hossein Hosseini, Mehmet Akçakaya

    Abstract: Deep learning (DL) has emerged as a powerful tool for accelerated MRI reconstruction, but often necessitates a database of fully-sampled measurements for training. Recent self-supervised and unsupervised learning approaches enable training without fully-sampled data. However, a database of undersampled measurements may not be available in many scenarios, especially for scans involving contrast or… ▽ More

    Submitted 28 November, 2023; v1 submitted 15 February, 2021; originally announced February 2021.

    Journal ref: International Conference on Learning Representations (ICLR), 2022

  46. arXiv:2010.13868  [pdf, other

    eess.IV cs.CV cs.LG eess.SP physics.med-ph

    Improved Supervised Training of Physics-Guided Deep Learning Image Reconstruction with Multi-Masking

    Authors: Burhaneddin Yaman, Seyed Amir Hossein Hosseini, Steen Moeller, Mehmet Akçakaya

    Abstract: Physics-guided deep learning (PG-DL) via algorithm unrolling has received significant interest for improved image reconstruction, including MRI applications. These methods unroll an iterative optimization algorithm into a series of regularizer and data consistency units. The unrolled networks are typically trained end-to-end using a supervised approach. Current supervised PG-DL approaches use all… ▽ More

    Submitted 26 October, 2020; originally announced October 2020.

    Journal ref: Proceedings of IEEE ICASSP, 2021

  47. arXiv:2009.11693  [pdf, other

    cs.LG physics.geo-ph stat.ML

    A Variational Auto-Encoder for Reservoir Monitoring

    Authors: Kristian Gundersen, Seyyed A. Hosseini, Anna Oleynik, Guttorm Alendal

    Abstract: Carbon dioxide Capture and Storage (CCS) is an important strategy in mitigating anthropogenic CO$_2$ emissions. In order for CCS to be successful, large quantities of CO$_2$ must be stored and the storage site conformance must be monitored. Here we present a deep learning method to reconstruct pressure fields and classify the flux out of the storage formation based on the pressure data from Above… ▽ More

    Submitted 2 October, 2020; v1 submitted 23 September, 2020; originally announced September 2020.

  48. Design and Implementation of a Maxi-Sized Mobile Robot (Karo) for Rescue Missions

    Authors: Soheil Habibian, Mehdi Dadvar, Behzad Peykari, Alireza Hosseini, M. Hossein Salehzadeh, Alireza H. M. Hosseini, Farshid Najafi

    Abstract: Rescue robots are expected to carry out reconnaissance and dexterity operations in unknown environments comprising unstructured obstacles. Although a wide variety of designs and implementations have been presented within the field of rescue robotics, embedding all mobility, dexterity, and reconnaissance capabilities in a single robot remains a challenging problem. This paper explains the design an… ▽ More

    Submitted 9 January, 2021; v1 submitted 23 July, 2020; originally announced August 2020.

    Journal ref: Robomech J 8, 1 (2021)

  49. arXiv:2008.06029  [pdf

    eess.IV cs.CV cs.LG eess.SP physics.med-ph

    Multi-Mask Self-Supervised Learning for Physics-Guided Neural Networks in Highly Accelerated MRI

    Authors: Burhaneddin Yaman, Hongyi Gu, Seyed Amir Hossein Hosseini, Omer Burak Demirel, Steen Moeller, Jutta Ellermann, Kâmil Uğurbil, Mehmet Akçakaya

    Abstract: Self-supervised learning has shown great promise due to its capability to train deep learning MRI reconstruction methods without fully-sampled data. Current self-supervised learning methods for physics-guided reconstruction networks split acquired undersampled data into two disjoint sets, where one is used for data consistency (DC) in the unrolled network and the other to define the training loss.… ▽ More

    Submitted 8 June, 2022; v1 submitted 13 August, 2020; originally announced August 2020.

    Journal ref: NMR in Biomedicine, 2022

  50. arXiv:2006.09450  [pdf, other

    eess.IV cs.CV cs.LG eess.SP physics.med-ph

    Noise2Inpaint: Learning Referenceless Denoising by Inpainting Unrolling

    Authors: Burhaneddin Yaman, Seyed Amir Hossein Hosseini, Mehmet Akçakaya

    Abstract: Deep learning based image denoising methods have been recently popular due to their improved performance. Traditionally, these methods are trained in a supervised manner, requiring a set of noisy input and clean target image pairs. More recently, self-supervised approaches have been proposed to learn denoising from only noisy images. These methods assume that noise across pixels is statistically i… ▽ More

    Submitted 19 November, 2020; v1 submitted 16 June, 2020; originally announced June 2020.