Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 165 results for author: Anwar, S

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.04429  [pdf, ps, other

    cs.CV

    UBLLIE: Unified Backlight and Low-Light Image Enhancement

    Authors: Yasmin Yasin, Muhammad Usman, Ibrahim Radwan, Saeed Anwar

    Abstract: Backlit and low-light images often suffer from severe exposure imbalance or global underexposure, presenting significant challenges for both visual perception and downstream computer vision tasks. In this paper, we propose a unified, unsupervised enhancement framework that addresses both types of degradation without relying on paired ground-truth data. Our approach builds on CLIP-guided prompt lea… ▽ More

    Submitted 5 August, 2026; originally announced August 2026.

  2. arXiv:2607.24791  [pdf

    cs.IR cs.AI cs.CL

    From Naive RAG to Deep Agentic Retrieval: An Evolving Context Engineering Pipeline for Regulatory Compliance

    Authors: Mishca de Costa, Muhammad Saleh Anwar, Dave Mercier, Issam Hammad

    Abstract: Retrieval-augmented generation (RAG) is the dominant paradigm for applying large language models (LLMs) to enterprise document corpora, yet naive implementations encounter hard limits as corpus scale and query complexity grow. This paper traces the evolution of a production retrieval pipeline at Ontario Power Generation (OPG) for regulatory compliance and rate case analysis under Ontario Energy Bo… ▽ More

    Submitted 28 June, 2026; originally announced July 2026.

    Comments: Accepted at the 2026 IEEE the 14th International Conference on Smart Energy Grid Engineering (SEGE 2026)

  3. arXiv:2607.17551  [pdf, ps, other

    cs.CV cs.AI

    Hierarchy-Aware and Anatomy-Guided Learning for Lung Ultrasound Video Classification

    Authors: Alya Almsouti, Lotfi Mecharbat, Noha Aboukhater, Yousef Alabrach, Siddiq Anwar, Andre Kumar, Ibrahim Almakky, Mohammad Yaqub

    Abstract: Lung ultrasound (LUS) is a bedside tool for assessing pulmonary edema in patients at risk due to heart failure or impaired kidney function. However, automated LUS analysis remains challenging because of speckle noise, imaging artifacts, and operator-dependent acquisition variability. In this work, we present a deep learning framework for multi-class LUS video classification that explores two compo… ▽ More

    Submitted 20 July, 2026; originally announced July 2026.

  4. arXiv:2607.16283  [pdf, ps, other

    cs.CV cs.AI

    GenSyn10: A Multi-Generative AI Dataset For Benchmarking Image Classification

    Authors: Md Faraz Kabir Khan, Saeed Anwar, Ghulam Mubashar Hassan

    Abstract: The rapid advancement of generative AI has outpaced our ability to reliably detect its outputs, particularly when detectors encounter generators they have not seen before. We introduce GenSyn10, a CIFAR-10-aligned synthetic image dataset of 60,000 images (10 classes, 32$\times$32, 50k/10k split) generated using three architecturally diverse state-of-the-art models: FLUX.2-dev (Rectified Flow Trans… ▽ More

    Submitted 10 July, 2026; originally announced July 2026.

  5. arXiv:2606.01608  [pdf, ps, other

    cs.CV

    Exploiting Semantic and Pixel Representations for Ultra-Low Bitrate Image Compression

    Authors: Hao Wei, Yanhui Zhou, Chenyang Ge, Saeed Anwar, Ajmal Mian

    Abstract: Most existing extreme compression methods fail to achieve an optimal rate-distortion-perception trade-off, as they typically prioritize perceptual fidelity and visual realism over pixel-level accuracy. Consequently, the resulting reconstructions often deviate noticeably from the originals. Ultra-low bitrate image compression is therefore crucial-not only for producing extremely compact representat… ▽ More

    Submitted 31 May, 2026; originally announced June 2026.

  6. arXiv:2605.18880  [pdf, ps, other

    cs.LG cs.CV q-bio.QM

    A Multi-Dimensional Clustering Approach for Identifying Inborn Errors of Immunity

    Authors: Nishad Kulkarni, Alexandra K. Martinson, Nicholas L. Rider, Michael Keller, Syed Muhammad Anwar

    Abstract: Rare diseases such as inborn errors of immunity (IEI) require early diagnosis to prevent end organ damage and improve quality of life. Hurdles in accessing and curating large scale electronic health record (EHR) data limit routine data driven analyses to remain on the forefront of IEI and other rare disease trends. Development of machine learning (ML) algorithms in IEI for pattern recognition as w… ▽ More

    Submitted 15 May, 2026; originally announced May 2026.

    Comments: Accepted at EMBC 2026

  7. arXiv:2605.16775  [pdf, ps, other

    cs.CV cs.AI cs.LG

    VolTA-3D: Self-Supervised Learning for Brain MRI using 3D Volumetric Token Alignment

    Authors: Amy Makawana, Abhijeet Parida, Marius George Linguraru, Julia Ive, Syed Muhammad Anwar

    Abstract: Self-supervised learning (SSL) has advanced medical image analysis be enabling learning form large unlabelled data. However, in brain magnetic resonance imaging (MRI), most 3D models remain specialized for either segmentation of classification, limiting their ability to generalize across datasets, imaging protocols,, and downstream tasks. This lack of transferability constrains the clinical utilit… ▽ More

    Submitted 15 May, 2026; originally announced May 2026.

    Comments: Accepted at EMBC 2026

  8. arXiv:2605.00605  [pdf, ps, other

    cs.CV

    Faithful Extreme Image Rescaling with Learnable Reversible Transformation and Semantic Priors

    Authors: Hao Wei, Yanhui Zhou, Chenyang Ge, Saeed Anwar, Ajmal Mian

    Abstract: Most recent extreme rescaling methods struggle to preserve semantically consistent structures and produce realistic details, due to the severely ill-posed nature of low- to high-resolution mapping under scaling factors of $16\times$ or higher. To alleviate the above problems, we propose FaithEIR, a diffusion-based framework for extreme image rescaling. Inspired by singular value decomposition, we… ▽ More

    Submitted 1 May, 2026; originally announced May 2026.

  9. arXiv:2604.06817  [pdf, ps, other

    cs.CL

    SemEval-2026 Task 9: Detecting Multilingual, Multicultural and Multievent Online Polarization

    Authors: Usman Naseem, Robert Geislinger, Juan Ren, Sarah Kohail, Rudy Garrido Veliz, P Sam Sahil, Yiran Zhang, Marco Antonio Stranisci, Idris Abdulmumin, Özge Alaçam, Cengiz Acartürk, Aisha Jabr, Saba Anwar, Abinew Ali Ayele, Elena Tutubalina, Aung Kyaw Htet, Xintong Wang, Surendrabikram Thapa, Tanmoy Chakraborty, Dheeraj Kodati, Sahar Moradizeyveh, Firoj Alam, Ye Kyaw Thu, Shantipriya Parida, Ihsan Ayyub Qazi , et al. (9 additional authors not shown)

    Abstract: We present SemEval-2026 Task 9, a shared task on online polarization detection, covering 22 languages and comprising over 110K annotated instances. Each data instance is multi-labeled with the presence of polarization, polarization type, and polarization manifestation. Participants were asked to predict labels in three sub-tasks: (1) detecting the presence of polarization, (2) identifying the type… ▽ More

    Submitted 1 July, 2026; v1 submitted 8 April, 2026; originally announced April 2026.

  10. arXiv:2603.18418  [pdf, ps, other

    cs.CV cs.AI

    Mind the Rarities: Can Rare Skin Diseases Be Reliably Diagnosed via Diagnostic Reasoning?

    Authors: Yang Liu, Jiyao Yang, Hongjin Zhao, Xiaoyong Li, Yanzhe Ji, Xingjian Li, Runmin Jiang, Tianyang Wang, Saeed Anwar, Dongwoo Kim, Yue Yao, Zhenyue Qin, Min Xu

    Abstract: Large vision-language models (LVLMs) demonstrate strong performance in dermatology; however, evaluating diagnostic reasoning for rare conditions remains largely unexplored. Existing benchmarks focus on common diseases and assess only final accuracy, overlooking the clinical reasoning process, which is critical for complex cases. We address this gap by constructing DermCase, a long-context benchmar… ▽ More

    Submitted 18 March, 2026; originally announced March 2026.

  11. arXiv:2603.10436  [pdf, ps, other

    cs.RO cs.DC

    COHORT: Hybrid RL for Collaborative Large DNN Inference on Multi-Robot Systems Under Real-Time Constraints

    Authors: Mohammad Saeid Anwar, Anuradha Ravi, Indrajeet Ghosh, Gaurav Shinde, Carl Busart, Nirmalya Roy

    Abstract: Large deep neural networks (DNNs), especially transformer-based and multimodal architectures, are computationally demanding and challenging to deploy on resource-constrained edge platforms like field robots. These challenges intensify in mission-critical scenarios (e.g., disaster response), where robots must collaborate under tight constraints on bandwidth, latency, and battery life, often without… ▽ More

    Submitted 11 March, 2026; originally announced March 2026.

    Comments: Recently accepted at 27th IEEE International Symposium on a World of Wireless, Mobile and Multimedia Networks ( IEEE WoWMoM 2026)

  12. arXiv:2602.21142  [pdf, ps, other

    cs.CV cs.LG

    LUMEN: Longitudinal Multi-Modal Radiology Model for Prognosis and Diagnosis

    Authors: Zhifan Jiang, Dong Yang, Vishwesh Nath, Abhijeet Parida, Nishad P. Kulkarni, Ziyue Xu, Daguang Xu, Syed Muhammad Anwar, Holger R. Roth, Marius George Linguraru

    Abstract: Large vision-language models (VLMs) have evolved from general-purpose applications to specialized use cases such as in the clinical domain, demonstrating potential for decision support in radiology. One promising application is assisting radiologists in decision-making by the analysis of radiology imaging data such as chest X-rays (CXR) via a visual and natural language question-answering (VQA) in… ▽ More

    Submitted 24 February, 2026; originally announced February 2026.

    Comments: Accepted to IEEE International Symposium on Biomedical Imaging (ISBI) 2026

  13. arXiv:2601.16302  [pdf, ps, other

    cs.CV

    FeTTL: Federated Template and Task Learning for Multi-Institutional Medical Imaging

    Authors: Abhijeet Parida, Antonia Alomar, Zhifan Jiang, Pooneh Roshanitabrizi, Austin Tapp, Ziyue Xu, Syed Muhammad Anwar, Maria J. Ledesma-Carbayo, Holger R. Roth, Marius George Linguraru

    Abstract: Federated learning enables collaborative model training across geographically distributed medical centers while preserving data privacy. However, domain shifts and heterogeneity in data often lead to a degradation in model performance. Medical imaging applications are particularly affected by variations in acquisition protocols, scanner types, and patient populations. To address these issues, we i… ▽ More

    Submitted 22 January, 2026; originally announced January 2026.

  14. arXiv:2601.01406  [pdf, ps, other

    cs.CV cs.AI

    SwinIFS: Landmark Guided Swin Transformer For Identity Preserving Face Super Resolution

    Authors: Habiba Kausar, Saeed Anwar, Omar Jamal Hammad, Ibrahim Radwan, Abdul Bais

    Abstract: Face super-resolution aims to recover high-quality facial images from severely degraded low-resolution inputs, but remains challenging due to the loss of fine structural details and identity-specific features. This work introduces SwinIFS, a landmark-guided super-resolution framework that integrates structural priors with hierarchical attention mechanisms to achieve identity-preserving reconstruct… ▽ More

    Submitted 10 July, 2026; v1 submitted 4 January, 2026; originally announced January 2026.

  15. arXiv:2512.23894  [pdf, ps, other

    cs.CV

    MRI-to-CT Synthesis With Cranial Suture Segmentations Using A Variational Autoencoder Framework

    Authors: Krithika Iyer, Austin Tapp, Athelia Paulli, Gabrielle Dickerson, Syed Muhammad Anwar, Natasha Lepore, Marius George Linguraru

    Abstract: Quantifying normative pediatric cranial development and suture ossification is crucial for diagnosing and treating growth-related cephalic disorders. Computed tomography (CT) is widely used to evaluate cranial and sutural deformities; however, its ionizing radiation is contraindicated in children without significant abnormalities. Magnetic resonance imaging (MRI) offers radiation free scans with s… ▽ More

    Submitted 29 December, 2025; originally announced December 2025.

  16. arXiv:2512.19350  [pdf, ps, other

    cs.AI

    PENDULUM: A Benchmark for Assessing Sycophancy in Multimodal Large Language Models

    Authors: A. B. M. Ashikur Rahman, Saeed Anwar, Muhammad Usman, Irfan Ahmad, Ajmal Mian

    Abstract: Sycophancy, an excessive tendency of AI models to agree with user input at the expense of factual accuracy or in contradiction of visual evidence, poses a critical and underexplored challenge for multimodal large language models (MLLMs). While prior studies have examined this behavior in text-only settings of large language models, existing research on visual or multimodal counterparts remains lim… ▽ More

    Submitted 22 December, 2025; originally announced December 2025.

  17. Elevating Intrusion Detection and Security Fortification in Intelligent Networks through Cutting-Edge Machine Learning Paradigms

    Authors: Md Minhazul Islam Munna, Md Mahbubur Rahman, Jaroslav Frnda, Muhammad Shahid Anwar, Alpamis Kutlimuratov

    Abstract: The proliferation of IoT devices and their reliance on Wi-Fi networks have introduced significant security vulnerabilities, particularly the KRACK and Kr00k attacks, which exploit weaknesses in WPA2 encryption to intercept and manipulate sensitive data. Traditional IDS using classifiers face challenges such as model overfitting, incomplete feature extraction, and high false positive rates, limitin… ▽ More

    Submitted 22 December, 2025; originally announced December 2025.

    Journal ref: Sci Rep 15, 39989 (2025)

  18. arXiv:2512.15808  [pdf

    q-bio.QM cs.AI cs.CV cs.LG

    Foundation Models in Biomedical Imaging: Turning Hype into Reality

    Authors: Amgad Muneer, Kai Zhang, Ibraheem Hamdi, Rizwan Qureshi, Muhammad Waqas, Shereen Fouad, Hazrat Ali, Syed Muhammad Anwar, Jia Wu

    Abstract: Foundation models (FMs) are driving a prominent shift in biomedical imaging from task-specific models to unified backbone models for diverse tasks. This opens an avenue to integrate imaging, pathology, clinical records, and genomics data into a composite system. However, this vision contrasts sharply with modern medicine's trajectory toward more granular sub-specialization. This tension, coupled w… ▽ More

    Submitted 21 April, 2026; v1 submitted 17 December, 2025; originally announced December 2025.

    Comments: 9 figures and 3 tables

    Journal ref: Nature Biomedical Engineering 10, 1557-1575 (2026)

  19. Improving Pre-trained Adult Glioma Segmentation Models Using only Post-processing Techniques

    Authors: Abhijeet Parida, Daniel Capellán-Martín, Zhifan Jiang, Nishad Kulkarni, Krithika Iyer, Austin Tapp, Syed Muhammad Anwar, María J. Ledesma-Carbayo, Marius George Linguraru

    Abstract: Gliomas are the most common malignant brain tumors in adults and are among the most lethal. Despite aggressive treatment, the median survival rate is less than 15 months. Accurate multiparametric MRI (mpMRI) tumor segmentation is critical for surgical planning, radiotherapy, and disease monitoring. While deep learning models have improved the accuracy of automated segmentation, large-scale pre-tra… ▽ More

    Submitted 11 June, 2026; v1 submitted 16 December, 2025; originally announced December 2025.

  20. Adaptable Segmentation Pipeline for Diverse Brain Tumors with Radiomic-Guided Subtyping and Lesion-Wise Model Ensemble

    Authors: Daniel Capellán-Martín, Abhijeet Parida, Zhifan Jiang, Nishad Kulkarni, Krithika Iyer, Austin Tapp, Syed Muhammad Anwar, María J. Ledesma-Carbayo, Marius George Linguraru

    Abstract: Robust and generalizable segmentation of brain tumors on multi-parametric magnetic resonance imaging (MRI) remains difficult because tumor types differ widely. The BraTS 2025 Lighthouse Challenge benchmarks segmentation methods on diverse high-quality datasets of adult and pediatric tumors: multi-consortium international pediatric brain tumor segmentation (PED), preoperative meningioma tumor segme… ▽ More

    Submitted 11 June, 2026; v1 submitted 16 December, 2025; originally announced December 2025.

    Comments: 12 pages, 5 figures, 3 tables. Algorithm presented at MICCAI BraTS 2025

  21. arXiv:2511.20065  [pdf, ps, other

    cs.CV

    FLaTEC: Frequency-Disentangled Latent Triplanes for Efficient Compression of LiDAR Point Clouds

    Authors: Xiaoge Zhang, Zijie Wu, Mingtao Feng, Zichen Geng, Mehwish Nasim, Saeed Anwar, Ajmal Mian

    Abstract: Point cloud compression methods jointly optimize bitrates and reconstruction distortion. However, balancing compression ratio and reconstruction quality is difficult because low-frequency and high-frequency components contribute differently at the same resolution. To address this, we propose FLaTEC, a frequency-aware compression model that enables the compression of a full scan with high compressi… ▽ More

    Submitted 25 November, 2025; originally announced November 2025.

  22. arXiv:2511.12810  [pdf, ps, other

    cs.CV cs.AI eess.IV

    MSRNet: A Multi-Scale Recursive Network for Camouflaged Object Detection

    Authors: Leena Alghamdi, Muhammad Usman, Hafeez Anwar, Abdul Bais, Saeed Anwar

    Abstract: Camouflaged object detection is an emerging and challenging computer vision task that requires identifying and segmenting objects that blend seamlessly into their environments due to high similarity in color, texture, and size. This task is further complicated by low-light conditions, partial occlusion, small object size, intricate background patterns, and multiple objects. While many sophisticate… ▽ More

    Submitted 8 July, 2026; v1 submitted 16 November, 2025; originally announced November 2025.

  23. arXiv:2511.12627  [pdf, ps, other

    cs.CV cs.AI eess.IV

    C3Net: Context-Contrast Network for Camouflaged Object Detection

    Authors: Baber Jan, Aiman H. El-Maleh, Abdul Jabbar Siddiqui, Abdul Bais, Saeed Anwar

    Abstract: Camouflaged object detection identifies objects that blend seamlessly with their surroundings through similar colors, textures, and patterns. This task challenges both traditional segmentation methods and modern foundation models, which fail dramatically on camouflaged objects. We identify six fundamental challenges in COD: Intrinsic Similarity, Edge Disruption, Extreme Scale Variation, Environmen… ▽ More

    Submitted 16 November, 2025; originally announced November 2025.

  24. Post-Processing Methods for Improving Accuracy in MRI Inpainting

    Authors: Nishad Kulkarni, Krithika Iyer, Austin Tapp, Abhijeet Parida, Daniel Capellán-Martín, Zhifan Jiang, María J. Ledesma-Carbayo, Syed Muhammad Anwar, Marius George Linguraru

    Abstract: Magnetic Resonance Imaging (MRI) is the primary imaging modality used in the diagnosis, assessment, and treatment planning for brain pathologies. However, most automated MRI analysis tools, such as segmentation and registration pipelines, are optimized for healthy anatomies and often fail when confronted with large lesions such as tumors. To overcome this, image inpainting techniques aim to locall… ▽ More

    Submitted 11 April, 2026; v1 submitted 16 October, 2025; originally announced October 2025.

  25. arXiv:2510.04472  [pdf, ps, other

    cs.CV cs.AI cs.LG eess.IV

    SPEGNet: Synergistic Perception-Guided Network for Camouflaged Object Detection

    Authors: Baber Jan, Saeed Anwar, Aiman H. El-Maleh, Abdul Jabbar Siddiqui, Abdul Bais

    Abstract: Camouflaged object detection segments objects with intrinsic similarity and edge disruption. Current detection methods rely on accumulated complex components. Each approach adds components such as boundary modules, attention mechanisms, and multi-scale processors independently. This accumulation creates a computational burden without proportional gains. To manage this complexity, they process at r… ▽ More

    Submitted 6 October, 2025; originally announced October 2025.

  26. arXiv:2509.11714  [pdf, ps, other

    eess.IV cs.LG

    EMeRALDS: Electronic Medical Record Driven Automated Lung Nodule Detection and Classification in Thoracic CT Images

    Authors: Hafza Eman, Furqan Shaukat, Muhammad Hamza Zafar, Syed Muhammad Anwar

    Abstract: Objective: Lung cancer is a leading cause of cancer-related mortality worldwide, primarily due to delayed diagnosis and poor early detection. This study aims to develop a computer-aided diagnosis (CAD) system that leverages large vision-language models (VLMs) for the accurate detection and classification of pulmonary nodules in computed tomography (CT) scans. Methods: We propose an end-to-end CA… ▽ More

    Submitted 15 September, 2025; originally announced September 2025.

  27. arXiv:2508.18161  [pdf, ps, other

    quant-ph cs.LG

    Hybrid Quantum-Classical Learning for Multiclass Image Classification

    Authors: Shuchismita Anwar, Sowmitra Das, Muhammad Iqbal Hossain, Jishnu Mahmud

    Abstract: This study explores the challenge of improving multiclass image classification through quantum machine-learning techniques. It explores how the discarded qubit states of Noisy Intermediate-Scale Quantum (NISQ) quantum convolutional neural networks (QCNNs) can be leveraged alongside a classical classifier to improve classification performance. Current QCNNs discard qubit states after pooling; yet,… ▽ More

    Submitted 25 August, 2025; originally announced August 2025.

    Comments: 13 pages, 8 figures

  28. arXiv:2508.13068  [pdf, ps, other

    cs.CV cs.LG

    Eyes on the Image: Gaze Supervised Multimodal Learning for Chest X-ray Diagnosis and Report Generation

    Authors: Tanjim Islam Riju, Shuchismita Anwar, Saman Sarker Joy, Farig Sadeque, Swakkhar Shatabda

    Abstract: Medical vision-language models still struggle to match radiologists' attention and to verbalize findings with explicit spatial grounding. We address this gap with a two-stage multimodal framework for chest X-ray interpretation built on the MIMIC-Eye dataset. In the first stage introduces a gaze-token classifier that fuses image patches, bounding-box masks, transcription embeddings, and radiologist… ▽ More

    Submitted 18 August, 2026; v1 submitted 18 August, 2025; originally announced August 2025.

  29. arXiv:2507.19304  [pdf, ps, other

    cs.CV cs.AI

    Multistream Network for LiDAR and Camera-based 3D Object Detection in Outdoor Scenes

    Authors: Muhammad Ibrahim, Naveed Akhtar, Haitian Wang, Saeed Anwar, Ajmal Mian

    Abstract: Fusion of LiDAR and RGB data has the potential to enhance outdoor 3D object detection accuracy. To address real-world challenges in outdoor 3D object detection, fusion of LiDAR and RGB input has started gaining traction. However, effective integration of these modalities for precise object detection task still remains a largely open problem. To address that, we propose a MultiStream Detection (MuS… ▽ More

    Submitted 25 July, 2025; originally announced July 2025.

    Comments: This paper has been accepted by IEEE/RSJ IROS 2025 for oral presentation on 19 Oct. 2025

  30. arXiv:2506.20302  [pdf, ps, other

    cs.CV

    TDiR: Transformer based Diffusion for Image Restoration Tasks

    Authors: Abbas Anwar, Ibrahim Radwan, Ali Arshad Nasir, Mudassir Masood, Saeed Anwar

    Abstract: Images captured in challenging environments often experience various types of degradation, such as noise, color cast, blur, and light scattering. These issues significantly lower image quality, thereby reducing their usefulness in downstream tasks such as object detection, mapping, and classification. Our transformer-based diffusion model was developed to address image restoration challenges and e… ▽ More

    Submitted 24 July, 2026; v1 submitted 25 June, 2025; originally announced June 2025.

  31. arXiv:2506.13925  [pdf, ps, other

    cs.CV cs.AI

    Segmenting Visuals With Querying Words: Language Anchors For Semi-Supervised Image Segmentation

    Authors: Numair Nadeem, Saeed Anwar, Muhammad Hamza Asad, Abdul Bais

    Abstract: Vision Language Models (VLMs) provide rich semantic priors but are underexplored in Semi supervised Semantic Segmentation. Recent attempts to integrate VLMs to inject high level semantics overlook the semantic misalignment between visual and textual representations that arises from using domain invariant text embeddings without adapting them to dataset and image specific contexts. This lack of dom… ▽ More

    Submitted 21 March, 2026; v1 submitted 16 June, 2025; originally announced June 2025.

  32. arXiv:2506.01214  [pdf, ps, other

    cs.CV cs.AI

    A Review on Coarse to Fine-Grained Animal Action Recognition

    Authors: Ali Zia, Renuka Sharma, Abdelwahed Khamis, Xuesong Li, Muhammad Husnain, Numan Shafi, Saeed Anwar, Sabine Schmoelzl, Eric Stone, Lars Petersson, Vivien Rolland

    Abstract: This review provides an in-depth exploration of the field of animal action recognition, focusing on coarse-grained (CG) and fine-grained (FG) techniques. The primary aim is to examine the current state of research in animal behaviour recognition and to elucidate the unique challenges associated with recognising subtle animal actions in outdoor environments. These challenges differ significantly fr… ▽ More

    Submitted 1 June, 2025; originally announced June 2025.

  33. arXiv:2505.24026  [pdf, other

    cs.CV cs.AI

    MaskAdapt: Unsupervised Geometry-Aware Domain Adaptation Using Multimodal Contextual Learning and RGB-Depth Masking

    Authors: Numair Nadeem, Muhammad Hamza Asad, Saeed Anwar, Abdul Bais

    Abstract: Semantic segmentation of crops and weeds is crucial for site-specific farm management; however, most existing methods depend on labor intensive pixel-level annotations. A further challenge arises when models trained on one field (source domain) fail to generalize to new fields (target domain) due to domain shifts, such as variations in lighting, camera setups, soil composition, and crop growth sta… ▽ More

    Submitted 29 May, 2025; originally announced May 2025.

    Comments: 11 pages, 5 figures, presented at the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025. Reviewer comments available upon request

  34. arXiv:2505.20624  [pdf, ps, other

    cs.CL

    POLAR: A Benchmark for Multilingual, Multicultural, and Multi-Event Online Polarization

    Authors: Usman Naseem, Robert Geislinger, Juan Ren, Sarah Kohail, Rudy Garrido Veliz, P Sam Sahil, Yiran Zhang, Marco Antonio Stranisci, Idris Abdulmumin, Özge Alacam, Cengiz Acartürk, Aisha Jabr, Saba Anwar, Abinew Ali Ayele, Simona Frenda, Alessandra Teresa Cignarella, Elena Tutubalina, Oleg Rogov, Aung Kyaw Htet, Xintong Wang, Surendrabikram Thapa, Kritesh Rauniyar, Tanmoy Chakraborty, Arfeen Zeeshan, Dheeraj Kodati , et al. (18 additional authors not shown)

    Abstract: Online polarization poses a growing challenge for democratic discourse, yet most computational social science research remains monolingual, culturally narrow, or event-specific. We introduce POLAR, a multilingual, multicultural, and multi-event dataset with over 110K instances in 22 languages drawn from diverse online platforms and real-world events. Polarization is annotated along three axes, nam… ▽ More

    Submitted 5 February, 2026; v1 submitted 26 May, 2025; originally announced May 2025.

    Comments: Preprint

  35. arXiv:2503.06094  [pdf, other

    cs.CV

    PointDiffuse: A Dual-Conditional Diffusion Model for Enhanced Point Cloud Semantic Segmentation

    Authors: Yong He, Hongshan Yu, Mingtao Feng, Tongjia Chen, Zechuan Li, Anwaar Ulhaq, Saeed Anwar, Ajmal Saeed Mian

    Abstract: Diffusion probabilistic models are traditionally used to generate colors at fixed pixel positions in 2D images. Building on this, we extend diffusion models to point cloud semantic segmentation, where point positions also remain fixed, and the diffusion model generates point labels instead of colors. To accelerate the denoising process in reverse diffusion, we introduce a noisy label embedding mec… ▽ More

    Submitted 11 March, 2025; v1 submitted 8 March, 2025; originally announced March 2025.

    Comments: 8 pages, 3 figures, 7 tables

  36. arXiv:2503.04509  [pdf, other

    cs.LG cs.AI

    STX-Search: Explanation Search for Continuous Dynamic Spatio-Temporal Models

    Authors: Saif Anwar, Nathan Griffiths, Thomas Popham, Abhir Bhalerao

    Abstract: Recent improvements in the expressive power of spatio-temporal models have led to performance gains in many real-world applications, such as traffic forecasting and social network modelling. However, understanding the predictions from a model is crucial to ensure reliability and trustworthiness, particularly for high-risk applications, such as healthcare and transport. Few existing methods are abl… ▽ More

    Submitted 6 March, 2025; originally announced March 2025.

  37. arXiv:2502.15198  [pdf

    cs.LG eess.SP q-bio.NC

    Graph-Based Deep Learning on Stereo EEG for Predicting Seizure Freedom in Epilepsy Patients

    Authors: Artur Agaronyan, Syeda Abeera Amir, Nunthasiri Wittayanakorn, John Schreiber, Marius G. Linguraru, William Gaillard, Chima Oluigbo, Syed Muhammad Anwar

    Abstract: Predicting seizure freedom is essential for tailoring epilepsy treatment. But accurate prediction remains challenging with traditional methods, especially with diverse patient populations. This study developed a deep learning-based graph neural network (GNN) model to predict seizure freedom from stereo electroencephalography (sEEG) data in patients with refractory epilepsy. We utilized high-qualit… ▽ More

    Submitted 20 February, 2025; originally announced February 2025.

  38. arXiv:2502.14476  [pdf, other

    cs.CL

    Argument-Based Comparative Question Answering Evaluation Benchmark

    Authors: Irina Nikishina, Saba Anwar, Nikolay Dolgov, Maria Manina, Daria Ignatenko, Viktor Moskvoretskii, Artem Shelmanov, Tim Baldwin, Chris Biemann

    Abstract: In this paper, we aim to solve the problems standing in the way of automatic comparative question answering. To this end, we propose an evaluation framework to assess the quality of comparative question answering summaries. We formulate 15 criteria for assessing comparative answers created using manual annotation and annotation from 6 large language models and two comparative question asnwering da… ▽ More

    Submitted 20 February, 2025; originally announced February 2025.

    Comments: 8 pages, 7 Tables, 13 Figures, 18 pages with Appendix

  39. arXiv:2501.16346  [pdf, other

    cs.LG cs.AI

    Self-supervised Graph Transformer with Contrastive Learning for Brain Connectivity Analysis towards Improving Autism Detection

    Authors: Yicheng Leng, Syed Muhammad Anwar, Islem Rekik, Sen He, Eung-Joo Lee

    Abstract: Functional Magnetic Resonance Imaging (fMRI) provides useful insights into the brain function both during task or rest. Representing fMRI data using correlation matrices is found to be a reliable method of analyzing the inherent connectivity of the brain in the resting and active states. Graph Neural Networks (GNNs) have been widely used for brain network analysis due to their inherent explainabil… ▽ More

    Submitted 18 January, 2025; originally announced January 2025.

  40. arXiv:2501.15737  [pdf

    eess.IV cs.LG

    Geometric Deep Learning for Automated Landmarking of Maxillary Arches on 3D Oral Scans from Newborns with Cleft Lip and Palate

    Authors: Artur Agaronyan, HyeRan Choo, Marius Linguraru, Syed Muhammad Anwar

    Abstract: Rapid advances in 3D model scanning have enabled the mass digitization of dental clay models. However, most clinicians and researchers continue to use manual morphometric analysis methods on these models such as landmarking. This is a significant step in treatment planning for craniomaxillofacial conditions. We aimed to develop and test a geometric deep learning model that would accurately and rel… ▽ More

    Submitted 26 January, 2025; originally announced January 2025.

    Comments: ISBI 2025

  41. arXiv:2501.02822  [pdf, other

    cs.CV cs.AI cs.RO

    RDD4D: 4D Attention-Guided Road Damage Detection And Classification

    Authors: Asma Alkalbani, Muhammad Saqib, Ahmed Salim Alrawahi, Abbas Anwar, Chandarnath Adak, Saeed Anwar

    Abstract: Road damage detection and assessment are crucial components of infrastructure maintenance. However, current methods often struggle with detecting multiple types of road damage in a single image, particularly at varying scales. This is due to the lack of road datasets with various damage types having varying scales. To overcome this deficiency, first, we present a novel dataset called Diverse Road… ▽ More

    Submitted 6 January, 2025; originally announced January 2025.

  42. arXiv:2412.10595  [pdf, ps, other

    cs.IR cs.CY cs.GT

    Recommendation and Temptation

    Authors: Md Sanzeed Anwar, Paramveer S. Dhillon, Grant Schoenebeck

    Abstract: Traditional recommender systems based on revealed preferences often fail to capture the fundamental duality in user behavior, where consumption choices are driven by both inherent value (enrichment) and instant appeal (temptation). Consequently, these systems may generate recommendations that prioritize short-term engagement over long-lasting user satisfaction. We propose a novel recommender desig… ▽ More

    Submitted 23 July, 2025; v1 submitted 13 December, 2024; originally announced December 2024.

    Comments: Published in Proceedings of the 19th ACM Conference on Recommender Systems (RecSys 2025)

  43. arXiv:2412.04111  [pdf, other

    eess.IV cs.CV

    Adult Glioma Segmentation in Sub-Saharan Africa using Transfer Learning on Stratified Finetuning Data

    Authors: Abhijeet Parida, Daniel Capellán-Martín, Zhifan Jiang, Austin Tapp, Xinyang Liu, Syed Muhammad Anwar, María J. Ledesma-Carbayo, Marius George Linguraru

    Abstract: Gliomas, a kind of brain tumor characterized by high mortality, present substantial diagnostic challenges in low- and middle-income countries, particularly in Sub-Saharan Africa. This paper introduces a novel approach to glioma segmentation using transfer learning to address challenges in resource-limited regions with minimal and low-quality MRI data. We leverage pre-trained deep learning models,… ▽ More

    Submitted 20 December, 2024; v1 submitted 5 December, 2024; originally announced December 2024.

    Comments: 10 pages, 3 figures, 3 tables. BraTS-Africa 2024 winning algorithm presented at MICCAI BraTS 2024

  44. arXiv:2412.04094  [pdf, other

    eess.IV cs.CV

    Magnetic Resonance Imaging Feature-Based Subtyping and Model Ensemble for Enhanced Brain Tumor Segmentation

    Authors: Zhifan Jiang, Daniel Capellán-Martín, Abhijeet Parida, Austin Tapp, Xinyang Liu, María J. Ledesma-Carbayo, Syed Muhammad Anwar, Marius George Linguraru

    Abstract: Accurate and automatic segmentation of brain tumors in multi-parametric magnetic resonance imaging (mpMRI) is essential for quantitative measurements, which play an increasingly important role in clinical diagnosis and prognosis. The International Brain Tumor Segmentation (BraTS) Challenge 2024 offers a unique benchmarking opportunity, including various types of brain tumors in both adult and pedi… ▽ More

    Submitted 20 December, 2024; v1 submitted 5 December, 2024; originally announced December 2024.

    Comments: 11 pages, 4 figures, 3 tables. This paper was accepted at MICCAI-BraTS 2024

  45. arXiv:2412.02695  [pdf, other

    cs.CY cs.LG eess.SP

    An ADHD Diagnostic Interface Based on EEG Spectrograms and Deep Learning Techniques

    Authors: Medha Pappula, Syed Muhammad Anwar

    Abstract: This paper introduces an innovative approach to Attention-deficit/hyperactivity disorder (ADHD) diagnosis by employing deep learning (DL) techniques on electroencephalography (EEG) signals. This method addresses the limitations of current behavior-based diagnostic methods, which often lead to misdiagnosis and gender bias. By utilizing a publicly available EEG dataset and converting the signals int… ▽ More

    Submitted 3 December, 2024; originally announced December 2024.

    Comments: Presented at SIPAIM 2024

  46. arXiv:2411.17392  [pdf, ps, other

    cs.CV

    NumGrad-Pull: Numerical Gradient Guided Tri-plane Representation for Surface Reconstruction from Point Clouds

    Authors: Ruikai Cui, Binzhu Xie, Shi Qiu, Jiawei Liu, Saeed Anwar, Nick Barnes

    Abstract: Reconstructing continuous surfaces from unoriented and unordered 3D points is a fundamental challenge in computer vision and graphics. Recent advancements address this problem by training neural signed distance functions to pull 3D location queries to their closest points on a surface, following the predicted signed distances and the analytical gradients computed by the network. In this paper, we… ▽ More

    Submitted 7 July, 2026; v1 submitted 26 November, 2024; originally announced November 2024.

    Comments: Accepted in IEEE TVCG

  47. arXiv:2411.13860  [pdf, ps, other

    cs.CV eess.IV

    DiffCom: Decoupled Sparse Priors Guided Diffusion Compression for Point Clouds

    Authors: Xiaoge Zhang, Zijie Wu, Mehwish Nasim, Mingtao Feng, Saeed Anwar, Ajmal Mian

    Abstract: Lossy compression relies on an autoencoder to transform a point cloud into latent points for storage, leaving the inherent redundancy of latent representations unexplored. To reduce redundancy in latent points, we propose a diffusion-based framework guided by sparse priors that achieves high reconstruction quality, especially at low bitrates. Our approach features an efficient dual-density data fl… ▽ More

    Submitted 6 October, 2025; v1 submitted 21 November, 2024; originally announced November 2024.

  48. A New Logic For Pediatric Brain Tumor Segmentation

    Authors: Max Bengtsson, Elif Keles, Gorkem Durak, Syed Anwar, Yuri S. Velichko, Marius G. Linguraru, Angela J. Waanders, Ulas Bagci

    Abstract: In this paper, we present a novel approach for segmenting pediatric brain tumors using a deep learning architecture, inspired by expert radiologists' segmentation strategies. Our model delineates four distinct tumor labels and is benchmarked on a held-out PED BraTS 2024 test set (i.e., pediatric brain tumor datasets introduced by BraTS). Furthermore, we evaluate our model's performance against the… ▽ More

    Submitted 18 February, 2025; v1 submitted 2 November, 2024; originally announced November 2024.

  49. arXiv:2410.09831  [pdf, other

    cs.CV cs.AI cs.CE

    LoLI-Street: Benchmarking Low-Light Image Enhancement and Beyond

    Authors: Md Tanvir Islam, Inzamamul Alam, Simon S. Woo, Saeed Anwar, IK Hyun Lee, Khan Muhammad

    Abstract: Low-light image enhancement (LLIE) is essential for numerous computer vision tasks, including object detection, tracking, segmentation, and scene understanding. Despite substantial research on improving low-quality images captured in underexposed conditions, clear vision remains critical for autonomous vehicles, which often struggle with low-light scenarios, signifying the need for continuous rese… ▽ More

    Submitted 13 October, 2024; originally announced October 2024.

    Comments: Accepted by the Asian Conference on Computer Vision (ACCV 2024)

  50. arXiv:2410.05771  [pdf, other

    cs.CV

    Cefdet: Cognitive Effectiveness Network Based on Fuzzy Inference for Action Detection

    Authors: Zhe Luo, Weina Fu, Shuai Liu, Saeed Anwar, Muhammad Saqib, Sambit Bakshi, Khan Muhammad

    Abstract: Action detection and understanding provide the foundation for the generation and interaction of multimedia content. However, existing methods mainly focus on constructing complex relational inference networks, overlooking the judgment of detection effectiveness. Moreover, these methods frequently generate detection results with cognitive abnormalities. To solve the above problems, this study propo… ▽ More

    Submitted 16 October, 2024; v1 submitted 8 October, 2024; originally announced October 2024.

    Comments: The paper has been accepted by ACM MM. If you find this work helpful, please consider citing our paper. Zhe Luo, Weina Fu, Shuai Liu, Saeed Anwar, Muhammad Saqib, Sambit Bakshi, Khan Muhammad (2024) Cefdet: Cognitive Effectiveness Network Based on Fuzzy Inference for Action Detection, 32nd ACM International Conference on Multimedia, online first, 10.1145/3664647.3681226