Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 60 results for author: Jaiswal, S

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.01462  [pdf, ps, other

    cs.AI

    Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI

    Authors: Quang Bui, Shlok Jaiswal, Samuel Paik-Heintz, Kevin Zhou, Kaushik Madapati, Krittaphas Chaisutyakorn, Noah Dane Hebdon, Dimitrios Proios, Sebastián Andrés Cajas Ordóñez, Kacper Dobek, Boya Zhang, Aly Dhedhi, Ahram Han, Kushul Reddy Palakala, Rahul Gorijavolu, Jacques Kpodonu, Leo Anthony Celi

    Abstract: Multimodal clinical models are usually judged on accuracy with every modality present, but deployment removes modalities; an echocardiogram is often unavailable where an ECG is routine. Two questions then matter beyond the size of the accuracy loss: which modality was responsible, and whether the model fails loudly or silently once that modality is dropped. The distinction is per-example and modal… ▽ More

    Submitted 2 August, 2026; originally announced August 2026.

  2. arXiv:2607.04500  [pdf, ps, other

    cs.CV

    Geographic Diversity Beats Data Volume for Cross-Domain Generalization in Zero-Label JEPA Driving World Models

    Authors: Santosh Jaiswal

    Abstract: Self-supervised latent world models can assign a surprise score to driving scenarios without any human labels. A natural follow-up question is whether such a model, trained on driving data from one geographic region, can generalize its notion of complexity to unseen cities and sensor configurations. We study this question through a controlled transfer experiment: we train JEPA-based world models o… ▽ More

    Submitted 5 July, 2026; originally announced July 2026.

    Comments: 9 pages, 3 figures

  3. arXiv:2606.28383  [pdf, ps, other

    cs.CV cs.LG

    Zero-Label Driving Scenario Complexity Detection via Joint Embedding Predictive Architecture

    Authors: Santosh Jaiswal

    Abstract: Identifying complex and safety-critical driving scenarios in large unlabelled datasets is an important but expensive problem. Existing approaches rely on human annotators, supervised classifiers, or carefully engineered rule sets, all of which require substantial prior knowledge about what constitutes a difficult scenario. We ask whether a model can discover scenario complexity on its own, with no… ▽ More

    Submitted 21 June, 2026; originally announced June 2026.

    Comments: 12 pages, 6 figures

  4. arXiv:2605.08093  [pdf, ps, other

    cs.CY cs.AI cs.HC

    Playing Games with My Heart: An Evaluation of AI Companion Apps

    Authors: Maribeth Rauh, Dick A. H. Blankvoort, Matias Duran, Caoilfhionn Ní Dheoráin, Harshvardhan J. Pandit, Syrine Enneifer, Siddharth D. Jaiswal, Anthony Ventresque, Abeba Birhane

    Abstract: The use of chatbots for various forms of companionship is growing rapidly, raising a myriad of questions about simulated relationships, emotional dependence, and psychological harm. While major platforms such as ChatGPT, Grok, and Character AI are the subject of a growing body of research and legal inquiries, apps explicitly built for simulating intimate interpersonal relationships remain under-ex… ▽ More

    Submitted 7 August, 2026; v1 submitted 8 April, 2026; originally announced May 2026.

    ACM Class: K.4.2; I.2.1

  5. arXiv:2604.23094  [pdf, ps, other

    cs.CV cs.GR cs.LG

    FusionRelight: Relighting Portraits in Real Time via Hybrid Domain Knowledge Fusion

    Authors: Qian Huang, Mayoore Selvarasa Jaiswal, Zhen Zhong, Rochelle Pereira, Jianyuan Min

    Abstract: Portrait relighting is a low-level vision problem in which physically plausible illumination transfer, identity preservation, and compact real-time inference must be considered together. Iterative diffusion-style methods can synthesize fine detail, but stochastic inference and cost complicate deterministic live video creation; physically grounded relighting preserves identity, but controlled synth… ▽ More

    Submitted 7 August, 2026; v1 submitted 24 April, 2026; originally announced April 2026.

  6. arXiv:2602.06917  [pdf, ps, other

    eess.AS cs.LG

    Automatic Detection and Analysis of Singing Mistakes for Music Pedagogy

    Authors: Sumit Kumar, Suraj Jaiswal, Parampreet Singh, Vipul Arora

    Abstract: The advancement of machine learning in audio analysis has opened new possibilities for technology-enhanced music education. This paper introduces a framework for automatic singing mistake detection in the context of music pedagogy, supported by a newly curated dataset. The dataset comprises synchronized teacher learner vocal recordings, with annotations marking different types of mistakes made by… ▽ More

    Submitted 6 February, 2026; originally announced February 2026.

    Comments: Under Review at Transactions of Audio Speech and Language Processing

  7. arXiv:2601.15286  [pdf, ps, other

    cs.CV cs.AI cs.LG cs.RO

    Iterative Refinement Improves Compositional Image Generation

    Authors: Shantanu Jaiswal, Mihir Prabhudesai, Nikash Bhardwaj, Zheyang Qin, Amir Zadeh, Chuan Li, Katerina Fragkiadaki, Deepak Pathak

    Abstract: Text-to-image (T2I) models have achieved remarkable progress, yet they continue to struggle with complex prompts that require simultaneously handling multiple objects, relations, and attributes. Existing inference-time strategies, such as parallel sampling with verifiers or simply increasing denoising steps, can improve prompt alignment but remain inadequate for richly compositional settings where… ▽ More

    Submitted 21 January, 2026; originally announced January 2026.

    Comments: Project webpage: https://iterative-img-gen.github.io/

  8. arXiv:2511.22880  [pdf, ps, other

    cs.DC cs.AI cs.LG

    Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems

    Authors: Shashwat Jaiswal, Shrikara Arun, Anjaly Parayil, Ankur Mallick, Spyros Mastorakis, Alind Khare, Chloi Alverti, Renee St Amant, Chetan Bansal, Victor Rühle, Josep Torrellas

    Abstract: Low-Rank Adaptation (LoRA) has become the de facto method for parameter-efficient fine-tuning of large language models (LLMs), enabling rapid adaptation to diverse domains. In production, LoRA-based models are served at scale, creating multi-tenant environments with hundreds of adapters sharing a base model. However, state-of-the-art serving systems co-batch heterogeneous adapters without accounti… ▽ More

    Submitted 28 November, 2025; originally announced November 2025.

  9. arXiv:2510.00088  [pdf, ps, other

    cs.AI cs.CY

    Judging by Appearances? Auditing and Intervening Vision-Language Models for Bail Prediction

    Authors: Sagnik Basu, Shubham Prakash, Ashish Maruti Barge, Siddharth D Jaiswal, Abhisek Dash, Saptarshi Ghosh, Animesh Mukherjee

    Abstract: Large language models (LLMs) have been extensively used for legal judgment prediction tasks based on case reports and crime history. However, with a surge in the availability of large vision language models (VLMs), legal judgment prediction systems can now be made to leverage the images of the criminals in addition to the textual case reports/crime history. Applications built in this way could lea… ▽ More

    Submitted 30 September, 2025; originally announced October 2025.

  10. arXiv:2509.03518  [pdf, ps, other

    cs.LG

    Can LLMs Lie? Investigation beyond Hallucination

    Authors: Haoran Huan, Mihir Prabhudesai, Mengning Wu, Shantanu Jaiswal, Deepak Pathak

    Abstract: Large language models (LLMs) have demonstrated impressive capabilities across a variety of tasks, but their increasing autonomy in real-world applications raises concerns about their trustworthiness. While hallucinations-unintentional falsehoods-have been widely studied, the phenomenon of lying, where an LLM knowingly generates falsehoods to achieve an ulterior objective, remains underexplored. In… ▽ More

    Submitted 3 September, 2025; originally announced September 2025.

    Comments: Website at https://llm-liar.github.io/

  11. arXiv:2508.07124  [pdf, ps, other

    cs.DC cs.DB

    AerialDB: A Federated Peer-to-Peer Spatio-temporal Edge Datastore for Drone Fleets

    Authors: Shashwat Jaiswal, Suman Raj, Subhajit Sidhanta, Yogesh Simmhan

    Abstract: Recent years have seen an unprecedented growth in research that leverages the newest computing paradigm of Internet of Drones, comprising a fleet of connected Unmanned Aerial Vehicles (UAVs) used for a wide range of tasks such as monitoring and analytics in highly mobile and changing environments characteristic of disaster regions. Given that the typical data (i.e., videos and images) collected by… ▽ More

    Submitted 9 August, 2025; originally announced August 2025.

  12. arXiv:2507.12373  [pdf, ps, other

    cs.ET eess.SP

    Emerging Paradigms in the Energy Sector: Forecasting and System Control Optimisation

    Authors: Dariush Pourkeramati, Gareth Wadge, Rachel Hassall, Charlotte Mitchell, Anish Khadka, Shiwang Jaiswal, Andrew Duncan, Rossella Arcucci

    Abstract: The energy sector is experiencing rapid transformation due to increasing renewable energy integration, decentralisation of power systems, and a heightened focus on efficiency and sustainability. With energy demand becoming increasingly dynamic and generation sources more variable, advanced forecasting and optimisation strategies are crucial for maintaining grid stability, cost-effectiveness, and e… ▽ More

    Submitted 16 July, 2025; originally announced July 2025.

  13. arXiv:2505.14226  [pdf, ps, other

    cs.CL cs.AI

    Phonetic Perturbations Reveal Tokenizer-Rooted Safety Gaps in LLMs

    Authors: Darpan Aswal, Siddharth D Jaiswal

    Abstract: Safety-aligned LLMs remain vulnerable to digital phenomena like textese that introduce non-canonical perturbations to words but preserve the phonetics. We introduce CMP-RT (code-mixed phonetic perturbations for red-teaming), a novel diagnostic probe that pinpoints tokenization as the root cause of this vulnerability. A mechanistic analysis reveals that phonetic perturbations fragment safety-critic… ▽ More

    Submitted 7 April, 2026; v1 submitted 20 May, 2025; originally announced May 2025.

  14. arXiv:2505.05043  [pdf, ps, other

    cs.CV

    xTrace: A Facial Expressive Behaviour Analysis Tool for Continuous Affect Recognition

    Authors: Mani Kumar Tellamekala, Shashank Jaiswal, Thomas Smith, Timur Alamev, Gary McKeown, Anthony Brown, Michel Valstar

    Abstract: Recognising expressive behaviours in face videos is a long-standing challenge in Affective Computing. Despite significant advancements in recent years, it still remains a challenge to build a robust and reliable system for naturalistic and in-the-wild facial expressive behaviour analysis in real time. This paper addresses two key challenges in building such a system: (1). The paucity of large-scal… ▽ More

    Submitted 12 October, 2025; v1 submitted 8 May, 2025; originally announced May 2025.

  15. arXiv:2503.19501   

    cs.CV cs.AI

    Pose-Based Fall Detection System: Efficient Monitoring on Standard CPUs

    Authors: Vinayak Mali, Saurabh Jaiswal

    Abstract: Falls among elderly residents in assisted living homes pose significant health risks, often leading to injuries and a decreased quality of life. Current fall detection solutions typically rely on sensor-based systems that require dedicated hardware, or on video-based models that demand high computational resources and GPUs for real-time processing. In contrast, this paper presents a robust fall de… ▽ More

    Submitted 28 June, 2026; v1 submitted 25 March, 2025; originally announced March 2025.

    Comments: Misleading Results

  16. arXiv:2503.14138  [pdf, other

    cs.CV cs.AI cs.CY

    Exploring Disparity-Accuracy Trade-offs in Face Recognition Systems: The Role of Datasets, Architectures, and Loss Functions

    Authors: Siddharth D Jaiswal, Sagnik Basu, Sandipan Sikdar, Animesh Mukherjee

    Abstract: Automated Face Recognition Systems (FRSs), developed using deep learning models, are deployed worldwide for identity verification and facial attribute analysis. The performance of these models is determined by a complex interdependence among the model architecture, optimization/loss function and datasets. Although FRSs have surpassed human-level accuracy, they continue to be disparate against cert… ▽ More

    Submitted 18 March, 2025; originally announced March 2025.

    Comments: This work has been accepted for publication at AAAI ICWSM 2025

  17. SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling

    Authors: Shashwat Jaiswal, Kunal Jain, Yogesh Simmhan, Anjaly Parayil, Ankur Mallick, Rujia Wang, Renee St. Amant, Chetan Bansal, Victor Rühle, Anoop Kulkarni, Steve Kofsky, Saravan Rajmohan

    Abstract: Global cloud service providers handle inference workloads for Large Language Models (LLMs) that span latency-sensitive (e.g., chatbots) and insensitive (e.g., report writing) tasks, resulting in diverse and often conflicting Service Level Agreement (SLA) requirements. Managing such mixed workloads is challenging due to the complexity of the inference serving stack, which encompasses multiple model… ▽ More

    Submitted 12 November, 2025; v1 submitted 20 February, 2025; originally announced February 2025.

    Comments: 25 pages, 16 figures, 2 tables. The workload traces, our simulator harness and the SageServe scheduler are available at https://github.com/shashwatj07/SageServe

    Journal ref: Proceedings of the ACM on Measurement and Analysis of Computing Systems, Vol. 9, No. 3, Article 61. December 2025

  18. arXiv:2412.04065  [pdf, other

    cs.LG

    Space to Policy: Scalable Brick Kiln Detection and Automatic Compliance Monitoring with Geospatial Data

    Authors: Zeel B Patel, Rishabh Mondal, Shataxi Dubey, Suraj Jaiswal, Sarath Guttikunda, Nipun Batra

    Abstract: Air pollution kills 7 million people annually. The brick kiln sector significantly contributes to economic development but also accounts for 8-14\% of air pollution in India. Policymakers have implemented compliance measures to regulate brick kilns. Emission inventories are critical for air quality modeling and source apportionment studies. However, the largely unorganized nature of the brick kiln… ▽ More

    Submitted 10 April, 2025; v1 submitted 5 December, 2024; originally announced December 2024.

  19. arXiv:2411.13754  [pdf, other

    cs.LG cs.AI cs.CV

    Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios

    Authors: Shantanu Jaiswal, Debaditya Roy, Basura Fernando, Cheston Tan

    Abstract: Complex visual reasoning and question answering (VQA) is a challenging task that requires compositional multi-step processing and higher-level reasoning capabilities beyond the immediate recognition and localization of objects and events. Here, we introduce a fully neural Iterative and Parallel Reasoning Mechanism (IPRM) that combines two distinct forms of computation -- iterative and parallel --… ▽ More

    Submitted 20 November, 2024; originally announced November 2024.

    Comments: NeurIPS 2024 camera ready; source code to be released at: https://github.com/shantanuj/IPRM_Iterative_and_Parallel_Reasoning_Mechanism

  20. arXiv:2410.16712  [pdf, other

    cs.SD cs.CL eess.AS

    DENOASR: Debiasing ASRs through Selective Denoising

    Authors: Anand Kumar Rai, Siddharth D Jaiswal, Shubham Prakash, Bendi Pragnya Sree, Animesh Mukherjee

    Abstract: Automatic Speech Recognition (ASR) systems have been examined and shown to exhibit biases toward particular groups of individuals, influenced by factors such as demographic traits, accents, and speech styles. Noise can disproportionately impact speakers with certain accents, dialects, or speaking styles, leading to biased error rates. In this work, we introduce a novel framework DENOASR, which is… ▽ More

    Submitted 22 October, 2024; originally announced October 2024.

    Comments: Paper accepted at IEEE ICKG 2024

  21. arXiv:2409.00106  [pdf, other

    cs.CL cs.AI cs.CV cs.LG

    Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis

    Authors: Aishik Nagar, Shantanu Jaiswal, Cheston Tan

    Abstract: Vision-language models (VLMs) have shown impressive zero- and few-shot performance on real-world visual question answering (VQA) benchmarks, alluding to their capabilities as visual reasoning engines. However, the benchmarks being used conflate "pure" visual reasoning with world knowledge, and also have questions that involve a limited number of reasoning steps. Thus, it remains unclear whether a… ▽ More

    Submitted 27 August, 2024; originally announced September 2024.

    Comments: 21 pages

  22. arXiv:2407.15810  [pdf, other

    cs.CV cs.CY

    Breaking the Global North Stereotype: A Global South-centric Benchmark Dataset for Auditing and Mitigating Biases in Facial Recognition Systems

    Authors: Siddharth D Jaiswal, Animesh Ganai, Abhisek Dash, Saptarshi Ghosh, Animesh Mukherjee

    Abstract: Facial Recognition Systems (FRSs) are being developed and deployed globally at unprecedented rates. Most platforms are designed in a limited set of countries but deployed in worldwide, without adequate checkpoints. This is especially problematic for Global South countries which lack strong legislation to safeguard persons facing disparate performance of these systems. A combination of unavailabili… ▽ More

    Submitted 26 July, 2024; v1 submitted 22 July, 2024; originally announced July 2024.

    Comments: This work has been accepted for publication at AAAI/ACM AIES 2024

  23. arXiv:2407.14650  [pdf, other

    cs.CY cs.HC cs.IR

    Auditing the Grid-Based Placement of Private Label Products on E-commerce Search Result Pages

    Authors: Siddharth D Jaiswal, Abhisek Dash, Nitika Shroff, Yashwanth Babu Vunnam, Saptarshi Ghosh, Animesh Mukherjee

    Abstract: E-commerce platforms support the needs and livelihoods of their two most important stakeholders -- customers and producers/sellers. Multiple algorithmic systems, like ``search'' systems mediate the interactions between these stakeholders by connecting customers to producers with relevant items. Search results include (i) private label (PL) products that are manufactured/sold by the platform itself… ▽ More

    Submitted 19 July, 2024; originally announced July 2024.

  24. arXiv:2406.15635  [pdf, other

    cs.LG cs.CR cs.CV

    DataFreeShield: Defending Adversarial Attacks without Training Data

    Authors: Hyeyoon Lee, Kanghyun Choi, Dain Kwon, Sunjong Park, Mayoore Selvarasa Jaiswal, Noseong Park, Jonghyun Choi, Jinho Lee

    Abstract: Recent advances in adversarial robustness rely on an abundant set of training data, where using external or additional datasets has become a common setting. However, in real life, the training data is often kept private for security and privacy issues, while only the pretrained weight is available to the public. In such scenarios, existing methods that assume accessibility to the original data bec… ▽ More

    Submitted 21 June, 2024; originally announced June 2024.

    Comments: ICML 2024

  25. arXiv:2406.10723   

    cs.CV

    Eye in the Sky: Detection and Compliance Monitoring of Brick Kilns using Satellite Imagery

    Authors: Rishabh Mondal, Shataxi Dubey, Vannsh Jani, Shrimay Shah, Suraj Jaiswal, Zeel B Patel, Nipun Batra

    Abstract: Air pollution kills 7 million people annually. The brick manufacturing industry accounts for 8%-14% of air pollution in the densely populated Indo-Gangetic plain. Due to the unorganized nature of brick kilns, policy violation detection, such as proximity to human habitats, remains challenging. While previous studies have utilized computer vision-based machine learning methods for brick kiln detect… ▽ More

    Submitted 16 September, 2024; v1 submitted 15 June, 2024; originally announced June 2024.

    Comments: The PI was not in favor of making the work public on arXiv as the content is not yet ready to be released

  26. arXiv:2402.13771  [pdf, other

    cs.CV cs.AI cs.CY cs.HC

    Mask-up: Investigating Biases in Face Re-identification for Masked Faces

    Authors: Siddharth D Jaiswal, Ankit Kr. Verma, Animesh Mukherjee

    Abstract: AI based Face Recognition Systems (FRSs) are now widely distributed and deployed as MLaaS solutions all over the world, moreso since the COVID-19 pandemic for tasks ranging from validating individuals' faces while buying SIM cards to surveillance of citizens. Extensive biases have been reported against marginalized groups in these systems and have led to highly discriminatory outcomes. The post-pa… ▽ More

    Submitted 21 February, 2024; originally announced February 2024.

    Comments: This work has been submitted to the IEEE for possible publication

  27. arXiv:2310.06061  [pdf, other

    cs.CY cs.CL

    Auditing Gender Analyzers on Text Data

    Authors: Siddharth D Jaiswal, Ankit Kumar Verma, Animesh Mukherjee

    Abstract: AI models have become extremely popular and accessible to the general public. However, they are continuously under the scanner due to their demonstrable biases toward various sections of the society like people of color and non-binary people. In this study, we audit three existing gender analyzers -- uClassify, Readable and HackerFactor, for biases against non-binary individuals. These tools are d… ▽ More

    Submitted 9 October, 2023; originally announced October 2023.

    Comments: This work has been accepted at IEEE/ACM ASONAM 2023. Please cite the version appearing in the ASONAM proceedings

  28. arXiv:2309.14046  [pdf, other

    cs.IR cs.LG

    Diversify and Conquer: Bandits and Diversity for an Enhanced E-commerce Homepage Experience

    Authors: Sangeet Jaiswal, Korah T Malayil, Saif Jawaid, Sreekanth Vempati

    Abstract: In the realm of e-commerce, popular platforms utilize widgets to recommend advertisements and products to their users. However, the prevalence of mobile device usage on these platforms introduces a unique challenge due to the limited screen real estate available. Consequently, the positioning of relevant widgets becomes pivotal in capturing and maintaining customer engagement. Given the restricted… ▽ More

    Submitted 25 September, 2023; originally announced September 2023.

    Comments: Accepted in Proceedings of Fashionxrecys Workshop, 17th ACM Conference on Recommender Systems, 2023

  29. arXiv:2308.05390  [pdf, other

    cs.CV cs.IR cs.LG

    Product Review Image Ranking for Fashion E-commerce

    Authors: Sangeet Jaiswal, Dhruv Patel, Sreekanth Vempati, Konduru Saiswaroop

    Abstract: In a fashion e-commerce platform where customers can't physically examine the products on their own, being able to see other customers' text and image reviews of the product is critical while making purchase decisions. Given the high reliance on these reviews, over the years we have observed customers proactively sharing their reviews. With an increase in the coverage of User Generated Content (UG… ▽ More

    Submitted 10 August, 2023; originally announced August 2023.

    Comments: Accepted in Proceedings of ACM SIGIR Workshop on eCommerce (SIGIR eCom'22)

  30. arXiv:2307.10587  [pdf, other

    cs.CL cs.HC

    A Deep Dive into the Disparity of Word Error Rates Across Thousands of NPTEL MOOC Videos

    Authors: Anand Kumar Rai, Siddharth D Jaiswal, Animesh Mukherjee

    Abstract: Automatic speech recognition (ASR) systems are designed to transcribe spoken language into written text and find utility in a variety of applications including voice assistants and transcription services. However, it has been observed that state-of-the-art ASR systems which deliver impressive benchmark results, struggle with speakers of certain regions or demographics due to variation in their spe… ▽ More

    Submitted 20 July, 2023; originally announced July 2023.

  31. arXiv:2306.08889  [pdf, other

    cs.CV cs.AI

    Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion

    Authors: Ishaan Singh Rawal, Alexander Matyasko, Shantanu Jaiswal, Basura Fernando, Cheston Tan

    Abstract: While VideoQA Transformer models demonstrate competitive performance on standard benchmarks, the reasons behind their success are not fully understood. Do these models capture the rich multimodal structures and dynamics from video and text jointly? Or are they achieving high scores by exploiting biases and spurious features? Hence, to provide insights, we design $\textit{QUAG}$ (QUadrant AveraGe),… ▽ More

    Submitted 7 June, 2024; v1 submitted 15 June, 2023; originally announced June 2023.

    Comments: Accepted at ICML 2024

  32. arXiv:2305.02660  [pdf, other

    eess.IV cs.CV

    Expanding Synthetic Real-World Degradations for Blind Video Super Resolution

    Authors: Mehran Jeelani, Sadbhawna, Noshaba Cheema, Klaus Illgner-Fehns, Philipp Slusallek, Sunil Jaiswal

    Abstract: Video super-resolution (VSR) techniques, especially deep-learning-based algorithms, have drastically improved over the last few years and shown impressive performance on synthetic data. However, their performance on real-world video data suffers because of the complexity of real-world degradations and misaligned video frames. Since obtaining a synthetic dataset consisting of low-resolution (LR) an… ▽ More

    Submitted 4 May, 2023; originally announced May 2023.

  33. arXiv:2305.02645  [pdf, other

    cs.CV

    Edge-aware Consistent Stereo Video Depth Estimation

    Authors: Elena Kosheleva, Sunil Jaiswal, Faranak Shamsafar, Noshaba Cheema, Klaus Illgner-Fehns, Philipp Slusallek

    Abstract: Video depth estimation is crucial in various applications, such as scene reconstruction and augmented reality. In contrast to the naive method of estimating depths from images, a more sophisticated approach uses temporal information, thereby eliminating flickering and geometrical inconsistencies. We propose a consistent method for dense video depth estimation; however, unlike the existing monocula… ▽ More

    Submitted 4 May, 2023; originally announced May 2023.

  34. arXiv:2305.01732  [pdf, other

    cs.CV

    High-Resolution Synthetic RGB-D Datasets for Monocular Depth Estimation

    Authors: Aakash Rajpal, Noshaba Cheema, Klaus Illgner-Fehns, Philipp Slusallek, Sunil Jaiswal

    Abstract: Accurate depth maps are essential in various applications, such as autonomous driving, scene reconstruction, point-cloud creation, etc. However, monocular-depth estimation (MDE) algorithms often fail to provide enough texture & sharpness, and also are inconsistent for homogeneous scenes. These algorithms mostly use CNN or vision transformer-based architectures requiring large datasets for supervis… ▽ More

    Submitted 2 May, 2023; originally announced May 2023.

  35. arXiv:2304.08111  [pdf, other

    cs.CV

    Leveraging Multi-view Data for Improved Detection Performance: An Industrial Use Case

    Authors: Faranak Shamsafar, Sunil Jaiswal, Benjamin Kelkel, Kireeti Bodduna, Klaus Illgner-Fehns

    Abstract: Printed circuit boards (PCBs) are essential components of electronic devices, and ensuring their quality is crucial in their production. However, the vast variety of components and PCBs manufactured by different companies makes it challenging to adapt to production lines with speed demands. To address this challenge, we present a multi-view object detection framework that offers a fast and precise… ▽ More

    Submitted 17 April, 2023; originally announced April 2023.

  36. mlpack 4: a fast, header-only C++ machine learning library

    Authors: Ryan R. Curtin, Marcus Edel, Omar Shrit, Shubham Agrawal, Suryoday Basak, James J. Balamuta, Ryan Birmingham, Kartik Dutt, Dirk Eddelbuettel, Rishabh Garg, Shikhar Jaiswal, Aakash Kaushik, Sangyeon Kim, Anjishnu Mukherjee, Nanubala Gnana Sai, Nippun Sharma, Yashwant Singh Parihar, Roshan Swain, Conrad Sanderson

    Abstract: For over 15 years, the mlpack machine learning library has served as a "swiss army knife" for C++-based machine learning. Its efficient implementations of common and cutting-edge machine learning algorithms have been used in a wide variety of scientific and industrial applications. This paper overviews mlpack 4, a significant upgrade over its predecessor. The library has been significantly refacto… ▽ More

    Submitted 1 February, 2023; originally announced February 2023.

    Journal ref: Journal of Open Source Software, Vol. 8, No. 82, 2023

  37. arXiv:2212.03470  [pdf, other

    eess.AS cs.SD

    Improving trajectory localization accuracy via direction-of-arrival derivative estimation

    Authors: Ruchi Pandey, Shreyas Jaiswal, Huy Phan, Santosh Nannuru

    Abstract: Sound source localization is crucial in acoustic sensing and monitoring-related applications. In this paper, we do a comprehensive analysis of improvement in sound source localization by combining the direction of arrivals (DOAs) with their derivatives which quantify the changes in the positions of sources over time. This study uses the SALSA-Lite feature with a convolutional recurrent neural netw… ▽ More

    Submitted 10 December, 2022; v1 submitted 7 December, 2022; originally announced December 2022.

  38. arXiv:2211.16822  [pdf, other

    cs.CL

    A Probabilistic-Logic based Commonsense Representation Framework for Modelling Inferences with Multiple Antecedents and Varying Likelihoods

    Authors: Shantanu Jaiswal, Liu Yan, Dongkyu Choi, Kenneth Kwok

    Abstract: Commonsense knowledge-graphs (CKGs) are important resources towards building machines that can 'reason' on text or environmental inputs and make inferences beyond perception. While current CKGs encode world knowledge for a large number of concepts and have been effectively utilized for incorporating commonsense in neural models, they primarily encode declarative or single-condition inferential kno… ▽ More

    Submitted 15 December, 2022; v1 submitted 30 November, 2022; originally announced November 2022.

  39. arXiv:2211.12850  [pdf, other

    cs.LG cs.IR

    OOD-DiskANN: Efficient and Scalable Graph ANNS for Out-of-Distribution Queries

    Authors: Shikhar Jaiswal, Ravishankar Krishnaswamy, Ankit Garg, Harsha Vardhan Simhadri, Sheshansh Agrawal

    Abstract: State-of-the-art algorithms for Approximate Nearest Neighbor Search (ANNS) such as DiskANN, FAISS-IVF, and HNSW build data dependent indices that offer substantially better accuracy and search efficiency over data-agnostic indices by overfitting to the index data distribution. When the query data is drawn from a different distribution - e.g., when index represents image embeddings and query repres… ▽ More

    Submitted 30 November, 2022; v1 submitted 22 October, 2022; originally announced November 2022.

  40. arXiv:2211.01603  [pdf, other

    q-bio.GN cs.LG eess.SP

    Using Signal Processing in Tandem With Adapted Mixture Models for Classifying Genomic Signals

    Authors: Saish Jaiswal, Shreya Nema, Hema A Murthy, Manikandan Narayanan

    Abstract: Genomic signal processing has been used successfully in bioinformatics to analyze biomolecular sequences and gain varied insights into DNA structure, gene organization, protein binding, sequence evolution, etc. But challenges remain in finding the appropriate spectral representation of a biomolecular sequence, especially when multiple variable-length sequences need to be handled consistently. In t… ▽ More

    Submitted 3 November, 2022; originally announced November 2022.

  41. arXiv:2210.16556  [pdf, other

    cs.LG cs.PL

    MinUn: Accurate ML Inference on Microcontrollers

    Authors: Shikhar Jaiswal, Rahul Kiran Kranti Goli, Aayan Kumar, Vivek Seshadri, Rahul Sharma

    Abstract: Running machine learning inference on tiny devices, known as TinyML, is an emerging research area. This task requires generating inference code that uses memory frugally, a task that standard ML frameworks are ill-suited for. A deployment framework for TinyML must be a) parametric in the number representation to take advantage of the emerging representations like posits, b) carefully assign high-p… ▽ More

    Submitted 30 November, 2022; v1 submitted 29 October, 2022; originally announced October 2022.

  42. Marching with the Pink Parade: Evaluating Visual Search Recommendations for Non-binary Clothing Items

    Authors: Siddharth D Jaiswal, Animesh Mukherjee

    Abstract: Fashion, a highly subjective topic is interpreted differently by all individuals. E-commerce platforms, despite these diverse requirements, tend to cater to the average buyer instead of focusing on edge cases like non-binary shoppers. This case study, through participant surveys, shows that visual search on e-commerce platforms like Amazon, Beagle.Vision and Lykdat, is particularly poor for non-bi… ▽ More

    Submitted 4 December, 2021; originally announced December 2021.

    Comments: This work has been accepted for publication at ACM CHI 2022 (Case Studies) as an Extended Abstract

  43. arXiv:2111.13470  [pdf, other

    cs.CV

    TDAM: Top-Down Attention Module for Contextually Guided Feature Selection in CNNs

    Authors: Shantanu Jaiswal, Basura Fernando, Cheston Tan

    Abstract: Attention modules for Convolutional Neural Networks (CNNs) are an effective method to enhance performance on multiple computer-vision tasks. While existing methods appropriately model channel-, spatial- and self-attention, they primarily operate in a feedforward bottom-up manner. Consequently, the attention mechanism strongly depends on the local information of a single input feature map and does… ▽ More

    Submitted 21 October, 2022; v1 submitted 26 November, 2021; originally announced November 2021.

    Comments: ECCV 2022 Camera Ready

  44. arXiv:2111.09137  [pdf, other

    cs.CV

    Two-Face: Adversarial Audit of Commercial Face Recognition Systems

    Authors: Siddharth D Jaiswal, Karthikeya Duggirala, Abhisek Dash, Animesh Mukherjee

    Abstract: Computer vision applications like automated face detection are used for a variety of purposes ranging from unlocking smart devices to tracking potential persons of interest for surveillance. Audits of these applications have revealed that they tend to be biased against minority groups which result in unfair and concerning societal and political outcomes. Despite multiple studies over time, these b… ▽ More

    Submitted 17 November, 2021; originally announced November 2021.

    Comments: This work has been accepted for publication at ICWSM 2022

  45. arXiv:2110.13570  [pdf, other

    cs.CV

    Learning Graph Representation of Person-specific Cognitive Processes from Audio-visual Behaviours for Automatic Personality Recognition

    Authors: Siyang Song, Zilong Shao, Shashank Jaiswal, Linlin Shen, Michel Valstar, Hatice Gunes

    Abstract: This approach builds on two following findings in cognitive science: (i) human cognition partially determines expressed behaviour and is directly linked to true personality traits; and (ii) in dyadic interactions individuals' nonverbal behaviours are influenced by their conversational partner behaviours. In this context, we hypothesise that during a dyadic interaction, a target subject's facial re… ▽ More

    Submitted 27 October, 2021; v1 submitted 26 October, 2021; originally announced October 2021.

    Comments: Submitted to IJCV

    MSC Class: 68T40 ACM Class: I.2.1

  46. arXiv:2106.01400  [pdf, other

    eess.AS cs.LG cs.SD

    Dual Script E2E framework for Multilingual and Code-Switching ASR

    Authors: Mari Ganesh Kumar, Jom Kuriakose, Anand Thyagachandran, Arun Kumar A, Ashish Seth, Lodagala Durga Prasad, Saish Jaiswal, Anusha Prakash, Hema Murthy

    Abstract: India is home to multiple languages, and training automatic speech recognition (ASR) systems for languages is challenging. Over time, each language has adopted words from other languages, such as English, leading to code-mixing. Most Indian languages also have their own unique scripts, which poses a major limitation in training multilingual and code-switching ASR systems. Inspired by results in… ▽ More

    Submitted 2 June, 2021; originally announced June 2021.

    Comments: Accepted for publication at Interspeech 2021

  47. arXiv:2105.10427  [pdf, other

    cs.AR

    Prefetcher-based DRAM Architecture

    Authors: Saurabh Jaiswal, Shailendra Kumar Gupta, Soumya Soubhagya Dandapat

    Abstract: Advancement in Processor technology has made it easy to handle data-intensive workloads, but limiting main memory advances has created performance bottlenecks. In DRAM, there have been improvements in DRAM access latency as well as reduction in cost-per-bit with the increase in cell density. But still DRAM data transfer rate lags behind the processing speed of the current generation processors. As… ▽ More

    Submitted 21 May, 2021; originally announced May 2021.

  48. arXiv:2011.10608  [pdf, other

    cs.CV

    Large Scale Neural Architecture Search with Polyharmonic Splines

    Authors: Ulrich Finkler, Michele Merler, Rameswar Panda, Mayoore S. Jaiswal, Hui Wu, Kandan Ramakrishnan, Chun-Fu Chen, Minsik Cho, David Kung, Rogerio Feris, Bishwaranjan Bhattacharjee

    Abstract: Neural Architecture Search (NAS) is a powerful tool to automatically design deep neural networks for many tasks, including image classification. Due to the significant computational burden of the search phase, most NAS methods have focused so far on small, balanced datasets. All attempts at conducting NAS at large scale have employed small proxy sets, and then transferred the learned architectures… ▽ More

    Submitted 20 November, 2020; originally announced November 2020.

  49. CNN-based driving of block partitioning for intra slices encoding

    Authors: Franck Galpin, Fabien Racapé, Sunil Jaiswal, Philippe Bordes, Fabrice Le Léannec, Edouard François

    Abstract: This paper provides a technical overview of a deep-learning-based encoder method aiming at optimizing next generation hybrid video encoders for driving the block partitioning in intra slices. An encoding approach based on Convolutional Neural Networks is explored to partly substitute classical heuristics-based encoder speed-ups by a systematic and automatic process. The solution allows controlling… ▽ More

    Submitted 12 November, 2020; originally announced November 2020.

    Comments: 10 pages

    Journal ref: 2019 Data Compression Conference (DCC)

  50. arXiv:2007.10546  [pdf, ps, other

    cs.CY cs.AI cs.LG

    Ideas for Improving the Field of Machine Learning: Summarizing Discussion from the NeurIPS 2019 Retrospectives Workshop

    Authors: Shagun Sodhani, Mayoore S. Jaiswal, Lauren Baker, Koustuv Sinha, Carl Shneider, Peter Henderson, Joel Lehman, Ryan Lowe

    Abstract: This report documents ideas for improving the field of machine learning, which arose from discussions at the ML Retrospectives workshop at NeurIPS 2019. The goal of the report is to disseminate these ideas more broadly, and in turn encourage continuing discussion about how the field could improve along these axes. We focus on topics that were most discussed at the workshop: incentives for encourag… ▽ More

    Submitted 20 July, 2020; originally announced July 2020.