Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–30 of 30 results for author: Patel, H L

Searching in archive cs. Search in all archives.
.
  1. arXiv:2605.24702  [pdf, ps, other

    cs.CV

    Do Image-Text Metrics Respect Semantic Invariances?

    Authors: Amit Agarwal, Hitesh Laxmichand Patel, Meizhu Liu, Jyotika Singh, Karan Dua, Hansa Meghwani, Matthew Rowe, Michael Avendi, Yassi Abbasi, Tao Sheng, Sujith Ravi, Dan Roth

    Abstract: Reference-free image-to-text evaluators are now standard for scoring image-caption alignment, yet it is unclear whether they respect semantic invariances. We present an invariance probe on five popular evaluators (CLIPScore, PAC-S, UMIC, FLEUR, and a deterministic LLM judge) under semantics-preserving perturbations along three axes -- spatial (flips, context-preserving repositioning, light rotatio… ▽ More

    Submitted 23 May, 2026; originally announced May 2026.

  2. arXiv:2605.07053  [pdf, ps, other

    cs.CL cs.AI

    GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations

    Authors: Jyotika Singh, Fang Tu, Aziza Mirsaidova, Amit Agarwal, Hitesh Laxmichand Patel, Sandip Ghoshal, Miguel Ballesteros, Karan Dua, Yassine Benajiba, Weiyi Sun, Tao Sheng, Graham Horwood, Sujith Ravi, Dan Roth

    Abstract: Benchmarks like GSM8K are popular measures of mathematical reasoning, but leaderboard gains can overstate true capability due to memorization of fixed test sets. Most robustness variants apply surface-level perturbations (paraphrases, renamings, number swaps, distractors) that largely preserve the underlying facts, and static releases can themselves become memorization targets over time. We introd… ▽ More

    Submitted 26 May, 2026; v1 submitted 7 May, 2026; originally announced May 2026.

  3. arXiv:2604.23323  [pdf, ps, other

    cs.CL cs.SD

    Robust Audio-Text Retrieval via Cross-Modal Attention and Hybrid Loss

    Authors: Meizhu Liu, Matthew Rowe, Amit Agarwal, Michael Avendi, Yassi Abbasi, Hitesh Laxmichand Patel, Paul Li, Kyu J. Han, Tao Sheng, Sujith Ravi, Dan Roth

    Abstract: Audio-text retrieval enables semantic alignment between audio content and natural language queries, supporting applications in multimedia search, accessibility, and surveillance. However, current state-of-the-art approaches struggle with long, noisy, and weakly labeled audio due to their reliance on contrastive learning and large-batch training. We propose a novel multimodal retrieval framework th… ▽ More

    Submitted 25 April, 2026; originally announced April 2026.

  4. arXiv:2604.17771  [pdf, ps, other

    cs.CL cs.AI cs.DB

    SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks

    Authors: Mohammadtaher Safarzadeh, Hitesh Laxmichand Patel, Afshin Orojlooyjadid, Graham Horwood, Dan Roth

    Abstract: Large language models (LLMs) have achieved strong performance on natural language to SQL (NL2SQL) benchmarks, yet their reported accuracy may be inflated by contamination from benchmark queries or structurally similar patterns seen during training. We introduce SPENCE (Syntactic Probing and Evaluation of NL2SQL Contamination Effects), a controlled syntactic probing framework for detecting and quan… ▽ More

    Submitted 19 April, 2026; originally announced April 2026.

    Comments: ACL 2026 Main Conference

  5. arXiv:2604.11490  [pdf, ps, other

    cs.AI cs.CL cs.CV

    Anthropogenic Regional Adaptation in Multimodal Vision-Language Model

    Authors: Samuel Cahyawijaya, Peerat Limkonchotiwat, Tack Hwa Wong, Hitesh Laxmichand Patel, Amit Agarwal, Manuel Antonio Rufino, Carlos Rafael Catalan, Muhammad Reza Qorib, Vicky Feliren, Holy Lovenia, Aye Hninn Khine, Frederikus Hudi, David Anugraha, Alham Fikri Aji, Romrawin Chumpu, Viet-Thanh Pham, Minghan Wang, Mohamed Fazli Imam, Ruochen Zhang, Joseph Marvin Imperial, Khumaisa Nur'aini, Do Xuan Long, Musa Izzanardi Wijanarko, Joel Ruben Antony Moniz, Patrick Amadeus Irawan , et al. (23 additional authors not shown)

    Abstract: While the field of vision-language (VL) has achieved remarkable success in integrating visual and textual information across multiple languages and domains, there is still no dedicated framework for assessing human-centric alignment in vision-language systems. We offer two contributions to address this gap. First, we introduce Anthropogenic Regional Adaptation: a novel paradigm that aims to optimi… ▽ More

    Submitted 16 April, 2026; v1 submitted 13 April, 2026; originally announced April 2026.

  6. arXiv:2604.00018  [pdf, ps, other

    cs.CL cs.AI

    Think Twice Before You Write -- an Entropy-based Decoding Strategy to Enhance LLM Reasoning

    Authors: Jiashu He, Meizhu Liu, Olaitan P Olaleye, Amit Agarwal, M. Avendi, Yassi Abbasi, Matthew Rowe, Hitesh Laxmichand Patel, Paul Li, Tao Sheng, Sujith Ravi, Dan Roth

    Abstract: Decoding strategies play a central role in shaping the reasoning ability of large language models (LLMs). Traditional methods such as greedy decoding and beam search often suffer from error propagation, while sampling-based approaches introduce randomness without adequate robustness. Self-consistency improves reliability by aggregating multiple rollouts, but incurs significant computational overhe… ▽ More

    Submitted 10 March, 2026; originally announced April 2026.

  7. arXiv:2602.06291  [pdf, ps, other

    cs.CL

    Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math

    Authors: Guijin Son, Donghun Yang, Hitesh Laxmichand Patel, Hyunwoo Ko, Amit Agarwal, Sunghee Ahn, Kyong-Ha Lee, Youngjae Yu

    Abstract: Recent progress in reasoning models suggests that generating plausible attempts for research-level mathematics may be within reach, but verification remains a bottleneck, consuming scarce expert time. We hypothesize that a meaningful solution should contain enough method-level information that, when applied to a neighborhood of related questions, it should yield better downstream performance than… ▽ More

    Submitted 5 February, 2026; originally announced February 2026.

    Comments: Preprint

  8. arXiv:2601.18026  [pdf, ps, other

    cs.CL

    CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data

    Authors: Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera-Gómez, Sara Hincapie-Monsalve, Thom Vaughan, Damian Stewart, Malte Ostendorff, Idris Abdulmumin, Vukosi Marivate, Shamsuddeen Hassan Muhammad, Atnafu Lambebo Tonja, Hend Al-Khalifa, Nadia Ghezaiel Hammouda, Verrah Otiende, Tack Hwa Wong, Jakhongir Saydaliev, Melika Nobakhtian, Muhammad Ravi Shulthan Habibi, Chalamalasetti Kranti, Carol Muchemi, Khang Nguyen, Faisal Muhammad Adam, Luis Frentzen Salim, Reem Alqifari , et al. (72 additional authors not shown)

    Abstract: Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heterogeneous web data often used to train multilingual language models. In this paper, we introduce CommonLID, a community-driven, human-annotated LID benchmark for the web domain, covering 109 languages. Many of the include… ▽ More

    Submitted 8 June, 2026; v1 submitted 25 January, 2026; originally announced January 2026.

    Comments: 18 pages, 8 tables, 5 figures

  9. arXiv:2601.05461  [pdf, ps, other

    cs.IR

    RECOR: Reasoning-focused Multi-turn Conversational Retrieval Benchmark

    Authors: Mohammed Ali, Abdelrahman Abdallah, Amit Agarwal, Hitesh Laxmichand Patel, Adam Jatowt

    Abstract: Existing benchmarks treat multi-turn conversation and reasoning-intensive retrieval separately, yet real-world information seeking requires both. To bridge this gap, we present a benchmark for reasoning-based conversational information retrieval comprising 707 conversations (2,971 turns) across eleven domains. To ensure quality, our Decomposition-and-Verification framework transforms complex queri… ▽ More

    Submitted 8 January, 2026; originally announced January 2026.

  10. arXiv:2601.04388  [pdf, ps, other

    cs.AI

    LLM-Guided Lifecycle-Aware Clustering of Multi-Turn Customer Support Conversations

    Authors: Priyaranjan Pattnayak, Sanchari Chowdhuri, Amit Agarwal, Hitesh Laxmichand Patel

    Abstract: Clustering customer chat data is vital for cloud providers handling multi service queries. Traditional methods struggle with overlapping concerns and create broad, static clusters that degrade over time. Reclustering disrupts continuity, making issue tracking difficult. We propose an adaptive system that segments multi turn chats into service specific concerns and incrementally refines clusters as… ▽ More

    Submitted 7 January, 2026; originally announced January 2026.

    Comments: Accepted in AACL 2025 Main Conference

  11. arXiv:2511.22787  [pdf, ps, other

    cs.CV

    World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language Models

    Authors: Eunsu Kim, Junyeong Park, Na Min An, Junseong Kim, Hitesh Laxmichand Patel, Jiho Jin, Julia Kruk, Amit Agarwal, Srikant Panda, Fenal Ashokbhai Ilasariya, Hyunjung Shim, Alice Oh

    Abstract: In a globalized world, cultural elements from diverse origins frequently appear together within a single visual scene. We refer to these as culture mixing scenarios, yet how Large Vision-Language Models (LVLMs) perceive them remains underexplored. We investigate culture mixing as a critical challenge for LVLMs and examine how current models behave when cultural items from multiple regions appear t… ▽ More

    Submitted 10 December, 2025; v1 submitted 27 November, 2025; originally announced November 2025.

  12. arXiv:2510.24081  [pdf, ps, other

    cs.CL

    Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures

    Authors: Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah, Abdelrahman Eldesokey, Abeer Kashar, Abolade Daud, Abosede Grace Olanihun, Adamu Labaran Mohammed, Adeyemi Praise, Adhikarimayum Meerajita Sharma, Aditi Gupta, Adril Putra Merin, Adwoa Bremang, Afitab Iyigun, Afonso Simplício, Ahmed Essouaied, Aicha Chorana, Akhil Eppa, Akintunde Oladipo, Akriti Kuri, Akshay Ramesh, Aleksei Dorkin, Alfred Malengo Kondoro, Alham Fikri Aji, Ali Eren Çetintaş , et al. (355 additional authors not shown)

    Abstract: To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we present Global PIQA, a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world. The 141 language varieties in Global PIQA cov… ▽ More

    Submitted 29 May, 2026; v1 submitted 28 October, 2025; originally announced October 2025.

    Comments: Preprint

  13. arXiv:2510.04230  [pdf, ps, other

    cs.CL

    Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought

    Authors: Guijin Son, Donghun Yang, Hitesh Laxmichand Patel, Amit Agarwal, Hyunwoo Ko, Chanuk Lim, Srikant Panda, Minhyuk Kim, Nikunj Drolia, Dasol Choi, Kyong-Ha Lee, Youngjae Yu

    Abstract: Recent frontier models employ long chain-of-thought reasoning to explore solution spaces in context and achieve stonger performance. While many works study distillation to build smaller yet capable models, most focus on English and little is known about language-specific reasoning. To bridge this gap, we first introduct **Language-Mixed CoT**, a reasoning schema that switches between English and a… ▽ More

    Submitted 13 January, 2026; v1 submitted 5 October, 2025; originally announced October 2025.

    Comments: Work in Progress

  14. arXiv:2510.02133  [pdf, ps, other

    cs.AI cs.LG

    FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models

    Authors: Karan Dua, Hitesh Laxmichand Patel, Puneet Mittal, Ranjeet Gupta, Amit Agarwal, Praneet Pabolu, Srikant Panda, Hansa Meghwani, Graham Horwood, Fahad Shah

    Abstract: Developing document understanding models at enterprise scale requires large, diverse, and well-annotated datasets spanning a wide range of document types. However, collecting such data is prohibitively expensive due to privacy constraints, legal restrictions, and the sheer volume of manual annotation needed - costs that can scale into millions of dollars. We introduce FlexDoc, a scalable synthetic… ▽ More

    Submitted 2 October, 2025; originally announced October 2025.

    Comments: Accepted at EMNLP 2025

    ACM Class: I.2.7; I.2.10; I.4.8; I.4.9

  15. arXiv:2509.23879  [pdf, ps, other

    cs.CV cs.AI cs.CL cs.MM

    PCRI: Measuring Context Robustness in Multimodal Models for Enterprise Applications

    Authors: Hitesh Laxmichand Patel, Amit Agarwal, Srikant Panda, Hansa Meghwani, Karan Dua, Paul Li, Tao Sheng, Sujith Ravi, Dan Roth

    Abstract: The reliability of Multimodal Large Language Models (MLLMs) in real-world settings is often undermined by sensitivity to irrelevant or distracting visual context, an aspect not captured by existing evaluation metrics. We introduce the \textbf{Patch Context Robustness Index (PCRI)}, the first systematic and interpretable score for quantifying MLLM robustness to variations in visual context granular… ▽ More

    Submitted 28 September, 2025; originally announced September 2025.

    Comments: Accepted in EMNLP 2025

    MSC Class: 68T50; 68T45 ACM Class: I.2.7; I.2.10; I.4.8; I.4.10; I.4.0

  16. arXiv:2509.23673  [pdf, ps, other

    cs.CV cs.AI cs.CL cs.MM

    RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks

    Authors: Amit Agarwal, Hitesh Laxmichand Patel, Srikant Panda, Hansa Meghwani, Jyotika Singh, Karan Dua, Paul Li, Tao Sheng, Sujith Ravi, Dan Roth

    Abstract: Multimodal Large Language Models (MLLMs) have achieved impressive results on vision-language benchmarks, yet it remains unclear whether these benchmarks assess genuine global reasoning or allow success via localized visual cues. Existing evaluation methods do not explicitly measure this distinction, hindering effective dataset curation and real-world focused model development. We introduce Regio… ▽ More

    Submitted 28 September, 2025; originally announced September 2025.

    Comments: Accepted in EMNLP 2025

    MSC Class: 68T45; 68T50 ACM Class: I.2.7; I.2.10; I.4.7; I.4.8

  17. arXiv:2509.23659  [pdf, ps, other

    cs.CL cs.AI

    Aligning LLMs for Multilingual Consistency in Enterprise Applications

    Authors: Amit Agarwal, Hansa Meghwani, Hitesh Laxmichand Patel, Tao Sheng, Sujith Ravi, Dan Roth

    Abstract: Large language models (LLMs) remain unreliable for global enterprise applications due to substantial performance gaps between high-resource and mid/low-resource languages, driven by English-centric pretraining and internal reasoning biases. This inconsistency undermines customer experience and operational reliability in multilingual settings such as customer support, content moderation, and inform… ▽ More

    Submitted 25 October, 2025; v1 submitted 28 September, 2025; originally announced September 2025.

    Comments: Accepted at EMNLP 2025

    MSC Class: 68T05; 68T50; 68Q25 ACM Class: I.2.7; I.5.1; I.2.8

  18. arXiv:2509.22703  [pdf, ps, other

    cs.CL cs.AI cs.CY

    AccessEval: Benchmarking Disability Bias in Large Language Models

    Authors: Srikant Panda, Amit Agarwal, Hitesh Laxmichand Patel

    Abstract: Large Language Models (LLMs) are increasingly deployed across diverse domains but often exhibit disparities in how they handle real-life queries. To systematically investigate these effects within various disability contexts, we introduce \textbf{AccessEval (Accessibility Evaluation)}, a benchmark evaluating 21 closed- and open-source LLMs across 6 real-world domains and 9 disability types using p… ▽ More

    Submitted 22 September, 2025; originally announced September 2025.

  19. arXiv:2509.14270  [pdf, ps, other

    cs.CL cs.AI cs.LG cs.MM cs.SD eess.AS

    SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech Models

    Authors: Karan Dua, Puneet Mittal, Ranjeet Gupta, Hitesh Laxmichand Patel

    Abstract: High-quality Text-to-Speech (TTS) model training requires extensive and diverse text and speech data. It is challenging to procure such data from real sources due to issues of domain specificity, licensing, and scalability. Large language models (LLMs) can certainly generate textual data, but they create repetitive text with insufficient variation in the prompt during the generation process. Anoth… ▽ More

    Submitted 1 October, 2025; v1 submitted 15 September, 2025; originally announced September 2025.

    Comments: Accepted at ACL 2025

    ACM Class: I.2.7

    Journal ref: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track) - 2025

  20. arXiv:2508.15831  [pdf, ps, other

    cs.CL cs.AI cs.CY

    Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs

    Authors: Vishnu Hari, Kalpana Panda, Srikant Panda, Amit Agarwal, Hitesh Laxmichand Patel

    Abstract: Large Language Models (LLMs) routinely infer users demographic traits from phrasing alone, which can result in biased responses, even when no explicit demographic information is provided. The role of disability cues in shaping these inferences remains largely uncharted. Thus, we present the first systematic audit of disability-conditioned demographic bias across eight state-of-the-art instruction-… ▽ More

    Submitted 21 October, 2025; v1 submitted 18 August, 2025; originally announced August 2025.

    Comments: Accepted at ICCV 2025

    MSC Class: 68T50; 68T07; 68T05 ACM Class: I.2.7; I.2.6; K.4.2

  21. arXiv:2508.15830  [pdf, ps, other

    cs.CL cs.AI

    DAIQ: Auditing Demographic Attribute Inference from Question in LLMs

    Authors: Srikant Panda, Hitesh Laxmichand Patel, Shahad Al-Khalifa, Amit Agarwal, Hend Al-Khalifa, Sharefah Al-Ghamdi

    Abstract: Recent evaluations of Large language models (LLMs) audit social bias primarily through prompts that explicitly reference demographic attributes, overlooking whether models infer sensitive demographics from neutral questions. Such inference constitutes epistemic overreach and raises concerns for privacy. We introduce Demographic Attribute Inference from Questions (DAIQ), a diagnostic audit framewor… ▽ More

    Submitted 24 January, 2026; v1 submitted 18 August, 2025; originally announced August 2025.

    Comments: Preprint

  22. arXiv:2506.02097  [pdf, ps, other

    cs.AI

    Hybrid AI for Responsive Multi-Turn Online Conversations with Novel Dynamic Routing and Feedback Adaptation

    Authors: Priyaranjan Pattnayak, Amit Agarwal, Hansa Meghwani, Hitesh Laxmichand Patel, Srikant Panda

    Abstract: Retrieval-Augmented Generation (RAG) systems and large language model (LLM)-powered chatbots have significantly advanced conversational AI by combining generative capabilities with external knowledge retrieval. Despite their success, enterprise-scale deployments face critical challenges, including diverse user queries, high latency, hallucinations, and difficulty integrating frequently updated dom… ▽ More

    Submitted 25 June, 2025; v1 submitted 2 June, 2025; originally announced June 2025.

    Comments: Proceedings of the 4th International Workshop on Knowledge Augmented Methods for Natural Language Processing in NAACL 2025, pages 215 to 229, Albuquerque, New Mexico, USA. Association for Computational Linguistics

    Journal ref: Proceedings of the 4th International Workshop on Knowledge-Augmented Methods for Natural Language Processing (KnowledgeNLP 2025), pp. 215 to 229, Association for Computational Linguistics, Albuquerque, New Mexico, May 2025

  23. arXiv:2505.18366  [pdf, ps, other

    cs.IR cs.AI cs.CL cs.LG

    Hard Negative Mining for Domain-Specific Retrieval in Enterprise Systems

    Authors: Hansa Meghwani, Amit Agarwal, Priyaranjan Pattnayak, Hitesh Laxmichand Patel, Srikant Panda

    Abstract: Enterprise search systems often struggle to retrieve accurate, domain-specific information due to semantic mismatches and overlapping terminologies. These issues can degrade the performance of downstream applications such as knowledge management, customer support, and retrieval-augmented generation agents. To address this challenge, we propose a scalable hard-negative mining framework tailored spe… ▽ More

    Submitted 23 May, 2025; originally announced May 2025.

    Comments: Accepted to ACL 2025

    ACM Class: H.3.3; I.2.6; I.2.7

  24. arXiv:2505.17332  [pdf, ps, other

    cs.CL cs.AI cs.LG cs.MA

    SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use

    Authors: Hitesh Laxmichand Patel, Amit Agarwal, Arion Das, Bhargava Kumar, Srikant Panda, Priyaranjan Pattnayak, Taki Hasan Rafi, Tejaswini Kumar, Dong-Kyu Chae

    Abstract: Enterprise customers are increasingly adopting Large Language Models (LLMs) for critical communication tasks, such as drafting emails, crafting sales pitches, and composing casual messages. Deploying such models across different regions requires them to understand diverse cultural and linguistic contexts and generate safe and respectful responses. For enterprise applications, it is crucial to miti… ▽ More

    Submitted 22 May, 2025; originally announced May 2025.

    Comments: Published in the Proceedings of the 2025 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL 2025), Industry Track, pages 558-582

    ACM Class: I.2.7; I.2.6

  25. arXiv:2504.16977  [pdf, other

    cs.CL cs.AI

    Tokenization Matters: Improving Zero-Shot NER for Indic Languages

    Authors: Priyaranjan Pattnayak, Hitesh Laxmichand Patel, Amit Agarwal

    Abstract: Tokenization is a critical component of Natural Language Processing (NLP), especially for low resource languages, where subword segmentation influences vocabulary structure and downstream task accuracy. Although Byte Pair Encoding (BPE) is a standard tokenization method in multilingual language models, its suitability for Named Entity Recognition (NER) in low resource Indic languages remains under… ▽ More

    Submitted 23 April, 2025; originally announced April 2025.

  26. arXiv:2503.07920  [pdf, other

    cs.CV cs.AI cs.CL

    Crowdsource, Crawl, or Generate? Creating SEA-VL, a Multicultural Vision-Language Dataset for Southeast Asia

    Authors: Samuel Cahyawijaya, Holy Lovenia, Joel Ruben Antony Moniz, Tack Hwa Wong, Mohammad Rifqi Farhansyah, Thant Thiri Maung, Frederikus Hudi, David Anugraha, Muhammad Ravi Shulthan Habibi, Muhammad Reza Qorib, Amit Agarwal, Joseph Marvin Imperial, Hitesh Laxmichand Patel, Vicky Feliren, Bahrul Ilmi Nasution, Manuel Antonio Rufino, Genta Indra Winata, Rian Adam Rajagede, Carlos Rafael Catalan, Mohamed Fazli Imam, Priyaranjan Pattnayak, Salsabila Zahirah Pranida, Kevin Pratama, Yeshil Bangera, Adisai Na-Thalang , et al. (67 additional authors not shown)

    Abstract: Southeast Asia (SEA) is a region of extraordinary linguistic and cultural diversity, yet it remains significantly underrepresented in vision-language (VL) research. This often results in artificial intelligence (AI) models that fail to capture SEA cultural nuances. To fill this gap, we present SEA-VL, an open-source initiative dedicated to developing high-quality, culturally relevant data for SEA… ▽ More

    Submitted 18 March, 2025; v1 submitted 10 March, 2025; originally announced March 2025.

    Comments: [SEA-VL Dataset] https://huggingface.co/collections/SEACrowd/sea-vl-multicultural-vl-dataset-for-southeast-asia-67cf223d0c341d4ba2b236e7 [Appendix J] https://github.com/SEACrowd/seacrowd.github.io/blob/master/docs/SEA_VL_Appendix_J.pdf

  27. arXiv:2502.13108  [pdf, other

    cs.CL cs.AI cs.LG

    Clinical QA 2.0: Multi-Task Learning for Answer Extraction and Categorization

    Authors: Priyaranjan Pattnayak, Hitesh Laxmichand Patel, Amit Agarwal, Bhargava Kumar, Srikant Panda, Tejaswini Kumar

    Abstract: Clinical Question Answering (CQA) plays a crucial role in medical decision-making, enabling physicians to extract relevant information from Electronic Medical Records (EMRs). While transformer-based models such as BERT, BioBERT, and ClinicalBERT have demonstrated state-of-the-art performance in CQA, existing models lack the ability to categorize extracted answers, which is critical for structured… ▽ More

    Submitted 23 April, 2025; v1 submitted 18 February, 2025; originally announced February 2025.

  28. arXiv:2412.17759  [pdf, other

    cs.AI cs.CV cs.LG

    Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy

    Authors: Priyaranjan Pattnayak, Hitesh Laxmichand Patel, Bhargava Kumar, Amit Agarwal, Ishan Banerjee, Srikant Panda, Tejaswini Kumar

    Abstract: Multimodal learning, a rapidly evolving field in artificial intelligence, seeks to construct more versatile and robust systems by integrating and analyzing diverse types of data, including text, images, audio, and video. Inspired by the human ability to assimilate information through many senses, this method enables applications such as text-to-video conversion, visual question answering, and imag… ▽ More

    Submitted 23 December, 2024; originally announced December 2024.

  29. arXiv:2411.14962  [pdf, other

    cs.CL cs.AI cs.CR

    LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents

    Authors: Hitesh Laxmichand Patel, Amit Agarwal, Bhargava Kumar, Karan Gupta, Priyaranjan Pattnayak

    Abstract: Accurate barcode detection and decoding in Identity documents is crucial for applications like security, healthcare, and education, where reliable data extraction and verification are essential. However, building robust detection models is challenging due to the lack of diverse, realistic datasets an issue often tied to privacy concerns and the wide variety of document formats. Traditional tools l… ▽ More

    Submitted 23 December, 2024; v1 submitted 22 November, 2024; originally announced November 2024.

    Comments: 5 pages, 1 figures

  30. arXiv:2404.01897  [pdf, ps, other

    cs.NE cs.AI cs.LG

    Continuous Spiking Graph Neural Networks

    Authors: Nan Yin, Mengzhu Wan, Li Shen, Hitesh Laxmichand Patel, Baopu Li, Bin Gu, Huan Xiong

    Abstract: Continuous graph neural networks (CGNNs) have garnered significant attention due to their ability to generalize existing discrete graph neural networks (GNNs) by introducing continuous dynamics. They typically draw inspiration from diffusion-based methods to introduce a novel propagation scheme, which is analyzed using ordinary differential equations (ODE). However, the implementation of CGNNs req… ▽ More

    Submitted 12 July, 2025; v1 submitted 2 April, 2024; originally announced April 2024.