Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–6 of 6 results for author: Gohil, R

Searching in archive cs. Search in all archives.
.
  1. Bridging the Gap: Converting Read Text to Conversational Dialogue

    Authors: Parshav Singla, Agnik Banerjee, Aaditya Arora, Shruti Aggarwal, Anil Kumar Verma, Vikram C M, Raj Prakash Gohil, Gopal Kumar Agarwal

    Abstract: In recent advancements within speech processing, converting read speech to conversational speech has gained significant attention. The primary challenge in this domain is maintaining naturalness and intelligibility while minimizing computational overhead for real-time applications. Traditional read speech often lacks the nuanced prosodic variation essential for natural conversational interactions,… ▽ More

    Submitted 18 May, 2026; originally announced May 2026.

    Comments: 11 pages, 4 figures. Published in ICICC 2025, Springer Lecture Notes in Networks and Systems

    Journal ref: Innovative Computing and Communications (ICICC 2025), Lecture Notes in Networks and Systems, Springer Nature, 2025, pp. 543-556

  2. arXiv:2602.09043  [pdf, ps, other

    eess.AS cs.LG cs.SD

    Windowed SummaryMixing: An Efficient Fine-Tuning of Self-Supervised Learning Models for Low-resource Speech Recognition

    Authors: Aditya Srinivas Menon, Kumud Tripathi, Raj Gohil, Pankaj Wasnik

    Abstract: Self-supervised learning (SSL) has advanced speech processing but suffers from quadratic complexity due to self-attention. To address this, SummaryMixing (SM) has been proposed as a linear-time alternative that summarizes entire utterances using mean pooling but lacks sufficient local context. In this work, we introduce Windowed SummaryMixing (WSM), which enhances SM by integrating local neighborh… ▽ More

    Submitted 4 February, 2026; originally announced February 2026.

    Comments: The paper has been accepted at ICASSP 2026, Barcelona, Spain

  3. arXiv:2511.14219  [pdf, ps, other

    cs.AI cs.SD

    Listen Like a Teacher: Mitigating Whisper Hallucinations using Adaptive Layer Attention and Knowledge Distillation

    Authors: Kumud Tripathi, Aditya Srinivas Menon, Aman Gaurav, Raj Prakash Gohil, Pankaj Wasnik

    Abstract: The Whisper model, an open-source automatic speech recognition system, is widely adopted for its strong performance across multilingual and zero-shot settings. However, it frequently suffers from hallucination errors, especially under noisy acoustic conditions. Previous works to reduce hallucinations in Whisper-style ASR systems have primarily focused on audio preprocessing or post-processing of t… ▽ More

    Submitted 18 November, 2025; originally announced November 2025.

    Comments: Accepted at AAAI 2026 - Main Technical Track

  4. arXiv:2506.02083  [pdf, ps, other

    cs.SD cs.AI cs.LG cs.MM

    LASPA: Language Agnostic Speaker Disentanglement with Prefix-Tuned Cross-Attention

    Authors: Aditya Srinivas Menon, Raj Prakash Gohil, Kumud Tripathi, Pankaj Wasnik

    Abstract: Speaker recognition models face challenges in multi-lingual settings due to the entanglement of linguistic information within speaker embeddings. The overlap between vocal traits such as accent, vocal anatomy, and a language's phonetic structure complicates separating linguistic and speaker information. Disentangling these components can significantly improve speaker recognition accuracy. To this… ▽ More

    Submitted 2 June, 2025; originally announced June 2025.

    Comments: Accepted at Interspeech 2025, Netherlands

  5. SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction

    Authors: Saurabh Agrawal, Raj Gohil, Gopal Kumar Agrawal, Vikram C M, Kushal Verma

    Abstract: Speech quality assessment is a critical process in selecting text-to-speech synthesis (TTS) or voice conversion models. Evaluation of voice synthesis can be done using objective metrics or subjective metrics. Although there are many objective metrics like the Perceptual Evaluation of Speech Quality (PESQ), Perceptual Objective Listening Quality Assessment (POLQA) or Short-Time Objective Intelligib… ▽ More

    Submitted 2 June, 2025; originally announced June 2025.

    Journal ref: 2024 International Conference on Signal Processing and Communications (SPCOM), 2024}, pages 1-5, 10631576

  6. arXiv:2504.09074  [pdf, other

    cs.AR

    A Case for Kolmogorov-Arnold Networks in Prefetching: Towards Low-Latency, Generalizable ML-Based Prefetchers

    Authors: Dhruv Kulkarni, Bharat Bhammar, Henil Thaker, Pranav Dhobi, R. P. Gohil, Sai Manoj Pudukotai Dinkarrao

    Abstract: The memory wall problem arises due to the disparity between fast processors and slower memory, causing significant delays in data access, even more so on edge devices. Data prefetching is a key strategy to address this, with traditional methods evolving to incorporate Machine Learning (ML) for improved accuracy. Modern prefetchers must balance high accuracy with low latency to further practicality… ▽ More

    Submitted 12 April, 2025; originally announced April 2025.