Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 51 results for author: Arora, V

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.20326  [pdf, ps, other

    eess.AS cs.LG

    $TCP_α$: Margin-Controlled Confidence estimation for reliable Music Information Retrieval

    Authors: Parampreet Singh, Anushka Singh, Sumit Kumar, Vipul Arora

    Abstract: Deep neural networks are often overconfident, assigning high confidence even to incorrect predictions. Consequently, users lack a reliable signal for deciding when a prediction can be trusted. Post-hoc confidence estimation addresses this by training a lightweight auxiliary head over a frozen classifier. Existing targets, however, suffer from inherent ambiguity: they assign overlapping confidence… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

  2. arXiv:2606.26824  [pdf, ps, other

    cs.SD eess.AS

    wav2tok 2.0: Scalable Audio Tokenization Maintaining Explicit Pairwise Token Alignment for Efficient Audio Retrieval

    Authors: Adhiraj Banerjee, Vipul Arora

    Abstract: Learning discrete speech representations that preserve similarity across variable-length utterances is central to query-by-example spoken term detection (QbE-STD). While wav2tok introduced CTC-based sequence alignment to enforce token consistency, its tightly coupled clustering and alignment training recipe limits scalability. We propose wav2tok 2.0, a scalable alignment-aware speech tokenizer bui… ▽ More

    Submitted 25 June, 2026; originally announced June 2026.

    Comments: Accepted at INTERSPEECH 2026

  3. arXiv:2606.23911  [pdf, ps, other

    cs.IR

    Scaling Dense Retrieval with LLM-Annotated Training Data: Structured Mining and Progressive Curriculum for E-Commerce Sponsored Search

    Authors: Md Omar Faruk Rokon, Shasvat Desai, Jhalak Nilesh Acharya, Isha Shah, Kumar Priyam, Brahanyaa Somasundaram, Vamsee Tangirala, Minuteresa Thomas, Vivek Arora, Vijay Manchi, Hong Yao, Kuang-chih Lee

    Abstract: How can we generate high-quality training data for dense retrieval models at production scale, without relying on click signals or manual annotation? This question is critical for e-commerce sponsored search, where click-based training suffers from position bias and tail-query sparsity, and manual labeling at the scale of hundreds of millions of query-item pairs is economically infeasible. Our wor… ▽ More

    Submitted 22 June, 2026; originally announced June 2026.

    Comments: Accepted at E-Commerce Workshop, SIGIR 2026

  4. arXiv:2605.06582  [pdf, ps, other

    cs.LG cs.CL cs.SD eess.AS

    PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization

    Authors: Adhiraj Banerjee, Vipul Arora

    Abstract: Modern learning systems represent perceptual signals with continuous vectors, but comparison, retrieval, memory, alignment, and reasoning are often naturally symbolic. In language, this interface is given by tokens; for speech and audio, it must be learned. Existing audio tokenizers use local quantization, clustering, or reconstruction, leaving sequence consistency, compactness, length control, te… ▽ More

    Submitted 21 June, 2026; v1 submitted 7 May, 2026; originally announced May 2026.

    Comments: 29 pages main content, 50 total pages, 6 Figures, pre-print, Under Review

  5. arXiv:2603.28061  [pdf, ps, other

    cs.DS

    Testing Sparse Functions over the Reals

    Authors: Vipul Arora, Arnab Bhattacharyya, Philips George John, Sayantan Sen

    Abstract: Over the last three decades, function testing has been extensively studied over Boolean, finite fields, and discrete settings. However, to encode the real-world applications more succinctly, function testing over the reals (where the domain and range, both are reals) is of prime importance. Recently, there have been some works in the direction of testing for algebraic representations of such funct… ▽ More

    Submitted 30 March, 2026; originally announced March 2026.

    Comments: 43 pages

  6. arXiv:2602.06917  [pdf, ps, other

    eess.AS cs.LG

    Automatic Detection and Analysis of Singing Mistakes for Music Pedagogy

    Authors: Sumit Kumar, Suraj Jaiswal, Parampreet Singh, Vipul Arora

    Abstract: The advancement of machine learning in audio analysis has opened new possibilities for technology-enhanced music education. This paper introduces a framework for automatic singing mistake detection in the context of music pedagogy, supported by a newly curated dataset. The dataset comprises synchronized teacher learner vocal recordings, with annotations marking different types of mistakes made by… ▽ More

    Submitted 6 February, 2026; originally announced February 2026.

    Comments: Under Review at Transactions of Audio Speech and Language Processing

  7. arXiv:2601.18766  [pdf, ps, other

    eess.AS cs.LG

    Learning to Discover: A Generalized Framework for Raga Identification without Forgetting

    Authors: Parampreet Singh, Somya Kumar, Chaitanya Shailendra Nitawe, Vipul Arora

    Abstract: Raga identification in Indian Art Music (IAM) remains challenging due to the presence of numerous rarely performed Ragas that are not represented in available training datasets. Traditional classification models struggle in this setting, as they assume a closed set of known categories and therefore fail to recognise or meaningfully group previously unseen Ragas. Recent works have tried categorizin… ▽ More

    Submitted 26 January, 2026; originally announced January 2026.

    Comments: Accepted at NCC 2026 conference

  8. arXiv:2601.03626  [pdf, ps, other

    eess.AS cs.LG

    Learning from Limited Labels: Transductive Graph Label Propagation for Indian Music Analysis

    Authors: Parampreet Singh, Akshay Raina, Sayeedul Islam Sheikh, Vipul Arora

    Abstract: Supervised machine learning frameworks rely on extensive labeled datasets for robust performance on real-world tasks. However, there is a lack of large annotated datasets in audio and music domains, as annotating such recordings is resource-intensive, laborious, and often require expert domain knowledge. In this work, we explore the use of label propagation (LP), a graph-based semi-supervised lear… ▽ More

    Submitted 7 January, 2026; originally announced January 2026.

    Comments: Published at Journal of Acoustical Society of India, 2025

    Journal ref: Journal of Acoustical Society of India, Vol. 52, No. 3, pp. 145-154, 2025

  9. arXiv:2512.02432  [pdf, ps, other

    cs.SD

    Continual Learning for Singing Voice Separation with Human in the Loop Adaptation

    Authors: Ankur Gupta, Anshul Rai, Archit Bansal, Vipul Arora

    Abstract: Deep learning-based works for singing voice separation have performed exceptionally well in the recent past. However, most of these works do not focus on allowing users to interact with the model to improve performance. This can be crucial when deploying the model in real-world scenarios where music tracks can vary from the original training data in both genre and instruments. In this paper, we pr… ▽ More

    Submitted 2 December, 2025; originally announced December 2025.

    Comments: Proceedings of the 26th International Symposium on Frontiers of Research in Speech and Music, 2021

  10. arXiv:2512.01199  [pdf, ps, other

    cs.LG q-bio.NC

    Know Thyself by Knowing Others: Learning Neuron Identity from Population Context

    Authors: Vinam Arora, Divyansha Lachi, Ian J. Knight, Mehdi Azabou, Blake Richards, Cole L. Hurwitz, Josh Siegle, Eva L. Dyer

    Abstract: Neurons process information in ways that depend on their cell type, connectivity, and the brain region in which they are embedded. However, inferring these factors from neural activity remains a significant challenge. To build general-purpose representations that allow for resolving information about a neuron's identity, we introduce NuCLR, a self-supervised framework that aims to learn representa… ▽ More

    Submitted 30 November, 2025; originally announced December 2025.

    Comments: Accepted at Neurips 2025

  11. arXiv:2511.04557  [pdf, ps, other

    cs.LG cs.AI

    Integrating Temporal and Structural Context in Graph Transformers for Relational Deep Learning

    Authors: Divyansha Lachi, Mahmoud Mohammadi, Joe Meyer, Vinam Arora, Tom Palczewski, Eva L. Dyer

    Abstract: In domains such as healthcare, finance, and e-commerce, the temporal dynamics of relational data emerge from complex interactions-such as those between patients and providers, or users and products across diverse categories. To be broadly useful, models operating on these data must integrate long-range spatial and temporal dependencies across diverse types of entities, while also supporting multip… ▽ More

    Submitted 6 November, 2025; originally announced November 2025.

  12. arXiv:2510.21330  [pdf, ps, other

    cs.LG hep-lat quant-ph

    SCORENF: Score-based Normalizing Flows for Sampling Unnormalized distributions

    Authors: Vikas Kanaujia, Vipul Arora

    Abstract: Unnormalized probability distributions are central to modeling complex physical systems across various scientific domains. Traditional sampling methods, such as Markov Chain Monte Carlo (MCMC), often suffer from slow convergence, critical slowing down, poor mode mixing, and high autocorrelation. In contrast, likelihood-based and adversarial machine learning models, though effective, are heavily da… ▽ More

    Submitted 24 October, 2025; originally announced October 2025.

    Comments: \c{opyright} 20XX IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

  13. arXiv:2510.01574  [pdf, ps, other

    cs.IR cs.AI cs.CL cs.LG

    Synthetic Prefixes to Mitigate Bias in Real-Time Neural Query Autocomplete

    Authors: Adithya Rajan, Xiaoyu Liu, Prateek Verma, Vibhu Arora

    Abstract: We introduce a data-centric approach for mitigating presentation bias in real-time neural query autocomplete systems through the use of synthetic prefixes. These prefixes are generated from complete user queries collected during regular search sessions where autocomplete was not active. This allows us to enrich the training data for learning to rank models with more diverse and less biased example… ▽ More

    Submitted 1 October, 2025; originally announced October 2025.

    Comments: Accepted to the Proceedings of the ACM SIGIR Asia Pacific Conference on Information Retrieval (SIGIR-AP 2025), December 7-10, 2025, Xi'an, China

  14. arXiv:2509.11717  [pdf, ps, other

    cs.SD cs.LG eess.AS

    CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents

    Authors: Adhiraj Banerjee, Vipul Arora

    Abstract: Text-guided sound separation enables flexible audio editing, assistive listening, and open-domain source extraction, but systems such as AudioSep remain too expensive for low-latency edge or codec-mediated deployment. Existing neural audio codec separators are efficient, yet largely restricted to fixed stems or closed taxonomies. We introduce CodecSep, a prompt-driven universal sound separation fr… ▽ More

    Submitted 21 June, 2026; v1 submitted 15 September, 2025; originally announced September 2025.

    Comments: main content- 27 pages, total - 53 pages, 12 figure, Accepted by Transactions on Machine Learning Research (TMLR), 2026

    Journal ref: Transactions on Machine Learning Research, 2026. ISSN 2835-8856

  15. arXiv:2505.04419  [pdf, other

    eess.AS cs.AI cs.LG cs.SD

    Recognizing Ornaments in Vocal Indian Art Music with Active Annotation

    Authors: Sumit Kumar, Parampreet Singh, Vipul Arora

    Abstract: Ornamentations, embellishments, or microtonal inflections are essential to melodic expression across many musical traditions, adding depth, nuance, and emotional impact to performances. Recognizing ornamentations in singing voices is key to MIR, with potential applications in music pedagogy, singer identification, genre classification, and controlled singing voice generation. However, the lack of… ▽ More

    Submitted 7 May, 2025; originally announced May 2025.

  16. arXiv:2501.12384  [pdf, other

    cs.CV cs.LG eess.IV

    CCESAR: Coastline Classification-Extraction From SAR Images Using CNN-U-Net Combination

    Authors: Vidhu Arora, Shreyan Gupta, Ananthakrishna Kudupu, Aditya Priyadarshi, Aswathi Mundayatt, Jaya Sreevalsan-Nair

    Abstract: In this article, we improve the deep learning solution for coastline extraction from Synthetic Aperture Radar (SAR) images by proposing a two-stage model involving image classification followed by segmentation. We hypothesize that a single segmentation model usually used for coastline detection is insufficient to characterize different coastline types. We demonstrate that the need for a two-stage… ▽ More

    Submitted 21 January, 2025; originally announced January 2025.

  17. arXiv:2501.00465  [pdf, ps, other

    cs.LG

    Dementia Detection using Multi-modal Methods on Audio Data

    Authors: Saugat Kannojia, Anirudh Praveen, Danish Vasdev, Saket Nandedkar, Divyansh Mittal, Sarthak Kalankar, Shaurya Johari, Vipul Arora

    Abstract: Dementia is a neurodegenerative disease that causes gradual cognitive impairment, which is very common in the world and undergoes a lot of research every year to prevent and cure it. It severely impacts the patient's ability to remember events and communicate clearly, where most variations of it have no known cure, but early detection can help alleviate symptoms before they become worse. One of th… ▽ More

    Submitted 7 July, 2025; v1 submitted 31 December, 2024; originally announced January 2025.

    Comments: 4 pages

  18. arXiv:2411.14431  [pdf, ps, other

    cs.CC cs.DS

    On Optimal Testing of Linearity

    Authors: Vipul Arora, Esty Kelman, Uri Meir

    Abstract: Linearity testing has been a focal problem in property testing of functions. We combine different known techniques and observations about linearity testing in order to resolve two recent versions of this task. First, we focus on the online manipulations model introduced by Kalemaj, Raskhodnikova and Varma (ITCS 2022 \& Theory of Computing 2023). In this model, up to $t$ data entries are adversar… ▽ More

    Submitted 21 November, 2024; originally announced November 2024.

    Comments: To appear at SOSA 2025

  19. arXiv:2411.14100  [pdf, other

    eess.AS cs.CL cs.IR

    BEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken Term Detection

    Authors: Anup Singh, Kris Demuynck, Vipul Arora

    Abstract: Spoken term detection (STD) is often hindered by reliance on frame-level features and the computationally intensive DTW-based template matching, limiting its practicality. To address these challenges, we propose a novel approach that encodes speech into discrete, speaker-agnostic semantic tokens. This facilitates fast retrieval using text-based search algorithms and effectively handles out-of-voca… ▽ More

    Submitted 21 December, 2024; v1 submitted 21 November, 2024; originally announced November 2024.

    Comments: Accepted at ICASSP 2025

  20. arXiv:2408.01725  [pdf, other

    cs.CY

    The Drama Machine: Simulating Character Development with LLM Agents

    Authors: Liam Magee, Vanicka Arora, Gus Gollings, Norma Lam-Saw

    Abstract: This paper explores use of multiple large language model (LLM) agents to simulate complex, dynamic characters in dramatic scenarios. We introduce a drama machine framework that coordinates interactions between LLM agents playing different 'Ego' and 'Superego' psychological roles. In roleplay simulations, this design allows intersubjective dialogue and intra-subjective internal monologue to develop… ▽ More

    Submitted 31 August, 2024; v1 submitted 3 August, 2024; originally announced August 2024.

    Comments: 28 pages, 2 figures

    ACM Class: J.4; J.5; K.4.2

  21. arXiv:2407.11907  [pdf, ps, other

    cs.LG cs.SI

    GraphFM: A generalist graph transformer that learns transferable representations across diverse domains

    Authors: Divyansha Lachi, Mehdi Azabou, Vinam Arora, Eva Dyer

    Abstract: Graph neural networks (GNNs) are often trained on individual datasets, requiring specialized models and significant hyperparameter tuning due to the unique structures and features of each dataset. This approach limits the scalability and generalizability of GNNs, as models must be tailored for each specific graph type. To address these challenges, we introduce GraphFM, a scalable multi-graph pretr… ▽ More

    Submitted 14 February, 2026; v1 submitted 16 July, 2024; originally announced July 2024.

    Journal ref: Transactions on Machine Learning Research, 2025

  22. Explainable Deep Learning Analysis for Raga Identification in Indian Art Music

    Authors: Parampreet Singh, Vipul Arora

    Abstract: Raga identification is an important problem within the domain of Indian Art music, as Ragas are fundamental to its composition and performance, playing a crucial role in music retrieval, preservation, and education. Few studies that have explored this task employ approaches such as signal processing, Machine Learning (ML), and more recently, Deep Learning (DL) based methods. However, a key questio… ▽ More

    Submitted 21 December, 2024; v1 submitted 4 June, 2024; originally announced June 2024.

    Journal ref: IEEE Transactions on Audio, Speech and Language Processing, vol. 33, ISSN: 2998-4173, pp. 2302-2311, 2025

  23. arXiv:2405.09734  [pdf

    cs.CY

    Attention is All You Want: Machinic Gaze and the Anthropocene

    Authors: Liam Magee, Vanicka Arora

    Abstract: This chapter experiments with ways computational vision interprets and synthesises representations of the Anthropocene. Text-to-image systems such as MidJourney and StableDiffusion, trained on large data sets of harvested images and captions, yield often striking compositions that serve, alternately, as banal reproduction, alien imaginary and refracted commentary on the preoccupations of Internet… ▽ More

    Submitted 15 May, 2024; originally announced May 2024.

    Comments: 19 pages

    ACM Class: K.4.2; J.5

  24. arXiv:2403.09465  [pdf, other

    cs.DS cs.LG stat.ML

    Outlier Robust Multivariate Polynomial Regression

    Authors: Vipul Arora, Arnab Bhattacharyya, Mathews Boban, Venkatesan Guruswami, Esty Kelman

    Abstract: We study the problem of robust multivariate polynomial regression: let $p\colon\mathbb{R}^n\to\mathbb{R}$ be an unknown $n$-variate polynomial of degree at most $d$ in each variable. We are given as input a set of random samples $(\mathbf{x}_i,y_i) \in [-1,1]^n \times \mathbb{R}$ that are noisy versions of $(\mathbf{x}_i,p(\mathbf{x}_i))$. More precisely, each $\mathbf{x}_i$ is sampled independent… ▽ More

    Submitted 14 March, 2024; originally announced March 2024.

  25. arXiv:2402.07599  [pdf, other

    eess.AS cs.SD

    Interactive singing melody extraction based on active adaptation

    Authors: Kavya Ranjan Saxena, Vipul Arora

    Abstract: Extraction of predominant pitch from polyphonic audio is one of the fundamental tasks in the field of music information retrieval and computational musicology. To accomplish this task using machine learning, a large amount of labeled audio data is required to train the model. However, a classical model pre-trained on data from one domain (source), e.g., songs of a particular singer or genre, may n… ▽ More

    Submitted 12 February, 2024; originally announced February 2024.

  26. arXiv:2401.15948  [pdf, other

    cs.LG cond-mat.stat-mech physics.comp-ph

    AdvNF: Reducing Mode Collapse in Conditional Normalising Flows using Adversarial Learning

    Authors: Vikas Kanaujia, Mathias S. Scheurer, Vipul Arora

    Abstract: Deep generative models complement Markov-chain-Monte-Carlo methods for efficiently sampling from high-dimensional distributions. Among these methods, explicit generators, such as Normalising Flows (NFs), in combination with the Metropolis Hastings algorithm have been extensively applied to get unbiased samples from target distributions. We systematically study central problems in conditional NFs,… ▽ More

    Submitted 11 April, 2024; v1 submitted 29 January, 2024; originally announced January 2024.

    Comments: 29 pages, submitted to Scipost Physics

    Journal ref: SciPost Phys. 16, 132 (2024)

  27. arXiv:2401.03251  [pdf, other

    eess.AS cs.LG cs.SD stat.ML

    TeLeS: Temporal Lexeme Similarity Score to Estimate Confidence in End-to-End ASR

    Authors: Nagarathna Ravi, Thishyan Raj T, Vipul Arora

    Abstract: Confidence estimation of predictions from an End-to-End (E2E) Automatic Speech Recognition (ASR) model benefits ASR's downstream and upstream tasks. Class-probability-based confidence scores do not accurately represent the quality of overconfident ASR predictions. An ancillary Confidence Estimation Model (CEM) calibrates the predictions. State-of-the-art (SOTA) solutions use binary target scores f… ▽ More

    Submitted 6 January, 2024; originally announced January 2024.

    Comments: Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing

  28. arXiv:2310.16046  [pdf, other

    cs.LG q-bio.NC

    A Unified, Scalable Framework for Neural Population Decoding

    Authors: Mehdi Azabou, Vinam Arora, Venkataramana Ganesh, Ximeng Mao, Santosh Nachimuthu, Michael J. Mendelson, Blake Richards, Matthew G. Perich, Guillaume Lajoie, Eva L. Dyer

    Abstract: Our ability to use deep learning approaches to decipher neural activity would likely benefit from greater scale, in terms of both model size and datasets. However, the integration of many neural recordings into one unified model is challenging, as each recording contains the activity of different neurons from different individual animals. In this paper, we introduce a training framework and archit… ▽ More

    Submitted 24 October, 2023; originally announced October 2023.

    Comments: Accepted at NeurIPS 2023

  29. arXiv:2310.04628  [pdf, other

    cs.CY

    (Re)framing Built Heritage through the Machinic Gaze

    Authors: Vanicka Arora, Liam Magee, Luke Munn

    Abstract: Built heritage has been both subject and product of a gaze that has been sustained through moments of colonial fixation on ruins and monuments, technocratic examination and representation, and fetishisation by aglobal tourist industry. We argue that the recent proliferation of machine learning and vision technologies create new scopic regimes for heritage: storing and retrieving existing images fr… ▽ More

    Submitted 6 October, 2023; originally announced October 2023.

    Comments: 18 pages, 5 figures

    ACM Class: J.5; K.4.2

  30. arXiv:2307.09753  [pdf, other

    cs.CY

    Unmaking AI Imagemaking: A Methodological Toolkit for Critical Investigation

    Authors: Luke Munn, Liam Magee, Vanicka Arora

    Abstract: AI image models are rapidly evolving, disrupting aesthetic production in many industries. However, understanding of their underlying archives, their logic of image reproduction, and their persistent biases remains limited. What kind of methods and approaches could open up these black boxes? In this paper, we provide three methodological approaches for investigating AI image models and apply them t… ▽ More

    Submitted 19 July, 2023; originally announced July 2023.

    Comments: 14 pages, 4 figures

    ACM Class: K.4.1; K.2; J.5

  31. arXiv:2306.12190  [pdf, other

    cs.LG

    Quantifying lottery tickets under label noise: accuracy, calibration, and complexity

    Authors: Viplove Arora, Daniele Irto, Sebastian Goldt, Guido Sanguinetti

    Abstract: Pruning deep neural networks is a widely used strategy to alleviate the computational burden in machine learning. Overwhelming empirical evidence suggests that pruned models retain very high accuracy even with a tiny fraction of parameters. However, relatively little work has gone into characterising the small pruned networks obtained, beyond a measure of their accuracy. In this paper, we use the… ▽ More

    Submitted 21 June, 2023; originally announced June 2023.

    Journal ref: Proceedings of the Thirty-Ninth Conference on Uncertainty in Artificial Intelligence, PMLR 216:88-98 (2023)

  32. arXiv:2304.07239   

    cs.CR

    Separating Key Agreement and Computational Differential Privacy

    Authors: Vipul Arora, Eldon Chung, Zeyong Li, Thomas Tan

    Abstract: Two party differential privacy allows two parties who do not trust each other, to come together and perform a joint analysis on their data whilst maintaining individual-level privacy. We show that any efficient, computationally differentially private protocol that has black-box access to key agreement (and nothing stronger), is also an efficient, information-theoretically differentially private pr… ▽ More

    Submitted 28 August, 2023; v1 submitted 14 April, 2023; originally announced April 2023.

    Comments: A key step in relating the probability that can be computed by the PSPACE algorithm to the statistical distinguishing probability is missing and not yet shown. Our arguments in this work so far have not yet been able to show this step. Thus the final conclusion that key agreement is black-box insufficient for CDP is not yet proven

  33. arXiv:2304.06733  [pdf, ps, other

    cs.LG cs.DS cs.IT math.ST

    Near-Optimal Degree Testing for Bayes Nets

    Authors: Vipul Arora, Arnab Bhattacharyya, Clément L. Canonne, Joy Qiping Yang

    Abstract: This paper considers the problem of testing the maximum in-degree of the Bayes net underlying an unknown probability distribution $P$ over $\{0,1\}^n$, given sample access to $P$. We show that the sample complexity of the problem is $\tildeΘ(2^{n/2}/\varepsilon^2)$. Our algorithm relies on a testing-by-learning framework, previously used to obtain sample-optimal testers; in order to apply this fra… ▽ More

    Submitted 12 April, 2023; originally announced April 2023.

  34. arXiv:2301.12066  [pdf, other

    cs.CY cs.AI

    Truth Machines: Synthesizing Veracity in AI Language Models

    Authors: Luke Munn, Liam Magee, Vanicka Arora

    Abstract: As AI technologies are rolled out into healthcare, academia, human resources, law, and a multitude of other domains, they become de-facto arbiters of truth. But truth is highly contested, with many different definitions and approaches. This article discusses the struggle for truth in AI systems and the general responses to date. It then investigates the production of truth in InstructGPT, a large… ▽ More

    Submitted 27 January, 2023; originally announced January 2023.

    Comments: 20 pages, 3 figures

  35. arXiv:2212.05058  [pdf, other

    cs.CY cs.AI

    Structured Like a Language Model: Analysing AI as an Automated Subject

    Authors: Liam Magee, Vanicka Arora, Luke Munn

    Abstract: Drawing from the resources of psychoanalysis and critical media studies, in this paper we develop an analysis of Large Language Models (LLMs) as automated subjects. We argue the intentional fictional projection of subjectivity onto LLMs can yield an alternate frame through which AI behaviour, including its productions of bias and harm, can be analysed. First, we introduce language models, discuss… ▽ More

    Submitted 8 December, 2022; originally announced December 2022.

  36. arXiv:2211.11060   

    eess.AS cs.LG cs.SD

    Simultaneously Learning Robust Audio Embeddings and balanced Hash codes for Query-by-Example

    Authors: Anup Singh, Kris Demuynck, Vipul Arora

    Abstract: Audio fingerprinting systems must efficiently and robustly identify query snippets in an extensive database. To this end, state-of-the-art systems use deep learning to generate compact audio fingerprints. These systems deploy indexing methods, which quantize fingerprints to hash codes in an unsupervised manner to expedite the search. However, these methods generate imbalanced hash codes, leading t… ▽ More

    Submitted 18 January, 2023; v1 submitted 20 November, 2022; originally announced November 2022.

    Comments: We need to rewrite the subsection 'Efficiency' section under section 4 to make it more easy to follow for the readers and appreciate our results

  37. arXiv:2211.10408  [pdf, other

    cs.CV

    CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical Flow

    Authors: Philippe Weinzaepfel, Thomas Lucas, Vincent Leroy, Yohann Cabon, Vaibhav Arora, Romain Brégier, Gabriela Csurka, Leonid Antsfeld, Boris Chidlovskii, Jérôme Revaud

    Abstract: Despite impressive performance for high-level downstream tasks, self-supervised pre-training methods have not yet fully delivered on dense geometric vision tasks such as stereo matching or optical flow. The application of self-supervised concepts, such as instance discrimination or masked image modeling, to geometric tasks is an active area of research. In this work, we build on the recent cross-v… ▽ More

    Submitted 18 August, 2023; v1 submitted 18 November, 2022; originally announced November 2022.

    Comments: ICCV 2023

  38. arXiv:2211.09376  [pdf, other

    cs.SD cs.LG eess.AS

    Balanced Deep CCA for Bird Vocalization Detection

    Authors: Sumit Kumar, B. Anshuman, Linus Ruettimann, Richard H. R. Hahnloser, Vipul Arora

    Abstract: Event detection improves when events are captured by two different modalities rather than just one. But to train detection systems on multiple modalities is challenging, in particular when there is abundance of unlabelled data but limited amounts of labeled data. We develop a novel self-supervised learning technique for multi-modal data that learns (hidden) correlations between simultaneously reco… ▽ More

    Submitted 17 November, 2022; originally announced November 2022.

  39. arXiv:2210.12532   

    eess.AS cs.SD

    Deep domain adaptation for polyphonic melody extraction

    Authors: Kavya Ranjan Saxena, Vipul Arora

    Abstract: Extraction of the predominant pitch from polyphonic audio is one of the fundamental tasks in the field of music information retrieval and computational musicology. To accomplish this task using machine learning, a large amount of labeled audio data is required to train the model that predicts the pitch contour. But a classical model pre-trained on data from one domain (source), e.g, songs of a par… ▽ More

    Submitted 5 April, 2023; v1 submitted 22 October, 2022; originally announced October 2022.

    Comments: Want to withdraw this paper because few concepts of domain adaptation are not clear in the paper

  40. arXiv:2210.10716  [pdf, other

    cs.CV

    CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion

    Authors: Philippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Romain Brégier, Yohann Cabon, Vaibhav Arora, Leonid Antsfeld, Boris Chidlovskii, Gabriela Csurka, Jérôme Revaud

    Abstract: Masked Image Modeling (MIM) has recently been established as a potent pre-training paradigm. A pretext task is constructed by masking patches in an input image, and this masked content is then predicted by a neural network using visible patches as sole input. This pre-training leads to state-of-the-art performance when finetuned for high-level semantic tasks, e.g. image classification and object d… ▽ More

    Submitted 12 January, 2023; v1 submitted 19 October, 2022; originally announced October 2022.

    Comments: NeurIPS 2022

  41. Attention-Based Audio Embeddings for Query-by-Example

    Authors: Anup Singh, Kris Demuynck, Vipul Arora

    Abstract: An ideal audio retrieval system efficiently and robustly recognizes a short query snippet from an extensive database. However, the performance of well-known audio fingerprinting systems falls short at high signal distortion levels. This paper presents an audio retrieval system that generates noise and reverberation robust audio fingerprints using the contrastive learning framework. Using these fin… ▽ More

    Submitted 21 November, 2024; v1 submitted 16 October, 2022; originally announced October 2022.

    Comments: Accepted in ISMIR 2022

  42. arXiv:2210.00521  [pdf, other

    cs.LG eess.SP

    Leveraging unsupervised data and domain adaptation for deep regression in low-cost sensor calibration

    Authors: Swapnil Dey, Vipul Arora, Sachchida Nand Tripathi

    Abstract: Air quality monitoring is becoming an essential task with rising awareness about air quality. Low cost air quality sensors are easy to deploy but are not as reliable as the costly and bulky reference monitors. The low quality sensors can be calibrated against the reference monitors with the help of deep learning. In this paper, we translate the task of sensor calibration into a semi-supervised dom… ▽ More

    Submitted 2 October, 2022; originally announced October 2022.

    Comments: submitted to IEEE Trans. on Neural Networks and Learning Systems as a regular article

  43. arXiv:2207.01942  [pdf, other

    cs.SE

    An Exploratory Study on Regression Vulnerabilities

    Authors: Larissa Braz, Enrico Fregnan, Vivek Arora, Alberto Bacchelli

    Abstract: Background: Security regressions are vulnerabilities introduced in a previously unaffected software system. They often happen as a result of source code changes (e.g., a bug fix) and can have severe effects. Aims: To increase the understanding of security regressions. This is an important step in developing secure software engineering. Method: We perform an exploratory, mixed-method case study… ▽ More

    Submitted 5 July, 2022; originally announced July 2022.

    Comments: This paper has been accepted at ESEM 2022 (16th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement)

  44. arXiv:2204.08404  [pdf, ps, other

    cs.DS

    Low Degree Testing over the Reals

    Authors: Vipul Arora, Arnab Bhattacharyya, Noah Fleming, Esty Kelman, Yuichi Yoshida

    Abstract: We study the problem of testing whether a function $f: \mathbb{R}^n \to \mathbb{R}$ is a polynomial of degree at most $d$ in the \emph{distribution-free} testing model. Here, the distance between functions is measured with respect to an unknown distribution $\mathcal{D}$ over $\mathbb{R}^n$ from which we can draw samples. In contrast to previous work, we do not assume that $\mathcal{D}$ has finite… ▽ More

    Submitted 18 April, 2022; originally announced April 2022.

  45. arXiv:2108.00640  [pdf, ps, other

    cs.LG eess.SP

    Few-shot calibration of low-cost air pollution (PM2.5) sensors using meta-learning

    Authors: Kalpit Yadav, Vipul Arora, Sonu Kumar Jha, Mohit Kumar, Sachchida Nand Tripathi

    Abstract: Low-cost particulate matter sensors are transforming air quality monitoring because they have lower costs and greater mobility as compared to reference monitors. Calibration of these low-cost sensors requires training data from co-deployed reference monitors. Machine Learning based calibration gives better performance than conventional techniques, but requires a large amount of training data from… ▽ More

    Submitted 2 August, 2021; originally announced August 2021.

    Comments: 3+1 pages, submitted to IEEE sensors conference 2021

  46. arXiv:2104.03476  [pdf, other

    cs.SE

    Secure Software Engineering in the Financial Services: A Practitioners' Perspective

    Authors: Vivek Arora, Enrique Larios Vargas, Maurício Aniche, Arie van Deursen

    Abstract: Secure software engineering is a fundamental activity in modern software development. However, while the field of security research has been advancing quite fast, in practice, there is still a vast knowledge gap between the security experts and the software development teams. After all, we cannot expect developers and other software practitioners to be security experts. Understanding how software… ▽ More

    Submitted 7 April, 2021; originally announced April 2021.

  47. arXiv:2011.10337  [pdf, other

    cs.LG cs.CL cs.IR

    Finding Prerequisite Relations between Concepts using Textbook

    Authors: Shivam Pal, Vipul Arora, Pawan Goyal

    Abstract: A prerequisite is anything that you need to know or understand first before attempting to learn or understand something new. In the current work, we present a method of finding prerequisite relations between concepts using related textbooks. Previous researchers have focused on finding these relations using Wikipedia link structure through unsupervised and supervised learning approaches. In the cu… ▽ More

    Submitted 20 November, 2020; originally announced November 2020.

  48. arXiv:2007.04950  [pdf, other

    cs.CV cs.AI cs.HC

    AI Assisted Apparel Design

    Authors: Alpana Dubey, Nitish Bhardwaj, Kumar Abhinav, Suma Mani Kuriakose, Sakshi Jain, Veenu Arora

    Abstract: Fashion is a fast-changing industry where designs are refreshed at large scale every season. Moreover, it faces huge challenge of unsold inventory as not all designs appeal to customers. This puts designers under significant pressure. Firstly, they need to create innumerous fresh designs. Secondly, they need to create designs that appeal to customers. Although we see advancements in approaches to… ▽ More

    Submitted 10 July, 2020; v1 submitted 9 July, 2020; originally announced July 2020.

  49. arXiv:1812.09693  [pdf

    cs.CV cs.IR

    Image Processing on IOPA Radiographs: A comprehensive case study on Apical Periodontitis

    Authors: Diganta Misra, Vanshika Arora

    Abstract: With the recent advancements in Image Processing Techniques and development of new robust computer vision algorithms, new areas of research within Medical Diagnosis and Biomedical Engineering are picking up pace. This paper provides a comprehensive in-depth case study of Image Processing, Feature Extraction and Analysis of Apical Periodontitis diagnostic cases in IOPA (Intra Oral Peri-Apical) Radi… ▽ More

    Submitted 22 March, 2019; v1 submitted 23 December, 2018; originally announced December 2018.

    Comments: 15 pages, 42 figures and Submitted at ICIAP 2019: 21st International Conference on Image Analysis and Processing

  50. arXiv:1504.07278  [pdf, ps, other

    cs.NE

    Optimal Convergence Rate in Feed Forward Neural Networks using HJB Equation

    Authors: Vipul Arora, Laxmidhar Behera, Ajay Pratap Yadav

    Abstract: A control theoretic approach is presented in this paper for both batch and instantaneous updates of weights in feed-forward neural networks. The popular Hamilton-Jacobi-Bellman (HJB) equation has been used to generate an optimal weight update law. The remarkable contribution in this paper is that closed form solutions for both optimal cost and weight update can be achieved for any feed-forward net… ▽ More

    Submitted 27 April, 2015; originally announced April 2015.

    Comments: 9 pages, journal