Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–24 of 24 results for author: Dey, L

Searching in archive cs. Search in all archives.
.
  1. arXiv:2605.30018  [pdf, ps, other

    cs.CL cs.LG

    Latent Performance Profiling of Large Language Models

    Authors: Tanmoy Chakraborty, Ayan Sengupta, Suparna Bhattacharya, Partha Pratim Chakrabarti, Amlan Chakrabarti, Supratik Chakraborty, Partha Pratim Das, Lipika Dey, Richa Singh, Mayank Vatsa

    Abstract: Large language models (LLMs) frequently achieve impressive scores on standardized benchmarks, yet accuracy alone offers a limited view of their capabilities. Evaluating open-source LLMs through leaderboards faces persistent issues like data contamination, narrow task scope, and weak alignment with real-world reliability. Benchmark-based evaluations such as MMLU PRO, BBH, or IFEval primarily captur… ▽ More

    Submitted 29 May, 2026; v1 submitted 28 May, 2026; originally announced May 2026.

  2. arXiv:2605.24034  [pdf, ps, other

    q-bio.GN cs.AI

    WTKO-CNN: Deep Learning Reveals Sequence Motifs Distinguishing Wild-Type and Knockout ATAC-seq Peaks

    Authors: Lopamudra Dey

    Abstract: Chromatin regulators can alter transcriptional programs by modifying the accessibility of regulatory DNA elements. Understanding how regulatory sequences differ between wild-type (WT) and knockout (KO) conditions is crucial for deciphering transcriptional control. Here, we applied a convolutional neural network, \textbf{WTKO-CNN} with an attention mechanism to classify DNA sequences as WT or KO, a… ▽ More

    Submitted 21 May, 2026; originally announced May 2026.

  3. arXiv:2509.16286  [pdf, ps, other

    cs.CY

    What's Not on the Plate? Rethinking Food Computing through Indigenous Indian Datasets

    Authors: Pamir Gogoi, Neha Joshi, Ayushi Pandey, Deepthi Sudharsan, Saransh Kumar Gupta, Lipika Dey, Partha Pratim Das, Kalika Bali, Vivek Seshadri

    Abstract: This paper presents a multimodal dataset of 1,000 indigenous recipes from remote regions of India, collected through a participatory model involving first-time digital workers from rural areas. The project covers ten endangered language communities in six states. Documented using a dedicated mobile app, the data set includes text, images, and audio, capturing traditional food practices along with… ▽ More

    Submitted 19 September, 2025; originally announced September 2025.

    MSC Class: I.2.4; K.4.0; J.5

  4. Towards an Action-Centric Ontology for Cooking Procedures Using Temporal Graphs

    Authors: Aarush Kumbhakern, Saransh Kumar Gupta, Lipika Dey, Partha Pratim Das

    Abstract: Formalizing cooking procedures remains a challenging task due to their inherent complexity and ambiguity. We introduce an extensible domain-specific language for representing recipes as directed action graphs, capturing processes, transfers, environments, concurrency, and compositional structure. Our approach enables precise, modular modeling of complex culinary workflows. Initial manual evaluatio… ▽ More

    Submitted 4 September, 2025; originally announced September 2025.

    Comments: 6 pages, 3 figures, 1 table, 11 references, ACM International Conference on Multimedia 2025 - Multi-modal Food Computing Workshop

  5. arXiv:2508.16117  [pdf, ps, other

    cs.AI cs.CL cs.IR

    Extending FKG.in: Towards a Food Claim Traceability Network

    Authors: Saransh Kumar Gupta, Rizwan Gulzar Mir, Lipika Dey, Partha Pratim Das, Anirban Sen, Ramesh Jain

    Abstract: The global food landscape is rife with scientific, cultural, and commercial claims about what foods are, what they do, what they should not do, or should not do. These range from rigorously studied health benefits (probiotics improve gut health) and misrepresentations (soaked almonds make one smarter) to vague promises (superfoods boost immunity) and culturally rooted beliefs (cold foods cause cou… ▽ More

    Submitted 4 September, 2025; v1 submitted 22 August, 2025; originally announced August 2025.

    Comments: 10 pages, 3 figures, 1 table, 45 references, ACM International Conference on Multimedia 2025 - Multi-modal Food Computing Workshop

  6. Enhancing FKG.in: automating Indian food composition analysis

    Authors: Saransh Kumar Gupta, Lipika Dey, Partha Pratim Das, Geeta Trilok-Kumar, Ramesh Jain

    Abstract: This paper presents a novel approach to compute food composition data for Indian recipes using a knowledge graph for Indian food (FKG[.]in) and LLMs. The primary focus is to provide a broad overview of an automated food composition analysis workflow and describe its core functionalities: nutrition data aggregation, food composition analysis, and LLM-augmented information resolution. This workflow… ▽ More

    Submitted 4 September, 2025; v1 submitted 6 December, 2024; originally announced December 2024.

    Comments: 15 pages, 5 figures, 30 references, International Conference on Pattern Recognition 2024 - Multimedia Assisted Dietary Management Workshop

  7. arXiv:2409.00830  [pdf, other

    cs.AI cs.CL cs.IR

    Building FKG.in: a Knowledge Graph for Indian Food

    Authors: Saransh Kumar Gupta, Lipika Dey, Partha Pratim Das, Ramesh Jain

    Abstract: This paper presents an ontology design along with knowledge engineering, and multilingual semantic reasoning techniques to build an automated system for assimilating culinary information for Indian food in the form of a knowledge graph. The main focus is on designing intelligent methods to derive ontology designs and capture all-encompassing knowledge about food, recipes, ingredients, cooking char… ▽ More

    Submitted 1 September, 2024; originally announced September 2024.

    Comments: 14 pages, 3 figures, 25 references, Formal Ontology in Information Systems Conference 2024 - Integrated Food Ontology Workshop

  8. arXiv:2408.15550  [pdf, other

    cs.AI

    Trustworthy and Responsible AI for Human-Centric Autonomous Decision-Making Systems

    Authors: Farzaneh Dehghani, Mahsa Dibaji, Fahim Anzum, Lily Dey, Alican Basdemir, Sayeh Bayat, Jean-Christophe Boucher, Steve Drew, Sarah Elaine Eaton, Richard Frayne, Gouri Ginde, Ashley Harris, Yani Ioannou, Catherine Lebel, John Lysack, Leslie Salgado Arzuaga, Emma Stanley, Roberto Souza, Ronnie de Souza Santos, Lana Wells, Tyler Williamson, Matthias Wilms, Zaman Wahid, Mark Ungrin, Marina Gavrilova , et al. (1 additional authors not shown)

    Abstract: Artificial Intelligence (AI) has paved the way for revolutionary decision-making processes, which if harnessed appropriately, can contribute to advancements in various sectors, from healthcare to economics. However, its black box nature presents significant ethical challenges related to bias and transparency. AI applications are hugely impacted by biases, presenting inconsistent and unreliable fin… ▽ More

    Submitted 2 September, 2024; v1 submitted 28 August, 2024; originally announced August 2024.

    Comments: 44 pages, 2 figures

  9. arXiv:2407.21039  [pdf, other

    cs.CL cs.AI cs.LG

    Mapping Patient Trajectories: Understanding and Visualizing Sepsis Prognostic Pathways from Patients Clinical Narratives

    Authors: Sudeshna Jana, Tirthankar Dasgupta, Lipika Dey

    Abstract: In recent years, healthcare professionals are increasingly emphasizing on personalized and evidence-based patient care through the exploration of prognostic pathways. To study this, structured clinical variables from Electronic Health Records (EHRs) data have traditionally been employed by many researchers. Presently, Natural Language Processing models have received great attention in clinical res… ▽ More

    Submitted 20 July, 2024; originally announced July 2024.

    Comments: preprint, 8 pages, 6 figures

  10. Generating insights about financial asks from Reddit posts and user interactions

    Authors: Sachin Thukral, Suyash Sangwan, Vipul Chauhan, Arnab Chatterjee, Lipika Dey

    Abstract: As an increasingly large number of people turn to platforms like Reddit, YouTube, Twitter, Instagram, etc. for financial advice, generating insights about the content generated and interactions taking place within these platforms have become a key research question. This study proposes content and interaction analysis techniques for a large repository created from social media content, where peopl… ▽ More

    Submitted 12 March, 2024; v1 submitted 7 March, 2024; originally announced March 2024.

    Comments: 6 pages, 3 figures, 2 tables; In ASONAM 2023 (The 2023 IEEE/ACM International Conference on Advances in Social Network Analysis and Mining)

  11. Understanding how social discussion platforms like Reddit are influencing financial behavior

    Authors: Sachin Thukral, Suyash Sangwan, Arnab Chatterjee, Lipika Dey, Aaditya Agrawal, Pramit Kumar Chandra, Animesh Mukherjee

    Abstract: This study proposes content and interaction analysis techniques for a large repository created from social media content. Though we have presented our study for a large platform dedicated to discussions around financial topics, the proposed methods are generic and applicable to all platforms. Along with an extension of topic extraction method using Latent Dirichlet Allocation, we propose a few mea… ▽ More

    Submitted 12 March, 2024; v1 submitted 7 March, 2024; originally announced March 2024.

    Comments: 8 pages, 8 figures, 3 tables, and 1 algorithm; Published in WI-IAT 2022 (The 21st IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology)

    Journal ref: IEEE/WIC/ACM International Joint Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT) 2022 (pp. 612-619)

  12. arXiv:2308.00005  [pdf

    cs.IR

    Detection and Classification of Novel Attacks and Anomaly in IoT Network using Rule based Deep Learning Model

    Authors: Sanjay Chakraborty, Saroj Kumar Pandey, Saikat Maity, Lopamudra Dey

    Abstract: Attackers are now using sophisticated techniques, like polymorphism, to change the attack pattern for each new attack. Thus, the detection of novel attacks has become the biggest challenge for cyber experts and researchers. Recently, anomaly and hybrid approaches are used for the detection of network attacks. Detecting novel attacks, on the other hand, is a key enabler for a wide range of IoT appl… ▽ More

    Submitted 29 July, 2023; originally announced August 2023.

  13. arXiv:2012.00333  [pdf, other

    cs.SI cs.CY

    Identifying pandemic-related stress factors from social-media posts -- effects on students and young-adults

    Authors: Sachin Thukral, Suyash Sangwan, Arnab Chatterjee, Lipika Dey

    Abstract: The COVID-19 pandemic has thrown natural life out of gear across the globe. Strict measures are deployed to curb the spread of the virus that is causing it, and the most effective of them have been social isolation. This has led to wide-spread gloom and depression across society but more so among the young and the elderly. There are currently more than 200 million college students in 186 countries… ▽ More

    Submitted 1 December, 2020; originally announced December 2020.

    Comments: 10 pages, 5 figures

  14. arXiv:2007.13520  [pdf, other

    cs.DL cs.CL cs.IR cs.LG cs.SI

    Identification, Tracking and Impact: Understanding the trade secret of catchphrases

    Authors: Jagriti Jalal, Mayank Singh, Arindam Pal, Lipika Dey, Animesh Mukherjee

    Abstract: Understanding the topical evolution in industrial innovation is a challenging problem. With the advancement in the digital repositories in the form of patent documents, it is becoming increasingly more feasible to understand the innovation secrets -- "catchphrases" of organizations. However, searching and understanding this enormous textual information is a natural bottleneck. In this paper, we pr… ▽ More

    Submitted 20 July, 2020; originally announced July 2020.

    Comments: To be published in the proceedings of the ACM/IEEE Joint Conference on Digital Libraries (JCDL 2020)

  15. arXiv:2004.09715  [pdf, other

    cs.DL

    Innovation and Revenue: Deep Diving into the Temporal Rank-shifts of Fortune 500 Companies

    Authors: Mayank Singh, Arindam Pal, Lipika Dey, Animesh Mukherjee

    Abstract: Research and innovation is important agenda for any company to remain competitive in the market. The relationship between innovation and revenue is a key metric for companies to decide on the amount to be invested for future research. Two important parameters to evaluate innovation are the quantity and quality of scientific papers and patents. Our work studies the relationship between innovation a… ▽ More

    Submitted 20 April, 2020; originally announced April 2020.

    Comments: Accepted at CODS-COMAD 2020

  16. arXiv:1911.02771  [pdf, other

    cs.SI cs.CY physics.soc-ph

    Characterizing behavioral trends in a community driven discussion platform

    Authors: Sachin Thukral, Arnab Chatterjee, Hardik Meisheri, Tushar Kataria, Aman Agarwal, Ishan Verma, Lipika Dey

    Abstract: This article presents a systematic analysis of the patterns of behavior of individuals as well as groups observed in community-driven platforms for discussion like Reddit, where users usually exchange information and viewpoints on their topics of interest. We perform a statistical analysis of the behavior of posts and model the users' interactions around them. A platform like Reddit which has grow… ▽ More

    Submitted 7 November, 2019; originally announced November 2019.

    Comments: 19 pages. Extended version of arxiv:1809.07087. Springer Lecture Notes Format, to be published in Lecture Notes in Social Networks (Springer)

  17. arXiv:1809.07087  [pdf, other

    cs.SI cs.CY cs.MA physics.soc-ph

    Analyzing behavioral trends in community driven discussion platforms like Reddit

    Authors: Sachin Thukral, Hardik Meisheri, Tushar Kataria, Aman Agarwal, Ishan Verma, Arnab Chatterjee, Lipika Dey

    Abstract: The aim of this paper is to present methods to systematically analyze individual and group behavioral patterns observed in community driven discussion platforms like Reddit where users exchange information and views on various topics of current interest. We conduct this study by analyzing the statistical behavior of posts and modeling user interactions around them. We have chosen Reddit as an exam… ▽ More

    Submitted 19 September, 2018; originally announced September 2018.

    Comments: 8 pages, 9 figs, ASONAM 2018

  18. arXiv:1801.10080  [pdf, other

    cs.DL cs.CL

    A Machine Learning Approach to Quantitative Prosopography

    Authors: Aayushee Gupta, Haimonti Dutta, Srikanta Bedathur, Lipika Dey

    Abstract: Prosopography is an investigation of the common characteristics of a group of people in history, by a collective study of their lives. It involves a study of biographies to solve historical problems. If such biographies are unavailable, surviving documents and secondary biographical data are used. Quantitative prosopography involves analysis of information from a wide variety of sources about "ord… ▽ More

    Submitted 30 January, 2018; originally announced January 2018.

  19. arXiv:1710.02745  [pdf, other

    cs.CL

    Multi-Document Summarization using Distributed Bag-of-Words Model

    Authors: Kaustubh Mani, Ishan Verma, Hardik Meisheri, Lipika Dey

    Abstract: As the number of documents on the web is growing exponentially, multi-document summarization is becoming more and more important since it can provide the main ideas in a document set in short time. In this paper, we present an unsupervised centroid-based document-level reconstruction framework using distributed bag of words model. Specifically, our approach selects summary sentences in order to mi… ▽ More

    Submitted 11 June, 2018; v1 submitted 7 October, 2017; originally announced October 2017.

  20. Sentiment Analysis of Review Datasets Using Naive Bayes and K-NN Classifier

    Authors: Lopamudra Dey, Sanjay Chakraborty, Anuraag Biswas, Beepa Bose, Sweta Tiwari

    Abstract: The advent of Web 2.0 has led to an increase in the amount of sentimental content available in the Web. Such content is often found in social media web sites in the form of movie or product reviews, user comments, testimonials, messages in discussion forums etc. Timely discovery of the sentimental or opinionated web content has a number of advantages, the most important of all being monetization.… ▽ More

    Submitted 31 October, 2016; originally announced October 2016.

    Comments: Volume-8, Issue-4, pp.54-62, 2016

  21. arXiv:1501.06456  [pdf

    cs.DB

    Weather forecasting using Convex hull & K-Means Techniques An Approach

    Authors: Ratul Dey Sanjay Chakraborty Lopamudra Dey

    Abstract: Data mining is a popular concept of mined necessary data from a large set of data. Data mining using clustering is a powerful way to analyze data and gives prediction. In this paper non structural time series data is used to forecast daily average temperature, humidity and overall weather conditions of Kolkata city. The air pollution data have been taken from West Bengal Pollution Control Board to… ▽ More

    Submitted 26 January, 2015; originally announced January 2015.

    Comments: 1st International Science & Technology Congress(IEMCON-2015) Elsevier

  22. arXiv:1411.7469  [pdf

    cs.DB

    Canonical PSO Based k-Means Clustering Approach for Real Datasets

    Authors: Lopamudra Dey, Sanjay Chakraborty

    Abstract: "Clustering" the significance and application of this technique is spread over various fields. Clustering is an unsupervised process in data mining, that is why the proper evaluation of the results and measuring the compactness and separability of the clusters are important issues.The procedure of evaluating the results of a clustering algorithm is known as cluster validity measure. Different type… ▽ More

    Submitted 26 November, 2014; originally announced November 2014.

  23. arXiv:1406.4756  [pdf

    cs.CY

    Weather Forecasting using Incremental K-means Clustering

    Authors: Sanjay Chakraborty, N. K. Nagwani, Lopamudra Dey

    Abstract: Clustering is a powerful tool which has been used in several forecasting works, such as time series forecasting, real time storm detection, flood forecasting and so on. In this paper, a generic methodology for weather forecasting is proposed by the help of incremental K-means clustering algorithm. Weather forecasting plays an important role in day to day applications.Weather forecasting of this pa… ▽ More

    Submitted 18 June, 2014; originally announced June 2014.

  24. arXiv:1406.4751  [pdf

    cs.DB cs.IR

    Performance Comparison of Incremental K-means and Incremental DBSCAN Algorithms

    Authors: Sanjay Chakraborty, N. K. Nagwani, Lopamudra Dey

    Abstract: Incremental K-means and DBSCAN are two very important and popular clustering techniques for today's large dynamic databases (Data warehouses, WWW and so on) where data are changed at random fashion. The performance of the incremental K-means and the incremental DBSCAN are different with each other based on their time analysis characteristics. Both algorithms are efficient compare to their existing… ▽ More

    Submitted 18 June, 2014; originally announced June 2014.