Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 53 results for author: Vyas, S

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.17723  [pdf, ps, other

    cs.CV

    Vision-Language Models for Analog Gauge Reading: An Empirical Study of Specialization, Transfer and Reliability

    Authors: Abdul Mueez, Aaditya Baranwal, Junior Chaj-Mejia, Guneet Bhatia, Jason T. Voelker, Shruti Vyas

    Abstract: Analog gauges remain common in industrial environments where manual inspection is costly or hazardous. The engineering application addressed here is direct numerical reading of single-target analog-gauge images, while the artificial-intelligence contribution is a systematic evaluation of specialization, transfer, robustness and reliability for a general-purpose vision-language model (VLM) without… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

    Comments: Submitted to Engineering Applications of Artificial Intelligence

  2. A Generative Approach for Improving Multi-Label Defect Classification in Photovoltaic Modules

    Authors: Abdul Mueez, Yogesh S. Rawat, Shruti Vyas

    Abstract: This paper addresses the challenge of multi-label defect classification in electroluminescence (EL) images of photovoltaic (PV) cells. Training models on images where multiple defects co-occur creates learning ambiguity, making it difficult to disentangle visual features for specific defect types, a problem compounded by the scarcity of examples for individual classes. To tackle this, we introduce… ▽ More

    Submitted 12 August, 2026; originally announced August 2026.

    Journal ref: Solar Energy 317 (2026) 114943

  3. arXiv:2606.26734  [pdf, ps, other

    cs.CV cs.AI

    Robust Onion: Peeling Open Vocab Object Detectors Under Noise

    Authors: Priyank Pathak, Mukilan Karuppasamy, Aaditya Baranwal, Shruti Vyas, Yogesh S Rawat

    Abstract: The impact of real-world noise on Open Vocabulary Object Detectors (OV-ODs) remains poorly understood due to their architectural complexity. We present our comprehensive analysis Robust Onion, an empirical study that uses controlled synthetic visual degradations to peel OV-ODs layer-by-layer, revealing how, why, and where robustness degrades, systematically analyzing feature collapse. Our findings… ▽ More

    Submitted 26 June, 2026; v1 submitted 25 June, 2026; originally announced June 2026.

    Comments: Accepted at The 19th European Conference on Computer Vision (ECCV)

  4. arXiv:2606.22890  [pdf, ps, other

    cs.CV

    PHOEBI: An Open-World Benchmark for Bacterial Identification in Phase-Contrast Microscopy

    Authors: Aaditya Baranwal, Md Jahid Hasan, Shruti Vyas

    Abstract: Optical microscopy enables rapid, label-free imaging of live bacteria and is the standard instrument for species identification across clinical, environmental, and industrial microbiology. Yet field samples are routinely polymicrobial and may contain organisms that were never seen during system training, and no computer-vision benchmark tests multi-label species identification from phase-contrast… ▽ More

    Submitted 22 June, 2026; originally announced June 2026.

  5. arXiv:2605.10157  [pdf, ps, other

    cs.CV cs.CL

    MolSight: Molecular Property Prediction with Images

    Authors: Aaditya Baranwal, Akshaj Gupta, Yogesh S Rawat, Shruti Vyas

    Abstract: Every molecule ever synthesised can be drawn as a 2D skeletal diagram, yet in modern property prediction this universally available representation has received less focus in favour of molecular graphs, 3D conformers, or billion-parameter language models, each imposing its own computational and data-engineering overhead. We present $\textbf{MolSight}$, the first systematic large-scale study of visi… ▽ More

    Submitted 14 June, 2026; v1 submitted 11 May, 2026; originally announced May 2026.

  6. arXiv:2605.04607  [pdf, ps, other

    cs.RO

    Right Model, Right Time: Real-Time Cascaded-Fidelity MPC for Bipedal Walking

    Authors: Franek Stark, Felix Wiebe, Shubham Vyas, Dennis Mronga, Frank Kirchner

    Abstract: This paper presents a multi-phase whole-body model predictive control (MPC) approach for bipedal walking, combining a detailed whole-body model in the near horizon with a simplified single-rigid-body model in the later prediction steps. This reduces computational complexity while retaining prediction capabilities. The resulting nonlinear optimal control problem is solved entirely within the genera… ▽ More

    Submitted 3 June, 2026; v1 submitted 6 May, 2026; originally announced May 2026.

    Comments: Presented at IEEE ICRA 2026 Workshop "2cnd Workshop on Frontiers of Optimization for Robotics"

    Journal ref: Proceedings of the 2nd ICRA Workshop on Frontiers of Optimization for Robotics, 2026

  7. arXiv:2604.18957  [pdf, ps, other

    cs.CV

    Bridging Foundation Models and ASTM Metallurgical Standards for Automated Grain Size Estimation from Microscopy Images

    Authors: Abdul Mueez, Shruti Vyas

    Abstract: Extracting standardized metallurgical metrics from microscopy images remains challenging due to complex grain morphology and the data demands of supervised segmentation. To bridge foundational computer vision with practical metallurgical evaluation, we propose an automated pipeline for dense instance segmentation and grain size estimation that adapts Cellpose-SAM to microstructures and integrates… ▽ More

    Submitted 20 April, 2026; originally announced April 2026.

    Comments: Accepted at the 11th IEEE Workshop on Computer Vision for Multimodal Microscopy Image Analysis (CVMI), CVPR Workshops 2026

  8. arXiv:2604.16248  [pdf, ps, other

    cs.CV

    Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization

    Authors: Siddhant Bharadwaj, Ashish Vashist, Fahimul Aleem, Shruti Vyas

    Abstract: Image geolocalization has traditionally been addressed through retrieval-based place recognition or geometry-based visual localization pipelines. Recent advances in Vision-Language Models (VLMs) have demonstrated strong zero-shot reasoning capabilities across multimodal tasks, yet their performance in geographic inference remains underexplored. In this work, we present a systematic evaluation of m… ▽ More

    Submitted 17 April, 2026; originally announced April 2026.

    Comments: Accepted to the CVPR EarthVision 2026 Workshop

  9. arXiv:2604.09881  [pdf, ps, other

    eess.AS cs.HC

    Toward using Speech to Sense Student Emotion in Remote Learning Environments

    Authors: Sargam Vyas, Bogdan Vlasenko, André Mayoraz, Egon Werlen, Per Bergamin, Mathew Magimai. -Doss

    Abstract: With advancements in multimodal communication technologies, remote learning environments such as, distance universities are increasing. Remote learning typically happens asynchronously. As a consequence, unlike face-to-face in-person classroom teaching, this lacks availability of sufficient emotional cues for making learning a pleasant experience. Motivated by advances made in the paralinguistic s… ▽ More

    Submitted 10 April, 2026; originally announced April 2026.

  10. Mixed-Integer vs. Continuous Model Predictive Control for Binary Thrusters: A Comparative Study

    Authors: Franek Stark, Jakob Middelberg, Shubham Vyas

    Abstract: Binary on/off thrusters are commonly used for spacecraft attitude and position control during proximity operations. However, their discrete nature poses challenges for conventional continuous control methods. The control of these discrete actuators is either explicitly formulated as a mixed-integer optimization problem or handled in a two-layer approach, where a continuous controller's output is c… ▽ More

    Submitted 14 April, 2026; v1 submitted 20 March, 2026; originally announced March 2026.

    Comments: Accepted to CEAS EuroGNC 2026

  11. arXiv:2602.08690  [pdf, ps, other

    cs.LG cs.CR

    SoK: The Pitfalls of Deep Reinforcement Learning for Cybersecurity

    Authors: Shae McFadden, Myles Foley, Elizabeth Bates, Ilias Tsingenopoulos, Sanyam Vyas, Vasilios Mavroudis, Chris Hicks, Fabio Pierazzi

    Abstract: Deep Reinforcement Learning (DRL) has achieved remarkable success in domains requiring sequential decision-making, motivating its application to cybersecurity problems. However, transitioning DRL from laboratory simulations to bespoke cyber environments can introduce numerous issues. This is further exacerbated by the often adversarial, non-stationary, and partially-observable nature of most cyber… ▽ More

    Submitted 22 June, 2026; v1 submitted 9 February, 2026; originally announced February 2026.

    Comments: Accepted at USENIX Security 2026

  12. ChemPro: A Progressive Chemistry Benchmark for Large Language Models

    Authors: Aaditya Baranwal, Shruti Vyas

    Abstract: We introduce ChemPro, a progressive benchmark with 4100 natural language question-answer pairs in Chemistry, across 4 coherent sections of difficulty designed to assess the proficiency of Large Language Models (LLMs) in a broad spectrum of general chemistry topics. We include Multiple Choice Questions and Numerical Questions spread across fine-grained information recall, long-horizon reasoning, mu… ▽ More

    Submitted 20 April, 2026; v1 submitted 3 February, 2026; originally announced February 2026.

    Comments: Accepted at Artificial Intelligence Chemistry Journal

    ACM Class: I.2.7; I.2.0; J.2

    Journal ref: Artif. Intell. Chem. 4(1), 100118 (2026)

  13. arXiv:2510.18600  [pdf, ps, other

    cs.RO

    Quadrupeds for Planetary Exploration: Field Testing Control Algorithms on an Active Volcano

    Authors: Shubham Vyas, Franek Stark, Rohit Kumar, Hannah Isermann, Jonas Haack, Mihaela Popescu, Jakob Middelberg, Dennis Mronga, Frank Kirchner

    Abstract: Missions such as the Ingenuity helicopter have shown the advantages of using novel locomotion modes to increase the scientific return of planetary exploration missions. Legged robots can further expand the reach and capability of future planetary missions by traversing more difficult terrain than wheeled rovers, such as jumping over cracks on the ground or traversing rugged terrain with boulders.… ▽ More

    Submitted 21 October, 2025; originally announced October 2025.

    Comments: Presented at 18th Symposium on Advanced Space Technologies in Robotics and Automation (ASTRA)

    Journal ref: 18th Symposium on Advanced Space Technologies in Robotics and Automation (ASTRA), 2025

  14. arXiv:2510.17249  [pdf, ps, other

    cs.RO

    An adaptive hierarchical control framework for quadrupedal robots in planetary exploration

    Authors: Franek Stark, Rohit Kumar, Shubham Vyas, Hannah Isermann, Jonas Haack, Mihaela Popescu, Jakob Middelberg, Dennis Mronga, Frank Kirchner

    Abstract: Planetary exploration missions require robots capable of navigating extreme and unknown environments. While wheeled rovers have dominated past missions, their mobility is limited to traversable surfaces. Legged robots, especially quadrupeds, can overcome these limitations by handling uneven, obstacle-rich, and deformable terrains. However, deploying such robots in unknown conditions is challenging… ▽ More

    Submitted 20 October, 2025; originally announced October 2025.

    Comments: Presented at 18th Symposium on Advanced Space Technologies in Robotics and Automation (ASTRA)

  15. SynSpill: Improved Industrial Spill Detection With Synthetic Data

    Authors: Aaditya Baranwal, Abdul Mueez, Jason Voelker, Guneet Bhatia, Shruti Vyas

    Abstract: Large-scale Vision-Language Models (VLMs) have transformed general-purpose visual recognition through strong zero-shot capabilities. However, their performance degrades significantly in niche, safety-critical domains such as industrial spill detection, where hazardous events are rare, sensitive, and difficult to annotate. This scarcity -- driven by privacy concerns, data sensitivity, and the infre… ▽ More

    Submitted 13 August, 2025; originally announced August 2025.

    Comments: Accepted at ICCV (VISION'25 Workshop) 2025

    Journal ref: 2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), pp. 1425-1434

  16. Re:Verse -- Can Your VLM Read a Manga?

    Authors: Aaditya Baranwal, Madhav Kataria, Naitik Agrawal, Yogesh S Rawat, Shruti Vyas

    Abstract: Current Vision Language Models (VLMs) demonstrate a critical gap between surface-level recognition and deep narrative reasoning when processing sequential visual storytelling. Through a comprehensive investigation of manga narrative understanding, we reveal that while recent large multimodal models excel at individual panel interpretation, they systematically fail at temporal causality and cross-p… ▽ More

    Submitted 18 August, 2025; v1 submitted 11 August, 2025; originally announced August 2025.

    Comments: Accepted (oral) at ICCV (AISTORY Workshop) 2025

    Journal ref: 2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), pp. 3820-3830

  17. arXiv:2508.00399  [pdf, ps, other

    cs.CV

    iSafetyBench: A video-language benchmark for safety in industrial environment

    Authors: Raiyaan Abdullah, Yogesh Singh Rawat, Shruti Vyas

    Abstract: Recent advances in vision-language models (VLMs) have enabled impressive generalization across diverse video understanding tasks under zero-shot settings. However, their capabilities in high-stakes industrial domains-where recognizing both routine operations and safety-critical anomalies is essential-remain largely underexplored. To address this gap, we introduce iSafetyBench, a new video-language… ▽ More

    Submitted 13 August, 2025; v1 submitted 1 August, 2025; originally announced August 2025.

    Comments: Accepted to VISION'25 - ICCV 2025 workshop

  18. arXiv:2507.04883  [pdf, ps, other

    cs.LG cs.AI cs.CR

    Beyond Training-time Poisoning: Component-level and Post-training Backdoors in Deep Reinforcement Learning

    Authors: Sanyam Vyas, Alberto Caron, Chris Hicks, Pete Burnap, Vasilios Mavroudis

    Abstract: Deep Reinforcement Learning (DRL) systems are increasingly used in safety-critical applications, yet their security remains severely underexplored. This work investigates backdoor attacks, which implant hidden triggers that cause malicious actions only when specific inputs appear in the observation space. Existing DRL backdoor research focuses solely on training-time attacks requiring unrealistic… ▽ More

    Submitted 7 July, 2025; originally announced July 2025.

  19. arXiv:2507.03283  [pdf, ps, other

    cs.CV

    MolVision: Molecular Property Prediction with Vision Language Models

    Authors: Deepan Adak, Yogesh Singh Rawat, Shruti Vyas

    Abstract: Molecular property prediction is a fundamental task in computational chemistry with critical applications in drug discovery and materials science. While recent works have explored Large Language Models (LLMs) for this task, they primarily rely on textual molecular representations such as SMILES/SELFIES, which can be ambiguous and structurally less informative. In this work, we introduce MolVision,… ▽ More

    Submitted 4 July, 2025; originally announced July 2025.

  20. arXiv:2506.13432  [pdf, ps, other

    cs.RO

    Adaptive Model-Base Control of Quadrupeds via Online System Identification using Kalman Filter

    Authors: Jonas Haack, Franek Stark, Shubham Vyas, Frank Kirchner, Shivesh Kumar

    Abstract: Many real-world applications require legged robots to be able to carry variable payloads. Model-based controllers such as model predictive control (MPC) have become the de facto standard in research for controlling these systems. However, most model-based control architectures use fixed plant models, which limits their applicability to different tasks. In this paper, we present a Kalman filter (KF… ▽ More

    Submitted 16 June, 2025; originally announced June 2025.

    Comments: 6 pages, 5 figures, 1 table, accepted for IEEE IROS 2025

  21. Better Semi-supervised Learning for Multi-domain ASR Through Incremental Retraining and Data Filtering

    Authors: Andres Carofilis, Pradeep Rangappa, Srikanth Madikeri, Shashi Kumar, Sergio Burdisso, Jeena Prakash, Esau Villatoro-Tello, Petr Motlicek, Bidisha Sharma, Kadri Hacioglu, Shankar Venkatesan, Saurabh Vyas, Andreas Stolcke

    Abstract: Fine-tuning pretrained ASR models for specific domains is challenging when labeled data is scarce. But unlabeled audio and labeled data from related domains are often available. We propose an incremental semi-supervised learning pipeline that first integrates a small in-domain labeled set and an auxiliary dataset from a closely related domain, achieving a relative improvement of 4% over no auxilia… ▽ More

    Submitted 5 June, 2025; originally announced June 2025.

    Comments: Accepted at Interspeech 2025, Netherlands

    Journal ref: Proc. Interspeech 2025, 3618-3622

  22. Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering

    Authors: Pradeep Rangappa, Andres Carofilis, Jeena Prakash, Shashi Kumar, Sergio Burdisso, Srikanth Madikeri, Esau Villatoro-Tello, Bidisha Sharma, Petr Motlicek, Kadri Hacioglu, Shankar Venkatesan, Saurabh Vyas, Andreas Stolcke

    Abstract: Fine-tuning pretrained ASR models for specific domains is challenging for small organizations with limited labeled data and computational resources. Here, we explore different data selection pipelines and propose a robust approach that improves ASR adaptation by filtering pseudo-labels generated using Whisper (encoder-decoder) and Zipformer (transducer) models. Our approach integrates multiple sel… ▽ More

    Submitted 4 June, 2025; originally announced June 2025.

    Comments: Accepted at Interspeech 2025, Netherlands

    Journal ref: Proc. Interspeech 2025, pp. 4928-4932

  23. arXiv:2503.15290  [pdf, other

    cs.RO

    Reinforcement Learning for Robust Athletic Intelligence: Lessons from the 2nd 'AI Olympics with RealAIGym' Competition

    Authors: Felix Wiebe, Niccolò Turcato, Alberto Dalla Libera, Jean Seong Bjorn Choe, Bumkyu Choi, Tim Lukas Faust, Habib Maraqten, Erfan Aghadavoodi, Marco Cali, Alberto Sinigaglia, Giulio Giacomuzzo, Diego Romeres, Jong-kook Kim, Gian Antonio Susto, Shubham Vyas, Dennis Mronga, Boris Belousov, Jan Peters, Frank Kirchner, Shivesh Kumar

    Abstract: In the field of robotics many different approaches ranging from classical planning over optimal control to reinforcement learning (RL) are developed and borrowed from other fields to achieve reliable control in diverse tasks. In order to get a clear understanding of their individual strengths and weaknesses and their applicability in real world robotic scenarios is it important to benchmark and co… ▽ More

    Submitted 19 March, 2025; originally announced March 2025.

    Comments: 8 pages, 7 figures

  24. arXiv:2502.03950  [pdf, other

    cs.CV

    LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models

    Authors: Priyank Pathak, Shyam Marjit, Shruti Vyas, Yogesh S Rawat

    Abstract: Visual-language foundation Models (FMs) exhibit remarkable zero-shot generalization across diverse tasks, largely attributed to extensive pre-training on largescale datasets. However, their robustness on low-resolution/pixelated (LR) images, a common challenge in real-world scenarios, remains underexplored. We introduce LR0.FM, a comprehensive benchmark evaluating the impact of low resolution on t… ▽ More

    Submitted 18 May, 2025; v1 submitted 6 February, 2025; originally announced February 2025.

    Comments: Accepted to ICLR 2025

  25. arXiv:2502.01329  [pdf, other

    cs.RO

    Benchmarking Different QP Formulations and Solvers for Dynamic Quadrupedal Walking

    Authors: Franek Stark, Jakob Middelberg, Dennis Mronga, Shubham Vyas, Frank Kirchner

    Abstract: Quadratic Programs (QPs) are widely used in the control of walking robots, especially in Model Predictive Control (MPC) and Whole-Body Control (WBC). In both cases, the controller design requires the formulation of a QP and the selection of a suitable QP solver, both requiring considerable time and expertise. While computational performance benchmarks exist for QP solvers, studies comparing optima… ▽ More

    Submitted 3 February, 2025; originally announced February 2025.

    Comments: Preprint; Accepted for ICRA 2025

  26. arXiv:2412.11194  [pdf, ps, other

    cs.SE cs.AI

    Direction for Detection: A Survey of Automated Vulnerability Detection and all of its Pain Points

    Authors: Dan Ristea, Shae McFadden, Ezzeldin Shereen, Madeleine Dwyer, Sanyam Vyas, Chris Hicks, Vasilios Mavroudis

    Abstract: Security vulnerabilities in software can have severe consequences; however, manual vulnerability detection is costly and does not scale, especially as agentic coding frameworks increase the rate of code production. Over the last decade, a large body of research has applied machine learning machine learning to automate vulnerability detection (ML4AVD), yet self-reported performance on the most popu… ▽ More

    Submitted 7 May, 2026; v1 submitted 15 December, 2024; originally announced December 2024.

  27. arXiv:2410.17518  [pdf, other

    physics.comp-ph cs.LG

    Univariate Conditional Variational Autoencoder for Morphogenic Patterns Design in Frontal Polymerization-Based Manufacturing

    Authors: Qibang Liu, Pengfei Cai, Diab Abueidda, Sagar Vyas, Seid Koric, Rafael Gomez-Bombarelli, Philippe Geubelle

    Abstract: Under some initial and boundary conditions, the rapid reaction-thermal diffusion process taking place during frontal polymerization (FP) destabilizes the planar mode of front propagation, leading to spatially varying, complex hierarchical patterns in thermoset polymeric materials. Although modern reaction-diffusion models can predict the patterns resulting from unstable FP, the inverse design of p… ▽ More

    Submitted 31 October, 2024; v1 submitted 22 October, 2024; originally announced October 2024.

  28. Model Predictive Parkour Control of a Monoped Hopper in Dynamically Changing Environments

    Authors: Maximilian Albracht, Shivesh Kumar, Shubham Vyas, Frank Kirchner

    Abstract: A great advantage of legged robots is their ability to operate on particularly difficult and obstructed terrain, which demands dynamic, robust, and precise movements. The study of obstacle courses provides invaluable insights into the challenges legged robots face, offering a controlled environment to assess and enhance their capabilities. Traversing it with a one-legged hopper introduces intricat… ▽ More

    Submitted 26 August, 2024; originally announced August 2024.

    Comments: Published in: IEEE Robotics and Automation Letters

  29. Mitigating Deep Reinforcement Learning Backdoors in the Neural Activation Space

    Authors: Sanyam Vyas, Chris Hicks, Vasilios Mavroudis

    Abstract: This paper investigates the threat of backdoors in Deep Reinforcement Learning (DRL) agent policies and proposes a novel method for their detection at runtime. Our study focuses on elusive in-distribution backdoor triggers. Such triggers are designed to induce a deviation in the behaviour of a backdoored agent while blending into the expected data distribution to evade detection. Through experimen… ▽ More

    Submitted 21 July, 2024; originally announced July 2024.

    Comments: 11 Pages, 12 figures

    Journal ref: 2024 IEEE Security and Privacy Workshops (SPW), pp. 76-86, 2024

  30. arXiv:2406.09929  [pdf, other

    cs.RO eess.SY

    AUV trajectory optimization with hydrodynamic forces for Icy Moon Exploration

    Authors: Lukas Rust, Shubham Vyas, Bilal Wehbe

    Abstract: To explore oceans on ice-covered moons in the solar system, energy-efficient Autonomous Underwater Vehicles (AUVs) with long ranges must cover enough distance to record and collect enough data. These usually underactuated vehicles are hard to control when performing tasks such as vertical docking or the inspection of vertical walls. This paper introduces a control strategy for DeepLeng to navigate… ▽ More

    Submitted 14 June, 2024; originally announced June 2024.

    Comments: 7 pages, 8 figures

    Journal ref: In 17th Symposium on Advanced Space Technologies in Robotics and Automation, 18-20 October 2023. 2023

  31. arXiv:2404.13693  [pdf, ps, other

    eess.IV cs.CV

    Advancing Automatic Photovoltaic Defect Detection using Semi-Supervised Semantic Segmentation of Electroluminescence Images

    Authors: Abhishek Jha, Yogesh Rawat, Shruti Vyas

    Abstract: Photovoltaic (PV) systems allow us to tap into all abundant solar energy, however they require regular maintenance for high efficiency and to prevent degradation. Traditional manual health check, using Electroluminescence (EL) imaging, is expensive and logistically challenging which makes automated defect detection essential. Current automation approaches require extensive manual expert labeling,… ▽ More

    Submitted 14 July, 2025; v1 submitted 21 April, 2024; originally announced April 2024.

    Comments: 19 pages, 10 figures

  32. arXiv:2312.10788  [pdf, other

    cs.RO

    Linear Model Predictive Control for a planar free-floating platform: A comparison of binary input constraint formulations

    Authors: Franek Stark, Shubham Vyas, Georg Schildbach, Frank Kirchner

    Abstract: This work develops a first Model Predictive Control for European Space Agencies 3-dof free-floating platform. The challenges of the platform are the on/off thrusters, which cannot be actuated continuously and which are subject to certain timing constraints. This work compares penalty-term, Linear Complementarity Constraints, and classical Mixed Integer formulations in order to develop a controller… ▽ More

    Submitted 17 December, 2023; originally announced December 2023.

    Comments: 17th Symposium on Advanced Space Technologies in Robotics and Automation (ASTRA 2023)

  33. arXiv:2312.07169  [pdf, other

    cs.CV

    Semi-supervised Active Learning for Video Action Detection

    Authors: Ayush Singh, Aayush J Rana, Akash Kumar, Shruti Vyas, Yogesh Singh Rawat

    Abstract: In this work, we focus on label efficient learning for video action detection. We develop a novel semi-supervised active learning approach which utilizes both labeled as well as unlabeled data along with informative sample selection for action detection. Video action detection requires spatio-temporal localization along with classification, which poses several challenges for both active learning i… ▽ More

    Submitted 3 April, 2024; v1 submitted 12 December, 2023; originally announced December 2023.

    Comments: AAAI Conference on Artificial Intelligence, Main Technical Track (AAAI), 2024, Code: https://github.com/AKASH2907/semi-sup-active-learning

  34. arXiv:2310.07380  [pdf, other

    cs.LG cs.AI

    Histopathological Image Classification and Vulnerability Analysis using Federated Learning

    Authors: Sankalp Vyas, Amar Nath Patra, Raj Mani Shukla

    Abstract: Healthcare is one of the foremost applications of machine learning (ML). Traditionally, ML models are trained by central servers, which aggregate data from various distributed devices to forecast the results for newly generated data. This is a major concern as models can access sensitive user information, which raises privacy concerns. A federated learning (FL) approach can help address this issue… ▽ More

    Submitted 11 October, 2023; originally announced October 2023.

    Comments: Accepted in IEEE International Conference on Trust, Security and Privacy in Computing and Communications (TrustCom)

  35. AcroMonk: A Minimalist Underactuated Brachiating Robot

    Authors: Mahdi Javadi, Daniel Harnack, Paula Stocco, Shivesh Kumar, Shubham Vyas, Daniel Pizzutilo, Frank Kirchner

    Abstract: Brachiation is a dynamic, coordinated swinging maneuver of body and arms used by monkeys and apes to move between branches. As a unique underactuated mode of locomotion, it is interesting to study from a robotics perspective since it can broaden the deployment scenarios for humanoids and animaloids. While several brachiating robots of varying complexity have been proposed in the past, this paper p… ▽ More

    Submitted 15 May, 2023; originally announced May 2023.

    Comments: The open-source implementation is available at https://github.com/dfki-ric-underactuated-lab/acromonk and a video demonstration of the experiments can be accessed at https://youtu.be/FIcDNtJo9Jc}

    Journal ref: journal={IEEE Robotics and Automation Letters}, year={2023}, volume={8}, number={6}, pages={3637-3644}

  36. arXiv:2303.04926  [pdf, other

    cs.CR cs.AI

    Automated Cyber Defence: A Review

    Authors: Sanyam Vyas, John Hannay, Andrew Bolton, Professor Pete Burnap

    Abstract: Within recent times, cybercriminals have curated a variety of organised and resolute cyber attacks within a range of cyber systems, leading to consequential ramifications to private and governmental institutions. Current security-based automation and orchestrations focus on automating fixed purpose and hard-coded solutions, which are easily surpassed by modern-day cyber attacks. Research within Au… ▽ More

    Submitted 8 March, 2023; originally announced March 2023.

  37. arXiv:2207.10693  [pdf, other

    cs.RO

    Trajectory Optimization and Following for a Three Degrees of Freedom Overactuated Floating Platform

    Authors: Anton Bredenbeck, Shubham Vyas, Martin Zwick, Dorit Borrmann, Miguel Olivares-Mendez, Andreas Nüchter

    Abstract: Space robotics applications, such as Active Space Debris Removal (ASDR), require representative testing before launch. A commonly used approach to emulate the microgravity environment in space is air-bearing based platforms on flat-floors, such as the European Space Agency's Orbital Robotics and GNC Lab (ORGL). This work proposes a control architecture for a floating platform at the ORGL, equipped… ▽ More

    Submitted 21 July, 2022; originally announced July 2022.

    Comments: Accepted to IROS2022, code at https://gitlab.com/anton.bredenbeck/ff-trajectories

  38. arXiv:2207.02431  [pdf, other

    cs.CV cs.LG

    GAMa: Cross-view Video Geo-localization

    Authors: Shruti Vyas, Chen Chen, Mubarak Shah

    Abstract: The existing work in cross-view geo-localization is based on images where a ground panorama is matched to an aerial image. In this work, we focus on ground videos instead of images which provides additional contextual cues which are important for this task. There are no existing datasets for this problem, therefore we propose GAMa dataset, a large-scale dataset with ground videos and corresponding… ▽ More

    Submitted 6 July, 2022; originally announced July 2022.

    Journal ref: ECCV 2022

  39. arXiv:2207.02159  [pdf, other

    cs.CV cs.MM

    Robustness Analysis of Video-Language Models Against Visual and Language Perturbations

    Authors: Madeline C. Schiappa, Shruti Vyas, Hamid Palangi, Yogesh S. Rawat, Vibhav Vineet

    Abstract: Joint visual and language modeling on large-scale datasets has recently shown good progress in multi-modal tasks when compared to single modal learning. However, robustness of these approaches against real-world perturbations has not been studied. In this work, we perform the first extensive robustness study of video-language models against various real-world perturbations. We focus on text-to-vid… ▽ More

    Submitted 18 July, 2023; v1 submitted 5 July, 2022; originally announced July 2022.

    Comments: NeurIPS 2022 Datasets and Benchmarks Track. This projects webpage is located at https://bit.ly/3CNOly4

    Journal ref: Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (2022)

  40. arXiv:2207.01398  [pdf, other

    cs.CV eess.IV

    Large-scale Robustness Analysis of Video Action Recognition Models

    Authors: Madeline Chantry Schiappa, Naman Biyani, Prudvi Kamtam, Shruti Vyas, Hamid Palangi, Vibhav Vineet, Yogesh Rawat

    Abstract: We have seen a great progress in video action recognition in recent years. There are several models based on convolutional neural network (CNN) and some recent transformer based approaches which provide top performance on existing benchmarks. In this work, we perform a large-scale robustness analysis of these existing models for video action recognition. We focus on robustness against real-world d… ▽ More

    Submitted 7 April, 2023; v1 submitted 4 July, 2022; originally announced July 2022.

    Comments: Accepted in 2023 Conference on Computer Vision and Pattern Recognition (CVPR)

  41. arXiv:2206.03993  [pdf, other

    cs.RO

    Finding and Following Optimal Trajectories for an Overactuated Floating Robotic Platform

    Authors: Anton Bredenbeck, Shubham Vyas, Willem Suter, Martin Zwick, Dorit Borrmann, Miguel Olivares-Mendez, Andreas Nüchter

    Abstract: The recent increase in yearly spacecraft launches and the high number of planned launches have raised questions about maintaining accessibility to space for all interested parties. A key to sustaining the future of space-flight is the ability to service malfunctioning - and actively remove dysfunctional spacecraft from orbit. Robotic platforms that autonomously perform these tasks are a topic of o… ▽ More

    Submitted 19 July, 2022; v1 submitted 8 June, 2022; originally announced June 2022.

    Comments: 16th Symposium on Advanced Space Technologies in Robotics and Automation 2022

  42. arXiv:2205.08109  [pdf

    cs.LG cs.AI eess.SP

    Forecasting Solar Power Generation on the basis of Predictive and Corrective Maintenance Activities

    Authors: Soham Vyas, Yuvraj Goyal, Neel Bhatt, Sanskar Bhuwania, Hardik Patel, Shakti Mishra, Brijesh Tripathi

    Abstract: Solar energy forecasting has seen tremendous growth in the last decade using historical time series collected from a weather station, such as weather variables wind speed and direction, solar radiance, and temperature. It helps in the overall management of solar power plants. However, the solar power plant regularly requires preventive and corrective maintenance activities that further impact ener… ▽ More

    Submitted 17 May, 2022; originally announced May 2022.

  43. arXiv:2204.07892  [pdf, other

    cs.CV

    Video Action Detection: Analysing Limitations and Challenges

    Authors: Rajat Modi, Aayush Jung Rana, Akash Kumar, Praveen Tirupattur, Shruti Vyas, Yogesh Singh Rawat, Mubarak Shah

    Abstract: Beyond possessing large enough size to feed data hungry machines (eg, transformers), what attributes measure the quality of a dataset? Assuming that the definitions of such attributes do exist, how do we quantify among their relative existences? Our work attempts to explore these questions for video action detection. The task aims to spatio-temporally localize an actor and assign a relevant action… ▽ More

    Submitted 16 April, 2022; originally announced April 2022.

    Comments: CVPRW'22

  44. arXiv:2110.10899  [pdf, other

    cs.CV

    LARNet: Latent Action Representation for Human Action Synthesis

    Authors: Naman Biyani, Aayush J Rana, Shruti Vyas, Yogesh S Rawat

    Abstract: We present LARNet, a novel end-to-end approach for generating human action videos. A joint generative modeling of appearance and dynamics to synthesize a video is very challenging and therefore recent works in video synthesis have proposed to decompose these two factors. However, these methods require a driving video to model the video dynamics. In this work, we propose a generative approach inste… ▽ More

    Submitted 26 October, 2021; v1 submitted 21 October, 2021; originally announced October 2021.

    Comments: British Machine Vision Conference (BMVC) 2021

  45. arXiv:2110.07993  [pdf, other

    cs.CV

    Pose-guided Generative Adversarial Net for Novel View Action Synthesis

    Authors: Xianhang Li, Junhao Zhang, Kunchang Li, Shruti Vyas, Yogesh S Rawat

    Abstract: We focus on the problem of novel-view human action synthesis. Given an action video, the goal is to generate the same action from an unseen viewpoint. Naturally, novel view video synthesis is more challenging than image synthesis. It requires the synthesis of a sequence of realistic frames with temporal coherency. Besides, transferring the different actions to a novel target view requires awarenes… ▽ More

    Submitted 8 December, 2021; v1 submitted 15 October, 2021; originally announced October 2021.

    Comments: Accepted by WACV2022

  46. arXiv:2110.06827  [pdf, other

    cs.MM cs.CV cs.LG

    NoisyActions2M: A Multimedia Dataset for Video Understanding from Noisy Labels

    Authors: Mohit Sharma, Raj Patra, Harshal Desai, Shruti Vyas, Yogesh Rawat, Rajiv Ratn Shah

    Abstract: Deep learning has shown remarkable progress in a wide range of problems. However, efficient training of such models requires large-scale datasets, and getting annotations for such datasets can be challenging and costly. In this work, we explore the use of user-generated freely available labels from web videos for video understanding. We create a benchmark dataset consisting of around 2 million vid… ▽ More

    Submitted 13 October, 2021; originally announced October 2021.

    Comments: Accepted at ACM Multimedia Asia 2021

  47. arXiv:2107.11494  [pdf, other

    cs.CV

    TinyAction Challenge: Recognizing Real-world Low-resolution Activities in Videos

    Authors: Praveen Tirupattur, Aayush J Rana, Tushar Sangam, Shruti Vyas, Yogesh S Rawat, Mubarak Shah

    Abstract: This paper summarizes the TinyAction challenge which was organized in ActivityNet workshop at CVPR 2021. This challenge focuses on recognizing real-world low-resolution activities present in videos. Action recognition task is currently focused around classifying the actions from high-quality videos where the actors and the action is clearly visible. While various approaches have been shown effecti… ▽ More

    Submitted 23 July, 2021; originally announced July 2021.

    Comments: 8 pages. arXiv admin note: text overlap with arXiv:2007.07355

  48. arXiv:2106.03956  [pdf, other

    cs.CV

    Novel View Video Prediction Using a Dual Representation

    Authors: Sarah Shiraz, Krishna Regmi, Shruti Vyas, Yogesh S. Rawat, Mubarak Shah

    Abstract: We address the problem of novel view video prediction; given a set of input video clips from a single/multiple views, our network is able to predict the video from a novel view. The proposed approach does not require any priors and is able to predict the video from wider angular distances, upto 45 degree, as compared to the recent studies predicting small variations in viewpoint. Moreover, our met… ▽ More

    Submitted 7 June, 2021; originally announced June 2021.

    Comments: Accepted in ICIP 2021

  49. Multilingual and code-switching ASR challenges for low resource Indian languages

    Authors: Anuj Diwan, Rakesh Vaideeswaran, Sanket Shah, Ankita Singh, Srinivasa Raghavan, Shreya Khare, Vinit Unni, Saurabh Vyas, Akash Rajpuria, Chiranjeevi Yarra, Ashish Mittal, Prasanta Kumar Ghosh, Preethi Jyothi, Kalika Bali, Vivek Seshadri, Sunayana Sitaram, Samarth Bharadwaj, Jai Nanavati, Raoul Nanavati, Karthik Sankaranarayanan, Tejaswi Seeram, Basil Abraham

    Abstract: Recently, there is increasing interest in multilingual automatic speech recognition (ASR) where a speech recognition system caters to multiple low resource languages by taking advantage of low amounts of labeled corpora in multiple languages. With multilingualism becoming common in today's world, there has been increasing interest in code-switching ASR as well. In code-switching, multiple language… ▽ More

    Submitted 31 March, 2021; originally announced April 2021.

    Comments: 6 pages

  50. View-invariant action recognition

    Authors: Yogesh S Rawat, Shruti Vyas

    Abstract: Human action recognition is an important problem in computer vision. It has a wide range of applications in surveillance, human-computer interaction, augmented reality, video indexing, and retrieval. The varying pattern of spatio-temporal appearance generated by human action is key for identifying the performed action. We have seen a lot of research exploring this dynamics of spatio-temporal appea… ▽ More

    Submitted 1 September, 2020; originally announced September 2020.