Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–9 of 9 results for author: Juvekar, A

Searching in archive cs. Search in all archives.
.
  1. arXiv:2607.21580  [pdf, ps, other

    cs.CV cs.AI

    GraphVid: Interactive Graph-Controllable Video Generation

    Authors: Vedant Shah, Onkar Susladkar, Tushar Prakash, Kiet Nguyen, Tianjio Yu, Adheesh Juvekar, Muntasir Waheed, Ismini Lourentzou

    Abstract: Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap… ▽ More

    Submitted 23 July, 2026; originally announced July 2026.

  2. arXiv:2604.08536  [pdf, ps, other

    cs.CV cs.AI

    RewardFlow: Generate Images by Optimizing What You Reward

    Authors: Onkar Susladkar, Dong-Hwan Jang, Tushar Prakash, Adheesh Juvekar, Vedant Shah, Ayush Barik, Nabeel Bashir, Muntasir Wahed, Ritish Shrirao, Ismini Lourentzou

    Abstract: We introduce RewardFlow, an inversion-free framework that steers pretrained diffusion and flow-matching models at inference time through multi-reward Langevin dynamics. RewardFlow unifies complementary differentiable rewards for semantic alignment, perceptual fidelity, localized grounding, object consistency, and human preference, and further introduces a differentiable VQA-based reward that provi… ▽ More

    Submitted 9 April, 2026; originally announced April 2026.

    Comments: CVPR 2026. Project page: https://plan-lab.github.io/rewardflow

  3. arXiv:2602.12221  [pdf, ps, other

    cs.CV

    Best of Both Worlds: Multimodal Reasoning and Generation via Unified Discrete Flow Matching

    Authors: Onkar Susladkar, Tushar Prakash, Gayatri Deshmukh, Kiet A. Nguyen, Jiaxun Zhang, Adheesh Juvekar, Tianshu Bao, Lin Chai, Sparsh Mittal, Inderjit S Dhillon, Ismini Lourentzou

    Abstract: We propose UniDFlow, a unified discrete flow-matching framework for multimodal understanding, generation, and editing. It decouples understanding and generation via task-specific low-rank adapters, avoiding objective interference and representation entanglement, while a novel reference-based multimodal preference alignment optimizes relative outcomes under identical conditioning, improving faithfu… ▽ More

    Submitted 2 June, 2026; v1 submitted 12 February, 2026; originally announced February 2026.

  4. arXiv:2601.16210  [pdf, ps, other

    cs.CV cs.AI

    PyraTok: Language-Aligned Pyramidal Tokenizer for Video Understanding and Generation

    Authors: Onkar Susladkar, Tushar Prakash, Adheesh Juvekar, Kiet A. Nguyen, Dong-Hwan Jang, Inderjit S Dhillon, Ismini Lourentzou

    Abstract: Discrete video VAEs underpin modern text-to-video generation and video understanding systems, yet existing tokenizers typically learn visual codebooks at a single scale with limited vocabularies and shallow language supervision, leading to poor cross-modal alignment and zero-shot transfer. We introduce PyraTok, a language-aligned pyramidal tokenizer that learns semantically structured discrete lat… ▽ More

    Submitted 23 February, 2026; v1 submitted 22 January, 2026; originally announced January 2026.

  5. arXiv:2506.21546  [pdf, ps, other

    cs.CV cs.AI cs.CL cs.LG

    Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination

    Authors: Xinzhuo Li, Adheesh Juvekar, Jiaxun Zhang, Xingyou Liu, Muntasir Wahed, Kiet A. Nguyen, Yifan Shen, Tianjiao Yu, Ismini Lourentzou

    Abstract: Segmentation Vision-Language Models (VLMs) have significantly advanced grounded visual understanding, yet they remain prone to pixel-grounding hallucinations, producing masks for incorrect objects or for objects that are entirely absent. Existing evaluations rely almost entirely on text- or label-based perturbations, which check only whether the predicted mask matches the queried label. Such evalu… ▽ More

    Submitted 23 April, 2026; v1 submitted 26 June, 2025; originally announced June 2025.

    Comments: Project webpage: https://plan-lab.github.io/hallusegbench/

  6. arXiv:2503.10628  [pdf, other

    cs.AI cs.LG

    Uncertainty in Action: Confidence Elicitation in Embodied Agents

    Authors: Tianjiao Yu, Vedant Shah, Muntasir Wahed, Kiet A. Nguyen, Adheesh Juvekar, Tal August, Ismini Lourentzou

    Abstract: Expressing confidence is challenging for embodied agents navigating dynamic multimodal environments, where uncertainty arises from both perception and decision-making processes. We present the first work investigating embodied confidence elicitation in open-ended multimodal environments. We introduce Elicitation Policies, which structure confidence assessment across inductive, deductive, and abduc… ▽ More

    Submitted 13 March, 2025; originally announced March 2025.

    Comments: Project page: https://plan-lab.github.io/ece/

  7. arXiv:2412.19331  [pdf, other

    cs.CV cs.AI cs.LG

    CALICO: Part-Focused Semantic Co-Segmentation with Large Vision-Language Models

    Authors: Kiet A. Nguyen, Adheesh Juvekar, Tianjiao Yu, Muntasir Wahed, Ismini Lourentzou

    Abstract: Recent advances in Large Vision-Language Models (LVLMs) have enabled general-purpose vision tasks through visual instruction tuning. While existing LVLMs can generate segmentation masks from text prompts for single images, they struggle with segmentation-grounded reasoning across images, especially at finer granularities such as object parts. In this paper, we introduce the new task of part-focuse… ▽ More

    Submitted 3 April, 2025; v1 submitted 26 December, 2024; originally announced December 2024.

    Comments: Accepted to CVPR 2025. Project page: https://plan-lab.github.io/calico/

  8. arXiv:2412.15209  [pdf, ps, other

    cs.CV cs.AI cs.LG

    PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation

    Authors: Muntasir Wahed, Kiet A. Nguyen, Adheesh Sunil Juvekar, Xinzhuo Li, Xiaona Zhou, Vedant Shah, Tianjiao Yu, Pinar Yanardag, Ismini Lourentzou

    Abstract: Despite significant advancements in Large Vision-Language Models (LVLMs)' capabilities, existing pixel-grounding models operate in single-image settings, limiting their ability to perform detailed, fine-grained comparisons across multiple images. Conversely, current multi-image understanding models lack pixel-level grounding. Our work addresses this gap by introducing the task of multi-image pixel… ▽ More

    Submitted 27 November, 2025; v1 submitted 19 December, 2024; originally announced December 2024.

    Comments: Project page: https://plan-lab.github.io/prima

  9. arXiv:1811.06194  [pdf

    cs.CV

    Face Verification and Forgery Detection for Ophthalmic Surgery Images

    Authors: Kaushal Bhogale, Nishant Shankar, Adheesh Juvekar, Asutosh Padhi

    Abstract: Although modern face verification systems are accessible and accurate, they are not always robust to pose variance and occlusions. Moreover, accurate models require a large amount of data to train. We structure our experiments to operate on small amounts of data obtained from an NGO that funds ophthalmic surgeries. We set up our face verification task as that of verifying pre-operation and post-op… ▽ More

    Submitted 15 November, 2018; originally announced November 2018.