Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–7 of 7 results for author: Sushko, P

Searching in archive cs. Search in all archives.
.
  1. arXiv:2605.02881  [pdf, ps, other

    cs.RO

    MolmoAct2: Action Reasoning Models for Real-world Deployment

    Authors: Haoquan Fang, Jiafei Duan, Donovan Clay, Sam Wang, Shuo Liu, Weikai Huang, Xiang Fan, Wei-Chuan Tsai, Shirui Chen, Yi Ru Wang, Shanli Xing, Jaemin Cho, Jae Sung Park, Ainaz Eftekhar, Peter Sushko, Karen Farley, Angad Wadhwa, Cole Harrison, Winson Han, Ying-Chun Lee, Eli VanderBilt, Rose Hendrix, Suveen Ellawela, Lucas Ngoo, Joyce Chai , et al. (4 additional authors not shown)

    Abstract: Vision-Language-Action (VLA) models aim to provide a single generalist controller for robots, but today's systems fall short on the criteria that matter for real-world deployment. Frontier models are closed, open-weight alternatives are tied to expensive hardware, reasoning-augmented policies pay prohibitive latency for their grounding, and fine-tuned success rates remain below the threshold for d… ▽ More

    Submitted 8 May, 2026; v1 submitted 4 May, 2026; originally announced May 2026.

    Comments: 31 pages, project page: https://allenai.org/blog/molmoact2

  2. arXiv:2604.08516  [pdf, ps, other

    cs.CV

    MolmoWeb: Open Visual Web Agent and Open Data for the Open Web

    Authors: Tanmay Gupta, Piper Wolters, Zixian Ma, Peter Sushko, Rock Yuren Pang, Diego Llanes, Yue Yang, Taira Anderson, Boyuan Zheng, Zhongzheng Ren, Harsh Trivedi, Taylor Blanton, Caleb Ouellette, Winson Han, Ali Farhadi, Ranjay Krishna

    Abstract: Web agents--autonomous systems that navigate and execute tasks on the web on behalf of users--have the potential to transform how people interact with the digital world. However, the most capable web agents today rely on proprietary models with undisclosed training data and recipes, limiting scientific understanding, reproducibility, and community-driven progress. We believe agents for the open… ▽ More

    Submitted 9 April, 2026; originally announced April 2026.

    Comments: https://allenai.org/blog/molmoweb

  3. arXiv:2512.10940  [pdf, ps, other

    cs.CV cs.AI

    OmniView: An All-Seeing Diffusion Model for 3D and 4D View Synthesis

    Authors: Xiang Fan, Sharath Girish, Vivek Ramanujan, Chaoyang Wang, Ashkan Mirzaei, Petr Sushko, Aliaksandr Siarohin, Sergey Tulyakov, Ranjay Krishna

    Abstract: Prior approaches injecting camera control into diffusion models have focused on specific subsets of 4D consistency tasks: novel view synthesis, text-to-video with camera control, image-to-video, amongst others. Therefore, these fragmented approaches are trained on disjoint slices of available 3D/4D data. We introduce OmniView, a unified framework that generalizes across a wide range of 4D consiste… ▽ More

    Submitted 21 January, 2026; v1 submitted 11 December, 2025; originally announced December 2025.

    Comments: Project page: https://snap-research.github.io/OmniView/

  4. arXiv:2511.05924  [pdf, ps, other

    cs.LG

    DiScoFormer: Plug-In Density and Score Estimation with Transformers

    Authors: Vasily Ilin, Peter Sushko, Ranjay Krishna

    Abstract: Estimating probability density and its score from samples remains a core problem in generative modeling, Bayesian inference, and kinetic theory. Existing methods are bifurcated: classical kernel density estimators (KDE) generalize across distributions but suffer from the curse of dimensionality, while modern neural score models achieve high precision but require retraining for every target distrib… ▽ More

    Submitted 2 June, 2026; v1 submitted 8 November, 2025; originally announced November 2025.

    Comments: Accepted in ICML 2026 (oral)

    MSC Class: 68T07; 62G07 ACM Class: I.2.6; G.3

  5. arXiv:2508.06905  [pdf, ps, other

    cs.CV

    MultiRef: Controllable Image Generation with Multiple Visual References

    Authors: Ruoxi Chen, Dongping Chen, Siyuan Wu, Sinan Wang, Shiyun Lang, Petr Sushko, Gaoyang Jiang, Yao Wan, Ranjay Krishna

    Abstract: Visual designers naturally draw inspiration from multiple visual references, combining diverse elements and aesthetic principles to create artwork. However, current image generative frameworks predominantly rely on single-source inputs -- either text prompts or individual reference images. In this paper, we focus on the task of controllable image generation using multiple visual references. We int… ▽ More

    Submitted 26 August, 2025; v1 submitted 9 August, 2025; originally announced August 2025.

    Comments: Accepted to ACM MM 2025 Datasets

  6. arXiv:2504.18130  [pdf, ps, other

    cs.LG math.PR math.ST

    Score-based deterministic density sampling

    Authors: Vasily Ilin, Peter Sushko, Jingwei Hu

    Abstract: We propose a deterministic sampling framework using Score-Based Transport Modeling for sampling an unnormalized target density $π$ given only its score $\nabla \log π$. Our method approximates the Wasserstein gradient flow on $\mathrm{KL}(f_t\|π)$ by learning the time-varying score $\nabla \log f_t$ on the fly using score matching. While having the same marginal distribution as Langevin dynamics,… ▽ More

    Submitted 20 October, 2025; v1 submitted 25 April, 2025; originally announced April 2025.

    Comments: 13 pages, 2 tables, 11 figures. Key words: Deterministic sampling; score-based transport modeling; Wasserstein gradient flow; relative entropy; Fisher information; annealing; neural network; neural tangent kernel

    MSC Class: Primary 65C05; 35Q84; 49Q22; Secondary 60H10; 68T07

  7. arXiv:2502.03629  [pdf, other

    cs.CV cs.AI cs.CL cs.LG

    REALEDIT: Reddit Edits As a Large-scale Empirical Dataset for Image Transformations

    Authors: Peter Sushko, Ayana Bharadwaj, Zhi Yang Lim, Vasily Ilin, Ben Caffee, Dongping Chen, Mohammadreza Salehi, Cheng-Yu Hsieh, Ranjay Krishna

    Abstract: Existing image editing models struggle to meet real-world demands. Despite excelling in academic benchmarks, they have yet to be widely adopted for real user needs. Datasets that power these models use artificial edits, lacking the scale and ecological validity necessary to address the true diversity of user requests. We introduce REALEDIT, a large-scale image editing dataset with authentic user r… ▽ More

    Submitted 28 April, 2025; v1 submitted 5 February, 2025; originally announced February 2025.

    Comments: Published at CVPR 2025