Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 53 results for author: Benaim, S

Searching in archive cs. Search in all archives.
.
  1. arXiv:2607.05376  [pdf, ps, other

    cs.CV cs.GR

    MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing

    Authors: Gal Fiebelman, Hadar Averbuch-Elor, Sagie Benaim

    Abstract: Recent advances in video diffusion models have enabled either long single-view generation through temporal autoregression, or short multi-view synthesis through bidirectional attention. However, generating long, multi-view consistent videos of dynamic scenes remains unsolved. In this work, we present MV-Forcing, a framework that composes temporal and view-wise autoregression within a single diffus… ▽ More

    Submitted 6 July, 2026; originally announced July 2026.

    Comments: Accepted to ECCV 2026. Project webpage: https://galfiebelman.github.io/mv-forcing/

  2. arXiv:2607.03875  [pdf, ps, other

    cs.CV

    MACRO: Training-free Multi-plane Attention for Closeup Render Optimization

    Authors: Nitzan Hodos, Roy Amoyal, Lior Fritz, Ianir Ideses, Sagie Benaim, Netalee Efrat

    Abstract: Close-up rendering, zooming into a scene well beyond any training camera, is important for virtual production and interactive 3D content, yet remains an open challenge. 3D Gaussian splatting (3DGS) enables high-fidelity, real-time novel view synthesis, but its rendering quality degrades at close range. Recent diffusion-based methods that enhance the rendering by conditioning on reference images fr… ▽ More

    Submitted 4 July, 2026; originally announced July 2026.

    Comments: Project page: https://nitzanhod.github.io/MACRO

  3. arXiv:2607.00861  [pdf, ps, other

    cs.CV cs.GR

    TrajLoc: Trajectory-Attention Localization for Multi-Object Motion Control

    Authors: Omer Sela, Inbar Huberman-Spiegelglas, Michael Rotman, Sagie Benaim, Avi Ben-Cohen

    Abstract: Controlling the motion of multiple objects in image-to-video (I2V) generation requires preserving object identities while enforcing adherence to distinct target trajectories. This becomes particularly challenging as the number of objects increases and their paths intersect or occlude one another. Existing approaches entangle multiple trajectories within a shared, dense conditioning signal, making… ▽ More

    Submitted 1 July, 2026; originally announced July 2026.

    Comments: Project page: https://sela-omer.github.io/traj-loc/ Code: https://github.com/Sela-Omer/traj-loc

    ACM Class: I.3.3; I.2.10

  4. arXiv:2606.32033  [pdf, ps, other

    cs.CV

    SpheRoPE: Zero-Shot Optimization-Free 360 Panorama Generation with Spherical RoPE

    Authors: Or Hirschorn, Aaron Olender, Eli Alshan, Ianir Ideses, Lior Fritz, Sagie Benaim

    Abstract: We present a zero-shot, training-free and optimization-free framework for generating 360 panoramic images and videos by directly injecting spherical priors into pre-trained diffusion transformers. Existing methods either rely on costly fine-tuning on scarce panoramic data that limits generalization, or leverage multi-step optimization that incurs prohibitive inference latency. We observe that cont… ▽ More

    Submitted 30 June, 2026; originally announced June 2026.

  5. arXiv:2605.30332  [pdf, ps, other

    cs.CV

    Colored Noise Diffusion Sampling

    Authors: Hadar Davidson, Noam Issachar, Sagie Benaim

    Abstract: Diffusion models achieve state-of-the-art image synthesis, with their generative trajectories fundamentally exhibiting a spectral bias, resolving low-frequency global structures early and high-frequency fine details later. Conventional stochastic differential equation (SDE) solvers fail to account for this dynamic, naively injecting uniform white noise throughout the entire process and misusing th… ▽ More

    Submitted 28 May, 2026; originally announced May 2026.

  6. arXiv:2605.30268  [pdf, ps, other

    cs.CV cs.AI

    PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions

    Authors: Omer Benishu, Gal Fiebelman, Sagie Benaim

    Abstract: We address the task of generating physically accurate and visually faithful 4D Human-Object Interaction (HOI). Given a static 3D human and target object represented as 3D Gaussian Splats (3DGS), our goal is to synthesize dynamic scenes where the human actively engages with the object through actions, such as punching or kicking, in accordance with a given input text. To this end, we introduce PhyG… ▽ More

    Submitted 28 May, 2026; originally announced May 2026.

  7. arXiv:2604.15284  [pdf, ps, other

    cs.CV

    GlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene Tokens

    Authors: Roni Itkin, Noam Issachar, Yehonatan Keypur, Xingyu Chen, Anpei Chen, Sagie Benaim

    Abstract: The efficient spatial allocation of primitives serves as the foundation of 3D Gaussian Splatting, as it directly dictates the synergy between representation compactness, reconstruction speed, and rendering fidelity. Previous solutions, whether based on iterative optimization or feed-forward inference, suffer from significant trade-offs between these goals, mainly due to the reliance on local, heur… ▽ More

    Submitted 17 April, 2026; v1 submitted 16 April, 2026; originally announced April 2026.

  8. arXiv:2602.09532  [pdf, ps, other

    cs.CV

    RAD: Retrieval-Augmented Monocular Metric Depth Estimation for Underrepresented Classes

    Authors: Michael Baltaxe, Dan Levi, Sagie Benaim

    Abstract: Monocular Metric Depth Estimation (MMDE) is essential for physically intelligent systems, yet accurate depth estimation for underrepresented classes in complex scenes remains a persistent challenge. To address this, we propose RAD, a retrieval-augmented framework that approximates the benefits of multi-view stereo by utilizing retrieved neighbors as structural geometric proxies. Our method first e… ▽ More

    Submitted 4 April, 2026; v1 submitted 10 February, 2026; originally announced February 2026.

  9. arXiv:2602.09146  [pdf, ps, other

    cs.CV

    SemanticMoments: Training-Free Motion Similarity via Third Moment Features

    Authors: Saar Huberman, Kfir Goldberg, Or Patashnik, Sagie Benaim, Ron Mokady

    Abstract: Retrieving videos based on semantic motion is a fundamental, yet unsolved, problem. Existing video representation approaches overly rely on static appearance and scene context rather than motion dynamics, a bias inherited from their training data and objectives. Conversely, traditional motion-centric inputs like optical flow lack the semantic grounding needed to understand high-level motion. To de… ▽ More

    Submitted 9 February, 2026; originally announced February 2026.

  10. arXiv:2602.06032  [pdf, ps, other

    cs.CV

    Splat and Distill: Augmenting Teachers with Feed-Forward 3D Reconstruction For 3D-Aware Distillation

    Authors: David Shavin, Sagie Benaim

    Abstract: Vision Foundation Models (VFMs) have achieved remarkable success when applied to various downstream 2D tasks. Despite their effectiveness, they often exhibit a critical lack of 3D awareness. To this end, we introduce Splat and Distill, a framework that instills robust 3D awareness into 2D VFMs by augmenting the teacher model with a fast, feed-forward 3D reconstruction pipeline. Given 2D features p… ▽ More

    Submitted 11 February, 2026; v1 submitted 5 February, 2026; originally announced February 2026.

    Comments: Accepted to ICLR 2026

    MSC Class: 68U10 (Image processing); 68T45 (Computer vision) ACM Class: I.2.10; I.4.8

    Journal ref: ICLR 2026

  11. arXiv:2512.07807  [pdf, ps, other

    cs.CV cs.GR

    Lang3D-XL: Language Embedded 3D Gaussians for Large-scale Scenes

    Authors: Shai Krakovsky, Gal Fiebelman, Sagie Benaim, Hadar Averbuch-Elor

    Abstract: Embedding a language field in a 3D representation enables richer semantic understanding of spatial environments by linking geometry with descriptive meaning. This allows for a more intuitive human-computer interaction, enabling querying or editing scenes using natural language, and could potentially improve tasks like scene retrieval, navigation, and multimodal reasoning. While such capabilities c… ▽ More

    Submitted 8 December, 2025; originally announced December 2025.

    Comments: Accepted to SIGGRAPH Asia 2025. Project webpage: https://tau-vailab.github.io/Lang3D-XL

  12. arXiv:2510.20766  [pdf, ps, other

    cs.CV

    DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion

    Authors: Noam Issachar, Guy Yariv, Sagie Benaim, Yossi Adi, Dani Lischinski, Raanan Fattal

    Abstract: Diffusion Transformer models can generate images with remarkable fidelity and detail, yet training them at ultra-high resolutions remains extremely costly due to the self-attention mechanism's quadratic scaling with the number of image tokens. In this paper, we introduce Dynamic Position Extrapolation (DyPE), a novel, training-free method that enables pre-trained diffusion transformers to synthesi… ▽ More

    Submitted 29 January, 2026; v1 submitted 23 October, 2025; originally announced October 2025.

  13. arXiv:2508.16577  [pdf, ps, other

    cs.CV cs.AI

    MV-RAG: Retrieval Augmented Multiview Diffusion

    Authors: Yosef Dayani, Omer Benishu, Sagie Benaim

    Abstract: Text-to-3D generation approaches have advanced significantly by leveraging pretrained 2D diffusion priors, producing high-quality and 3D-consistent outputs. However, they often fail to produce out-of-domain (OOD) or rare concepts, yielding inconsistent or inaccurate results. To this end, we propose MV-RAG, a novel text-to-3D pipeline that first retrieves relevant 2D images from a large in-the-wild… ▽ More

    Submitted 22 August, 2025; originally announced August 2025.

    Comments: Project page: https://yosefdayani.github.io/MV-RAG

  14. arXiv:2504.05296  [pdf, ps, other

    cs.GR cs.CV

    Let it Snow! Animating 3D Gaussian Scenes with Dynamic Weather Effects via Physics-Guided Score Distillation

    Authors: Gal Fiebelman, Hadar Averbuch-Elor, Sagie Benaim

    Abstract: 3D Gaussian Splatting has recently enabled fast and photorealistic reconstruction of static 3D scenes. However, dynamic editing of such scenes remains a significant challenge. We introduce a novel framework, Physics-Guided Score Distillation, to address a fundamental conflict: physics simulation provides a strong motion prior that is insufficient for photorealism , while video-based Score Distilla… ▽ More

    Submitted 25 March, 2026; v1 submitted 7 April, 2025; originally announced April 2025.

    Comments: Accepted to CVPR 2026. Project webpage: https://galfiebelman.github.io/let-it-snow/

  15. arXiv:2503.09601  [pdf, other

    cs.CV

    RewardSDS: Aligning Score Distillation via Reward-Weighted Sampling

    Authors: Itay Chachy, Guy Yariv, Sagie Benaim

    Abstract: Score Distillation Sampling (SDS) has emerged as an effective technique for leveraging 2D diffusion priors for tasks such as text-to-3D generation. While powerful, SDS struggles with achieving fine-grained alignment to user intent. To overcome this, we introduce RewardSDS, a novel approach that weights noise samples based on alignment scores from a reward model, producing a weighted SDS loss. This… ▽ More

    Submitted 13 March, 2025; v1 submitted 12 March, 2025; originally announced March 2025.

  16. arXiv:2502.14789  [pdf, other

    cs.CV

    Structurally Disentangled Feature Fields Distillation for 3D Understanding and Editing

    Authors: Yoel Levy, David Shavin, Itai Lang, Sagie Benaim

    Abstract: Recent work has demonstrated the ability to leverage or distill pre-trained 2D features obtained using large pre-trained 2D models into 3D features, enabling impressive 3D editing and understanding capabilities using only 2D supervision. Although impressive, models assume that 3D features are captured using a single feature field and often make a simplifying assumption that features are view-indep… ▽ More

    Submitted 20 February, 2025; originally announced February 2025.

  17. arXiv:2502.09611  [pdf, other

    cs.LG cs.CV

    Designing a Conditional Prior Distribution for Flow-Based Generative Models

    Authors: Noam Issachar, Mohammad Salama, Raanan Fattal, Sagie Benaim

    Abstract: Flow-based generative models have recently shown impressive performance for conditional generation tasks, such as text-to-image generation. However, current methods transform a general unimodal noise distribution to a specific mode of the target data distribution. As such, every point in the initial source distribution can be mapped to every point in the target distribution, resulting in long aver… ▽ More

    Submitted 13 February, 2025; originally announced February 2025.

  18. arXiv:2501.03059  [pdf, other

    cs.CV cs.AI cs.LG

    Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation

    Authors: Guy Yariv, Yuval Kirstain, Amit Zohar, Shelly Sheynin, Yaniv Taigman, Yossi Adi, Sagie Benaim, Adam Polyak

    Abstract: We consider the task of Image-to-Video (I2V) generation, which involves transforming static images into realistic video sequences based on a textual description. While recent advancements produce photorealistic outputs, they frequently struggle to create videos with accurate and consistent object motion, especially in multi-object scenarios. To address these limitations, we propose a two-stage com… ▽ More

    Submitted 6 January, 2025; originally announced January 2025.

  19. arXiv:2410.09792  [pdf, other

    cs.CV

    Generating Intermediate Representations for Compositional Text-To-Image Generation

    Authors: Ran Galun, Sagie Benaim

    Abstract: Text-to-image diffusion models have demonstrated an impressive ability to produce high-quality outputs. However, they often struggle to accurately follow fine-grained spatial information in an input text. To this end, we propose a compositional approach for text-to-image generation based on two stages. In the first stage, we design a diffusion-based generative model to produce one or more aligned… ▽ More

    Submitted 20 October, 2024; v1 submitted 13 October, 2024; originally announced October 2024.

    Comments: Accepted to NeurIPS 2024 Workshop on Compositional Learning: Perspectives, Methods, and Paths Forward

  20. arXiv:2406.13621  [pdf, ps, other

    cs.CL cs.CV cs.LG

    LaMI: Augmenting Large Language Models via Late Multi-Image Fusion

    Authors: Guy Yariv, Idan Schwartz, Yossi Adi, Sagie Benaim

    Abstract: Commonsense reasoning often requires both textual and visual knowledge, yet Large Language Models (LLMs) trained solely on text lack visual grounding (e.g., "what color is an emperor penguin's belly?"). Visual Language Models (VLMs) perform better on visually grounded tasks but face two limitations: (i) often reduced performance on text-only commonsense reasoning compared to text-trained LLMs, and… ▽ More

    Submitted 11 April, 2026; v1 submitted 19 June, 2024; originally announced June 2024.

    Comments: Accepted to ACL 2026

  21. arXiv:2406.04332  [pdf, other

    cs.CV cs.LG

    Coarse-To-Fine Tensor Trains for Compact Visual Representations

    Authors: Sebastian Loeschcke, Dan Wang, Christian Leth-Espensen, Serge Belongie, Michael J. Kastoryano, Sagie Benaim

    Abstract: The ability to learn compact, high-quality, and easy-to-optimize representations for visual data is paramount to many applications such as novel view synthesis and 3D reconstruction. Recent work has shown substantial success in using tensor networks to design such compact and high-quality representations. However, the ability to optimize tensor-based representations, and in particular, the highly… ▽ More

    Submitted 6 June, 2024; originally announced June 2024.

    Comments: Project webpage: https://sebulo.github.io/PuTT_website/

  22. arXiv:2405.19321  [pdf, other

    cs.CV

    DGD: Dynamic 3D Gaussians Distillation

    Authors: Isaac Labe, Noam Issachar, Itai Lang, Sagie Benaim

    Abstract: We tackle the task of learning dynamic 3D semantic radiance fields given a single monocular video as input. Our learned semantic radiance field captures per-point semantics as well as color and geometric properties for a dynamic 3D scene, enabling the generation of novel views and their corresponding semantics. This enables the segmentation and tracking of a diverse set of 3D semantic entities, sp… ▽ More

    Submitted 29 May, 2024; originally announced May 2024.

  23. arXiv:2310.19080  [pdf, other

    cs.CV

    Reward Finetuning for Faster and More Accurate Unsupervised Object Discovery

    Authors: Katie Z Luo, Zhenzhen Liu, Xiangyu Chen, Yurong You, Sagie Benaim, Cheng Perng Phoo, Mark Campbell, Wen Sun, Bharath Hariharan, Kilian Q. Weinberger

    Abstract: Recent advances in machine learning have shown that Reinforcement Learning from Human Feedback (RLHF) can improve machine learning models and align them with human preferences. Although very successful for Large Language Models (LLMs), these advancements have not had a comparable impact in research for autonomous vehicles -- where alignment with human expectations can be imperative. In this paper,… ▽ More

    Submitted 5 November, 2023; v1 submitted 29 October, 2023; originally announced October 2023.

  24. arXiv:2309.16429  [pdf, other

    cs.LG cs.AI

    Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation

    Authors: Guy Yariv, Itai Gat, Sagie Benaim, Lior Wolf, Idan Schwartz, Yossi Adi

    Abstract: We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio: globally, the input audio is semantically associated with the entire output video, and temporally, each segment of the input audio is associated with a corresp… ▽ More

    Submitted 28 September, 2023; originally announced September 2023.

    Comments: 9 pages, 6 figures

  25. arXiv:2303.17155  [pdf, other

    cs.CV cs.AI

    Discriminative Class Tokens for Text-to-Image Diffusion Models

    Authors: Idan Schwartz, Vésteinn Snæbjarnarson, Hila Chefer, Ryan Cotterell, Serge Belongie, Lior Wolf, Sagie Benaim

    Abstract: Recent advances in text-to-image diffusion models have enabled the generation of diverse and high-quality images. While impressive, the images often fall short of depicting subtle details and are susceptible to errors due to ambiguity in the input text. One way of alleviating these issues is to train diffusion models on class-labeled datasets. This approach has two disadvantages: (i) supervised da… ▽ More

    Submitted 9 January, 2025; v1 submitted 30 March, 2023; originally announced March 2023.

    Comments: ICCV 2023

  26. arXiv:2302.04862  [pdf, other

    cs.CV cs.LG

    Polynomial Neural Fields for Subband Decomposition and Manipulation

    Authors: Guandao Yang, Sagie Benaim, Varun Jampani, Kyle Genova, Jonathan T. Barron, Thomas Funkhouser, Bharath Hariharan, Serge Belongie

    Abstract: Neural fields have emerged as a new paradigm for representing signals, thanks to their ability to do it compactly while being easy to optimize. In most applications, however, neural fields are treated like black boxes, which precludes many signal manipulation tasks. In this paper, we propose a new class of neural fields called polynomial neural fields (PNFs). The key advantage of a PNF is that it… ▽ More

    Submitted 9 February, 2023; originally announced February 2023.

    Comments: Accepted to NeurIPS 2022

  27. arXiv:2211.09782  [pdf, other

    cs.CV cs.CR cs.LG

    Assessing Neural Network Robustness via Adversarial Pivotal Tuning

    Authors: Peter Ebert Christensen, Vésteinn Snæbjarnarson, Andrea Dittadi, Serge Belongie, Sagie Benaim

    Abstract: The robustness of image classifiers is essential to their deployment in the real world. The ability to assess this resilience to manipulations or deviations from the training data is thus crucial. These modifications have traditionally consisted of minimal changes that still manage to fool classifiers, and modern approaches are increasingly robust to them. Semantic manipulations that modify elemen… ▽ More

    Submitted 6 January, 2024; v1 submitted 17 November, 2022; originally announced November 2022.

    Comments: Major changes include new experiments in Table 1 on page 5 and Table 2-4 on page 6, new figure 5 on page 8. Paper accepted at WACV (oral)

  28. arXiv:2207.11226  [pdf, other

    cs.CV cs.LG

    FewGAN: Generating from the Joint Distribution of a Few Images

    Authors: Lior Ben-Moshe, Sagie Benaim, Lior Wolf

    Abstract: We introduce FewGAN, a generative model for generating novel, high-quality and diverse images whose patch distribution lies in the joint patch distribution of a small number of N>1 training samples. The method is, in essence, a hierarchical patch-GAN that applies quantization at the first coarse scale, in a similar fashion to VQ-GAN, followed by a pyramid of residual fully convolutional GANs at fi… ▽ More

    Submitted 18 July, 2022; originally announced July 2022.

  29. arXiv:2206.12396  [pdf, other

    cs.CV

    Text-Driven Stylization of Video Objects

    Authors: Sebastian Loeschcke, Serge Belongie, Sagie Benaim

    Abstract: We tackle the task of stylizing video objects in an intuitive and semantic manner following a user-specified text prompt. This is a challenging task as the resulting video must satisfy multiple properties: (1) it has to be temporally consistent and avoid jittering or similar artifacts, (2) the resulting stylization must preserve both the global semantics of the object and its fine-grained details,… ▽ More

    Submitted 27 June, 2022; v1 submitted 24 June, 2022; originally announced June 2022.

  30. arXiv:2206.02776  [pdf, other

    cs.CV

    Volumetric Disentanglement for 3D Scene Manipulation

    Authors: Sagie Benaim, Frederik Warburg, Peter Ebert Christensen, Serge Belongie

    Abstract: Recently, advances in differential volumetric rendering enabled significant breakthroughs in the photo-realistic and fine-detailed reconstruction of complex 3D scenes, which is key for many virtual reality applications. However, in the context of augmented reality, one may also wish to effect semantic manipulations or augmentations of objects within a scene. To this end, we propose a volumetric fr… ▽ More

    Submitted 6 June, 2022; originally announced June 2022.

  31. arXiv:2205.02673  [pdf, other

    cs.LG cs.AI

    On Disentangled and Locally Fair Representations

    Authors: Yaron Gurovich, Sagie Benaim, Lior Wolf

    Abstract: We study the problem of performing classification in a manner that is fair for sensitive groups, such as race and gender. This problem is tackled through the lens of disentangled and locally fair representations. We learn a locally fair representation, such that, under the learned representation, the neighborhood of each sample is balanced in terms of the sensitive attribute. For instance, when a… ▽ More

    Submitted 5 May, 2022; originally announced May 2022.

  32. arXiv:2112.05080  [pdf, other

    cs.CV cs.AI

    Locally Shifted Attention With Early Global Integration

    Authors: Shelly Sheynin, Sagie Benaim, Adam Polyak, Lior Wolf

    Abstract: Recent work has shown the potential of transformers for computer vision applications. An image is first partitioned into patches, which are then used as input tokens for the attention mechanism. Due to the expensive quadratic cost of the attention mechanism, either a large patch size is used, resulting in coarse-grained global interactions, or alternatively, attention is applied only on a local re… ▽ More

    Submitted 22 December, 2021; v1 submitted 9 December, 2021; originally announced December 2021.

  33. arXiv:2112.03221  [pdf, other

    cs.CV cs.CL cs.GR

    Text2Mesh: Text-Driven Neural Stylization for Meshes

    Authors: Oscar Michel, Roi Bar-On, Richard Liu, Sagie Benaim, Rana Hanocka

    Abstract: In this work, we develop intuitive controls for editing the style of 3D objects. Our framework, Text2Mesh, stylizes a 3D mesh by predicting color and local geometric details which conform to a target text prompt. We consider a disentangled representation of a 3D object using a fixed mesh input (content) coupled with a learned neural network, which we term neural style field network. In order to mo… ▽ More

    Submitted 6 December, 2021; originally announced December 2021.

    Comments: project page: https://threedle.github.io/text2mesh/

  34. arXiv:2110.12427  [pdf, other

    cs.CV

    Image-Based CLIP-Guided Essence Transfer

    Authors: Hila Chefer, Sagie Benaim, Roni Paiss, Lior Wolf

    Abstract: We make the distinction between (i) style transfer, in which a source image is manipulated to match the textures and colors of a target image, and (ii) essence transfer, in which one edits the source image to include high-level semantic attributes from the target. Crucially, the semantic attributes that constitute the essence of an image may differ from image to image. Our blending operator combin… ▽ More

    Submitted 11 October, 2022; v1 submitted 24 October, 2021; originally announced October 2021.

    Comments: To appear in ECCV'22

  35. arXiv:2106.09679  [pdf, other

    cs.CV

    JOKR: Joint Keypoint Representation for Unsupervised Cross-Domain Motion Retargeting

    Authors: Ron Mokady, Rotem Tzaban, Sagie Benaim, Amit H. Bermano, Daniel Cohen-Or

    Abstract: The task of unsupervised motion retargeting in videos has seen substantial advancements through the use of deep neural networks. While early works concentrated on specific object priors such as a human face or body, recent work considered the unsupervised case. When the source and target videos, however, are of different shapes, current methods fail. To alleviate this problem, we introduce JOKR -… ▽ More

    Submitted 17 June, 2021; originally announced June 2021.

  36. arXiv:2105.14609  [pdf, other

    cs.CV

    Identity and Attribute Preserving Thumbnail Upscaling

    Authors: Noam Gat, Sagie Benaim, Lior Wolf

    Abstract: We consider the task of upscaling a low resolution thumbnail image of a person, to a higher resolution image, which preserves the person's identity and other attributes. Since the thumbnail image is of low resolution, many higher resolution versions exist. Previous approaches produce solutions where the person's identity is not preserved, or biased solutions, such as predominantly Caucasian faces.… ▽ More

    Submitted 30 May, 2021; originally announced May 2021.

    Comments: ICIP 2021

  37. arXiv:2104.14535  [pdf, other

    cs.CV cs.LG

    A Hierarchical Transformation-Discriminating Generative Model for Few Shot Anomaly Detection

    Authors: Shelly Sheynin, Sagie Benaim, Lior Wolf

    Abstract: Anomaly detection, the task of identifying unusual samples in data, often relies on a large set of training samples. In this work, we consider the setting of few-shot anomaly detection in images, where only a few images are given at training. We devise a hierarchical generative model that captures the multi-scale patch distribution of each training image. We further enhance the representation of o… ▽ More

    Submitted 29 April, 2021; originally announced April 2021.

  38. arXiv:2010.05785  [pdf, other

    cs.CV

    Permuted AdaIN: Reducing the Bias Towards Global Statistics in Image Classification

    Authors: Oren Nuriel, Sagie Benaim, Lior Wolf

    Abstract: Recent work has shown that convolutional neural network classifiers overly rely on texture at the expense of shape cues. We make a similar but different distinction between shape and local image cues, on the one hand, and global image statistics, on the other. Our method, called Permuted Adaptive Instance Normalization (pAdaIN), reduces the representation of global statistics in the hidden layers… ▽ More

    Submitted 23 June, 2021; v1 submitted 9 October, 2020; originally announced October 2020.

    Comments: 8 pages, 3 figures

    ACM Class: I.4.0

  39. arXiv:2006.12226  [pdf, other

    cs.CV cs.LG

    Hierarchical Patch VAE-GAN: Generating Diverse Videos from a Single Sample

    Authors: Shir Gur, Sagie Benaim, Lior Wolf

    Abstract: We consider the task of generating diverse and novel videos from a single video sample. Recently, new hierarchical patch-GAN based approaches were proposed for generating diverse images, given only a single sample at training time. Moving to videos, these approaches fail to generate diverse samples, and often collapse into generating samples similar to the training video. We introduce a novel patc… ▽ More

    Submitted 22 October, 2020; v1 submitted 22 June, 2020; originally announced June 2020.

  40. arXiv:2004.12361  [pdf, other

    cs.CV cs.LG eess.IV

    Evaluation Metrics for Conditional Image Generation

    Authors: Yaniv Benny, Tomer Galanti, Sagie Benaim, Lior Wolf

    Abstract: We present two new metrics for evaluating generative models in the class-conditional image generation setting. These metrics are obtained by generalizing the two most popular unconditional metrics: the Inception Score (IS) and the Fre'chet Inception Distance (FID). A theoretical analysis shows the motivation behind each proposed metric and links the novel metrics to their unconditional counterpart… ▽ More

    Submitted 8 February, 2021; v1 submitted 26 April, 2020; originally announced April 2020.

    Comments: To be published in "INTERNATIONAL JOURNAL OF COMPUTER VISION"

  41. arXiv:2004.06130  [pdf, other

    cs.CV

    SpeedNet: Learning the Speediness in Videos

    Authors: Sagie Benaim, Ariel Ephrat, Oran Lang, Inbar Mosseri, William T. Freeman, Michael Rubinstein, Michal Irani, Tali Dekel

    Abstract: We wish to automatically predict the "speediness" of moving objects in videos---whether they move faster, at, or slower than their "natural" speed. The core component in our approach is SpeedNet---a novel deep network trained to detect if a video is playing at normal rate, or if it is sped up. SpeedNet is trained on a large corpus of natural videos in a self-supervised manner, without requiring an… ▽ More

    Submitted 26 July, 2020; v1 submitted 13 April, 2020; originally announced April 2020.

    Comments: Accepted to CVPR 2020 (oral). Project webpage: http://speednet-cvpr20.github.io

  42. Structural-analogy from a Single Image Pair

    Authors: Sagie Benaim, Ron Mokady, Amit Bermano, Daniel Cohen-Or, Lior Wolf

    Abstract: The task of unsupervised image-to-image translation has seen substantial advancements in recent years through the use of deep neural networks. Typically, the proposed solutions learn the characterizing distribution of two large, unpaired collections of images, and are able to alter the appearance of a given image, while keeping its geometry intact. In this paper, we explore the capabilities of neu… ▽ More

    Submitted 6 January, 2021; v1 submitted 5 April, 2020; originally announced April 2020.

    Comments: Published in 'Computer Graphics Forum'

  43. arXiv:2001.05026  [pdf, other

    cs.LG stat.ML

    Unsupervised Learning of the Set of Local Maxima

    Authors: Lior Wolf, Sagie Benaim, Tomer Galanti

    Abstract: This paper describes a new form of unsupervised learning, whose input is a set of unlabeled points that are assumed to be local maxima of an unknown value function v in an unknown subset of the vector space. Two functions are learned: (i) a set indicator c, which is a binary classifier, and (ii) a comparator function h that given two nearby samples, predicts which sample has the higher value of th… ▽ More

    Submitted 14 January, 2020; originally announced January 2020.

    Comments: ICLR 2019

  44. arXiv:2001.05017  [pdf, other

    cs.CV cs.LG

    Emerging Disentanglement in Auto-Encoder Based Unsupervised Image Content Transfer

    Authors: Ori Press, Tomer Galanti, Sagie Benaim, Lior Wolf

    Abstract: We study the problem of learning to map, in an unsupervised way, between domains A and B, such that the samples b in B contain all the information that exists in samples a in A and some additional information. For example, ignoring occlusions, B can be people with glasses, A people without, and the glasses, would be the added information. When mapping a sample a from the first domain to the other… ▽ More

    Submitted 14 January, 2020; originally announced January 2020.

    Journal ref: ICLR 2019

  45. arXiv:1908.11628  [pdf, other

    cs.CV

    Domain Intersection and Domain Difference

    Authors: Sagie Benaim, Michael Khaitov, Tomer Galanti, Lior Wolf

    Abstract: We present a method for recovering the shared content between two visual domains as well as the content that is unique to each domain. This allows us to map from one domain to the other, in a way in which the content that is specific for the first domain is removed and the content that is specific for the second is imported from any image in the second domain. In addition, our method enables gener… ▽ More

    Submitted 30 August, 2019; originally announced August 2019.

    Journal ref: ICCV 2019

  46. arXiv:1906.06558  [pdf, other

    cs.CV

    Mask Based Unsupervised Content Transfer

    Authors: Ron Mokady, Sagie Benaim, Lior Wolf, Amit Bermano

    Abstract: We consider the problem of translating, in an unsupervised manner, between two domains where one contains some additional information compared to the other. The proposed method disentangles the common and separate parts of these domains and, through the generation of a mask, focuses the attention of the underlying network to the desired augmentation alone, without wastefully reconstructing the ent… ▽ More

    Submitted 13 January, 2020; v1 submitted 15 June, 2019; originally announced June 2019.

  47. arXiv:1812.06087  [pdf, other

    cs.SD cs.LG eess.AS stat.ML

    Semi-Supervised Monaural Singing Voice Separation With a Masking Network Trained on Synthetic Mixtures

    Authors: Michael Michelashvili, Sagie Benaim, Lior Wolf

    Abstract: We study the problem of semi-supervised singing voice separation, in which the training data contains a set of samples of mixed music (singing and instrumental) and an unmatched set of instrumental music. Our solution employs a single mapping function g, which, applied to a mixed sample, recovers the underlying instrumental music, and, applied to an instrumental sample, returns the same sample. Th… ▽ More

    Submitted 6 May, 2019; v1 submitted 14 December, 2018; originally announced December 2018.

  48. arXiv:1807.08501  [pdf, other

    cs.LG stat.ML

    Risk Bounds for Unsupervised Cross-Domain Mapping with IPMs

    Authors: Tomer Galanti, Sagie Benaim, Lior Wolf

    Abstract: The recent empirical success of unsupervised cross-domain mapping algorithms, between two domains that share common characteristics, is not well-supported by theoretical justifications. This lacuna is especially troubling, given the clear ambiguity in such mappings. We work with adversarial training methods based on IPMs and derive a novel risk bound, which upper bounds the risk between the lear… ▽ More

    Submitted 2 November, 2020; v1 submitted 23 July, 2018; originally announced July 2018.

    Comments: arXiv admin note: text overlap with arXiv:1709.00074

  49. arXiv:1806.06029  [pdf, other

    cs.CV

    One-Shot Unsupervised Cross Domain Translation

    Authors: Sagie Benaim, Lior Wolf

    Abstract: Given a single image x from domain A and a set of images from domain B, our task is to generate the analogous of x in B. We argue that this task could be a key AI capability that underlines the ability of cognitive agents to act in the world and present empirical evidence that the existing unsupervised domain translation methods fail on this task. Our method follows a two step process. First, a va… ▽ More

    Submitted 23 October, 2018; v1 submitted 15 June, 2018; originally announced June 2018.

    Comments: Published at NIPS 2018

  50. arXiv:1712.07886  [pdf, other

    cs.LG

    Estimating the Success of Unsupervised Image to Image Translation

    Authors: Sagie Benaim, Tomer Galanti, Lior Wolf

    Abstract: While in supervised learning, the validation error is an unbiased estimator of the generalization (test) error and complexity-based generalization bounds are abundant, no such bounds exist for learning a mapping in an unsupervised way. As a result, when training GANs and specifically when using GANs for learning to map between domains in a completely unsupervised way, one is forced to select the h… ▽ More

    Submitted 22 March, 2018; v1 submitted 21 December, 2017; originally announced December 2017.

    Comments: The first and second authors contributed equally