Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 85 results for author: Satoh, S

.
  1. arXiv:2608.05592  [pdf, ps, other

    cs.CV

    Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs

    Authors: Ziling Huang, Shin'ichi Satoh

    Abstract: Multimodal Large Language Models (MLLMs) have achieved strong progress in video understanding, yet it remains challenging because the token limitation makes MLLMs difficult to capture temporally sparse evidence. Existing methods typically rely on uniform sampling, or frame selection, but these strategies usually optimize either broad temporal coverage or local relevance, making it difficult to pre… ▽ More

    Submitted 6 August, 2026; originally announced August 2026.

  2. arXiv:2607.05199  [pdf, ps, other

    cs.AI

    Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models

    Authors: Raj Jaiswal, Dhruv Jain, Rishabh Dhawan, Sree Krishna Uppalapati, Shin'ichi Satoh, Tanuja Ganu, Rajiv Ratn Shah

    Abstract: Physics reasoning fails structurally in small language models: an error at any step propagates forward, corrupting every inference that follows. Limited domain knowledge, hallucination under multi-step derivation, and distributional sensitivity compound this failure. We propose a step-level reward framework that identifies the first reasoning error, generates targeted structured feedback, and trai… ▽ More

    Submitted 6 July, 2026; originally announced July 2026.

  3. arXiv:2603.18392  [pdf, ps, other

    math.GT

    A note on Fox colorings of virtual tangles

    Authors: Takuji Nakamura, Yasutaka Nakanishi, Shin Satoh, Kodai Wada

    Abstract: We study Fox colorings of tangle diagrams by $R=\mathbb{Z}$ or $\mathbb{Z}/p\mathbb{Z}$, where $p\geq3$ is an odd integer. For an $R$-colored $m$-string tangle diagram, the colors at the $2m$ boundary points form a vector $v\in R^{2m}$. We show that for classical tangle diagrams, such vectors are completely characterized by the alternating sum condition $Δ(v)=0$. We then investigate how this restr… ▽ More

    Submitted 18 March, 2026; originally announced March 2026.

    Comments: 11 pages

    MSC Class: 57K10; 57K12

  4. DiverXplorer: Stock Image Exploration via Diversity Adjustment for Graphic Design

    Authors: Antonio Tejero-de-Pablos, Sichao Song, Naoto Ohsaka, Mayu Otani, Shin'ichi Satoh

    Abstract: Graphic designers explore large stock image collections during open-ended or early-stage design tasks, yet common tools emphasize relevance and similarity, limiting designers' ability to overview the design space or discover visual patterns. We present an image exploration prototype that enables stepwise adjustment of diversity, allowing users to transition from diverse overviews to increasingly f… ▽ More

    Submitted 9 March, 2026; originally announced March 2026.

    Comments: Accepted at CHI EA 2026

  5. arXiv:2601.15634  [pdf, ps, other

    math.GT

    The $V_1$- and $V_2$-polynomials of a long virtual knot

    Authors: Shin Satoh, Kodai Wada

    Abstract: We introduce two polynomial invariants $V_1(K;t)$ and $V_2(K;t)$ of a long virtual knot $K$, which generalize the degree-two finite type invariants $v_{2,1}$ and $v_{2,2}$ of Goussarov, Polyak, and Viro. We establish their fundamental properties and show that any pair of Laurent polynomials can be realized as $(V_1(K;t),V_2(K;t))$ for some long virtual knot $K$. While these polynomials are not fin… ▽ More

    Submitted 21 January, 2026; originally announced January 2026.

    Comments: 16 pages

    MSC Class: 57K12; 57K14; 57K16

  6. arXiv:2512.05369  [pdf, ps, other

    math.GT

    The intersection polynomials of a long virtual knot II: Two supporting genera and characterizations

    Authors: Takuji Nakamura, Yasutaka Nakanishi, Shin Satoh, Kodai Wada

    Abstract: We develop the study of the twelve intersection polynomials of long virtual knots, previously introduced in our preceding paper. We define two geometric invariants, the $1$- and $2$-supporting genera, using two distinct surface realizations. These genera yield a natural filtration of the set of long virtual knots, and we analyze the behavior of the intersection polynomials for long virtual knots w… ▽ More

    Submitted 4 December, 2025; originally announced December 2025.

    Comments: 20 pages

    MSC Class: Primary 57K12; Secondary 57K10; 57K14

  7. arXiv:2512.05366  [pdf, ps, other

    math.GT

    The intersection polynomials of a long virtual knot I: Definitions and properties

    Authors: Takuji Nakamura, Yasutaka Nakanishi, Shin Satoh, Kodai Wada

    Abstract: We introduce twelve polynomial invariants for long virtual knots, called intersection polynomials, extending and refining the three intersection polynomials for virtual knots. They are defined via intersection numbers of cycles on a closed surface, considering the order of over- and under-crossings. We study their fundamental properties including behavior under symmetries, crossing changes, and co… ▽ More

    Submitted 4 December, 2025; originally announced December 2025.

    Comments: 25 pages

    MSC Class: Primary 57K12; Secondary 57K14; 57K16

  8. arXiv:2511.20250  [pdf, ps, other

    cs.CV cs.AI cs.LG

    Uplifting Table Tennis: A Robust, Real-World Application for 3D Trajectory and Spin Estimation

    Authors: Daniel Kienzle, Katja Ludwig, Julian Lorenz, Shin'ichi Satoh, Rainer Lienhart

    Abstract: Obtaining the precise 3D motion of a table tennis ball from standard monocular videos is a challenging problem, as existing methods trained on synthetic data struggle to generalize to the noisy, imperfect ball and table detections of the real world. This is primarily due to the inherent lack of 3D ground truth trajectories and spin annotations for real-world video. To overcome this, we propose a n… ▽ More

    Submitted 25 November, 2025; originally announced November 2025.

    ACM Class: I.2.6; I.2.10; I.4.5

  9. arXiv:2510.10191  [pdf, ps, other

    cs.CV

    Fairness Without Labels: Pseudo-Balancing for Bias Mitigation in Face Gender Classification

    Authors: Haohua Dong, Ana Manzano Rodríguez, Camille Guinaudeau, Shin'ichi Satoh

    Abstract: Face gender classification models often reflect and amplify demographic biases present in their training data, leading to uneven performance across gender and racial subgroups. We introduce pseudo-balancing, a simple and effective strategy for mitigating such biases in semi-supervised learning. Our method enforces demographic balance during pseudo-label selection, using only unlabeled images from… ▽ More

    Submitted 11 October, 2025; originally announced October 2025.

    Comments: 8 pages. Accepted for publication in the ICCV 2025 Workshop Proceedings (2nd FAILED Workshop). Also available on HAL (hal-05210445v1)

    MSC Class: 68T07 ACM Class: I.2.10; I.4.8; I.5.4

  10. arXiv:2506.15180  [pdf, ps, other

    cs.CV

    ReSeDis: A Dataset for Referring-based Object Search across Large-Scale Image Collections

    Authors: Ziling Huang, Yidan Zhang, Shin'ichi Satoh

    Abstract: Large-scale visual search engines are expected to solve a dual problem at once: (i) locate every image that truly contains the object described by a sentence and (ii) identify the object's bounding box or exact pixels within each hit. Existing techniques address only one side of this challenge. Visual grounding yields tight boxes and masks but rests on the unrealistic assumption that the object is… ▽ More

    Submitted 18 June, 2025; originally announced June 2025.

  11. arXiv:2506.14473  [pdf, ps, other

    cs.CV cs.LG

    Foundation Model Insights and a Multi-Model Approach for Superior Fine-Grained One-shot Subset Selection

    Authors: Zhijing Wan, Zhixiang Wang, Zheng Wang, Xin Xu, Shin'ichi Satoh

    Abstract: One-shot subset selection serves as an effective tool to reduce deep learning training costs by identifying an informative data subset based on the information extracted by an information extractor (IE). Traditional IEs, typically pre-trained on the target dataset, are inherently dataset-dependent. Foundation models (FMs) offer a promising alternative, potentially mitigating this limitation. This… ▽ More

    Submitted 27 June, 2025; v1 submitted 17 June, 2025; originally announced June 2025.

    Comments: 18 pages, 10 figures, accepted by ICML 2025

  12. arXiv:2506.07032  [pdf, ps, other

    cs.CL cs.CV

    A Culturally-diverse Multilingual Multimodal Video Benchmark & Model

    Authors: Bhuiyan Sanjid Shafique, Ashmal Vayani, Muhammad Maaz, Hanoona Abdul Rasheed, Dinura Dissanayake, Mohammed Irfan Kurpath, Yahya Hmaiti, Go Inoue, Jean Lahoud, Md. Safirur Rashid, Shadid Intisar Quasem, Maheen Fatima, Franco Vidal, Mykola Maslych, Ketan Pravin More, Sanoojan Baliah, Hasindri Watawana, Yuhao Li, Fabian Farestam, Leon Schaller, Roman Tymtsiv, Simon Weber, Hisham Cholakkal, Ivan Laptev, Shin'ichi Satoh , et al. (4 additional authors not shown)

    Abstract: Large multimodal models (LMMs) have recently gained attention due to their effectiveness to understand and generate descriptions of visual content. Most existing LMMs are in English language. While few recent works explore multilingual image LMMs, to the best of our knowledge, moving beyond the English language for cultural and linguistic inclusivity is yet to be investigated in the context of vid… ▽ More

    Submitted 29 September, 2025; v1 submitted 8 June, 2025; originally announced June 2025.

  13. arXiv:2506.02547  [pdf, ps, other

    cs.CV cs.ET

    Probabilistic Online Event Downsampling

    Authors: Andreu Girbau-Xalabarder, Jun Nagata, Shinichi Sumiyoshi, Ricard Marsal, Shin'ichi Satoh

    Abstract: Event cameras capture scene changes asynchronously on a per-pixel basis, enabling extremely high temporal resolution. However, this advantage comes at the cost of high bandwidth, memory, and computational demands. To address this, prior work has explored event downsampling, but most approaches rely on fixed heuristics or threshold-based strategies, limiting their adaptability. Instead, we propose… ▽ More

    Submitted 23 September, 2025; v1 submitted 3 June, 2025; originally announced June 2025.

    Comments: Best paper award finalist at CVPR 2025 Event-Vision workshop

  14. arXiv:2504.19863  [pdf, other

    cs.CV cs.AI cs.LG

    Towards Ball Spin and Trajectory Analysis in Table Tennis Broadcast Videos via Physically Grounded Synthetic-to-Real Transfer

    Authors: Daniel Kienzle, Robin Schön, Rainer Lienhart, Shin'Ichi Satoh

    Abstract: Analyzing a player's technique in table tennis requires knowledge of the ball's 3D trajectory and spin. While, the spin is not directly observable in standard broadcasting videos, we show that it can be inferred from the ball's trajectory in the video. We present a novel method to infer the initial spin and 3D trajectory from the corresponding 2D trajectory in a video. Without ground truth labels… ▽ More

    Submitted 28 April, 2025; originally announced April 2025.

    Comments: To be published in 2025 IEEE/CVF International Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)

  15. arXiv:2504.05001  [pdf, ps, other

    astro-ph.IM eess.SY gr-qc physics.ins-det

    SILVIA: Ultra-precision formation flying demonstration for space-based interferometry

    Authors: Takahiro Ito, Kiwamu Izumi, Isao Kawano, Ikkoh Funaki, Shuichi Sato, Tomotada Akutsu, Kentaro Komori, Mitsuru Musha, Yuta Michimura, Satoshi Satoh, Takuya Iwaki, Kentaro Yokota, Kenta Goto, Katsumi Furukawa, Taro Matsuo, Toshihiro Tsuzuki, Katsuhiko Yamada, Takahiro Sasaki, Taisei Nishishita, Yuki Matsumoto, Chikako Hirose, Wataru Torii, Satoshi Ikari, Koji Nagano, Masaki Ando , et al. (4 additional authors not shown)

    Abstract: We propose SILVIA (Space Interferometer Laboratory Voyaging towards Innovative Applications), a mission concept designed to demonstrate ultra-precision formation flying between three spacecraft separated by 100 m. SILVIA aims to achieve sub-micrometer precision in relative distance control by integrating spacecraft sensors, laser interferometry, low-thrust and low-noise micro-propulsion for real-t… ▽ More

    Submitted 3 September, 2025; v1 submitted 7 April, 2025; originally announced April 2025.

    Comments: 10 pages, 6 figures, accepted for publication in Publications of the Astronomical Society of Japan

  16. arXiv:2410.11599  [pdf, ps, other

    math.GT

    Hurwitz equivalence in the universal dihedral quandle

    Authors: Takuji Nakamura, Yasutaka Nakanishi, Shin Satoh, Kodai Wada

    Abstract: We investigate the Hurwitz action of the $m$-braid group on the $m$-fold Cartesian product of the universal dihedral quandle. We introduce three computable invariants and prove that they give a complete classification of the orbits under this action. As a consequence, we describe an explicit complete system of orbit representatives. We further obtain analogous classifications for the corresponding… ▽ More

    Submitted 18 March, 2026; v1 submitted 15 October, 2024; originally announced October 2024.

    Comments: 22 pages. The original manuscript has been split into two separate papers. This version (v2) contains the first part; the second part will appear as a separate submission

    MSC Class: Primary 20F36; 05E18; Secondary 57K12; 57M10

  17. Matting by Generation

    Authors: Zhixiang Wang, Baiang Li, Jian Wang, Yu-Lun Liu, Jinwei Gu, Yung-Yu Chuang, Shin'ichi Satoh

    Abstract: This paper introduces an innovative approach for image matting that redefines the traditional regression-based task as a generative modeling challenge. Our method harnesses the capabilities of latent diffusion models, enriched with extensive pre-trained knowledge, to regularize the matting process. We present novel architectural innovations that empower our model to produce mattes with superior re… ▽ More

    Submitted 30 July, 2024; originally announced July 2024.

    Comments: SIGGRAPH'24, Project page: https://lightchaserx.github.io/matting-by-generation/

  18. arXiv:2405.17188  [pdf, other

    cs.CV

    The SkatingVerse Workshop & Challenge: Methods and Results

    Authors: Jian Zhao, Lei Jin, Jianshu Li, Zheng Zhu, Yinglei Teng, Jiaojiao Zhao, Sadaf Gulshad, Zheng Wang, Bo Zhao, Xiangbo Shu, Yunchao Wei, Xuecheng Nie, Xiaojie Jin, Xiaodan Liang, Shin'ichi Satoh, Yandong Guo, Cewu Lu, Junliang Xing, Jane Shen Shengmei

    Abstract: The SkatingVerse Workshop & Challenge aims to encourage research in developing novel and accurate methods for human action understanding. The SkatingVerse dataset used for the SkatingVerse Challenge has been publicly released. There are two subsets in the dataset, i.e., the training subset and testing subset. The training subsets consists of 19,993 RGB video sequences, and the testing subsets cons… ▽ More

    Submitted 27 May, 2024; originally announced May 2024.

  19. TC-OCR: TableCraft OCR for Efficient Detection & Recognition of Table Structure & Content

    Authors: Avinash Anand, Raj Jaiswal, Pijush Bhuyan, Mohit Gupta, Siddhesh Bangar, Md. Modassir Imam, Rajiv Ratn Shah, Shin'ichi Satoh

    Abstract: The automatic recognition of tabular data in document images presents a significant challenge due to the diverse range of table styles and complex structures. Tables offer valuable content representation, enhancing the predictive capabilities of various systems such as search engines and Knowledge Graphs. Addressing the two main problems, namely table detection (TD) and table structure recognition… ▽ More

    Submitted 19 April, 2024; v1 submitted 16 April, 2024; originally announced April 2024.

    Comments: 8 pages, 2 figures, Workshop of 1st MMIR Deep Multimodal Learning for Information Retrieval

  20. RanLayNet: A Dataset for Document Layout Detection used for Domain Adaptation and Generalization

    Authors: Avinash Anand, Raj Jaiswal, Mohit Gupta, Siddhesh S Bangar, Pijush Bhuyan, Naman Lal, Rajeev Singh, Ritika Jha, Rajiv Ratn Shah, Shin'ichi Satoh

    Abstract: Large ground-truth datasets and recent advances in deep learning techniques have been useful for layout detection. However, because of the restricted layout diversity of these datasets, training on them requires a sizable number of annotated instances, which is both expensive and time-consuming. As a result, differences between the source and target domains may significantly impact how well these… ▽ More

    Submitted 19 April, 2024; v1 submitted 15 April, 2024; originally announced April 2024.

    Comments: 8 pages, 6 figures, MMAsia 2023 Proceedings of the 5th ACM International Conference on Multimedia in Asia

    Journal ref: In Proceedings of the 5th ACM International Conference on Multimedia in Asia 2023. Association for Computing Machinery, NY, USA, Article 74, pp. 1-6

  21. arXiv:2403.18158  [pdf, other

    cs.CV

    The Effects of Short Video-Sharing Services on Video Copy Detection

    Authors: Rintaro Yanagi, Yamato Okamoto, Shuhei Yokoo, Shin'ichi Satoh

    Abstract: The short video-sharing services that allow users to post 10-30 second videos (e.g., YouTube Shorts and TikTok) have attracted a lot of attention in recent years. However, conventional video copy detection (VCD) methods mainly focus on general video-sharing services (e.g., YouTube and Bilibili), and the effects of short video-sharing services on video copy detection are still unclear. Considering… ▽ More

    Submitted 26 March, 2024; originally announced March 2024.

  22. arXiv:2403.03917  [pdf, other

    math.GT

    On wen knots

    Authors: Celeste Damiani, Shin Satoh

    Abstract: We introduce the notion of wen knots, and prove that the set of wen knots is a proper subset of the set of extended welded knots. Furthermore we prove that the complementary subset consists of welded knots up to horizontal mirror reflections. This allow us to characterise completely extended welded knots by the parity of their number of wens, that we can always reduce to 0 or 1.

    Submitted 6 March, 2024; originally announced March 2024.

    Comments: All comments are welcome!

    MSC Class: Primary 57K12; Secondary 57K45

  23. arXiv:2401.16193  [pdf, other

    cs.LG cs.DB

    Contributing Dimension Structure of Deep Feature for Coreset Selection

    Authors: Zhijing Wan, Zhixiang Wang, Yuran Wang, Zheng Wang, Hongyuan Zhu, Shin'ichi Satoh

    Abstract: Coreset selection seeks to choose a subset of crucial training samples for efficient learning. It has gained traction in deep learning, particularly with the surge in training dataset sizes. Sample selection hinges on two main aspects: a sample's representation in enhancing performance and the role of sample diversity in averting overfitting. Existing methods typically measure both the representat… ▽ More

    Submitted 2 March, 2024; v1 submitted 29 January, 2024; originally announced January 2024.

    Comments: 13 pages,11 figures, to be published in AAAI2024

  24. arXiv:2401.13195  [pdf, other

    math.GT

    Virtualized Delta, sharp, and pass moves for oriented virtual knots and links

    Authors: Takuji Nakamura, Yasutaka Nakanishi, Shin Satoh, Kodai Wada

    Abstract: We study virtualized Delta, sharp, and pass moves for oriented virtual links, and give necessary and sufficient conditions for two oriented virtual links to be related by the local moves. In particular, they are unknotting operations for oriented virtual knots. We provide lower bounds for the unknotting numbers and prove that they are best possible.

    Submitted 23 January, 2024; originally announced January 2024.

    Comments: 20 pages

    MSC Class: Primary 57K12; Secondary 57K10

  25. arXiv:2401.12506  [pdf, other

    math.GT

    Virtualized Delta moves for virtual knots and links

    Authors: Takuji Nakamura, Yasutaka Nakanishi, Shin Satoh, Kodai Wada

    Abstract: We introduce a local deformation called the virtualized $Δ$-move for virtual knots and links. We prove that the virtualized $Δ$-move is an unknotting operation for virtual knots. Furthermore we give a necessary and sufficient condition for two virtual links to be related by a finite sequence of virtualized $Δ$-moves.

    Submitted 23 January, 2024; originally announced January 2024.

    Comments: 10 pages

    MSC Class: Primary 57K12; Secondary 57K10

  26. arXiv:2310.16377  [pdf, other

    eess.SY math.OC

    Nonlinear steering control under input magnitude and rate constraints with exponential convergence

    Authors: Rin Suyama, Satoshi Satoh, Atsuo Maki

    Abstract: A ship steering control is designed for a nonlinear maneuvering model whose rudder manipulation is constrained in both magnitude and rate. In our method, the tracking problem of the target heading angle with input constraints is converted into the tracking problem for a strict-feedback system without any input constraints. To derive this system, hyperbolic tangent ($\tanh$) function and auxiliary… ▽ More

    Submitted 25 October, 2023; originally announced October 2023.

    Comments: 12 pages, 6 figures, a preprint submitted to the Journal of Marine Science and Technology

  27. arXiv:2309.08372  [pdf, other

    cs.CV cs.MM

    Beyond Domain Gap: Exploiting Subjectivity in Sketch-Based Person Retrieval

    Authors: Kejun Lin, Zhixiang Wang, Zheng Wang, Yinqiang Zheng, Shin'ichi Satoh

    Abstract: Person re-identification (re-ID) requires densely distributed cameras. In practice, the person of interest may not be captured by cameras and, therefore, needs to be retrieved using subjective information (e.g., sketches from witnesses). Previous research defines this case using the sketch as sketch re-identification (Sketch re-ID) and focuses on eliminating the domain gap. Actually, subjectivity… ▽ More

    Submitted 15 September, 2023; originally announced September 2023.

    Comments: ACM Multimedia 2023

  28. arXiv:2306.11528  [pdf, other

    cs.CV

    TransRef: Multi-Scale Reference Embedding Transformer for Reference-Guided Image Inpainting

    Authors: Taorong Liu, Liang Liao, Delin Chen, Jing Xiao, Zheng Wang, Chia-Wen Lin, Shin'ichi Satoh

    Abstract: Image inpainting for completing complicated semantic environments and diverse hole patterns of corrupted images is challenging even for state-of-the-art learning-based inpainting methods trained on large-scale data. A reference image capturing the same scene of a corrupted image offers informative guidance for completing the corrupted image as it shares similar texture and structure priors to that… ▽ More

    Submitted 11 February, 2025; v1 submitted 20 June, 2023; originally announced June 2023.

    Comments: Neurocomputing 2025

  29. arXiv:2305.07290  [pdf, other

    cs.CV

    The 3rd Anti-UAV Workshop & Challenge: Methods and Results

    Authors: Jian Zhao, Jianan Li, Lei Jin, Jiaming Chu, Zhihao Zhang, Jun Wang, Jiangqiang Xia, Kai Wang, Yang Liu, Sadaf Gulshad, Jiaojiao Zhao, Tianyang Xu, Xuefeng Zhu, Shihan Liu, Zheng Zhu, Guibo Zhu, Zechao Li, Zheng Wang, Baigui Sun, Yandong Guo, Shin ichi Satoh, Junliang Xing, Jane Shen Shengmei

    Abstract: The 3rd Anti-UAV Workshop & Challenge aims to encourage research in developing novel and accurate methods for multi-scale object tracking. The Anti-UAV dataset used for the Anti-UAV Challenge has been publicly released. There are two main differences between this year's competition and the previous two. First, we have expanded the existing dataset, and for the first time, released a training set s… ▽ More

    Submitted 15 July, 2023; v1 submitted 12 May, 2023; originally announced May 2023.

    Comments: Technical report for 3rd Anti-UAV Workshop and Challenge. arXiv admin note: text overlap with arXiv:2108.09909

  30. arXiv:2304.06430  [pdf, other

    cs.CV cs.AI

    Certified Zeroth-order Black-Box Defense with Robust UNet Denoiser

    Authors: Astha Verma, A V Subramanyam, Siddhesh Bangar, Naman Lal, Rajiv Ratn Shah, Shin'ichi Satoh

    Abstract: Certified defense methods against adversarial perturbations have been recently investigated in the black-box setting with a zeroth-order (ZO) perspective. However, these methods suffer from high model variance with low performance on high-dimensional datasets due to the ineffective design of the denoiser and are limited in their utilization of ZO techniques. To this end, we propose a certified ZO… ▽ More

    Submitted 6 July, 2024; v1 submitted 13 April, 2023; originally announced April 2023.

  31. arXiv:2304.01816  [pdf, other

    cs.CV

    Toward Verifiable and Reproducible Human Evaluation for Text-to-Image Generation

    Authors: Mayu Otani, Riku Togashi, Yu Sawai, Ryosuke Ishigami, Yuta Nakashima, Esa Rahtu, Janne Heikkilä, Shin'ichi Satoh

    Abstract: Human evaluation is critical for validating the performance of text-to-image generative models, as this highly cognitive process requires deep comprehension of text and images. However, our survey of 37 recent papers reveals that many works rely solely on automatic measures (e.g., FID) or perform poorly described human evaluations that are not reliable or repeatable. This paper proposes a standard… ▽ More

    Submitted 4 April, 2023; originally announced April 2023.

    Comments: CVPR 2023

  32. arXiv:2302.12253  [pdf, other

    cs.CV

    DisCO: Portrait Distortion Correction with Perspective-Aware 3D GANs

    Authors: Zhixiang Wang, Yu-Lun Liu, Jia-Bin Huang, Shin'ichi Satoh, Sizhuo Ma, Gurunandan Krishnan, Jian Wang

    Abstract: Close-up facial images captured at short distances often suffer from perspective distortion, resulting in exaggerated facial features and unnatural/unattractive appearances. We propose a simple yet effective method for correcting perspective distortions in a single close-up face. We first perform GAN inversion using a perspective-distorted input facial image by jointly optimizing the camera intrin… ▽ More

    Submitted 8 December, 2023; v1 submitted 23 February, 2023; originally announced February 2023.

    Comments: Project website: https://portrait-disco.github.io/

  33. arXiv:2212.05709  [pdf, other

    cs.CV

    HOTCOLD Block: Fooling Thermal Infrared Detectors with a Novel Wearable Design

    Authors: Hui Wei, Zhixiang Wang, Xuemei Jia, Yinqiang Zheng, Hao Tang, Shin'ichi Satoh, Zheng Wang

    Abstract: Adversarial attacks on thermal infrared imaging expose the risk of related applications. Estimating the security of these systems is essential for safely deploying them in the real world. In many cases, realizing the attacks in the physical space requires elaborate special perturbations. These solutions are often \emph{impractical} and \emph{attention-grabbing}. To address the need for a physicall… ▽ More

    Submitted 12 December, 2022; originally announced December 2022.

    Comments: Accepted to AAAI 2023

  34. arXiv:2211.07566  [pdf, other

    cs.CV

    Self-distillation with Online Diffusion on Batch Manifolds Improves Deep Metric Learning

    Authors: Zelong Zeng, Fan Yang, Hong Liu, Shin'ichi Satoh

    Abstract: Recent deep metric learning (DML) methods typically leverage solely class labels to keep positive samples far away from negative ones. However, this type of method normally ignores the crucial knowledge hidden in the data (e.g., intra-class information variation), which is harmful to the generalization of the trained model. To alleviate this problem, in this paper we propose Online Batch Diffusion… ▽ More

    Submitted 14 November, 2022; originally announced November 2022.

    Comments: 14 pages

  35. arXiv:2210.03355  [pdf, other

    cs.CV

    Multiple Object Tracking from appearance by hierarchically clustering tracklets

    Authors: Andreu Girbau, Ferran Marqués, Shin'ichi Satoh

    Abstract: Current approaches in Multiple Object Tracking (MOT) rely on the spatio-temporal coherence between detections combined with object appearance to match objects from consecutive frames. In this work, we explore MOT using object appearances as the main source of association between objects in a video, using spatial and temporal priors as weighting factors. We form initial tracklets by leveraging on t… ▽ More

    Submitted 7 October, 2022; originally announced October 2022.

    Comments: To be published in BMVC 2022

  36. Physical Adversarial Attack meets Computer Vision: A Decade Survey

    Authors: Hui Wei, Hao Tang, Xuemei Jia, Zhixiang Wang, Hanxun Yu, Zhubo Li, Shin'ichi Satoh, Luc Van Gool, Zheng Wang

    Abstract: Despite the impressive achievements of Deep Neural Networks (DNNs) in computer vision, their vulnerability to adversarial attacks remains a critical concern. Extensive research has demonstrated that incorporating sophisticated perturbations into input images can lead to a catastrophic degradation in DNNs' performance. This perplexing phenomenon not only exists in the digital space but also in the… ▽ More

    Submitted 20 November, 2024; v1 submitted 29 September, 2022; originally announced September 2022.

    Comments: Published at IEEE TPAMI. GitHub:https://github.com/weihui1308/PAA

  37. arXiv:2207.14498  [pdf, other

    cs.CV

    Reference-Guided Texture and Structure Inference for Image Inpainting

    Authors: Taorong Liu, Liang Liao, Zheng Wang, Shin'ichi Satoh

    Abstract: Existing learning-based image inpainting methods are still in challenge when facing complex semantic environments and diverse hole patterns. The prior information learned from the large scale training data is still insufficient for these situations. Reference images captured covering the same scenes share similar texture and structure priors with the corrupted images, which offers new prospects fo… ▽ More

    Submitted 29 July, 2022; originally announced July 2022.

    Comments: IEEE International Conference on Image Processing(ICIP 2022)

  38. arXiv:2206.08880  [pdf, other

    cs.CV cs.LG

    Improving Generalization of Metric Learning via Listwise Self-distillation

    Authors: Zelong Zeng, Fan Yang, Zheng Wang, Shin'ichi Satoh

    Abstract: Most deep metric learning (DML) methods employ a strategy that forces all positive samples to be close in the embedding space while keeping them away from negative ones. However, such a strategy ignores the internal relationships of positive (negative) samples and often leads to overfitting, especially in the presence of hard samples and mislabeled samples. In this work, we propose a simple yet ef… ▽ More

    Submitted 17 June, 2022; originally announced June 2022.

    Comments: 11 pages, 7 figures

  39. Unsupervised Foggy Scene Understanding via Self Spatial-Temporal Label Diffusion

    Authors: Liang Liao, Wenyi Chen, Jing Xiao, Zheng Wang, Chia-Wen Lin, Shin'ichi Satoh

    Abstract: Understanding foggy image sequence in the driving scenes is critical for autonomous driving, but it remains a challenging task due to the difficulty in collecting and annotating real-world images of adverse weather. Recently, the self-training strategy has been considered a powerful solution for unsupervised domain adaptation, which iteratively adapts the model from the source domain to the target… ▽ More

    Submitted 10 June, 2022; originally announced June 2022.

    Comments: IEEE Transactions on Image Processing 2022

  40. Geo-Localization via Ground-to-Satellite Cross-View Image Retrieval

    Authors: Zelong Zeng, Zheng Wang, Fan Yang, Shin'ichi Satoh

    Abstract: The large variation of viewpoint and irrelevant content around the target always hinder accurate image retrieval and its subsequent tasks. In this paper, we investigate an extremely challenging task: given a ground-view image of a landmark, we aim to achieve cross-view geo-localization by searching out its corresponding satellite-view images. Specifically, the challenge comes from the gap between… ▽ More

    Submitted 22 May, 2022; originally announced May 2022.

    Comments: 13 pages, 10 figures

    Journal ref: IEEE Transactions on Multimedia (2022)

  41. arXiv:2204.00974  [pdf, other

    cs.CV

    Neural Global Shutter: Learn to Restore Video from a Rolling Shutter Camera with Global Reset Feature

    Authors: Zhixiang Wang, Xiang Ji, Jia-Bin Huang, Shin'ichi Satoh, Xiao Zhou, Yinqiang Zheng

    Abstract: Most computer vision systems assume distortion-free images as inputs. The widely used rolling-shutter (RS) image sensors, however, suffer from geometric distortion when the camera and object undergo motion during capture. Extensive researches have been conducted on correcting RS distortions. However, most of the existing work relies heavily on the prior assumptions of scenes or motions. Besides, t… ▽ More

    Submitted 25 June, 2022; v1 submitted 2 April, 2022; originally announced April 2022.

    Comments: CVPR2022, https://github.com/lightChaserX/neural-global-shutter

  42. arXiv:2203.14438  [pdf, other

    cs.CV

    Optimal Correction Cost for Object Detection Evaluation

    Authors: Mayu Otani, Riku Togashi, Yuta Nakashima, Esa Rahtu, Janne Heikkilä, Shin'ichi Satoh

    Abstract: Mean Average Precision (mAP) is the primary evaluation measure for object detection. Although object detection has a broad range of applications, mAP evaluates detectors in terms of the performance of ranked instance retrieval. Such the assumption for the evaluation task does not suit some downstream tasks. To alleviate the gap between downstream tasks and the evaluation scenario, we propose Optim… ▽ More

    Submitted 27 March, 2022; originally announced March 2022.

    Comments: CVPR 2022

  43. Improving Camouflaged Object Detection with the Uncertainty of Pseudo-edge Labels

    Authors: Nobukatsu Kajiura, Hong Liu, Shin'ichi Satoh

    Abstract: This paper focuses on camouflaged object detection (COD), which is a task to detect objects hidden in the background. Most of the current COD models aim to highlight the target object directly while outputting ambiguous camouflaged boundaries. On the other hand, the performance of the models considering edge information is not yet satisfactory. To this end, we propose a new framework that makes fu… ▽ More

    Submitted 29 October, 2021; originally announced October 2021.

    Comments: Accepted to ACM Multimedia Asia 2021

  44. Scalable Personalised Item Ranking through Parametric Density Estimation

    Authors: Riku Togashi, Masahiro Kato, Mayu Otani, Tetsuya Sakai, Shin'ichi Satoh

    Abstract: Learning from implicit feedback is challenging because of the difficult nature of the one-class problem: we can observe only positive examples. Most conventional methods use a pairwise ranking approach and negative samplers to cope with the one-class problem. However, such methods have two main drawbacks particularly in large-scale applications; (1) the pairwise approach is severely inefficient du… ▽ More

    Submitted 10 May, 2021; originally announced May 2021.

    Comments: Accepted by SIGIR'21

  45. arXiv:2102.12067  [pdf, other

    math.GT

    The intersection polynomials of a virtual knot I: Definitions and calculations

    Authors: Ryuji Higa, Takuji Nakamura, Yasutaka Nakanishi, Shin Satoh

    Abstract: We introduce three kinds of invariants of a virtual knot called the first, second, and third intersection polynomials. The definition is based on the intersection number of a pair of curves on a closed surface. The calculations of intersection polynomials are given up to crossing number four. We also study several properties of intersection polynomials.

    Submitted 25 January, 2022; v1 submitted 23 February, 2021; originally announced February 2021.

    Comments: 27 pages, 22 figures. We divide the original version into four papers

    MSC Class: 57K12 (Primary) 57K10; 57K14 (Secondary)

  46. arXiv:2101.07481  [pdf, other

    cs.IR

    Density-Ratio Based Personalised Ranking from Implicit Feedback

    Authors: Riku Togashi, Masahiro Kato, Mayu Otani, Shin'ichi Satoh

    Abstract: Learning from implicit user feedback is challenging as we can only observe positive samples but never access negative ones. Most conventional methods cope with this issue by adopting a pairwise ranking approach with negative sampling. However, the pairwise ranking approach has a severe disadvantage in the convergence time owing to the quadratically increasing computational cost with respect to the… ▽ More

    Submitted 19 January, 2021; originally announced January 2021.

    Comments: Accepted by WWW 2021

  47. arXiv:2012.08054  [pdf, other

    cs.CV

    Image Inpainting Guided by Coherence Priors of Semantics and Textures

    Authors: Liang Liao, Jing Xiao, Zheng Wang, Chia-Wen Lin, Shin'ichi Satoh

    Abstract: Existing inpainting methods have achieved promising performance in recovering defected images of specific scenes. However, filling holes involving multiple semantic categories remains challenging due to the obscure semantic boundaries and the mixture of different semantic textures. In this paper, we introduce coherence priors between the semantics and textures which make it possible to concentrate… ▽ More

    Submitted 14 December, 2020; originally announced December 2020.

  48. arXiv:2011.05061  [pdf, other

    cs.IR

    Alleviating Cold-Start Problems in Recommendation through Pseudo-Labelling over Knowledge Graph

    Authors: Riku Togashi, Mayu Otani, Shin'ichi Satoh

    Abstract: Solving cold-start problems is indispensable to provide meaningful recommendation results for new users and items. Under sparsely observed data, unobserved user-item pairs are also a vital source for distilling latent users' information needs. Most present works leverage unobserved samples for extracting negative signals. However, such an optimisation strategy can lead to biased results toward alr… ▽ More

    Submitted 10 November, 2020; originally announced November 2020.

    Comments: WSDM 2021

  49. arXiv:2008.05383  [pdf, other

    cs.CV

    Towards Unsupervised Crowd Counting via Regression-Detection Bi-knowledge Transfer

    Authors: Yuting Liu, Zheng Wang, Miaojing Shi, Shin'ichi Satoh, Qijun Zhao, Hongyu Yang

    Abstract: Unsupervised crowd counting is a challenging yet not largely explored task. In this paper, we explore it in a transfer learning setting where we learn to detect and count persons in an unlabeled target set by transferring bi-knowledge learnt from regression- and detection-based models in a labeled source set. The dual source knowledge of the two models is heterogeneous and complementary as they ca… ▽ More

    Submitted 27 September, 2020; v1 submitted 12 August, 2020; originally announced August 2020.

    Comments: This paper has been accepted by ACM MM 2020(Oral)

  50. arXiv:2007.13559  [pdf, other

    cs.CV cs.LG eess.IV

    MADGAN: unsupervised Medical Anomaly Detection GAN using multiple adjacent brain MRI slice reconstruction

    Authors: Changhee Han, Leonardo Rundo, Kohei Murao, Tomoyuki Noguchi, Yuki Shimahara, Zoltan Adam Milacski, Saori Koshino, Evis Sala, Hideki Nakayama, Shinichi Satoh

    Abstract: Unsupervised learning can discover various unseen abnormalities, relying on large-scale unannotated medical images of healthy subjects. Towards this, unsupervised methods reconstruct a 2D/3D single medical image to detect outliers either in the learned feature space or from high reconstruction loss. However, without considering continuity between multiple adjacent slices, they cannot directly disc… ▽ More

    Submitted 12 October, 2020; v1 submitted 24 July, 2020; originally announced July 2020.

    Comments: 23 pages, 11 figures, submitted to BMC Bioinformatics. Extended version of arXiv:1906.06114