-
Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs
Authors:
Ziling Huang,
Shin'ichi Satoh
Abstract:
Multimodal Large Language Models (MLLMs) have achieved strong progress in video understanding, yet it remains challenging because the token limitation makes MLLMs difficult to capture temporally sparse evidence. Existing methods typically rely on uniform sampling, or frame selection, but these strategies usually optimize either broad temporal coverage or local relevance, making it difficult to pre…
▽ More
Multimodal Large Language Models (MLLMs) have achieved strong progress in video understanding, yet it remains challenging because the token limitation makes MLLMs difficult to capture temporally sparse evidence. Existing methods typically rely on uniform sampling, or frame selection, but these strategies usually optimize either broad temporal coverage or local relevance, making it difficult to preserve both global storyline context and fine-grained evidence. We propose VideoRouter(VR) that rethinks long-video understanding as coordinating complementary evidence views rather than selecting a single subset of frames. It first organizes each video into a question-agnostic temporal hierarchy, which partition the video into coarse-to-fine temporally coherent segments. In this hierarchy, upper-level nodes capture broad storyline context and event progression, while lower-level nodes preserve fine-grained local details and evidence-bearing moments. This naturally gives rise to two complementary views: a global view for coverage-oriented reasoning and a local view for detail-oriented evidence recovery. We further introduce a verification-guided router to determine which view is better supported by the selected evidence and select the final answer. We validate the effectiveness of the proposed design through extensive experiments, showing that the verification-guided router effectively coordinates global and local reasoning, and that, on VideoMME, our method outperforms state-of-the-art frame selection methods by 2.9 points, respectively, under the LLaVA-Video-7B backbone. We will release the code.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models
Authors:
Raj Jaiswal,
Dhruv Jain,
Rishabh Dhawan,
Sree Krishna Uppalapati,
Shin'ichi Satoh,
Tanuja Ganu,
Rajiv Ratn Shah
Abstract:
Physics reasoning fails structurally in small language models: an error at any step propagates forward, corrupting every inference that follows. Limited domain knowledge, hallucination under multi-step derivation, and distributional sensitivity compound this failure. We propose a step-level reward framework that identifies the first reasoning error, generates targeted structured feedback, and trai…
▽ More
Physics reasoning fails structurally in small language models: an error at any step propagates forward, corrupting every inference that follows. Limited domain knowledge, hallucination under multi-step derivation, and distributional sensitivity compound this failure. We propose a step-level reward framework that identifies the first reasoning error, generates targeted structured feedback, and trains the model to revise its solution via policy gradient with KL regularization, without exposing it to ground truth solutions as generation targets. Unlike annotation-dependent step-level methods, no preference data construction is required and the external verifier operates exclusively at training time. Across five physics benchmarks, our framework delivers accuracy gains of 17-20% over CoT prompting and 10-16% over the strongest baseline, reduces calculation errors from 56.9% to 23.5%, and reduces miscomprehension errors from 22.3% to 12.0% in the best observed cases. Conceptual errors reduce from 89.7% to 68.7%, yet persist as the hardest failure mode across all conditions.
△ Less
Submitted 6 July, 2026;
originally announced July 2026.
-
A note on Fox colorings of virtual tangles
Authors:
Takuji Nakamura,
Yasutaka Nakanishi,
Shin Satoh,
Kodai Wada
Abstract:
We study Fox colorings of tangle diagrams by $R=\mathbb{Z}$ or $\mathbb{Z}/p\mathbb{Z}$, where $p\geq3$ is an odd integer. For an $R$-colored $m$-string tangle diagram, the colors at the $2m$ boundary points form a vector $v\in R^{2m}$. We show that for classical tangle diagrams, such vectors are completely characterized by the alternating sum condition $Δ(v)=0$. We then investigate how this restr…
▽ More
We study Fox colorings of tangle diagrams by $R=\mathbb{Z}$ or $\mathbb{Z}/p\mathbb{Z}$, where $p\geq3$ is an odd integer. For an $R$-colored $m$-string tangle diagram, the colors at the $2m$ boundary points form a vector $v\in R^{2m}$. We show that for classical tangle diagrams, such vectors are completely characterized by the alternating sum condition $Δ(v)=0$. We then investigate how this restriction changes in the virtual setting. For $R=\mathbb{Z}$, the realizability of $v$ is determined by a divisibility condition on $Δ(v)$. For $R=\mathbb{Z}/p\mathbb{Z}$, every vector is realizable by a virtual tangle diagram.
△ Less
Submitted 18 March, 2026;
originally announced March 2026.
-
DiverXplorer: Stock Image Exploration via Diversity Adjustment for Graphic Design
Authors:
Antonio Tejero-de-Pablos,
Sichao Song,
Naoto Ohsaka,
Mayu Otani,
Shin'ichi Satoh
Abstract:
Graphic designers explore large stock image collections during open-ended or early-stage design tasks, yet common tools emphasize relevance and similarity, limiting designers' ability to overview the design space or discover visual patterns. We present an image exploration prototype that enables stepwise adjustment of diversity, allowing users to transition from diverse overviews to increasingly f…
▽ More
Graphic designers explore large stock image collections during open-ended or early-stage design tasks, yet common tools emphasize relevance and similarity, limiting designers' ability to overview the design space or discover visual patterns. We present an image exploration prototype that enables stepwise adjustment of diversity, allowing users to transition from diverse overviews to increasingly focused subsets during exploration. Our approach implements diversity control via determinantal point process (DPP)-based sampling and exposes diversity-similarity tradeoffs through interaction rather than static ranking. We report findings from a pilot study with professional graphic designers comparing our technique to baselines inspired by current tools in open-ended image selection tasks. Results suggest that stepwise diversity control supports early-stage sensemaking and comparison of visual patterns, while revealing important tradeoffs: diversity aids discovery and reduces backtracking, but becomes less desirable as exploration progresses. We aim to provide a novel perspective on how to implement transitions between diversity and similarity. Our code is available at https://github.com/CyberAgentAILab/DiverXplorer.
△ Less
Submitted 9 March, 2026;
originally announced March 2026.
-
The $V_1$- and $V_2$-polynomials of a long virtual knot
Authors:
Shin Satoh,
Kodai Wada
Abstract:
We introduce two polynomial invariants $V_1(K;t)$ and $V_2(K;t)$ of a long virtual knot $K$, which generalize the degree-two finite type invariants $v_{2,1}$ and $v_{2,2}$ of Goussarov, Polyak, and Viro. We establish their fundamental properties and show that any pair of Laurent polynomials can be realized as $(V_1(K;t),V_2(K;t))$ for some long virtual knot $K$. While these polynomials are not fin…
▽ More
We introduce two polynomial invariants $V_1(K;t)$ and $V_2(K;t)$ of a long virtual knot $K$, which generalize the degree-two finite type invariants $v_{2,1}$ and $v_{2,2}$ of Goussarov, Polyak, and Viro. We establish their fundamental properties and show that any pair of Laurent polynomials can be realized as $(V_1(K;t),V_2(K;t))$ for some long virtual knot $K$. While these polynomials are not finite type invariants of any degree with respect to virtualizations, their first derivatives at $t=1$ define finite type invariants of degree three. As an application, we obtain an explicit Gauss diagram formula for the $α_3$-invariant.
△ Less
Submitted 21 January, 2026;
originally announced January 2026.
-
The intersection polynomials of a long virtual knot II: Two supporting genera and characterizations
Authors:
Takuji Nakamura,
Yasutaka Nakanishi,
Shin Satoh,
Kodai Wada
Abstract:
We develop the study of the twelve intersection polynomials of long virtual knots, previously introduced in our preceding paper. We define two geometric invariants, the $1$- and $2$-supporting genera, using two distinct surface realizations. These genera yield a natural filtration of the set of long virtual knots, and we analyze the behavior of the intersection polynomials for long virtual knots w…
▽ More
We develop the study of the twelve intersection polynomials of long virtual knots, previously introduced in our preceding paper. We define two geometric invariants, the $1$- and $2$-supporting genera, using two distinct surface realizations. These genera yield a natural filtration of the set of long virtual knots, and we analyze the behavior of the intersection polynomials for long virtual knots with small supporting genera. Moreover, we investigate virtual $2$-string tangles, analyzing how their sums with long virtual knots affect the intersection polynomials through right closures. As an application, we provide complete realizability criteria for all twelve intersection polynomials.
△ Less
Submitted 4 December, 2025;
originally announced December 2025.
-
The intersection polynomials of a long virtual knot I: Definitions and properties
Authors:
Takuji Nakamura,
Yasutaka Nakanishi,
Shin Satoh,
Kodai Wada
Abstract:
We introduce twelve polynomial invariants for long virtual knots, called intersection polynomials, extending and refining the three intersection polynomials for virtual knots. They are defined via intersection numbers of cycles on a closed surface, considering the order of over- and under-crossings. We study their fundamental properties including behavior under symmetries, crossing changes, and co…
▽ More
We introduce twelve polynomial invariants for long virtual knots, called intersection polynomials, extending and refining the three intersection polynomials for virtual knots. They are defined via intersection numbers of cycles on a closed surface, considering the order of over- and under-crossings. We study their fundamental properties including behavior under symmetries, crossing changes, and concatenation products. All are finite-type invariants of degree two under crossing changes, but not under virtualizations, and we examine their relation to the closure and the values at $t=1$ of their derivatives.
△ Less
Submitted 4 December, 2025;
originally announced December 2025.
-
Uplifting Table Tennis: A Robust, Real-World Application for 3D Trajectory and Spin Estimation
Authors:
Daniel Kienzle,
Katja Ludwig,
Julian Lorenz,
Shin'ichi Satoh,
Rainer Lienhart
Abstract:
Obtaining the precise 3D motion of a table tennis ball from standard monocular videos is a challenging problem, as existing methods trained on synthetic data struggle to generalize to the noisy, imperfect ball and table detections of the real world. This is primarily due to the inherent lack of 3D ground truth trajectories and spin annotations for real-world video. To overcome this, we propose a n…
▽ More
Obtaining the precise 3D motion of a table tennis ball from standard monocular videos is a challenging problem, as existing methods trained on synthetic data struggle to generalize to the noisy, imperfect ball and table detections of the real world. This is primarily due to the inherent lack of 3D ground truth trajectories and spin annotations for real-world video. To overcome this, we propose a novel two-stage pipeline that divides the problem into a front-end perception task and a back-end 2D-to-3D uplifting task. This separation allows us to train the front-end components with abundant 2D supervision from our newly created TTHQ dataset, while the back-end uplifting network is trained exclusively on physically-correct synthetic data. We specifically re-engineer the uplifting model to be robust to common real-world artifacts, such as missing detections and varying frame rates. By integrating a ball detector and a table keypoint detector, our approach transforms a proof-of-concept uplifting method into a practical, robust, and high-performing end-to-end application for 3D table tennis trajectory and spin analysis.
△ Less
Submitted 25 November, 2025;
originally announced November 2025.
-
Fairness Without Labels: Pseudo-Balancing for Bias Mitigation in Face Gender Classification
Authors:
Haohua Dong,
Ana Manzano Rodríguez,
Camille Guinaudeau,
Shin'ichi Satoh
Abstract:
Face gender classification models often reflect and amplify demographic biases present in their training data, leading to uneven performance across gender and racial subgroups. We introduce pseudo-balancing, a simple and effective strategy for mitigating such biases in semi-supervised learning. Our method enforces demographic balance during pseudo-label selection, using only unlabeled images from…
▽ More
Face gender classification models often reflect and amplify demographic biases present in their training data, leading to uneven performance across gender and racial subgroups. We introduce pseudo-balancing, a simple and effective strategy for mitigating such biases in semi-supervised learning. Our method enforces demographic balance during pseudo-label selection, using only unlabeled images from a race-balanced dataset without requiring access to ground-truth annotations.
We evaluate pseudo-balancing under two conditions: (1) fine-tuning a biased gender classifier using unlabeled images from the FairFace dataset, and (2) stress-testing the method with intentionally imbalanced training data to simulate controlled bias scenarios. In both cases, models are evaluated on the All-Age-Faces (AAF) benchmark, which contains a predominantly East Asian population. Our results show that pseudo-balancing consistently improves fairness while preserving or enhancing accuracy. The method achieves 79.81% overall accuracy - a 6.53% improvement over the baseline - and reduces the gender accuracy gap by 44.17%. In the East Asian subgroup, where baseline disparities exceeded 49%, the gap is narrowed to just 5.01%. These findings suggest that even in the absence of label supervision, access to a demographically balanced or moderately skewed unlabeled dataset can serve as a powerful resource for debiasing existing computer vision models.
△ Less
Submitted 11 October, 2025;
originally announced October 2025.
-
ReSeDis: A Dataset for Referring-based Object Search across Large-Scale Image Collections
Authors:
Ziling Huang,
Yidan Zhang,
Shin'ichi Satoh
Abstract:
Large-scale visual search engines are expected to solve a dual problem at once: (i) locate every image that truly contains the object described by a sentence and (ii) identify the object's bounding box or exact pixels within each hit. Existing techniques address only one side of this challenge. Visual grounding yields tight boxes and masks but rests on the unrealistic assumption that the object is…
▽ More
Large-scale visual search engines are expected to solve a dual problem at once: (i) locate every image that truly contains the object described by a sentence and (ii) identify the object's bounding box or exact pixels within each hit. Existing techniques address only one side of this challenge. Visual grounding yields tight boxes and masks but rests on the unrealistic assumption that the object is present in every test image, producing a flood of false alarms when applied to web-scale collections. Text-to-image retrieval excels at sifting through massive databases to rank relevant images, yet it stops at whole-image matches and offers no fine-grained localization. We introduce Referring Search and Discovery (ReSeDis), the first task that unifies corpus-level retrieval with pixel-level grounding. Given a free-form description, a ReSeDis model must decide whether the queried object appears in each image and, if so, where it is, returning bounding boxes or segmentation masks. To enable rigorous study, we curate a benchmark in which every description maps uniquely to object instances scattered across a large, diverse corpus, eliminating unintended matches. We further design a task-specific metric that jointly scores retrieval recall and localization precision. Finally, we provide a straightforward zero-shot baseline using a frozen vision-language model, revealing significant headroom for future study. ReSeDis offers a realistic, end-to-end testbed for building the next generation of robust and scalable multimodal search systems.
△ Less
Submitted 18 June, 2025;
originally announced June 2025.
-
Foundation Model Insights and a Multi-Model Approach for Superior Fine-Grained One-shot Subset Selection
Authors:
Zhijing Wan,
Zhixiang Wang,
Zheng Wang,
Xin Xu,
Shin'ichi Satoh
Abstract:
One-shot subset selection serves as an effective tool to reduce deep learning training costs by identifying an informative data subset based on the information extracted by an information extractor (IE). Traditional IEs, typically pre-trained on the target dataset, are inherently dataset-dependent. Foundation models (FMs) offer a promising alternative, potentially mitigating this limitation. This…
▽ More
One-shot subset selection serves as an effective tool to reduce deep learning training costs by identifying an informative data subset based on the information extracted by an information extractor (IE). Traditional IEs, typically pre-trained on the target dataset, are inherently dataset-dependent. Foundation models (FMs) offer a promising alternative, potentially mitigating this limitation. This work investigates two key questions: (1) Can FM-based subset selection outperform traditional IE-based methods across diverse datasets? (2) Do all FMs perform equally well as IEs for subset selection? Extensive experiments uncovered surprising insights: FMs consistently outperform traditional IEs on fine-grained datasets, whereas their advantage diminishes on coarse-grained datasets with noisy labels. Motivated by these finding, we propose RAM-APL (RAnking Mean-Accuracy of Pseudo-class Labels), a method tailored for fine-grained image datasets. RAM-APL leverages multiple FMs to enhance subset selection by exploiting their complementary strengths. Our approach achieves state-of-the-art performance on fine-grained datasets, including Oxford-IIIT Pet, Food-101, and Caltech-UCSD Birds-200-2011.
△ Less
Submitted 27 June, 2025; v1 submitted 17 June, 2025;
originally announced June 2025.
-
A Culturally-diverse Multilingual Multimodal Video Benchmark & Model
Authors:
Bhuiyan Sanjid Shafique,
Ashmal Vayani,
Muhammad Maaz,
Hanoona Abdul Rasheed,
Dinura Dissanayake,
Mohammed Irfan Kurpath,
Yahya Hmaiti,
Go Inoue,
Jean Lahoud,
Md. Safirur Rashid,
Shadid Intisar Quasem,
Maheen Fatima,
Franco Vidal,
Mykola Maslych,
Ketan Pravin More,
Sanoojan Baliah,
Hasindri Watawana,
Yuhao Li,
Fabian Farestam,
Leon Schaller,
Roman Tymtsiv,
Simon Weber,
Hisham Cholakkal,
Ivan Laptev,
Shin'ichi Satoh
, et al. (4 additional authors not shown)
Abstract:
Large multimodal models (LMMs) have recently gained attention due to their effectiveness to understand and generate descriptions of visual content. Most existing LMMs are in English language. While few recent works explore multilingual image LMMs, to the best of our knowledge, moving beyond the English language for cultural and linguistic inclusivity is yet to be investigated in the context of vid…
▽ More
Large multimodal models (LMMs) have recently gained attention due to their effectiveness to understand and generate descriptions of visual content. Most existing LMMs are in English language. While few recent works explore multilingual image LMMs, to the best of our knowledge, moving beyond the English language for cultural and linguistic inclusivity is yet to be investigated in the context of video LMMs. In pursuit of more inclusive video LMMs, we introduce a multilingual Video LMM benchmark, named ViMUL-Bench, to evaluate Video LMMs across 14 languages, including both low- and high-resource languages: English, Chinese, Spanish, French, German, Hindi, Arabic, Russian, Bengali, Urdu, Sinhala, Tamil, Swedish, and Japanese. Our ViMUL-Bench is designed to rigorously test video LMMs across 15 categories including eight culturally diverse categories, ranging from lifestyles and festivals to foods and rituals and from local landmarks to prominent cultural personalities. ViMUL-Bench comprises both open-ended (short and long-form) and multiple-choice questions spanning various video durations (short, medium, and long) with 8k samples that are manually verified by native language speakers. In addition, we also introduce a machine translated multilingual video training set comprising 1.2 million samples and develop a simple multilingual video LMM, named ViMUL, that is shown to provide a better tradeoff between high-and low-resource languages for video understanding. We hope our ViMUL-Bench and multilingual video LMM along with a large-scale multilingual video training set will help ease future research in developing cultural and linguistic inclusive multilingual video LMMs. Our proposed benchmark, video LMM and training data will be publicly released at https://mbzuai-oryx.github.io/ViMUL/.
△ Less
Submitted 29 September, 2025; v1 submitted 8 June, 2025;
originally announced June 2025.
-
Probabilistic Online Event Downsampling
Authors:
Andreu Girbau-Xalabarder,
Jun Nagata,
Shinichi Sumiyoshi,
Ricard Marsal,
Shin'ichi Satoh
Abstract:
Event cameras capture scene changes asynchronously on a per-pixel basis, enabling extremely high temporal resolution. However, this advantage comes at the cost of high bandwidth, memory, and computational demands. To address this, prior work has explored event downsampling, but most approaches rely on fixed heuristics or threshold-based strategies, limiting their adaptability. Instead, we propose…
▽ More
Event cameras capture scene changes asynchronously on a per-pixel basis, enabling extremely high temporal resolution. However, this advantage comes at the cost of high bandwidth, memory, and computational demands. To address this, prior work has explored event downsampling, but most approaches rely on fixed heuristics or threshold-based strategies, limiting their adaptability. Instead, we propose a probabilistic framework, POLED, that models event importance through an event-importance probability density function (ePDF), which can be arbitrarily defined and adapted to different applications. Our approach operates in a purely online setting, estimating event importance on-the-fly from raw event streams, enabling scene-specific adaptation. Additionally, we introduce zero-shot event downsampling, where downsampled events must remain usable for models trained on the original event stream, without task-specific adaptation. We design a contour-preserving ePDF that prioritizes structurally important events and evaluate our method across four datasets and tasks--object classification, image interpolation, surface normal estimation, and object detection--demonstrating that intelligent sampling is crucial for maintaining performance under event-budget constraints. Code available.
△ Less
Submitted 23 September, 2025; v1 submitted 3 June, 2025;
originally announced June 2025.
-
Towards Ball Spin and Trajectory Analysis in Table Tennis Broadcast Videos via Physically Grounded Synthetic-to-Real Transfer
Authors:
Daniel Kienzle,
Robin Schön,
Rainer Lienhart,
Shin'Ichi Satoh
Abstract:
Analyzing a player's technique in table tennis requires knowledge of the ball's 3D trajectory and spin. While, the spin is not directly observable in standard broadcasting videos, we show that it can be inferred from the ball's trajectory in the video. We present a novel method to infer the initial spin and 3D trajectory from the corresponding 2D trajectory in a video. Without ground truth labels…
▽ More
Analyzing a player's technique in table tennis requires knowledge of the ball's 3D trajectory and spin. While, the spin is not directly observable in standard broadcasting videos, we show that it can be inferred from the ball's trajectory in the video. We present a novel method to infer the initial spin and 3D trajectory from the corresponding 2D trajectory in a video. Without ground truth labels for broadcast videos, we train a neural network solely on synthetic data. Due to the choice of our input data representation, physically correct synthetic training data, and using targeted augmentations, the network naturally generalizes to real data. Notably, these simple techniques are sufficient to achieve generalization. No real data at all is required for training. To the best of our knowledge, we are the first to present a method for spin and trajectory prediction in simple monocular broadcast videos, achieving an accuracy of 92.0% in spin classification and a 2D reprojection error of 0.19% of the image diagonal.
△ Less
Submitted 28 April, 2025;
originally announced April 2025.
-
SILVIA: Ultra-precision formation flying demonstration for space-based interferometry
Authors:
Takahiro Ito,
Kiwamu Izumi,
Isao Kawano,
Ikkoh Funaki,
Shuichi Sato,
Tomotada Akutsu,
Kentaro Komori,
Mitsuru Musha,
Yuta Michimura,
Satoshi Satoh,
Takuya Iwaki,
Kentaro Yokota,
Kenta Goto,
Katsumi Furukawa,
Taro Matsuo,
Toshihiro Tsuzuki,
Katsuhiko Yamada,
Takahiro Sasaki,
Taisei Nishishita,
Yuki Matsumoto,
Chikako Hirose,
Wataru Torii,
Satoshi Ikari,
Koji Nagano,
Masaki Ando
, et al. (4 additional authors not shown)
Abstract:
We propose SILVIA (Space Interferometer Laboratory Voyaging towards Innovative Applications), a mission concept designed to demonstrate ultra-precision formation flying between three spacecraft separated by 100 m. SILVIA aims to achieve sub-micrometer precision in relative distance control by integrating spacecraft sensors, laser interferometry, low-thrust and low-noise micro-propulsion for real-t…
▽ More
We propose SILVIA (Space Interferometer Laboratory Voyaging towards Innovative Applications), a mission concept designed to demonstrate ultra-precision formation flying between three spacecraft separated by 100 m. SILVIA aims to achieve sub-micrometer precision in relative distance control by integrating spacecraft sensors, laser interferometry, low-thrust and low-noise micro-propulsion for real-time measurement and control of distances and relative orientations between spacecraft. A 100-meter-scale mission in a near-circular low Earth orbit has been identified as an ideal, cost-effective setting for demonstrating SILVIA, as this configuration maintains a good balance between small relative perturbations and low risk for collision. This mission will fill the current technology gap towards future missions, including gravitational wave observatories such as DECIGO (DECihertz Interferometer Gravitational wave Observatory), designed to detect the primordial gravitational wave background, and high-contrast nulling infrared interferometers like LIFE (Large Interferometer for Exoplanets), designed for direct imaging of thermal emissions from nearby terrestrial planet candidates. The mission concept and its key technologies are outlined, paving the way for the next generation of high-precision space-based observatories.
△ Less
Submitted 3 September, 2025; v1 submitted 7 April, 2025;
originally announced April 2025.
-
Hurwitz equivalence in the universal dihedral quandle
Authors:
Takuji Nakamura,
Yasutaka Nakanishi,
Shin Satoh,
Kodai Wada
Abstract:
We investigate the Hurwitz action of the $m$-braid group on the $m$-fold Cartesian product of the universal dihedral quandle. We introduce three computable invariants and prove that they give a complete classification of the orbits under this action. As a consequence, we describe an explicit complete system of orbit representatives. We further obtain analogous classifications for the corresponding…
▽ More
We investigate the Hurwitz action of the $m$-braid group on the $m$-fold Cartesian product of the universal dihedral quandle. We introduce three computable invariants and prove that they give a complete classification of the orbits under this action. As a consequence, we describe an explicit complete system of orbit representatives. We further obtain analogous classifications for the corresponding Hurwitz actions of the pure $m$-braid group, the virtual $m$-braid group, and the virtual pure $m$-braid group.
△ Less
Submitted 18 March, 2026; v1 submitted 15 October, 2024;
originally announced October 2024.
-
Matting by Generation
Authors:
Zhixiang Wang,
Baiang Li,
Jian Wang,
Yu-Lun Liu,
Jinwei Gu,
Yung-Yu Chuang,
Shin'ichi Satoh
Abstract:
This paper introduces an innovative approach for image matting that redefines the traditional regression-based task as a generative modeling challenge. Our method harnesses the capabilities of latent diffusion models, enriched with extensive pre-trained knowledge, to regularize the matting process. We present novel architectural innovations that empower our model to produce mattes with superior re…
▽ More
This paper introduces an innovative approach for image matting that redefines the traditional regression-based task as a generative modeling challenge. Our method harnesses the capabilities of latent diffusion models, enriched with extensive pre-trained knowledge, to regularize the matting process. We present novel architectural innovations that empower our model to produce mattes with superior resolution and detail. The proposed method is versatile and can perform both guidance-free and guidance-based image matting, accommodating a variety of additional cues. Our comprehensive evaluation across three benchmark datasets demonstrates the superior performance of our approach, both quantitatively and qualitatively. The results not only reflect our method's robust effectiveness but also highlight its ability to generate visually compelling mattes that approach photorealistic quality. The project page for this paper is available at https://lightchaserx.github.io/matting-by-generation/
△ Less
Submitted 30 July, 2024;
originally announced July 2024.
-
The SkatingVerse Workshop & Challenge: Methods and Results
Authors:
Jian Zhao,
Lei Jin,
Jianshu Li,
Zheng Zhu,
Yinglei Teng,
Jiaojiao Zhao,
Sadaf Gulshad,
Zheng Wang,
Bo Zhao,
Xiangbo Shu,
Yunchao Wei,
Xuecheng Nie,
Xiaojie Jin,
Xiaodan Liang,
Shin'ichi Satoh,
Yandong Guo,
Cewu Lu,
Junliang Xing,
Jane Shen Shengmei
Abstract:
The SkatingVerse Workshop & Challenge aims to encourage research in developing novel and accurate methods for human action understanding. The SkatingVerse dataset used for the SkatingVerse Challenge has been publicly released. There are two subsets in the dataset, i.e., the training subset and testing subset. The training subsets consists of 19,993 RGB video sequences, and the testing subsets cons…
▽ More
The SkatingVerse Workshop & Challenge aims to encourage research in developing novel and accurate methods for human action understanding. The SkatingVerse dataset used for the SkatingVerse Challenge has been publicly released. There are two subsets in the dataset, i.e., the training subset and testing subset. The training subsets consists of 19,993 RGB video sequences, and the testing subsets consists of 8,586 RGB video sequences. Around 10 participating teams from the globe competed in the SkatingVerse Challenge. In this paper, we provide a brief summary of the SkatingVerse Workshop & Challenge including brief introductions to the top three methods. The submission leaderboard will be reopened for researchers that are interested in the human action understanding challenge. The benchmark dataset and other information can be found at: https://skatingverse.github.io/.
△ Less
Submitted 27 May, 2024;
originally announced May 2024.
-
TC-OCR: TableCraft OCR for Efficient Detection & Recognition of Table Structure & Content
Authors:
Avinash Anand,
Raj Jaiswal,
Pijush Bhuyan,
Mohit Gupta,
Siddhesh Bangar,
Md. Modassir Imam,
Rajiv Ratn Shah,
Shin'ichi Satoh
Abstract:
The automatic recognition of tabular data in document images presents a significant challenge due to the diverse range of table styles and complex structures. Tables offer valuable content representation, enhancing the predictive capabilities of various systems such as search engines and Knowledge Graphs. Addressing the two main problems, namely table detection (TD) and table structure recognition…
▽ More
The automatic recognition of tabular data in document images presents a significant challenge due to the diverse range of table styles and complex structures. Tables offer valuable content representation, enhancing the predictive capabilities of various systems such as search engines and Knowledge Graphs. Addressing the two main problems, namely table detection (TD) and table structure recognition (TSR), has traditionally been approached independently. In this research, we propose an end-to-end pipeline that integrates deep learning models, including DETR, CascadeTabNet, and PP OCR v2, to achieve comprehensive image-based table recognition. This integrated approach effectively handles diverse table styles, complex structures, and image distortions, resulting in improved accuracy and efficiency compared to existing methods like Table Transformers. Our system achieves simultaneous table detection (TD), table structure recognition (TSR), and table content recognition (TCR), preserving table structures and accurately extracting tabular data from document images. The integration of multiple models addresses the intricacies of table recognition, making our approach a promising solution for image-based table understanding, data extraction, and information retrieval applications. Our proposed approach achieves an IOU of 0.96 and an OCR Accuracy of 78%, showcasing a remarkable improvement of approximately 25% in the OCR Accuracy compared to the previous Table Transformer approach.
△ Less
Submitted 19 April, 2024; v1 submitted 16 April, 2024;
originally announced April 2024.
-
RanLayNet: A Dataset for Document Layout Detection used for Domain Adaptation and Generalization
Authors:
Avinash Anand,
Raj Jaiswal,
Mohit Gupta,
Siddhesh S Bangar,
Pijush Bhuyan,
Naman Lal,
Rajeev Singh,
Ritika Jha,
Rajiv Ratn Shah,
Shin'ichi Satoh
Abstract:
Large ground-truth datasets and recent advances in deep learning techniques have been useful for layout detection. However, because of the restricted layout diversity of these datasets, training on them requires a sizable number of annotated instances, which is both expensive and time-consuming. As a result, differences between the source and target domains may significantly impact how well these…
▽ More
Large ground-truth datasets and recent advances in deep learning techniques have been useful for layout detection. However, because of the restricted layout diversity of these datasets, training on them requires a sizable number of annotated instances, which is both expensive and time-consuming. As a result, differences between the source and target domains may significantly impact how well these models function. To solve this problem, domain adaptation approaches have been developed that use a small quantity of labeled data to adjust the model to the target domain. In this research, we introduced a synthetic document dataset called RanLayNet, enriched with automatically assigned labels denoting spatial positions, ranges, and types of layout elements. The primary aim of this endeavor is to develop a versatile dataset capable of training models with robustness and adaptability to diverse document formats. Through empirical experimentation, we demonstrate that a deep layout identification model trained on our dataset exhibits enhanced performance compared to a model trained solely on actual documents. Moreover, we conduct a comparative analysis by fine-tuning inference models using both PubLayNet and IIIT-AR-13K datasets on the Doclaynet dataset. Our findings emphasize that models enriched with our dataset are optimal for tasks such as achieving 0.398 and 0.588 mAP95 score in the scientific document domain for the TABLE class.
△ Less
Submitted 19 April, 2024; v1 submitted 15 April, 2024;
originally announced April 2024.
-
The Effects of Short Video-Sharing Services on Video Copy Detection
Authors:
Rintaro Yanagi,
Yamato Okamoto,
Shuhei Yokoo,
Shin'ichi Satoh
Abstract:
The short video-sharing services that allow users to post 10-30 second videos (e.g., YouTube Shorts and TikTok) have attracted a lot of attention in recent years. However, conventional video copy detection (VCD) methods mainly focus on general video-sharing services (e.g., YouTube and Bilibili), and the effects of short video-sharing services on video copy detection are still unclear. Considering…
▽ More
The short video-sharing services that allow users to post 10-30 second videos (e.g., YouTube Shorts and TikTok) have attracted a lot of attention in recent years. However, conventional video copy detection (VCD) methods mainly focus on general video-sharing services (e.g., YouTube and Bilibili), and the effects of short video-sharing services on video copy detection are still unclear. Considering that illegally copied videos in short video-sharing services have service-distinctive characteristics, especially in those time lengths, the pros and cons of VCD in those services are required to be analyzed. In this paper, we examine the effects of short video-sharing services on VCD by constructing a dataset that has short video-sharing service characteristics. Our novel dataset is automatically constructed from the publicly available dataset to have reference videos and fixed short-time-length query videos, and such automation procedures assure the reproducibility and data privacy preservation of this paper. From the experimental results focusing on segment-level and video-level situations, we can see that three effects: "Segment-level VCD in short video-sharing services is more difficult than those in general video-sharing services", "Video-level VCD in short video-sharing services is easier than those in general video-sharing services", "The video alignment component mainly suppress the detection performance in short video-sharing services".
△ Less
Submitted 26 March, 2024;
originally announced March 2024.
-
On wen knots
Authors:
Celeste Damiani,
Shin Satoh
Abstract:
We introduce the notion of wen knots, and prove that the set of wen knots is a proper subset of the set of extended welded knots. Furthermore we prove that the complementary subset consists of welded knots up to horizontal mirror reflections. This allow us to characterise completely extended welded knots by the parity of their number of wens, that we can always reduce to 0 or 1.
We introduce the notion of wen knots, and prove that the set of wen knots is a proper subset of the set of extended welded knots. Furthermore we prove that the complementary subset consists of welded knots up to horizontal mirror reflections. This allow us to characterise completely extended welded knots by the parity of their number of wens, that we can always reduce to 0 or 1.
△ Less
Submitted 6 March, 2024;
originally announced March 2024.
-
Contributing Dimension Structure of Deep Feature for Coreset Selection
Authors:
Zhijing Wan,
Zhixiang Wang,
Yuran Wang,
Zheng Wang,
Hongyuan Zhu,
Shin'ichi Satoh
Abstract:
Coreset selection seeks to choose a subset of crucial training samples for efficient learning. It has gained traction in deep learning, particularly with the surge in training dataset sizes. Sample selection hinges on two main aspects: a sample's representation in enhancing performance and the role of sample diversity in averting overfitting. Existing methods typically measure both the representat…
▽ More
Coreset selection seeks to choose a subset of crucial training samples for efficient learning. It has gained traction in deep learning, particularly with the surge in training dataset sizes. Sample selection hinges on two main aspects: a sample's representation in enhancing performance and the role of sample diversity in averting overfitting. Existing methods typically measure both the representation and diversity of data based on similarity metrics, such as L2-norm. They have capably tackled representation via distribution matching guided by the similarities of features, gradients, or other information between data. However, the results of effectively diverse sample selection are mired in sub-optimality. This is because the similarity metrics usually simply aggregate dimension similarities without acknowledging disparities among the dimensions that significantly contribute to the final similarity. As a result, they fall short of adequately capturing diversity. To address this, we propose a feature-based diversity constraint, compelling the chosen subset to exhibit maximum diversity. Our key lies in the introduction of a novel Contributing Dimension Structure (CDS) metric. Different from similarity metrics that measure the overall similarity of high-dimensional features, our CDS metric considers not only the reduction of redundancy in feature dimensions, but also the difference between dimensions that contribute significantly to the final similarity. We reveal that existing methods tend to favor samples with similar CDS, leading to a reduced variety of CDS types within the coreset and subsequently hindering model performance. In response, we enhance the performance of five classical selection methods by integrating the CDS constraint. Our experiments on three datasets demonstrate the general effectiveness of the proposed method in boosting existing methods.
△ Less
Submitted 2 March, 2024; v1 submitted 29 January, 2024;
originally announced January 2024.
-
Virtualized Delta, sharp, and pass moves for oriented virtual knots and links
Authors:
Takuji Nakamura,
Yasutaka Nakanishi,
Shin Satoh,
Kodai Wada
Abstract:
We study virtualized Delta, sharp, and pass moves for oriented virtual links, and give necessary and sufficient conditions for two oriented virtual links to be related by the local moves. In particular, they are unknotting operations for oriented virtual knots. We provide lower bounds for the unknotting numbers and prove that they are best possible.
We study virtualized Delta, sharp, and pass moves for oriented virtual links, and give necessary and sufficient conditions for two oriented virtual links to be related by the local moves. In particular, they are unknotting operations for oriented virtual knots. We provide lower bounds for the unknotting numbers and prove that they are best possible.
△ Less
Submitted 23 January, 2024;
originally announced January 2024.
-
Virtualized Delta moves for virtual knots and links
Authors:
Takuji Nakamura,
Yasutaka Nakanishi,
Shin Satoh,
Kodai Wada
Abstract:
We introduce a local deformation called the virtualized $Δ$-move for virtual knots and links. We prove that the virtualized $Δ$-move is an unknotting operation for virtual knots. Furthermore we give a necessary and sufficient condition for two virtual links to be related by a finite sequence of virtualized $Δ$-moves.
We introduce a local deformation called the virtualized $Δ$-move for virtual knots and links. We prove that the virtualized $Δ$-move is an unknotting operation for virtual knots. Furthermore we give a necessary and sufficient condition for two virtual links to be related by a finite sequence of virtualized $Δ$-moves.
△ Less
Submitted 23 January, 2024;
originally announced January 2024.
-
Nonlinear steering control under input magnitude and rate constraints with exponential convergence
Authors:
Rin Suyama,
Satoshi Satoh,
Atsuo Maki
Abstract:
A ship steering control is designed for a nonlinear maneuvering model whose rudder manipulation is constrained in both magnitude and rate. In our method, the tracking problem of the target heading angle with input constraints is converted into the tracking problem for a strict-feedback system without any input constraints. To derive this system, hyperbolic tangent ($\tanh$) function and auxiliary…
▽ More
A ship steering control is designed for a nonlinear maneuvering model whose rudder manipulation is constrained in both magnitude and rate. In our method, the tracking problem of the target heading angle with input constraints is converted into the tracking problem for a strict-feedback system without any input constraints. To derive this system, hyperbolic tangent ($\tanh$) function and auxiliary variables are introduced to deal with the input constraints. Furthermore, using the feature of the derivative of $\tanh$ function, auxiliary systems are successfully derived in the strict-feedback form. The backstepping method is utilized to construct the feedback control law for the resulting cascade system. The proposed steering control is verified in numerical experiments, and the result shows that the tracking of the target heading angle is successful using the proposed control law.
△ Less
Submitted 25 October, 2023;
originally announced October 2023.
-
Beyond Domain Gap: Exploiting Subjectivity in Sketch-Based Person Retrieval
Authors:
Kejun Lin,
Zhixiang Wang,
Zheng Wang,
Yinqiang Zheng,
Shin'ichi Satoh
Abstract:
Person re-identification (re-ID) requires densely distributed cameras. In practice, the person of interest may not be captured by cameras and, therefore, needs to be retrieved using subjective information (e.g., sketches from witnesses). Previous research defines this case using the sketch as sketch re-identification (Sketch re-ID) and focuses on eliminating the domain gap. Actually, subjectivity…
▽ More
Person re-identification (re-ID) requires densely distributed cameras. In practice, the person of interest may not be captured by cameras and, therefore, needs to be retrieved using subjective information (e.g., sketches from witnesses). Previous research defines this case using the sketch as sketch re-identification (Sketch re-ID) and focuses on eliminating the domain gap. Actually, subjectivity is another significant challenge. We model and investigate it by posing a new dataset with multi-witness descriptions. It features two aspects. 1) Large-scale. It contains over 4,763 sketches and 32,668 photos, making it the largest Sketch re-ID dataset. 2) Multi-perspective and multi-style. Our dataset offers multiple sketches for each identity. Witnesses' subjective cognition provides multiple perspectives on the same individual, while different artists' drawing styles provide variation in sketch styles. We further have two novel designs to alleviate the challenge of subjectivity. 1) Fusing subjectivity. We propose a non-local (NL) fusion module that gathers sketches from different witnesses for the same identity. 2) Introducing objectivity. An AttrAlign module utilizes attributes as an implicit mask to align cross-domain features. To push forward the advance of Sketch re-ID, we set three benchmarks (large-scale, multi-style, cross-style). Extensive experiments demonstrate our leading performance in these benchmarks. Dataset and Codes are publicly available at: https://github.com/Lin-Kayla/subjectivity-sketch-reid
△ Less
Submitted 15 September, 2023;
originally announced September 2023.
-
TransRef: Multi-Scale Reference Embedding Transformer for Reference-Guided Image Inpainting
Authors:
Taorong Liu,
Liang Liao,
Delin Chen,
Jing Xiao,
Zheng Wang,
Chia-Wen Lin,
Shin'ichi Satoh
Abstract:
Image inpainting for completing complicated semantic environments and diverse hole patterns of corrupted images is challenging even for state-of-the-art learning-based inpainting methods trained on large-scale data. A reference image capturing the same scene of a corrupted image offers informative guidance for completing the corrupted image as it shares similar texture and structure priors to that…
▽ More
Image inpainting for completing complicated semantic environments and diverse hole patterns of corrupted images is challenging even for state-of-the-art learning-based inpainting methods trained on large-scale data. A reference image capturing the same scene of a corrupted image offers informative guidance for completing the corrupted image as it shares similar texture and structure priors to that of the holes of the corrupted image. In this work, we propose a transformer-based encoder-decoder network, named TransRef, for reference-guided image inpainting. Specifically, the guidance is conducted progressively through a reference embedding procedure, in which the referencing features are subsequently aligned and fused with the features of the corrupted image. For precise utilization of the reference features for guidance, a reference-patch alignment (Ref-PA) module is proposed to align the patch features of the reference and corrupted images and harmonize their style differences, while a reference-patch transformer (Ref-PT) module is proposed to refine the embedded reference feature. Moreover, to facilitate the research of reference-guided image restoration tasks, we construct a publicly accessible benchmark dataset containing 50K pairs of input and reference images. Both quantitative and qualitative evaluations demonstrate the efficacy of the reference information and the proposed method over the state-of-the-art methods in completing complex holes. Code and dataset can be accessed at https://github.com/Cameltr/TransRef.
△ Less
Submitted 11 February, 2025; v1 submitted 20 June, 2023;
originally announced June 2023.
-
The 3rd Anti-UAV Workshop & Challenge: Methods and Results
Authors:
Jian Zhao,
Jianan Li,
Lei Jin,
Jiaming Chu,
Zhihao Zhang,
Jun Wang,
Jiangqiang Xia,
Kai Wang,
Yang Liu,
Sadaf Gulshad,
Jiaojiao Zhao,
Tianyang Xu,
Xuefeng Zhu,
Shihan Liu,
Zheng Zhu,
Guibo Zhu,
Zechao Li,
Zheng Wang,
Baigui Sun,
Yandong Guo,
Shin ichi Satoh,
Junliang Xing,
Jane Shen Shengmei
Abstract:
The 3rd Anti-UAV Workshop & Challenge aims to encourage research in developing novel and accurate methods for multi-scale object tracking. The Anti-UAV dataset used for the Anti-UAV Challenge has been publicly released. There are two main differences between this year's competition and the previous two. First, we have expanded the existing dataset, and for the first time, released a training set s…
▽ More
The 3rd Anti-UAV Workshop & Challenge aims to encourage research in developing novel and accurate methods for multi-scale object tracking. The Anti-UAV dataset used for the Anti-UAV Challenge has been publicly released. There are two main differences between this year's competition and the previous two. First, we have expanded the existing dataset, and for the first time, released a training set so that participants can focus on improving their models. Second, we set up two tracks for the first time, i.e., Anti-UAV Tracking and Anti-UAV Detection & Tracking. Around 76 participating teams from the globe competed in the 3rd Anti-UAV Challenge. In this paper, we provide a brief summary of the 3rd Anti-UAV Workshop & Challenge including brief introductions to the top three methods in each track. The submission leaderboard will be reopened for researchers that are interested in the Anti-UAV challenge. The benchmark dataset and other information can be found at: https://anti-uav.github.io/.
△ Less
Submitted 15 July, 2023; v1 submitted 12 May, 2023;
originally announced May 2023.
-
Certified Zeroth-order Black-Box Defense with Robust UNet Denoiser
Authors:
Astha Verma,
A V Subramanyam,
Siddhesh Bangar,
Naman Lal,
Rajiv Ratn Shah,
Shin'ichi Satoh
Abstract:
Certified defense methods against adversarial perturbations have been recently investigated in the black-box setting with a zeroth-order (ZO) perspective. However, these methods suffer from high model variance with low performance on high-dimensional datasets due to the ineffective design of the denoiser and are limited in their utilization of ZO techniques. To this end, we propose a certified ZO…
▽ More
Certified defense methods against adversarial perturbations have been recently investigated in the black-box setting with a zeroth-order (ZO) perspective. However, these methods suffer from high model variance with low performance on high-dimensional datasets due to the ineffective design of the denoiser and are limited in their utilization of ZO techniques. To this end, we propose a certified ZO preprocessing technique for removing adversarial perturbations from the attacked image in the black-box setting using only model queries. We propose a robust UNet denoiser (RDUNet) that ensures the robustness of black-box models trained on high-dimensional datasets. We propose a novel black-box denoised smoothing (DS) defense mechanism, ZO-RUDS, by prepending our RDUNet to the black-box model, ensuring black-box defense. We further propose ZO-AE-RUDS in which RDUNet followed by autoencoder (AE) is prepended to the black-box model. We perform extensive experiments on four classification datasets, CIFAR-10, CIFAR-10, Tiny Imagenet, STL-10, and the MNIST dataset for image reconstruction tasks. Our proposed defense methods ZO-RUDS and ZO-AE-RUDS beat SOTA with a huge margin of $35\%$ and $9\%$, for low dimensional (CIFAR-10) and with a margin of $20.61\%$ and $23.51\%$ for high-dimensional (STL-10) datasets, respectively.
△ Less
Submitted 6 July, 2024; v1 submitted 13 April, 2023;
originally announced April 2023.
-
Toward Verifiable and Reproducible Human Evaluation for Text-to-Image Generation
Authors:
Mayu Otani,
Riku Togashi,
Yu Sawai,
Ryosuke Ishigami,
Yuta Nakashima,
Esa Rahtu,
Janne Heikkilä,
Shin'ichi Satoh
Abstract:
Human evaluation is critical for validating the performance of text-to-image generative models, as this highly cognitive process requires deep comprehension of text and images. However, our survey of 37 recent papers reveals that many works rely solely on automatic measures (e.g., FID) or perform poorly described human evaluations that are not reliable or repeatable. This paper proposes a standard…
▽ More
Human evaluation is critical for validating the performance of text-to-image generative models, as this highly cognitive process requires deep comprehension of text and images. However, our survey of 37 recent papers reveals that many works rely solely on automatic measures (e.g., FID) or perform poorly described human evaluations that are not reliable or repeatable. This paper proposes a standardized and well-defined human evaluation protocol to facilitate verifiable and reproducible human evaluation in future works. In our pilot data collection, we experimentally show that the current automatic measures are incompatible with human perception in evaluating the performance of the text-to-image generation results. Furthermore, we provide insights for designing human evaluation experiments reliably and conclusively. Finally, we make several resources publicly available to the community to facilitate easy and fast implementations.
△ Less
Submitted 4 April, 2023;
originally announced April 2023.
-
DisCO: Portrait Distortion Correction with Perspective-Aware 3D GANs
Authors:
Zhixiang Wang,
Yu-Lun Liu,
Jia-Bin Huang,
Shin'ichi Satoh,
Sizhuo Ma,
Gurunandan Krishnan,
Jian Wang
Abstract:
Close-up facial images captured at short distances often suffer from perspective distortion, resulting in exaggerated facial features and unnatural/unattractive appearances. We propose a simple yet effective method for correcting perspective distortions in a single close-up face. We first perform GAN inversion using a perspective-distorted input facial image by jointly optimizing the camera intrin…
▽ More
Close-up facial images captured at short distances often suffer from perspective distortion, resulting in exaggerated facial features and unnatural/unattractive appearances. We propose a simple yet effective method for correcting perspective distortions in a single close-up face. We first perform GAN inversion using a perspective-distorted input facial image by jointly optimizing the camera intrinsic/extrinsic parameters and face latent code. To address the ambiguity of joint optimization, we develop starting from a short distance, optimization scheduling, reparametrizations, and geometric regularization. Re-rendering the portrait at a proper focal length and camera distance effectively corrects perspective distortions and produces more natural-looking results. Our experiments show that our method compares favorably against previous approaches qualitatively and quantitatively. We showcase numerous examples validating the applicability of our method on in-the-wild portrait photos. We will release our code and the evaluation protocol to facilitate future work.
△ Less
Submitted 8 December, 2023; v1 submitted 23 February, 2023;
originally announced February 2023.
-
HOTCOLD Block: Fooling Thermal Infrared Detectors with a Novel Wearable Design
Authors:
Hui Wei,
Zhixiang Wang,
Xuemei Jia,
Yinqiang Zheng,
Hao Tang,
Shin'ichi Satoh,
Zheng Wang
Abstract:
Adversarial attacks on thermal infrared imaging expose the risk of related applications. Estimating the security of these systems is essential for safely deploying them in the real world. In many cases, realizing the attacks in the physical space requires elaborate special perturbations. These solutions are often \emph{impractical} and \emph{attention-grabbing}. To address the need for a physicall…
▽ More
Adversarial attacks on thermal infrared imaging expose the risk of related applications. Estimating the security of these systems is essential for safely deploying them in the real world. In many cases, realizing the attacks in the physical space requires elaborate special perturbations. These solutions are often \emph{impractical} and \emph{attention-grabbing}. To address the need for a physically practical and stealthy adversarial attack, we introduce \textsc{HotCold} Block, a novel physical attack for infrared detectors that hide persons utilizing the wearable Warming Paste and Cooling Paste. By attaching these readily available temperature-controlled materials to the body, \textsc{HotCold} Block evades human eyes efficiently. Moreover, unlike existing methods that build adversarial patches with complex texture and structure features, \textsc{HotCold} Block utilizes an SSP-oriented adversarial optimization algorithm that enables attacks with pure color blocks and explores the influence of size, shape, and position on attack performance. Extensive experimental results in both digital and physical environments demonstrate the performance of our proposed \textsc{HotCold} Block. \emph{Code is available: \textcolor{magenta}{https://github.com/weihui1308/HOTCOLDBlock}}.
△ Less
Submitted 12 December, 2022;
originally announced December 2022.
-
Self-distillation with Online Diffusion on Batch Manifolds Improves Deep Metric Learning
Authors:
Zelong Zeng,
Fan Yang,
Hong Liu,
Shin'ichi Satoh
Abstract:
Recent deep metric learning (DML) methods typically leverage solely class labels to keep positive samples far away from negative ones. However, this type of method normally ignores the crucial knowledge hidden in the data (e.g., intra-class information variation), which is harmful to the generalization of the trained model. To alleviate this problem, in this paper we propose Online Batch Diffusion…
▽ More
Recent deep metric learning (DML) methods typically leverage solely class labels to keep positive samples far away from negative ones. However, this type of method normally ignores the crucial knowledge hidden in the data (e.g., intra-class information variation), which is harmful to the generalization of the trained model. To alleviate this problem, in this paper we propose Online Batch Diffusion-based Self-Distillation (OBD-SD) for DML. Specifically, we first propose a simple but effective Progressive Self-Distillation (PSD), which distills the knowledge progressively from the model itself during training. The soft distance targets achieved by PSD can present richer relational information among samples, which is beneficial for the diversity of embedding representations. Then, we extend PSD with an Online Batch Diffusion Process (OBDP), which is to capture the local geometric structure of manifolds in each batch, so that it can reveal the intrinsic relationships among samples in the batch and produce better soft distance targets. Note that our OBDP is able to restore the insufficient manifold relationships obtained by the original PSD and achieve significant performance improvement. Our OBD-SD is a flexible framework that can be integrated into state-of-the-art (SOTA) DML methods. Extensive experiments on various benchmarks, namely CUB200, CARS196, and Stanford Online Products, demonstrate that our OBD-SD consistently improves the performance of the existing DML methods on multiple datasets with negligible additional training time, achieving very competitive results. Code: \url{https://github.com/ZelongZeng/OBD-SD_Pytorch}
△ Less
Submitted 14 November, 2022;
originally announced November 2022.
-
Multiple Object Tracking from appearance by hierarchically clustering tracklets
Authors:
Andreu Girbau,
Ferran Marqués,
Shin'ichi Satoh
Abstract:
Current approaches in Multiple Object Tracking (MOT) rely on the spatio-temporal coherence between detections combined with object appearance to match objects from consecutive frames. In this work, we explore MOT using object appearances as the main source of association between objects in a video, using spatial and temporal priors as weighting factors. We form initial tracklets by leveraging on t…
▽ More
Current approaches in Multiple Object Tracking (MOT) rely on the spatio-temporal coherence between detections combined with object appearance to match objects from consecutive frames. In this work, we explore MOT using object appearances as the main source of association between objects in a video, using spatial and temporal priors as weighting factors. We form initial tracklets by leveraging on the idea that instances of an object that are close in time should be similar in appearance, and build the final object tracks by fusing the tracklets in a hierarchical fashion. We conduct extensive experiments that show the effectiveness of our method over three different MOT benchmarks, MOT17, MOT20, and DanceTrack, being competitive in MOT17 and MOT20 and establishing state-of-the-art results in DanceTrack.
△ Less
Submitted 7 October, 2022;
originally announced October 2022.
-
Physical Adversarial Attack meets Computer Vision: A Decade Survey
Authors:
Hui Wei,
Hao Tang,
Xuemei Jia,
Zhixiang Wang,
Hanxun Yu,
Zhubo Li,
Shin'ichi Satoh,
Luc Van Gool,
Zheng Wang
Abstract:
Despite the impressive achievements of Deep Neural Networks (DNNs) in computer vision, their vulnerability to adversarial attacks remains a critical concern. Extensive research has demonstrated that incorporating sophisticated perturbations into input images can lead to a catastrophic degradation in DNNs' performance. This perplexing phenomenon not only exists in the digital space but also in the…
▽ More
Despite the impressive achievements of Deep Neural Networks (DNNs) in computer vision, their vulnerability to adversarial attacks remains a critical concern. Extensive research has demonstrated that incorporating sophisticated perturbations into input images can lead to a catastrophic degradation in DNNs' performance. This perplexing phenomenon not only exists in the digital space but also in the physical world. Consequently, it becomes imperative to evaluate the security of DNNs-based systems to ensure their safe deployment in real-world scenarios, particularly in security-sensitive applications. To facilitate a profound understanding of this topic, this paper presents a comprehensive overview of physical adversarial attacks. Firstly, we distill four general steps for launching physical adversarial attacks. Building upon this foundation, we uncover the pervasive role of artifacts carrying adversarial perturbations in the physical world. These artifacts influence each step. To denote them, we introduce a new term: adversarial medium. Then, we take the first step to systematically evaluate the performance of physical adversarial attacks, taking the adversarial medium as a first attempt. Our proposed evaluation metric, hiPAA, comprises six perspectives: Effectiveness, Stealthiness, Robustness, Practicability, Aesthetics, and Economics. We also provide comparative results across task categories, together with insightful observations and suggestions for future research directions.
△ Less
Submitted 20 November, 2024; v1 submitted 29 September, 2022;
originally announced September 2022.
-
Reference-Guided Texture and Structure Inference for Image Inpainting
Authors:
Taorong Liu,
Liang Liao,
Zheng Wang,
Shin'ichi Satoh
Abstract:
Existing learning-based image inpainting methods are still in challenge when facing complex semantic environments and diverse hole patterns. The prior information learned from the large scale training data is still insufficient for these situations. Reference images captured covering the same scenes share similar texture and structure priors with the corrupted images, which offers new prospects fo…
▽ More
Existing learning-based image inpainting methods are still in challenge when facing complex semantic environments and diverse hole patterns. The prior information learned from the large scale training data is still insufficient for these situations. Reference images captured covering the same scenes share similar texture and structure priors with the corrupted images, which offers new prospects for the image inpainting tasks. Inspired by this, we first build a benchmark dataset containing 10K pairs of input and reference images for reference-guided inpainting. Then we adopt an encoder-decoder structure to separately infer the texture and structure features of the input image considering their pattern discrepancy of texture and structure during inpainting. A feature alignment module is further designed to refine these features of the input image with the guidance of a reference image. Both quantitative and qualitative evaluations demonstrate the superiority of our method over the state-of-the-art methods in terms of completing complex holes.
△ Less
Submitted 29 July, 2022;
originally announced July 2022.
-
Improving Generalization of Metric Learning via Listwise Self-distillation
Authors:
Zelong Zeng,
Fan Yang,
Zheng Wang,
Shin'ichi Satoh
Abstract:
Most deep metric learning (DML) methods employ a strategy that forces all positive samples to be close in the embedding space while keeping them away from negative ones. However, such a strategy ignores the internal relationships of positive (negative) samples and often leads to overfitting, especially in the presence of hard samples and mislabeled samples. In this work, we propose a simple yet ef…
▽ More
Most deep metric learning (DML) methods employ a strategy that forces all positive samples to be close in the embedding space while keeping them away from negative ones. However, such a strategy ignores the internal relationships of positive (negative) samples and often leads to overfitting, especially in the presence of hard samples and mislabeled samples. In this work, we propose a simple yet effective regularization, namely Listwise Self-Distillation (LSD), which progressively distills a model's own knowledge to adaptively assign a more appropriate distance target to each sample pair in a batch. LSD encourages smoother embeddings and information mining within positive (negative) samples as a way to mitigate overfitting and thus improve generalization. Our LSD can be directly integrated into general DML frameworks. Extensive experiments show that LSD consistently boosts the performance of various metric learning methods on multiple datasets.
△ Less
Submitted 17 June, 2022;
originally announced June 2022.
-
Unsupervised Foggy Scene Understanding via Self Spatial-Temporal Label Diffusion
Authors:
Liang Liao,
Wenyi Chen,
Jing Xiao,
Zheng Wang,
Chia-Wen Lin,
Shin'ichi Satoh
Abstract:
Understanding foggy image sequence in the driving scenes is critical for autonomous driving, but it remains a challenging task due to the difficulty in collecting and annotating real-world images of adverse weather. Recently, the self-training strategy has been considered a powerful solution for unsupervised domain adaptation, which iteratively adapts the model from the source domain to the target…
▽ More
Understanding foggy image sequence in the driving scenes is critical for autonomous driving, but it remains a challenging task due to the difficulty in collecting and annotating real-world images of adverse weather. Recently, the self-training strategy has been considered a powerful solution for unsupervised domain adaptation, which iteratively adapts the model from the source domain to the target domain by generating target pseudo labels and re-training the model. However, the selection of confident pseudo labels inevitably suffers from the conflict between sparsity and accuracy, both of which will lead to suboptimal models. To tackle this problem, we exploit the characteristics of the foggy image sequence of driving scenes to densify the confident pseudo labels. Specifically, based on the two discoveries of local spatial similarity and adjacent temporal correspondence of the sequential image data, we propose a novel Target-Domain driven pseudo label Diffusion (TDo-Dif) scheme. It employs superpixels and optical flows to identify the spatial similarity and temporal correspondence, respectively and then diffuses the confident but sparse pseudo labels within a superpixel or a temporal corresponding pair linked by the flow. Moreover, to ensure the feature similarity of the diffused pixels, we introduce local spatial similarity loss and temporal contrastive loss in the model re-training stage. Experimental results show that our TDo-Dif scheme helps the adaptive model achieve 51.92% and 53.84% mean intersection-over-union (mIoU) on two publicly available natural foggy datasets (Foggy Zurich and Foggy Driving), which exceeds the state-of-the-art unsupervised domain adaptive semantic segmentation methods. Models and data can be found at https://github.com/velor2012/TDo-Dif.
△ Less
Submitted 10 June, 2022;
originally announced June 2022.
-
Geo-Localization via Ground-to-Satellite Cross-View Image Retrieval
Authors:
Zelong Zeng,
Zheng Wang,
Fan Yang,
Shin'ichi Satoh
Abstract:
The large variation of viewpoint and irrelevant content around the target always hinder accurate image retrieval and its subsequent tasks. In this paper, we investigate an extremely challenging task: given a ground-view image of a landmark, we aim to achieve cross-view geo-localization by searching out its corresponding satellite-view images. Specifically, the challenge comes from the gap between…
▽ More
The large variation of viewpoint and irrelevant content around the target always hinder accurate image retrieval and its subsequent tasks. In this paper, we investigate an extremely challenging task: given a ground-view image of a landmark, we aim to achieve cross-view geo-localization by searching out its corresponding satellite-view images. Specifically, the challenge comes from the gap between ground-view and satellite-view, which includes not only large viewpoint changes (some parts of the landmark may be invisible from front view to top view) but also highly irrelevant background (the target landmark tend to be hidden in other surrounding buildings), making it difficult to learn a common representation or a suitable mapping.
To address this issue, we take advantage of drone-view information as a bridge between ground-view and satellite-view domains. We propose a Peer Learning and Cross Diffusion (PLCD) framework. PLCD consists of three parts: 1) a peer learning across ground-view and drone-view to find visible parts to benefit ground-drone cross-view representation learning; 2) a patch-based network for satellite-drone cross-view representation learning; 3) a cross diffusion between ground-drone space and satellite-drone space. Extensive experiments conducted on the University-Earth and University-Google datasets show that our method outperforms state-of-the-arts significantly.
△ Less
Submitted 22 May, 2022;
originally announced May 2022.
-
Neural Global Shutter: Learn to Restore Video from a Rolling Shutter Camera with Global Reset Feature
Authors:
Zhixiang Wang,
Xiang Ji,
Jia-Bin Huang,
Shin'ichi Satoh,
Xiao Zhou,
Yinqiang Zheng
Abstract:
Most computer vision systems assume distortion-free images as inputs. The widely used rolling-shutter (RS) image sensors, however, suffer from geometric distortion when the camera and object undergo motion during capture. Extensive researches have been conducted on correcting RS distortions. However, most of the existing work relies heavily on the prior assumptions of scenes or motions. Besides, t…
▽ More
Most computer vision systems assume distortion-free images as inputs. The widely used rolling-shutter (RS) image sensors, however, suffer from geometric distortion when the camera and object undergo motion during capture. Extensive researches have been conducted on correcting RS distortions. However, most of the existing work relies heavily on the prior assumptions of scenes or motions. Besides, the motion estimation steps are either oversimplified or computationally inefficient due to the heavy flow warping, limiting their applicability. In this paper, we investigate using rolling shutter with a global reset feature (RSGR) to restore clean global shutter (GS) videos. This feature enables us to turn the rectification problem into a deblur-like one, getting rid of inaccurate and costly explicit motion estimation. First, we build an optic system that captures paired RSGR/GS videos. Second, we develop a novel algorithm incorporating spatial and temporal designs to correct the spatial-varying RSGR distortion. Third, we demonstrate that existing image-to-image translation algorithms can recover clean GS videos from distorted RSGR inputs, yet our algorithm achieves the best performance with the specific designs. Our rendered results are not only visually appealing but also beneficial to downstream tasks. Compared to the state-of-the-art RS solution, our RSGR solution is superior in both effectiveness and efficiency. Considering it is easy to realize without changing the hardware, we believe our RSGR solution can potentially replace the RS solution in taking distortion-free videos with low noise and low budget.
△ Less
Submitted 25 June, 2022; v1 submitted 2 April, 2022;
originally announced April 2022.
-
Optimal Correction Cost for Object Detection Evaluation
Authors:
Mayu Otani,
Riku Togashi,
Yuta Nakashima,
Esa Rahtu,
Janne Heikkilä,
Shin'ichi Satoh
Abstract:
Mean Average Precision (mAP) is the primary evaluation measure for object detection. Although object detection has a broad range of applications, mAP evaluates detectors in terms of the performance of ranked instance retrieval. Such the assumption for the evaluation task does not suit some downstream tasks. To alleviate the gap between downstream tasks and the evaluation scenario, we propose Optim…
▽ More
Mean Average Precision (mAP) is the primary evaluation measure for object detection. Although object detection has a broad range of applications, mAP evaluates detectors in terms of the performance of ranked instance retrieval. Such the assumption for the evaluation task does not suit some downstream tasks. To alleviate the gap between downstream tasks and the evaluation scenario, we propose Optimal Correction Cost (OC-cost), which assesses detection accuracy at image level. OC-cost computes the cost of correcting detections to ground truths as a measure of accuracy. The cost is obtained by solving an optimal transportation problem between the detections and the ground truths. Unlike mAP, OC-cost is designed to penalize false positive and false negative detections properly, and every image in a dataset is treated equally. Our experimental result validates that OC-cost has better agreement with human preference than a ranking-based measure, i.e., mAP for a single image. We also show that detectors' rankings by OC-cost are more consistent on different data splits than mAP. Our goal is not to replace mAP with OC-cost but provide an additional tool to evaluate detectors from another aspect. To help future researchers and developers choose a target measure, we provide a series of experiments to clarify how mAP and OC-cost differ.
△ Less
Submitted 27 March, 2022;
originally announced March 2022.
-
Improving Camouflaged Object Detection with the Uncertainty of Pseudo-edge Labels
Authors:
Nobukatsu Kajiura,
Hong Liu,
Shin'ichi Satoh
Abstract:
This paper focuses on camouflaged object detection (COD), which is a task to detect objects hidden in the background. Most of the current COD models aim to highlight the target object directly while outputting ambiguous camouflaged boundaries. On the other hand, the performance of the models considering edge information is not yet satisfactory. To this end, we propose a new framework that makes fu…
▽ More
This paper focuses on camouflaged object detection (COD), which is a task to detect objects hidden in the background. Most of the current COD models aim to highlight the target object directly while outputting ambiguous camouflaged boundaries. On the other hand, the performance of the models considering edge information is not yet satisfactory. To this end, we propose a new framework that makes full use of multiple visual cues, i.e., saliency as well as edges, to refine the predicted camouflaged map. This framework consists of three key components, i.e., a pseudo-edge generator, a pseudo-map generator, and an uncertainty-aware refinement module. In particular, the pseudo-edge generator estimates the boundary that outputs the pseudo-edge label, and the conventional COD method serves as the pseudo-map generator that outputs the pseudo-map label. Then, we propose an uncertainty-based module to reduce the uncertainty and noise of such two pseudo labels, which takes both pseudo labels as input and outputs an edge-accurate camouflaged map. Experiments on various COD datasets demonstrate the effectiveness of our method with superior performance to the existing state-of-the-art methods.
△ Less
Submitted 29 October, 2021;
originally announced October 2021.
-
Scalable Personalised Item Ranking through Parametric Density Estimation
Authors:
Riku Togashi,
Masahiro Kato,
Mayu Otani,
Tetsuya Sakai,
Shin'ichi Satoh
Abstract:
Learning from implicit feedback is challenging because of the difficult nature of the one-class problem: we can observe only positive examples. Most conventional methods use a pairwise ranking approach and negative samplers to cope with the one-class problem. However, such methods have two main drawbacks particularly in large-scale applications; (1) the pairwise approach is severely inefficient du…
▽ More
Learning from implicit feedback is challenging because of the difficult nature of the one-class problem: we can observe only positive examples. Most conventional methods use a pairwise ranking approach and negative samplers to cope with the one-class problem. However, such methods have two main drawbacks particularly in large-scale applications; (1) the pairwise approach is severely inefficient due to the quadratic computational cost; and (2) even recent model-based samplers (e.g. IRGAN) cannot achieve practical efficiency due to the training of an extra model.
In this paper, we propose a learning-to-rank approach, which achieves convergence speed comparable to the pointwise counterpart while performing similarly to the pairwise counterpart in terms of ranking effectiveness. Our approach estimates the probability densities of positive items for each user within a rich class of distributions, viz. \emph{exponential family}. In our formulation, we derive a loss function and the appropriate negative sampling distribution based on maximum likelihood estimation. We also develop a practical technique for risk approximation and a regularisation scheme. We then discuss that our single-model approach is equivalent to an IRGAN variant under a certain condition. Through experiments on real-world datasets, our approach outperforms the pointwise and pairwise counterparts in terms of effectiveness and efficiency.
△ Less
Submitted 10 May, 2021;
originally announced May 2021.
-
The intersection polynomials of a virtual knot I: Definitions and calculations
Authors:
Ryuji Higa,
Takuji Nakamura,
Yasutaka Nakanishi,
Shin Satoh
Abstract:
We introduce three kinds of invariants of a virtual knot called the first, second, and third intersection polynomials. The definition is based on the intersection number of a pair of curves on a closed surface. The calculations of intersection polynomials are given up to crossing number four. We also study several properties of intersection polynomials.
We introduce three kinds of invariants of a virtual knot called the first, second, and third intersection polynomials. The definition is based on the intersection number of a pair of curves on a closed surface. The calculations of intersection polynomials are given up to crossing number four. We also study several properties of intersection polynomials.
△ Less
Submitted 25 January, 2022; v1 submitted 23 February, 2021;
originally announced February 2021.
-
Density-Ratio Based Personalised Ranking from Implicit Feedback
Authors:
Riku Togashi,
Masahiro Kato,
Mayu Otani,
Shin'ichi Satoh
Abstract:
Learning from implicit user feedback is challenging as we can only observe positive samples but never access negative ones. Most conventional methods cope with this issue by adopting a pairwise ranking approach with negative sampling. However, the pairwise ranking approach has a severe disadvantage in the convergence time owing to the quadratically increasing computational cost with respect to the…
▽ More
Learning from implicit user feedback is challenging as we can only observe positive samples but never access negative ones. Most conventional methods cope with this issue by adopting a pairwise ranking approach with negative sampling. However, the pairwise ranking approach has a severe disadvantage in the convergence time owing to the quadratically increasing computational cost with respect to the sample size; it is problematic, particularly for large-scale datasets and complex models such as neural networks. By contrast, a pointwise approach does not directly solve a ranking problem, and is therefore inferior to a pairwise counterpart in top-K ranking tasks; however, it is generally advantageous in regards to the convergence time. This study aims to establish an approach to learn personalised ranking from implicit feedback, which reconciles the training efficiency of the pointwise approach and ranking effectiveness of the pairwise counterpart. The key idea is to estimate the ranking of items in a pointwise manner; we first reformulate the conventional pointwise approach based on density ratio estimation and then incorporate the essence of ranking-oriented approaches (e.g. the pairwise approach) into our formulation. Through experiments on three real-world datasets, we demonstrate that our approach not only dramatically reduces the convergence time (one to two orders of magnitude faster) but also significantly improving the ranking performance.
△ Less
Submitted 19 January, 2021;
originally announced January 2021.
-
Image Inpainting Guided by Coherence Priors of Semantics and Textures
Authors:
Liang Liao,
Jing Xiao,
Zheng Wang,
Chia-Wen Lin,
Shin'ichi Satoh
Abstract:
Existing inpainting methods have achieved promising performance in recovering defected images of specific scenes. However, filling holes involving multiple semantic categories remains challenging due to the obscure semantic boundaries and the mixture of different semantic textures. In this paper, we introduce coherence priors between the semantics and textures which make it possible to concentrate…
▽ More
Existing inpainting methods have achieved promising performance in recovering defected images of specific scenes. However, filling holes involving multiple semantic categories remains challenging due to the obscure semantic boundaries and the mixture of different semantic textures. In this paper, we introduce coherence priors between the semantics and textures which make it possible to concentrate on completing separate textures in a semantic-wise manner. Specifically, we adopt a multi-scale joint optimization framework to first model the coherence priors and then accordingly interleavingly optimize image inpainting and semantic segmentation in a coarse-to-fine manner. A Semantic-Wise Attention Propagation (SWAP) module is devised to refine completed image textures across scales by exploring non-local semantic coherence, which effectively mitigates mix-up of textures. We also propose two coherence losses to constrain the consistency between the semantics and the inpainted image in terms of the overall structure and detailed textures. Experimental results demonstrate the superiority of our proposed method for challenging cases with complex holes.
△ Less
Submitted 14 December, 2020;
originally announced December 2020.
-
Alleviating Cold-Start Problems in Recommendation through Pseudo-Labelling over Knowledge Graph
Authors:
Riku Togashi,
Mayu Otani,
Shin'ichi Satoh
Abstract:
Solving cold-start problems is indispensable to provide meaningful recommendation results for new users and items. Under sparsely observed data, unobserved user-item pairs are also a vital source for distilling latent users' information needs. Most present works leverage unobserved samples for extracting negative signals. However, such an optimisation strategy can lead to biased results toward alr…
▽ More
Solving cold-start problems is indispensable to provide meaningful recommendation results for new users and items. Under sparsely observed data, unobserved user-item pairs are also a vital source for distilling latent users' information needs. Most present works leverage unobserved samples for extracting negative signals. However, such an optimisation strategy can lead to biased results toward already popular items by frequently handling new items as negative instances. In this study, we tackle the cold-start problems for new users/items by appropriately leveraging unobserved samples. We propose a knowledge graph (KG)-aware recommender based on graph neural networks, which augments labelled samples through pseudo-labelling. Our approach aggressively employs unobserved samples as positive instances and brings new items into the spotlight. To avoid exhaustive label assignments to all possible pairs of users and items, we exploit a KG for selecting probably positive items for each user. We also utilise an improved negative sampling strategy and thereby suppress the exacerbation of popularity biases. Through experiments, we demonstrate that our approach achieves improvements over the state-of-the-art KG-aware recommenders in a variety of scenarios; in particular, our methodology successfully improves recommendation performance for cold-start users/items.
△ Less
Submitted 10 November, 2020;
originally announced November 2020.
-
Towards Unsupervised Crowd Counting via Regression-Detection Bi-knowledge Transfer
Authors:
Yuting Liu,
Zheng Wang,
Miaojing Shi,
Shin'ichi Satoh,
Qijun Zhao,
Hongyu Yang
Abstract:
Unsupervised crowd counting is a challenging yet not largely explored task. In this paper, we explore it in a transfer learning setting where we learn to detect and count persons in an unlabeled target set by transferring bi-knowledge learnt from regression- and detection-based models in a labeled source set. The dual source knowledge of the two models is heterogeneous and complementary as they ca…
▽ More
Unsupervised crowd counting is a challenging yet not largely explored task. In this paper, we explore it in a transfer learning setting where we learn to detect and count persons in an unlabeled target set by transferring bi-knowledge learnt from regression- and detection-based models in a labeled source set. The dual source knowledge of the two models is heterogeneous and complementary as they capture different modalities of the crowd distribution. We formulate the mutual transformations between the outputs of regression- and detection-based models as two scene-agnostic transformers which enable knowledge distillation between the two models. Given the regression- and detection-based models and their mutual transformers learnt in the source, we introduce an iterative self-supervised learning scheme with regression-detection bi-knowledge transfer in the target. Extensive experiments on standard crowd counting benchmarks, ShanghaiTech, UCF\_CC\_50, and UCF\_QNRF demonstrate a substantial improvement of our method over other state-of-the-arts in the transfer learning setting.
△ Less
Submitted 27 September, 2020; v1 submitted 12 August, 2020;
originally announced August 2020.
-
MADGAN: unsupervised Medical Anomaly Detection GAN using multiple adjacent brain MRI slice reconstruction
Authors:
Changhee Han,
Leonardo Rundo,
Kohei Murao,
Tomoyuki Noguchi,
Yuki Shimahara,
Zoltan Adam Milacski,
Saori Koshino,
Evis Sala,
Hideki Nakayama,
Shinichi Satoh
Abstract:
Unsupervised learning can discover various unseen abnormalities, relying on large-scale unannotated medical images of healthy subjects. Towards this, unsupervised methods reconstruct a 2D/3D single medical image to detect outliers either in the learned feature space or from high reconstruction loss. However, without considering continuity between multiple adjacent slices, they cannot directly disc…
▽ More
Unsupervised learning can discover various unseen abnormalities, relying on large-scale unannotated medical images of healthy subjects. Towards this, unsupervised methods reconstruct a 2D/3D single medical image to detect outliers either in the learned feature space or from high reconstruction loss. However, without considering continuity between multiple adjacent slices, they cannot directly discriminate diseases composed of the accumulation of subtle anatomical anomalies, such as Alzheimer's Disease (AD). Moreover, no study has shown how unsupervised anomaly detection is associated with either disease stages, various (i.e., more than two types of) diseases, or multi-sequence Magnetic Resonance Imaging (MRI) scans. Therefore, we propose unsupervised Medical Anomaly Detection Generative Adversarial Network (MADGAN), a novel two-step method using GAN-based multiple adjacent brain MRI slice reconstruction to detect brain anomalies at different stages on multi-sequence structural MRI: (Reconstruction) Wasserstein loss with Gradient Penalty + 100 L1 loss-trained on 3 healthy brain axial MRI slices to reconstruct the next 3 ones-reconstructs unseen healthy/abnormal scans; (Diagnosis) Average L2 loss per scan discriminates them, comparing the ground truth/reconstructed slices. For training, we use two different datasets composed of 1,133 healthy T1-weighted (T1) and 135 healthy contrast-enhanced T1 (T1c) brain MRI scans for detecting AD and brain metastases/various diseases, respectively. Our Self-Attention MADGAN can detect AD on T1 scans at a very early stage, Mild Cognitive Impairment (MCI), with Area Under the Curve (AUC) 0.727, and AD at a late stage with AUC 0.894, while detecting brain metastases on T1c scans with AUC 0.921.
△ Less
Submitted 12 October, 2020; v1 submitted 24 July, 2020;
originally announced July 2020.