-
Auditing Instruction-Trajectory Mismatches in Multimodal Robot Demonstrations
Authors:
Simon Holk,
Ryosuke Takanami,
Tatsuya Matsushima,
Yusuke Iwasawa,
Yutaka Matsuo,
Yueh-Hua Wu,
Kei Ota
Abstract:
Robot demonstration datasets used to train vision-language-action policies can contain a subtle but harmful failure mode: trajectories that are behaviorally correct but paired with the wrong language instruction. We study post-hoc auditing of these Instruction-Trajectory Mismatches (ITMs). Unlike failed rollouts, ITMs often look plausible, and can corrupt the language-behavior mapping learned by t…
▽ More
Robot demonstration datasets used to train vision-language-action policies can contain a subtle but harmful failure mode: trajectories that are behaviorally correct but paired with the wrong language instruction. We study post-hoc auditing of these Instruction-Trajectory Mismatches (ITMs). Unlike failed rollouts, ITMs often look plausible, and can corrupt the language-behavior mapping learned by the policy. We propose Multimodal Probabilistic Fusion (MMPF), a training-free auditing framework that treats each modality as an expert, estimates a task-label distribution from local neighborhood agreement and global prototype similarity, and then fuses modalities with predictive-entropy weighting in a product of experts. Across LIBERO benchmarks with injected instruction mismatches and noisy real-robot data, MMPF achieves the strongest overall ITM detection and label correction accuracy. We also show that auditing improves most downstream policy learning in settings where language is needed to disambiguate the task. We demonstrate in real robot experiments that our method can achieve improved policy performance and show the trade-off of filtering demonstrations compared to relabeling.
△ Less
Submitted 7 August, 2026;
originally announced August 2026.
-
Chandra X-Ray Imaging and Spatially Resolved Spectroscopy of SN 1987A: Energy-Dependent Morphology of the Equatorial Ring
Authors:
Yusuke Sakai,
Shinya Yamada,
Koji Mori,
Hiromasa Suzuki,
Haruka Sakemi,
Tsukasa Matsushima,
Shintaro Kaneko,
Kai Matsunaga,
Shogo B. Kobayashi,
Haruto Aoki,
Toshiki Sato
Abstract:
We present a systematic imaging and spatially resolved spectral study of SN 1987A using Chandra observations obtained between 1999 and 2025. By combining multiepoch ACIS and HETG data, we investigate the long-term evolution of the remnant in both the soft and hard X-ray bands. To characterize the radial structure, we model the projected emission with a torus profile and derive its radius and width…
▽ More
We present a systematic imaging and spatially resolved spectral study of SN 1987A using Chandra observations obtained between 1999 and 2025. By combining multiepoch ACIS and HETG data, we investigate the long-term evolution of the remnant in both the soft and hard X-ray bands. To characterize the radial structure, we model the projected emission with a torus profile and derive its radius and width on the image plane. We find an energy dependence in the ring morphology: while the soft and hard bands exhibit similar structures at early epochs, the soft-band emission becomes systematically broader than the hard-band emission after the early 2010s. Furthermore, when considering the radius and width together, the soft-band emission shows an inward extension, suggesting an increasing contribution from interior and/or high-latitude emission components. The flux evolution of the Fe K line is consistent with previous XMM-Newton results, and we detect its presence already in earlier epochs (~2007-2009) using combined Chandra spectra. Spatially resolved analysis further indicates that the Fe K emission is enhanced in the eastern region. These results provide a unified view of the long-term morphological and spectral evolution of SN 1987A and highlight the emergence of energy-dependent radial structure as a key feature in its late-time evolution.
△ Less
Submitted 7 July, 2026;
originally announced July 2026.
-
FlexLAM: Resolving the Bottleneck Trade-off in Latent Action Learning
Authors:
Takanori Yoshimoto,
Yang Hu,
Naruya Kondo,
Tatsuya Matsushima
Abstract:
Latent actions provide a compact interface between action-free video and downstream decision-making, yet existing Latent Action Models (LAMs) force every transition through a fixed-capacity bottleneck. We identify a bottleneck trade-off: overly tight codes can discard transition cues needed for action alignment, while overly loose codes preserve additional transition variation that must be resolve…
▽ More
Latent actions provide a compact interface between action-free video and downstream decision-making, yet existing Latent Action Models (LAMs) force every transition through a fixed-capacity bottleneck. We identify a bottleneck trade-off: overly tight codes can discard transition cues needed for action alignment, while overly loose codes preserve additional transition variation that must be resolved when alignment labels are scarce or narrowly distributed. FlexLAM replaces this fixed capacity with variable-length latent actions trained by nested dropout, yielding prefix-valid codes that capture compact transition structure first and add detail only when needed, without new architectures or losses. A single FlexLAM matches or surpasses separately trained fixed-capacity LAMs at every evaluated token budget under standard scarce-label supervision and under a low-return single-task alignment stress test, indicating that FlexLAM is not merely adjustable at inference time but learns a better latent-action interface at the same token budgets. The same model supports inference-time token-budget adjustment without retraining, and FlexLAM improves Ego4D transition reconstruction. These results suggest that variable-length latent actions are an architecture-free, drop-in upgrade to the fixed-capacity bottleneck in latent action models, latent-action world models, and video-pretrained action interfaces.
△ Less
Submitted 17 June, 2026;
originally announced June 2026.
-
YUBI: Yielding Universal Bidigital Interface for Bimanual Dexterous Manipulation at Scale
Authors:
Takehiko Ohkawa,
Jumpei Arima,
Yuki Noguchi,
Masatoshi Tateno,
Makoto Sugiura,
Takuya Okubo,
Kengo Ikeuchi,
Yuma Shin,
Hiroki Nishizawa,
Naoaki Kanazawa,
Yuki Wakayama,
Daiki Fukunaga,
Koshi Makihara,
Tomohiro Motoda,
Floris Erich,
Yukiyasu Domae,
Tatsuya Matsushima,
Yohishiro Okumatsu,
Kei Ota
Abstract:
We introduce Yielding Universal Bidigital Interface (YUBI), a finger-aligned gripper designed to enable intuitive, ergonomic, and scalable data collection for bimanual dexterous manipulation. While handheld data collection systems such as Universal Manipulation Interface (UMI) enable affordable data collection, their bulky pistol-grip designs can pose ergonomic and usability challenges for fine-gr…
▽ More
We introduce Yielding Universal Bidigital Interface (YUBI), a finger-aligned gripper designed to enable intuitive, ergonomic, and scalable data collection for bimanual dexterous manipulation. While handheld data collection systems such as Universal Manipulation Interface (UMI) enable affordable data collection, their bulky pistol-grip designs can pose ergonomic and usability challenges for fine-grained, dexterous manipulation tasks. To address this, YUBI presents a distinct design principle: yielding, finger-driven actuation that directly maps human finger movements to gripper jaw motion. Using the YUBI devices, we set up a data collection system with integrated VR-based 6 DoF tracking of the gripper, ensuring high-fidelity trajectory data acquisition. We curate a UMI-based dataset of unprecedented scale: 8,434 hours across 1.20M episodes and 119 tasks. Experiments show that YUBI offers advantages over the UMI gripper in versatility for complex bimanual tasks, dexterity, and operational efficiency. A single policy trained on the YUBI dataset transfers across multiple bimanual robots (UR, Franka, and ELEY) simply by mounting the gripper on each platform, confirming that the collected data are directly executable as policy supervision. We release the gripper hardware, data-collection software, and dataset as one integrated stack, offering the open community a reproducible path to large-scale data acquisition for advancing robotic foundation models.
△ Less
Submitted 8 June, 2026;
originally announced June 2026.
-
See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs
Authors:
Yueh-Hua Wu,
Tatsuya Matsushima,
Kei Ota
Abstract:
Generalization remains a central bottleneck for vision-language-action (VLA) models: under distractors, appearance shifts, and semantically similar tasks, the policy must often infer local execution details from coarse instructions while also deciding which parts of the image matter for control. We present S2 (See Less, Specify More), a framework for improving VLA generalization by training the ex…
▽ More
Generalization remains a central bottleneck for vision-language-action (VLA) models: under distractors, appearance shifts, and semantically similar tasks, the policy must often infer local execution details from coarse instructions while also deciding which parts of the image matter for control. We present S2 (See Less, Specify More), a framework for improving VLA generalization by training the executor under a cleaner interface.
Specify More preserves the original instruction as a stable high-level goal while relabeling each trajectory into refined trajectory- and subtask-level language that disambiguates the current execution mode. Unlike native attention, See Less imposes an explicit visual evidence budget, training the executor to act from task-sufficient evidence rather than unconstrained visual context, without any region or mask annotation.
This interface lets the executor follow detailed guidance without relying on distracting visual patches or resolving avoidable ambiguity on its own, and it remains compatible with off-the-shelf VLM planners through in-context learning. Across our main evaluation settings, S2 improves overall generalization metrics by changing the executor's learning problem: coarse instructions induce avoidable supervision aliasing, goal-preserving local guidance outperforms instruction replacement in our main ablations, and explicit evidence budgeting reduces dependence on broad visual context beyond efficiency considerations.
Across eight real-robot tasks on TX-G2 (an AgiBot G2-compatible variant) and HSR, S2 raises mean subtask success from 54.2% to 79.0% over pi0.5. Together, these results suggest that VLA generalization improves when the executor is trained to act from informative local guidance and task-sufficient visual evidence, rather than recovering both from weak supervision.
△ Less
Submitted 8 June, 2026; v1 submitted 1 June, 2026;
originally announced June 2026.
-
Continuous Reasoning for Vision-Language-Action
Authors:
Yueh-Hua Wu,
Tatsuya Matsushima,
Kei Ota
Abstract:
Natural language is a powerful reasoning medium for language and vision-language models, but it is mismatched to the granularity of continuous control. Text and explicit subgoals operate at task-level granularity, whereas vision-language-action (VLA) policies must choose actions at a much finer temporal scale; a single reasoning step can therefore span many action chunks while remaining only weakl…
▽ More
Natural language is a powerful reasoning medium for language and vision-language models, but it is mismatched to the granularity of continuous control. Text and explicit subgoals operate at task-level granularity, whereas vision-language-action (VLA) policies must choose actions at a much finer temporal scale; a single reasoning step can therefore span many action chunks while remaining only weakly coupled to the action needed now. This suggests a different question for VLA: what should play the role of language? We argue that a useful VLA reasoning medium must be shareable across model instances, verifiable through downstream action improvement, and aligned with temporally extended control structure.
Based on this view, we propose Continuous Reasoning for Vision-Language-Action. Our model first predicts continuous reasoning in the form of a structured set of continuous thoughts, then reuses them as shared context for chunk-structured action generation. Better action prediction alone does not certify good reasoning: if the same internal medium cannot be shared across model instances and independently verified through improved downstream control, the added latent may simply become a model-private shortcut that helps on seen behaviors without supporting generalizable control. We therefore instantiate continuous reasoning as a shared Gaussian latent interface and train it with a self-verification objective in which an exponential-moving-average teacher must successfully consume the student's reasoning when predicting target actions.
Empirically, Continuous Reasoning improves LIBERO-PRO robustness and performs strongly on real robots, raising mean subtask success over π0.5 by 40.4% on TX-G2, an AgiBot G2-compatible variant, and 26.3% on HSR. This suggests that reasoning in VLA is less about extra tokens than about a shareable, verifiable internal language for action.
△ Less
Submitted 8 June, 2026; v1 submitted 29 May, 2026;
originally announced June 2026.
-
A Mixture Autoregressive Image Generative Model on Quadtree Regions for Gaussian Noise Removal via Variational Bayes and Gradient Methods
Authors:
Shota Saito,
Yuta Nakahara,
Kohei Horinouchi,
Naoki Ichijo,
Manabu Kobayashi,
Toshiyasu Matsushima
Abstract:
This paper addresses the problem of image denoising for grayscale images. We propose a probabilistic image generative model that combines a quadtree region-partitioning model with a mixture autoregressive model, and propose a framework that reduces MAP (maximum a posteriori)-estimation-based denoising to the maximization of a variational lower bound. To maximize this lower bound, we develop an alg…
▽ More
This paper addresses the problem of image denoising for grayscale images. We propose a probabilistic image generative model that combines a quadtree region-partitioning model with a mixture autoregressive model, and propose a framework that reduces MAP (maximum a posteriori)-estimation-based denoising to the maximization of a variational lower bound. To maximize this lower bound, we develop an algorithm that alternately applies variational Bayes and gradient methods. We particularly demonstrate that the gradient-based update rule can be computed analytically without numerical computation or approximation. We carried out some experiments to verify that the proposed algorithm actually removes image noise and to identify directions for future improvement.
△ Less
Submitted 12 May, 2026;
originally announced May 2026.
-
Surface-localized topological superconductivity in nodal-loop materials: BdG analysis
Authors:
Takeru Matsushima,
Hiroki Tsuchiura
Abstract:
We theoretically study surface superconductivity in a nodal-line semimetal by combining a minimal tight-binding model with a layer-resolved Bogoliubov-de Gennes approach. In the normal state, the model realizes a bulk nodal loop and an associated drumhead surface band in a slab geometry with open boundaries in the $z$ direction: the central layers reproduce the bulk-like density of states, whereas…
▽ More
We theoretically study surface superconductivity in a nodal-line semimetal by combining a minimal tight-binding model with a layer-resolved Bogoliubov-de Gennes approach. In the normal state, the model realizes a bulk nodal loop and an associated drumhead surface band in a slab geometry with open boundaries in the $z$ direction: the central layers reproduce the bulk-like density of states, whereas the surface layer exhibits a sharp zero-energy peak originating from the drumhead states. On top of this band structure we introduce chiral $p$-wave and $d_{x^2-y^2}$-wave superconducting channels and determine the layer-dependent gap amplitudes self-consistently. The chiral $p$-wave order parameter is strongly enhanced at the outermost layers and decays within only a few layers towards the interior, while the $d$-wave order parameter is more than an order of magnitude smaller on all layers. The quasiparticle dispersion and surface local density of states in the chiral $p$-wave state show that the drumhead band is efficiently gapped out and that the zero-energy peak in the normal surface spectrum is split into two coherence peaks, directly reflecting the induced superconducting gap. These results demonstrate that superconductivity driven by drumhead surface states is naturally biased toward a surface-localized chiral $p$-wave pairing symmetry and may offer qualitative guidance for interpreting surface-sensitive experiments on Pd-doped CaAgP.
△ Less
Submitted 20 March, 2026; v1 submitted 26 February, 2026;
originally announced February 2026.
-
Variable Splitting Binary Tree Models Based on Bayesian Context Tree Models for Time Series Segmentation
Authors:
Yuta Nakahara,
Shota Saito,
Kohei Horinouchi,
Koshi Shimada,
Naoki Ichijo,
Manabu Kobayashi,
Toshiyasu Matsushima
Abstract:
We propose a variable splitting binary tree (VSBT) model based on Bayesian context tree (BCT) models for time series segmentation. Unlike previous applications of BCT models, the tree structure in our model represents interval partitioning on the time domain. Moreover, interval partitioning is represented by recursive logistic regression models. By adjusting logistic regression coefficients, our m…
▽ More
We propose a variable splitting binary tree (VSBT) model based on Bayesian context tree (BCT) models for time series segmentation. Unlike previous applications of BCT models, the tree structure in our model represents interval partitioning on the time domain. Moreover, interval partitioning is represented by recursive logistic regression models. By adjusting logistic regression coefficients, our model can represent split positions at arbitrary locations within each interval. This enables more compact tree representations. For simultaneous estimation of both split positions and tree depth, we develop an effective inference algorithm that combines local variational approximation for logistic regression with the context tree weighting (CTW) algorithm. We present numerical examples on synthetic data demonstrating the effectiveness of our model and algorithm.
△ Less
Submitted 22 January, 2026;
originally announced January 2026.
-
Soft Bayesian Context Tree Models for Real-Valued Time Series
Authors:
Shota Saito,
Yuta Nakahara,
Toshiyasu Matsushima
Abstract:
This paper proposes the soft Bayesian context tree model (Soft-BCT), which is a novel BCT model for real-valued time series. The Soft-BCT considers soft (probabilistic) splits of the context space, instead of hard (deterministic) splits of the context space as in the previous BCT for real-valued time series. A learning algorithm of the Soft-BCT is proposed based on the variational inference. The r…
▽ More
This paper proposes the soft Bayesian context tree model (Soft-BCT), which is a novel BCT model for real-valued time series. The Soft-BCT considers soft (probabilistic) splits of the context space, instead of hard (deterministic) splits of the context space as in the previous BCT for real-valued time series. A learning algorithm of the Soft-BCT is proposed based on the variational inference. The results of experiments demonstrate the superiority of the Soft-BCT compared to the previous BCT for some datasets.
△ Less
Submitted 21 May, 2026; v1 submitted 16 January, 2026;
originally announced January 2026.
-
Anomalous Enhancement of Yield Strength due to Static Friction
Authors:
Ryudo Suzuki,
Takashi Matsushima,
Tetsuo Yamaguchi,
Marie Tani,
Shin-ichi Sasa
Abstract:
Friction is fundamental to mechanical stability across scales, from geological faults and architectural structures to granular materials and animal feet. We study the mechanical stability of a minimal friction-stabilized structure composed of three cylindrical particles arranged in a triangular stack on a floor under gravity. We analyze the yield force, defined as the threshold compressive force a…
▽ More
Friction is fundamental to mechanical stability across scales, from geological faults and architectural structures to granular materials and animal feet. We study the mechanical stability of a minimal friction-stabilized structure composed of three cylindrical particles arranged in a triangular stack on a floor under gravity. We analyze the yield force, defined as the threshold compressive force applied quasi-statically from above at which the structure collapses due to sliding at the floor contact. Using singular perturbation analysis, we derive an expression which quantitatively predicts the yield force as a function of the static friction coefficient and a small dimensionless parameter $ε$ characterizing elastic deformation.
△ Less
Submitted 29 May, 2026; v1 submitted 10 November, 2025;
originally announced November 2025.
-
Necessary and Sufficient Conditions for Capacity-Achieving Private Information Retrieval with Adversarial Servers
Authors:
Atsushi Miki,
Toshiyasu Matsushima
Abstract:
Private information retrieval (PIR) is a mechanism for efficiently downloading messages while keeping the index of the desired message secret from the servers. PIR schemes have been extended to various scenarios with adversarial servers: PIR schemes where some servers are unresponsive or return noisy responses are called robust PIR and Byzantine PIR, respectively; PIR schemes where some servers co…
▽ More
Private information retrieval (PIR) is a mechanism for efficiently downloading messages while keeping the index of the desired message secret from the servers. PIR schemes have been extended to various scenarios with adversarial servers: PIR schemes where some servers are unresponsive or return noisy responses are called robust PIR and Byzantine PIR, respectively; PIR schemes where some servers collude to reveal the index are called colluding PIR. The information-theoretic upper bound on the download efficiency of these PIR schemes has been proved in previous studies. However, systematic ways to construct PIR schemes that achieve the upper bound are not known. In order to construct a capacity-achieving PIR schemes systematically, it is necessary to clarify the conditions that the queries should satisfy. This paper proves the necessary and sufficient conditions for capacity-achieving PIR schemes.
△ Less
Submitted 21 January, 2026; v1 submitted 8 November, 2025;
originally announced November 2025.
-
AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation
Authors:
Ryosuke Takanami,
Petr Khrapchenkov,
Shu Morikuni,
Jumpei Arima,
Yuta Takaba,
Shunsuke Maeda,
Takuya Okubo,
Genki Sano,
Satoshi Sekioka,
Aoi Kadoya,
Motonari Kambara,
Naoya Nishiura,
Haruto Suzuki,
Takanori Yoshimoto,
Koya Sakamoto,
Shinnosuke Ono,
Hu Yang,
Daichi Yashima,
Aoi Horo,
Tomohiro Motoda,
Kensuke Chiyoma,
Hiroshi Ito,
Koki Fukuda,
Akihito Goto,
Kazumi Morinaga
, et al. (10 additional authors not shown)
Abstract:
As robots transition from controlled settings to unstructured human environments, building generalist agents that can reliably follow natural language instructions remains a central challenge. Progress in robust mobile manipulation requires large-scale multimodal datasets that capture contact-rich and long-horizon tasks, yet existing resources lack synchronized force-torque sensing, hierarchical a…
▽ More
As robots transition from controlled settings to unstructured human environments, building generalist agents that can reliably follow natural language instructions remains a central challenge. Progress in robust mobile manipulation requires large-scale multimodal datasets that capture contact-rich and long-horizon tasks, yet existing resources lack synchronized force-torque sensing, hierarchical annotations, and explicit failure cases. We address this gap with the AIRoA MoMa Dataset, a large-scale real-world multimodal dataset for mobile manipulation. It includes synchronized RGB images, joint states, six-axis wrist force-torque signals, and internal robot states, together with a novel two-layer annotation schema of sub-goals and primitive actions for hierarchical learning and error analysis. The initial dataset comprises 25,469 episodes (approx. 94 hours) collected with the Human Support Robot (HSR) and is fully standardized in the LeRobot v2.1 format. By uniquely integrating mobile manipulation, contact-rich interaction, and long-horizon structure, AIRoA MoMa provides a critical benchmark for advancing the next generation of Vision-Language-Action models. The first version of our dataset is now available at https://huggingface.co/datasets/airoa-org/airoa-moma .
△ Less
Submitted 29 September, 2025;
originally announced September 2025.
-
Leave No Observation Behind: Real-time Correction for VLA Action Chunks
Authors:
Kohei Sendai,
Maxime Alvarez,
Tatsuya Matsushima,
Yutaka Matsuo,
Yusuke Iwasawa
Abstract:
To improve efficiency and temporal coherence, Vision-Language-Action (VLA) models often predict action chunks; however, this action chunking harms reactivity under inference delay and long horizons. We introduce Asynchronous Action Chunk Correction (A2C2), which is a lightweight real-time chunk correction head that runs every control step and adds a time-aware correction to any off-the-shelf VLA's…
▽ More
To improve efficiency and temporal coherence, Vision-Language-Action (VLA) models often predict action chunks; however, this action chunking harms reactivity under inference delay and long horizons. We introduce Asynchronous Action Chunk Correction (A2C2), which is a lightweight real-time chunk correction head that runs every control step and adds a time-aware correction to any off-the-shelf VLA's action chunk. The module combines the latest observation, the predicted action from VLA (base action), a positional feature that encodes the index of the base action within the chunk, and some features from the base policy, then outputs a per-step correction. This preserves the base model's competence while restoring closed-loop responsiveness. The approach requires no retraining of the base policy and is orthogonal to asynchronous execution schemes such as Real Time Chunking (RTC). On the dynamic Kinetix task suite (12 tasks) and LIBERO Spatial, our method yields consistent success rate improvements across increasing delays and execution horizons (+23% point and +7% point respectively, compared to RTC), and also improves robustness for long horizons even with zero injected delay. Since the correction head is small and fast, there is minimal overhead compared to the inference of large VLA models. These results indicate that A2C2 is an effective, plug-in mechanism for deploying high-capacity chunking policies in real-time control.
△ Less
Submitted 27 September, 2025;
originally announced September 2025.
-
Retrieve-Augmented Generation for Speeding up Diffusion Policy without Additional Training
Authors:
Sodtavilan Odonchimed,
Tatsuya Matsushima,
Simon Holk,
Yusuke Iwasawa,
Yutaka Matsuo
Abstract:
Diffusion Policies (DPs) have attracted attention for their ability to achieve significant accuracy improvements in various imitation learning tasks. However, DPs depend on Diffusion Models, which require multiple noise removal steps to generate a single action, resulting in long generation times. To solve this problem, knowledge distillation-based methods such as Consistency Policy (CP) have been…
▽ More
Diffusion Policies (DPs) have attracted attention for their ability to achieve significant accuracy improvements in various imitation learning tasks. However, DPs depend on Diffusion Models, which require multiple noise removal steps to generate a single action, resulting in long generation times. To solve this problem, knowledge distillation-based methods such as Consistency Policy (CP) have been proposed. However, these methods require a significant amount of training time, especially for difficult tasks. In this study, we propose RAGDP (Retrieve-Augmented Generation for Diffusion Policies) as a novel framework that eliminates the need for additional training using a knowledge base to expedite the inference of pre-trained DPs. In concrete, RAGDP encodes observation-action pairs through the DP encoder to construct a vector database of expert demonstrations. During inference, the current observation is embedded, and the most similar expert action is extracted. This extracted action is combined with an intermediate noise removal step to reduce the number of steps required compared to the original diffusion step. We show that by using RAGDP with the base model and existing acceleration methods, we improve the accuracy and speed trade-off with no additional training. Even when accelerating the models 20 times, RAGDP maintains an advantage in accuracy, with a 7% increase over distillation models such as CP.
△ Less
Submitted 28 July, 2025;
originally announced July 2025.
-
A Lower Bound for the Number of Linear Regions of Ternary ReLU Regression Neural Networks
Authors:
Yuta Nakahara,
Manabu Kobayashi,
Toshiyasu Matsushima
Abstract:
With the advancement of deep learning, reducing computational complexity and memory consumption has become a critical challenge, and ternary neural networks (NNs) that restrict parameters to $\{-1, 0, +1\}$ have attracted attention as a promising approach. While ternary NNs demonstrate excellent performance in practical applications such as image recognition and natural language processing, their…
▽ More
With the advancement of deep learning, reducing computational complexity and memory consumption has become a critical challenge, and ternary neural networks (NNs) that restrict parameters to $\{-1, 0, +1\}$ have attracted attention as a promising approach. While ternary NNs demonstrate excellent performance in practical applications such as image recognition and natural language processing, their theoretical understanding remains insufficient. In this paper, we theoretically analyze the expressivity of ternary NNs from the perspective of the number of linear regions. Specifically, we evaluate the number of linear regions of ternary regression NNs with Rectified Linear Unit (ReLU) for activation functions and prove that the number of linear regions increases polynomially with respect to network width and exponentially with respect to depth, similar to standard NNs. Moreover, we show that it suffices to first double the width, then either square the width or double the depth of ternary NNs with alternating ReLU and identity layers to achieve a lower bound on the maximum number of linear regions comparable to that of general ReLU regression NNs. When using ReLU in all the layers, a similar bound is obtained by further doubling the width. This provides a theoretical explanation, in some sense, for the practical success of ternary NNs.
△ Less
Submitted 25 April, 2026; v1 submitted 21 July, 2025;
originally announced July 2025.
-
SPARK: Graph-Based Online Semantic Integration System for Robot Task Planning
Authors:
Mimo Shirasaka,
Yuya Ikeda,
Tatsuya Matsushima,
Yutaka Matsuo,
Yusuke Iwasawa
Abstract:
The ability to update information acquired through various means online during task execution is crucial for a general-purpose service robot. This information includes geometric and semantic data. While SLAM handles geometric updates on 2D maps or 3D point clouds, online updates of semantic information remain unexplored. We attribute the challenge to the online scene graph representation, for its…
▽ More
The ability to update information acquired through various means online during task execution is crucial for a general-purpose service robot. This information includes geometric and semantic data. While SLAM handles geometric updates on 2D maps or 3D point clouds, online updates of semantic information remain unexplored. We attribute the challenge to the online scene graph representation, for its utility and scalability. Building on previous works regarding offline scene graph representations, we study online graph representations of semantic information in this work. We introduce SPARK: Spatial Perception and Robot Knowledge Integration. This framework extracts semantic information from environment-embedded cues and updates the scene graph accordingly, which is then used for subsequent task planning. We demonstrate that graph representations of spatial relationships enhance the robot system's ability to perform tasks in dynamic environments and adapt to unconventional spatial cues, like gestures.
△ Less
Submitted 25 June, 2025;
originally announced June 2025.
-
Implementing van der Waals forces for polytope particles in DEM simulations of clay
Authors:
Dominik Krengel,
Jian Chen,
Zhipeng Yu,
Hans-Georg Matuttis,
Takashi Matsushima
Abstract:
Clay minerals are non-spherical nano-scale particles that usually form flocculated, house-of-card like structures under the influence of inter-molecular forces. Numerical modeling of clays is still in its infancy as the required inter-particle forces are available only for spherical particles. A polytope approach would allow shape-accurate forces and torques while simultaneously being more perform…
▽ More
Clay minerals are non-spherical nano-scale particles that usually form flocculated, house-of-card like structures under the influence of inter-molecular forces. Numerical modeling of clays is still in its infancy as the required inter-particle forces are available only for spherical particles. A polytope approach would allow shape-accurate forces and torques while simultaneously being more performant. The Anandarajah solution provides an analytical formulation for van der Waals forces for cuboid particles but in its original form is not suitable for implementation in DEM simulations. In this work, we discuss the necessary changes for a functional implementation of the Anandarajah solution in a DEM simulation of rectangular particles and their extension to cuboid particles.
△ Less
Submitted 15 June, 2025;
originally announced June 2025.
-
A Comprehensive Survey on Physical Risk Control in the Era of Foundation Model-enabled Robotics
Authors:
Takeshi Kojima,
Yaonan Zhu,
Yusuke Iwasawa,
Toshinori Kitamura,
Gang Yan,
Shu Morikuni,
Ryosuke Takanami,
Alfredo Solano,
Tatsuya Matsushima,
Akiko Murakami,
Yutaka Matsuo
Abstract:
Recent Foundation Model-enabled robotics (FMRs) display greatly improved general-purpose skills, enabling more adaptable automation than conventional robotics. Their ability to handle diverse tasks thus creates new opportunities to replace human labor. However, unlike general foundation models, FMRs interact with the physical world, where their actions directly affect the safety of humans and surr…
▽ More
Recent Foundation Model-enabled robotics (FMRs) display greatly improved general-purpose skills, enabling more adaptable automation than conventional robotics. Their ability to handle diverse tasks thus creates new opportunities to replace human labor. However, unlike general foundation models, FMRs interact with the physical world, where their actions directly affect the safety of humans and surrounding objects, requiring careful deployment and control. Based on this proposition, our survey comprehensively summarizes robot control approaches to mitigate physical risks by covering all the lifespan of FMRs ranging from pre-deployment to post-accident stage. Specifically, we broadly divide the timeline into the following three phases: (1) pre-deployment phase, (2) pre-incident phase, and (3) post-incident phase. Throughout this survey, we find that there is much room to study (i) pre-incident risk mitigation strategies, (ii) research that assumes physical interaction with humans, and (iii) essential issues of foundation models themselves. We hope that this survey will be a milestone in providing a high-resolution analysis of the physical risks of FMRs and their control, contributing to the realization of a good human-robot relationship.
△ Less
Submitted 30 May, 2025; v1 submitted 18 May, 2025;
originally announced May 2025.
-
Self-organization, detailed balance, and stress-structure correlations in 2D granular dynamics
Authors:
Raphael Blumenfeld,
Takashi Matsushima,
Jie Zhang
Abstract:
We argue that a number of recent experimental and numerical observations point to an ongoing cooperative stress-structure self-organisation (SO) in quasi-static granular dynamics. These observations include: a) detail-insensitive collapses of certain quantities; b) correlations between stress and structure and evidence of entropy-stability competition in settled packings, which cast doubt on most…
▽ More
We argue that a number of recent experimental and numerical observations point to an ongoing cooperative stress-structure self-organisation (SO) in quasi-static granular dynamics. These observations include: a) detail-insensitive collapses of certain quantities; b) correlations between stress and structure and evidence of entropy-stability competition in settled packings, which cast doubt on most linear stress theories of granular materials; c) detailed balanced steady states, which seem contradictory to the common belief that only systems in thermal equilibrium satisfy detailed balance, but are not, as we explain. We then propose a new statistical mechanical formulation that takes into account the cooperative SO.
△ Less
Submitted 8 April, 2025;
originally announced April 2025.
-
In-orbit Performance of the Soft X-ray Imaging Telescope Xtend aboard XRISM
Authors:
Hiroyuki Uchida,
Koji Mori,
Hiroshi Tomida,
Hiroshi Nakajima,
Hirofumi Noda,
Takaaki Tanaka,
Hiroshi Murakami,
Hiromasa Suzuki,
Shogo Benjamin Kobayashi,
Tomokage Yoneyama,
Kouichi Hagino,
Kumiko Kawabata Nobukawa,
Hideki Uchiyama,
Masayoshi Nobukawa,
Hironori Matsumoto,
Takeshi Go Tsuru,
Makoto Yamauchi,
Isamu Hatsukade,
Hirokazu Odaka,
Takayoshi Kohmura,
Kazutaka Yamaoka,
Tessei Yoshida,
Yoshiaki Kanemaru,
Daiki Ishi,
Tadayasu Dotani
, et al. (40 additional authors not shown)
Abstract:
We present a summary of the in-orbit performance of the soft X-ray imaging telescope Xtend onboard the XRISM mission, based on in-flight observation data, including first-light celestial objects, calibration sources, and results from the cross-calibration campaign with other currently-operating X-ray observatories. XRISM/Xtend has a large field of view of $38.5'\times38.5'$, covering an energy ran…
▽ More
We present a summary of the in-orbit performance of the soft X-ray imaging telescope Xtend onboard the XRISM mission, based on in-flight observation data, including first-light celestial objects, calibration sources, and results from the cross-calibration campaign with other currently-operating X-ray observatories. XRISM/Xtend has a large field of view of $38.5'\times38.5'$, covering an energy range of 0.4--13 keV, as demonstrated by the first-light observation of the galaxy cluster Abell 2319. It also features an energy resolution of 170--180 eV at 6 keV, which meets the mission requirement and enables to resolve He-like and H-like Fe K$α$ lines. Throughout the observation during the performance verification phase, we confirm that two issues identified in SXI onboard the previous Hitomi mission -- light leakage and crosstalk events -- are addressed and suppressed in the case of Xtend. A joint cross-calibration observation of the bright quasar 3C273 results in an effective area measured to be $\sim420$ cm$^{2}$@1.5 keV and $\sim310$ cm$^{2}$@6.0 keV, which matches values obtained in ground tests. We also continuously monitor the health of Xtend by analyzing overclocking data, calibration source spectra, and day-Earth observations: the readout noise is stable and low, and contamination is negligible even one year after launch. A low background level compared to other major X-ray instruments onboard satellites, combined with the largest grasp ($Ω_{\rm eff}\sim60$ ${\rm cm^2~degree^2}$) of Xtend, will not only support Resolve analysis, but also enable significant scientific results on its own. This includes near future follow-up observations and transient searches in the context of time-domain and multi-messenger astrophysics.
△ Less
Submitted 25 March, 2025;
originally announced March 2025.
-
Soft X-ray Imager of the Xtend system onboard XRISM
Authors:
Hirofumi Noda,
Koji Mori,
Hiroshi Tomida,
Hiroshi Nakajima,
Takaaki Tanaka,
Hiroshi Murakami,
Hiroyuki Uchida,
Hiromasa Suzuki,
Shogo Benjamin Kobayashi,
Tomokage Yoneyama,
Kouichi Hagino,
Kumiko Nobukawa,
Hideki Uchiyama,
Masayoshi Nobukawa,
Hironori Matsumoto,
Takeshi Go Tsuru,
Makoto Yamauchi,
Isamu Hatsukade,
Hirokazu Odaka,
Takayoshi Kohmura,
Kazutaka Yamaoka,
Tessei Yoshida,
Yoshiaki Kanemaru,
Junko Hiraga,
Tadayasu Dotani
, et al. (35 additional authors not shown)
Abstract:
The Soft X-ray Imager (SXI) is the X-ray charge-coupled device (CCD) camera for the soft X-ray imaging telescope Xtend installed on the X-ray Imaging and Spectroscopy Mission (XRISM), which was adopted as a recovery mission for the Hitomi X-ray satellite and was successfully launched on 2023 September 7 (JST). In order to maximize the science output of XRISM, we set the requirements for Xtend and…
▽ More
The Soft X-ray Imager (SXI) is the X-ray charge-coupled device (CCD) camera for the soft X-ray imaging telescope Xtend installed on the X-ray Imaging and Spectroscopy Mission (XRISM), which was adopted as a recovery mission for the Hitomi X-ray satellite and was successfully launched on 2023 September 7 (JST). In order to maximize the science output of XRISM, we set the requirements for Xtend and find that the CCD set employed in the Hitomi/SXI or similar, i.e., a $2 \times 2$ array of back-illuminated CCDs with a $200~μ$m-thick depletion layer, would be practically best among available choices, when used in combination with the X-ray mirror assembly. We design the XRISM/SXI, based on the Hitomi/SXI, to have a wide field of view of $38' \times 38'$ in the $0.4-13$ keV energy range. We incorporated several significant improvements from the Hitomi/SXI into the CCD chip design to enhance the optical-light blocking capability and to increase the cosmic-ray tolerance, reducing the degradation of charge-transfer efficiency in orbit. By the time of the launch of XRISM, the imaging and spectroscopic capabilities of the SXI has been extensively studied in on-ground experiments with the full flight-model configuration or equivalent setups and confirmed to meet the requirements. The optical blocking capability, the cooling and temperature control performance, and the transmissivity and quantum efficiency to incident X-rays of the CCDs are also all confirmed to meet the requirements. Thus, we successfully complete the pre-flight development of the SXI for XRISM.
△ Less
Submitted 11 February, 2025;
originally announced February 2025.
-
Effects of particle angularity on granular self-organization
Authors:
Dominik Krengel,
Haoran Jiang,
Takashi Matsushima,
Raphael Blumenfeld
Abstract:
Recent studies of two-dimensional poly-disperse disc systems revealed a coordinated self-organisation of cell stresses and shapes, with certain distributions collapsing onto a master form for many processes, size distributions, friction coefficients, and cell orders. Here we examine the effects of grain angularity on the indicators of self-organisation, using simulations of bi-disperse regular…
▽ More
Recent studies of two-dimensional poly-disperse disc systems revealed a coordinated self-organisation of cell stresses and shapes, with certain distributions collapsing onto a master form for many processes, size distributions, friction coefficients, and cell orders. Here we examine the effects of grain angularity on the indicators of self-organisation, using simulations of bi-disperse regular $N$-polygons and varying $N$ systematically. We find that: the strong correlation between local cell stresses and orientations, as well as the collapses of the conditional distributions of scaled cell stress ratios to a master Weibull form for all cell orders $k$, are independent of angularity and friction coefficient. In contrast, increasing angularity makes the collapses of the conditional distributions sensitive to changes in the friction coefficient.
△ Less
Submitted 24 August, 2025; v1 submitted 9 February, 2025;
originally announced February 2025.
-
Enhancing and Exploring Mild Cognitive Impairment Detection with W2V-BERT-2.0
Authors:
Yueguan Wang,
Tatsunari Matsushima,
Soichiro Matsushima,
Toshimitsu Sakai
Abstract:
This study explores a multi-lingual audio self-supervised learning model for detecting mild cognitive impairment (MCI) using the TAUKADIAL cross-lingual dataset. While speech transcription-based detection with BERT models is effective, limitations exist due to a lack of transcriptions and temporal information. To address these issues, the study utilizes features directly from speech utterances wit…
▽ More
This study explores a multi-lingual audio self-supervised learning model for detecting mild cognitive impairment (MCI) using the TAUKADIAL cross-lingual dataset. While speech transcription-based detection with BERT models is effective, limitations exist due to a lack of transcriptions and temporal information. To address these issues, the study utilizes features directly from speech utterances with W2V-BERT-2.0. We propose a visualization method to detect essential layers of the model for MCI classification and design a specific inference logic considering the characteristics of MCI. The experiment shows competitive results, and the proposed inference logic significantly contributes to the improvements from the baseline. We also conduct detailed analysis which reveals the challenges related to speaker bias in the features and the sensitivity of MCI classification accuracy to the data split, providing valuable insights for future research.
△ Less
Submitted 27 January, 2025;
originally announced January 2025.
-
Status of Xtend telescope onboard X-Ray Imaging and Spectroscopy Mission (XRISM)
Authors:
Koji Mori,
Hiroshi Tomida,
Hiroshi Nakajima,
Takashi Okajima,
Hirofumi Noda,
Hiroyuki Uchida,
Hiromasa Suzuki,
Shogo Benjamin Kobayashi,
Tomokage Yoneyama,
Kouichi Hagino,
Kumiko Nobukawa,
Takaaki Tanaka,
Hiroshi Murakami,
Hideki Uchiyama,
Masayoshi Nobukawa,
Hironori Matsumoto,
Takeshi Tsuru,
Makoto Yamauchi,
Isamu Hatsukade,
Hirokazu Odaka,
Takayoshi Kohmura,
Kazutaka Yamaoka,
Manabu Ishida,
Yoshitomo Maeda,
Takayuki Hayashi
, et al. (38 additional authors not shown)
Abstract:
Xtend is one of the two telescopes onboard the X-ray imaging and spectroscopy mission (XRISM), which was launched on September 7th, 2023. Xtend comprises the Soft X-ray Imager (SXI), an X-ray CCD camera, and the X-ray Mirror Assembly (XMA), a thin-foil-nested conically approximated Wolter-I optics. A large field of view of $38^{\prime}\times38^{\prime}$ over the energy range from 0.4 to 13 keV is…
▽ More
Xtend is one of the two telescopes onboard the X-ray imaging and spectroscopy mission (XRISM), which was launched on September 7th, 2023. Xtend comprises the Soft X-ray Imager (SXI), an X-ray CCD camera, and the X-ray Mirror Assembly (XMA), a thin-foil-nested conically approximated Wolter-I optics. A large field of view of $38^{\prime}\times38^{\prime}$ over the energy range from 0.4 to 13 keV is realized by the combination of the SXI and XMA with a focal length of 5.6 m. The SXI employs four P-channel, back-illuminated type CCDs with a thick depletion layer of 200 $μ$m. The four CCD chips are arranged in a 2$\times$2 grid and cooled down to $-110$ $^{\circ}$C with a single-stage Stirling cooler. Before the launch of XRISM, we conducted a month-long spacecraft thermal vacuum test. The performance verification of the SXI was successfully carried out in a course of multiple thermal cycles of the spacecraft. About a month after the launch of XRISM, the SXI was carefully activated and the soundness of its functionality was checked by a step-by-step process. Commissioning observations followed the initial operation. We here present pre- and post-launch results verifying the Xtend performance. All the in-orbit performances are consistent with those measured on ground and satisfy the mission requirement. Extensive calibration studies are ongoing.
△ Less
Submitted 28 June, 2024;
originally announced June 2024.
-
Initial operations of the Soft X-ray Imager onboard XRISM
Authors:
Hiromasa Suzuki,
Tomokage Yoneyama,
Shogo B. Kobayashi,
Hirofumi Noda,
Hiroyuki Uchida,
Kumiko K. Nobukawa,
Kouichi Hagino,
Koji Mori,
Hiroshi Tomida,
Hiroshi Nakajima,
Takaaki Tanaka,
Hiroshi Murakami,
Hideki Uchiyama,
Masayoshi Nobukawa,
Yoshiaki Kanemaru,
Yoshinori Otsuka,
Haruhiko Yokosu,
Wakana Yonemaru,
Hanako Nakano,
Kazuhiro Ichikawa,
Reo Takemoto,
Tsukasa Matsushima,
Marina Yoshimoto,
Mio Aoyagi,
Kohei Shima
, et al. (30 additional authors not shown)
Abstract:
XRISM (X-Ray Imaging and Spectroscopy Mission) is an astronomical satellite with the capability of high-resolution spectroscopy with the X-ray microcalorimeter, Resolve, and wide field-of-view imaging with the CCD camera, Xtend. Xtend consists of the mirror assembly (XMA: X-ray Mirror Assembly) and detector (SXI: Soft X-ray Imager). The SXI is composed of CCDs, analog and digital electronics, and…
▽ More
XRISM (X-Ray Imaging and Spectroscopy Mission) is an astronomical satellite with the capability of high-resolution spectroscopy with the X-ray microcalorimeter, Resolve, and wide field-of-view imaging with the CCD camera, Xtend. Xtend consists of the mirror assembly (XMA: X-ray Mirror Assembly) and detector (SXI: Soft X-ray Imager). The SXI is composed of CCDs, analog and digital electronics, and a mechanical cooler. After the successful launch on September 6th, 2023 (UT) and subsequent critical operations, the mission instruments were turned on and set up. The CCDs have been kept at the designed operating temperature of $-110^\circ$C after the electronics and cooling system were successfully set up. During the initial operation phase, which continued for more than a month after the critical operations, we verified the observation procedure, stability of the cooling system, all the observation options with different imaging areas and/or timing resolutions, and time-tagged and automated operations including those for South Atlantic Anomaly passages. We optimized the operation procedure and observation parameters including the cooler settings, imaging areas for the small window modes, and event selection algorithm. We summarize our policy and procedure of the initial operations for the SXI. We also report on a couple of issues we faced during the initial operations and lessons learned from them.
△ Less
Submitted 14 February, 2025; v1 submitted 28 June, 2024;
originally announced June 2024.
-
Necessary and Sufficient Conditions for Capacity-Achieving Private Information Retrieval with Non-Colluding and Colluding Servers
Authors:
Atsushi Miki,
Yusuke Morishita,
Toshiyasu Matsushima
Abstract:
Private Information Retrieval (PIR) is a mechanism for efficiently downloading messages while keeping the index secret. Here, PIRs in which servers do not communicate with each other are called standard PIRs, and PIRs in which some servers communicate with each other are called colluding PIRs. The information-theoretic upper bound on efficiency has been given in previous studies. However, the cond…
▽ More
Private Information Retrieval (PIR) is a mechanism for efficiently downloading messages while keeping the index secret. Here, PIRs in which servers do not communicate with each other are called standard PIRs, and PIRs in which some servers communicate with each other are called colluding PIRs. The information-theoretic upper bound on efficiency has been given in previous studies. However, the conditions for PIRs to keep privacy, to decode the desired message, and to achieve that upper bound have not been clarified in matrix form. In this paper, we prove the necessary and sufficient conditions for the properties of standard PIR and colluding PIR. Further, we represent the properties in matrix form.
△ Less
Submitted 12 October, 2024; v1 submitted 21 April, 2024;
originally announced April 2024.
-
An Algorithmic Framework for Constructing Multiple Decision Trees by Evaluating Their Combination Performance Throughout the Construction Process
Authors:
Keito Tajima,
Naoki Ichijo,
Yuta Nakahara,
Toshiyasu Matsushima
Abstract:
Predictions using a combination of decision trees are known to be effective in machine learning. Typical ideas for constructing a combination of decision trees for prediction are bagging and boosting. Bagging independently constructs decision trees without evaluating their combination performance and averages them afterward. Boosting constructs decision trees sequentially, only evaluating a combin…
▽ More
Predictions using a combination of decision trees are known to be effective in machine learning. Typical ideas for constructing a combination of decision trees for prediction are bagging and boosting. Bagging independently constructs decision trees without evaluating their combination performance and averages them afterward. Boosting constructs decision trees sequentially, only evaluating a combination performance of a new decision tree and the fixed past decision trees at each step. Therefore, neither method directly constructs nor evaluates a combination of decision trees for the final prediction. When the final prediction is based on a combination of decision trees, it is natural to evaluate the appropriateness of the combination when constructing them. In this study, we propose a new algorithmic framework that constructs decision trees simultaneously and evaluates their combination performance throughout the construction process. Our framework repeats two procedures. In the first procedure, we construct new candidates of combinations of decision trees to find a proper combination of decision trees. In the second procedure, we evaluate each combination performance of decision trees under some criteria and select a better combination. To confirm the performance of the proposed framework, we perform experiments on synthetic and benchmark data.
△ Less
Submitted 9 February, 2024;
originally announced February 2024.
-
Boosting-Based Sequential Meta-Tree Ensemble Construction for Improved Decision Trees
Authors:
Ryota Maniwa,
Naoki Ichijo,
Yuta Nakahara,
Toshiyasu Matsushima
Abstract:
A decision tree is one of the most popular approaches in machine learning fields. However, it suffers from the problem of overfitting caused by overly deepened trees. Then, a meta-tree is recently proposed. It solves the problem of overfitting caused by overly deepened trees. Moreover, the meta-tree guarantees statistical optimality based on Bayes decision theory. Therefore, the meta-tree is expec…
▽ More
A decision tree is one of the most popular approaches in machine learning fields. However, it suffers from the problem of overfitting caused by overly deepened trees. Then, a meta-tree is recently proposed. It solves the problem of overfitting caused by overly deepened trees. Moreover, the meta-tree guarantees statistical optimality based on Bayes decision theory. Therefore, the meta-tree is expected to perform better than the decision tree. In contrast to a single decision tree, it is known that ensembles of decision trees, which are typically constructed boosting algorithms, are more effective in improving predictive performance. Thus, it is expected that ensembles of meta-trees are more effective in improving predictive performance than a single meta-tree, and there are no previous studies that construct multiple meta-trees in boosting. Therefore, in this study, we propose a method to construct multiple meta-trees using a boosting approach. Through experiments with synthetic and benchmark datasets, we conduct a performance comparison between the proposed methods and the conventional methods using ensembles of decision trees. Furthermore, while ensembles of decision trees can cause overfitting as well as a single decision tree, experiments confirmed that ensembles of meta-trees can prevent overfitting due to the tree depth.
△ Less
Submitted 9 February, 2024;
originally announced February 2024.
-
Real-World Robot Applications of Foundation Models: A Review
Authors:
Kento Kawaharazuka,
Tatsuya Matsushima,
Andrew Gambardella,
Jiaxian Guo,
Chris Paxton,
Andy Zeng
Abstract:
Recent developments in foundation models, like Large Language Models (LLMs) and Vision-Language Models (VLMs), trained on extensive data, facilitate flexible application across different tasks and modalities. Their impact spans various fields, including healthcare, education, and robotics. This paper provides an overview of the practical application of foundation models in real-world robotics, wit…
▽ More
Recent developments in foundation models, like Large Language Models (LLMs) and Vision-Language Models (VLMs), trained on extensive data, facilitate flexible application across different tasks and modalities. Their impact spans various fields, including healthcare, education, and robotics. This paper provides an overview of the practical application of foundation models in real-world robotics, with a primary emphasis on the replacement of specific components within existing robot systems. The summary encompasses the perspective of input-output relationships in foundation models, as well as their role in perception, motion planning, and control within the field of robotics. This paper concludes with a discussion of future challenges and implications for practical robot applications.
△ Less
Submitted 22 October, 2024; v1 submitted 8 February, 2024;
originally announced February 2024.
-
Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Authors:
Open X-Embodiment Collaboration,
Abby O'Neill,
Abdul Rehman,
Abhinav Gupta,
Abhiram Maddukuri,
Abhishek Gupta,
Abhishek Padalkar,
Abraham Lee,
Acorn Pooley,
Agrim Gupta,
Ajay Mandlekar,
Ajinkya Jain,
Albert Tung,
Alex Bewley,
Alex Herzog,
Alex Irpan,
Alexander Khazatsky,
Anant Rai,
Anchit Gupta,
Andrew Wang,
Andrey Kolobov,
Anikait Singh,
Animesh Garg,
Aniruddha Kembhavi,
Annie Xie
, et al. (269 additional authors not shown)
Abstract:
Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for many applications. Can such a consolidation happen in robotics? Conventionally, robotic learning method…
▽ More
Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for many applications. Can such a consolidation happen in robotics? Conventionally, robotic learning methods train a separate model for every application, every robot, and even every environment. Can we instead train generalist X-robot policy that can be adapted efficiently to new robots, tasks, and environments? In this paper, we provide datasets in standardized data formats and models to make it possible to explore this possibility in the context of robotic manipulation, alongside experimental results that provide an example of effective X-robot policies. We assemble a dataset from 22 different robots collected through a collaboration between 21 institutions, demonstrating 527 skills (160266 tasks). We show that a high-capacity model trained on this data, which we call RT-X, exhibits positive transfer and improves the capabilities of multiple robots by leveraging experience from other platforms. More details can be found on the project website https://robotics-transformer-x.github.io.
△ Less
Submitted 14 May, 2025; v1 submitted 13 October, 2023;
originally announced October 2023.
-
TRAIL Team Description Paper for RoboCup@Home 2023
Authors:
Chikaha Tsuji,
Dai Komukai,
Mimo Shirasaka,
Hikaru Wada,
Tsunekazu Omija,
Aoi Horo,
Daiki Furuta,
Saki Yamaguchi,
So Ikoma,
Soshi Tsunashima,
Masato Kobayashi,
Koki Ishimoto,
Yuya Ikeda,
Tatsuya Matsushima,
Yusuke Iwasawa,
Yutaka Matsuo
Abstract:
Our team, TRAIL, consists of AI/ML laboratory members from The University of Tokyo. We leverage our extensive research experience in state-of-the-art machine learning to build general-purpose in-home service robots. We previously participated in two competitions using Human Support Robot (HSR): RoboCup@Home Japan Open 2020 (DSPL) and World Robot Summit 2020, equivalent to RoboCup World Tournament.…
▽ More
Our team, TRAIL, consists of AI/ML laboratory members from The University of Tokyo. We leverage our extensive research experience in state-of-the-art machine learning to build general-purpose in-home service robots. We previously participated in two competitions using Human Support Robot (HSR): RoboCup@Home Japan Open 2020 (DSPL) and World Robot Summit 2020, equivalent to RoboCup World Tournament. Throughout the competitions, we showed that a data-driven approach is effective for performing in-home tasks. Aiming for further development of building a versatile and fast-adaptable system, in RoboCup @Home 2023, we unify three technologies that have recently been evaluated as components in the fields of deep learning and robot learning into a real household robot system. In addition, to stimulate research all over the RoboCup@Home community, we build a platform that manages data collected from each site belonging to the community around the world, taking advantage of the characteristics of the community.
△ Less
Submitted 5 October, 2023;
originally announced October 2023.
-
Self-Recovery Prompting: Promptable General Purpose Service Robot System with Foundation Models and Self-Recovery
Authors:
Mimo Shirasaka,
Tatsuya Matsushima,
Soshi Tsunashima,
Yuya Ikeda,
Aoi Horo,
So Ikoma,
Chikaha Tsuji,
Hikaru Wada,
Tsunekazu Omija,
Dai Komukai,
Yutaka Matsuo Yusuke Iwasawa
Abstract:
A general-purpose service robot (GPSR), which can execute diverse tasks in various environments, requires a system with high generalizability and adaptability to tasks and environments. In this paper, we first developed a top-level GPSR system for worldwide competition (RoboCup@Home 2023) based on multiple foundation models. This system is both generalizable to variations and adaptive by prompting…
▽ More
A general-purpose service robot (GPSR), which can execute diverse tasks in various environments, requires a system with high generalizability and adaptability to tasks and environments. In this paper, we first developed a top-level GPSR system for worldwide competition (RoboCup@Home 2023) based on multiple foundation models. This system is both generalizable to variations and adaptive by prompting each model. Then, by analyzing the performance of the developed system, we found three types of failure in more realistic GPSR application settings: insufficient information, incorrect plan generation, and plan execution failure. We then propose the self-recovery prompting pipeline, which explores the necessary information and modifies its prompts to recover from failure. We experimentally confirm that the system with the self-recovery mechanism can accomplish tasks by resolving various failure cases. Supplementary videos are available at https://sites.google.com/view/srgpsr .
△ Less
Submitted 26 September, 2023; v1 submitted 25 September, 2023;
originally announced September 2023.
-
GenDOM: Generalizable One-shot Deformable Object Manipulation with Parameter-Aware Policy
Authors:
So Kuroki,
Jiaxian Guo,
Tatsuya Matsushima,
Takuya Okubo,
Masato Kobayashi,
Yuya Ikeda,
Ryosuke Takanami,
Paul Yoo,
Yutaka Matsuo,
Yusuke Iwasawa
Abstract:
Due to the inherent uncertainty in their deformability during motion, previous methods in deformable object manipulation, such as rope and cloth, often required hundreds of real-world demonstrations to train a manipulation policy for each object, which hinders their applications in our ever-changing world. To address this issue, we introduce GenDOM, a framework that allows the manipulation policy…
▽ More
Due to the inherent uncertainty in their deformability during motion, previous methods in deformable object manipulation, such as rope and cloth, often required hundreds of real-world demonstrations to train a manipulation policy for each object, which hinders their applications in our ever-changing world. To address this issue, we introduce GenDOM, a framework that allows the manipulation policy to handle different deformable objects with only a single real-world demonstration. To achieve this, we augment the policy by conditioning it on deformable object parameters and training it with a diverse range of simulated deformable objects so that the policy can adjust actions based on different object parameters. At the time of inference, given a new object, GenDOM can estimate the deformable object parameters with only a single real-world demonstration by minimizing the disparity between the grid density of point clouds of real-world demonstrations and simulations in a differentiable physics simulator. Empirical validations on both simulated and real-world object manipulation setups clearly show that our method can manipulate different objects with a single demonstration and significantly outperforms the baseline in both environments (a 62% improvement for in-domain ropes and a 15% improvement for out-of-distribution ropes in simulation, as well as a 26% improvement for ropes and a 50% improvement for cloths in the real world), demonstrating the effectiveness of our approach in one-shot deformable object manipulation.
△ Less
Submitted 27 January, 2025; v1 submitted 16 September, 2023;
originally announced September 2023.
-
GenORM: Generalizable One-shot Rope Manipulation with Parameter-Aware Policy
Authors:
So Kuroki,
Jiaxian Guo,
Tatsuya Matsushima,
Takuya Okubo,
Masato Kobayashi,
Yuya Ikeda,
Ryosuke Takanami,
Paul Yoo,
Yutaka Matsuo,
Yusuke Iwasawa
Abstract:
Due to the inherent uncertainty in their deformability during motion, previous methods in rope manipulation often require hundreds of real-world demonstrations to train a manipulation policy for each rope, even for simple tasks such as rope goal reaching, which hinder their applications in our ever-changing world. To address this issue, we introduce GenORM, a framework that allows the manipulation…
▽ More
Due to the inherent uncertainty in their deformability during motion, previous methods in rope manipulation often require hundreds of real-world demonstrations to train a manipulation policy for each rope, even for simple tasks such as rope goal reaching, which hinder their applications in our ever-changing world. To address this issue, we introduce GenORM, a framework that allows the manipulation policy to handle different deformable ropes with a single real-world demonstration. To achieve this, we augment the policy by conditioning it on deformable rope parameters and training it with a diverse range of simulated deformable ropes so that the policy can adjust actions based on different rope parameters. At the time of inference, given a new rope, GenORM estimates the deformable rope parameters by minimizing the disparity between the grid density of point clouds of real-world demonstrations and simulations. With the help of a differentiable physics simulator, we require only a single real-world demonstration. Empirical validations on both simulated and real-world rope manipulation setups clearly show that our method can manipulate different ropes with a single demonstration and significantly outperforms the baseline in both environments (62% improvement in in-domain ropes, and 15% improvement in out-of-distribution ropes in simulation, 26% improvement in real-world), demonstrating the effectiveness of our approach in one-shot rope manipulation.
△ Less
Submitted 27 January, 2025; v1 submitted 13 June, 2023;
originally announced June 2023.
-
Prediction Algorithms Achieving Bayesian Decision Theoretical Optimality Based on Decision Trees as Data Observation Processes
Authors:
Yuta Nakahara,
Shota Saito,
Naoki Ichijo,
Koki Kazama,
Toshiyasu Matsushima
Abstract:
In the field of decision trees, most previous studies have difficulty ensuring the statistical optimality of a prediction of new data and suffer from overfitting because trees are usually used only to represent prediction functions to be constructed from given data. In contrast, some studies, including this paper, used the trees to represent stochastic data observation processes behind given data.…
▽ More
In the field of decision trees, most previous studies have difficulty ensuring the statistical optimality of a prediction of new data and suffer from overfitting because trees are usually used only to represent prediction functions to be constructed from given data. In contrast, some studies, including this paper, used the trees to represent stochastic data observation processes behind given data. Moreover, they derived the statistically optimal prediction, which is robust against overfitting, based on the Bayesian decision theory by assuming a prior distribution for the trees. However, these studies still have a problem in computing this Bayes optimal prediction because it involves an infeasible summation for all division patterns of a feature space, which is represented by the trees and some parameters. In particular, an open problem is a summation with respect to combinations of division axes, i.e., the assignment of features to inner nodes of the tree. We solve this by a Markov chain Monte Carlo method, whose step size is adaptively tuned according to a posterior distribution for the trees.
△ Less
Submitted 12 June, 2023;
originally announced June 2023.
-
Batch Updating of a Posterior Tree Distribution over a Meta-Tree
Authors:
Yuta Nakahara,
Toshiyasu Matsushima
Abstract:
Previously, we proposed a probabilistic data generation model represented by an unobservable tree and a sequential updating method to calculate a posterior distribution over a set of trees. The set is called a meta-tree. In this paper, we propose a more efficient batch updating method.
Previously, we proposed a probabilistic data generation model represented by an unobservable tree and a sequential updating method to calculate a posterior distribution over a set of trees. The set is called a meta-tree. In this paper, we propose a more efficient batch updating method.
△ Less
Submitted 16 July, 2023; v1 submitted 16 March, 2023;
originally announced March 2023.
-
Collective Intelligence for 2D Push Manipulations with Mobile Robots
Authors:
So Kuroki,
Tatsuya Matsushima,
Jumpei Arima,
Hiroki Furuta,
Yutaka Matsuo,
Shixiang Shane Gu,
Yujin Tang
Abstract:
While natural systems often present collective intelligence that allows them to self-organize and adapt to changes, the equivalent is missing in most artificial systems. We explore the possibility of such a system in the context of cooperative 2D push manipulations using mobile robots. Although conventional works demonstrate potential solutions for the problem in restricted settings, they have com…
▽ More
While natural systems often present collective intelligence that allows them to self-organize and adapt to changes, the equivalent is missing in most artificial systems. We explore the possibility of such a system in the context of cooperative 2D push manipulations using mobile robots. Although conventional works demonstrate potential solutions for the problem in restricted settings, they have computational and learning difficulties. More importantly, these systems do not possess the ability to adapt when facing environmental changes. In this work, we show that by distilling a planner derived from a differentiable soft-body physics simulator into an attention-based neural network, our multi-robot push manipulation system achieves better performance than baselines. In addition, our system also generalizes to configurations not seen during training and is able to adapt toward task completions when external turbulence and environmental changes are applied. Supplementary videos can be found on our project website: https://sites.google.com/view/ciom/home
△ Less
Submitted 27 January, 2025; v1 submitted 28 November, 2022;
originally announced November 2022.
-
Coordinated Stress-Structure Self-Organization in Granular Packing
Authors:
Xiaoyu Jiang,
Raphael Blumenfeld,
Takashi Matsushima
Abstract:
During quasi-static dynamics of granular systems, the stress and structure self-organise, but there is currently no quantitative measure or understanding of this phenomenon. Such an understanding is essential because local structural properties of the settled material are then correlated with the local stress, which calls into question existing linear theories of stress transmission in granular me…
▽ More
During quasi-static dynamics of granular systems, the stress and structure self-organise, but there is currently no quantitative measure or understanding of this phenomenon. Such an understanding is essential because local structural properties of the settled material are then correlated with the local stress, which calls into question existing linear theories of stress transmission in granular media. A method to quantify the local stress-structure correlations is necessary for addressing this issue and we present here such a method for planar systems. We then use it to analyze numerically several different systems, compressed quasi-statically by two different procedures. We define cells, cell orders, cell orientations, and cell stresses and report the following results. 1. Cells orient along the local stress major principal axes. 2. The mean ratio of cell principal stresses decreases with cell order and increases with friction. 3. The ratio distributions collapse onto a single curve under a simple scaling, for all packing protocols and friction coefficients. 4. A constructed model explains the correlations between the local cell and stress principal axis orientations. 5. The collapse of the stress ratios onto a Weibull distribution is explained theoretically. Our results quantify the cooperative stress-structure self-organization and provide a way to relate quantitatively the stress-structure coupling to different process parameters and particle characteristics. Significantly, the strong stress-structure correlation, driven by structural re-organization upon application of external stress, suggests that current stress theories of granular matter need to be revisited.
△ Less
Submitted 27 October, 2024; v1 submitted 13 August, 2022;
originally announced August 2022.
-
World Robot Challenge 2020 -- Partner Robot: A Data-Driven Approach for Room Tidying with Mobile Manipulator
Authors:
Tatsuya Matsushima,
Yuki Noguchi,
Jumpei Arima,
Toshiki Aoki,
Yuki Okita,
Yuya Ikeda,
Koki Ishimoto,
Shohei Taniguchi,
Yuki Yamashita,
Shoichi Seto,
Shixiang Shane Gu,
Yusuke Iwasawa,
Yutaka Matsuo
Abstract:
Tidying up a household environment using a mobile manipulator poses various challenges in robotics, such as adaptation to large real-world environmental variations, and safe and robust deployment in the presence of humans.The Partner Robot Challenge in World Robot Challenge (WRC) 2020, a global competition held in September 2021, benchmarked tidying tasks in the real home environments, and importa…
▽ More
Tidying up a household environment using a mobile manipulator poses various challenges in robotics, such as adaptation to large real-world environmental variations, and safe and robust deployment in the presence of humans.The Partner Robot Challenge in World Robot Challenge (WRC) 2020, a global competition held in September 2021, benchmarked tidying tasks in the real home environments, and importantly, tested for full system performances.For this challenge, we developed an entire household service robot system, which leverages a data-driven approach to adapt to numerous edge cases that occur during the execution, instead of classical manual pre-programmed solutions. In this paper, we describe the core ingredients of the proposed robot system, including visual recognition, object manipulation, and motion planning. Our robot system won the second prize, verifying the effectiveness and potential of data-driven robot systems for mobile manipulation in home environments.
△ Less
Submitted 21 July, 2022; v1 submitted 20 July, 2022;
originally announced July 2022.
-
An Algorithm for Computing the Stratonovich's Value of Information
Authors:
Akira Kamatsuka,
Takahiro Yoshida,
Koki Kazama,
Toshiyasu Matsushima
Abstract:
We propose an algorithm for computing Stratonovich's value of information (VoI) that can be regarded as an analogue of the distortion-rate function. We construct an alternating optimization algorithm for VoI under a general information leakage constraint and derive a convergence condition. Furthermore, we discuss algorithms for computing VoI under specific information leakage constraints, such as…
▽ More
We propose an algorithm for computing Stratonovich's value of information (VoI) that can be regarded as an analogue of the distortion-rate function. We construct an alternating optimization algorithm for VoI under a general information leakage constraint and derive a convergence condition. Furthermore, we discuss algorithms for computing VoI under specific information leakage constraints, such as Shannon's mutual information (MI), $f$-leakage, Arimoto's MI, Sibson's MI, and Csiszar's MI.
△ Less
Submitted 8 May, 2022; v1 submitted 5 May, 2022;
originally announced May 2022.
-
Stochastic 2D Signal Generative Model with Wavelet Packets Basis Regarded as a Random Variable and Bayes Optimal Processing
Authors:
Ryohei Oka,
Yuta Nakahara,
Toshiyasu Matsushima
Abstract:
This study deals with two-dimensional (2D) signal processing using the wavelet packet transform. When the basis is unknown the candidate of basis increases in exponential order with respect to the signal size. Previous studies do not consider the basis as a random vaiables. Therefore, the cost function needs to be used to select a basis. However, this method is often a heuristic and a greedy searc…
▽ More
This study deals with two-dimensional (2D) signal processing using the wavelet packet transform. When the basis is unknown the candidate of basis increases in exponential order with respect to the signal size. Previous studies do not consider the basis as a random vaiables. Therefore, the cost function needs to be used to select a basis. However, this method is often a heuristic and a greedy search because it is impossible to search all the candidates for a huge number of bases. Therefore, it is difficult to evaluate the entire signal processing under a criterion and also it does not always gurantee the optimality of the entire signal processing. In this study, we propose a stochastic generative model in which the basis is regarded as a random variable. This makes it possible to evaluate entire signal processing under a unified criterion i.e. Bayes criterion. Moreover we can derive an optimal signal processing scheme that achieves the theoretical limit. This derived scheme shows that all the bases should be combined according to the posterior in stead of selecting a single basis. Although exponential order calculations is required for this scheme, we have derived a recursive algorithm for this scheme, which successfully reduces the computational complexity from the exponential order to the polynomial order.
△ Less
Submitted 1 May, 2022; v1 submitted 26 January, 2022;
originally announced February 2022.
-
A Generalization of the Stratonovich's Value of Information and Application to Privacy-Utility Trade-off
Authors:
Akira Kamatsuka,
Takahiro Yoshida,
Toshiyasu Matsushima
Abstract:
The Stratonovich's value of information (VoI) is quantity that measure how much inferential gain is obtained from a perturbed sample under information leakage constraint. In this paper, we introduce a generalized VoI for a general loss function and general information leakage. Then we derive an upper bound of the generalized VoI. Moreover, for a classical loss function, we provide a achievable con…
▽ More
The Stratonovich's value of information (VoI) is quantity that measure how much inferential gain is obtained from a perturbed sample under information leakage constraint. In this paper, we introduce a generalized VoI for a general loss function and general information leakage. Then we derive an upper bound of the generalized VoI. Moreover, for a classical loss function, we provide a achievable condition of the upper bound which is weaker than that of in previous studies. Since VoI can be viewed as a formulation of a privacy-utility trade-off (PUT) problem, we provide an interpretation of the achievable condition in the PUT context.
△ Less
Submitted 27 January, 2022;
originally announced January 2022.
-
Probability Distribution on Rooted Trees
Authors:
Yuta Nakahara,
Shota Saito,
Akira Kamatsuka,
Toshiyasu Matsushima
Abstract:
The hierarchical and recursive expressive capability of rooted trees is applicable to represent statistical models in various areas, such as data compression, image processing, and machine learning. On the other hand, such hierarchical expressive capability causes a problem in tree selection to avoid overfitting. One unified approach to solve this is a Bayesian approach, on which the rooted tree i…
▽ More
The hierarchical and recursive expressive capability of rooted trees is applicable to represent statistical models in various areas, such as data compression, image processing, and machine learning. On the other hand, such hierarchical expressive capability causes a problem in tree selection to avoid overfitting. One unified approach to solve this is a Bayesian approach, on which the rooted tree is regarded as a random variable and a direct loss function can be assumed on the selected model or the predicted value for a new data point. However, all the previous studies on this approach are based on the probability distribution on full trees, to the best of our knowledge. In this paper, we propose a generalized probability distribution for any rooted trees in which only the maximum number of child nodes and the maximum depth are fixed. Furthermore, we derive recursive methods to evaluate the characteristics of the probability distribution without any approximations.
△ Less
Submitted 24 January, 2022;
originally announced January 2022.
-
Tool as Embodiment for Recursive Manipulation
Authors:
Yuki Noguchi,
Tatsuya Matsushima,
Yutaka Matsuo,
Shixiang Shane Gu
Abstract:
Humans and many animals exhibit a robust capability to manipulate diverse objects, often directly with their bodies and sometimes indirectly with tools. Such flexibility is likely enabled by the fundamental consistency in underlying physics of object manipulation such as contacts and force closures. Inspired by viewing tools as extensions of our bodies, we present Tool-As-Embodiment (TAE), a param…
▽ More
Humans and many animals exhibit a robust capability to manipulate diverse objects, often directly with their bodies and sometimes indirectly with tools. Such flexibility is likely enabled by the fundamental consistency in underlying physics of object manipulation such as contacts and force closures. Inspired by viewing tools as extensions of our bodies, we present Tool-As-Embodiment (TAE), a parameterization for tool-based manipulation policies that treat hand-object and tool-object interactions in the same representation space. The result is a single policy that can be applied recursively on robots to use end effectors to manipulate objects, and use objects as tools, i.e. new end-effectors, to manipulate other objects. By sharing experiences across different embodiments for grasping or pushing, our policy exhibits higher performance than if separate policies were trained. Our framework could utilize all experiences from different resolutions of tool-enabled embodiments to a single generic policy for each manipulation skill. Videos at https://sites.google.com/view/recursivemanipulation
△ Less
Submitted 1 December, 2021;
originally announced December 2021.
-
Probability Distribution on Full Rooted Trees
Authors:
Yuta Nakahara,
Shota Saito,
Akira Kamatsuka,
Toshiyasu Matsushima
Abstract:
The recursive and hierarchical structure of full rooted trees is applicable to represent statistical models in various areas, such as data compression, image processing, and machine learning. In most of these cases, the full rooted tree is not a random variable; as such, model selection to avoid overfitting becomes problematic. A method to solve this problem is to assume a prior distribution on th…
▽ More
The recursive and hierarchical structure of full rooted trees is applicable to represent statistical models in various areas, such as data compression, image processing, and machine learning. In most of these cases, the full rooted tree is not a random variable; as such, model selection to avoid overfitting becomes problematic. A method to solve this problem is to assume a prior distribution on the full rooted trees. This enables the optimal model selection based on the Bayes decision theory. For example, by assigning a low prior probability to a complex model, the maximum a posteriori estimator prevents the selection of the complex one. Furthermore, we can average all the models weighted by their posteriors. In this paper, we propose a probability distribution on a set of full rooted trees. Its parametric representation is suitable for calculating the properties of our distribution using recursive functions, such as the mode, expectation, and posterior distribution. Although such distributions have been proposed in previous studies, they are only applicable to specific applications. Therefore, we extract their mathematically essential components and derive new generalized methods to calculate the expectation, posterior distribution, etc.
△ Less
Submitted 23 January, 2022; v1 submitted 27 September, 2021;
originally announced September 2021.
-
A Stochastic Model for Block Segmentation of Images Based on the Quadtree and the Bayes Code for It
Authors:
Yuta Nakahara,
Toshiyasu Matsushima
Abstract:
In information theory, lossless compression of general data is based on an explicit assumption of a stochastic generative model on target data. However, in lossless image compression, the researchers have mainly focused on the coding procedure that outputs the coded sequence from the input image, and the assumption of the stochastic generative model is implicit. In these studies, there is a diffic…
▽ More
In information theory, lossless compression of general data is based on an explicit assumption of a stochastic generative model on target data. However, in lossless image compression, the researchers have mainly focused on the coding procedure that outputs the coded sequence from the input image, and the assumption of the stochastic generative model is implicit. In these studies, there is a difficulty in confirming the information-theoretical optimality of the coding procedure to the stochastic generative model. Hence, in this paper, we propose a novel stochastic generative model of images by redefining the implicit stochastic generative model in a previous coding procedure. That is based on the quadtree so that our model effectively represents the variable block size segmentation of images. Then, we construct the Bayes code optimal for the proposed stochastic generative model. In general, the computational cost to calculate the posterior distribution required in the Bayes code increases exponentially for the image size. However, we introduce an efficient algorithm to calculate it in the polynomial order of the image size without loss of the optimality. Some experiments are performed to confirm the flexibility of the proposed stochastic model and the efficiency of the introduced algorithm.
△ Less
Submitted 7 June, 2021;
originally announced June 2021.
-
An Efficient Bayes Coding Algorithm for the Non-Stationary Source in Which Context Tree Model Varies from Interval to Interval
Authors:
Koshi Shimada,
Shota Saito,
Toshiyasu Matsushima
Abstract:
The context tree source is a source model in which the occurrence probability of symbols is determined from a finite past sequence, and is a broader class of sources that includes i.i.d. and Markov sources. The proposed source model in this paper represents that a subsequence in each interval is generated from a different context tree model. The Bayes code for such sources requires weighting of th…
▽ More
The context tree source is a source model in which the occurrence probability of symbols is determined from a finite past sequence, and is a broader class of sources that includes i.i.d. and Markov sources. The proposed source model in this paper represents that a subsequence in each interval is generated from a different context tree model. The Bayes code for such sources requires weighting of the posterior probability distributions for the change patterns of the context tree source and for all possible context tree models. Therefore, the challenge is how to reduce this exponential order computational complexity. In this paper, we assume a special class of prior probability distribution of change patterns and context tree models, and propose an efficient Bayes coding algorithm whose computational complexity is the polynomial order.
△ Less
Submitted 13 May, 2021; v1 submitted 11 May, 2021;
originally announced May 2021.
-
Co-Adaptation of Algorithmic and Implementational Innovations in Inference-based Deep Reinforcement Learning
Authors:
Hiroki Furuta,
Tadashi Kozuno,
Tatsuya Matsushima,
Yutaka Matsuo,
Shixiang Shane Gu
Abstract:
Recently many algorithms were devised for reinforcement learning (RL) with function approximation. While they have clear algorithmic distinctions, they also have many implementation differences that are algorithm-independent and sometimes under-emphasized. Such mixing of algorithmic novelty and implementation craftsmanship makes rigorous analyses of the sources of performance improvements across a…
▽ More
Recently many algorithms were devised for reinforcement learning (RL) with function approximation. While they have clear algorithmic distinctions, they also have many implementation differences that are algorithm-independent and sometimes under-emphasized. Such mixing of algorithmic novelty and implementation craftsmanship makes rigorous analyses of the sources of performance improvements across algorithms difficult. In this work, we focus on a series of off-policy inference-based actor-critic algorithms -- MPO, AWR, and SAC -- to decouple their algorithmic innovations and implementation decisions. We present unified derivations through a single control-as-inference objective, where we can categorize each algorithm as based on either Expectation-Maximization (EM) or direct Kullback-Leibler (KL) divergence minimization and treat the rest of specifications as implementation details. We performed extensive ablation studies, and identified substantial performance drops whenever implementation details are mismatched for algorithmic choices. These results show which implementation or code details are co-adapted and co-evolved with algorithms, and which are transferable across algorithms: as examples, we identified that tanh Gaussian policy and network sizes are highly adapted to algorithmic types, while layer normalization and ELU are critical for MPO's performances but also transfer to noticeable gains in SAC. We hope our work can inspire future work to further demystify sources of performance improvements across multiple algorithms and allow researchers to build on one another's both algorithmic and implementational innovations.
△ Less
Submitted 25 October, 2021; v1 submitted 31 March, 2021;
originally announced March 2021.
-
Policy Information Capacity: Information-Theoretic Measure for Task Complexity in Deep Reinforcement Learning
Authors:
Hiroki Furuta,
Tatsuya Matsushima,
Tadashi Kozuno,
Yutaka Matsuo,
Sergey Levine,
Ofir Nachum,
Shixiang Shane Gu
Abstract:
Progress in deep reinforcement learning (RL) research is largely enabled by benchmark task environments. However, analyzing the nature of those environments is often overlooked. In particular, we still do not have agreeable ways to measure the difficulty or solvability of a task, given that each has fundamentally different actions, observations, dynamics, rewards, and can be tackled with diverse R…
▽ More
Progress in deep reinforcement learning (RL) research is largely enabled by benchmark task environments. However, analyzing the nature of those environments is often overlooked. In particular, we still do not have agreeable ways to measure the difficulty or solvability of a task, given that each has fundamentally different actions, observations, dynamics, rewards, and can be tackled with diverse RL algorithms. In this work, we propose policy information capacity (PIC) -- the mutual information between policy parameters and episodic return -- and policy-optimal information capacity (POIC) -- between policy parameters and episodic optimality -- as two environment-agnostic, algorithm-agnostic quantitative metrics for task difficulty. Evaluating our metrics across toy environments as well as continuous control benchmark tasks from OpenAI Gym and DeepMind Control Suite, we empirically demonstrate that these information-theoretic metrics have higher correlations with normalized task solvability scores than a variety of alternatives. Lastly, we show that these metrics can also be used for fast and compute-efficient optimizations of key design parameters such as reward shaping, policy architectures, and MDP properties for better solvability by RL algorithms without ever running full RL experiments.
△ Less
Submitted 31 May, 2021; v1 submitted 23 March, 2021;
originally announced March 2021.