-
StreamSoccer: Event-Driven Memory for Streaming Soccer Commentary
Authors:
Chenxi Shao,
Bozhong Wang,
Jiaxin Huang,
Zhao Liu,
Sunwei Zhu,
Tianxin Hang,
Gaoqi He,
Yang Li,
Changbo Wang
Abstract:
Streaming video understanding requires models to causally update state as video arrives and organize growing history into semantic units that can evolve, persist, and be recalled under bounded computation and memory. This challenge is pronounced in live soccer commentary, where a system must describe completed events, summarize recent play, recall earlier events, or remain silent using only inform…
▽ More
Streaming video understanding requires models to causally update state as video arrives and organize growing history into semantic units that can evolve, persist, and be recalled under bounded computation and memory. This challenge is pronounced in live soccer commentary, where a system must describe completed events, summarize recent play, recall earlier events, or remain silent using only information available before each utterance. We present StreamSoccer, an event-driven system that uses event memory as its intermediate representation. A fixed-budget active memory integrates the stream; completed event states are retained locally and consolidated into retrievable historical records. A unified generator uses current, recent, and historical context to produce three commentary modes, while a rule-assisted scheduler selects a mode or silence. Unlike streaming video-language models organized around frames, visual tokens, or caches, and soccer-commentary methods based on predefined clips or output timestamps, StreamSoccer explicitly models event lifecycles. We construct a three-track streaming soccer commentary dataset and a layered evaluation protocol. At common reference anchors, StreamSoccer obtains CIDEr scores of 38.62, 23.96, and 17.39 for current-event, recent-window, and historical-memory commentary, ranking first on the current-event and historical-memory tracks and second on recent-window. Controlled ablations show that local completed events improve all tracks and that the full system performs best on all three. Across 174 raw-video runs on 58 matches, per-minute RTF p95 ranges from 0.10 to 0.22 without sustained growth with match history. These results indicate that event memory supports streaming soccer commentary across temporal scopes while controlling long-history computation.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
V-numbers of powers of cover ideals of unimodular hypergraphs
Authors:
Nguyen Thu Hang,
Thanh Vu
Abstract:
Let $H$ be a unimodular hypergraph with cover ideal $J(H)$. We prove that the local $v$-numbers of $J(H)^t$ are linear in $t$ for all $t\ge1$. We further show that the global $v$-number of $J(H)^t$ is linear in $t$ for all $t\ge n-1$. Finally, we prove that the global $v$-number of the powers of the cover ideal of any tree is linear in $t$ for all $t\ge1$.
Let $H$ be a unimodular hypergraph with cover ideal $J(H)$. We prove that the local $v$-numbers of $J(H)^t$ are linear in $t$ for all $t\ge1$. We further show that the global $v$-number of $J(H)^t$ is linear in $t$ for all $t\ge n-1$. Finally, we prove that the global $v$-number of the powers of the cover ideal of any tree is linear in $t$ for all $t\ge1$.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Ordered alternating paths and the depth of symbolic powers of cover ideals of graphs
Authors:
Nguyen Thu Hang,
Nguyen Thi Thanh Tam,
Thanh Vu
Abstract:
Let $G$ be a simple graph with cover ideal $J(G)$ in a polynomial ring $S$ in $|V(G)|$ variables. For a matching $M$ of $G$, we denote by $\ell(M)$ the length of the longest $M$-alternating path in $G$. We define $α_t(G)$ to be the maximum size of an ordered matching $M$ of $G$ such that $\ell(M) \le 2t-1$. We then prove that $$\operatorname{depth}(S/J(G)^{(t)}) \le |V(G)| - 1 - α_t(G)$$ for all…
▽ More
Let $G$ be a simple graph with cover ideal $J(G)$ in a polynomial ring $S$ in $|V(G)|$ variables. For a matching $M$ of $G$, we denote by $\ell(M)$ the length of the longest $M$-alternating path in $G$. We define $α_t(G)$ to be the maximum size of an ordered matching $M$ of $G$ such that $\ell(M) \le 2t-1$. We then prove that $$\operatorname{depth}(S/J(G)^{(t)}) \le |V(G)| - 1 - α_t(G)$$ for all $t \ge 1$, where $J(G)^{(t)}$ denotes the $t$-th symbolic power of $J(G)$, and that equality holds when $G$ is a forest.
△ Less
Submitted 5 July, 2026;
originally announced July 2026.
-
Critical subgraphs and the regularity of symbolic powers of cover ideals of graphs
Authors:
Nguyen Thu Hang,
Thanh Vu
Abstract:
Let $G$ be a simple graph. We demonstrate a method for using $t$-admissible subgraphs of $G$ to determine the regularity of the $t$-th symbolic power of the cover ideal of $G$. As an application, we compute the regularity of powers of cover ideals of bipartite unicyclic graphs.
Let $G$ be a simple graph. We demonstrate a method for using $t$-admissible subgraphs of $G$ to determine the regularity of the $t$-th symbolic power of the cover ideal of $G$. As an application, we compute the regularity of powers of cover ideals of bipartite unicyclic graphs.
△ Less
Submitted 19 May, 2026;
originally announced May 2026.
-
Admissible subgraphs and the depth of symbolic powers of cover ideals of graphs
Authors:
Tran Duc Dung,
Nguyen Thu Hang,
Thanh Vu
Abstract:
Let $G$ be a simple graph. We introduce the notion of $t$-admissible subgraphs of $G$ and show how to use them to compute the depth of the $t$-th symbolic powers of the cover ideal of $G$. As an application, we prove that \[ \depth\big(S/J(C_n)^{(t)}\big) = n - 1 - \left\lfloor \frac{tn}{2t+1} \right\rfloor \] for all $t \ge 2$ and $n \ge 3$, where $S = K[x_1,\ldots,x_n]$ and $J(C_n)$ is the cover…
▽ More
Let $G$ be a simple graph. We introduce the notion of $t$-admissible subgraphs of $G$ and show how to use them to compute the depth of the $t$-th symbolic powers of the cover ideal of $G$. As an application, we prove that \[ \depth\big(S/J(C_n)^{(t)}\big) = n - 1 - \left\lfloor \frac{tn}{2t+1} \right\rfloor \] for all $t \ge 2$ and $n \ge 3$, where $S = K[x_1,\ldots,x_n]$ and $J(C_n)$ is the cover ideal of the cycle on $n$ vertices.
△ Less
Submitted 19 May, 2026; v1 submitted 5 May, 2026;
originally announced May 2026.
-
The depth function of powers of cover ideals of path graphs
Authors:
Tran Duc Dung,
Nguyen Thu Hang,
Pham Hong Nam,
Nguyen Thi Thanh Tam
Abstract:
Let $G=P_n$ be a path graph with cover ideal $J(P_n)$. By using Hochster's depth formula, we prove the explicit formulae to compute the depth functions of powers of cover ideals of paths.
Let $G=P_n$ be a path graph with cover ideal $J(P_n)$. By using Hochster's depth formula, we prove the explicit formulae to compute the depth functions of powers of cover ideals of paths.
△ Less
Submitted 5 May, 2026;
originally announced May 2026.
-
Meta-CoT: Enhancing Granularity and Generalization in Image Editing
Authors:
Shiyi Zhang,
Yiji Cheng,
Tiankai Hang,
Zijin Yin,
Runze He,
Yu Xu,
Wenxun Dai,
Yunlong Lin,
Chunyu Wang,
Qinglin Lu,
Yansong Tang
Abstract:
Unified multi-modal understanding/generative models have shown improved image editing performance by incorporating fine-grained understanding into their Chain-of-Thought (CoT) process. However, a critical question remains underexplored: what forms of CoT and training strategy can jointly enhance both the understanding granularity and generalization? To address this, we propose Meta-CoT, a paradigm…
▽ More
Unified multi-modal understanding/generative models have shown improved image editing performance by incorporating fine-grained understanding into their Chain-of-Thought (CoT) process. However, a critical question remains underexplored: what forms of CoT and training strategy can jointly enhance both the understanding granularity and generalization? To address this, we propose Meta-CoT, a paradigm that performs a two-level decomposition of any single-image editing operation with two key properties: (1) Decomposability. We observe that any editing intention can be represented as a triplet - (task, target, required understanding ability). Inspired by this, Meta-CoT decomposes both the editing task and the target, generating task-specific CoT and traversing editing operations on all targets. This decomposition enhances the model's understanding granularity of editing operations and guides it to learn each element of the triplet during training, substantially improving the editing capability. (2) Generalizability. In the second decomposition level, we further break down editing tasks into five fundamental meta-tasks. We find that training on these five meta-tasks, together with the other two elements of the triplet, is sufficient to achieve strong generalization across diverse, unseen editing tasks. To further align the model's editing behavior with its CoT reasoning, we introduce the CoT-Editing Consistency Reward, which encourages more accurate and effective utilization of CoT information during editing. Experiments demonstrate that our method achieves an overall 15.8% improvement across 21 editing tasks, and generalizes effectively to unseen editing tasks when trained on only a small set of meta-tasks. Our code, benchmark, and model are released at https://shiyi-zh0408.github.io/projectpages/Meta-CoT/
△ Less
Submitted 27 April, 2026;
originally announced April 2026.
-
LPM 1.0: Video-based Character Performance Model
Authors:
Ailing Zeng,
Casper Yang,
Chauncey Ge,
Eddie Zhang,
Garvey Xu,
Gavin Lin,
Gilbert Gu,
Jeremy Pi,
Leo Li,
Mingyi Shi,
Shawn Wang,
Sheng Bi,
Steven Tang,
Thorn Hang,
Tobey Guo,
Vincent Li,
Xin Tong,
Yikang Li,
Yuchen Sun,
Yue Zhao,
Yuhan Lu,
Yuwei Li,
Zane Zhang,
Zeshi Yang,
Zi Ye
Abstract:
Performance, the externalization of intent, emotion, and personality through visual, vocal, and temporal behavior, is what makes a character alive. Learning such performance from video is a promising alternative to traditional 3D pipelines. However, existing video models struggle to jointly achieve high expressiveness, real-time inference, and long-horizon identity stability, a tension we call the…
▽ More
Performance, the externalization of intent, emotion, and personality through visual, vocal, and temporal behavior, is what makes a character alive. Learning such performance from video is a promising alternative to traditional 3D pipelines. However, existing video models struggle to jointly achieve high expressiveness, real-time inference, and long-horizon identity stability, a tension we call the performance trilemma. Conversation is the most comprehensive performance scenario, as characters simultaneously speak, listen, react, and emote while maintaining identity over time. To address this, we present LPM 1.0 (Large Performance Model), focusing on single-person full-duplex audio-visual conversational performance. Concretely, we build a multimodal human-centric dataset through strict filtering, speaking-listening audio-video pairing, performance understanding, and identity-aware multi-reference extraction; train a 17B-parameter Diffusion Transformer (Base LPM) for highly controllable, identity-consistent performance through multimodal conditioning; and distill it into a causal streaming generator (Online LPM) for low-latency, infinite-length interaction. At inference, given a character image with identity-aware references, LPM 1.0 generates listening videos from user audio and speaking videos from synthesized audio, with text prompts for motion control, all at real-time speed with identity-stable, infinite-length generation. LPM 1.0 thus serves as a visual engine for conversational agents, live streaming characters, and game NPCs. To systematically evaluate this setting, we propose LPM-Bench, the first benchmark for interactive character performance. LPM 1.0 achieves state-of-the-art results across all evaluated dimensions while maintaining real-time inference.
△ Less
Submitted 14 April, 2026; v1 submitted 9 April, 2026;
originally announced April 2026.
-
Hierarchical SVG Tokenization: Learning Compact Visual Programs for Scalable Vector Graphics Modeling
Authors:
Ximing Xing,
Ziteng Xue,
Zhenxi Li,
Weicong Liang,
Linqing Wang,
Zhantao Yang,
Tiankai Hang,
Zijin Yin,
Qinglin Lu,
Chunyu Wang,
Qian Yu
Abstract:
Recent large language models have shifted SVG generation from differentiable rendering optimization to autoregressive program synthesis. However, existing approaches still rely on generic byte-level tokenization inherited from natural language processing, which poorly reflects the geometric structure of vector graphics. Numerical coordinates are fragmented into discrete symbols, destroying spatial…
▽ More
Recent large language models have shifted SVG generation from differentiable rendering optimization to autoregressive program synthesis. However, existing approaches still rely on generic byte-level tokenization inherited from natural language processing, which poorly reflects the geometric structure of vector graphics. Numerical coordinates are fragmented into discrete symbols, destroying spatial relationships and introducing severe token redundancy, often leading to coordinate hallucination and inefficient long-sequence generation. To address these challenges, we propose HiVG, a hierarchical SVG tokenization framework tailored for autoregressive vector graphics generation. HiVG decomposes raw SVG strings into structured \textit{atomic tokens} and further compresses executable command--parameter groups into geometry-constrained \textit{segment tokens}, substantially improving sequence efficiency while preserving syntactic validity. To further mitigate spatial mismatch, we introduce a Hierarchical Mean--Noise (HMN) initialization strategy that injects numerical ordering signals and semantic priors into new token embeddings. Combined with a curriculum training paradigm that progressively increases program complexity, HiVG enables more stable learning of executable SVG programs. Extensive experiments on both text-to-SVG and image-to-SVG tasks demonstrate improved generation fidelity, spatial consistency, and sequence efficiency compared with conventional tokenization schemes. Our code is publicly available at https://github.com/ximinng/HiVG
△ Less
Submitted 10 April, 2026; v1 submitted 6 April, 2026;
originally announced April 2026.
-
HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising
Authors:
Kai Zou,
Dian Zheng,
Hongbo Liu,
Tiankai Hang,
Bin Liu,
Nenghai Yu
Abstract:
Autoregressive (AR) diffusion offers a promising framework for generating videos of theoretically infinite length. However, a major challenge is maintaining temporal continuity while preventing the progressive quality degradation caused by error accumulation. To ensure continuity, existing methods typically condition on highly denoised contexts; yet, this practice propagates prediction errors with…
▽ More
Autoregressive (AR) diffusion offers a promising framework for generating videos of theoretically infinite length. However, a major challenge is maintaining temporal continuity while preventing the progressive quality degradation caused by error accumulation. To ensure continuity, existing methods typically condition on highly denoised contexts; yet, this practice propagates prediction errors with high certainty, thereby exacerbating degradation. In this paper, we argue that a highly clean context is unnecessary. Drawing inspiration from bidirectional diffusion models, which denoise frames at a shared noise level while maintaining coherence, we propose that conditioning on context at the same noise level as the current block provides sufficient signal for temporal consistency while effectively mitigating error propagation. Building on this insight, we propose HiAR, a hierarchical denoising framework that reverses the conventional generation order: instead of completing each block sequentially, it performs causal generation across all blocks at every denoising step, so that each block is always conditioned on context at the same noise level. This hierarchy naturally admits pipelined parallel inference, yielding a 1.8 wall-clock speedup in our 4-step setting. We further observe that self-rollout distillation under this paradigm amplifies a low-motion shortcut inherent to the mode-seeking reverse-KL objective. To counteract this, we introduce a forward-KL regulariser in bidirectional-attention mode, which preserves motion diversity for causal inference without interfering with the distillation loss. On VBench (20s generation), HiAR achieves the best overall score and the lowest temporal drift among all compared methods.
△ Less
Submitted 9 March, 2026;
originally announced March 2026.
-
Generative Visual Chain-of-Thought for Image Editing
Authors:
Zijin Yin,
Tiankai Hang,
Yiji Cheng,
Shiyi Zhang,
Runze He,
Yu Xu,
Chunyu Wang,
Bing Li,
Zheng Chang,
Kongming Liang,
Qinglin Lu,
Zhanyu Ma
Abstract:
Existing image editing methods struggle to perceive where to edit, especially under complex scenes and nuanced spatial instructions. To address this issue, we propose Generative Visual Chain-of-Thought (GVCoT), a unified framework that performs native visual reasoning by first generating spatial cues to localize the target region and then executing the edit. Unlike prior text-only CoT or tool-depe…
▽ More
Existing image editing methods struggle to perceive where to edit, especially under complex scenes and nuanced spatial instructions. To address this issue, we propose Generative Visual Chain-of-Thought (GVCoT), a unified framework that performs native visual reasoning by first generating spatial cues to localize the target region and then executing the edit. Unlike prior text-only CoT or tool-dependent visual CoT paradigms, GVCoT jointly optimizes visual tokens generated during the reasoning and editing phases in an end-to-end manner. This way fosters the emergence of innate spatial reasoning ability and enables more effective utilization of visual-domain cues. The main challenge of training GCVoT lies in the scarcity of large-scale editing data with precise edit region annotations; to this end, we construct GVCoT-Edit-Instruct, a dataset of 1.8M high-quality samples spanning 19 tasks. We adopt a progressive training strategy: supervised fine-tuning to build foundational localization ability in reasoning trace before final editing, followed by reinforcement learning to further improve reasoning and editing quality. Finally, we introduce SREdit-Bench, a new benchmark designed to comprehensively stress-test models under sophisticated scenes and fine-grained referring expressions. Experiments demonstrate that GVCoT consistently outperforms state-of-the-art models on SREdit-Bench and ImgEdit. We hope our GVCoT will inspire future research toward interpretable and precise image editing.
△ Less
Submitted 16 March, 2026; v1 submitted 2 March, 2026;
originally announced March 2026.
-
TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts
Authors:
Yu Xu,
Hongbin Yan,
Juan Cao,
Yiji Cheng,
Tiankai Hang,
Runze He,
Zijin Yin,
Shiyi Zhang,
Yuxin Zhang,
Jintao Li,
Chunyu Wang,
Qinglin Lu,
Tong-Yee Lee,
Fan Tang
Abstract:
Unified image generation and editing models suffer from severe task interference in dense diffusion transformers architectures, where a shared parameter space must compromise between conflicting objectives (e.g., local editing v.s. subject-driven generation). While the sparse Mixture-of-Experts (MoE) paradigm is a promising solution, its gating networks remain task-agnostic, operating based on loc…
▽ More
Unified image generation and editing models suffer from severe task interference in dense diffusion transformers architectures, where a shared parameter space must compromise between conflicting objectives (e.g., local editing v.s. subject-driven generation). While the sparse Mixture-of-Experts (MoE) paradigm is a promising solution, its gating networks remain task-agnostic, operating based on local features, unaware of global task intent. This task-agnostic nature prevents meaningful specialization and fails to resolve the underlying task interference. In this paper, we propose a novel framework to inject semantic intent into MoE routing. We introduce a Hierarchical Task Semantic Annotation scheme to create structured task descriptors (e.g., scope, type, preservation). We then design Predictive Alignment Regularization to align internal routing decisions with the task's high-level semantics. This regularization evolves the gating network from a task-agnostic executor to a dispatch center. Our model effectively mitigates task interference, outperforming dense baselines in fidelity and quality, and our analysis shows that experts naturally develop clear and semantically correlated specializations.
△ Less
Submitted 25 March, 2026; v1 submitted 12 January, 2026;
originally announced January 2026.
-
Re-Align: Structured Reasoning-guided Alignment for In-Context Image Generation and Editing
Authors:
Runze He,
Yiji Cheng,
Tiankai Hang,
Zhimin Li,
Yu Xu,
Zijin Yin,
Shiyi Zhang,
Wenxun Dai,
Penghui Du,
Ao Ma,
Chunyu Wang,
Qinglin Lu,
Jizhong Han,
Jiao Dai
Abstract:
In-context image generation and editing (ICGE) enables users to specify visual concepts through interleaved image-text prompts, demanding precise understanding and faithful execution of user intent. Although recent unified multimodal models exhibit promising understanding capabilities, these strengths often fail to transfer effectively to image generation. We introduce Re-Align, a unified framewor…
▽ More
In-context image generation and editing (ICGE) enables users to specify visual concepts through interleaved image-text prompts, demanding precise understanding and faithful execution of user intent. Although recent unified multimodal models exhibit promising understanding capabilities, these strengths often fail to transfer effectively to image generation. We introduce Re-Align, a unified framework that bridges the gap between understanding and generation through structured reasoning-guided alignment. At its core lies the In-Context Chain-of-Thought (IC-CoT), a structured reasoning paradigm that decouples semantic guidance and reference association, providing clear textual target and mitigating confusion among reference images. Furthermore, Re-Align introduces an effective RL training scheme that leverages a surrogate reward to measure the alignment between structured reasoning text and the generated image, thereby improving the model's overall performance on ICGE tasks. Extensive experiments verify that Re-Align outperforms competitive methods of comparable model scale and resources on both in-context image generation and editing tasks.
△ Less
Submitted 8 January, 2026;
originally announced January 2026.
-
Discontinuous Strongly Quasiconvex Functions
Authors:
Nguyen Thi Van Hang,
Felipe Lara,
Nguyen Dong Yen
Abstract:
A fundamental open question asking whether all real-valued strongly quasiconvex functions defined on $\mathbb R^n$ are necessarily continuous, akin to their convex counterparts, is answered in detail in this paper. Among other things, we show that such functions can have infinitely many points of discontinuity. The failure of lower semicontinuity together with the lack of upper semicontinuity at i…
▽ More
A fundamental open question asking whether all real-valued strongly quasiconvex functions defined on $\mathbb R^n$ are necessarily continuous, akin to their convex counterparts, is answered in detail in this paper. Among other things, we show that such functions can have infinitely many points of discontinuity. The failure of lower semicontinuity together with the lack of upper semicontinuity at infinitely many points of certain real-valued strongly quasiconvex functions are also shown.
△ Less
Submitted 3 December, 2025;
originally announced December 2025.
-
Gaussian fluctuations for stochastic Volterra equations with small noise
Authors:
N. T. Dung,
N. T. Hang
Abstract:
In this paper, we consider a general class of stochastic Volterra equations with small noise. Our aim is to study the fluctuation of the solution around its deterministic limit. We use the techniques of Malliavin calculus to show that the fluctuation process satisfies central limit theorem and provide an optimal estimate for the rate of convergence. An application to stochastic Volterra equations…
▽ More
In this paper, we consider a general class of stochastic Volterra equations with small noise. Our aim is to study the fluctuation of the solution around its deterministic limit. We use the techniques of Malliavin calculus to show that the fluctuation process satisfies central limit theorem and provide an optimal estimate for the rate of convergence. An application to stochastic Volterra equations with fractional Brownian motion kernel is given to illustrate the theory.
△ Less
Submitted 4 April, 2026; v1 submitted 14 November, 2025;
originally announced November 2025.
-
The vertex covers, Betti numbers and projective dimensions of perfect binary trees
Authors:
Nguyen Thu Hang,
Tran Duc Dung,
Do Van Kien
Abstract:
Let $T$ be a perfect binary tree and $I$ be its edge ideal in the polynomial ring $S$. We determine the vertex cover number, independent number, and establish the recursive formula to compute the number of minimal vertex covers. As a consequence, we compute the depth and projective dimension of $S/I$ and show that the total Betti number of $S/I$ at the highest homological degree always equals one.
Let $T$ be a perfect binary tree and $I$ be its edge ideal in the polynomial ring $S$. We determine the vertex cover number, independent number, and establish the recursive formula to compute the number of minimal vertex covers. As a consequence, we compute the depth and projective dimension of $S/I$ and show that the total Betti number of $S/I$ at the highest homological degree always equals one.
△ Less
Submitted 23 September, 2025;
originally announced September 2025.
-
PromptEnhancer: A Simple Approach to Enhance Text-to-Image Models via Chain-of-Thought Prompt Rewriting
Authors:
Linqing Wang,
Ximing Xing,
Yiji Cheng,
Zhiyuan Zhao,
Donghao Li,
Tiankai Hang,
Jiale Tao,
Qixun Wang,
Ruihuang Li,
Comi Chen,
Xin Li,
Mingrui Wu,
Xinchi Deng,
Shuyang Gu,
Chunyu Wang,
Qinglin Lu
Abstract:
Recent advancements in text-to-image (T2I) diffusion models have demonstrated remarkable capabilities in generating high-fidelity images. However, these models often struggle to faithfully render complex user prompts, particularly in aspects like attribute binding, negation, and compositional relationships. This leads to a significant mismatch between user intent and the generated output. To addre…
▽ More
Recent advancements in text-to-image (T2I) diffusion models have demonstrated remarkable capabilities in generating high-fidelity images. However, these models often struggle to faithfully render complex user prompts, particularly in aspects like attribute binding, negation, and compositional relationships. This leads to a significant mismatch between user intent and the generated output. To address this challenge, we introduce PromptEnhancer, a novel and universal prompt rewriting framework that enhances any pretrained T2I model without requiring modifications to its weights. Unlike prior methods that rely on model-specific fine-tuning or implicit reward signals like image-reward scores, our framework decouples the rewriter from the generator. We achieve this by training a Chain-of-Thought (CoT) rewriter through reinforcement learning, guided by a dedicated reward model we term the AlignEvaluator. The AlignEvaluator is trained to provide explicit and fine-grained feedback based on a systematic taxonomy of 24 key points, which are derived from a comprehensive analysis of common T2I failure modes. By optimizing the CoT rewriter to maximize the reward from our AlignEvaluator, our framework learns to generate prompts that are more precisely interpreted by T2I models. Extensive experiments on the HunyuanImage 2.1 model demonstrate that PromptEnhancer significantly improves image-text alignment across a wide range of semantic and compositional challenges. Furthermore, we introduce a new, high-quality human preference benchmark to facilitate future research in this direction.
△ Less
Submitted 23 September, 2025; v1 submitted 4 September, 2025;
originally announced September 2025.
-
Incorporating Pre-trained Diffusion Models in Solving the Schrödinger Bridge Problem
Authors:
Zhicong Tang,
Tiankai Hang,
Shuyang Gu,
Dong Chen,
Baining Guo
Abstract:
This paper aims to unify Score-based Generative Models (SGMs), also known as Diffusion models, and the Schrödinger Bridge (SB) problem through three reparameterization techniques: Iterative Proportional Mean-Matching (IPMM), Iterative Proportional Terminus-Matching (IPTM), and Iterative Proportional Flow-Matching (IPFM). These techniques significantly accelerate and stabilize the training of SB-ba…
▽ More
This paper aims to unify Score-based Generative Models (SGMs), also known as Diffusion models, and the Schrödinger Bridge (SB) problem through three reparameterization techniques: Iterative Proportional Mean-Matching (IPMM), Iterative Proportional Terminus-Matching (IPTM), and Iterative Proportional Flow-Matching (IPFM). These techniques significantly accelerate and stabilize the training of SB-based models. Furthermore, the paper introduces novel initialization strategies that use pre-trained SGMs to effectively train SB-based models. By using SGMs as initialization, we leverage the advantages of both SB-based models and SGMs, ensuring efficient training of SB-based models and further improving the performance of SGMs. Extensive experiments demonstrate the significant effectiveness and improvements of the proposed methods. We believe this work contributes to and paves the way for future research on generative models.
△ Less
Submitted 25 August, 2025;
originally announced August 2025.
-
Monitoring in the Dark: Privacy-Preserving Runtime Verification of Cyber-Physical Systems
Authors:
Charles Koll,
Preston Tan Hang,
Mike Rosulek,
Houssam Abbas
Abstract:
In distributed Cyber-Physical Systems and Internet-of-Things applications, the nodes of the system send measurements to a monitor that checks whether these measurements satisfy given formal specifications. For instance in Urban Air Mobility, a local traffic authority will be monitoring drone traffic to evaluate its flow and detect emerging problematic patterns. Certain applications require both th…
▽ More
In distributed Cyber-Physical Systems and Internet-of-Things applications, the nodes of the system send measurements to a monitor that checks whether these measurements satisfy given formal specifications. For instance in Urban Air Mobility, a local traffic authority will be monitoring drone traffic to evaluate its flow and detect emerging problematic patterns. Certain applications require both the specification and the measurements to be private -- i.e. known only to their owners. Examples include traffic monitoring, testing of integrated circuit designs, and medical monitoring by wearable or implanted devices. In this paper we propose a protocol that enables privacy-preserving robustness monitoring. By following our protocol, both system (e.g. drone) and monitor (e.g. traffic authority) only learn the robustness of the measured trace w.r.t. the specification. But the system learns nothing about the formula, and the monitor learns nothing about the signal monitored. We do this using garbled circuits, for specifications in Signal Temporal Logic interpreted over timed state sequences. We analyze the runtime and memory overhead of privacy preservation, the size of the circuits, and their practicality for three different usage scenarios: design testing, offline monitoring, and online monitoring of Cyber-Physical Systems.
△ Less
Submitted 21 May, 2025;
originally announced May 2025.
-
Fast Autoregressive Models for Continuous Latent Generation
Authors:
Tiankai Hang,
Jianmin Bao,
Fangyun Wei,
Dong Chen
Abstract:
Autoregressive models have demonstrated remarkable success in sequential data generation, particularly in NLP, but their extension to continuous-domain image generation presents significant challenges. Recent work, the masked autoregressive model (MAR), bypasses quantization by modeling per-token distributions in continuous spaces using a diffusion head but suffers from slow inference due to the h…
▽ More
Autoregressive models have demonstrated remarkable success in sequential data generation, particularly in NLP, but their extension to continuous-domain image generation presents significant challenges. Recent work, the masked autoregressive model (MAR), bypasses quantization by modeling per-token distributions in continuous spaces using a diffusion head but suffers from slow inference due to the high computational cost of the iterative denoising process. To address this, we propose the Fast AutoRegressive model (FAR), a novel framework that replaces MAR's diffusion head with a lightweight shortcut head, enabling efficient few-step sampling while preserving autoregressive principles. Additionally, FAR seamlessly integrates with causal Transformers, extending them from discrete to continuous token generation without requiring architectural modifications. Experiments demonstrate that FAR achieves $2.3\times$ faster inference than MAR while maintaining competitive FID and IS scores. This work establishes the first efficient autoregressive paradigm for high-fidelity continuous-space image generation, bridging the critical gap between quality and scalability in visual autoregressive modeling.
△ Less
Submitted 24 April, 2025;
originally announced April 2025.
-
On the set of associated radicals of powers of monomial ideals
Authors:
Nguyen Thu Hang,
Truong Thi Hien
Abstract:
Let $I$ be a monomial ideal in a polynomial ring. In this paper, we study the asymptotic behavior of the set of associated radical ideals of the (symbolic) powers of $I$. We show that both $\asr(I^s)$ and $\asr(I^{(s)})$ need not stabilize for large value of $s$. In the case $I$ is a square-free monomial ideal, we prove that $\asr(I^{(s)})$ is constant for $s$ large enough. Finally, if $I$ is the…
▽ More
Let $I$ be a monomial ideal in a polynomial ring. In this paper, we study the asymptotic behavior of the set of associated radical ideals of the (symbolic) powers of $I$. We show that both $\asr(I^s)$ and $\asr(I^{(s)})$ need not stabilize for large value of $s$. In the case $I$ is a square-free monomial ideal, we prove that $\asr(I^{(s)})$ is constant for $s$ large enough. Finally, if $I$ is the cover ideal of a balanced hypergraph, then $\asr(I^s)$ monotonically increases in $s$.
△ Less
Submitted 19 December, 2024;
originally announced December 2024.
-
Annihilator of local cohomology modules under localization and completion
Authors:
Nguyen Thi Anh Hang,
Le Thanh Nhan
Abstract:
Let $(R, \frak m)$ be a Noetherian local ring. This paper deals with the annihilator of Artinian local cohomology modules $H^i_{\frak m}(M)$ in the relation with the structure of the base ring $R$, for non negative integers $i$ and finitely generated $R$-modules $M$. Firstly, the catenarity and the unmixedness of local rings are characterized via the compatibility of annihilator of top local cohom…
▽ More
Let $(R, \frak m)$ be a Noetherian local ring. This paper deals with the annihilator of Artinian local cohomology modules $H^i_{\frak m}(M)$ in the relation with the structure of the base ring $R$, for non negative integers $i$ and finitely generated $R$-modules $M$. Firstly, the catenarity and the unmixedness of local rings are characterized via the compatibility of annihilator of top local cohomology modules under localization and completion, respectively. Secondly, some necessary and sufficient conditions for a local ring being a quotient of a Cohen-Macaulay local ring are given in term of the annihilator of all local cohomology modules under localization and completion.
△ Less
Submitted 29 August, 2025; v1 submitted 9 December, 2024;
originally announced December 2024.
-
A fresh look into variational analysis of $\mathcal C^2$-partly smooth functions
Authors:
Nguyen T. V. Hang,
Ebrahim Sarabi
Abstract:
$\mathcal C^2$-partial smoothness of functions has been an important subject of research in optimization, on both theoretical and algorithmic aspects, since it was first introduced by Lewis in 2002. Our work aims at providing a fresh variational analysis viewpoint on the class of $\mathcal C^2$-partly smooth functions. Namely, we explore the relationship between $\mathcal C^2…
▽ More
$\mathcal C^2$-partial smoothness of functions has been an important subject of research in optimization, on both theoretical and algorithmic aspects, since it was first introduced by Lewis in 2002. Our work aims at providing a fresh variational analysis viewpoint on the class of $\mathcal C^2$-partly smooth functions. Namely, we explore the relationship between $\mathcal C^2$-partial smoothness and strict twice epi-differentiability and demonstrate that functions from the latter class are always strictly twice epi-differentiable. On the other hand, we provide two examples to show that the opposite conclusion does not hold in general. As a consequence of our analysis, we calculate the second subderivative of $\mathcal C^2$-partly smooth functions. Applications to stability analysis of related generalized equations involving a general perturbation and to asymptotic analysis of the well-known sample average approximation method for stochastic programs with $\mathcal C^2$-partly smooth regularizers are also given.
△ Less
Submitted 9 March, 2026; v1 submitted 1 November, 2024;
originally announced November 2024.
-
Regularity of Powers and symbolic powers of edge ideals of cubic circulant graphs
Authors:
Nguyen Thu Hang,
My Hanh Pham,
Thanh Vu
Abstract:
We compute the regularity of powers and symbolic powers of edge ideals of all cubic circulant graphs. In particular, we establish Conjecture of Minh for cubic circulant graphs.
We compute the regularity of powers and symbolic powers of edge ideals of all cubic circulant graphs. In particular, we establish Conjecture of Minh for cubic circulant graphs.
△ Less
Submitted 30 September, 2024;
originally announced September 2024.
-
Fisher information bounds and applications to SDEs with small noise
Authors:
Nguyen Tien Dung,
Nguyen Thu Hang
Abstract:
In this paper, we first establish general bounds on the Fisher information distance to the class of normal distributions of Malliavin differentiable random variables. We then study the rate of Fisher information convergence in the central limit theorem for the solution of small noise stochastic differential equations and its additive functionals. We also show that the convergence rate is of optima…
▽ More
In this paper, we first establish general bounds on the Fisher information distance to the class of normal distributions of Malliavin differentiable random variables. We then study the rate of Fisher information convergence in the central limit theorem for the solution of small noise stochastic differential equations and its additive functionals. We also show that the convergence rate is of optimal order.
△ Less
Submitted 19 August, 2024;
originally announced August 2024.
-
Improved Noise Schedule for Diffusion Training
Authors:
Tiankai Hang,
Shuyang Gu,
Xin Geng,
Baining Guo
Abstract:
Diffusion models have emerged as the de facto choice for generating high-quality visual signals across various domains. However, training a single model to predict noise across various levels poses significant challenges, necessitating numerous iterations and incurring significant computational costs. Various approaches, such as loss weighting strategy design and architectural refinements, have be…
▽ More
Diffusion models have emerged as the de facto choice for generating high-quality visual signals across various domains. However, training a single model to predict noise across various levels poses significant challenges, necessitating numerous iterations and incurring significant computational costs. Various approaches, such as loss weighting strategy design and architectural refinements, have been introduced to expedite convergence and improve model performance. In this study, we propose a novel approach to design the noise schedule for enhancing the training of diffusion models. Our key insight is that the importance sampling of the logarithm of the Signal-to-Noise ratio ($\log \text{SNR}$), theoretically equivalent to a modified noise schedule, is particularly beneficial for training efficiency when increasing the sample frequency around $\log \text{SNR}=0$. This strategic sampling allows the model to focus on the critical transition point between signal dominance and noise dominance, potentially leading to more robust and accurate predictions.We empirically demonstrate the superiority of our noise schedule over the standard cosine schedule.Furthermore, we highlight the advantages of our noise schedule design on the ImageNet benchmark, showing that the designed schedule consistently benefits different prediction targets. Our findings contribute to the ongoing efforts to optimize diffusion models, potentially paving the way for more efficient and effective training paradigms in the field of generative AI.
△ Less
Submitted 27 November, 2024; v1 submitted 3 July, 2024;
originally announced July 2024.
-
Existence and asymptotic autonomous robustness of random attractors for three-dimensional stochastic globally modified Navier-Stokes equations on unbounded domains
Authors:
Bui Kim My,
Ho Thi Hang,
Kush Kinra,
Manil T. Mohan,
Pham Tri Nguyen
Abstract:
In this article, we discuss the existence and asymptotically autonomous robustness (AAR) (almost surely) of random attractors for 3D stochastic globally modified Navier-Stokes equations (SGMNSE) on Poincaré domains (which may be bounded or unbounded). Our aim is to investigate the existence and AAR of random attractors for 3D SGMNSE when the time-dependent forcing converges to a time-independent f…
▽ More
In this article, we discuss the existence and asymptotically autonomous robustness (AAR) (almost surely) of random attractors for 3D stochastic globally modified Navier-Stokes equations (SGMNSE) on Poincaré domains (which may be bounded or unbounded). Our aim is to investigate the existence and AAR of random attractors for 3D SGMNSE when the time-dependent forcing converges to a time-independent function under the perturbation of linear multiplicative noise as well as additive noise. The main approach is to provide a way to justify that, on some uniformly tempered universe, the usual pullback asymptotic compactness of the solution operators is uniform across an infinite time-interval $(-\infty,τ]$. The backward uniform ``tail-smallness'' and ``flattening-property'' of the solutions over $(-\infty,τ]$ have been demonstrated to achieve this goal. To the best of our knowledge, this is the first attempt to establish the existence as well as AAR of random attractors for 3D SGMNSE on unbounded domains.
△ Less
Submitted 13 February, 2026; v1 submitted 11 June, 2024;
originally announced June 2024.
-
Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization
Authors:
Zhanhao Liang,
Yuhui Yuan,
Shuyang Gu,
Bohan Chen,
Tiankai Hang,
Mingxi Cheng,
Ji Li,
Liang Zheng
Abstract:
Generating visually appealing images is fundamental to modern text-to-image generation models. A potential solution to better aesthetics is direct preference optimization (DPO), which has been applied to diffusion models to improve general image quality including prompt alignment and aesthetics. Popular DPO methods propagate preference labels from clean image pairs to all the intermediate steps al…
▽ More
Generating visually appealing images is fundamental to modern text-to-image generation models. A potential solution to better aesthetics is direct preference optimization (DPO), which has been applied to diffusion models to improve general image quality including prompt alignment and aesthetics. Popular DPO methods propagate preference labels from clean image pairs to all the intermediate steps along the two generation trajectories. However, preference labels provided in existing datasets are blended with layout and aesthetic opinions, which would disagree with aesthetic preference. Even if aesthetic labels were provided (at substantial cost), it would be hard for the two-trajectory methods to capture nuanced visual differences at different steps. To improve aesthetics economically, this paper uses existing generic preference data and introduces step-by-step preference optimization (SPO) that discards the propagation strategy and allows fine-grained image details to be assessed. Specifically, at each denoising step, we 1) sample a pool of candidates by denoising from a shared noise latent, 2) use a step-aware preference model to find a suitable win-lose pair to supervise the diffusion model, and 3) randomly select one from the pool to initialize the next denoising step. This strategy ensures that diffusion models focus on the subtle, fine-grained visual differences instead of layout aspect. We find that aesthetics can be significantly enhanced by accumulating these improved minor differences. When fine-tuning Stable Diffusion v1.5 and SDXL, SPO yields significant improvements in aesthetics compared with existing DPO methods while not sacrificing image-text alignment compared with vanilla models. Moreover, SPO converges much faster than DPO methods due to the use of more correct preference labels provided by the step-aware preference model.
△ Less
Submitted 25 March, 2025; v1 submitted 6 June, 2024;
originally announced June 2024.
-
Simplified Diffusion Schrödinger Bridge
Authors:
Zhicong Tang,
Tiankai Hang,
Shuyang Gu,
Dong Chen,
Baining Guo
Abstract:
This paper introduces a novel theoretical simplification of the Diffusion Schrödinger Bridge (DSB) that facilitates its unification with Score-based Generative Models (SGMs), addressing the limitations of DSB in complex data generation and enabling faster convergence and enhanced performance. By employing SGMs as an initial solution for DSB, our approach capitalizes on the strengths of both framew…
▽ More
This paper introduces a novel theoretical simplification of the Diffusion Schrödinger Bridge (DSB) that facilitates its unification with Score-based Generative Models (SGMs), addressing the limitations of DSB in complex data generation and enabling faster convergence and enhanced performance. By employing SGMs as an initial solution for DSB, our approach capitalizes on the strengths of both frameworks, ensuring a more efficient training process and improving the performance of SGM. We also propose a reparameterization technique that, despite theoretical approximations, practically improves the network's fitting capabilities. Our extensive experimental evaluations confirm the effectiveness of the simplified DSB, demonstrating its significant improvements. We believe the contributions of this work pave the way for advanced generative modeling.
△ Less
Submitted 28 October, 2024; v1 submitted 21 March, 2024;
originally announced March 2024.
-
Projective dimension and regularity of 3-path ideals of unicyclic graphs
Authors:
Nguyen Thu Hang,
Thanh Vu
Abstract:
We compute the projective dimension and regularity of $3$-path ideals of arbitrary graphs with at most one cycle.
We compute the projective dimension and regularity of $3$-path ideals of arbitrary graphs with at most one cycle.
△ Less
Submitted 25 February, 2024;
originally announced February 2024.
-
Elastic energy driven multivariant selection in martensites via quantum annealing
Authors:
Lara C. P. dos Santos,
Tian Hang,
Roland Sandt,
Martin Finsterbusch,
Yann Le Bouar,
Robert Spatschek
Abstract:
We demonstrate the use of quantum annealing for the selection of multiple martensite variants in a microstructure with long-range coherency stresses and external mechanical load. The general approach is illustrated for martensites with four different variants, based on the minimization of the linear elastic energy. The equilibrium variant distribution is then analysed under application of tensile…
▽ More
We demonstrate the use of quantum annealing for the selection of multiple martensite variants in a microstructure with long-range coherency stresses and external mechanical load. The general approach is illustrated for martensites with four different variants, based on the minimization of the linear elastic energy. The equilibrium variant distribution is then analysed under application of tensile and shear strains and for different values of the considered shear and tetragonal contributions of the different martensite variants. The interface orientations between different domains of variants can be explained using the perspective of the elastic energy anisotropy for regular stripe patterns. For random grain orientations, the response to an external elastic strain is weaker and variants changes can be interpreted based on the rotated eigenstrain tensor.
△ Less
Submitted 31 January, 2024;
originally announced January 2024.
-
CCA: Collaborative Competitive Agents for Image Editing
Authors:
Tiankai Hang,
Shuyang Gu,
Dong Chen,
Xin Geng,
Baining Guo
Abstract:
This paper presents a novel generative model, Collaborative Competitive Agents (CCA), which leverages the capabilities of multiple Large Language Models (LLMs) based agents to execute complex tasks. Drawing inspiration from Generative Adversarial Networks (GANs), the CCA system employs two equal-status generator agents and a discriminator agent. The generators independently process user instructio…
▽ More
This paper presents a novel generative model, Collaborative Competitive Agents (CCA), which leverages the capabilities of multiple Large Language Models (LLMs) based agents to execute complex tasks. Drawing inspiration from Generative Adversarial Networks (GANs), the CCA system employs two equal-status generator agents and a discriminator agent. The generators independently process user instructions and generate results, while the discriminator evaluates the outputs, and provides feedback for the generator agents to further reflect and improve the generation results. Unlike the previous generative model, our system can obtain the intermediate steps of generation. This allows each generator agent to learn from other successful executions due to its transparency, enabling a collaborative competition that enhances the quality and robustness of the system's results. The primary focus of this study is image editing, demonstrating the CCA's ability to handle intricate instructions robustly. The paper's main contributions include the introduction of a multi-agent-based generative model with controllable intermediate steps and iterative optimization, a detailed examination of agent relationships, and comprehensive experiments on image editing. Code is available at \href{https://github.com/TiankaiHang/CCA}{https://github.com/TiankaiHang/CCA}.
△ Less
Submitted 15 February, 2025; v1 submitted 23 January, 2024;
originally announced January 2024.
-
Study change of the performance of airfoil of small wind turbine under low wind speed by CFD simulation
Authors:
Le Quang Sang,
Dinh Van Thin,
Nguyen Huu Duc,
Nguyen Duc Minh,
Doan Hong Quan,
Le Thi Thuy Hang
Abstract:
Renewable energy has received strong attention and investment to replace fossil energy sources and reduce greenhouse gas emissions. Quite good and good wind speed areas have been invested in building large-capacity wind farms for many years. The low wind speed region occupies a very large on the world, which has been interested in the exploitation of wind energy in recent years. In this study, the…
▽ More
Renewable energy has received strong attention and investment to replace fossil energy sources and reduce greenhouse gas emissions. Quite good and good wind speed areas have been invested in building large-capacity wind farms for many years. The low wind speed region occupies a very large on the world, which has been interested in the exploitation of wind energy in recent years. In this study, the original airfoil of S1010 operated at low wind speed was redesigned to increase the aerodynamic efficiency of the airfoil by using XFLR5 software. After, the new VAST-EPU-S1010 airfoil model was adjusted to the maximum thickness and the maximum thickness position. It was simulated in low wind speed conditions of 4-6 m/s by CFD simulation. The lift coefficient, drag coefficient and $C_{L}$/$C_{D}$ coefficient ratio were evaluated under the effect of the angle of attack and the maximum thickness by using the $k-ε$ model. Simulation results show that the VAST-EPU-S1010 airfoil achieved the greatest aerodynamic efficiency at the angle of attack of $3\,^{\circ}$, the maximum thickness of 8\% and the maximum thickness position of 20.32\%. The maximum value of $C_{L}$/$C_{D}$ of the new airfoil at 6 m/s is higher than at the 4 m/s by about 6.25\%.
△ Less
Submitted 20 November, 2023;
originally announced November 2023.
-
Smoothness of Subgradient Mappings and Its Applications in Parametric Optimization
Authors:
Nguyen T. V. Hang,
Ebrahim Sarabi
Abstract:
We demonstrate that the concept of strict proto-differentiability of subgradient mappings can play a similar role as smoothness of the gradient mapping of a function in the study of subgradient mappings of prox-regular functions. We then show that metric regularity and strong metric regularity are equivalent for a class of generalized equations when this condition is satisfied. For a class of comp…
▽ More
We demonstrate that the concept of strict proto-differentiability of subgradient mappings can play a similar role as smoothness of the gradient mapping of a function in the study of subgradient mappings of prox-regular functions. We then show that metric regularity and strong metric regularity are equivalent for a class of generalized equations when this condition is satisfied. For a class of composite functions, called C2-decomposable, we argue that strict proto-differentiability can be characterized via a simple relative interior condition. Leveraging this observation, we present a characterization of the continuous differentiability of the proximal mapping for this class of function via a certain relative interior condition. Applications to the study of strong metric regularity of the KKT system of a class of composite optimization problems are also provided.
△ Less
Submitted 28 October, 2024; v1 submitted 10 November, 2023;
originally announced November 2023.
-
Depth of powers of edge ideals of Cohen-Macaulay trees
Authors:
Nguyen Thu Hang,
Truong Thi Hien,
Thanh Vu
Abstract:
Let $I$ be the edge ideal of a Cohen-Macaulay tree of dimension $d$ over a polynomial ring $S = \mathrm{k}[x_1,\ldots,x_{d},y_1,\ldots,y_d]$. We prove that for all $t \ge 1$,
$$\operatorname{depth} (S/I^t) = \operatorname{max} \{d -t + 1, 1 \}.$$
Let $I$ be the edge ideal of a Cohen-Macaulay tree of dimension $d$ over a polynomial ring $S = \mathrm{k}[x_1,\ldots,x_{d},y_1,\ldots,y_d]$. We prove that for all $t \ge 1$,
$$\operatorname{depth} (S/I^t) = \operatorname{max} \{d -t + 1, 1 \}.$$
△ Less
Submitted 10 September, 2023;
originally announced September 2023.
-
InstructDiffusion: A Generalist Modeling Interface for Vision Tasks
Authors:
Zigang Geng,
Binxin Yang,
Tiankai Hang,
Chen Li,
Shuyang Gu,
Ting Zhang,
Jianmin Bao,
Zheng Zhang,
Han Hu,
Dong Chen,
Baining Guo
Abstract:
We present InstructDiffusion, a unifying and generic framework for aligning computer vision tasks with human instructions. Unlike existing approaches that integrate prior knowledge and pre-define the output space (e.g., categories and coordinates) for each vision task, we cast diverse vision tasks into a human-intuitive image-manipulating process whose output space is a flexible and interactive pi…
▽ More
We present InstructDiffusion, a unifying and generic framework for aligning computer vision tasks with human instructions. Unlike existing approaches that integrate prior knowledge and pre-define the output space (e.g., categories and coordinates) for each vision task, we cast diverse vision tasks into a human-intuitive image-manipulating process whose output space is a flexible and interactive pixel space. Concretely, the model is built upon the diffusion process and is trained to predict pixels according to user instructions, such as encircling the man's left shoulder in red or applying a blue mask to the left car. InstructDiffusion could handle a variety of vision tasks, including understanding tasks (such as segmentation and keypoint detection) and generative tasks (such as editing and enhancement). It even exhibits the ability to handle unseen tasks and outperforms prior methods on novel datasets. This represents a significant step towards a generalist modeling interface for vision tasks, advancing artificial general intelligence in the field of computer vision.
△ Less
Submitted 7 September, 2023;
originally announced September 2023.
-
Investigation of CP-even Higgs bosons decays $H \rightarrow μτ$ within constraints of $l_a \rightarrow l_b γ$ in a 3-3-1 model with inverse seesaw neutrinos
Authors:
H. V. Quyet,
T. T. Hieu,
N. T. Tham,
N. T. T. Hang,
H. T. Hung
Abstract:
In a 3-3-1 model with inverse seesaw neutrinos, we use a simple form of Higgs potential to give four CP-even Higgs bosons ($H \equiv h^0_1,h^0_2,h^0_3,h^0_4$). We investigate $H \rightarrow μτ$ decays in the parameter space regions satisfying the experimental limits of $l_a \rightarrow l_b γ$ with running parameters being the mass of the charged Higgs boson ($m_{H_1^\pm}$) and the mixing matrix of…
▽ More
In a 3-3-1 model with inverse seesaw neutrinos, we use a simple form of Higgs potential to give four CP-even Higgs bosons ($H \equiv h^0_1,h^0_2,h^0_3,h^0_4$). We investigate $H \rightarrow μτ$ decays in the parameter space regions satisfying the experimental limits of $l_a \rightarrow l_b γ$ with running parameters being the mass of the charged Higgs boson ($m_{H_1^\pm}$) and the mixing matrix of the heavy neutrinos ($M_R$). We show that there exist regions of parameter space where all partial widths $Γ(H \rightarrow μτ)$ are less than the current experimental limit ($4.1 \times 10^{-6} GeV$). Analyzing the contributing components to $Γ(H \rightarrow μτ)$, we also compare the mass of the SM-like Higgs boson with the corresponding ones of the other CP-even Higgs bosons in this model.
△ Less
Submitted 9 August, 2023;
originally announced August 2023.
-
Convergence of Augmented Lagrangian Methods for Composite Optimization Problems
Authors:
Nguyen T. V. Hang,
Ebrahim Sarabi
Abstract:
Local convergence analysis of the augmented Lagrangian method (ALM) is established for a large class of composite optimization problems with nonunique Lagrange multipliers under a second-order sufficient condition. We present a new second-order variational property, called the semi-stability of second subderivatives, and demonstrate that it is widely satisfied for numerous classes of functions, im…
▽ More
Local convergence analysis of the augmented Lagrangian method (ALM) is established for a large class of composite optimization problems with nonunique Lagrange multipliers under a second-order sufficient condition. We present a new second-order variational property, called the semi-stability of second subderivatives, and demonstrate that it is widely satisfied for numerous classes of functions, important for applications in constrained and composite optimization problems. Using the latter condition and a certain second-order sufficient condition, we are able to establish Q-linear convergence of the primal-dual sequence for an inexact version of the ALM for composite programs.
△ Less
Submitted 20 October, 2023; v1 submitted 28 July, 2023;
originally announced July 2023.
-
The affine cones over Fano-Mukai fourfold of genus $7$ are flexible
Authors:
Nguyen Thi Anh Hang,
Hoang Le Truong
Abstract:
In this paper, we will show that the affine cones over any smooth Fano-Mukai fourfold of genus $7$ are flexible.
In this paper, we will show that the affine cones over any smooth Fano-Mukai fourfold of genus $7$ are flexible.
△ Less
Submitted 5 April, 2023;
originally announced April 2023.
-
Efficient Diffusion Training via Min-SNR Weighting Strategy
Authors:
Tiankai Hang,
Shuyang Gu,
Chen Li,
Jianmin Bao,
Dong Chen,
Han Hu,
Xin Geng,
Baining Guo
Abstract:
Denoising diffusion models have been a mainstream approach for image generation, however, training these models often suffers from slow convergence. In this paper, we discovered that the slow convergence is partly due to conflicting optimization directions between timesteps. To address this issue, we treat the diffusion training as a multi-task learning problem, and introduce a simple yet effectiv…
▽ More
Denoising diffusion models have been a mainstream approach for image generation, however, training these models often suffers from slow convergence. In this paper, we discovered that the slow convergence is partly due to conflicting optimization directions between timesteps. To address this issue, we treat the diffusion training as a multi-task learning problem, and introduce a simple yet effective approach referred to as Min-SNR-$γ$. This method adapts loss weights of timesteps based on clamped signal-to-noise ratios, which effectively balances the conflicts among timesteps. Our results demonstrate a significant improvement in converging speed, 3.4$\times$ faster than previous weighting strategies. It is also more effective, achieving a new record FID score of 2.06 on the ImageNet $256\times256$ benchmark using smaller architectures than that employed in previous state-of-the-art. The code is available at https://github.com/TiankaiHang/Min-SNR-Diffusion-Training.
△ Less
Submitted 11 March, 2024; v1 submitted 16 March, 2023;
originally announced March 2023.
-
On generalized Gauss maps of minimal surfaces sharing hypersurfaces in a projective variety
Authors:
Si Duc Quang,
Do Thi Thuy Hang
Abstract:
In this article, we study the uniqueness problem for the generalized gauss maps of minimal surfaces (with the same base) immersed in $\mathbb R^{n+1}$ which have the same inverse image of some hypersurfaces in a projective subvariety $V\subset\mathbb P^n(\mathbb C)$. As we know, this is the first time the unicity of generalized gauss maps on minimal surfaces sharing hypersurfaces in a projective v…
▽ More
In this article, we study the uniqueness problem for the generalized gauss maps of minimal surfaces (with the same base) immersed in $\mathbb R^{n+1}$ which have the same inverse image of some hypersurfaces in a projective subvariety $V\subset\mathbb P^n(\mathbb C)$. As we know, this is the first time the unicity of generalized gauss maps on minimal surfaces sharing hypersurfaces in a projective varieties is studied. Our results generalize and improve the previous results in this field.
△ Less
Submitted 21 January, 2024; v1 submitted 13 November, 2022;
originally announced November 2022.
-
Depth stability of cover ideals
Authors:
Mai Phuoc Binh,
Nguyen Thu Hang,
Truong Thi Hien,
Tran Nam Trung
Abstract:
Let R = K[x1,...,xr] be a polynomial ring over a field K. Let G be a graph with vertex set {1,...,r} and let J be the cover ideal of G. We give a sharp bound for the stability index of symbolic depth function sdstab(J). In the case G is bipartite, it yields a sharp bound for the stability index of depth function dstab(J) and this bound is exact if G is a forest.
Let R = K[x1,...,xr] be a polynomial ring over a field K. Let G be a graph with vertex set {1,...,r} and let J be the cover ideal of G. We give a sharp bound for the stability index of symbolic depth function sdstab(J). In the case G is bipartite, it yields a sharp bound for the stability index of depth function dstab(J) and this bound is exact if G is a forest.
△ Less
Submitted 17 October, 2022;
originally announced October 2022.
-
A Chain Rule for Strict Twice Epi-Differentiability and its Applications
Authors:
N. T. V. Hang,
M. E. Sarabi
Abstract:
The presence of second-order smoothness for objective functions of optimization problems can provide valuable information about their stability properties and help us design efficient numerical algorithms for solving these problems. Such second-order information, however, cannot be expected in various constrained and composite optimization problems since we often have to express their objective fu…
▽ More
The presence of second-order smoothness for objective functions of optimization problems can provide valuable information about their stability properties and help us design efficient numerical algorithms for solving these problems. Such second-order information, however, cannot be expected in various constrained and composite optimization problems since we often have to express their objective functions in terms of extended-real-valued functions for which the classical second derivative may not exist. One powerful geometrical tool to use for dealing with such functions is the concept of twice epi-differentiability. In this paper, we are going to study a stronger version of this concept, called strict twice epi-differentiability. We characterize this concept for certain composite functions and use it to establish the equivalence of metric regularity and strong metric regularity for a class of generalized equations at their nondegenerate solutions. Finally, we present a characterization of continuous differentiability of the proximal mapping of our composite functions.
△ Less
Submitted 2 August, 2023; v1 submitted 3 September, 2022;
originally announced September 2022.
-
Language-Guided Face Animation by Recurrent StyleGAN-based Generator
Authors:
Tiankai Hang,
Huan Yang,
Bei Liu,
Jianlong Fu,
Xin Geng,
Baining Guo
Abstract:
Recent works on language-guided image manipulation have shown great power of language in providing rich semantics, especially for face images. However, the other natural information, motions, in language is less explored. In this paper, we leverage the motion information and study a novel task, language-guided face animation, that aims to animate a static face image with the help of languages. To…
▽ More
Recent works on language-guided image manipulation have shown great power of language in providing rich semantics, especially for face images. However, the other natural information, motions, in language is less explored. In this paper, we leverage the motion information and study a novel task, language-guided face animation, that aims to animate a static face image with the help of languages. To better utilize both semantics and motions from languages, we propose a simple yet effective framework. Specifically, we propose a recurrent motion generator to extract a series of semantic and motion information from the language and feed it along with visual information to a pre-trained StyleGAN to generate high-quality frames. To optimize the proposed framework, three carefully designed loss functions are proposed including a regularization loss to keep the face identity, a path length regularization loss to ensure motion smoothness, and a contrastive loss to enable video synthesis with various language guidance in one single model. Extensive experiments with both qualitative and quantitative evaluations on diverse domains (\textit{e.g.,} human face, anime face, and dog face) demonstrate the superiority of our model in generating high-quality and realistic videos from one still image with the guidance of language. Code will be available at https://github.com/TiankaiHang/language-guided-animation.git.
△ Less
Submitted 3 July, 2024; v1 submitted 10 August, 2022;
originally announced August 2022.
-
Role of Subgradients in Variational Analysis of Polyhedral Functions
Authors:
N. T. V. Hang,
W. Jung,
M. E. Sarabi
Abstract:
Understanding the role that subgradients play in various second-order variational analysis constructions can help us uncover new properties of important classes of functions in variational analysis. Focusing mainly on the behavior of the second subderivative and subgradient proto-derivative of polyhedral functions, functions with polyhedral epigraphs, we demonstrate that choosing the underlying su…
▽ More
Understanding the role that subgradients play in various second-order variational analysis constructions can help us uncover new properties of important classes of functions in variational analysis. Focusing mainly on the behavior of the second subderivative and subgradient proto-derivative of polyhedral functions, functions with polyhedral epigraphs, we demonstrate that choosing the underlying subgradient, utilized in the definitions of these concepts, from the relative interior of the subdifferential of polyhedral functions ensures stronger second-order variational properties such as strict twice epi-differentiability and strict subgradient proto-differentiability. This allows us to characterize continuous differentiability of the proximal mapping and twice continuous differentiability of the Moreau envelope of polyhedral functions. We close the paper with proving the equivalence of metric regularity and strong metric regularity of a class of generalized equations at their nondegenerate solutions.
△ Less
Submitted 10 January, 2023; v1 submitted 15 July, 2022;
originally announced July 2022.
-
Advancing High-Resolution Video-Language Representation with Large-Scale Video Transcriptions
Authors:
Hongwei Xue,
Tiankai Hang,
Yanhong Zeng,
Yuchong Sun,
Bei Liu,
Huan Yang,
Jianlong Fu,
Baining Guo
Abstract:
We study joint video and language (VL) pre-training to enable cross-modality learning and benefit plentiful downstream VL tasks. Existing works either extract low-quality video features or learn limited text embedding, while neglecting that high-resolution videos and diversified semantics can significantly improve cross-modality learning. In this paper, we propose a novel High-resolution and Diver…
▽ More
We study joint video and language (VL) pre-training to enable cross-modality learning and benefit plentiful downstream VL tasks. Existing works either extract low-quality video features or learn limited text embedding, while neglecting that high-resolution videos and diversified semantics can significantly improve cross-modality learning. In this paper, we propose a novel High-resolution and Diversified VIdeo-LAnguage pre-training model (HD-VILA) for many visual tasks. In particular, we collect a large dataset with two distinct properties: 1) the first high-resolution dataset including 371.5k hours of 720p videos, and 2) the most diversified dataset covering 15 popular YouTube categories. To enable VL pre-training, we jointly optimize the HD-VILA model by a hybrid Transformer that learns rich spatiotemporal features, and a multimodal Transformer that enforces interactions of the learned video features with diversified texts. Our pre-training model achieves new state-of-the-art results in 10 VL understanding tasks and 2 more novel text-to-visual generation tasks. For example, we outperform SOTA models with relative increases of 40.4% R@1 in zero-shot MSR-VTT text-to-video retrieval task and 55.4% in high-resolution dataset LSMDC. The learned VL embedding is also effective in generating visually pleasing and semantically relevant results in text-to-visual editing and super-resolution tasks.
△ Less
Submitted 8 July, 2022; v1 submitted 19 November, 2021;
originally announced November 2021.
-
Density estimates for the exponential functionals of fractional Brownian motion
Authors:
Nguyen Tien Dung,
Nguyen Thu Hang,
Pham Thi Phuong Thuy
Abstract:
In this note, we investigate the density of the exponential functional of the fractional Brownian motion. Based on the techniques of Malliavin's calculus, we provide a log-normal upper bound for the density.
In this note, we investigate the density of the exponential functional of the fractional Brownian motion. Based on the techniques of Malliavin's calculus, we provide a log-normal upper bound for the density.
△ Less
Submitted 21 September, 2021;
originally announced September 2021.
-
Canonical stretched rings
Authors:
Nguyen Thi Anh Hang,
Do Van Kien,
Hoang Le Truong
Abstract:
In this paper, we introduce the concept of canonical stretched rings, sparse stretched rings and maximum sparse ideals. Then we give characterizations of canonical stretched rings and sparse stretched rings; and a characterization of Gorenstein rings in terms of their maximum sparse ideals. Several explicit examples are provided along the paper to illustrate such rings.
In this paper, we introduce the concept of canonical stretched rings, sparse stretched rings and maximum sparse ideals. Then we give characterizations of canonical stretched rings and sparse stretched rings; and a characterization of Gorenstein rings in terms of their maximum sparse ideals. Several explicit examples are provided along the paper to illustrate such rings.
△ Less
Submitted 25 August, 2021;
originally announced August 2021.
-
An Efficient Shock Advice Algorithm based on K-nearest Neighbors for Automated External Defibrillators
Authors:
Dao Thanh Hai,
Nguyen Minh Tuan,
Nguyen Thi Thu Hang,
Le Hai Chau
Abstract:
Shockable rhythms, namely ventricular fibrillation and ventricular tachycardia, are the main cause of sudden cardiac arrests, which can be detected quickly by the automated external defibrillator (AED) devices. In this paper, a simple but effective algorithm is proposed as the shock advice algorithm applied in AED. The proposed algorithm consists of K-nearest neighbor classifier and an optimal set…
▽ More
Shockable rhythms, namely ventricular fibrillation and ventricular tachycardia, are the main cause of sudden cardiac arrests, which can be detected quickly by the automated external defibrillator (AED) devices. In this paper, a simple but effective algorithm is proposed as the shock advice algorithm applied in AED. The proposed algorithm consists of K-nearest neighbor classifier and an optimal set of 36 features, which are extracted from original ECG and shockable, non-shockable signals using modified variational mode decomposition technique. Cross-validation procedure and sequential forward feature selection are carefully applied to select an optimal set from entire feature space. The performance results show that the MVMD is the key element for SCA detection performance, and the proposed algorithm is simpler while remaining relatively high detection performance compared to previous publications.
△ Less
Submitted 25 September, 2021; v1 submitted 29 June, 2021;
originally announced July 2021.
-
Contribution of heavy neutrinos to decay of standard-model-like Higgs boson $h\rightarrow μτ$ in a 3-3-1 model with additional gauge singlets
Authors:
H. T. Hung,
N. T. Tham,
T. T. Hieu,
N. T. T. Hang
Abstract:
In the framework of the improved version of the 3-3-1 models with right-handed neutrinos, which is added to the Majorana neutrinos as new gauge singlets, the recent experimental neutrino oscillation data is completely explained through the inverse seesaw mechanism. We show that the major contributions to $Br(μ\rightarrow eγ)$ are derived from corrections at 1-loop order of heavy neutrinos and boso…
▽ More
In the framework of the improved version of the 3-3-1 models with right-handed neutrinos, which is added to the Majorana neutrinos as new gauge singlets, the recent experimental neutrino oscillation data is completely explained through the inverse seesaw mechanism. We show that the major contributions to $Br(μ\rightarrow eγ)$ are derived from corrections at 1-loop order of heavy neutrinos and bosons. But, these contributions are sometimes mutually destructive, creating regions of parametric spaces where the experimental limits of $Br(μ\rightarrow eγ)$ are satisfied. In these regions, we find that $Br(τ\rightarrow μγ)$ can achieve values of $10 ^{- 10}$ and $Br(τ\rightarrow eγ)$ may even reach values of $10 ^{- 9}$ very close to the upper bound of the current experimental limits. Those are ideal areas to study lepton-flavor-violating decays of the standard- model- like Higgs boson ($h_1^0$). We also pointed out that the contributions of heavy neutrinos play an important role to change $Br(h_1^0\rightarrow μτ)$, this is presented through different forms of mass mixing matrices ($M_R$) of heavy neutrinos. When $M_R \sim diag(1,1,1)$, $Br(h_1^0\rightarrow μτ)$ can get a greater value than the cases $M_R \sim diag(1,2,3)$ and $M_R \sim diag(3,2,1)$ and the largest that $Br(h_1^0\rightarrow μτ)$ can reach is very close $10 ^{-3}$.
△ Less
Submitted 13 August, 2021; v1 submitted 29 March, 2021;
originally announced March 2021.