-
Difficulty-Calibrated Interpolation Paths for Conditional Flow Matching
Authors:
Airin Akter Tania,
Md Raihan Khan
Abstract:
Conditional Flow Matching trains generative models by regressing a network onto the velocity of a prescribed noise-to-data interpolation path. The interpolation schedule that shapes this path is known to affect convergence and sample quality, yet it is invariably fixed in advance, independent of both the data and the model. We show that the regression difficulty of Conditional Flow Matching varies…
▽ More
Conditional Flow Matching trains generative models by regressing a network onto the velocity of a prescribed noise-to-data interpolation path. The interpolation schedule that shapes this path is known to affect convergence and sample quality, yet it is invariably fixed in advance, independent of both the data and the model. We show that the regression difficulty of Conditional Flow Matching varies systematically along the path, and we propose Difficulty-Calibrated Flow Matching, which derives the schedule from the model itself: a short pilot run with the linear path records the per-time loss, and the schedule is set to the quantile function of this difficulty profile, so the trajectory lingers where the velocity is hardest to learn. The method has a single hyperparameter, leaves the training objective and its gradient equivalence intact, composes with classifier-free guidance, and adds about two percent training overhead. In controlled experiments on CIFAR-10, MNIST, and Fashion-MNIST with an identical compact U-Net, the calibrated path attains the best FID on CIFAR-10 at full sampling budget and clearly outperforms all fixed schedules in the large-batch, few-update regime, precisely the setting where compute is scarcest.
△ Less
Submitted 21 August, 2026;
originally announced August 2026.
-
New Limits on $n \rightarrow n'$ Transformation from HFIR Cold Neutron Beam
Authors:
James M. Rogers,
Leah J. Broussard,
Christopher B. Crawford,
Lisa DeBeer-Schmitt,
Matthew J. Frost,
Francisco M. Gonzalez,
Carolyn O. Haviland,
Lawrence Heilbronn,
Erik B. Iverson,
Yuri Kamyshkov,
Mubasshir Khan,
Andrew Mullins,
David Milstead,
Linus B. Persson,
Cary Rock,
Valentina Santoro,
Alexander Saunders,
Shaun Vavra,
Nathan D. Whittington
Abstract:
Hypothetical neutron $n$ to sterile neutron $n'$ transformations would violate baryon number $\mathcal{B}$ and point to the nature of Dark Matter. We performed a new search for $n \rightarrow n'$ using an intense cold neutron beam from the High Flux Isotope Reactor at Oak Ridge National Laboratory. We used a theoretical model that describes the transformation $n \rightarrow n'$ with two parameters…
▽ More
Hypothetical neutron $n$ to sterile neutron $n'$ transformations would violate baryon number $\mathcal{B}$ and point to the nature of Dark Matter. We performed a new search for $n \rightarrow n'$ using an intense cold neutron beam from the High Flux Isotope Reactor at Oak Ridge National Laboratory. We used a theoretical model that describes the transformation $n \rightarrow n'$ with two parameters: a small mass difference $Δ{m}$ between the interaction states $n$ and $n'$ and a mixing vacuum angle $θ_0$. A thin absorbing cadmium wafer was used in the center of the superconducting 6.6 T magnet which provided a large gradient for the non-adiabatic $n \rightarrow n'$ transition. No signal was observed above background in the $^{3}\text{He}$ neutron detector 20 meters downstream of the magnet. This result gives an order of magnitude improvement in the lower limit for the probability $2θ_0^2$ of the $n \rightarrow n'$ transformation in vacuum in the range of $Δ{m}$ between $0.1$ neV and $1000$ neV.
△ Less
Submitted 21 August, 2026;
originally announced August 2026.
-
Thermodynamically Consistent Merging of Multidimensional QCD Equations of State
Authors:
Prachi Garella,
Yumu Yang,
Musa R. Khan,
Tulio E. Restrepo,
Joaquin Grefa,
Johannes Jahan,
Mauricio Hippert,
Jorge Noronha,
Claudia Ratti,
Romulo Rougemont
Abstract:
We present a thermodynamically consistent framework for merging complementary models into a multidimensional QCD equation of state. An internal mixing variable is determined by minimizing a single grand potential at fixed temperature and baryon chemical potential, ensuring thermodynamic consistency and stability. Interactions between the components allow for a crossover, a critical endpoint, and a…
▽ More
We present a thermodynamically consistent framework for merging complementary models into a multidimensional QCD equation of state. An internal mixing variable is determined by minimizing a single grand potential at fixed temperature and baryon chemical potential, ensuring thermodynamic consistency and stability. Interactions between the components allow for a crossover, a critical endpoint, and a first-order transition. As a proof of principle, we merge a quantum van der Waals hadron-resonance-gas model with a holographic Einstein--Maxwell--Dilaton model. The resulting equation of state reproduces the appropriate description in each regime, agrees well with available lattice-QCD results, and is suitable for heavy-ion phenomenology over a broad range of temperature and baryon chemical potential.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
AutoLumNet: Monotone Optimal Transport for Single-Shot Exposure Correction
Authors:
Airin Akter Tania,
Md Raihan Khan,
Mohiuddin Ahmad
Abstract:
Single-shot exposure correction aims to map an arbitrarily degraded image---whether under-exposed, over-exposed, or a spatial mixture of both---to a well-exposed output from a single capture. We present AutoLumNet, a framework that decomposes this task into a global monotone tone curve and a bounded local residual, making the global component the locus of formal guarantees. The tone curve is param…
▽ More
Single-shot exposure correction aims to map an arbitrarily degraded image---whether under-exposed, over-exposed, or a spatial mixture of both---to a well-exposed output from a single capture. We present AutoLumNet, a framework that decomposes this task into a global monotone tone curve and a bounded local residual, making the global component the locus of formal guarantees. The tone curve is parameterized as the normalized cumulative integral of a strictly positive density, ensuring strict monotonicity by construction rather than by penalty. We prove that this parameterization (i)~preserves the pairwise luminance ordering of all pixels and all spatial extrema unconditionally, and (ii)~is dense in the space of valid tone corrections, containing the one-dimensional optimal-transport map from the input to any target luminance distribution. A differentiable sorted-sample Wasserstein-2 objective drives the learned curve toward the OT optimum during training. Spatially varying effects that the global map provably cannot address---local shading, chrominance shifts, and clipped-region restoration---are handled by a bounded residual decoder with dual-branch convex fusion, for which we provide an explicit sufficient condition for local order preservation. Experiments on five benchmarks (MSEC, SICE, LCDP, LOL-v1, LOL-v2-real) show that AutoLumNet achieves state-of-the-art PSNR and SSIM across both under- and over-exposure regimes at 11.2\,ms per frame, and generalizes zero-shot to pure low-light benchmarks without retraining. To our knowledge, AutoLumNet is the first exposure-correction method to unite structural monotonicity, optimal-transport optimality, and bounded local adaptivity within a single trainable architecture. Code is available at https://github.com/kraihan/Autolumnet.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
COBALT: Column-swapping Optimized Bit-serial Accelerator for LSTM Tasks
Authors:
Mohd Tasleem Khan
Abstract:
Long Short-Term Memory (LSTM) networks continue to be widely deployed for real-time sequential tasks on edge devices, yet their computational demands challenge deployment on resource-constrained hardware. This work introduces COBALT, a bit-serial compressed LSTM accelerator built on a matrix (input)-vector (weight) reformulation of the standard circulant matrix-vector multiplication (MVM) with off…
▽ More
Long Short-Term Memory (LSTM) networks continue to be widely deployed for real-time sequential tasks on edge devices, yet their computational demands challenge deployment on resource-constrained hardware. This work introduces COBALT, a bit-serial compressed LSTM accelerator built on a matrix (input)-vector (weight) reformulation of the standard circulant matrix-vector multiplication (MVM) with offset-binary coding. A novel column-swapping scheme operates on partial products at the bit level when generated in pairs, systematically exposing redundancy across output rows to reduce the number of PP generators and selectors. This redundancy is further exploited using a lightweight correction unit that derives a row's output directly from its paired row. Additionally, for block-circulant MVMs, relocating the shift-accumulate and correction units of each sub-MVM further reduces resource usage. The compressed network achieves weight compression of up to 93.6% while maintaining accuracy on the TIMIT and LibriSpeech-100h benchmarks. On a field-programmable gate array, COBALT achieves superior overall efficiency relative to state-of-the-art LSTM accelerators.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Orthogonal Polynomial Approximation for Matrix Log Normalization in Global Covariance Pooling
Authors:
Md Rifat Ur Rahman,
Md Raihan Khan,
Md Sakib Hossain Shovon,
Pietro Liò,
Mohammad Ali Moni
Abstract:
Global Covariance Pooling (GCP) improves deep networks by capturing second-order feature statistics, and is especially effective for fine-grained recognition. Because covariance matrices live on the Symmetric Positive Definite (SPD) manifold, a normalization step is required before the Euclidean classifier. The faithful choice is the matrix logarithm (MLN-COV), which maps the SPD manifold to its t…
▽ More
Global Covariance Pooling (GCP) improves deep networks by capturing second-order feature statistics, and is especially effective for fine-grained recognition. Because covariance matrices live on the Symmetric Positive Definite (SPD) manifold, a normalization step is required before the Euclidean classifier. The faithful choice is the matrix logarithm (MLN-COV), which maps the SPD manifold to its tangent space; in practice it was abandoned in favour of the matrix square root because its eigendecomposition-based gradient is numerically unstable. We show that this instability is an artifact of computing the logarithm spectrally, not of the logarithm itself. Approximating the logarithm with finite polynomials in the covariance matrix removes the eigendecomposition from both passes: every operation becomes a General Matrix Multiplication (GEMM), the gradient stays bounded on the spectral support of the pre-normalized covariance, and the unstable 1/(lambda_i-lambda_j) term never appears. The key ingredient is a mean-eigenvalue pre-normalization that centres the spectrum near 1, away from the singularity of log, with a scalar post-compensation that returns the singular part of log(A) in closed form. Our recommended normalizer is a degree-8 Chebyshev expansion evaluated by a three-term matrix recurrence, with a matching reverse recurrence for the backward pass; Legendre, Laguerre, Taylor and Pade expansions are studied as controls that isolate the roles of the basis and of the target function. On three fine-grained benchmarks and ImageNet-1k the decomposition-free logarithm is both faster and more accurate than the spectral logarithm and than the square-root approximations it replaces, and at matched basis and degree the log target beats the square-root target, confirming that the gain comes from the faithful Riemannian map rather than from a better polynomial family.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
CauSec: Unboxing the Causal Drivers of Static Vulnerability Analysis Performance
Authors:
Md Akram Khan,
Daniel Rodriguez-Cardenas,
Alejandro Velasco Dimate,
Denys Poshyvanyk,
Adwait Nadkarni
Abstract:
Static Application Security Testing (SAST) tools are widely used in both industry and academia. Such tools often make design choices that sacrifice detection to achieve higher performance, i.e., increased precision, decreased runtime, or increased scalability. These design choices rely on certain assumptions regarding the target code or the analysis technique itself. Hence, the assumptions directl…
▽ More
Static Application Security Testing (SAST) tools are widely used in both industry and academia. Such tools often make design choices that sacrifice detection to achieve higher performance, i.e., increased precision, decreased runtime, or increased scalability. These design choices rely on certain assumptions regarding the target code or the analysis technique itself. Hence, the assumptions directly impact the detection outcome through the design choices they influence. This motivates a key question: do the sacrifices in the detection capabilities actually help tools achieve the expected performance gains? That is, are the underlying assumptions valid?
This paper seeks to address this question by relying on a key observation that the assumptions made by these tools are generally of a causal nature. We propose CAUSEC, a causal analysis framework that makes SAST assumptions testable and explains why the performance changes given certain assumptions, beyond simple correlations. CAUSEC formalizes the assumptions of the SAST tool into the abstraction of a security assumption and combines assumption-driven causal modeling with effect estimation and validation to test its validity and investigate the factors affecting it. To understand what security assumptions generally entail, we perform a systematic literature review of SASTs that detect crypto-API misuse, leading to the discovery and qualitative analysis of 57 assumptions. We then demonstrate the utility and robustness of CAUSEC by testing a popular assumption in four highly relevant tools, using a manually labeled ground truth dataset consisting of 57,038 alerts. Our analysis leads to several key findings that represent insights regarding assumptions and causal effects, which we distill into 3 takeaways for future work.
△ Less
Submitted 19 August, 2026; v1 submitted 19 August, 2026;
originally announced August 2026.
-
Transforming Heart Disease Prediction with Advanced Machine Learning Techniques
Authors:
Sami Ullah,
Muhammad Mohsin Khan
Abstract:
Heart disease remains the leading cause of mortality globally, necessitating early and accurate detection to improve patient outcomes. This research focuses on the predictive analysis of heart disease using machine learning (ML) techniques, comparing the performance of multiple classifiers to identify the most accurate and least error-prone method. Two datasets from UCI and Kaggle repositories wer…
▽ More
Heart disease remains the leading cause of mortality globally, necessitating early and accurate detection to improve patient outcomes. This research focuses on the predictive analysis of heart disease using machine learning (ML) techniques, comparing the performance of multiple classifiers to identify the most accurate and least error-prone method. Two datasets from UCI and Kaggle repositories were utilized, each containing 14 attributes related to heart health indicators. Techniques including J48, Naive Bayes, Logistic Regression, Simple Cart, Bagging, Decision Stump, AdaBoost, Artificial Neural Networks, and Support Vector Machine (SVM) were applied. Evaluation metrics such as Mean Absolute Error (MAE), Relative Absolute Error (RAE), accuracy, precision, recall, and F-measure were used for performance comparison. Results revealed that SVM achieved the highest performance on the UCI dataset, while Simple Cart performed best on the Kaggle dataset, offering the highest accuracy and lowest error rates. The research work concludes that ML models, when properly tuned and validated, can significantly assist in the early diagnosis of heart disease, offering critical support for clinical decision-making. Future work may involve hybrid approaches and the use of more recent datasets to further improve prediction accuracy.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Compiling WebAssembly Concolic Execution with Staging, Continuations, and Snapshots (Extended Version)
Authors:
Dinghong Zhong,
Alexander Bai,
Mikail Khan,
Guannan Wei
Abstract:
Concolic execution is a variant of symbolic execution that runs a program simultaneously with concrete and symbolic inputs. It records the symbolic constraints encountered along a concrete execution path, then solves those constraints to generate inputs that explore new paths. Existing concolic engines generally follow one of two implementation strategies: Interpreter-based systems are comparative…
▽ More
Concolic execution is a variant of symbolic execution that runs a program simultaneously with concrete and symbolic inputs. It records the symbolic constraints encountered along a concrete execution path, then solves those constraints to generate inputs that explore new paths. Existing concolic engines generally follow one of two implementation strategies: Interpreter-based systems are comparatively simple to build but incur substantial interpretation overhead, while instrumentation-based systems avoid this overhead but typically re-execute the program from the beginning for each new input.
In this paper, we develop a new approach that achieves the best of both worlds. Starting from the concrete semantics of the target language, we first develop a definitional concolic interpreter and stage it to compile away interpretation overhead while retaining the simplicity of an interpretation-based implementation. By expressing the staged interpreter in continuation-passing style, we can capture execution snapshots at branch points and resume from them when exploring alternative paths, avoiding repeated execution from the program entry. Because snapshot-reuse can itself incur overhead, we further develop a heuristic that favors snapshot-reuse only when it is expected to be beneficial. We instantiate this approach for WebAssembly and implement it in a new concolic-execution compiler GenWasym. Across 184 benchmarks, GenWasym with staging alone achieves a $29.4\times$ average speedup over the interpreter-based WASP; heuristic snapshot-reuse further increases the speedup to $44.9\times$.
△ Less
Submitted 20 August, 2026; v1 submitted 18 August, 2026;
originally announced August 2026.
-
Borel-Moore Homology of 2-Particle Unordered Configuration Spaces of Graphs via Discrete Morse Theory
Authors:
Munawar Khan
Abstract:
We define a discrete Morse function on the open regular cell complex $C_2(Γ)$, arising from the $2$-particle unordered configuration space of a graph $Γ$. We compute the homology of the associated Morse complex. Using an isomorphism established by Knudson and Scoville between the Borel-Moore homology of $C_2(Γ)$ and the homology of the Morse complex, we conclude that the Borel-Moore homology group…
▽ More
We define a discrete Morse function on the open regular cell complex $C_2(Γ)$, arising from the $2$-particle unordered configuration space of a graph $Γ$. We compute the homology of the associated Morse complex. Using an isomorphism established by Knudson and Scoville between the Borel-Moore homology of $C_2(Γ)$ and the homology of the Morse complex, we conclude that the Borel-Moore homology groups of $C_2(Γ)$ are completely determined by the first Betti number of $Γ$.
△ Less
Submitted 7 August, 2026;
originally announced August 2026.
-
When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice
Authors:
Muhammad Salar Khan,
Hamza Umer,
Hasan Mahmud,
Sandra Rothenberg
Abstract:
Large language models (LLMs) are increasingly integrated into financial advisory systems, yet their role in reproducing religious bias remains underexamined. This study provides systematic mixed-methods evidence of such bias across three LLMs (ChatGPT, Gemini, and Grok) using 432 simulated advisor-client interactions spanning 16 religious identity pairings (Christian, Muslim, Hindu, and non-religi…
▽ More
Large language models (LLMs) are increasingly integrated into financial advisory systems, yet their role in reproducing religious bias remains underexamined. This study provides systematic mixed-methods evidence of such bias across three LLMs (ChatGPT, Gemini, and Grok) using 432 simulated advisor-client interactions spanning 16 religious identity pairings (Christian, Muslim, Hindu, and non-religious) and three core household financial decisions: stock investment, house purchase, and life insurance. Combining regression and reflexive thematic analyses, we identify structural biases across models and decision contexts and the discursive mechanisms through which they are linguistically enacted. Unbiased advice appeared in only 12-18% of cases. Gemini consistently produced more bias than Grok, while ChatGPT's outputs were statistically comparable to Grok's. Religiously symmetric advisor-client pairings almost always triggered explicit religious framing, and non-religious clients often received advisor-centered religious appeals. Qualitative findings show that bias is linguistically manifested through religious anchoring, uneven cultural signaling, and tone modulation, varying by model and financial scenario. Stock investment prompts produced more financially technical responses, whereas life insurance advice triggered stronger religious language. The study develops a dual-dimensional framework linking structural bias rooted in model training and design with discursive bias expressed through language, advancing understanding of algorithmic bias in LLM-generated financial advice. It also shows that such advice adapts linguistically to identity cues, revealing a managerial dilemma between personalization and neutrality. Finally, it highlights implications for businesses, financial institutions, and regulators seeking to ensure neutrality, cultural sensitivity, and trust in AI-mediated advice.
△ Less
Submitted 11 July, 2026;
originally announced August 2026.
-
On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos
Authors:
Darakshan Rashid,
Raza Imam,
Ufaq Khan,
Muhammad Bilal,
Shazad Ashraf,
Dwarikanath Mahapatra,
Mohammad Yaqub,
Muhammad Haris Khan,
Imran Razzak,
Brejesh Lall,
Lena Maier-Hein,
Yutong Xie
Abstract:
Temporal vision-language models (TVLMs) offer a reusable, prompt-based interface for surgical video understanding, yet, their robustness under clinically realistic acquisition artifacts in endoscopy remains insufficiently characterized. In practice, degradations such as defocus, haze, motion blur, noise, cautery smoke, and packet loss introduce structured distribution shifts which may compromise v…
▽ More
Temporal vision-language models (TVLMs) offer a reusable, prompt-based interface for surgical video understanding, yet, their robustness under clinically realistic acquisition artifacts in endoscopy remains insufficiently characterized. In practice, degradations such as defocus, haze, motion blur, noise, cautery smoke, and packet loss introduce structured distribution shifts which may compromise video-text alignment. We study the robustness of temporal VLMs under such shifts caused by corruptions in clip frames. We introduce Endo-C6, a compact corruption benchmark of six endoscopy-realistic perturbations evaluated at a fixed high severity, and apply it to public Gastrointestinal (GI) endoscopy and laparoscopic cholecystectomy videos. Under a standardized prompt protocol, we benchmark 3 recent surgical TVLM baselines and analyze robustness in both mean and worst-case settings, spanning 294 dataset-level evaluations. Finally, we present RobustEndoCLIP, obtained by few-shot parameter-efficient tuning with VeRA, outperforming existing TVLM baselines. Our findings show that off-the-shelf TVLMs can exhibit severe worst-case collapse under endoscopy-specific corruptions, whereas lightweight few-shot adaptation can substantially improve corrupted performance and robustness without changing the prompt-based interface. We expect Endo-C6 to support standardized robustness reporting and promote more reliable clinical vision-language systems.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images
Authors:
Jennifer D'Souza,
Fahad Ahmed,
Cecilia Andrea Bustamante Andrade,
Lina Frolova,
Poorani Gnanasambandan,
Dilshad Hussain,
Muhammad Uzair Khan,
Nkembeng Kevin Nkengfoa,
Paul Praveen J.,
Fabio Priante,
Sjoerd Franciscus van der Werf,
Thomas Frederik Jan van Roeden
Abstract:
Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching Scientific Figures provide 1,951 figures from 205 publications, expert-annotated for classification, data table e…
▽ More
Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching Scientific Figures provide 1,951 figures from 205 publications, expert-annotated for classification, data table extraction, summarization, and visual question answering. In these companion proceedings, we present a forward-looking perspective on how the benchmark can guide future scientific-image challenges. We examine how its tasks probe capabilities from visual and quantitative reading to domain-grounded reasoning and evidential justification, and how Bloom-informed question design can support deeper scientific understanding. We propose "scientific conceptual understanding from images" as a long-term benchmark objective, with future directions including broader domains and figure types, contextual and cross-document synthesis, hypothesis evaluation, provenance, uncertainty, counterfactual grounding, and open-ended multimodal research. This perspective connects the ICDAR 2026 challenge to a broader agenda for machine-actionable scientific visual knowledge and verifiable multimodal scientific AI.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning
Authors:
Xulin Fan,
Jialu Li,
Mohammad Nur Hossain Khan,
Kexin Hu,
Bashima Islam,
Mark Hasegawa-Johnson,
Nancy L. McElwain
Abstract:
Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low signal-to-noise ratio, and cross-family domain shifts. We present a family-conditioned, multi-tier audio tagger that combines a LoRA-finetuned Whisper encoder with a lightweight, targe…
▽ More
Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low signal-to-noise ratio, and cross-family domain shifts. We present a family-conditioned, multi-tier audio tagger that combines a LoRA-finetuned Whisper encoder with a lightweight, target-speaker-aware Transformer for long-context inference and framewise prediction across tiers. To improve temporal coherence, we incorporate a simple sequence-level smoothing loss, and to enhance robustness across households, we introduce a factorized speaker-token design with a shared tier token and a learned family-specific offset, reducing family bias and promoting generalizable representations. Together, these choices enable efficient and effective infant-centered audio tagging of daylong audio recordings in home environments.
△ Less
Submitted 11 August, 2026;
originally announced August 2026.
-
Improving TensorSketch Using Complex Random Variables
Authors:
Amit Sharma,
Mohammad Azhar Khan,
Rameshwar Pratap,
Keegan Kang
Abstract:
\texttt{TensorSketch} by~\cite{pham2013fast,kar2012random} provides efficient sketching algorithms for high-dimensional polynomial kernels $\vec{x}^{\otimes p} \in \R^{d^p}$. \cite{kar2012random} uses dense Johnson-Lindenstrauss (JL)-type projections with computational cost $O(pDd)$, where $D$ denotes the sketch dimension, whereas~\cite{pham2013fast} extends the sparse \texttt{CountSketch}~\citep{…
▽ More
\texttt{TensorSketch} by~\cite{pham2013fast,kar2012random} provides efficient sketching algorithms for high-dimensional polynomial kernels $\vec{x}^{\otimes p} \in \R^{d^p}$. \cite{kar2012random} uses dense Johnson-Lindenstrauss (JL)-type projections with computational cost $O(pDd)$, where $D$ denotes the sketch dimension, whereas~\cite{pham2013fast} extends the sparse \texttt{CountSketch}~\citep{count_sketch} algorithm, yielding a faster algorithm for high-dimensional sparse inputs with running time $O\big(p(\nnz{\vec{x}} + D \log D)\big)$. However, the variance of both estimators grows exponentially with the polynomial degree $p$, scaling as $3^{p}/D$. Recent work by~\cite{pmlr-v206-wacker23a} showed that using complex-valued distribution reduces this dependence to $2^{p}/D$ for the approach of~\cite{kar2012random}. However, their method relies on dense JL-type projections with computational cost $O(pDd)$ and does not extend to the algorithm of~\cite{pham2013fast}.
In this work, we introduce a simple variant of \texttt{TensorSketch}~\citep{pham2013fast} that achieves the same variance bound as~\cite{pmlr-v206-wacker23a}, while retaining its advantage of the input-sparsity running time. We validate our results with supporting experiments on synthetic and real-world datasets.
△ Less
Submitted 11 August, 2026;
originally announced August 2026.
-
Autonomous Lindblad Realizability of Nonunitary Linear Dynamics with a Carleman Lattice Boltzmann Application
Authors:
Muhammad Idrees Khan
Abstract:
Carleman lifting converts nonlinear polynomial dynamics into finite linear systems, but the resulting truncations are generally nonunitary and need not correspond to physical quantum evolution. We prove that a finite linear endpoint admits an autonomous Gorini--Kossakowski--Sudarshan--Lindblad (GKSL) realization on vacuum coherences if and only if it is invertible and power bounded. The constructi…
▽ More
Carleman lifting converts nonlinear polynomial dynamics into finite linear systems, but the resulting truncations are generally nonunitary and need not correspond to physical quantum evolution. We prove that a finite linear endpoint admits an autonomous Gorini--Kossakowski--Sudarshan--Lindblad (GKSL) realization on vacuum coherences if and only if it is invertible and power bounded. The construction is explicit and realizes the nonunitary map directly as open-system dynamics, with no endpoint postselection and with one encoding and one decoding over repeated timesteps. We apply the result to the complete D2Q9 multiple-relaxation-time lattice Boltzmann (LB) timestep by compiling collision and periodic streaming into a single Carleman endpoint. The resulting GKSL evolution reproduces the classical Carleman trajectory over multiple timesteps, while the remaining discrepancy from nonlinear LB dynamics is the expected Carleman truncation error. The result establishes a general criterion for autonomous open-quantum realization of finite nonunitary dynamics, with Carleman--LB dynamics as a concrete example.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
Eikonal Regularisation in Physics-Informed Neural Networks for Three-Dimensional Level-Set Advection: Transferability of Two-Dimensional Design Principles
Authors:
Muhammad Akbar Khan
Abstract:
Physics-informed neural networks applied to the level-set formulation of interface advection commonly augment the residual and initial-condition losses with an eikonal regulariser, penalising the deviation of $\|\nablaφ\|$ from unity. A previous two-dimensional study identified this weight as the dominant hyperparameter and found its optimum shifts by four orders of magnitude between rigid-body an…
▽ More
Physics-informed neural networks applied to the level-set formulation of interface advection commonly augment the residual and initial-condition losses with an eikonal regulariser, penalising the deviation of $\|\nablaφ\|$ from unity. A previous two-dimensional study identified this weight as the dominant hyperparameter and found its optimum shifts by four orders of magnitude between rigid-body and deforming flows, but left open whether these principles transfer to three dimensions and whether single-seed results survive run-to-run variability. We answer both by repeating the weight selection across four 3D benchmarks (translating sphere, rotating sphere, slotted sphere, reversed vortex), sweeping six weights with three seeds at full training budget under a pre-registered selection rule. The ordering transfers: the selected weight tracks how far the exact solution departs from the signed-distance property, spanning four decades from $10^{-1}$ where it holds exactly to $10^{-5}$ where the interface is stretched. Values transfer only benchmark by benchmark; two of four carry over unchanged and two do not, so inheritance must be verified. The multi-seed protocol reveals that at small weights the seed-to-seed standard deviation equals the error itself, and the regulariser reduces it by more than an order of magnitude, buying reproducibility as well as accuracy. We benchmark against a fifth-order WENO solver on identical grids and error measures; the classical scheme is more accurate on all four problems, by two orders of magnitude on smooth rigid advection, with a margin that narrows with geometric difficulty and is smaller in volume conservation than in the field norm. Finally, we show that the relative $L_2$ error cannot certify the preservation of thin features, and report a feature-restricted measure that can.
△ Less
Submitted 8 August, 2026;
originally announced August 2026.
-
Active Brownian motion in a single-relaxation viscoelastic fluid
Authors:
Sanatan Halder,
Manas Khan
Abstract:
Active Brownian particles (ABPs) in viscoelastic (VE) media exhibit fascinating dynamical phenomena set by self-propulsion, thermal fluctuations, and fluid viscoelasticity. We extend our model, in which the Brownian dynamics within a slowly diffusing harmonic well emulates that in a single-relaxation VE fluid, to study active Brownian motion in such media. Consequently, the resultant dynamics is g…
▽ More
Active Brownian particles (ABPs) in viscoelastic (VE) media exhibit fascinating dynamical phenomena set by self-propulsion, thermal fluctuations, and fluid viscoelasticity. We extend our model, in which the Brownian dynamics within a slowly diffusing harmonic well emulates that in a single-relaxation VE fluid, to study active Brownian motion in such media. Consequently, the resultant dynamics is governed by the interplay of the characteristic timescales of the systems: the crossover and equilibration times of the VE fluid, $τ_k$ and $λ$, respectively, and the persistence time of the ABP, $τ_{\mathrm{R}}$. Following analytical predictions and simulations, we study two practically relevant regimes where the dynamics is dominated by the persistence of active motion and the elastic confinement of the VE fluid, with a phoretically active Pt-coated Janus colloid in a dynamic optical trap, and show quantitative agreement with the simulations. This approach provides a VE environment with tunable VE properties that remain unaffected by the strength of self-propulsion, allowing us to systematically investigate active Brownian motion in VE media in ways that are not otherwise possible with physical VE fluids.
△ Less
Submitted 7 August, 2026;
originally announced August 2026.
-
TRACE: A Multi-Layer Benchmark for Human AI Controller Coordination Under Drift and Failure
Authors:
Joshua Zuniga,
Srinivasan Subramanian,
Ramya Madhuri Narapureddy,
Md Abdullah Al Hafiz Khan
Abstract:
Modern cyber-physical and AI-assisted systems couple human operators, AI decision modules, and automated controllers in a single control loop, so trustworthiness depends on the whole loop, not any one model. Yet no standard benchmark captures time-aligned, multi-layer traces of how drift and failures propagate across these layers, so we cannot diagnose where coordination breaks down, why, or how t…
▽ More
Modern cyber-physical and AI-assisted systems couple human operators, AI decision modules, and automated controllers in a single control loop, so trustworthiness depends on the whole loop, not any one model. Yet no standard benchmark captures time-aligned, multi-layer traces of how drift and failures propagate across these layers, so we cannot diagnose where coordination breaks down, why, or how to recover. This paper targets one facet of that gap: drift, a deviation that can originate in any stack layer and that conventional single-modality monitoring cannot localize to a layer or pin to an onset time. We construct a benchmark by injecting controlled drift into traces derived from ALFRED, a grounded-instruction benchmark for everyday household tasks, yielding 1,918 drifted traces. Each trace is a time-aligned sequence of per-step records across five execution layers (state, observation, decision, rules, control), labeled with the drift type, affected layer, onset time, responsible actor, and causal mechanism, and validated by independent raters with inter-annotator agreement reported. We pair the dataset with a leak-aware protocol that removes a near-perfect onset leak, and a baseline study across classical, recurrent, and attention-based model families. Under this honest protocol, drift is identifiable and attributable well above random and majority baselines across every family (affected layer macro-F1 near 0.70, responsible actor near 0.85, causal mechanism near 0.49), and heavy attention offers no advantage over simpler models on this symbolic benchmark.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
Bypassing Krum: Selection-Aware Backdoor Attacks in Federated Learning
Authors:
Srinivasan Subramanian,
Md. Abdullah Al Hafiz Khan,
Kazi Aminul Islam
Abstract:
Robust aggregation methods are widely used in federated learning to mitigate the impact of adversarial client behavior. Distance-based aggregation rules, such as Krum and Multi-Krum, select updates that are closest to the majority under the assumption that benign updates form a compact cluster. However, these methods rely on geometric properties that can be exploited by adaptive adversaries. We in…
▽ More
Robust aggregation methods are widely used in federated learning to mitigate the impact of adversarial client behavior. Distance-based aggregation rules, such as Krum and Multi-Krum, select updates that are closest to the majority under the assumption that benign updates form a compact cluster. However, these methods rely on geometric properties that can be exploited by adaptive adversaries. We introduce the Krum-Proxy attack, a selection-aware backdoor injection strategy that consistently bypasses Byzantine-robust aggregation. Rather than relying on naive scaling or constraining, our method actively optimizes malicious updates to infiltrate the dense core of the benign distribution. The proposed method constructs adversarial updates that are not only similar to benign updates but are also optimized to lie in regions of the update space that are favored during aggregation. This is achieved through a two-stage optimization procedure that separates task-specific attack objectives from geometry-aware refinement, using a nearest-neighbor proxy, stochastic reference modeling, and anchor-guided alignment. To maintain stealth, we introduce a projection mechanism that constrains adversarial updates within realistic norm and variance bounds. Experiments on standard federated learning benchmarks show that Krum-Proxy achieves higher attack success while preserving clean accuracy, highlighting the vulnerability of distance-based aggregation to selection-aware adversaries.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Authors:
Xiaomin Li,
Yuexing Hao,
Jianheng Hou,
Jintao Huang,
Qianfeng Wen,
Shirley Huang,
Yifan Liu,
Xiaoyi Liu,
Yilan Fan,
Yijun Wang,
Koutian Wu,
Ruoqi Gao,
Muhammad Ahmed Mohsin,
Jing Tang,
Brihi Joshi,
Heming Liu,
Zheyuan Deng,
Zonglin Di,
Sankalp Jajee,
Jiuyao Lu,
Zhiwei Zhang,
Saksham Kapoor,
Ishan Gupta,
Yunhan Zhao,
Chanwoo Park
, et al. (68 additional authors not shown)
Abstract:
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First,…
▽ More
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.
△ Less
Submitted 4 August, 2026;
originally announced August 2026.
-
Similarity Weighted Aggregation with Global Differential Privacy for Federated Brain Lesion Segmentation
Authors:
Muhammad Irfan Khan,
Eero Lehtonen,
Joni Obradovic,
Elina Kontio,
Esa Alhoniemi,
Suleiman A. Khan,
Mojtaba Jafaritadi
Abstract:
Federated Learning (FL) enables collaborative training of machine learning models across multiple institutions without sharing sensitive data, making it particularly suitable for medical imaging applications. However, heterogeneous data distributions across institutions and potential information leakage through model updates remain important challenges. In this work, we propose DP-SimAgg, a privac…
▽ More
Federated Learning (FL) enables collaborative training of machine learning models across multiple institutions without sharing sensitive data, making it particularly suitable for medical imaging applications. However, heterogeneous data distributions across institutions and potential information leakage through model updates remain important challenges. In this work, we propose DP-SimAgg, a privacy-preserving federated learning framework that integrates similarity-weighted aggregation with a server-side differential privacy mechanism. The proposed method applies L2 clipping to bound collaborator updates, computes similarity-based aggregation weights to mitigate the effects of non-IID data distributions, and injects calibrated Gaussian noise at the central server, providing per-round privacy guarantees under the assumed sensitivity bound. The framework is implemented using Intel's OpenFL platform and evaluated on the FeTS 2022 dataset consisting of 1251 multi-modal MRI scans for brain tumor segmentation. Experimental results demonstrate that DP-SimAgg maintains competitive segmentation performance while providing privacy protection. Under a strict per-round privacy budget (epsilon = 1, cumulative epsilon_total = 20 over 20 rounds), the method achieves Dice scores of 0.6357, 0.5305, and 0.5274 for the enhancing tumor (ET), tumor core (TC), and whole tumor (WT) regions, respectively. With a more relaxed per-round budget (epsilon = 10, cumulative epsilon_total = 200), performance approaches that of the non-private baseline while incorporating a central Gaussian mechanism with per-round (epsilon, delta)-DP accounting under the assumed sensitivity bound. These results highlight the potential of DP-SimAgg for enabling privacy-preserving collaborative learning in medical imaging applications.
△ Less
Submitted 1 August, 2026;
originally announced August 2026.
-
Bridging the English-Arabic Medical Knowledge Gap: Targeted Low-Rank Adaptation via Causal Layer Selection
Authors:
Chaimae Abouzahir,
Musa Khan,
Hala Ali-Hassan,
Congbo Ma,
Khaled Saleh,
Yousra Sadqi,
Jihad Mallat,
Walid Al-Eisawi,
Nizar Habash,
Farah E. Shamout
Abstract:
Large Language Models (LLMs) perform strongly in English medical tasks but degrade substantially in Arabic, a gap widely attributed to limited training data. We systematically investigate this assumption via tuned lens probing and causal activation patching, and find that Arabic medical knowledge is present in intermediate model representations but fails to surface at the output. This mechanistic…
▽ More
Large Language Models (LLMs) perform strongly in English medical tasks but degrade substantially in Arabic, a gap widely attributed to limited training data. We systematically investigate this assumption via tuned lens probing and causal activation patching, and find that Arabic medical knowledge is present in intermediate model representations but fails to surface at the output. This mechanistic insight motivates a targeted adaptation strategy: rather than fine-tuning the full network, we propose Targeted Low-Rank Adaptation (TLoRA), restricted to the layer window where cross-lingual representations diverge, upstream of the output layers where the failure manifests. We evaluate TLoRA on multiple-choice medical QA, where our approach outperforms full-network LoRA, zero-shot, and few-shot baselines. We further evaluate it on short-answer generation and multi-turn clinical dialogue, where it performs competitively without the need for task-specific finetuning. We additionally introduce AraClinicDialog, a clinician-constructed Arabic medical dialogue benchmark in MSA with validated variants across four Arabic dialects. Together, these contributions demonstrate that mechanistic diagnosis can serve as a practical guide for targeted adaptation in underrepresented-language medical LLMs.
△ Less
Submitted 31 July, 2026;
originally announced August 2026.
-
From Code Review to Code Critique: Intent, Drift, and Spotlight for AI-Generated Diffs at Scale
Authors:
Chandra Maddila,
Mashrur Rashik,
Euna Mehnaz Khan,
Smriti Jha,
James Saindon,
Nachi Nagappan,
Peter C. Rigby
Abstract:
AI coding agents are generating code at volumes that exceed the capacity of traditional peer review. At the same time, existing AI code review tools over-index on low-value suggestions such as style and best practices while under-indexing on the concerns human reviewers prioritize most: correctness, security, and performance. We present ARCTIC, an AI-powered Code Critique system that reframes code…
▽ More
AI coding agents are generating code at volumes that exceed the capacity of traditional peer review. At the same time, existing AI code review tools over-index on low-value suggestions such as style and best practices while under-indexing on the concerns human reviewers prioritize most: correctness, security, and performance. We present ARCTIC, an AI-powered Code Critique system that reframes code review around three capabilities: intent prediction, which infers why a change was made from conversation logs and metadata; drift detection, which measures divergence between the developer's intent and the agent's output via backtranslation; and code spotlight, which ranks the regions of a diff most warranting human scrutiny. We ground these capabilities in a six-theme taxonomy derived from 18,000 code reviews. Offline evaluation shows that intent prediction achieves 0.86 F1, drift detection reaches near-perfect ordinal agreement with human annotators (QWK = 0.907), and spotlight outperforms the baseline AI reviewer by 2.4x on quality estimation at 5x fewer tokens. In the experimental rollout, the drift scores reduces code misalignment by an additional 5.76 points (p = 0.026), intent prediction receives 90.2% approval, and zero defects have been attributed to self-reviewed diffs since launch.
△ Less
Submitted 14 August, 2026; v1 submitted 31 July, 2026;
originally announced July 2026.
-
TongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment
Authors:
MD Wahiduzzaman Khan,
Mingshan Jia,
Xiaolin Zhang,
En Yu,
Kaska Musial-Gabrys
Abstract:
Modern face reenactment systems achieve impressive pose and expression transfer using geometry-driven representations. However, they largely ignore tongue dynamics, leading to anatomically inconsistent mouth interiors during speech and expressive motions. We introduce the first framework for cross-identity tongue dynamics transfer in face reenactment. We propose a foundation-model-assisted bootstr…
▽ More
Modern face reenactment systems achieve impressive pose and expression transfer using geometry-driven representations. However, they largely ignore tongue dynamics, leading to anatomically inconsistent mouth interiors during speech and expressive motions. We introduce the first framework for cross-identity tongue dynamics transfer in face reenactment. We propose a foundation-model-assisted bootstrapping pipeline that produces a dedicated tongue segmentation model for in-the-wild reenactment without curated annotations. We further introduce a spatially constrained latent masked diffusion model for realistic tongue synthesis, with adaptive mask dilation for seamless mouth boundary transitions. Extensive experiments demonstrate improvements of more than two times over all baselines on every tongue-specific metric. We additionally propose a VLM-based evaluation protocol that replicates expert annotation at scale, confirming perceptual superiority across all ablation variants.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head Avatars
Authors:
MD Wahiduzzaman Khan,
Mingshan Jia,
Xiaolin Zhang,
En Yu,
Kaska Musial-Gabrys
Abstract:
Creating photorealistic animatable head avatars from a single image remains a fundamental challenge in digital human synthesis. While recent 3D Gaussian Splatting methods have achieved promising results, they rely on external tracking pipelines whose latency is excluded from inference measurements. Furthermore, they adopt unified representations that entangle geometrically distinct facial regions,…
▽ More
Creating photorealistic animatable head avatars from a single image remains a fundamental challenge in digital human synthesis. While recent 3D Gaussian Splatting methods have achieved promising results, they rely on external tracking pipelines whose latency is excluded from inference measurements. Furthermore, they adopt unified representations that entangle geometrically distinct facial regions, limiting both expressiveness and rendering fidelity. We propose SpiD (Split and Drive), a single-image Gaussian head avatar framework built on two disentanglement axes. The compute axis internalizes per-frame driving, eliminating external tracking dependency at inference. The feature axis decomposes the avatar into three specialized Gaussian branches, each modeling a geometrically distinct facial domain. Extensive experiments demonstrate consistently strong performance against state-of-the-art methods while achieving the fastest inference speed among all compared methods on a single GPU with the complete driving pipeline included.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
Outflow Behavior from the Transonic Advective Disks: A Hydrodynamical Simulation Study
Authors:
Sanjit Debnath,
Indranil Chattopadhyay,
Raj Kishor Joshi,
Philippe Laurent,
Priyesh Kumar Tripathi,
M. Saleem Khan
Abstract:
We investigate the properties of outflows from the transonic advective accretion disk using hydrodynamical numerical simulations. We consider two different disk temperatures with an order-of-magnitude difference. For the hotter disk, we adopt initial conditions for velocity, specific angular momentum, and temperature from analytical solutions. In the colder disk case, the velocity and angular mome…
▽ More
We investigate the properties of outflows from the transonic advective accretion disk using hydrodynamical numerical simulations. We consider two different disk temperatures with an order-of-magnitude difference. For the hotter disk, we adopt initial conditions for velocity, specific angular momentum, and temperature from analytical solutions. In the colder disk case, the velocity and angular momentum profiles are kept identical, while an order of magnitude reduction in the temperature. The simulations are performed in the presence of viscosity and radiative cooling, considering bremsstrahlung and synchrotron processes. In both disk models, the outflow rate increases with viscosity. We also examine the poloidal velocity structures for both cases. We analyze the influence of viscosity on the mass flux-weighted energy and momentum fluxes of the outflows. Our results show that both energy and momentum fluxes increase with higher viscosity and may play a significant role in accretion feedback mechanisms.
△ Less
Submitted 30 July, 2026; v1 submitted 29 July, 2026;
originally announced July 2026.
-
Effect of Multi-Species Plasma on Fanaroff-Riley Radio Jets
Authors:
Priyesh Kumar Tripathi,
Indranil Chattopadhyay,
Raj Kishor Joshi,
Ritaban Chatterjee,
Sanjit Debnath,
M. Saleem Khan
Abstract:
The Fanaroff-Riley (FR) dichotomy observed in extragalactic radio jets has been attributed to a range of possible mechanisms, including intrinsic jet properties such as the presence of different species in the plasma. Jet material may span from a pure electron-positron pair plasma to mixed plasmas containing electrons, positrons, and protons, or even to hadronic jets made up of electrons and proto…
▽ More
The Fanaroff-Riley (FR) dichotomy observed in extragalactic radio jets has been attributed to a range of possible mechanisms, including intrinsic jet properties such as the presence of different species in the plasma. Jet material may span from a pure electron-positron pair plasma to mixed plasmas containing electrons, positrons, and protons, or even to hadronic jets made up of electrons and protons only. To investigate this aspect, we present results from three-dimensional simulations of low-power, supersonic, magnetized jets at kiloparsec scales in a magnetohydrodynamic framework. By varying the plasma composition, we show its impact on jet stability and on the development of diffuse structures typical of core-brightened FR type I sources. Our results indicate that the growth of non-axisymmetric instabilities plays a key role in disrupting the jet head.
△ Less
Submitted 30 July, 2026; v1 submitted 29 July, 2026;
originally announced July 2026.
-
A Physics-Informed Neural Operator for Thermal Ranking of Low-Cost Wall Materials in Hot-Dry Climates
Authors:
Muhammad Akbar Khan,
Fahim Raees,
Ubaida Fatima
Abstract:
Identifying cost-effective indigenous building materials that minimise heat penetration through walls is critical for indoor thermal comfort in low-income rural housing in hot-dry climates, where summer temperatures routinely exceed 45 C. We present a two-stage computational framework for thermal ranking of five low-cost indigenous wall materials: mud brick, clay-straw adobe, lime-stabilised bambo…
▽ More
Identifying cost-effective indigenous building materials that minimise heat penetration through walls is critical for indoor thermal comfort in low-income rural housing in hot-dry climates, where summer temperatures routinely exceed 45 C. We present a two-stage computational framework for thermal ranking of five low-cost indigenous wall materials: mud brick, clay-straw adobe, lime-stabilised bamboo panel, fired clay brick, and lime-mud composite. First, a validated Crank-Nicolson finite difference method (FDM) solves the one-dimensional transient heat equation with Robin boundary conditions under diurnal solar and outdoor air-temperature forcing, generating 1500 periodic-day solutions across a nine-dimensional parameter space by Latin Hypercube sampling. Second, a Physics-Informed Neural Operator (PINO) with a Fourier Neural Operator (FNO) backbone learns the parameter-to-solution operator mu -> T(x,t), enforcing both data fidelity and PDE consistency. The trained PINO attains a relative L2 field error of 5.14e-4 and a 0.201 K mean absolute error on the peak inner surface temperature, preserving the FDM material ranking exactly; PINO trained on 150 FDM samples matches a data-only FNO trained on twice as many, so the physics loss is most valuable when data are scarce. The periodic-day formulation also yields the ISO 13786 time lag and decrement factor, reproduced to within 0.99 h and 0.010. At nominal hot-dry summer conditions, clay-straw adobe achieves the best cost-performance index among widely available materials. A climate sweep, confirmed by FDM spot checks, reveals a regime boundary: under sub-ambient outdoor conditions the ranking inverts to conductive fired clay brick, delineating heat-exclusion and heat-rejection regimes. The framework supports evidence-based material selection for post-flood reconstruction in hot-dry regions.
△ Less
Submitted 28 July, 2026;
originally announced July 2026.
-
The effect of triaxial galaxy shapes on the dynamics of triple supermassive black holes in a cosmological context
Authors:
Navonil Saha,
Margarita Sobolenko,
Peter Berczik,
Andreas Just,
Fazeel Mahmood Khan
Abstract:
The hierarchical nature of galaxy formation in the Lambda cold dark matter ($Λ$CDM) cosmological framework model often leads to the presence of multiple supermassive black holes (SMBHs) in the galactic nuclei. The timescale over which galaxies merge plays a crucial role in shaping the dynamical evolution and the merger dynamics of their central SMBHs. While binary SMBH evolution has been extensive…
▽ More
The hierarchical nature of galaxy formation in the Lambda cold dark matter ($Λ$CDM) cosmological framework model often leads to the presence of multiple supermassive black holes (SMBHs) in the galactic nuclei. The timescale over which galaxies merge plays a crucial role in shaping the dynamical evolution and the merger dynamics of their central SMBHs. While binary SMBH evolution has been extensively studied, the long-term dynamics of triple SMBH systems, especially in realistic, nonspherical galactic potentials, still remain less understood. In this work, we investigated the role of triaxiality in shaping the dynamical evolution of three SMBH triple systems taken from the ROMULUS25 cosmological simulation embedded in triaxial stellar backgrounds to find common dynamical evolution patterns and estimate typical coalescence times using high-resolution gravitodynamical $\textit{N}$-body simulations. We explored a range of orbital configurations and host galaxy shapes with initial conditions from the ROMULUS25 data and tracked the orbital evolution from the galactic inspiral to the formation of hard binaries at sub-parsec separations and used the observed hardening rates to estimate the time of coalescence. In all cases, the two heaviest black holes form an efficiently hardening binary, which merges within the Hubble time, while the third black hole (BH) either forms a stable hierarchical triple system with the heavier binary or remains on a wide galactic orbit. Finally, we analyzed the triaxiality of the galactic remnant from our simulations and conclude that the initial triaxial shape of the galaxies does not significantly change the final dynamical outcome of the triple systems.
△ Less
Submitted 26 July, 2026;
originally announced July 2026.
-
Separating Clicks from Baits: Using Large Language Models to Detect Misleading YouTube Thumbnails
Authors:
Wajiha Naveed,
Muhammad Muneeb Pervez,
Zaeem Mohtashim Khan,
Zafar Ayyub Qazi,
Zartash Afzal Uzmi
Abstract:
Misleading video thumbnails on platforms like YouTube are a pervasive problem, undermining user trust and platform integrity. This paper proposes a novel multi-modal detection pipeline that uses Large Language Models (LLMs) to flag misleading thumbnails. We first construct a comprehensive dataset of 2,843 videos from eight countries, including 1,359 misleading thumbnail videos that collectively am…
▽ More
Misleading video thumbnails on platforms like YouTube are a pervasive problem, undermining user trust and platform integrity. This paper proposes a novel multi-modal detection pipeline that uses Large Language Models (LLMs) to flag misleading thumbnails. We first construct a comprehensive dataset of 2,843 videos from eight countries, including 1,359 misleading thumbnail videos that collectively amassed over 7.6 billion views, providing a unique cross-cultural perspective on this global issue. Our detection pipeline integrates video-to-text descriptions, thumbnail images, and subtitle transcripts to holistically analyze content and flag misleading thumbnails. Through extensive experimentation and prompt engineering, we evaluate the performance of four frontier-level LLMs, including GPT-4o, GPT-4o Mini, Claude 3.5 Sonnet, and Gemini-1.5 Flash. We further evaluate open-weight vision-language models, LLaVA-v1.5 and Qwen2.5-VL-7B-Instruct, to assess the generalizability of our approach beyond proprietary systems. Our findings show the effectiveness of LLMs in identifying misleading thumbnails, with Claude 3.5 Sonnet consistently showing strong performance, achieving an accuracy of 93.8%, precision over 92%, and recall exceeding 94% in certain scenarios. Beyond evaluating detection performance, we conducted a careful failure analysis to understand when LLMs fail in identifying misleading thumbnails. We discuss the implications of our findings for content moderation, user experience, and the ethical considerations of deploying such systems at scale. Our findings pave the way for more transparent, trustworthy video platforms and stronger content integrity for audiences worldwide.
△ Less
Submitted 26 July, 2026;
originally announced July 2026.
-
Gravitational acceleration encoded in Jaynes Cummings exchange frequencies: Quantum Fisher information, readout, and validity conditions
Authors:
Salman Sajad Wani,
Mughees Ahmed Khan,
Mohammad Haris Khan,
Fardeen Ahmad Sofi,
Abrar Ahmed Naqash,
Saif Al-Kuwari
Abstract:
We derive an effective trapped atom--cavity model in which a constant gravitational acceleration shifts the oscillator equilibrium and changes the local standing-wave coupling, thereby encoding the acceleration in the Jaynes Cummings exchange frequencies. With the atom and motion initially in their ground states and the cavity field initially coherent, we solve the closed-system carrier dynamics e…
▽ More
We derive an effective trapped atom--cavity model in which a constant gravitational acceleration shifts the oscillator equilibrium and changes the local standing-wave coupling, thereby encoding the acceleration in the Jaynes Cummings exchange frequencies. With the atom and motion initially in their ground states and the cavity field initially coherent, we solve the closed-system carrier dynamics exactly and derive the displaced-frame quantum Fisher information (QFI) of the joint atom-cavity state. This QFI is proportional to the square of the local coupling slope, grows quadratically with interrogation time, and scales linearly with mean photon number. At the node, phase-referenced Ramsey detection gives a sign-sensitive estimate of axial acceleration and locally saturates the joint QFI. Away from the node, photon counting and phase-optimized homodyne detection provide cavity readouts when the cavity state carries more QFI than the atomic state. At the off-node operating point studied, Lindblad simulations show that cavity loss produces a finite-time QFI optimum. Lamb Dicke and sideband-suppression conditions control the carrier approximation. In the closed-system benchmark, the carrier-model QFI agrees with the atom-cavity QFI obtained from the unexpanded model after tracing out motion.
△ Less
Submitted 25 July, 2026;
originally announced July 2026.
-
Identifying the signatures of residual activity in harmonically bound active Brownian dynamics
Authors:
Sanatan Halder,
Manas Khan
Abstract:
A confined self-propelled particle exhibits a range of intriguing dynamical phenomena dictated by the interplay between the intrinsic activity of the particle and the imposed confinement. This competition manifests as a crossover in the steady-state position distribution of a harmonically bound active Brownian particle (HBABP) from Boltzmann-like to bimodal, commonly recognized as the passive and…
▽ More
A confined self-propelled particle exhibits a range of intriguing dynamical phenomena dictated by the interplay between the intrinsic activity of the particle and the imposed confinement. This competition manifests as a crossover in the steady-state position distribution of a harmonically bound active Brownian particle (HBABP) from Boltzmann-like to bimodal, commonly recognized as the passive and active regimes, respectively, upon variations in activity and confinement strength. We present a comprehensive analysis of the resultant dynamics of an HBABP employing analytical calculations and numerical simulations, examining the variations in the position distribution, residual or resultant velocity, mean square displacement, power spectral density, and effective harmonic confinement at varying activities in the characteristic regimes across the crossover. These analyses provide a reliable identification of the signature of residual or remnant activity in ABP dynamics after being impeded by the harmonic confinement. Our results show that the resultant HBABP dynamics in the regime with a Boltzmann-like position distribution is dominated by residual activity, and the motion in the other regime, with a bimodal position distribution, is similar to that of a harmonically bound Brownian particle--devoid of residual activity--at a displaced position, where the activity is balanced by the restoring force field.
△ Less
Submitted 24 July, 2026;
originally announced July 2026.
-
A Dual Path Framework with Hotspot Guided Fusion for Three Dimensional CT to PET Synthesis in Head and Neck Cancer
Authors:
Mohd Maaz Khan,
Oluwaseyi Oderinde
Abstract:
18F-FDG PET/CT plays a central role in staging, treatment planning, and response assessment for head and neck cancer by providing functional information that complements anatomical CT imaging. However, PET acquisition requires radiotracer administration, specialized infrastructure, and additional cost, limiting its availability for repeated imaging. We present a proof of concept deep learning fram…
▽ More
18F-FDG PET/CT plays a central role in staging, treatment planning, and response assessment for head and neck cancer by providing functional information that complements anatomical CT imaging. However, PET acquisition requires radiotracer administration, specialized infrastructure, and additional cost, limiting its availability for repeated imaging. We present a proof of concept deep learning framework for synthesizing PET like images directly from routine CT scans with the goal of providing complementary metabolic information that may support imaging triage and clinical decision support rather than replace diagnostic PET. Forty-four patients from the publicly available QIN-HEADNECK dataset were retrospectively analyzed using five fold cross-validation. We propose a fully three dimensional dual path architecture consisting of (i) a regression U-Net optimized for voxel-wise quantitative SUV estimation and (ii) a conditional generative adversarial network optimized for realistic PET texture. Their outputs are integrated using hotspot guided Laplacian pyramid blending, allowing quantitative information from the regression pathway to be preserved within metabolically active regions while leveraging adversarial texture synthesis elsewhere. The proposed framework achieved a mean absolute error of 0.00395, PSNR of 39.19 dB, and SSIM of 0.9634 on reconstructed three dimensional PET volumes. Qualitative evaluation demonstrated accurate localization of many FDG-avid lesions while producing anatomically realistic background texture. Consistent with previous CT to PET synthesis studies, the principal limitation was systematic underestimation of SUV within highly metabolically active tumor regions.
△ Less
Submitted 23 July, 2026;
originally announced July 2026.
-
A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models
Authors:
Mohtashim Khan
Abstract:
Traditional benchmarks for LLMs primarily rely on static datasets and objective scoring metrics, which often fail to capture differences in response quality when multiple answers are acceptable. In such settings, correctness alone is insufficient to distinguish between responses that vary in clarity, completeness, and usefulness.
This paper introduces a consensus-based evaluation framework that…
▽ More
Traditional benchmarks for LLMs primarily rely on static datasets and objective scoring metrics, which often fail to capture differences in response quality when multiple answers are acceptable. In such settings, correctness alone is insufficient to distinguish between responses that vary in clarity, completeness, and usefulness.
This paper introduces a consensus-based evaluation framework that measures relative preference among model-generated responses rather than absolute correctness. Instead of evaluating outputs against a fixed ground truth, we assess how a panel of diverse LLMs ranks anonymized candidate responses to the same prompt. This approach treats aggregate inter-model agreement as a proxy for perceived response quality under blind conditions.
We conduct a controlled study using five state-of-the-art LLMs across multiple domains, including programming, general knowledge, safety, logical reasoning, and mathematics. Each model generates responses and independently ranks peer outputs through a structured voting process. Scores are aggregated into a Relative Intelligence Index (RII), representing how frequently a model's responses are preferred by other models.
Our findings reveal consistent preference patterns across domains, with certain models more frequently ranked highly by their peers. However, we emphasize that these results reflect inter-model preference alignment rather than objective correctness or human judgment. This framework provides a scalable, model-driven method for comparative evaluation, offering an alternative perspective on response quality in scenarios where multiple valid answers exist. While not directly aligned with human evaluation, prior work suggests that aggregated model preferences can partially correlate with human judgments, motivating this as a proxy signal.
△ Less
Submitted 19 July, 2026;
originally announced July 2026.
-
Explainable Deepfake Detection Challenge
Authors:
Abhijeet Narang,
Kartik Kuckreja,
Shreya Ghosh,
Muhammad Haris Khan,
Usman Tariq,
Jianfei Cai,
Abhinav Dhall
Abstract:
Deepfake detection is moving beyond binary classification decisions toward systems that can also explain the visual evidence supporting those decisions. This transition is important for real-world verification settings, where diverse users need to understand not only whether an image is manipulated, but also why it is considered suspicious. The Explainable Deepfake Detection Challenge at ACM Multi…
▽ More
Deepfake detection is moving beyond binary classification decisions toward systems that can also explain the visual evidence supporting those decisions. This transition is important for real-world verification settings, where diverse users need to understand not only whether an image is manipulated, but also why it is considered suspicious. The Explainable Deepfake Detection Challenge at ACM Multimedia 2026 is designed to benchmark this joint capability. Built on XPlainVerse, a million-scale benchmark for explainable deepfake detection, the challenge evaluates methods on image classification and grounded natural-language explanation generation. Participants submit a real/fake label together with two explanations for each image: a detailed complex explanation for technical users and a concise simple explanation for general users. The evaluation combines classification metrics with semantic similarity, simplicity, and intent-aware grounding metrics that assess whether explanations identify the relevant manipulated entities and supporting visual evidence. The methodologies developed through the challenge will contribute to the development of next-generation explainable deepfake detectors. Evaluation script, baseline models, and accompanying code are available on https://github.com/Abhijeet8901/XPlainVerse-ACMChallenge.
△ Less
Submitted 23 July, 2026;
originally announced July 2026.
-
The Blessing of Dimensionality: How Near-Orthogonality in High-Dimensional Spaces Explains Temporal Portability
Authors:
Abigail Woodring,
Adrian Chan,
Rana Muhammad Shahroz Khan,
Sukwon Yun,
Chau-Wai Wong,
Tianlong Chen
Abstract:
Fine-tuning has been widely used to adapt large language models (LLMs) for domain-specific tasks. Parameter efficient fine-tuning (PEFT) methods such as low-rank adaptation (LoRA) are frequently used to reduce computational costs. PortLLM is a training-free and data-free scheme used to adapt LLMs after continual pretraining. Although the initial PortLLM results show that LoRA patches exhibit short…
▽ More
Fine-tuning has been widely used to adapt large language models (LLMs) for domain-specific tasks. Parameter efficient fine-tuning (PEFT) methods such as low-rank adaptation (LoRA) are frequently used to reduce computational costs. PortLLM is a training-free and data-free scheme used to adapt LLMs after continual pretraining. Although the initial PortLLM results show that LoRA patches exhibit short-term temporal portability, the long-term performance of PortLLM across several updates of continual pretraining remains underexplored. Furthermore, the intriguing effectiveness of PortLLM is not well understood from a theoretical standpoint. We address these two open questions by (1) performing an extensive empirical study of the long-term temporal portability of PortLLM patches across 10 continual pretraining steps using base models Mistral, Gemma, and Qwen; and (2) offering two theoretical analyses to explain our observation that the simple PortLLM method achieves competitive performance. We find empirically that the portability persists across longer time duration, indicating that repeated fine-tuning is not required when the base model is periodically updated. We find theoretically that near-orthogonality of high-dimensional vectors is a key justification for temporal portability. Our analyses also demonstrate a geometric perspective of the loss landscape in facilitating the theoretical comparison of different adaptation options.
△ Less
Submitted 22 July, 2026;
originally announced July 2026.
-
ReFace: Reorganizing Facial Spatiotemporal Representations for Improved Pain Assessment
Authors:
Stefanos Gkikas,
Yu Fang,
Christian Arzate Cruz,
Muhammad Umar Khan,
Raul Fernandez Rojas
Abstract:
Automatic pain assessment from facial video remains challenging due to the spatial heterogeneity of pain-related facial cues. This study proposes ReFace, a spatial reorganization pipeline that divides facial input into four spatial quadrants before tokenization, rather than processing the entire face as a single region. Evaluated on the AI4Pain dataset, the proposed approach achieves $56.00\%$ acc…
▽ More
Automatic pain assessment from facial video remains challenging due to the spatial heterogeneity of pain-related facial cues. This study proposes ReFace, a spatial reorganization pipeline that divides facial input into four spatial quadrants before tokenization, rather than processing the entire face as a single region. Evaluated on the AI4Pain dataset, the proposed approach achieves $56.00\%$ accuracy on the test set using video only, achieving the highest reported accuracy under the fixed AI4Pain benchmark protocol among the compared methods. Notably, the four-quadrant configuration processes the same total pixel budget as the full-face input, yet achieves higher accuracy, suggesting that spatial reorganization can improve performance under the proposed tokenization design. A single quadrant region, processing just one quarter of those pixels, remains competitive at a fraction of the computational cost.
△ Less
Submitted 26 July, 2026; v1 submitted 21 July, 2026;
originally announced July 2026.
-
A Unified Tokenization Framework for Pain Recognition using Heterogeneous 3D Modalities
Authors:
Stefanos Gkikas,
Christian Arzate Cruz,
Valentina Becchetti,
Muhammad Umar Khan,
Alessandro Giuseppi,
Raul Fernandez Rojas
Abstract:
Pain is a complex and pervasive phenomenon affecting a large percentage of the population, and accurate assessment is essential for effective clinical management and intervention. Computational pain recognition systems enable continuous monitoring, support clinical decision-making, and help mitigate pain-related distress and functional decline. This study introduces a unified tokenization framewor…
▽ More
Pain is a complex and pervasive phenomenon affecting a large percentage of the population, and accurate assessment is essential for effective clinical management and intervention. Computational pain recognition systems enable continuous monitoring, support clinical decision-making, and help mitigate pain-related distress and functional decline. This study introduces a unified tokenization framework for heterogeneous 3D modalities in pain recognition that provides a single processing pipeline across behavioral and brain-activity 3D data, without requiring separate architectures for each modality or handcrafted inductive biases. The framework preserves spatial, temporal, and time--frequency structure while mapping diverse inputs into a shared token space. Extensive experiments show that the proposed approach effectively processes facial videos and fNIRS data in both raw-signal and spectrogram-based representations. On the AI4Pain benchmark dataset, the proposed framework achieves state-of-the-art performance while maintaining high computational efficiency and enabling real-time assessment on both GPU and CPU hardware.
△ Less
Submitted 26 July, 2026; v1 submitted 21 July, 2026;
originally announced July 2026.
-
FlexiAvatar: Unified 3D Gaussian Human Avatars Under Arbitrary Body Visibility
Authors:
Yihalem Yimolal Tiruneh,
Muhammad Salman Ali,
Uyoung Jeong,
Muneeb A. Khan,
MD Khalequzzaman Chowdhury Sayem,
Allanur Bayramgeldiyev,
Binod Bhattarai,
Seungryul Baek
Abstract:
Reconstructing animatable 3D human avatars from monocular video is a fundamental problem in computer vision with broad applications in AR/VR and digital content creation. Existing approaches typically couple parametric body models with neural rendering or 3D Gaussian splatting and optimize all body regions jointly from short videos, which often degrades fidelity in the visible areas. To overcome t…
▽ More
Reconstructing animatable 3D human avatars from monocular video is a fundamental problem in computer vision with broad applications in AR/VR and digital content creation. Existing approaches typically couple parametric body models with neural rendering or 3D Gaussian splatting and optimize all body regions jointly from short videos, which often degrades fidelity in the visible areas. To overcome this limitation, we introduce FlexiAvatar, a unified framework that explicitly optimizes only the visible body regions, effectively eliminating artifacts arising from unobserved limbs. Our method integrates occlusion-robust SMPL-X tracking with part-specific residual refinement to capture high-frequency geometric and appearance details. To complete entirely unseen regions (e.g., back views), we leverage a diffusion-based approach to generate texture consistent with the observed appearance. Experiments on full-body (NeuMan, ZJU-MoCap, WildAvatar), upper/half-body (talk-show clips), and head-only (INSTA) inputs show that FlexiAvatar delivers consistently higher reconstruction quality, outperforming state-of-the-art methods by an average PSNR improvement of approximately 3% across datasets. Finally, by restricting optimization to observed regions, our method reduces the effective number of Gaussians that must be optimized and rendered, leading to reduced runtime and memory overhead in partial-visibility scenarios.
△ Less
Submitted 21 July, 2026;
originally announced July 2026.
-
Depth-Regularized JEPA World Models Learn More Transferable Representations from Real Outdoor Robot Data
Authors:
Usman M. Khan
Abstract:
World models, especially based on JEPA architectures, have been shown to learn robust dynamics of various environments. However, learning from visually complex real-world data remains a challenge, especially in unpredictable outdoor environments. We introduce depth as a geometric prior during training in learning more robust latent dynamics directly from robot video data and handling visual comple…
▽ More
World models, especially based on JEPA architectures, have been shown to learn robust dynamics of various environments. However, learning from visually complex real-world data remains a challenge, especially in unpredictable outdoor environments. We introduce depth as a geometric prior during training in learning more robust latent dynamics directly from robot video data and handling visual complexity. This combines depth supervision with an isotropy-inducing latent regularizer (SIGReg), maximizing task-agnostic latent diversity while constraining how that diversity is organized, with the combined objective targeting the highest-entropy representation consistent with scene geometry. To satisfy this greater complexity without increasing inference time, we also add training-only overparameterization. Training an 18M-parameter model on video from a real agricultural robot, we evaluate with frozen-representation visual odometry probes, predictor-based surprise detection, and multi-step latent rollout fidelity. Compared to the baseline LeWM, our method lowers visual odometry probe error by 33%, substantially increases surprise-score separation both in-domain and on the out-of-domain TartanGround benchmark, and improves multi-step rollout fidelity under domain shift, with gains that grow with rollout horizon. Notably, we also see improvements in surprise-score separation on physics understanding that is not directly tied to 3D geometry, such as lighting and shadows. These results show that a lightweight training-time geometric prior makes a compact JEPA world model more useful and more transferable on real outdoor data with strong underlying representations, without adding inference overhead. Our work suggests that depth as a physically grounded prior can enhance world model generalization on a variety of tasks.
△ Less
Submitted 14 July, 2026;
originally announced July 2026.
-
GenSyn10: A Multi-Generative AI Dataset For Benchmarking Image Classification
Authors:
Md Faraz Kabir Khan,
Saeed Anwar,
Ghulam Mubashar Hassan
Abstract:
The rapid advancement of generative AI has outpaced our ability to reliably detect its outputs, particularly when detectors encounter generators they have not seen before. We introduce GenSyn10, a CIFAR-10-aligned synthetic image dataset of 60,000 images (10 classes, 32$\times$32, 50k/10k split) generated using three architecturally diverse state-of-the-art models: FLUX.2-dev (Rectified Flow Trans…
▽ More
The rapid advancement of generative AI has outpaced our ability to reliably detect its outputs, particularly when detectors encounter generators they have not seen before. We introduce GenSyn10, a CIFAR-10-aligned synthetic image dataset of 60,000 images (10 classes, 32$\times$32, 50k/10k split) generated using three architecturally diverse state-of-the-art models: FLUX.2-dev (Rectified Flow Transformer), HunyuanImage-3.0 (MoE Transformer), and Qwen-Image-2512 (Multimodal Diffusion Transformer), to advance research in AI-generated image detection. A central challenge in this domain is that detectors perform well on known generators but degrade on unseen ones. GenSyn10 addresses this limitation by curating data from multiple contemporary architectures under a standardized generation protocol, enabling controlled and systematic evaluation of out-of-distribution (OOD) generalization to novel generators. Images are generated using a template-based prompt engine and downsampled to ensure consistency. We evaluate 17 image classification models under a four-stage protocol: real-data baseline, zero-shot transfer, fine-tuning, and retention. Despite a measurable domain gap, CIFAR-10-trained models achieve up to 96.86\% zero-shot accuracy on GenSyn10, increasing to 99.88\% after fine-tuning. In binary real-vs-synthetic classification, fine-tuned models achieve 97-99.9\% accuracy on seen generators but drop to 79-96\% on images from an unseen generator, highlighting persistent limitations in OOD generalization. These results establish GenSyn10 as a controlled benchmark for studying synthetic image detection beyond single-generator settings, supporting research on robustness, domain adaptation, and cross-generator generalization.
△ Less
Submitted 10 July, 2026;
originally announced July 2026.
-
All Games Have Equilibria
Authors:
M. Ali Khan,
Arthur Paul Pedersen,
Maxwell B. Stinchcombe
Abstract:
Research on Nash equilibrium existence for infinite games has grown into a patchwork of technical preconditions and counterexamples. This paper presents a unified program in equilibrium theory by revising the predominant model of mixed strategies based on countable additivity. A game is specified by a nonempty set of players and, for each player, a nonempty action set and a bounded von Neumann-Mor…
▽ More
Research on Nash equilibrium existence for infinite games has grown into a patchwork of technical preconditions and counterexamples. This paper presents a unified program in equilibrium theory by revising the predominant model of mixed strategies based on countable additivity. A game is specified by a nonempty set of players and, for each player, a nonempty action set and a bounded von Neumann-Morgenstern utility function. Every such game is shown to admit a Nash equilibrium in finitely additive mixed strategies. In addition, the equilibrium correspondence for any such game is shown to be nonempty, compact-valued, and upper hemicontinuous, and the same is true for equilibria obtained as limits of finite approximations. Techniques developed in this paper show that infinite games long treated as intractable become amenable to direct equilibrium analysis.
△ Less
Submitted 4 August, 2026; v1 submitted 16 July, 2026;
originally announced July 2026.
-
Evaluating Health Misinformation in Low-Resource Languages: Integrating Small Language Models with a Culturally-Sensitive Responsible NLP Framework (Bangla as a Case Study)
Authors:
Farnaz Farid,
Raihan Alam,
Al Al-Areqi,
Farhad Ahamed,
Muhammad Hassan Khan,
Sadia Hossain,
Irena Veljanova,
Anika Tabassum Binte Hossain
Abstract:
Artificial Intelligence (AI) technologies, while serving as a foundational enabler for modern social media and digital health services, exert a bivalent effect by simultaneously acting as a combatant against and a spread vector for misinformation. A prevalent challenge in mitigating this issue arises in non-English contexts and low socioeconomic classes, where limited data hinders the training of…
▽ More
Artificial Intelligence (AI) technologies, while serving as a foundational enabler for modern social media and digital health services, exert a bivalent effect by simultaneously acting as a combatant against and a spread vector for misinformation. A prevalent challenge in mitigating this issue arises in non-English contexts and low socioeconomic classes, where limited data hinders the training of AI models for effective detection. Consequently, culturally and linguistically diverse (CALD) communities struggle to access trustworthy health information through AI-driven tools. Current AI tools underperform due to a lack of training data and are largely unable to consider language nuances and traditions in non-English contexts. This research addresses these gaps by proposing a CALD-friendly AI-based health misinformation detector and providing a dashboard for medical professionals to analyse this misinformation, a critical step toward mitigating a growing concern among CALD populations. To this end, we conduct a series of experiments using a Bangla-translated health misinformation dataset to evaluate the performance of various Small Language Models (SLMs). SLMs are particularly relevant in this context given the frequent underperformance of Large Language Models (LLMs), which often stems from insufficient domain-specific knowledge and the prohibitive costs of resource-intensive fine-tuning. The results demonstrate that Phi-4 is the superior model, achieving an ideal balance between precision and recall in claim extraction. Then, to mitigate the limitations of SLMs, we design and test a novel health misinformation detection framework grounded in Responsible Natural Language Processing (NLP), which incorporates cultural sensitivity, potential for harm, and communication quality, thereby providing a holistic lens for evaluating misinformation in low-resource languages.
△ Less
Submitted 14 July, 2026;
originally announced July 2026.
-
BAT-RM: A Boundary-Aware Transformer with Region-Aware Multi-Directional Mamba for Clinically Deployed Cervical Cancer Radiotherapy Auto-Contouring
Authors:
Istiak Ahmed,
Kazi Shahriar Sanjid,
Galib Ahmed,
Md. Tanzim Hossain,
Md. Anwarul Islam,
Shahrukh Khan,
Md. Ashrif Rahman Arian,
Md. Nishan Khan,
Md. Misbah Khan,
S M Hasibul Hoque,
Rahnuma Shahrin Rista,
Md. Jobairul Islam,
Sheikh Anisul Haque,
Md Arifur Rahman,
Syed Md. Akram Hussain,
Syeda Nashra,
Sayeed Shafayet Chowdhury,
Md. Mostafa Kamal Sarker,
M. Monir Uddin
Abstract:
We present a clinically deployed end-to-end auto-contouring system for cervical cancer radiotherapy planning, anchored by the Boundary-Aware Transformer with Region-Aware Mamba (BAT-RM), a hybrid architecture that integrates Sobel-gated boundary attention, a linear-time, multi-directional Mamba module for long-range context, and a boundary-skeleton-guided fusion gate. This design achieves linear-t…
▽ More
We present a clinically deployed end-to-end auto-contouring system for cervical cancer radiotherapy planning, anchored by the Boundary-Aware Transformer with Region-Aware Mamba (BAT-RM), a hybrid architecture that integrates Sobel-gated boundary attention, a linear-time, multi-directional Mamba module for long-range context, and a boundary-skeleton-guided fusion gate. This design achieves linear-time complexity for long-range context modeling, avoiding the quadratic cost of full spatial self-attention. The full pipeline spans multi-institutional data collection, rigorous inter-rater quality assurance, external validation in an independent cohort, and a web-based clinical interface natively compatible with Varian, RayStation, and Monaco. Against four baselines, BAT-RM achieves superior performance across seven anatomical classes, with statistically significant improvements in target volumes, including GTV and CTV, and in organs at risk such as the rectum and bladder. A prospective multi-center reader study involving 13 radiation oncologists demonstrated that AI assistance elevates junior oncologists' IoU from 0.899 to 0.965, approaching senior-level accuracy, while reducing contouring time by more than 80%. The system also reduced expert consultation rates and improved inter-reader consistency, reflecting gains in both efficiency and quality assurance. Following clinical deployment at a partner hospital, the system reduced patient wait times from days to hours without additional staffing, enabling same-day or next-day initiation of treatment for routine cases. BAT-RM demonstrates that a rigorous research pipeline, from data curation to clinical deployment, can translate directly into measurable patient benefit in resource-constrained settings where the demand for radiotherapy far exceeds specialist capacity.
△ Less
Submitted 11 July, 2026;
originally announced July 2026.
-
A Study Of Skew-Polycyclic Codes Over A Non-Chain Ring
Authors:
Seema Antil,
Seema Chahal,
Manju Khan,
Sugandha Maheshwary
Abstract:
For a prime \(p\) and a positive integer \(m\), let \(\mathbb{F}_{p^m}\) be the finite field of cardinality \(p^m\), and let
$
R_{u^2,v^2,p^m}
=\mathbb{F}_{p^m}+u\mathbb{F}_{p^m}+v\mathbb{F}_{p^m}
+uv\mathbb{F}_{p^m},
~ u^2=v^2=0,\ uv=vu,
$
be a finite non-chain ring. In this paper, we study skew polycyclic codes of length \(lj\) associated with \(f(x)^j\), where \(f(x)\) is a centra…
▽ More
For a prime \(p\) and a positive integer \(m\), let \(\mathbb{F}_{p^m}\) be the finite field of cardinality \(p^m\), and let
$
R_{u^2,v^2,p^m}
=\mathbb{F}_{p^m}+u\mathbb{F}_{p^m}+v\mathbb{F}_{p^m}
+uv\mathbb{F}_{p^m},
~ u^2=v^2=0,\ uv=vu,
$
be a finite non-chain ring. In this paper, we study skew polycyclic codes of length \(lj\) associated with \(f(x)^j\), where \(f(x)\) is a central polynomial of degree \(l\) in $R_{u^2, v^2, p^m}[x; Θ],$ where $Θ$ being an automorphism of \(R_{u^2,v^2,p^m}\). We describe these codes, characterize free skew polycyclic codes, and determine their ranks.
Under suitable centrality assumptions, we decompose the quotient ring associated with \(x^{np^s}-λ\), where \(\gcd(n,p)=1\) and \(Θ(λ)=λ\). This reduces the study of skew \((λ,Θ)\)-constacyclic codes of length \(np^s\) to the study of left ideals of
$\frac{R_{u^2,v^2,p^m}[x;Θ]}{\langle f(x)^j\rangle},
$ where \(f(x)\) is a central irreducible divisor of degree \(l\) of \(x^{np^s}-λ\), for an invertible element \(λ\in R_{u^2,v^2,p^m}\) and \(j\in\mathbb{N}\).
We then apply these results to skew \((λ,Θ)\)-constacyclic codes of length \(p^s\) for different classes of units \(λ\). Several examples are presented to illustrate the theory and to obtain optimal codes. Finally, when \(Θ\) is the identity automorphism, we study constacyclic codes of length \(np^s\) over \(R_{u^2,v^2,p^m}\), according as \(x^n-α_0\) is irreducible or reducible over \(\mathbb{F}_{p^m}\). These results extend the work of \cite{CCDF18} and \cite{ZTG18} on constacyclic codes of length \(np^s\) over \(\mathbb{F}_{p^m}+u\mathbb{F}_{p^m}\) to the finite non-chain ring \(R_{u^2,v^2,p^m}\).
△ Less
Submitted 9 July, 2026;
originally announced July 2026.
-
AI-Driven Thermal Mapping and Management in 3D Integrated Photonic Circuits
Authors:
Liton Kumar Biswas,
Katayoon Yahyaei,
Shajib Ghosh,
M Shafkat M Khan,
Himanandhan Reddy Kottur,
Rayhane Ghane-Motlagh,
Mahdi Nikdast,
Navid Asadizanjani
Abstract:
Photonic Integrated Circuits (PICs) are advancing high-performance computing, data centers, and sensing, yet three-dimensional (3D) PICs introduce critical thermal management challenges due to high-density bonding and heterogeneous materials. Traditional methods like thermal microscopes and in-package sensors yield sparse data, limiting full thermal profile visibility. This paper presents a dual-m…
▽ More
Photonic Integrated Circuits (PICs) are advancing high-performance computing, data centers, and sensing, yet three-dimensional (3D) PICs introduce critical thermal management challenges due to high-density bonding and heterogeneous materials. Traditional methods like thermal microscopes and in-package sensors yield sparse data, limiting full thermal profile visibility. This paper presents a dual-method solution combining an AI-driven thermal modeling framework with a design-based heuristic approach. The AI method integrates sparse sensor data with design layer and density information to predict multilayer temperature variations, while the heuristic approach uses localized material properties, design layout, component geometries, and sensor coordinates to refine thermal estimations in specific regions. A 2D thermal map of a 3D PIC is generated by interpolating sensor data and adjusting for local thermal resistivity using comparative analysis between design regions. The heuristic method complements the AI model, improving estimation accuracy without extensive training data. Together, these methods offer a scalable, accurate solution for real-time thermal mapping and design-time simulation, enabling reliable thermal management in next-generation 3D photonic systems.
△ Less
Submitted 24 June, 2026;
originally announced July 2026.
-
Fusion or Confusion? Potential and Challenges in Fusion of Onboard Sensors and V2X Data in Cooperative Perception
Authors:
Amir Mohammadisarab,
Miguel Sepulcre,
Luca Lusvarghi,
Sergei S. Avedisov,
Mohammad Irfan Khan,
Takayuki Shimizu,
Onur Altintas,
Javier Gozalvez
Abstract:
Connected Automated Vehicles (CAVs) utilize their onboard sensors to perceive the environment. The perception range and accuracy can be affected by adverse weather or non-line-of-sight conditions. Cooperative perception or sensor sharing can overcome these limitations by enabling CAVs to exchange sensor data, thus collectively enhancing their perception capabilities. Previous studies have shown th…
▽ More
Connected Automated Vehicles (CAVs) utilize their onboard sensors to perceive the environment. The perception range and accuracy can be affected by adverse weather or non-line-of-sight conditions. Cooperative perception or sensor sharing can overcome these limitations by enabling CAVs to exchange sensor data, thus collectively enhancing their perception capabilities. Previous studies have shown the potential of cooperative perception, but limited attention has been given to the fusion of V2X data received through cooperative perception messages with onboard sensor information. The fusion process can be influenced by the quantity and quality of the V2X data. An increased volume of V2X data can reduce uncertainty in the perceived environment; however, when the data is noisy, it may compromise the accuracy of the fusion results. This study investigates the fusion of onboard sensor and V2X data in cooperative perception, and demonstrates that while perception can significantly improve as the V2X penetration rate increases, it can introduce a significant number of false positives if V2X data is not highly accurate. False positives result in the detection of ghost objects that do not actually exist. These ghost objects can, in turn, compromise safety and driving efficiency. Our analysis found that false positives or ghost objects can appear even with accurate V2X data. These findings highlight the challenges in cooperative perception and the importance of developing robust data fusion methods to enhance the reliability of cooperative perception. This is particularly relevant in light of ongoing standardization efforts, such as ETSI TS 103 324 on collective perception.
△ Less
Submitted 7 July, 2026;
originally announced July 2026.
-
Prompt-to-Paper: Agentic AI System for Bioinformatics
Authors:
Ramsha Kamran,
Maheera Amjad,
Zartasha Mustansar,
Arsalan Shaukat,
Salma Sherbaz,
Muhammad U. S. Khan
Abstract:
While recent advances in large language models have enabled end-to-end automated manuscript generation, existing systems suffer from three critical deficiencies: (i) generated claims are not deterministically grounded in verifiable literature, (ii) experimental results are frequently fabricated rather than executed, and (iii) there exists no standardized, multi-dimensional framework to assess whet…
▽ More
While recent advances in large language models have enabled end-to-end automated manuscript generation, existing systems suffer from three critical deficiencies: (i) generated claims are not deterministically grounded in verifiable literature, (ii) experimental results are frequently fabricated rather than executed, and (iii) there exists no standardized, multi-dimensional framework to assess whether AI-generated manuscripts meet the quality and rigor required for real-world publication. We present Prompt-to-Paper, a multi-agent framework that directly addresses this evaluation gap through three integrated innovations. First, a deterministic retrieval-augmented generation pipeline with section-aware relevance scoring and snowball citation expansion grounds every claim in a verifiable corpus of 60--100 papers. Second, an autonomous coding agent executes real computational biology experiments replacing synthetic outputs with genuine numerical results. Third, an eight-dimensional automated quality scorer, benchmarked with approximate reference statistics from published papers and augmented with explicit hallucination penalties, provides standardized, reproducible quality assessments. The quality-driven improvement loop uses a context-rich reviser that routes each iteration to one of three researcher actions and fires a deep research cycle every ten iterations to re-run experiments and re-manuscript from stronger outputs. We validate the system on five bioinformatics case studies; all five cases compiled submission-formatted PDFs with zero out-of-range citations. The improvement loop raises manuscript quality by an average of +17.96 points on a 0--100 scale (maximum +26.04. As partial external checks, a human reviewer scored the five manuscripts at an average of 7.0 out of 10. Complete manuscripts are produced at approximately 0.31 USD per paper.
△ Less
Submitted 5 July, 2026;
originally announced July 2026.
-
Quadrature-Aware Complex-Linear Neural Operator for Boundary-to-Field Prediction in Resonant Acoustics
Authors:
Muhammad Idrees Khan,
Hua-Dong Yao
Abstract:
Repeated prediction of acoustic fields from spatially distributed boundary excitation is computationally expensive when each source realization requires a new wave simulation. This work introduces a quadrature-aware complex-linear boundary operator (CLBO) that maps complex normal velocity on a vibrating surface to complex pressure at receiver locations. The model couples learned source and receive…
▽ More
Repeated prediction of acoustic fields from spatially distributed boundary excitation is computationally expensive when each source realization requires a new wave simulation. This work introduces a quadrature-aware complex-linear boundary operator (CLBO) that maps complex normal velocity on a vibrating surface to complex pressure at receiver locations. The model couples learned source and receiver basis functions through an explicit complex surface-quadrature contraction, so the boundary excitation enters linearly by construction. This preserves complex superposition, homogeneity, and zero response to zero excitation, while representing the source through coordinates, normals, and quadrature weights rather than a fixed flattened input vector. Reference data were generated using a verified three-dimensional multiple-relaxation-time (MRT) lattice Boltzmann solver and stored in a solver-agnostic boundary-to-field format. CLBO was compared with a fixed-sensor complex DeepONet under matched case splits and optimization settings, with additional tests of structural consistency, receiver-coordinate interpolation, source discretization, source-family holdout, label efficiency, physics-informed ablations, unseen source mixtures, and computational cost. Across five training seeds, CLBO achieved a mean complex relative field error of 0.184 +/- 0.00771, compared with 0.367 +/- 0.00742 for DeepONet. Its measured source-superposition error was 1.31 x 10^-7, and its mean error on newly simulated mixed-source cases was 0.237, compared with 0.415 for DeepONet. Inference was 1.83 x 10^4 faster than the reference calculation for the reported query size. These results show that enforcing the known complex-linear boundary-to-field structure improves physical consistency and generalization under distributed acoustic excitation.
△ Less
Submitted 5 July, 2026;
originally announced July 2026.