-
Jais 2: A Family of Arabic-Centric Open Large Language Models
Authors:
Mohamed Anwar,
Abed Alhakim Freihat,
George Ibrahim,
Mostafa Awad,
Abdelrahman Sadallah,
Gurpreet Gosal,
Gokulakrishnan Ramakrishnan,
Sarath Chandran,
Biswajit Mishra,
Rituraj Joshi,
Ahmed Frikha,
Etienne Goffinet,
Abhishek Maiti,
Ali El Filali,
Sarah AlBarri,
Samujjwal Ghosh,
Rahul Pal,
Parvez Mullah,
Awantika Shukla,
Sajid siddiki,
Samta Kamboj,
Onkar Pandit,
Sunil Kumar Sahu,
AbdelRahman Elbadawy,
Amr Mohamed
, et al. (35 additional authors not shown)
Abstract:
Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong performance across the Arabic and culturally grounded benchmarks evaluated in this report. The family includes, to our knowledge, the largest open Arabic-centric LLM trained from scratch at 70B parameters, and a competiti…
▽ More
Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong performance across the Arabic and culturally grounded benchmarks evaluated in this report. The family includes, to our knowledge, the largest open Arabic-centric LLM trained from scratch at 70B parameters, and a competitive 8B-parameter variant among the evaluated open models. A custom Arabic-centric vocabulary enables efficient training and inference. In addition, an optimized architecture and training recipe yield highly compute-efficient training. With a substantially smaller token budget than comparable models, Jais 2 achieves strong Arabic performance on the benchmarks considered in this report and competitive English results. The models obtain leading results among the evaluated open models on OALL2 and AraGen. They also perform strongly on several culturally grounded Arabic benchmarks, including poetry, religion, cuisine, and dream interpretation, as well as in general tasks such as translation and summarization. We release the models in HuggingFace under a commercially permissive license. Jais 2 70B is also released as a chat app on the Web, iOS, and Android; it runs on Cerebras hardware, delivering up to 2,000 tokens per second, and enabling high-throughput Arabic-centric chat serving in our deployment setting. By uniting scale, linguistic diversity, cultural fidelity, openness, and speed, Jais 2 provides an open-weight foundation intended to support further research and development in Arabic-centric LLMs.
△ Less
Submitted 7 July, 2026;
originally announced August 2026.
-
CalTwin: Towards Calibrated, Shift-Robust Medical World Models via Fisher-Information Regularisation
Authors:
Behraj Khan,
Shabir Ahmad,
Syed Ahmad Chan Bukhari,
Tahir Qasim Syed
Abstract:
Medical world models aim to learn a latent state of patient or organ physiology and a transition function that forecasts how that state evolves under interventions, supporting downstream tasks from imaging-based diagnosis to digital-twin treatment planning. Two failure modes threaten the reliability of such models in clinical deployment: (i)~\emph{covariate shift}, because training data are fragme…
▽ More
Medical world models aim to learn a latent state of patient or organ physiology and a transition function that forecasts how that state evolves under interventions, supporting downstream tasks from imaging-based diagnosis to digital-twin treatment planning. Two failure modes threaten the reliability of such models in clinical deployment: (i)~\emph{covariate shift}, because training data are fragmented across hospitals, scanners, and time, so the feature distribution seen by the latent-dynamics predictor differs across fragments and from the distribution at deployment; and (ii)~\emph{confidence misalignment}, because multi-step forecasts are often overconfident exactly where clinical risk is highest. We argue that both problems admit a unified treatment via a single lightweight regularisation objective, \textbf{CalTwin}, which combines a Fisher-Information-based shift penalty adapted from our prior work on fragmented covariate-shift remediation~\cite{khan2025mitigating,khan2025causal} with a Confidence Misalignment Penalty adapted from our prior work on calibrated vision-language classification~\cite{khan2025confidence}, applied here to a GRU-based medical world model's latent transition predictor. We derive the combined objective, establish which proof steps transfer from the classification setting without modification and which require adaptation, and evaluate it on the PhysioNet 2019 Sepsis Challenge, treating the two hospital systems as sequential training fragments and the unseen system as an out-of-distribution test. CalTwin reduces OOD next-step latent-state MSE by 9.1\% relative to the no-penalty baseline (FIM penalty alone accounts for 7.0\%); the ECE reduction from the Confidence Misalignment Penalty is real but small (0.7\% for CalTwin, 1.3\% for CMP alone).
△ Less
Submitted 29 July, 2026;
originally announced July 2026.
-
UnDA: Unpaired Domain Alignment for Cross-Modal Knowledge Transfer in Medical Imaging
Authors:
Rafsan Jany,
Shadab Tanjeed Ahmad,
Ahsan Bulbul,
Tahsinul Islam,
Md Azam Hossain,
Abu Raihan Mostofa Kamal
Abstract:
Multimodal based approaches often outperform single modality approaches in downstream tasks as the different modalities provide complementary information, yet acquiring paired clinical data remains a significant challenge in real world scenarios. While cross-modal knowledge distillation addresses this, existing methods often struggle with large modality gaps and the propagation of noise from uncer…
▽ More
Multimodal based approaches often outperform single modality approaches in downstream tasks as the different modalities provide complementary information, yet acquiring paired clinical data remains a significant challenge in real world scenarios. While cross-modal knowledge distillation addresses this, existing methods often struggle with large modality gaps and the propagation of noise from uncertain source-domain predictions. To overcome these challenges, we propose UnDA, an anchor-guided framework for unpaired cross-modal distillation. Our approach introduces a backbone-agnostic Alignment Module that extracts semantically structured class tokens via an attention based pooling mechanism. To ensure robust knowledge transfer, we propose Uncertainty-Weighted Optimal Transport (UCT-OT), which dynamically weights feature-level alignment based on prediction confidence, effectively suppressing noisy supervision. Furthermore, a per-class ProtoNCE objective maintains stable prototype memories to enforce global discriminability across unpaired batches. Evaluations on representative segmentation tasks under strictly unpaired settings show consistent improvements in accuracy and boundary precision in the target modality, demonstrating that meaningful structural knowledge can be transferred across heterogeneous data sources without paired datasets.
△ Less
Submitted 23 July, 2026;
originally announced July 2026.
-
The Second LoViF 2026 Challenge on Real-World All-in-One Image Restoration: Methods and Results
Authors:
Xiang Chen,
Hao Li,
Jiangxin Dong,
Jinshan Pan,
Xin Li,
Hongbo Ding,
Junpeng Jiang,
Xingyu Qiu,
Yilian Zhong,
Yuxiang Chen,
Shibo Yin,
Zixuan Huang,
Yushun Fang,
Xilei Zhu,
Yahui Wang,
Chen Lu,
Xiaodong Zhou,
Qingyue Cao,
Changwei Gong,
Jingyun Liu,
Xingchen Yi,
Hansen Shi,
Ruiyi Liu,
Jirui Xie,
Tao Liu
, et al. (67 additional authors not shown)
Abstract:
This paper presents a review of the second LoViF Challenge on Real-World All-in-One Image Restoration. The challenge aims to advance unified image restoration under diverse real-world degradation conditions, including blur, low-light, haze, rain, and snow. It provides a common benchmark for evaluating the restoration accuracy, robustness, and generalization capability of models across multiple deg…
▽ More
This paper presents a review of the second LoViF Challenge on Real-World All-in-One Image Restoration. The challenge aims to advance unified image restoration under diverse real-world degradation conditions, including blur, low-light, haze, rain, and snow. It provides a common benchmark for evaluating the restoration accuracy, robustness, and generalization capability of models across multiple degradation categories within a unified framework. The competition attracted 158 registered participants, and 20 teams were included in the final ranking after their submitted results were successfully reproduced and verified. This report provides a comprehensive analysis of the submitted solutions and corresponding results, highlighting recent advances in real-world all-in-one image restoration. The summarized methods and empirical findings reveal effective design strategies and establish an updated benchmark for future research in real-world low-level vision.
△ Less
Submitted 23 July, 2026;
originally announced July 2026.
-
SpEmoC: A Balanced Speaker-Segment Multimodal Emotion Benchmark
Authors:
Sania Bano,
Shahzad Ahmad,
Santosh Kumar Vipparthi,
Sukalpa Chanda,
Subrahmanyam Murala
Abstract:
Understanding human emotions in spoken conversations is a key challenge in affective computing, with applications in empathetic AI, human computer interaction, and mental health monitoring. However, existing datasets vary in scale, emotion distribution, modality alignment, and data partitioning strategies, which can influence reliable cross-dataset generalization and minority-emotion modeling. We…
▽ More
Understanding human emotions in spoken conversations is a key challenge in affective computing, with applications in empathetic AI, human computer interaction, and mental health monitoring. However, existing datasets vary in scale, emotion distribution, modality alignment, and data partitioning strategies, which can influence reliable cross-dataset generalization and minority-emotion modeling. We introduce SpEmoC a Speaking segment Emotion for Conversations comprising 306,544 raw clips from 3,100 English language movies and TV series. From these, 30,000 high quality, class balanced clips are curated, featuring synchronized visual, audio, and textual modalities annotated for seven emotions through a hybrid pipeline that integrates pretrained models with human validation. SpEmoC uses strict movie- and series-level splits to prevent content overlap between split sets, allowing more reliable evaluation of model generalization. The dataset also maintains a near-balanced distribution across seven emotions, including minority classes such as Fear and Disgust, which supports more balanced learning across categories. Extensive experiments, including in-domain benchmarking, cross-dataset transfer, low-data training, class-imbalance analysis, and modality transfer show that balanced data and careful splitting lead to more stable performance across emotions when models are evaluated on other datasets. These results highlight the importance of dataset design for robust and transferable multimodal emotion recognition.
△ Less
Submitted 20 July, 2026;
originally announced July 2026.
-
The regional control of a fractional spatio-temporal SIR model
Authors:
Sofwah Ahmad
Abstract:
This paper investigates a regional optimal control problem for a nonlinear spatio-temporal epidemiological model involving fractional diffusion on a bounded domain. In this work, we consider a general form of disease transmission, and two types of control, vaccination and treatment. We establish the existence and uniqueness of global-in-time solutions to our proposed system. We also prove the exis…
▽ More
This paper investigates a regional optimal control problem for a nonlinear spatio-temporal epidemiological model involving fractional diffusion on a bounded domain. In this work, we consider a general form of disease transmission, and two types of control, vaccination and treatment. We establish the existence and uniqueness of global-in-time solutions to our proposed system. We also prove the existence of an optimal control that minimizes the infected individuals as well as the cost of vaccination and treatment. The necessary conditions for optimality are derived. We show that concentrating vaccination in a selected region of the domain can substantially mitigate the spread of the disease. Numerical simulations are given.
△ Less
Submitted 14 July, 2026;
originally announced July 2026.
-
Tellurium sublattice instability driven amorphization in the chalcogenide AgSbTe2 under pressure
Authors:
Baihong Sun,
Zihan Zhang,
Wei Luo,
Sergei Grazhdannikov,
Wenting Lu,
Shiyu Feng,
Haikai Zou,
Chenxin Wei,
Martin Kunz,
Hirokazu Kadobayashi,
Bihang Wang,
Azkar Saeed Ahmad,
Yaron Amouyal,
Rajeev Ahuja,
Elissaios Stavrou
Abstract:
Pressure provides a powerful thermodynamic route to access hidden structural states in functional materials, yet the microscopic origin of pressure-induced amorphization remains elusive in many complex chalcogenides. Here we report a detailed high-pressure structural study of AgSbTe2,combining synchrotron X-ray diffraction with density functional theory and molecular dynamics calculations up to 60…
▽ More
Pressure provides a powerful thermodynamic route to access hidden structural states in functional materials, yet the microscopic origin of pressure-induced amorphization remains elusive in many complex chalcogenides. Here we report a detailed high-pressure structural study of AgSbTe2,combining synchrotron X-ray diffraction with density functional theory and molecular dynamics calculations up to 60 GPa. We uncover a pressure-driven transformation from the ambient R-3m phase to a fully disordered cubic Im3m phase, through an extended intermediate amorphous state. Enthalpy calculations reveal a near-degeneracy between the R3m and Im3m structures over a broad pressure range, dictating amorphization. Contrary to previously speculated cation vacancies, the amorphization is governed by a pronounced displacement instability of the Te sublattice. Remarkably, the time dependent decompression pathway controls the final structural state, resulting in either amorphous (slow decompression) or fully crystalline (fast decompression) states, indicative of a strong counterintuitive kinetic effect.
△ Less
Submitted 5 July, 2026;
originally announced July 2026.
-
Length--Velocity Gauge Equivalence of Quantum Geometric Nonlinear Conductivity
Authors:
Shakeel Ahmad,
Fei Xue
Abstract:
Nonlinear transport has emerged as a sensitive probe of quantum geometry beyond the Berry-curvature physics of linear response. However, the intrinsic second-order dc response remains conceptually subtle: different quantum and semiclassical formulations can appear to give different static limits, with different assignments of Fermi sea and Fermi surface contributions. Here we resolve this ambiguit…
▽ More
Nonlinear transport has emerged as a sensitive probe of quantum geometry beyond the Berry-curvature physics of linear response. However, the intrinsic second-order dc response remains conceptually subtle: different quantum and semiclassical formulations can appear to give different static limits, with different assignments of Fermi sea and Fermi surface contributions. Here we resolve this ambiguity by developing a gauge-consistent density-matrix theory of intrinsic nonlinear conductivity in both the length gauge, where the electric field couples through the position operator, and the velocity gauge, where it enters through the vector potential. We show that the two gauges give the same adiabatic dc response when the same retarded continuation is used for all external frequencies and when the velocity gauge current includes all field-dependent vertices. The apparent Fermi sea terms cancel in the full expression, leaving a Fermi surface quantum geometric contribution determined by the band-normalized quantum metric. This result implies that a fully gapped insulator has no residual dc nonlinear Hall current in the adiabatic clean limit. The reactive part of the Fermi surface term agrees with the original semiclassical Berry-connection-polarizability response, while the dissipative Ohmic sector requires a more careful treatment of relaxation and impurity scattering. Our work establishes the length-velocity gauge equivalence for quantum geometric nonlinear response and provides a foundation for using nonlinear transport to probe magnetic quantum geometry, especially in PT-symmetric antiferromagnets.
△ Less
Submitted 28 June, 2026;
originally announced June 2026.
-
Female-RHINO: A Real-Time Scanner-Integrated Framework for Automated Quantitative Uterine MRI Analysis and Structured Reporting
Authors:
Deepak Bhatia,
Saad Ahmad,
Smiti Tripathy,
Maria Camila Bustos Vivas,
Lieselotte Kratzsch,
Anika Knupfer,
Jordina Aviles Verdera,
Susanne Schulz-Heise,
Matthias May,
Jana Hutter
Abstract:
Standardized assessment of uterine MRI remains challenging due to anatomical variability, observer dependence, and the lack of workflow-integrated automated analysis tools. This work presents Female-RHINO: (R)eproductive (H)ealth (I)maging A(N)alysis T(O)ol, a real-time AI-assisted framework for automated quantitative uterine MRI analysis and structured reporting during image acquisition. We prese…
▽ More
Standardized assessment of uterine MRI remains challenging due to anatomical variability, observer dependence, and the lack of workflow-integrated automated analysis tools. This work presents Female-RHINO: (R)eproductive (H)ealth (I)maging A(N)alysis T(O)ol, a real-time AI-assisted framework for automated quantitative uterine MRI analysis and structured reporting during image acquisition. We present an end-to-end system that integrates inline communication with the MRI scanner and deep learning-based analysis to derive quantitative uterine biomarkers from sagittal T2-weighted pelvic MRI. The framework combines segmentation and anatomical landmark detection models trained and evaluated on more than 500 multi-center datasets spanning diverse protocols, vendors, and patient populations. It performs volumetry, detects and quantifies common incidental findings such as fibroids and Nabothian cysts, and extracts six anatomical landmarks for biometric assessment. Results are compiled into a structured clinician-oriented report with integrated visualizations, without manual interaction. Evaluation on independent retrospective and prospective cohorts demonstrated robust performance across varying acquisition settings. Mean Dice similarity coefficients were 0.82 for the uterus and 0.80 for fibroids, with lower but consistent agreement for Nabothian cysts. Landmark detection achieved a mean radial error of 3.7 mm. End-to-end processing was completed in under 70 seconds, enabling availability of results during the ongoing scan. Prospective deployment yielded immediate, standardized, and reproducible analyses supported by inter-observer agreement. The proposed system enables real-time scanner-integrated AI for automated uterine MRI analysis and reporting, with potential to improve standardization, efficiency, and clinical workflow in pelvic imaging.
△ Less
Submitted 23 June, 2026;
originally announced June 2026.
-
A Mixed-Reality Testbed for Autonomous Vehicles
Authors:
H. M. Sabbir Ahmad,
Ehsan Sabouni,
Emrullah Celik,
Zean Wan,
Damola Ajeyemi,
Christos G. Cassandras,
Wenchao Li
Abstract:
We propose a mixed-reality, hardware-in-the-loop (HIL) testbed for autonomous vehicles that seamlessly integrates a physical testbed of mobile robots with a high-fidelity simulation environment. The virtual simulation enables the creation of diverse, safety-critical driving scenarios to validate state-of-the-art perception, planning, and control algorithms, while augmenting simulations with physic…
▽ More
We propose a mixed-reality, hardware-in-the-loop (HIL) testbed for autonomous vehicles that seamlessly integrates a physical testbed of mobile robots with a high-fidelity simulation environment. The virtual simulation enables the creation of diverse, safety-critical driving scenarios to validate state-of-the-art perception, planning, and control algorithms, while augmenting simulations with physical robots equipped with multimodal sensors in photorealistic virtual environments further facilitating rigorous validation. Our testbed also features vehicular connectivity using wireless communication and can accommodate a large number of agents through the combination of physical robots and virtual simulated agents, supporting research on multi-agent systems including Connected and Autonomous Vehicles (CAVs). Finally, we present a safety-guaranteed framework combining perception, planning and a novel online learning-based controller using Control Barrier Functions (CBFs) for CAVs. Experiments using the proposed framework are used to validate and demonstrate the key functionalities and the overall utility of the testbed to bridge the gap between simulation and real-world hardware deployment.
△ Less
Submitted 17 June, 2026;
originally announced June 2026.
-
SegTME-UNI2: A Foundation Model-Based Framework for Generalisable Multiclass Cell Segmentation and LLM-Driven Tumour Microenvironment Characterisation in Histopathology
Authors:
Wan Siti Halimatul Munirah Wan Ahmad,
Faris Syahmi Samidi,
Mohammad Badal Ahmmed,
Vimal Angela Thiviyanathan,
Selvam Thavaraj,
Anwar P. P. Abdul Majeed
Abstract:
Characterising the tumour microenvironment (TME) from routine H&E-stained histology images requires simultaneous cell segmentation, feature extraction, and interpretable clinical reporting. We present SegTME-UNI2, a unified framework addressing these requirements. Its core is UNI2-UperHoVeR, a dual-head segmentation model pairing the UNI2-h pathology foundation model (ViT-Giant, pretrained on >100…
▽ More
Characterising the tumour microenvironment (TME) from routine H&E-stained histology images requires simultaneous cell segmentation, feature extraction, and interpretable clinical reporting. We present SegTME-UNI2, a unified framework addressing these requirements. Its core is UNI2-UperHoVeR, a dual-head segmentation model pairing the UNI2-h pathology foundation model (ViT-Giant, pretrained on >100M tiles from 100K slides) with two parallel UperNet decoders: one for six-class semantic segmentation and one for horizontal-vertical gradient regression enabling watershed-based nuclear instance separation. To address the lack of pixel-level annotations in large real-world repositories, UNI2-UperHoVeR undergoes a three-stage progressive pseudo-label curriculum. Each stage trains a fresh model without weight transfer, driving improvement entirely via increased pseudo-label quality: Stage 1: Uses human-annotated PanNuke (7,901 images, 189,744 nuclei, 0.25 um/pixel). Stage 2: Uses entropy-filtered pseudo-labels from the Stage 1 model on 271,711 TCGA-UT scale-0 patches (0.5 um/pixel). Stage 3: Uses pseudo-labels from the Stage 2 model on all 1,608,060 TCGA-UT patches across six resolution scales (0.5-1.0 um/pixel). Segmentation outputs feed a structured TME feature extraction pipeline computing 20+ per-patch compositional, morphological, spatial entropy, and intercellular distance metrics. These are encoded as JSON and passed to a fine-tuned NVIDIA BioNeMo GPT model to generate clinically interpretable TME narratives. Preliminary validation on held-out PanNuke and TCGA-UT partitions demonstrates framework feasibility and internal consistency. The pseudo-labelled TCGA-UT dataset and UNI2-UperHoVeR checkpoint are publicly released to support large-scale TME profiling and spatial biology research.
△ Less
Submitted 21 June, 2026; v1 submitted 16 June, 2026;
originally announced June 2026.
-
SPARK: Spatial Policy-driven Adaptive Reinforcement learning for Knowledge distillation
Authors:
Mohamed Jismy Aashik Rasool,
Shabir Ahmad,
Gisong Oh,
Teag Kuen Whangbo
Abstract:
Low-bit quantization enables deployment of image restoration (IR) networks on resource-constrained devices, but introduces rounding noise that disproportionately degrades high-frequency regions such as edges and fine textures. Existing knowledge distillation (KD) methods apply distillation signals uniformly across all spatial locations, overlooking the varying reconstruction difficulty across imag…
▽ More
Low-bit quantization enables deployment of image restoration (IR) networks on resource-constrained devices, but introduces rounding noise that disproportionately degrades high-frequency regions such as edges and fine textures. Existing knowledge distillation (KD) methods apply distillation signals uniformly across all spatial locations, overlooking the varying reconstruction difficulty across image regions. To address this, we propose SPARK (Spatial Policy-driven Adaptive Reinforcement Learning for Knowledge Distillation), a framework that adaptively allocates distillation effort using a lightweight reinforcement learning (RL) policy network. At each training step, a difficulty feature extractor computes four signals, namely Laplacian variance, pixel variance, student reconstruction error, and teacher-student knowledge gap, which are fed into a compact policy CNN that produces a stochastic spatial weight map to modulate the KD loss during quantization-aware training (QAT). SPARK is IR task-agnostic, adds no inference cost, and integrates into any existing QAT pipeline without architectural changes. Experiments on benchmark datasets demonstrate that SPARK consistently outperforms PTQ, QAT, and state-of-the-art (SOTA) KD approaches across multiple student architectures, achieving reconstruction quality closest to the full-precision teacher under significant computational constraints.
△ Less
Submitted 13 June, 2026;
originally announced June 2026.
-
AI for Maritime Security: Comparative Evaluation of CNN and Vision Transformer Architectures for Maritime Object Detection
Authors:
Ismet Gocer,
Zakirul Bhuiayn,
Shakeel Ahmad,
Raza Hasan
Abstract:
This study aims to enhance maritime security by using advanced Artificial Intelligence (AI) and Computer Vision (CV) techniques. For this purpose, it was designed and assessed intelligent object detection systems that can detect the presence of ships on the sea surface under different real-time environments. To achieve this goal, a maritime image dataset with 6,468 images was used, covering differ…
▽ More
This study aims to enhance maritime security by using advanced Artificial Intelligence (AI) and Computer Vision (CV) techniques. For this purpose, it was designed and assessed intelligent object detection systems that can detect the presence of ships on the sea surface under different real-time environments. To achieve this goal, a maritime image dataset with 6,468 images was used, covering different weather conditions like cloudy, foggy, rainy, and sunny environments. Six deep learning architectures were evaluated, including a base Convolutional Neural Network (CNN) model, four transfer learning models (Xception, VGG16, MobileNetV2, and EfficientNetV2L), and a Vision Transformer (ViT) model. The models were compared using multiple performance indicators, including accuracy, Type I and Type II errors, model size, and video processing time. The results show that model performance varies depending on computational constraints and deployment conditions. While lightweight architectures are suitable for resource-limited devices, the ViT achieved the best overall performance, reaching 100% accuracy with the lowest error rates and the fastest video processing time. The findings highlight the potential of AI-driven computer vision systems for maritime surveillance, border protection, and autonomous navigation.
△ Less
Submitted 28 May, 2026;
originally announced June 2026.
-
Measurement of the muon neutrino charged-current cross section with SND@LHC
Authors:
The SND@LHC Collaboration,
:,
D. Abbaneo,
S. Ahmad,
R. Albanese,
A. Alexandrov,
F. Alicante,
F. Aloschi,
K. Androsov,
L. G. Arellano,
C. Asawatangtrakuldee,
M. A. Ayala Torres,
N. Bangaru,
C. Battilana,
A. Bay,
A. Bersani,
C. Betancourt,
D. Bick,
R. Biswas,
A. Blanco Castro,
V. Boccia,
M. Bogomilov,
D. Bonacorsi,
W. M. Bonivento,
P. Bordalo
, et al. (142 additional authors not shown)
Abstract:
We report a measurement of the muon neutrino charged-current (CC) interaction cross section on tungsten using the electronic detectors of the SND@LHC experiment at the CERN Large Hadron Collider. The analysis uses proton--proton collision data at a centre-of-mass energy of $\sqrt{s} = 13.6$ TeV, corresponding to an integrated luminosity of $68.6 ~\text{fb}^{-1}$ collected during LHC Run 3 in 2022…
▽ More
We report a measurement of the muon neutrino charged-current (CC) interaction cross section on tungsten using the electronic detectors of the SND@LHC experiment at the CERN Large Hadron Collider. The analysis uses proton--proton collision data at a centre-of-mass energy of $\sqrt{s} = 13.6$ TeV, corresponding to an integrated luminosity of $68.6 ~\text{fb}^{-1}$ collected during LHC Run 3 in 2022 and 2023. A total of 31 $ν_μ$ CC candidates are selected against an expected background of $5.0 \pm 1.1$ events, consistent with a signal expectation of $24^{+10}_{-9}$ events. The signal strength is measured to be $\hatμ = 1.09^{+0.72}_{-0.37}$, and the combined muon neutrino and anti-neutrino CC cross section on tungsten is determined to be $σ(ν_μ+ \barν_μ) = (37^{+24}_{-12})\times 10^{-35}~\text{cm}^2$ at a median energy of $228$ GeV. In addition, a calorimetric measurement of the hadronic energies of the neutrino candidate events is performed, making use of calibration data from dedicated test-beam campaigns.
△ Less
Submitted 12 June, 2026;
originally announced June 2026.
-
TABVERSE: Benchmarking Cross-Format Table Understanding in LLMs and VLMs
Authors:
Momina Ahsan,
Sarfraz Ahmad,
Ming Shan Hee,
Roy Ka-Wei Lee,
Preslav Nakov
Abstract:
Large Language Models (LLMs) and Vision-Language Models (VLMs) are increasingly evaluated on table reasoning tasks, but the role of table representation remains under-explored. In practice, the same table content may appear in different structural formats, such as HTML, Markdown, and LaTeX, or as rendered images. However, existing evaluations often let content, format, layout, and modality vary to…
▽ More
Large Language Models (LLMs) and Vision-Language Models (VLMs) are increasingly evaluated on table reasoning tasks, but the role of table representation remains under-explored. In practice, the same table content may appear in different structural formats, such as HTML, Markdown, and LaTeX, or as rendered images. However, existing evaluations often let content, format, layout, and modality vary together, making it difficult to isolate representation effects. We introduce TABVERSE, a controlled multimodal table benchmark that aligns the same table content across multiple structural formats and rendered images, with question category and difficulty tags. This design enables systematic evaluation of representation effects while holding table content fixed. We evaluate LLMs and VLMs across three tasks: Question Answering (QA), Structural Understanding Capability (SUC), and Structure Reconstruction (SR). Our results show that representation choice substantially affects table understanding. Models generally perform better with structured text than with rendered images, but the size of this gap depends on the task, model, and format. HTML is often the most robust text format, while row-sensitive structural tasks and syntactically usable LaTeX reconstruction remain challenging. These findings show that table representation is a key factor in reliable table evaluation.
△ Less
Submitted 8 June, 2026;
originally announced June 2026.
-
UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding
Authors:
Ahmer Tabassum,
Sarfraz Ahmad,
Hasan Iqbal,
Owais Aijaz,
Momina Ahsan,
Preslav Nakov
Abstract:
Meaningful multilingual evaluation must test models in the target language and educational context. Urdu, spoken by more than 230 million people, lacks a broad MMLU-style benchmark built from native educational sources. We introduce UrduMMLU, a benchmark of 26,431 Urdu MCQs across 26 subjects and five domains, collected from native Urdu MCQ banks and public examination PDFs. Unlike translation-bas…
▽ More
Meaningful multilingual evaluation must test models in the target language and educational context. Urdu, spoken by more than 230 million people, lacks a broad MMLU-style benchmark built from native educational sources. We introduce UrduMMLU, a benchmark of 26,431 Urdu MCQs across 26 subjects and five domains, collected from native Urdu MCQ banks and public examination PDFs. Unlike translation-based resources, UrduMMLU covers both standard academic subjects and Urdu- and region-specific content. We label the exam-derived portion through dual human annotation with strict consensus filtering. We evaluate 30 LLMs under English and Urdu prompts, yielding 60 zero-shot evaluations, and further evaluate four open-source LLMs under multiple few-shot settings across both prompt languages. Gemini-3.5-Flash performs best, reaching 90.20% and 90.34% accuracy, while no other model exceeds 85%. The strongest open-source model trails by 7.79 and 8.92 points, and many models lose 25 to 40 points on Urdu-centered Humanities subjects compared with STEM. Few-shot prompting yields only modest gains. UrduMMLU shows that Urdu knowledge remains uneven in current LLMs, especially for regionally grounded content.
△ Less
Submitted 5 June, 2026;
originally announced June 2026.
-
Quantum Algorithm for Distributed Reduction of Entanglements (QADR): A Trainable and Simulation-Efficient QML Framework
Authors:
Syed Farhan Ahmad,
Gregory T. Byrd
Abstract:
Training Variational Quantum Circuits (VQCs) under Noisy Intermediate-Scale Quantum (NISQ) constraints introduces severe computational limitations: classical statevector simulation memory scales exponentially ($\mathcal{O}(2^n)$), and global cost functions suffer from barren plateaus where gradient variance decays exponentially ($\mathcal{O}(1/2^n)$). This paper introduces and evaluates the Quantu…
▽ More
Training Variational Quantum Circuits (VQCs) under Noisy Intermediate-Scale Quantum (NISQ) constraints introduces severe computational limitations: classical statevector simulation memory scales exponentially ($\mathcal{O}(2^n)$), and global cost functions suffer from barren plateaus where gradient variance decays exponentially ($\mathcal{O}(1/2^n)$). This paper introduces and evaluates the Quantum Algorithm for Distributed Reduction of Entanglements (QADR), a hybrid quantum-classical machine learning framework that decomposes a global $n$-qubit VQC into localized sub-circuits operating approximately within the causal light cones of individual target qubits. QADR reduces classical simulation memory scaling from $\mathcal{O}(2^n)$ to $\mathcal{O}(n \cdot 2^{2d+1})$ for a light cone radius $d$, while naturally mitigating global barren plateaus. We benchmark QADR against standard global VQCs, Support Vector Machines (SVM), and two customized classical parameter-matched neural networks (CANN and PMNN) on the MNIST dataset and the high-dimensional NASA IMS wind turbine drivetrain diagnostic task. QADR demonstrates excellent scalability, operating successfully at $n_{\text{features}}=2000$ where standard global VQCs crash due to memory exhaustion, while matching or exceeding the performance of optimized classical architectures.
△ Less
Submitted 31 May, 2026;
originally announced June 2026.
-
A Novel Evaluation Metric for Unsupervised Learning in AIS-Based Maritime Anomaly Detection: MADQI
Authors:
Ismet Gocer,
Zakirul Bhuiyan,
Raza Hasan,
Shakeel Ahmad
Abstract:
This paper introduces a new systematic framework for detecting anomalies in maritime Automatic Identification System (AIS) datasets. These anomalies include abnormal vessel behaviours related to speed, position jumps, time gaps, and turn angles. Although unsupervised learning algorithms such as Isolation Forest are widely used for detecting anomalous vessel movements, they often lack systematic an…
▽ More
This paper introduces a new systematic framework for detecting anomalies in maritime Automatic Identification System (AIS) datasets. These anomalies include abnormal vessel behaviours related to speed, position jumps, time gaps, and turn angles. Although unsupervised learning algorithms such as Isolation Forest are widely used for detecting anomalous vessel movements, they often lack systematic and meaningful evaluation measures. To address this limitation, we propose a novel quality metric called Maritime Anomaly Detection Quality Index (MADQI). The prosed MADQI is a composite index designed to evaluate the anomaly detection performance of machine learning models without requiring labelled data. The proposed framework uses Haversine distance calculations to analyse AIS datasets and identify anomalies based on their spatial and behavioural characteristics. The proposed MADQI evaluation framework integrates four interconnected metrics: Anomaly Rate Consistency (ARC), Physical Plausibility Score (PPS), Score Distribution Separation (SDS), and Extreme Case Evidence (ECE). These metrics are combined through automatic normalisation using multi-chunk evaluation and adaptive scaling techniques. Experimental results on the AIS dataset show that the proposed framework achieved a MADQI score of 80.37%, demonstrating its effectiveness for unsupervised anomaly detection. In particular, the algorithm performed strongly in identifying abnormal vessel behaviour. Among the individual MADQI components, ECE and ARC achieved scores of 0.907 and 1.000, respectively, indicating excellent capability in detecting extreme anomalies and maintaining anomaly rate consistency. Overall, these results are encouraging and demonstrate that the proposed framework provides a reliable and meaningful approach for evaluating unsupervised anomaly detection in maritime AIS data.
△ Less
Submitted 28 May, 2026;
originally announced May 2026.
-
AfriScience-MT: Towards Decolonizing Science in Africa through Text Translation
Authors:
Idris Abdulmumin,
Tajuddeen Gwadabe,
Shamsuddeen Hassan Muhammad,
David Ifeoluwa Adelani,
Nomonde Khalo,
Ibrahim Said Ahmad,
Abiodun Modupe,
Anina Mumm,
Sibusiso Biyela,
Michelle Rabie,
Johanna Havemann,
Marek Rei,
Jade Abbott,
Vukosi Marivate
Abstract:
The dominance of colonial languages in African education and scientific communication limits how hundreds of millions of speakers of African languages access and produce scientific knowledge. A core obstacle is the lack of established scientific terminology in these languages. We introduce AfriScience-MT, a parallel corpus covering six African languages (Amharic, Hausa, Luganda, Northern Sotho, Yo…
▽ More
The dominance of colonial languages in African education and scientific communication limits how hundreds of millions of speakers of African languages access and produce scientific knowledge. A core obstacle is the lack of established scientific terminology in these languages. We introduce AfriScience-MT, a parallel corpus covering six African languages (Amharic, Hausa, Luganda, Northern Sotho, Yorùbá, and isiZulu) across 11 scientific domains. Professional translators, working with expert science communicators, translated plain-language summaries of scientific papers into each target language and created new terms where none existed. We benchmark machine translation systems and large language models in zero-shot, few-shot, and fine-tuned settings. Our results show that closed-source models outperform all open-source models at both the sentence and document levels: GPT-5.4 and Gemini-3.1-Flash-Lite lead with average sentence-level COMET scores of 68.3 and 68.0, respectively, and tie at an average document-level COMET of 48.3. Among open systems, fine-tuned NLLB-1.3B reaches 67.3 at the sentence level, and TranslateGemma-12B reaches 44.0 at the document level with 1-shot in-context learning. We release AfriScience-MT to support benchmarking and document-level scientific MT for African languages.
△ Less
Submitted 28 May, 2026;
originally announced May 2026.
-
Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South
Authors:
Charvi Rastogi,
Mukul Bhutani,
Minsuk Kahng,
Shamsuddeen Hassan Muhammad,
Evgeniia Razumovskaia,
Priyanka Suresh,
Ibrahim Said Ahmad,
Charu Kalia,
Yaaseen Mahomed,
Madhurima Maji,
Minjae Lee,
Alicia Parrish,
Jessica Quaye,
Vijay Janapa Reddi,
Aishwarya Verma,
Lora Aroyo
Abstract:
Despite the global deployment of text-to-image (T2I) models, their safety frameworks are largely calibrated to a Western-centric default, creating significant vulnerabilities for the rest of the world. To embrace cultural pluralism and bring historically under-represented perspectives in T2I safety, we conduct localised community-centered red teaming studies in the Global South. Our two-fold appro…
▽ More
Despite the global deployment of text-to-image (T2I) models, their safety frameworks are largely calibrated to a Western-centric default, creating significant vulnerabilities for the rest of the world. To embrace cultural pluralism and bring historically under-represented perspectives in T2I safety, we conduct localised community-centered red teaming studies in the Global South. Our two-fold approach prioritizes localization and participation, by focusing on secondary urban centers in these regions, and conducting community engagement and training workshops to contextualize local norms. As a result, we present PLACES, a dataset comprising over 26,000 examples of T2I model failures collected in partnership with universities in Ghana, Nigeria, and two regions of India (Karnataka and Punjab). Analysis of prompts collected reveals a wide-ranging diversity in socio-cultural and linguistic attributes, when compared to existing geography-agnostic crowdsourced red-teaming data. We observe unique adversarial patterns enabled by local cultural and linguistic nuances, and distinct clusters within region around specific themes, such as religion in India. Moreover, we uncover structural contextual gaps in existing safety frameworks by identifying novel harms showing normative dissonance (e.g., violating religious norms, ignoring local customs, and ominous symbolism). This work argues that expanding T2I safety requires moving beyond mere scale to incorporate deeply localised, participatory methodologies for data collection and contextualization. Content warning: This paper includes examples containing potentially harmful or offensive content.
△ Less
Submitted 18 May, 2026;
originally announced May 2026.
-
TriALS: Triphasic-Aided Liver Lesion Segmentation Benchmark in Non-Contrast CT
Authors:
Marawan Elbatel,
Mohamed Ghonim,
Jiaji Mao,
Zhuosheng Lin,
Katharina Eckstein,
Andrés Martínez Mora,
Jonathan Deissler,
Maximilian Rokuss,
Constantin Ulrich,
Zdravko Marinov,
Wenhui Deng,
Baoxun Li,
Huijun Hu,
Jun Shen,
Mohanad Ghonim,
Khadiga Omar Nassar,
Mariam Elbakry,
Menna Dyab,
Amr Muhammad Abdo Salem,
Nouran Elghitany,
Noha Elghitany,
Yi Qin,
Xuanqi Huang,
Haonan Wang,
Shao-Woo Yen
, et al. (40 additional authors not shown)
Abstract:
Automated segmentation of liver lesions on non-contrast computed tomography (NCCT) is clinically important but fundamentally challenging, particularly in low-resource settings across Africa and Asia where contrast agents are frequently unavailable. Progress has been limited by the absence of annotated NCCT benchmarks. Here we describe the TriALS challenge for automated liver lesion segmentation un…
▽ More
Automated segmentation of liver lesions on non-contrast computed tomography (NCCT) is clinically important but fundamentally challenging, particularly in low-resource settings across Africa and Asia where contrast agents are frequently unavailable. Progress has been limited by the absence of annotated NCCT benchmarks. Here we describe the TriALS challenge for automated liver lesion segmentation under contrast-limited conditions, supported by a multi-centre dataset of 150 cases with four-phase CT acquisitions (600 volumes) from Egyptian and Chinese institutions. Algorithms were evaluated on 70 cases from three institutions, including an independent external cohort. The top-performing method achieved a mean venous-phase Dice of 0.754, consistent with human-level performance, yet dropped to 0.57 on NCCT. On external validation, the leading method outperformed off-the-shelf models by up to 28% in Dice on NCCT. Algorithm performance was most strongly predicted by training data scale and pre-training strategy. A cross-year comparison exposed a persistent perceptual barrier on NCCT that scaling pre-training alone cannot overcome. Data, annotations, and code are available at https://github.com/xmed-lab/TriALS.
△ Less
Submitted 15 May, 2026;
originally announced May 2026.
-
On Some Properties of LCM-Lattices of Edge Ideals of k-Uniform Hypergraphs
Authors:
Muneeba Mansha,
Sarfraz Ahmad
Abstract:
In this article, we investigate the combinatorial and algebraic properties of the lcm-lattice associated with the edge ideal of a hypergraph. Let $\H$ be a hypergraph, $I(\H)$ its corresponding edge ideal in a polynomial ring in $n$ variables, and $\mathrm{Icm}(I(\H))$ the associated lcm-lattice. We establish conditions under which the lcm-lattice of an edge ideal is Boolean, modular, or complemen…
▽ More
In this article, we investigate the combinatorial and algebraic properties of the lcm-lattice associated with the edge ideal of a hypergraph. Let $\H$ be a hypergraph, $I(\H)$ its corresponding edge ideal in a polynomial ring in $n$ variables, and $\mathrm{Icm}(I(\H))$ the associated lcm-lattice. We establish conditions under which the lcm-lattice of an edge ideal is Boolean, modular, or complemented. Furthermore, we extend these results to the case of the product of lcm-lattices in the complemented case. Additionally, we study the effects of polarization on the lcm-lattices of $I(\H)$ and its polarized ideal.
△ Less
Submitted 13 May, 2026;
originally announced May 2026.
-
FIND: Toward Multimodal Financial Reasoning and Question Answering for Indic Languages
Authors:
Sarmistha Das,
Vaibhav Vishal,
Syed Ibrahim Ahmad,
Manish Gupta,
Sriparna Saha
Abstract:
Financial decision-making in multilingual settings demands accurate numerical reasoning grounded in diverse modalities, yet existing benchmarks largely overlook this high-stakes, real-world challenge, especially for Indic languages. We introduce FinVQA, a benchmark for evaluating financial numerical and multimodal reasoning in multilingual Indic contexts. FinVQA spans English, Hindi, Bengali, Mara…
▽ More
Financial decision-making in multilingual settings demands accurate numerical reasoning grounded in diverse modalities, yet existing benchmarks largely overlook this high-stakes, real-world challenge, especially for Indic languages. We introduce FinVQA, a benchmark for evaluating financial numerical and multimodal reasoning in multilingual Indic contexts. FinVQA spans English, Hindi, Bengali, Marathi, Gujarati, and Tamil, and comprises 18,900 samples across 14 financial domains. The dataset captures diverse reasoning paradigms under realistic constraints, and is structured across three difficulty levels (easy, moderate, hard) and four question formats: multiple choice, fill-in-the-blank, table matching, and true/false. To address these challenges, we propose FIND, a framework that combines supervised fine-tuning with constraint-aware decoding to promote faithful numerical reasoning, robust multimodal grounding, and structured decision-making. Together, FinVQA and FIND establish a rigorous evaluation and modeling paradigm for high-stakes multilingual multimodal financial reasoning.
△ Less
Submitted 13 May, 2026;
originally announced May 2026.
-
Beyond Majority Voting: Agreement-Based Clustering to Model Annotator Perspectives in Subjective NLP Tasks
Authors:
Tadesse Destaw Belay,
Ibrahim Said Ahmad,
Idris Abdulmumin,
Abinew Ali Ayele,
Alexander Gelbukh,
Eusebio Ricárdez-Vázquez,
Olga Kolesnikova,
Shamsuddeen Hassan Muhammad,
Seid Muhie Yimam
Abstract:
Disagreement in annotation is a common phenomenon in the development of NLP datasets and serves as a valuable source of insight. While majority voting remains the dominant strategy for aggregating labels, recent work has explored modeling individual annotators to preserve their perspectives. However, modeling each annotator is resource-intensive and remains underexplored across various NLP tasks.…
▽ More
Disagreement in annotation is a common phenomenon in the development of NLP datasets and serves as a valuable source of insight. While majority voting remains the dominant strategy for aggregating labels, recent work has explored modeling individual annotators to preserve their perspectives. However, modeling each annotator is resource-intensive and remains underexplored across various NLP tasks. We propose an agreement-based clustering technique to model the disagreement between the annotators. We conduct comprehensive experiments in 40 datasets in 18 typologically diverse languages, covering three subjective NLP tasks: sentiment analysis, emotion classification, and hate speech detection. We evaluate four aggregation approaches: majority vote, ensemble, multi-label, and multitask. The results demonstrate that agreement-based clustering can leverage the full spectrum of annotator perspectives and significantly enhance classification performance in subjective NLP tasks compared to majority voting and individual annotator modeling. Regarding the aggregation approach, the multi-label and multitask approaches are better for modeling clustered annotators than an ensemble and model majority vote.
△ Less
Submitted 11 May, 2026;
originally announced May 2026.
-
REPTILES: Repeated Tiles of Sargantana, a RISC-V multicore based on OpenPiton
Authors:
Noelia Oliete-Escuín,
Arnau Bigas,
Narcís Rodas,
Albert Aguilera,
Sajjad Ahmad,
Jonathan Balkind,
Xavier Carril,
Max Doblas,
Ivan Díaz,
Roger Figueras,
Alireza Foroodnia,
Cesar Fuguet,
Ignacio Genovese,
Raúl Gilabert,
Abbas Haghi,
Alexander Kropotov,
Neiel Leyva,
Oscar Lostes-Cazorla,
Lorién López-Villellas,
Davy Million,
Alireza Monemi,
Sérik Pérez,
Juan Antonio Rodríguez,
Víctor Soria-Pardos,
Behzad Salami
, et al. (4 additional authors not shown)
Abstract:
Chip industry continues advancing and expanding modern computing systems, resulting in more complex multi-core processors. Conversely, academic projects face scalability challenges due to limited resources, highlighting the need for open-source frameworks that enable innovation and knowledge sharing. Recently, several open-source proposals have emerged, offering flexible and scalable designs, but…
▽ More
Chip industry continues advancing and expanding modern computing systems, resulting in more complex multi-core processors. Conversely, academic projects face scalability challenges due to limited resources, highlighting the need for open-source frameworks that enable innovation and knowledge sharing. Recently, several open-source proposals have emerged, offering flexible and scalable designs, but fail to meet the performance demands of modern High-Performance Computing (HPC) applications. In this project, we present REPTILES, an open-source RISC-V multicore framework based on OpenPiton\thanks. REPTILES interconnects multiple Sargantana cores with the memory hierarchy of OpenPiton. Moreover, we present the new features incorporated in Sargantana and OpenPiton designs to improve the performance of HPC applications. We demonstrate that REPTILES presents suitable scalability, achieving a speedup of 3.1x on average with 4 cores. Additionally, we show that Sargantana's new features increase the performance of vector addition benchmark in a 9.3x.
△ Less
Submitted 6 May, 2026;
originally announced May 2026.
-
Plausible Deniability in Fully Homomorphic Computation
Authors:
Shahzad Ahmad,
Stefan Rass,
Zahra Seyedi
Abstract:
We introduce \emph{Plausible Deniability in Fully Homomorphic Computation} (PD-FHC), a framework enabling users to outsource Boolean computations to an untrusted cloud while maintaining both computational privacy against honest-but-curious providers and plausible deniability against coercive adversaries. We define the notion of a \emph{Deniable Computation Medium} (DCM) and a \emph{Deniable Comput…
▽ More
We introduce \emph{Plausible Deniability in Fully Homomorphic Computation} (PD-FHC), a framework enabling users to outsource Boolean computations to an untrusted cloud while maintaining both computational privacy against honest-but-curious providers and plausible deniability against coercive adversaries. We define the notion of a \emph{Deniable Computation Medium} (DCM) and a \emph{Deniable Computation Scheme} (DCS) as medium-independent abstractions, then instantiate them using RGB images with Fredkin-gate circuits. One real circuit and several decoys share a single fixed Fredkin-gate wiring. Embedded control bits decide what each gate computes at each pixel, so the same wiring evaluates the real function at the real positions and decoy functions elsewhere. The cloud applies this one wiring to every pixel identically, processing all circuits in a single pass. Under coercion, the user reveals a decoy with verifiable results while the real circuit stays hidden. We formalize multi-round coercion games with existence and circuit-discovery advantages. For the image instantiation, we prove \emph{information-theoretic position privacy} under a \emph{matched-marginal condition}: when the real, decoy, and fill bits are drawn from a common per-position law and placed at random, the embedded LSB plane is exchangeable, so an honest-but-curious provider gains no advantage over guessing at locating the real positions, for any such law and not only the uniform one. We are explicit that this is a condition Alice enforces, that it is distinct from steganalytic undetectability, and that the latter requires the embedded law to match the declared service's legitimate-input law.
△ Less
Submitted 10 July, 2026; v1 submitted 3 May, 2026;
originally announced May 2026.
-
PPO guided Agentic Pipeline for Adaptive Prompt Selection and Test Case Generation
Authors:
Gourisetty Venkata Sai Koushik,
Dama Aditya,
Mahankali Harish Sai,
Peddi Siddarhta,
Shadab Ahmad,
Vivek Yelleti
Abstract:
Developing effective test cases capable of thoroughly exercising large-scale software systems is inherently difficult, especially if such systems have voluminous, complex, and deeply nested source codes. In this work, we present a novel approach for generating test cases using a reinforcement learning-driven agentic framework where Proximal Policy Optimization (PPO) is coupled with an LLM engine t…
▽ More
Developing effective test cases capable of thoroughly exercising large-scale software systems is inherently difficult, especially if such systems have voluminous, complex, and deeply nested source codes. In this work, we present a novel approach for generating test cases using a reinforcement learning-driven agentic framework where Proximal Policy Optimization (PPO) is coupled with an LLM engine to guide prompt selection during test generation. Our approach consists of two phases. In Phase I, the ToT-guided optimization agent partitions and minimizes the source code by removing redundancies without changing the functional behavior of the source code. In Phase II, a PPO-based policy network is trained to solve the problem of selecting prompts among eight different prompting techniques, such as Boundary Value Analysis, Random Fuzzing, etc., based on the inputted 11-dimensional state vector representing the source code complexity metrics and live coverage metrics to direct the LLM engine towards exploring unvisited paths in the program. The PPO agent receives rewards based on a combination of increases in line and branch coverages, penalties for unexplored branches, and rewards for reducing source code length. From experiments conducted on twenty benchmark programs, it is evident that the proposed approach, PPO-LLM, outperforms CBMC, kS-LLM, and kS-LLM++ in terms of branch and line coverage in almost all cases, for various loop bound values ranging from BOUND~1 to BOUND~2000. While at BOUND~1, the coverage of branches is 100\% using PPO-LLM on the PALS suite, in comparison, it is around 86.8\% using kS-LLM++. This confirms that adaptive prompt selection driven by PPO substantially outperforms static prompting strategies on PALS type programs.
△ Less
Submitted 1 May, 2026;
originally announced May 2026.
-
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
Authors:
Muhammad Dehan Al Kautsar,
Saeed Almheiri,
Momina Ahsan,
Bilal Elbouardi,
Younes Samih,
Sarfraz Ahmad,
Amr Keleg,
Omar El Herraoui,
Kareem Elzeky,
Abed Alhakim Freihat,
Mohamed Anwar,
Zhuohan Xie,
Junhong Liang,
Mohammad Rustom Al Nasar,
Preslav Nakov,
Fajri Koto
Abstract:
There is a significant gap in evaluating cultural reasoning in LLMs using conversational datasets that capture culturally rich and dialectal contexts. Most Arabic benchmarks focus on short text snippets in Modern Standard Arabic (MSA), overlooking the cultural nuances that naturally arise in dialogues. To address this gap, we introduce ArabCulture-Dialogue, a culturally grounded conversational dat…
▽ More
There is a significant gap in evaluating cultural reasoning in LLMs using conversational datasets that capture culturally rich and dialectal contexts. Most Arabic benchmarks focus on short text snippets in Modern Standard Arabic (MSA), overlooking the cultural nuances that naturally arise in dialogues. To address this gap, we introduce ArabCulture-Dialogue, a culturally grounded conversational dataset covering 13 Arabic-speaking countries, in both MSA and each country's respective dialect, spanning 12 daily-life topics and 54 fine-grained subtopics. We utilize the dataset to form three benchmarking tasks: (i) multiple-choice cultural reasoning, (ii) machine translation between MSA and dialects, and (iii) dialect-steering generation. Our experiments indicate that the performance gap between MSA and Arabic dialects still exists, whereby the models perform worse on all three tasks in the dialectal setup, compared to the MSA one.
△ Less
Submitted 30 April, 2026;
originally announced May 2026.
-
SpecRLBench: A Benchmark for Generalization in Specification-Guided Reinforcement Learning
Authors:
Zijian Guo,
İlker Işık,
H. M. Sabbir Ahmad,
Wenchao Li
Abstract:
Specification-guided reinforcement learning (RL) provides a principled framework for encoding complex, temporally extended tasks using formal specifications such as linear temporal logic (LTL). While recent methods have shown promising results, their ability to generalize across unseen specifications and diverse environments remains insufficiently understood. In this work, we introduce SpecRLBench…
▽ More
Specification-guided reinforcement learning (RL) provides a principled framework for encoding complex, temporally extended tasks using formal specifications such as linear temporal logic (LTL). While recent methods have shown promising results, their ability to generalize across unseen specifications and diverse environments remains insufficiently understood. In this work, we introduce SpecRLBench, a benchmark designed to evaluate the generalization capabilities of LTL-based specification-guided RL methods. The benchmark spans multiple difficulty levels across navigation and manipulation domains, incorporating both static and dynamic environments, diverse robot dynamics, and varied observation modalities. Through extensive empirical evaluation, we characterize the strengths and limitations of existing approaches and reveal the challenges that emerge as specification and environment complexity increase. SpecRLBench provides a structured platform for systematic comparison and supports the development of more generalizable specification-guided RL methods. Code is available at https://github.com/BU-DEPEND-Lab/SpecRLBench.
△ Less
Submitted 27 April, 2026;
originally announced April 2026.
-
A Stackelberg Model for Hybridization in Cryptography
Authors:
Willie Kouam,
Stefan Rass,
Zahra Seyedi,
Shahzad Ahmad,
Eckhard Pfluegel
Abstract:
Similar to a strategic interaction between rational and intelligent agents, cryptography problems can be examined through the prism of game theory. In this setting, the agent aiming to protect a message is called the defender, while the one attempting to decrypt it, generally for malicious purposes, is the attacker. To strengthen security in cryptography, various strategies have been developed, am…
▽ More
Similar to a strategic interaction between rational and intelligent agents, cryptography problems can be examined through the prism of game theory. In this setting, the agent aiming to protect a message is called the defender, while the one attempting to decrypt it, generally for malicious purposes, is the attacker. To strengthen security in cryptography, various strategies have been developed, among which hybridization stands out as a key concept in modern cryptographic design. This strategy allows the defender to select among different encryption algorithms (classical, post-quantum, or hybrid) while carefully balancing security and operational costs. On the other side, the attacker, limited by available resources, chooses cryptanalysis methods capable of breaching the selected algorithm. We model this interaction as a Stackelberg cryptographic hybridization problem under resource constraints. Here, the defender randomizes over encryption algorithms, and the attacker observes the choice before selecting suitable cryptanalysis methods. The attacker's decision is framed as a conditional optimization problem, which we refer to as the ``attacker subgame''. We then propose a dynamic programming approach for the attacker's subgame, while the defender's Stackelberg optimization is formulated as a linear program.
△ Less
Submitted 27 April, 2026; v1 submitted 23 April, 2026;
originally announced April 2026.
-
SAHM: A Benchmark for Arabic Financial and Shari'ah-Compliant Reasoning
Authors:
Rania Elbadry,
Sarfraz Ahmad,
Ahmed Heakl,
Dani Bouch,
Momina Ahsan,
Muhra AlMahri,
Marwa Elsaid khalil,
Yuxia Wang,
Salem Lahlou,
Sophia Ananiadou,
Veselin Stoyanov,
Jimin Huang,
Xueqing Peng,
Preslav Nakov,
Zhuohan Xie
Abstract:
English financial NLP has advanced rapidly through benchmarks targeting earnings analysis, market sentiment, tabular reasoning, and financial question answering, yet Arabic financial NLP remains virtually nonexistent, despite 422 million speakers, $4.9 trillion in Gulf sovereign wealth, and a $4-5 trillion Islamic finance industry requiring specialized Shari'ah compliance over instruments like suk…
▽ More
English financial NLP has advanced rapidly through benchmarks targeting earnings analysis, market sentiment, tabular reasoning, and financial question answering, yet Arabic financial NLP remains virtually nonexistent, despite 422 million speakers, $4.9 trillion in Gulf sovereign wealth, and a $4-5 trillion Islamic finance industry requiring specialized Shari'ah compliance over instruments like sukuk, murabaha, and takaful. We introduce Sahm, the first Arabic financial benchmark spanning seven tasks: AAOIFI standards QA, fatwa-based QA/MCQ, accounting and business exams, financial sentiment analysis, extractive summarization, and event-cause reasoning, comprising 14,380 expert-verified instances from authentic regulatory, juristic, and corporate sources. Evaluating 20 LLMs, we find Arabic fluency does not imply financial reasoning: models achieving 91% on recognition tasks drop sharply on generation, and event-cause reasoning exposes the widest performance gap (1.89-9.84/10). We release the benchmark and dataset to support trustworthy Arabic financial assistants.
△ Less
Submitted 30 April, 2026; v1 submitted 21 April, 2026;
originally announced April 2026.
-
The Fourth Challenge on Image Super-Resolution ($\times$4) at NTIRE 2026: Benchmark Results and Method Overview
Authors:
Zheng Chen,
Kai Liu,
Jingkai Wang,
Xianglong Yan,
Jianze Li,
Ziqing Zhang,
Jue Gong,
Jiatong Li,
Lei Sun,
Xiaoyang Liu,
Radu Timofte,
Yulun Zhang,
Jihye Park,
Yoonjin Im,
Hyungju Chun,
Hyunhee Park,
MinKyu Park,
Zheng Xie,
Xiangyu Kong,
Weijun Yuan,
Zhan Li,
Qiurong Song,
Luen Zhu,
Fengkai Zhang,
Xinzhe Zhu
, et al. (128 additional authors not shown)
Abstract:
This paper presents the NTIRE 2026 image super-resolution ($\times$4) challenge, one of the associated competitions of the NTIRE 2026 Workshop at CVPR 2026. The challenge aims to reconstruct high-resolution (HR) images from low-resolution (LR) inputs generated through bicubic downsampling with a $\times$4 scaling factor. The objective is to develop effective super-resolution solutions and analyze…
▽ More
This paper presents the NTIRE 2026 image super-resolution ($\times$4) challenge, one of the associated competitions of the NTIRE 2026 Workshop at CVPR 2026. The challenge aims to reconstruct high-resolution (HR) images from low-resolution (LR) inputs generated through bicubic downsampling with a $\times$4 scaling factor. The objective is to develop effective super-resolution solutions and analyze recent advances in the field. To reflect the evolving objectives of image super-resolution, the challenge includes two tracks: (1) a restoration track, which emphasizes pixel-wise fidelity and ranks submissions based on PSNR; and (2) a perceptual track, which focuses on visual realism and evaluates results using a perceptual score. A total of 194 participants registered for the challenge, with 31 teams submitting valid entries. This report summarizes the challenge design, datasets, evaluation protocol, main results, and methods of participating teams. The challenge provides a unified benchmark and offers insights into current progress and future directions in image super-resolution.
△ Less
Submitted 15 April, 2026;
originally announced April 2026.
-
SemEval-2026 Task 9: Detecting Multilingual, Multicultural and Multievent Online Polarization
Authors:
Usman Naseem,
Robert Geislinger,
Juan Ren,
Sarah Kohail,
Rudy Garrido Veliz,
P Sam Sahil,
Yiran Zhang,
Marco Antonio Stranisci,
Idris Abdulmumin,
Özge Alaçam,
Cengiz Acartürk,
Aisha Jabr,
Saba Anwar,
Abinew Ali Ayele,
Elena Tutubalina,
Aung Kyaw Htet,
Xintong Wang,
Surendrabikram Thapa,
Tanmoy Chakraborty,
Dheeraj Kodati,
Sahar Moradizeyveh,
Firoj Alam,
Ye Kyaw Thu,
Shantipriya Parida,
Ihsan Ayyub Qazi
, et al. (9 additional authors not shown)
Abstract:
We present SemEval-2026 Task 9, a shared task on online polarization detection, covering 22 languages and comprising over 110K annotated instances. Each data instance is multi-labeled with the presence of polarization, polarization type, and polarization manifestation. Participants were asked to predict labels in three sub-tasks: (1) detecting the presence of polarization, (2) identifying the type…
▽ More
We present SemEval-2026 Task 9, a shared task on online polarization detection, covering 22 languages and comprising over 110K annotated instances. Each data instance is multi-labeled with the presence of polarization, polarization type, and polarization manifestation. Participants were asked to predict labels in three sub-tasks: (1) detecting the presence of polarization, (2) identifying the type of polarization, and (3) recognizing the polarization manifestation. The three tasks attracted over 1,000 participants worldwide and more than 10k submission on Codabench. We received final submissions from 67 teams and 73 system description papers. We report the baseline results and analyze the performance of the best-performing systems, highlighting the most common approaches and the most effective methods across different subtasks and languages. The dataset of this task is publicly available.
△ Less
Submitted 1 July, 2026; v1 submitted 8 April, 2026;
originally announced April 2026.
-
Multi-Robot Multi-Queue Control via Exhaustive Assignment Actor-Critic Learning
Authors:
Mohammad Merati,
H. M. Sabbir Ahmad,
Wenchao Li,
David Castañón
Abstract:
We study online task allocation for multi-robot, multi-queue systems with asymmetric stochastic arrivals and switching delays. We formulate the problem in discrete time: each location can host at most one robot per slot, servicing a task consumes one slot, switching between locations incurs a one-slot travel delay, and arrivals at locations are independent Bernoulli processes with heterogeneous ra…
▽ More
We study online task allocation for multi-robot, multi-queue systems with asymmetric stochastic arrivals and switching delays. We formulate the problem in discrete time: each location can host at most one robot per slot, servicing a task consumes one slot, switching between locations incurs a one-slot travel delay, and arrivals at locations are independent Bernoulli processes with heterogeneous rates. Building on our previous structural result that optimal policies are of exhaustive type, we formulate a discounted-cost Markov decision process and develop an exhaustive-assignment actor-critic policy architecture that enforces exhaustive service by construction and learns only the next-queue allocation for idle robots. Unlike the exhaustive-serve-longest (ESL) queue rule, whose optimality is known only under symmetry, the proposed policy adapts to asymmetry in arrival rates. Across different server-location ratios, loads, and asymmetric arrival profiles, the proposed policy consistently achieves lower discounted holding cost and smaller mean queue length than the ESL baseline, while remaining near-optimal on instances where an optimal benchmark is available. These results show that structure-aware actor-critic methods provide an effective approach for real-time multi-robot scheduling.
△ Less
Submitted 4 April, 2026;
originally announced April 2026.
-
Fanar 2.0: Arabic Generative AI Stack
Authors:
FANAR TEAM,
Ummar Abbas,
Mohammad Shahmeer Ahmad,
Minhaj Ahmad,
Abdulaziz Al-Homaid,
Anas Al-Nuaimi,
Enes Altinisik,
Ehsaneddin Asgari,
Sanjay Chawla,
Shammur Chowdhury,
Fahim Dalvi,
Kareem Darwish,
Nadir Durrani,
Mohamed Elfeky,
Ahmed Elmagarmid,
Mohamed Eltabakh,
Asim Ersoy,
Masoomali Fatehkia,
Mohammed Qusay Hashim,
Majd Hawasly,
Mohamed Hefeeda,
Mus'ab Husaini,
Keivin Isufaj,
Soon-Gyo Jung,
Houssam Lachemat
, et al. (12 additional authors not shown)
Abstract:
We present Fanar 2.0, the second generation of Qatar's Arabic-centric Generative AI platform. Sovereignty is a first-class design principle: every component, from data pipelines to deployment infrastructure, was designed and operated entirely at QCRI, Hamad Bin Khalifa University. Fanar 2.0 is a story of resource-constrained excellence: the effort ran on 256 NVIDIA H100 GPUs, with Arabic having on…
▽ More
We present Fanar 2.0, the second generation of Qatar's Arabic-centric Generative AI platform. Sovereignty is a first-class design principle: every component, from data pipelines to deployment infrastructure, was designed and operated entirely at QCRI, Hamad Bin Khalifa University. Fanar 2.0 is a story of resource-constrained excellence: the effort ran on 256 NVIDIA H100 GPUs, with Arabic having only ~0.5% of web data despite 400 million native speakers. Fanar 2.0 adopts a disciplined strategy of data quality over quantity, targeted continual pre-training, and model merging to achieve substantial gains within these constraints. At the core is Fanar-27B, continually pre-trained from a Gemma-3-27B backbone on a curated corpus of 120 billion high-quality tokens across three data recipes. Despite using 8x fewer pre-training tokens than Fanar 1.0, it delivers substantial benchmark improvements: Arabic knowledge (+9.1 pts), language (+7.3 pts), dialects (+3.5 pts), and English capability (+7.6 pts). Beyond the core LLM, Fanar 2.0 introduces a rich stack of new capabilities. FanarGuard is a state-of-the-art 4B bilingual moderation filter for Arabic safety and cultural alignment. The speech family Aura gains a long-form ASR model for hours-long audio. Oryx vision family adds Arabic-aware image and video understanding alongside culturally grounded image generation. An agentic tool-calling framework enables multi-step workflows. Fanar-Sadiq utilizes a multi-agent architecture for Islamic content. Fanar-Diwan provides classical Arabic poetry generation. FanarShaheen delivers LLM-powered bilingual translation. A redesigned multi-layer orchestrator coordinates all components through intent-aware routing and defense-in-depth safety validation. Taken together, Fanar 2.0 demonstrates that sovereign, resource-constrained AI development can produce systems competitive with those built at far greater scale.
△ Less
Submitted 17 March, 2026;
originally announced March 2026.
-
TinyNav: End-to-End TinyML for Real-Time Autonomous Navigation on Microcontrollers
Authors:
Pooria Roy,
Nourhan Jadallah. Tomer Lapid,
Shahzaib Ahmad,
Armita Afroushe,
Mete Bayrak
Abstract:
Autonomous navigation typically relies on power-intensive processors, limiting accessibility in low-cost robotics. Although microcontrollers offer a resource-efficient alternative, they impose strict constraints on model complexity. We present TinyNav, an end-to-end TinyML system for real-time autonomous navigation on an ESP32 microcontroller. A custom-trained, quantized 2D convolutional neural ne…
▽ More
Autonomous navigation typically relies on power-intensive processors, limiting accessibility in low-cost robotics. Although microcontrollers offer a resource-efficient alternative, they impose strict constraints on model complexity. We present TinyNav, an end-to-end TinyML system for real-time autonomous navigation on an ESP32 microcontroller. A custom-trained, quantized 2D convolutional neural network processes a 20-frame sliding window of depth data to predict steering and throttle commands. By avoiding 3D convolutions and recurrent layers, the 23k-parameter model achieves 30 ms inference latency. Correlation analysis and Grad-CAM validation indicate consistent spatial awareness and obstacle avoidance behavior. TinyNav demonstrates that responsive autonomous control can be deployed directly on highly constrained edge devices, reducing reliance on external compute resources.
△ Less
Submitted 10 March, 2026;
originally announced March 2026.
-
Post-Quantum Sanitizable Signatures from McEliece-Based Chameleon Hashing
Authors:
Shahzad Ahmad,
Stefan Rass,
Zahra Seyedi
Abstract:
We introduce a novel post-quantum sanitizable signature scheme constructed upon a chameleon hash function derived from the McEliece cryptosystem. In this design, the designated sanitizer possesses the inherent trapdoor of a Goppa code, which facilitates controlled collision-finding via Patterson decoding. This mechanism enables authorized modification of specific message blocks while ensuring all…
▽ More
We introduce a novel post-quantum sanitizable signature scheme constructed upon a chameleon hash function derived from the McEliece cryptosystem. In this design, the designated sanitizer possesses the inherent trapdoor of a Goppa code, which facilitates controlled collision-finding via Patterson decoding. This mechanism enables authorized modification of specific message blocks while ensuring all other content remains immutably bound. We provide formal security definitions and rigorous proofs of existential unforgeability and immutability, grounded in the hardness of syndrome decoding in the random-oracle model, where a robust random oracle thwarts trivial linear hash collisions. A key innovation lies in our precise characterization of the transparency property: by imposing a specific weight constraint on the randomizers generated by the signer, we achieve perfect transparency, rendering sanitized signatures indistinguishable from freshly signed ones. This work establishes the first transparent, code-based, post-quantum sanitizable signature scheme, offering strong theoretical guarantees and a pathway for practical deployment in long-term secure applications.
△ Less
Submitted 24 February, 2026;
originally announced February 2026.
-
VLM-UQBench: A Benchmark for Modality-Specific and Cross-Modality Uncertainties in Vision Language Models
Authors:
Chenyu Wang,
Tianle Chen,
H. M. Sabbir Ahmad,
Kayhan Batmanghelich,
Wenchao Li
Abstract:
Uncertainty quantification (UQ) is vital for ensuring that vision-language models (VLMs) behave safely and reliably. A central challenge is to localize uncertainty to its source, determining whether it arises from the image, the text, or misalignment between the two. We introduce VLM-UQBench, a benchmark for modality-specific and cross-modal data uncertainty in VLMs, It consists of 600 real-world…
▽ More
Uncertainty quantification (UQ) is vital for ensuring that vision-language models (VLMs) behave safely and reliably. A central challenge is to localize uncertainty to its source, determining whether it arises from the image, the text, or misalignment between the two. We introduce VLM-UQBench, a benchmark for modality-specific and cross-modal data uncertainty in VLMs, It consists of 600 real-world samples drawn from the VizWiz dataset, curated into clean, image-, text-, and cross-modal uncertainty subsets, and a scalable perturbation pipeline with 8 visual, 5 textual, and 3 cross-modal perturbations. We further propose two simple metrics that quantify the sensitivity of UQ scores to these perturbations and their correlation with hallucinations, and use them to evaluate a range of UQ methods across four VLMs and three datasets. Empirically, we find that: (i) existing UQ methods exhibit strong modality-specific specialization and substantial dependence on the underlying VLM, (ii) modality-specific uncertainty frequently co-occurs with hallucinations while current UQ scores provide only weak and inconsistent risk signals, and (iii) although UQ methods can rival reasoning-based chain-of-thought baselines on overt, group-level ambiguity, they largely fail to detect the subtle, instance-level ambiguity introduced by our perturbation pipeline. These results highlight a significant gap between current UQ practices and the fine-grained, modality-aware uncertainty required for reliable VLM deployment.
△ Less
Submitted 9 February, 2026;
originally announced February 2026.
-
Trojans in Artificial Intelligence (TrojAI) Final Report
Authors:
Kristopher W. Reese,
Taylor Kulp-McDowall,
Michael Majurski,
Tim Blattner,
Derek Juba,
Peter Bajcsy,
Antonio Cardone,
Philippe Dessauw,
Alden Dima,
Anthony J. Kearsley,
Melinda Kleczynski,
Joel Vasanth,
Walid Keyrouz,
Chace Ashcraft,
Neil Fendley,
Ted Staley,
Trevor Stout,
Josh Carney,
Greg Canal,
Will Redman,
Aurora Schmidt,
Cameron Hickert,
William Paul,
Jared Markowitz,
Nathan Drenkow
, et al. (46 additional authors not shown)
Abstract:
The Intelligence Advanced Research Projects Activity (IARPA) launched the TrojAI program to confront an emerging vulnerability in modern artificial intelligence: the threat of AI Trojans. These AI trojans are malicious, hidden backdoors intentionally embedded within an AI model that can cause a system to fail in unexpected ways, or allow a malicious actor to hijack the AI model at will. This multi…
▽ More
The Intelligence Advanced Research Projects Activity (IARPA) launched the TrojAI program to confront an emerging vulnerability in modern artificial intelligence: the threat of AI Trojans. These AI trojans are malicious, hidden backdoors intentionally embedded within an AI model that can cause a system to fail in unexpected ways, or allow a malicious actor to hijack the AI model at will. This multi-year initiative helped to map out the complex nature of the threat, pioneered foundational detection methods, and identified unsolved challenges that require ongoing attention by the burgeoning AI security field. This report synthesizes the program's key findings, including methodologies for detection through weight analysis and trigger inversion, as well as approaches for mitigating Trojan risks in deployed models. Comprehensive test and evaluation results highlight detector performance, sensitivity, and the prevalence of "natural" Trojans. The report concludes with lessons learned and recommendations for advancing AI security research.
△ Less
Submitted 27 February, 2026; v1 submitted 6 February, 2026;
originally announced February 2026.
-
ReFORM: Reflected Flows for On-support Offline RL via Noise Manipulation
Authors:
Songyuan Zhang,
Oswin So,
H. M. Sabbir Ahmad,
Eric Yang Yu,
Matthew Cleaveland,
Mitchell Black,
Chuchu Fan
Abstract:
Offline reinforcement learning (RL) aims to learn the optimal policy from a fixed dataset generated by behavior policies without additional environment interactions. One common challenge that arises in this setting is the out-of-distribution (OOD) error, which occurs when the policy leaves the training distribution. Prior methods penalize a statistical distance term to keep the policy close to the…
▽ More
Offline reinforcement learning (RL) aims to learn the optimal policy from a fixed dataset generated by behavior policies without additional environment interactions. One common challenge that arises in this setting is the out-of-distribution (OOD) error, which occurs when the policy leaves the training distribution. Prior methods penalize a statistical distance term to keep the policy close to the behavior policy, but this constrains policy improvement and may not completely prevent OOD actions. Another challenge is that the optimal policy distribution can be multimodal and difficult to represent. Recent works apply diffusion or flow policies to address this problem, but it is unclear how to avoid OOD errors while retaining policy expressiveness. We propose ReFORM, an offline RL method based on flow policies that enforces the less restrictive support constraint by construction. ReFORM learns a behavior cloning (BC) flow policy with a bounded source distribution to capture the support of the action distribution, then optimizes a reflected flow that generates bounded noise for the BC flow while keeping the support, to maximize the performance. Across 40 challenging tasks from the OGBench benchmark with datasets of varying quality and using a constant set of hyperparameters for all tasks, ReFORM dominates all baselines with hand-tuned hyperparameters on the performance profile curves.
△ Less
Submitted 4 February, 2026;
originally announced February 2026.
-
Vision-Language Models on the Edge for Real-Time Robotic Perception
Authors:
Sarat Ahmad,
Maryam Hafeez,
Syed Ali Raza Zaidi
Abstract:
Vision-Language Models (VLMs) enable multimodal reasoning for robotic perception and interaction, but their deployment in real-world systems remains constrained by latency, limited onboard resources, and privacy risks of cloud offloading. Edge intelligence within 6G, particularly Open RAN and Multi-access Edge Computing (MEC), offers a pathway to address these challenges by bringing computation cl…
▽ More
Vision-Language Models (VLMs) enable multimodal reasoning for robotic perception and interaction, but their deployment in real-world systems remains constrained by latency, limited onboard resources, and privacy risks of cloud offloading. Edge intelligence within 6G, particularly Open RAN and Multi-access Edge Computing (MEC), offers a pathway to address these challenges by bringing computation closer to the data source. This work investigates the deployment of VLMs on ORAN/MEC infrastructure using the Unitree G1 humanoid robot as an embodied testbed. We design a WebRTC-based pipeline that streams multimodal data to an edge node and evaluate LLaMA-3.2-11B-Vision-Instruct deployed at the edge versus in the cloud under real-time conditions. Our results show that edge deployment preserves near-cloud accuracy while reducing end-to-end latency by 5\%. We further evaluate Qwen2-VL-2B-Instruct, a compact model optimized for resource-constrained environments, which achieves sub-second responsiveness, cutting latency by more than half but at the cost of accuracy.
△ Less
Submitted 21 January, 2026;
originally announced January 2026.
-
Serverless AI Security: Attack Surface Analysis and Runtime Protection Mechanisms for FaaS-Based Machine Learning
Authors:
Chetan Pathade,
Vinod Dhimam,
Sheheryar Ahmad,
Ilsa Lareb
Abstract:
Serverless computing has achieved widespread adoption, with over 70% of AWS organizations using serverless solutions [1]. Meanwhile, machine learning inference workloads increasingly migrate to Function-as-a-Service (FaaS) platforms for their scalability and cost-efficiency [2], [3], [4]. However, this convergence introduces critical security challenges, with recent reports showing a 220% increase…
▽ More
Serverless computing has achieved widespread adoption, with over 70% of AWS organizations using serverless solutions [1]. Meanwhile, machine learning inference workloads increasingly migrate to Function-as-a-Service (FaaS) platforms for their scalability and cost-efficiency [2], [3], [4]. However, this convergence introduces critical security challenges, with recent reports showing a 220% increase in AI/ML vulnerabilities [5] and serverless computing's fragmented architecture raises new security concerns distinct from traditional cloud deployments [6], [7]. This paper presents the first comprehensive security analysis of machine learning workloads in serverless environments. We systematically characterize the attack surface across five categories: function-level vulnerabilities (cold start exploitation, dependency poisoning), model-specific threats (API-based extraction, adversarial inputs), infrastructure attacks (cross-function contamination, privilege escalation), supply chain risks (malicious layers, backdoored libraries), and IAM complexity (ephemeral nature, serverless functions). Through empirical assessments across AWS Lambda, Azure Functions, and Google Cloud Functions, we demonstrate real-world attack scenarios and quantify their security impact. We propose Serverless AI Shield (SAS), a multi-layered defense framework providing pre-deployment validation, runtime monitoring, and post-execution forensics. Our evaluation shows SAS achieves 94% detection rates while maintaining performance overhead below 9% for inference latency. We release an open-source security toolkit to enable practitioners to assess and harden their serverless AI deployments, advancing the field toward more resilient cloud-native machine learning systems.
△ Less
Submitted 15 January, 2026;
originally announced January 2026.
-
A Parallel Cross-Lingual Benchmark for Multimodal Idiomaticity Understanding
Authors:
Dilara Torunoğlu-Selamet,
Dogukan Arslan,
Rodrigo Wilkens,
Wei He,
Doruk Eryiğit,
Thomas Pickard,
Adriana S. Pagano,
Aline Villavicencio,
Gülşen Eryiğit,
Ágnes Abuczki,
Aida Cardoso,
Alesia Lazarenka,
Dina Almassova,
Amalia Mendes,
Anna Kanellopoulou,
Antoni Brosa-Rodríguez,
Baiba Saulite,
Beata Wojtowicz,
Bolette Pedersen,
Carlos Manuel Hidalgo-Ternero,
Chaya Liebeskind,
Danka Jokić,
Diego Alves,
Eleni Triantafyllidi,
Erik Velldal
, et al. (53 additional authors not shown)
Abstract:
Potentially idiomatic expressions (PIEs) construe meanings inherently tied to the everyday experience of a given language community. As such, they constitute an interesting challenge for assessing the linguistic (and to some extent cultural) capabilities of NLP systems. In this paper, we present XMPIE, a parallel multilingual and multimodal dataset of potentially idiomatic expressions. The dataset…
▽ More
Potentially idiomatic expressions (PIEs) construe meanings inherently tied to the everyday experience of a given language community. As such, they constitute an interesting challenge for assessing the linguistic (and to some extent cultural) capabilities of NLP systems. In this paper, we present XMPIE, a parallel multilingual and multimodal dataset of potentially idiomatic expressions. The dataset, containing 34 languages and over ten thousand items, allows comparative analyses of idiomatic patterns among language-specific realisations and preferences in order to gather insights about shared cultural aspects. This parallel dataset allows to evaluate model performance for a given PIE in different languages and whether idiomatic understanding in one language can be transferred to another. Moreover, the dataset supports the study of PIEs across textual and visual modalities, to measure to what extent PIE understanding in one modality transfers or implies in understanding in another modality (text vs. image). The data was created by language experts, with both textual and visual components crafted under multilingual guidelines, and each PIE is accompanied by five images representing a spectrum from idiomatic to literal meanings, including semantically related and random distractors. The result is a high-quality benchmark for evaluating multilingual and multimodal idiomatic language understanding.
△ Less
Submitted 24 February, 2026; v1 submitted 13 January, 2026;
originally announced January 2026.
-
Thermodynamic Driving Force Activated Phonon Scattering in InN
Authors:
Zaheer Ahmad,
Osama A. Rana,
Shakeel Ahmad,
Mark Vernon,
Brendan Cross,
Alexander Kozhanov
Abstract:
Defect related disorder during InN growth is a major challenge for making high performance electronic and optoelectronic devices. This is partly because film quality is often described using reactor specific settings instead of general physical variables. In this study, we show that plasma assisted MOCVD growth of InN can be described using a single thermodynamic driving force coordinate. This coo…
▽ More
Defect related disorder during InN growth is a major challenge for making high performance electronic and optoelectronic devices. This is partly because film quality is often described using reactor specific settings instead of general physical variables. In this study, we show that plasma assisted MOCVD growth of InN can be described using a single thermodynamic driving force coordinate. This coordinate brings together growth kinetics, defect sensitive Raman response and structural coherence across different process conditions. When we use this coordinate, the incorporation rate follows a universal activated trend with a kinetic scale of about 0.08 eV. Raman measurements show a clear crossover between a defect sparse and a defect rich regime, a disorder activated Raman metric increases quickly after the crossover, while an A1-LO control metric stays mostly the same. This suggests that short range lattice disorder, not long range polar coupling, dominates the defect activation process. X-ray diffraction shows that the out of plane coherence length stays the same for samples with the same driving force, even if reactor settings are very different. This supports the idea that structural coherence is organized by thermodynamics in this growth window. Finally, a simple kinetic Monte Carlo model using driving force biased incorporation and defect activation events matches the observed exponential trends and the two regimes, supporting the driving force approach. These results show that a transferable driving force coordinate can be used for plasma assisted InN growth and offer a quantitative way to achieve defect sparse growth conditions.
△ Less
Submitted 12 January, 2026;
originally announced January 2026.
-
Continuing past the inner horizon using WKB
Authors:
Shadi Ali Ahmad,
Ahmed Almheiri,
Simon Lin
Abstract:
Features of the black hole interior can be extracted from the analytic structure of boundary correlation functions. Working in the geodesic approximation, we find analytic continuations that probe the interior of rotating and charged black holes. These generate contributions from timelike geodesics that thread the interior and emerge in a future universe. We implement these continuations on the mo…
▽ More
Features of the black hole interior can be extracted from the analytic structure of boundary correlation functions. Working in the geodesic approximation, we find analytic continuations that probe the interior of rotating and charged black holes. These generate contributions from timelike geodesics that thread the interior and emerge in a future universe. We implement these continuations on the momentum space two-point function and exemplify this in several black hole backgrounds. We also identify position space analytic continuations achieving the same task that incorporate different continued momentum space correlators. These correspond to non-perturbative corrections to the WKB approximation. We demonstrate this explicitly in the rotating BTZ black hole by showing that the interior geodesics contribute to the continued position space correlator and motivate a picture for how these contributions arise in higher dimensions. For AdS Schwarzschild, we identify the analytically continued solution that captures the bouncing geodesic. We discuss the possibility of using these continuations to probe the instability of inner horizons from the boundary.
△ Less
Submitted 5 January, 2026;
originally announced January 2026.
-
High-pressure structural and lattice-dynamics study of Yttria-Stabilized Zirconia
Authors:
Shennan Hu,
Baihong Sun,
Wenting Lu,
Shiyu Feng,
Bihan Wang,
Hirokazu Kadobayashi,
Yuzhu Wang,
Xingya Wang,
Lili Zhang,
Bora Kalkan,
Azkar Saeed Ahmad,
Elissaios Stavrou
Abstract:
The structural evolution of two selected compositions of Yttria-Stabilized Zirconia (YSZ), with 3mol% (3YSZ) and 8mol% (8YSZ) of Y2O3, have been investigated under pressure using in-situ synchrotron X-ray diffraction (XRD) and Raman spectroscopy in a diamond anvil cell up to 40 GPa (at room temperature).The close crystallographic relation between the observed structures and the relatively large di…
▽ More
The structural evolution of two selected compositions of Yttria-Stabilized Zirconia (YSZ), with 3mol% (3YSZ) and 8mol% (8YSZ) of Y2O3, have been investigated under pressure using in-situ synchrotron X-ray diffraction (XRD) and Raman spectroscopy in a diamond anvil cell up to 40 GPa (at room temperature).The close crystallographic relation between the observed structures and the relatively large difference in the atomic numbers of Y/Zr and O, imposes the simultaneous study using both techniques, aiming to fully elucidate the structural evolution under pressure. The results, by combining both techniques, reveal that for both 3YSZ and 8YSZ, pressure promotes higher-symmetry structures. Under initial compression, the minority at ambient conditions monoclinic phase (m-phase) gradually transforms towards t-phase, a transition that is concluded for both 3YSZ/8YSZ at ~10 GPa. At higher pressures, the solely remaining t-phase of 3YSZ transforms to the t'', that in turns transforms to the c-phase above 28 GPa. Likewise, for 8YSZ the coexistence of t- and t''-phases continue up to 31 GPa, where both transforms towards c-phase, that remains stable up to the highest pressure of this study. Upon pressure release, all observed transitions are fully reversible with negligible hysteresis, with the exception of the practical disappearance of the monoclinic phase at ambient conditions. Our study underscores the significance of simultaneously performing and analyzing the results of both XRD and Raman spectroscopy studies in relevant crystallographic systems. Moreover, it provides a route towards a ``structural purification'' of YSZ through the elimination of the m-phase aiming to improve material properties.
△ Less
Submitted 1 January, 2026;
originally announced January 2026.
-
Universal Aging Dynamics and Scaling Laws in Three-Dimensional Driven Granular Gases
Authors:
Rameez Farooq Shah,
Syed Rashid Ahmad
Abstract:
We establish universal scaling laws and quantify aging in three-dimensional uniformly heated hard sphere granular gases through large-scale event-driven molecular dynamics ($N=500{,}000$). We report three primary quantitative discoveries: (i) The characteristic energy decay time exhibits a universal inverse scaling $τ_0 \propto ε^{-1.03 \pm 0.02}$ with the dissipation parameter $ε= 1 - e^2$. (ii)…
▽ More
We establish universal scaling laws and quantify aging in three-dimensional uniformly heated hard sphere granular gases through large-scale event-driven molecular dynamics ($N=500{,}000$). We report three primary quantitative discoveries: (i) The characteristic energy decay time exhibits a universal inverse scaling $τ_0 \propto ε^{-1.03 \pm 0.02}$ with the dissipation parameter $ε= 1 - e^2$. (ii) The steady-state temperature follows a precise power-law $T_{\mathrm{steady}} \propto ε^{-1.51 \pm 0.03}$, reflecting the non-linear balance between thermostat heating and collisional dissipation. (iii) The velocity autocorrelation function $\bar{A}(τ_w, τ)$ demonstrates pronounced aging, with decay rates $λ$ following a power-law slowing down $λ(τ_w) \propto τ_w^{-0.82 \pm 0.05}$. These results establish the first 3D quantitative benchmarks for aging in driven dissipative gases, where near-Gaussian statistics persist despite extreme structural clustering.
△ Less
Submitted 29 December, 2025;
originally announced December 2025.
-
Benchmark Success, Clinical Failure: When Reinforcement Learning Optimizes for Benchmarks, Not Patients
Authors:
Armin Berger,
Manuela Bergau,
Helen Schneider,
Saad Ahmad,
Tom Anglim Lagones,
Gianluca Brugnara,
Martha Foltyn-Dumitru,
Kai Schlamp,
Philipp Vollmuth,
Rafet Sifa
Abstract:
Recent Reinforcement Learning (RL) advances for Large Language Models (LLMs) have improved reasoning tasks, yet their resource-constrained application to medical imaging remains underexplored. We introduce ChexReason, a vision-language model trained via R1-style methodology (SFT followed by GRPO) using only 2,000 SFT samples, 1,000 RL samples, and a single A100 GPU. Evaluations on CheXpert and NIH…
▽ More
Recent Reinforcement Learning (RL) advances for Large Language Models (LLMs) have improved reasoning tasks, yet their resource-constrained application to medical imaging remains underexplored. We introduce ChexReason, a vision-language model trained via R1-style methodology (SFT followed by GRPO) using only 2,000 SFT samples, 1,000 RL samples, and a single A100 GPU. Evaluations on CheXpert and NIH benchmarks reveal a fundamental tension: GRPO recovers in-distribution performance (23% improvement on CheXpert, macro-F1 = 0.346) but degrades cross-dataset transferability (19% drop on NIH). This mirrors high-resource models like NV-Reason-CXR-3B, suggesting the issue stems from the RL paradigm rather than scale. We identify a generalization paradox where the SFT checkpoint uniquely improves on NIH before optimization, indicating teacher-guided reasoning captures more institution-agnostic features. Furthermore, cross-model comparisons show structured reasoning scaffolds benefit general-purpose VLMs but offer minimal gain for medically pre-trained models. Consequently, curated supervised fine-tuning may outperform aggressive RL for clinical deployment requiring robustness across diverse populations.
△ Less
Submitted 2 January, 2026; v1 submitted 28 December, 2025;
originally announced December 2025.
-
Clustered Federated Learning with Hierarchical Knowledge Distillation
Authors:
Sabtain Ahmad,
Meerzhan Kanatbekova,
Ivona Brandic,
Atakan Aral
Abstract:
Clustered Federated Learning (CFL) has emerged as a powerful approach for addressing data heterogeneity and ensuring privacy in large distributed IoT environments. By clustering clients and training cluster-specific models, CFL enables personalized models tailored to groups of heterogeneous clients. However, conventional CFL approaches suffer from fragmented learning for training independent globa…
▽ More
Clustered Federated Learning (CFL) has emerged as a powerful approach for addressing data heterogeneity and ensuring privacy in large distributed IoT environments. By clustering clients and training cluster-specific models, CFL enables personalized models tailored to groups of heterogeneous clients. However, conventional CFL approaches suffer from fragmented learning for training independent global models for each cluster and fail to take advantage of collective cluster insights. This paper advocates a shift to hierarchical CFL, allowing bi-level aggregation to train cluster-specific models at the edge and a unified global model at the cloud. This shift improves training efficiency yet might introduce communication challenges. To this end, we propose CFLHKD, a novel personalization scheme for integrating hierarchical cluster knowledge into CFL. Built upon multi-teacher knowledge distillation, CFLHKD enables inter-cluster knowledge sharing while preserving cluster-specific personalization. CFLHKD adopts a bi-level aggregation to bridge the gap between local and global learning. Extensive evaluations of standard benchmark datasets demonstrate that CFLHKD outperforms representative baselines in cluster-specific and global model accuracy and achieves a performance improvement of 3.32-7.57\%.
△ Less
Submitted 11 December, 2025;
originally announced December 2025.
-
Pushing Tensor Accelerators Beyond MatMul in a User-Schedulable Language
Authors:
Yihong Zhang,
Derek Gerstmann,
Andrew Adams,
Maaz Bin Safeer Ahmad
Abstract:
Tensor accelerators now represent a growing share of compute resources in modern CPUs and GPUs. However, they are hard to program, leading developers to use vendor-provided kernel libraries that support tensor accelerators. As a result, the usage of tensor accelerators is limited to the provided interface, mainly designed for traditional ML and scientific computing workloads.
In this paper, we s…
▽ More
Tensor accelerators now represent a growing share of compute resources in modern CPUs and GPUs. However, they are hard to program, leading developers to use vendor-provided kernel libraries that support tensor accelerators. As a result, the usage of tensor accelerators is limited to the provided interface, mainly designed for traditional ML and scientific computing workloads.
In this paper, we show that tensor accelerators can improve the performance of applications beyond simple variants of MatMul. For example, many image processing pipelines are linear transformations over matrices in disguise and can therefore utilize such specialized hardware. This is nonetheless hindered by the difficulties in programming tensor accelerators. We tackle this problem with compiler-based techniques. We use the Halide user-schedulable language and express operations as Halide algorithms succinctly. To this end, we implement a flexible tensor instruction selector based on equality saturation. The tensor instruction selector supports both CPU- and GPU-attached tensor accelerators and works with existing scheduling operations (e.g., producer-consumer fusion). Together, this enables developers to write diverse accelerator-leveraging applications in a few dozen lines.
Using our system, we demonstrate the potential of tensor accelerators beyond their traditional domains. We implement several image processing pipelines (e.g., filtering, resampling, and denoising) in our system and evaluate them against non-accelerator-leveraging baselines. We show that these pipelines can achieve significant speedups. For example, a downsampling routine is sped up by $6.1\times$ by utilizing Tensor Cores on an Nvidia RTX 4070 GPU.
△ Less
Submitted 10 February, 2026; v1 submitted 1 December, 2025;
originally announced December 2025.