-
Magnetic resonance imaging assessment of the suitability and consistency of radiotherapy treatment positioning achieved using intra-oral stents
Authors:
Tanya Kairn,
Philip Chan,
Benjamin Chua,
Susannah Cleland,
Jodi Dawes,
Lizbeth Kenny,
Charles Y. Lin,
William R. McDowall,
Tania Poroa,
Scott B. Crowe
Abstract:
As head-and-neck radiotherapy treatments grow more complex and precise, it becomes increasingly important to assess the anatomical separations that can be achieved using intra-oral stents. A series of twenty T2-weighted turbo spin echo magnetic resonance images (MRI) were acquired of one healthy participant, with a range of different wax and 3D printed intra-oral stents in situ. The resulting meas…
▽ More
As head-and-neck radiotherapy treatments grow more complex and precise, it becomes increasingly important to assess the anatomical separations that can be achieved using intra-oral stents. A series of twenty T2-weighted turbo spin echo magnetic resonance images (MRI) were acquired of one healthy participant, with a range of different wax and 3D printed intra-oral stents in situ. The resulting measurements showed that a 3D printed modular stent containing hard polylactic acid (PLA) and flexible thermoplastic polyurethane (TPU) components made the largest and most reproducible separation between the cheeks (70.8 +/- 0.3 mm), two hard PLA stents designed to exactly fit the participant's teeth produced the poorest positioning reproducibility (standard deviations of up to 3 mm between a range of landmarks measured in repeated images). Most stents were described as ``comfortable'' although the wax stents left small pieces of wax attached to the teeth after use. This MRI based comparison demonstrated that the materials and designs used for intra-oral stents can have substantial effects on the level of anatomical separation and positioning reproducibility that they produce.
△ Less
Submitted 23 June, 2026;
originally announced June 2026.
-
Improving Methodologies for LLM Evaluations Across Global Languages
Authors:
Akriti Vij,
Benjamin Chua,
Darshini Ramiah,
En Qi Ng,
Mahran Morsidi,
Naga Nikshith Gangarapu,
Sharmini Johnson,
Vanessa Wilfred,
Vikneswaran Kumaran,
Wan Sie Lee,
Wenzhuo Yang,
Yongsen Zheng,
Bill Black,
Boming Xia,
Frank Sun,
Hao Zhang,
Qinghua Lu,
Suyu Ma,
Yue Liu,
Chi-kiu Lo,
Fatemeh Azadi,
Isar Nejadgholi,
Sowmya Vajjala,
Agnes Delaborde,
Nicolas Rolin
, et al. (21 additional authors not shown)
Abstract:
As frontier AI models are deployed globally, it is essential that their behaviour remains safe and reliable across diverse linguistic and cultural contexts. To examine how current model safeguards hold up in such settings, participants from the International Network for Advanced AI Measurement, Evaluation and Science, including representatives from Singapore, Japan, Australia, Canada, the EU, Fran…
▽ More
As frontier AI models are deployed globally, it is essential that their behaviour remains safe and reliable across diverse linguistic and cultural contexts. To examine how current model safeguards hold up in such settings, participants from the International Network for Advanced AI Measurement, Evaluation and Science, including representatives from Singapore, Japan, Australia, Canada, the EU, France, Kenya, South Korea and the UK conducted a joint multilingual evaluation exercise. Led by Singapore AISI, two open-weight models were tested across ten languages spanning high and low resourced groups: Cantonese English, Farsi, French, Japanese, Korean, Kiswahili, Malay, Mandarin Chinese and Telugu. Over 6,000 newly translated prompts were evaluated across five harm categories (privacy, non-violent crime, violent crime, intellectual property and jailbreak robustness), using both LLM-as-a-judge and human annotation.
The exercise shows how safety behaviours can vary across languages. These include differences in safeguard robustness across languages and harm types and variation in evaluator reliability (LLM-as-judge vs. human review). Further, it also generated methodological insights for improving multilingual safety evaluations, such as the need for culturally contextualised translations, stress-tested evaluator prompts and clearer human annotation guidelines. This work represents an initial step toward a shared framework for multilingual safety testing of advanced AI systems and calls for continued collaboration with the wider research community and industry.
△ Less
Submitted 22 January, 2026;
originally announced January 2026.
-
Improving Methodologies for Agentic Evaluations Across Domains: Leakage of Sensitive Information, Fraud and Cybersecurity Threats
Authors:
Ee Wei Seah,
Yongsen Zheng,
Naga Nikshith,
Mahran Morsidi,
Gabriel Waikin Loh Matienzo,
Nigel Gay,
Akriti Vij,
Benjamin Chua,
En Qi Ng,
Sharmini Johnson,
Vanessa Wilfred,
Wan Sie Lee,
Anna Davidson,
Catherine Devine,
Erin Zorer,
Gareth Holvey,
Harry Coppock,
James Walpole,
Jerome Wynee,
Magda Dubois,
Michael Schmatz,
Patrick Keane,
Sam Deverett,
Bill Black,
Bo Yan
, et al. (45 additional authors not shown)
Abstract:
The rapid rise of autonomous AI systems and advancements in agent capabilities are introducing new risks due to reduced oversight of real-world interactions. Yet agent testing remains nascent and is still a developing science. As AI agents begin to be deployed globally, it is important that they handle different languages and cultures accurately and securely.
To address this, participants from T…
▽ More
The rapid rise of autonomous AI systems and advancements in agent capabilities are introducing new risks due to reduced oversight of real-world interactions. Yet agent testing remains nascent and is still a developing science. As AI agents begin to be deployed globally, it is important that they handle different languages and cultures accurately and securely.
To address this, participants from The International Network for Advanced AI Measurement, Evaluation and Science, including representatives from Singapore, Japan, Australia, Canada, the European Commission, France, Kenya, South Korea, and the United Kingdom have come together to align approaches to agentic evaluations.
This is the third exercise, building on insights from two earlier joint testing exercises conducted by the Network in November 2024 and February 2025. The objective is to further refine best practices for testing advanced AI systems.
The exercise was split into two strands: (1) common risks, including leakage of sensitive information and fraud, led by Singapore AISI; and (2) cybersecurity, led by UK AISI. A mix of open and closed-weight models were evaluated against tasks from various public agentic benchmarks. Given the nascency of agentic testing, our primary focus was on understanding methodological issues in conducting such tests, rather than examining test results or model capabilities. This collaboration marks an important step forward as participants work together to advance the science of agentic evaluations.
△ Less
Submitted 22 January, 2026;
originally announced January 2026.
-
Carathéodory number of homogeneous convex cones
Authors:
Chek Beng Chua
Abstract:
We study the Carathéodory number of homogeneous convex cones via their spectrahedral representations. A characterization of homogeneous convex cones whose ranks match their Carathéodory numbers is given. This characterization is then used to show that a homogeneous convex cone is selfdual if and only if its rank matches the Carathéodory numbers of both its closure and its dual cone. It is further…
▽ More
We study the Carathéodory number of homogeneous convex cones via their spectrahedral representations. A characterization of homogeneous convex cones whose ranks match their Carathéodory numbers is given. This characterization is then used to show that a homogeneous convex cone is selfdual if and only if its rank matches the Carathéodory numbers of both its closure and its dual cone. It is further used to show that the only sparse spectrahedral cones that are homogeneous convex cones are those described by homogeneous chordal graphs.
△ Less
Submitted 6 July, 2026; v1 submitted 21 November, 2025;
originally announced November 2025.
-
Trusted Media Challenge Dataset and User Study
Authors:
Weiling Chen,
Sheng Lun Benjamin Chua,
Stefan Winkler,
See-Kiong Ng
Abstract:
The development of powerful deep learning technologies has brought about some negative effects to both society and individuals. One such issue is the emergence of fake media. To tackle the issue, we have organized the Trusted Media Challenge (TMC) to explore how Artificial Intelligence (AI) technologies could be leveraged to combat fake media. To enable further research, we are releasing the datas…
▽ More
The development of powerful deep learning technologies has brought about some negative effects to both society and individuals. One such issue is the emergence of fake media. To tackle the issue, we have organized the Trusted Media Challenge (TMC) to explore how Artificial Intelligence (AI) technologies could be leveraged to combat fake media. To enable further research, we are releasing the dataset that we had prepared from the TMC challenge, consisting of 4,380 fake and 2,563 real videos, with various video and/or audio manipulation methods employed to produce different types of fake media. All the videos in the TMC dataset are accompanied with audios and have a minimum resolution of 360p. The videos have various durations, background, illumination, and may contain perturbations that mimic transmission errors and compression. We have also carried out a user study to demonstrate the quality of the TMC dataset and to compare the performance of humans and AI models. The results showed that the TMC dataset can fool human participants in many cases, and the winning AI models of the Trusted Media Challenge outperformed humans. The TMC dataset is available for research purpose upon request via tmc-dataset@aisingapore.org.
△ Less
Submitted 16 August, 2022; v1 submitted 12 January, 2022;
originally announced January 2022.
-
A global linear and local superlinear/quadratic inexact non-interior continuation method for variational inequalities
Authors:
Le Thi Khanh Hien,
Chek Beng Chua
Abstract:
We use the concept of barrier-based smoothing approximations introduced in [ C. B. Chua and Z. Li, A barrier-based smoothing proximal point algorithm for NCPs over closed convex cones, SIOPT 23(2), 2010] to extend the non-interior continuation method proposed in [B. Chen and N. Xiu, A global linear and local quadratic noninterior continuation method for nonlinear complementarity problems based on…
▽ More
We use the concept of barrier-based smoothing approximations introduced in [ C. B. Chua and Z. Li, A barrier-based smoothing proximal point algorithm for NCPs over closed convex cones, SIOPT 23(2), 2010] to extend the non-interior continuation method proposed in [B. Chen and N. Xiu, A global linear and local quadratic noninterior continuation method for nonlinear complementarity problems based on Chen-Mangasarian smoothing functions, SIOPT 9(3), 1999] to an inexact non-interior continuation method for variational inequalities over general closed convex sets. Newton equations involved in the method are solved inexactly to deal with high dimension problems. The method is proved to have global linear and local superlinear/quadratic convergence under suitable assumptions. We apply the method to non-negative orthants, positive semidefinite cones, polyhedral sets, epigraphs of matrix operator norm cone and epigraphs of matrix nuclear norm cone.
△ Less
Submitted 5 March, 2020; v1 submitted 7 February, 2018;
originally announced February 2018.
-
Examining Requirements Change Rework Effort: A Study
Authors:
Bee Bee Chua,
June Verner
Abstract:
Although software managers are generally good at new project estimation, their experience of scheduling rework tends to be poor. Inconsistent or incorrect effort estimation can increase the risk that the completion time for a project will be problematic. To continually alter software maintenance schedules during software maintenance is a daunting task. Our proposed framework, validated in a case s…
▽ More
Although software managers are generally good at new project estimation, their experience of scheduling rework tends to be poor. Inconsistent or incorrect effort estimation can increase the risk that the completion time for a project will be problematic. To continually alter software maintenance schedules during software maintenance is a daunting task. Our proposed framework, validated in a case study confirms that the variables resulting from requirements changes suffer from a number of problems, e.g., the coding used, end user involvement and user documentation. Our results clearly show a significant impact on rework effort as a result of unexpected errors that correlate with 1) weak characteristics and attributes as described in the program's source lines of code, especially in data declarations and data statements, 2) lack of communication between developers and users on a change effects, and 3) unavailability of user documentation. To keep rework effort under control, new criteria in change request forms are proposed. These criteria are shown in a proposed framework; the more case studies that are validated, the more reliable the result will be in determining the outcome of effort rework estimation.
△ Less
Submitted 29 July, 2010;
originally announced July 2010.
-
Co-Betweenness: A Pairwise Notion of Centrality
Authors:
Eric D. Kolaczyk,
David B. Chua,
Marc Barthelemy
Abstract:
Betweenness centrality is a metric that seeks to quantify a sense of the importance of a vertex in a network graph in terms of its "control" on the distribution of information along geodesic paths throughout that network. This quantity however does not capture how different vertices participate together in such control. In order to allow for the uncovering of finer details in this regard, we int…
▽ More
Betweenness centrality is a metric that seeks to quantify a sense of the importance of a vertex in a network graph in terms of its "control" on the distribution of information along geodesic paths throughout that network. This quantity however does not capture how different vertices participate together in such control. In order to allow for the uncovering of finer details in this regard, we introduce here an extension of betweenness centrality to pairs of vertices, which we term co-betweenness, that provides the basis for quantifying various analogous pairwise notions of importance and control. More specifically, we motivate and define a precise notion of co-betweenness, we present an efficient algorithm for its computation, extending the algorithm of Brandes in a natural manner, and we illustrate the utilization of this co-betweenness on a handful of different communication networks. From these real-world examples, we show that the co-betweenness allows one to identify certain vertices which are not the most central vertices but which, nevertheless, act as important actors in the relaying and dispatching of information in the network.
△ Less
Submitted 21 September, 2007;
originally announced September 2007.
-
Network Kriging
Authors:
David B. Chua,
Eric D. Kolaczyk,
Mark Crovella
Abstract:
Network service providers and customers are often concerned with aggregate performance measures that span multiple network paths. Unfortunately, forming such network-wide measures can be difficult, due to the issues of scale involved. In particular, the number of paths grows too rapidly with the number of endpoints to make exhaustive measurement practical. As a result, it is of interest to explo…
▽ More
Network service providers and customers are often concerned with aggregate performance measures that span multiple network paths. Unfortunately, forming such network-wide measures can be difficult, due to the issues of scale involved. In particular, the number of paths grows too rapidly with the number of endpoints to make exhaustive measurement practical. As a result, it is of interest to explore the feasibility of methods that dramatically reduce the number of paths measured in such situations while maintaining acceptable accuracy.
We cast the problem as one of statistical prediction--in the spirit of the so-called `kriging' problem in spatial statistics--and show that end-to-end network properties may be accurately predicted in many cases using a surprisingly small set of carefully chosen paths. More precisely, we formulate a general framework for the prediction problem, propose a class of linear predictors for standard quantities of interest (e.g., averages, totals, differences) and show that linear algebraic methods of subset selection may be used to effectively choose which paths to measure. We characterize the performance of the resulting methods, both analytically and numerically. The success of our methods derives from the low effective rank of routing matrices as encountered in practice, which appears to be a new observation in its own right with potentially broad implications on network measurement generally.
△ Less
Submitted 3 October, 2005; v1 submitted 1 October, 2005;
originally announced October 2005.
-
A Statistical Framework for Efficient Monitoring of End-to-End Network Properties
Authors:
David B. Chua,
Eric D. Kolaczyk,
Mark Crovella
Abstract:
Network service providers and customers are often concerned with aggregate performance measures that span multiple network paths. Unfortunately, forming such network-wide measures can be difficult, due to the issues of scale involved. In particular, the number of paths grows too rapidly with the number of endpoints to make exhaustive measurement practical. As a result, there is interest in the f…
▽ More
Network service providers and customers are often concerned with aggregate performance measures that span multiple network paths. Unfortunately, forming such network-wide measures can be difficult, due to the issues of scale involved. In particular, the number of paths grows too rapidly with the number of endpoints to make exhaustive measurement practical. As a result, there is interest in the feasibility of methods that dramatically reduce the number of paths measured in such situations while maintaining acceptable accuracy.
In previous work we proposed a statistical framework to efficiently address this problem, in the context of additive metrics such as delay and loss rate, for which the per-path metric is a sum of (possibly transformed) per-link measures. The key to our method lies in the observation and exploitation of significant redundancy in network paths (sharing of common links).
In this paper we make three contributions: (1) we generalize the framework to make it more immediately applicable to network measurements encountered in practice; (2) we demonstrate that the observed path redundancy upon which our method is based is robust to variation in key network conditions and characteristics, including link failures; and (3) we show how the framework may be applied to address three practical problems of interest to network providers and customers, using data from an operating network. In particular, we show how appropriate selection of small sets of path measurements can be used to accurately estimate network-wide averages of path delays, to reliably detect network anomalies, and to effectively make a choice between alternative sub-networks, as a customer choosing between two providers or two ingress points into a provider network.
△ Less
Submitted 8 December, 2004; v1 submitted 8 December, 2004;
originally announced December 2004.