Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–7 of 7 results for author: Wong, A I

Searching in archive cs. Search in all archives.
.
  1. arXiv:2509.19344  [pdf

    cs.CL

    Performance of Large Language Models in Answering Critical Care Medicine Questions

    Authors: Mahmoud Alwakeel, Aditya Nagori, An-Kwok Ian Wong, Neal Chaisson, Vijay Krishnamoorthy, Rishikesan Kamaleswaran

    Abstract: Large Language Models have been tested on medical student-level questions, but their performance in specialized fields like Critical Care Medicine (CCM) is less explored. This study evaluated Meta-Llama 3.1 models (8B and 70B parameters) on 871 CCM questions. Llama3.1:70B outperformed 8B by 30%, with 60% average accuracy. Performance varied across domains, highest in Research (68.4%) and lowest in… ▽ More

    Submitted 16 September, 2025; originally announced September 2025.

  2. arXiv:2509.14283  [pdf

    cs.CL

    Predicting Antibiotic Resistance Patterns Using Sentence-BERT: A Machine Learning Approach

    Authors: Mahmoud Alwakeel, Michael E. Yarrington, Rebekah H. Wrenn, Ethan Fang, Jian Pei, Anand Chowdhury, An-Kwok Ian Wong

    Abstract: Antibiotic resistance poses a significant threat in in-patient settings with high mortality. Using MIMIC-III data, we generated Sentence-BERT embeddings from clinical notes and applied Neural Networks and XGBoost to predict antibiotic susceptibility. XGBoost achieved an average F1 score of 0.86, while Neural Networks scored 0.84. This study is among the first to use document embeddings for predict… ▽ More

    Submitted 16 September, 2025; originally announced September 2025.

  3. arXiv:2503.21004  [pdf

    cs.CL

    Evaluating Large Language Models for Automated Clinical Abstraction in Pulmonary Embolism Registries: Performance Across Model Sizes, Versions, and Parameters

    Authors: Mahmoud Alwakeel, Emory Buck, Jonathan G. Martin, Imran Aslam, Sudarshan Rajagopal, Jian Pei, Mihai V. Podgoreanu, Christopher J. Lindsell, An-Kwok Ian Wong

    Abstract: Pulmonary embolism (PE) registries accelerate practice-improving research but depend on resource-intensive manual abstraction of radiology reports. We evaluated whether openly available large-language models (LLMs) can automate concept extraction from computed-tomography PE (CTPE) reports without sacrificing data quality. Four Llama-3 (L3) variants (3.0 8 B, 3.1 8 B, 3.1 70 B, 3.3 70 B) and two re… ▽ More

    Submitted 11 August, 2025; v1 submitted 26 March, 2025; originally announced March 2025.

  4. arXiv:2410.12722  [pdf, other

    cs.CL

    WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation

    Authors: João Matos, Shan Chen, Siena Placino, Yingya Li, Juan Carlos Climent Pardo, Daphna Idan, Takeshi Tohyama, David Restrepo, Luis F. Nakayama, Jose M. M. Pascual-Leone, Guergana Savova, Hugo Aerts, Leo A. Celi, A. Ian Wong, Danielle S. Bitterman, Jack Gallifant

    Abstract: Multimodal/vision language models (VLMs) are increasingly being deployed in healthcare settings worldwide, necessitating robust benchmarks to ensure their safety, efficacy, and fairness. Multiple-choice question and answer (QA) datasets derived from national medical examinations have long served as valuable evaluation tools, but existing datasets are largely text-only and available in a limited su… ▽ More

    Submitted 16 October, 2024; originally announced October 2024.

    Comments: submitted for review, total of 14 pages

  5. arXiv:2408.04396  [pdf, other

    cs.LG

    Evaluating the Impact of Pulse Oximetry Bias in Machine Learning under Counterfactual Thinking

    Authors: Inês Martins, João Matos, Tiago Gonçalves, Leo A. Celi, A. Ian Wong, Jaime S. Cardoso

    Abstract: Algorithmic bias in healthcare mirrors existing data biases. However, the factors driving unfairness are not always known. Medical devices capture significant amounts of data but are prone to errors; for instance, pulse oximeters overestimate the arterial oxygen saturation of darker-skinned individuals, leading to worse outcomes. The impact of this bias in machine learning (ML) models remains uncl… ▽ More

    Submitted 8 August, 2024; originally announced August 2024.

    Comments: 10 pages; accepted at MICCAI's Third Workshop on Applications of Medical AI (2024)

  6. arXiv:2407.00242  [pdf, other

    cs.CL

    EHRmonize: A Framework for Medical Concept Abstraction from Electronic Health Records using Large Language Models

    Authors: João Matos, Jack Gallifant, Jian Pei, A. Ian Wong

    Abstract: Electronic health records (EHRs) contain vast amounts of complex data, but harmonizing and processing this information remains a challenging and costly task requiring significant clinical expertise. While large language models (LLMs) have shown promise in various healthcare applications, their potential for abstracting medical concepts from EHRs remains largely unexplored. We introduce EHRmonize,… ▽ More

    Submitted 28 June, 2024; originally announced July 2024.

    Comments: submitted for review, total of 10 pages

  7. Benchmarking emergency department triage prediction models with machine learning and large public electronic health records

    Authors: Feng Xie, Jun Zhou, Jin Wee Lee, Mingrui Tan, Siqi Li, Logasan S/O Rajnthern, Marcel Lucas Chee, Bibhas Chakraborty, An-Kwok Ian Wong, Alon Dagan, Marcus Eng Hock Ong, Fei Gao, Nan Liu

    Abstract: The demand for emergency department (ED) services is increasing across the globe, particularly during the current COVID-19 pandemic. Clinical triage and risk assessment have become increasingly challenging due to the shortage of medical resources and the strain on hospital infrastructure caused by the pandemic. As a result of the widespread use of electronic health records (EHRs), we now have acce… ▽ More

    Submitted 20 March, 2022; v1 submitted 22 November, 2021; originally announced November 2021.