Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–13 of 13 results for author: Ke, K

Searching in archive cs. Search in all archives.
.
  1. arXiv:2607.17586  [pdf, ps, other

    cs.CR cs.AI cs.LG

    Detection, Attribution, Narration: An End-to-End Pipeline for Explainable Money Mule Identification

    Authors: Yuge Zhang, Yuanxing Zhang, Yichao Jin, Khairul Amsyar Mohd Razis, Nicholas Qi An Choo, Kai Yin Anders Wong, Xinyan Tang, Kenneth Zhu Ke, Wee Keong Dennis Lee, Jingyuan Zhao

    Abstract: Money mule accounts are critical facilitators of financial fraud, yet detecting them at scale remains challenging due to the heterogeneous nature of transactional and behavioural data. We present an end-to-end pipeline for customer-level mule detection comprising three stages: (1) a LightGBM classifier trained on 280 engineered features spanning transaction patterns, account demographics, network… ▽ More

    Submitted 20 July, 2026; originally announced July 2026.

  2. arXiv:2606.04301  [pdf, ps, other

    cs.CV

    XSSR: Cross-Domain Self-Supervised Representative Selection for Efficient Annotation in Medical Image Segmentation

    Authors: Byunghyun Ko, Aleksei Anisimov, Kobe Ke, Suhas Bharthepude, Jeongkyu Lee

    Abstract: Acquiring labeled medical image data is resource-intensive and a challenge further exacerbated in cross-domain scenarios where source and target datasets differ in imaging equipment, population, or clinical site. This study introduces XSSR (Cross-Domain Self-Supervised Representative Selection), a framework designed to minimize annotation effort in the target domain while maintaining robust segmen… ▽ More

    Submitted 2 June, 2026; originally announced June 2026.

    Comments: Accepted to the Third International Conference on AI in Healthcare (AIiH 2026). This is the preprint version of the paper

  3. arXiv:2605.28127  [pdf, ps, other

    cs.LG

    Adaptive Coarse-to-Fine Subgoal Refinement for Long-Horizon Offline Goal-Conditioned Reinforcement Learning

    Authors: Kaiqiang Ke, Shenghong He, Chengdong Xu, Yuheng Luo, Xiangyuan Lan, Chao Yu

    Abstract: Offline goal-conditioned reinforcement learning (GCRL) is challenging in long-horizon tasks, where distant state--goal pairs provide weak supervision and value estimates become vulnerable to accumulated bootstrapping errors. Hierarchical methods mitigate this difficulty by introducing intermediate subgoals, but fixed temporal abstractions or fixed hierarchy depths can be mismatched to state--goal… ▽ More

    Submitted 27 May, 2026; originally announced May 2026.

  4. arXiv:2605.20684  [pdf, ps, other

    cs.CL

    Beyond Semantic Similarity: A Two-Phase Non-Parametric Retrieval Workflow for Corporate Credit Underwriting

    Authors: Linus Ng Junjia, Ezekiel Tee Kongquan, Kelvin Heng, Kenneth Zhu Ke, Zhao Jing Yuan

    Abstract: Corporate credit underwriting requires analysts to extract actionable evidence from long, heterogeneous financial documents spanning hundreds of pages and multiple languages. Standard Retrieval-Augmented Generation (RAG) pipelines optimize for semantic similarity, which frequently surfaces passages that are topically related but lack decision utility, a problem we term the similarity-utility gap.… ▽ More

    Submitted 20 May, 2026; originally announced May 2026.

  5. arXiv:2605.08769  [pdf, ps, other

    cs.AI

    EvoMAS: Learning Execution-Time Workflows for Multi-Agent Systems

    Authors: Chengdong Xu, Kaiqiang Ke, Ziheng Liu, Jiaqi Wei, Zibo Shao, Weile Guo, Chao Yu

    Abstract: Large language model (LLM)-based multi-agent systems have shown strong potential on complex tasks through agent specialization, tool use, and collaborative reasoning. However, most automated multi-agent system design methods still follow a one-shot paradigm: a workflow is optimized or selected before execution and then reused unchanged throughout the task. This static coordination strategy is ill-… ▽ More

    Submitted 9 May, 2026; originally announced May 2026.

    Comments: 22 pages, 8 figures

  6. arXiv:2604.26462  [pdf, ps, other

    cs.CV

    A Multistage Extraction Pipeline for Long Scanned Financial Documents: An Empirical Study in Industrial KYC Workflows

    Authors: Yuxuan Han, Yuanxing Zhang, Yushuo Wang, Yichao Jin, Kenneth Zhu Ke, Jingyuan Zhao

    Abstract: Structured information extraction from long, multilingual scanned financial documents is a core requirement in industrial KYC and compliance workflows. These documents are typically non machine readable, noisy, and visually heterogeneous. They usually span dozens of pages while containing only sparse task relevant information. Although recent vision-language models achieve strong benchmark perform… ▽ More

    Submitted 29 April, 2026; originally announced April 2026.

  7. arXiv:2601.16518  [pdf

    cs.IT

    Noise-immune and AI-enhanced DNA storage via adaptive partition mapping of digital data

    Authors: Zimu Li, Bingyi Liu, Lei Zhao, Qian Zhang, Yang Liu, Jun Liu, Ke Ke, Huating Kong, Xiaolei Zuo, Chunhai Fan, Fei Wang

    Abstract: Encoding digital information into DNA sequences offers an attractive potential solution for storing rapidly growing data under the information age and the rise of artificial intelligence. However, practical implementations of DNA storage are constrained by errors introduced during synthesis, preservation, and sequencing processes, and traditional error-correcting codes remain vulnerable to noise l… ▽ More

    Submitted 23 January, 2026; originally announced January 2026.

  8. arXiv:2512.14465  [pdf, ps, other

    cs.AI

    Context-Picker: Dynamic context selection using multi-stage reinforcement learning

    Authors: Siyuan Zhu, Chengdong Xu, Kaiqiang Ke, Chao Yu

    Abstract: In long-context question answering, selecting the appropriate scope of context for a query remains a key and unresolved challenge. Insufficient context can lead to missing essential information, whereas excessive context often introduces noise and degrades answer quality. Conventional methods, such as retrieving a fixed number of passages or applying reranking, struggle to dynamically determine wh… ▽ More

    Submitted 20 January, 2026; v1 submitted 16 December, 2025; originally announced December 2025.

  9. arXiv:2511.21878  [pdf, ps, other

    cs.SE cs.PL

    Advancing Automated In-Isolation Validation in Repository-Level Code Translation

    Authors: Kaiyao Ke, Ali Reza Ibrahimzada, Rangeet Pan, Saurabh Sinha, Reyhaneh Jabbarvand

    Abstract: Repository-level code translation aims to migrate entire repositories across programming languages while preserving functionality automatically. Despite advancements in repository-level code translation, validating the translations remains challenging. This paper proposes TRAM, which combines context-aware type resolution with mock-based in-isolation validation to achieve high-quality translations… ▽ More

    Submitted 23 December, 2025; v1 submitted 26 November, 2025; originally announced November 2025.

  10. arXiv:2510.23066  [pdf, ps, other

    cs.IR

    Multi-Stage Field Extraction of Financial Documents with OCR and Compact Vision-Language Models

    Authors: Yichao Jin, Yushuo Wang, Qishuai Zhong, Kent Chiu Jin-Chun, Kenneth Zhu Ke, Donald MacDonald

    Abstract: Financial documents are essential sources of information for regulators, auditors, and financial institutions, particularly for assessing the wealth and compliance of Small and Medium-sized Businesses. However, SMB documents are often difficult to parse. They are rarely born digital and instead are distributed as scanned images that are none machine readable. The scans themselves are low in resolu… ▽ More

    Submitted 27 October, 2025; originally announced October 2025.

  11. arXiv:2509.12810  [pdf, ps, other

    cs.AI

    H$^2$R: Hierarchical Hindsight Reflection for Multi-Task LLM Agents

    Authors: Shicheng Ye, Chao Yu, Kaiqiang Ke, Chengdong Xu, Yinqi Wei

    Abstract: Large language model (LLM)-based agents have shown strong potential in multi-task scenarios, owing to their ability to transfer knowledge across diverse tasks. However, existing approaches often treat prior experiences and knowledge as monolithic units, leading to inefficient and coarse-grained knowledge transfer. In this work, we propose a novel hierarchical memory architecture that enables fine-… ▽ More

    Submitted 16 September, 2025; originally announced September 2025.

  12. arXiv:2508.06108  [pdf, ps, other

    cs.LG cs.AI

    GCHR : Goal-Conditioned Hindsight Regularization for Sample-Efficient Reinforcement Learning

    Authors: Xing Lei, Wenyan Yang, Kaiqiang Ke, Shentao Yang, Xuetao Zhang, Joni Pajarinen, Donglin Wang

    Abstract: Goal-conditioned reinforcement learning (GCRL) with sparse rewards remains a fundamental challenge in reinforcement learning. While hindsight experience replay (HER) has shown promise by relabeling collected trajectories with achieved goals, we argue that trajectory relabeling alone does not fully exploit the available experiences in off-policy GCRL methods, resulting in limited sample efficiency.… ▽ More

    Submitted 8 August, 2025; originally announced August 2025.

  13. arXiv:2410.24117  [pdf, ps, other

    cs.SE cs.LG

    AlphaTrans: A Neuro-Symbolic Compositional Approach for Repository-Level Code Translation and Validation

    Authors: Ali Reza Ibrahimzada, Kaiyao Ke, Mrigank Pawagi, Muhammad Salman Abid, Rangeet Pan, Saurabh Sinha, Reyhaneh Jabbarvand

    Abstract: Code translation transforms programs from one programming language (PL) to another. Several rule-based transpilers have been designed to automate code translation between different pairs of PLs. However, the rules can become obsolete as the PLs evolve and cannot generalize to other PLs. Recent studies have explored the automation of code translation using Large Language Models (LLMs). One key obse… ▽ More

    Submitted 19 June, 2025; v1 submitted 31 October, 2024; originally announced October 2024.

    Comments: Published in FSE 2025