Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–23 of 23 results for author: Acker, A

Searching in archive cs. Search in all archives.
.
  1. arXiv:2603.02057  [pdf, ps, other

    cs.DC

    Beyond Microservices: Testing Web-Scale RCA Methods on GPU-Driven LLM Workloads

    Authors: Dominik Scheinert, Alexander Acker, Thorsten Wittkopp, Soeren Becker, Hamza Yous, Karnakar Reddy, Ibrahim Farhat, Hakim Hacid, Odej Kao

    Abstract: Large language model (LLM) services have become an integral part of search, assistance, and decision-making applications. However, unlike traditional web or microservices, the hardware and software stack enabling LLM inference deployment is of higher complexity and far less field-tested, making it more susceptible to failures that are difficult to resolve. Keeping outage costs and quality of servi… ▽ More

    Submitted 2 March, 2026; originally announced March 2026.

    Comments: 13 pages, 8 figures, 1 table

  2. arXiv:2602.22760  [pdf, ps, other

    cs.DC cs.AI

    Distributed LLM Pretraining During Renewable Curtailment Windows: A Feasibility Study

    Authors: Philipp Wiesner, Soeren Becker, Brett Cornick, Dominik Scheinert, Alexander Acker, Odej Kao

    Abstract: Training large language models (LLMs) requires substantial compute and energy. At the same time, renewable energy sources regularly produce more electricity than the grid can absorb, leading to curtailment, the deliberate reduction of clean generation that would otherwise go to waste. These periods represent an opportunity: if training is aligned with curtailment windows, LLMs can be pretrained us… ▽ More

    Submitted 26 February, 2026; originally announced February 2026.

    Comments: Technical report

  3. arXiv:2511.13761  [pdf, ps, other

    cs.DC cs.AI cs.LG

    What happens when nanochat meets DiLoCo?

    Authors: Alexander Acker, Soeren Becker, Sasho Nedelkoski, Dominik Scheinert, Odej Kao, Philipp Wiesner

    Abstract: Although LLM training is typically centralized with high-bandwidth interconnects and large compute budgets, emerging methods target communication-constrained training in distributed environments. The model trade-offs introduced by this shift remain underexplored, and our goal is to study them. We use the open-source nanochat project, a compact 8K-line full-stack ChatGPT-like implementation conta… ▽ More

    Submitted 14 November, 2025; originally announced November 2025.

    Comments: 8pages, 3 figures, technical report

  4. arXiv:2510.08576  [pdf, ps, other

    cs.SE cs.AI cs.CL cs.HC

    Comparative Analysis of Large Language Models for the Machine-Assisted Resolution of User Intentions

    Authors: Justus Flerlage, Alexander Acker, Odej Kao

    Abstract: Large Language Models (LLMs) have emerged as transformative tools for natural language understanding and user intent resolution, enabling tasks such as translation, summarization, and, increasingly, the orchestration of complex workflows. This development signifies a paradigm shift from conventional, GUI-driven user interfaces toward intuitive, language-first interaction paradigms. Rather than man… ▽ More

    Submitted 11 November, 2025; v1 submitted 29 August, 2025; originally announced October 2025.

    Comments: Accepted at First International Workshop on Human-AI Collaborative Systems (HAIC), published in CEUR-WS.org Vol-4072 (2025). URN: urn:nbn:de:0074-4072-x

  5. arXiv:2510.03371  [pdf, ps, other

    cs.LG cs.AI cs.DC

    Distributed Low-Communication Training with Decoupled Momentum Optimization

    Authors: Sasho Nedelkoski, Alexander Acker, Odej Kao, Soeren Becker, Dominik Scheinert

    Abstract: The training of large models demands substantial computational resources, typically available only in data centers with high-bandwidth interconnects. However, reducing the reliance on high-bandwidth interconnects between nodes enables the use of distributed compute resources as an alternative to centralized data center training. Building on recent advances in distributed model training, we propose… ▽ More

    Submitted 3 October, 2025; originally announced October 2025.

    Comments: NeurIPS 2025 - DynaFront 2025: Dynamics at the Frontiers of Optimization, Sampling, and Games Workshop

  6. Data Work in Memory Institutions: Why and How Information Professionals Use Wikidata

    Authors: Riya Sinha, Amelia Acker, Hanlin Li

    Abstract: Wikidata, an open structured database and a sibling project to Wikipedia, has recently become an important platform for information professionals to share structured metadata from their memory institutions, organizations that maintain public knowledge and cultural heritage materials. While studies have investigated why and how peer producers contribute to Wikidata, the institutional motivations an… ▽ More

    Submitted 18 August, 2025; originally announced August 2025.

    Comments: 27 pages, 1 figure, 2 tables. Accepted to PACM HCI (CSCW 2025)

  7. arXiv:2312.14748  [pdf, other

    cs.LG cs.SE

    Progressing from Anomaly Detection to Automated Log Labeling and Pioneering Root Cause Analysis

    Authors: Thorsten Wittkopp, Alexander Acker, Odej Kao

    Abstract: The realm of AIOps is transforming IT landscapes with the power of AI and ML. Despite the challenge of limited labeled data, supervised models show promise, emphasizing the importance of leveraging labels for training, especially in deep learning contexts. This study enhances the field by introducing a taxonomy for log anomalies and exploring automated data labeling to mitigate labeling challenges… ▽ More

    Submitted 22 December, 2023; originally announced December 2023.

    Comments: accepted at AIOPS workshop @ICDM 2023

  8. arXiv:2312.00357  [pdf

    eess.IV cs.CV cs.LG

    A Generalizable Deep Learning System for Cardiac MRI

    Authors: Rohan Shad, Cyril Zakka, Dhamanpreet Kaur, Mrudang Mathur, Robyn Fong, Joseph Cho, Ross Warren Filice, John Mongan, Kimberly Kalianos, Nishith Khandwala, David Eng, Matthew Leipzig, Walter R. Witschey, Alejandro de Feria, Victor A. Ferrari, Euan A. Ashley, Michael A. Acker, Curtis Langlotz, William Hiesinger

    Abstract: Cardiac MRI allows for a comprehensive assessment of myocardial structure, function and tissue characteristics. Here we describe a foundational vision system for cardiac MRI, capable of representing the breadth of human cardiovascular disease and health. Our deep-learning model is trained via self-supervised contrastive learning, in which visual concepts in cine-sequence cardiac MRI scans are lear… ▽ More

    Submitted 25 March, 2026; v1 submitted 1 December, 2023; originally announced December 2023.

    Comments: Published in Nature Biomedical Engineering; Supplementary Appendix available on publisher website. Code: https://github.com/rohanshad/cmr_transformer

    ACM Class: I.2.10

    Journal ref: Nat. Biomed. Eng (2026)

  9. arXiv:2301.10681  [pdf, other

    cs.LG

    PULL: Reactive Log Anomaly Detection Based On Iterative PU Learning

    Authors: Thorsten Wittkopp, Dominik Scheinert, Philipp Wiesner, Alexander Acker, Odej Kao

    Abstract: Due to the complexity of modern IT services, failures can be manifold, occur at any stage, and are hard to detect. For this reason, anomaly detection applied to monitoring data such as logs allows gaining relevant insights to improve IT services steadily and eradicate failures. However, existing anomaly detection methods that provide high accuracy often rely on labeled training data, which are tim… ▽ More

    Submitted 25 January, 2023; originally announced January 2023.

    Comments: published in the proceedings of the 56th Hawaii International Conference on System Sciences (HICSS 2023)

  10. Data-Driven Approach for Log Instruction Quality Assessment

    Authors: Jasmin Bogatinovski, Sasho Nedelkoski, Alexander Acker, Jorge Cardoso, Odej Kao

    Abstract: In the current IT world, developers write code while system operators run the code mostly as a black box. The connection between both worlds is typically established with log messages: the developer provides hints to the (unknown) operator, where the cause of an occurred issue is, and vice versa, the operator can report bugs during operation. To fulfil this purpose, developers write log instructio… ▽ More

    Submitted 6 April, 2022; originally announced April 2022.

    Comments: This paper is accepted for publication at the 30th International Conference on Program Comprehension under doi: 10.1145/3524610.3527906. The copyrights are handled following the corresponding agreement between the author and publisher

  11. LogLAB: Attention-Based Labeling of Log Data Anomalies via Weak Supervision

    Authors: Thorsten Wittkopp, Philipp Wiesner, Dominik Scheinert, Alexander Acker

    Abstract: With increasing scale and complexity of cloud operations, automated detection of anomalies in monitoring data such as logs will be an essential part of managing future IT infrastructures. However, many methods based on artificial intelligence, such as supervised deep learning models, require large amounts of labeled training data to perform well. In practice, this data is rarely available because… ▽ More

    Submitted 25 November, 2021; v1 submitted 2 November, 2021; originally announced November 2021.

    Comments: Paper accepted on ICSOC 2021 and published on springer

    Journal ref: 19th International Conference on Service-Oriented Computing, 2021, 700-707

  12. arXiv:2109.09537  [pdf, other

    cs.LG

    A2Log: Attentive Augmented Log Anomaly Detection

    Authors: Thorsten Wittkopp, Alexander Acker, Sasho Nedelkoski, Jasmin Bogatinovski, Dominik Scheinert, Wu Fan, Odej Kao

    Abstract: Anomaly detection becomes increasingly important for the dependability and serviceability of IT services. As log lines record events during the execution of IT services, they are a primary source for diagnostics. Thereby, unsupervised methods provide a significant benefit since not all anomalies can be known at training time. Existing unsupervised methods need anomaly examples to obtain a suitable… ▽ More

    Submitted 20 September, 2021; originally announced September 2021.

    Comments: This paper has been accepted for HICSS 2022 and will appear in the conference proceedings

  13. Enel: Context-Aware Dynamic Scaling of Distributed Dataflow Jobs using Graph Propagation

    Authors: Dominik Scheinert, Houkun Zhu, Lauritz Thamsen, Morgan K. Geldenhuys, Jonathan Will, Alexander Acker, Odej Kao

    Abstract: Distributed dataflow systems like Spark and Flink enable the use of clusters for scalable data analytics. While runtime prediction models can be used to initially select appropriate cluster resources given target runtimes, the actual runtime performance of dataflow jobs depends on several factors and varies over time. Yet, in many situations, dynamic scaling can be used to meet formulated runtime… ▽ More

    Submitted 26 January, 2022; v1 submitted 27 August, 2021; originally announced August 2021.

    Comments: 8 pages, 5 figures, 3 tables

    Journal ref: IEEE IPCCC (2021) 1-8

  14. Bellamy: Reusing Performance Models for Distributed Dataflow Jobs Across Contexts

    Authors: Dominik Scheinert, Lauritz Thamsen, Houkun Zhu, Jonathan Will, Alexander Acker, Thorsten Wittkopp, Odej Kao

    Abstract: Distributed dataflow systems enable the use of clusters for scalable data analytics. However, selecting appropriate cluster resources for a processing job is often not straightforward. Performance models trained on historical executions of a concrete job are helpful in such situations, yet they are usually bound to a specific job execution context (e.g. node type, software versions, job parameters… ▽ More

    Submitted 17 October, 2021; v1 submitted 29 July, 2021; originally announced July 2021.

    Comments: 10 pages, 8 figures, 2 tables

    Journal ref: IEEE CLUSTER (2021) 261-270

  15. Learning Dependencies in Distributed Cloud Applications to Identify and Localize Anomalies

    Authors: Dominik Scheinert, Alexander Acker, Lauritz Thamsen, Morgan K. Geldenhuys, Odej Kao

    Abstract: Operation and maintenance of large distributed cloud applications can quickly become unmanageably complex, putting human operators under immense stress when problems occur. Utilizing machine learning for identification and localization of anomalies in such systems supports human experts and enables fast mitigation. However, due to the various inter-dependencies of system components, anomalies do n… ▽ More

    Submitted 9 September, 2021; v1 submitted 9 March, 2021; originally announced March 2021.

    Comments: 6 pages, 5 figures, 3 tables

    Journal ref: IEEE/ACM CloudIntelligence (2021) 7-12

  16. TELESTO: A Graph Neural Network Model for Anomaly Classification in Cloud Services

    Authors: Dominik Scheinert, Alexander Acker

    Abstract: Deployment, operation and maintenance of large IT systems becomes increasingly complex and puts human experts under extreme stress when problems occur. Therefore, utilization of machine learning (ML) and artificial intelligence (AI) is applied on IT system operation and maintenance - summarized in the term AIOps. One specific direction aims at the recognition of re-occurring anomaly types to enabl… ▽ More

    Submitted 29 July, 2021; v1 submitted 25 February, 2021; originally announced February 2021.

    Comments: 12 pages, 2 figures, 4 tables

    Journal ref: Springer ICSOC LNCS 12632 (2020) 214-227

  17. arXiv:2102.11570  [pdf, other

    cs.AI cs.CL cs.SE

    Robust and Transferable Anomaly Detection in Log Data using Pre-Trained Language Models

    Authors: Harold Ott, Jasmin Bogatinovski, Alexander Acker, Sasho Nedelkoski, Odej Kao

    Abstract: Anomalies or failures in large computer systems, such as the cloud, have an impact on a large number of users that communicate, compute, and store information. Therefore, timely and accurate anomaly detection is necessary for reliability, security, safe operation, and mitigation of losses in these increasingly important systems. Recently, the evolution of the software industry opens up several pro… ▽ More

    Submitted 23 February, 2021; originally announced February 2021.

  18. Towards AIOps in Edge Computing Environments

    Authors: Soeren Becker, Florian Schmidt, Anton Gulenko, Alexander Acker, Odej Kao

    Abstract: Edge computing was introduced as a technical enabler for the demanding requirements of new network technologies like 5G. It aims to overcome challenges related to centralized cloud computing environments by distributing computational resources to the edge of the network towards the customers. The complexity of the emerging infrastructures increases significantly, together with the ramifications of… ▽ More

    Submitted 12 February, 2021; originally announced February 2021.

  19. arXiv:2102.00880  [pdf, other

    cs.LG cs.NI

    Decentralized Federated Learning Preserves Model and Data Privacy

    Authors: Thorsten Wittkopp, Alexander Acker

    Abstract: The increasing complexity of IT systems requires solutions, that support operations in case of failure. Therefore, Artificial Intelligence for System Operations (AIOps) is a field of research that is becoming increasingly focused, both in academia and industry. One of the major issues of this area is the lack of access to adequately labeled data, which is majorly due to legal protection regulation… ▽ More

    Submitted 1 February, 2021; originally announced February 2021.

  20. arXiv:2101.06054  [pdf, other

    cs.LG cs.SE

    Artificial Intelligence for IT Operations (AIOPS) Workshop White Paper

    Authors: Jasmin Bogatinovski, Sasho Nedelkoski, Alexander Acker, Florian Schmidt, Thorsten Wittkopp, Soeren Becker, Jorge Cardoso, Odej Kao

    Abstract: Artificial Intelligence for IT Operations (AIOps) is an emerging interdisciplinary field arising in the intersection between the research areas of machine learning, big data, streaming analytics, and the management of IT operations. AIOps, as a field, is a candidate to produce the future standard for IT operation management. To that end, AIOps has several challenges. First, it needs to combine sep… ▽ More

    Submitted 15 January, 2021; originally announced January 2021.

    Comments: 8 pages, white paper for the AIOPS 2020 workshop at ICSOC 2020

  21. arXiv:2008.09340  [pdf, other

    cs.LG cs.IR stat.ML

    Self-Attentive Classification-Based Anomaly Detection in Unstructured Logs

    Authors: Sasho Nedelkoski, Jasmin Bogatinovski, Alexander Acker, Jorge Cardoso, Odej Kao

    Abstract: The detection of anomalies is essential mining task for the security and reliability in computer systems. Logs are a common and major data source for anomaly detection methods in almost every computer system. They collect a range of significant events describing the runtime system status. Recent studies have focused predominantly on one-class deep learning methods on predefined non-learnable numer… ▽ More

    Submitted 21 August, 2020; originally announced August 2020.

    Comments: 11 pages, 8 figures, Accepted at ICDM 2020: 20th IEEE International Conference on Data Mining

  22. arXiv:2007.03568  [pdf, other

    cs.LG eess.SY stat.ML

    Superiority of Simplicity: A Lightweight Model for Network Device Workload Prediction

    Authors: Alexander Acker, Thorsten Wittkopp, Sasho Nedelkoski, Jasmin Bogatinovski, Odej Kao

    Abstract: The rapid growth and distribution of IT systems increases their complexity and aggravates operation and maintenance. To sustain control over large sets of hosts and the connecting networks, monitoring solutions are employed and constantly enhanced. They collect diverse key performance indicators (KPIs) (e.g. CPU utilization, allocated memory, etc.) and provide detailed information about the system… ▽ More

    Submitted 7 July, 2020; originally announced July 2020.

  23. arXiv:2003.07905  [pdf, other

    cs.LG cs.SE

    Self-Supervised Log Parsing

    Authors: Sasho Nedelkoski, Jasmin Bogatinovski, Alexander Acker, Jorge Cardoso, Odej Kao

    Abstract: Logs are extensively used during the development and maintenance of software systems. They collect runtime events and allow tracking of code execution, which enables a variety of critical tasks such as troubleshooting and fault detection. However, large-scale software systems generate massive volumes of semi-structured log records, posing a major challenge for automated analysis. Parsing semi-stru… ▽ More

    Submitted 17 March, 2020; originally announced March 2020.