Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–10 of 10 results for author: Schouten, G

Searching in archive cs. Search in all archives.
.
  1. arXiv:2607.20149  [pdf, ps, other

    cs.CY

    Data Annotations as Pedagogical Hints: From Subjective Labels to Critical Thinking

    Authors: Ralf Raumanns, Theresa Elstner, Louis Ferger-Andrews, Louise M. Carlsen, Martin Potthast, Gerard Schouten, Josien P. W. Pluim, Veronika Cheplygina

    Abstract: Machine learning courses often use pre-labeled datasets, hiding the subjectivity of human annotation. This creates students with an overly trusting view of AI data and models, undervaluing interpretive diversity. We investigated whether manual data annotation tasks teach students about subjective labeling. Study Design: An annotation activity was implemented at two universities: Fontys (Netherla… ▽ More

    Submitted 22 July, 2026; originally announced July 2026.

    Comments: 24 pages, 6 figures, 6 tables

  2. arXiv:2606.03214  [pdf, ps, other

    cs.AI cs.CV cs.CY cs.LG

    Effect of Demographic Bias on Skin Lesion Classification

    Authors: Ralf Raumanns, Gerard Schouten, Veronika Cheplygina, Josien P. W. Pluim

    Abstract: In this study, we evaluate the performance of skin lesion classification using ResNet-based convolutional models, focusing on the impact of demographic bias in training data, particularly variations in patient sex and age. We use linear programming to generate datasets with controlled demographic characteristics, allowing systematic investigation of bias effects. Three learning strategies are eval… ▽ More

    Submitted 2 June, 2026; originally announced June 2026.

    Comments: Accepted for publication at the Journal of Machine Learning for Biomedical Imaging (MELBA) , 26 pages, 12 figures

    Journal ref: https://melba-journal.org/2026:011

  3. arXiv:2505.14707  [pdf, ps, other

    cs.MM cs.AI cs.CV

    CrypticBio: A Large Multimodal Dataset for Visually Confusing Biodiversity

    Authors: Georgiana Manolache, Gerard Schouten, Joaquin Vanschoren

    Abstract: We present CrypticBio, the largest publicly available multimodal dataset of visually confusing species, specifically curated to support the development of AI models in the context of biodiversity applications. Visually confusing or cryptic species are groups of two or more taxa that are nearly indistinguishable based on visual characteristics alone. While much existing work addresses taxonomic ide… ▽ More

    Submitted 16 May, 2025; originally announced May 2025.

    Comments: We present CrypticBio, the largest publicly available multimodal dataset of visually confusing species, specifically curated to support the development of AI models for biodiversity identification using images, language and spatiotemporal data

  4. Dataset Distribution Impacts Model Fairness: Single vs. Multi-Task Learning

    Authors: Ralf Raumanns, Gerard Schouten, Josien P. W. Pluim, Veronika Cheplygina

    Abstract: The influence of bias in datasets on the fairness of model predictions is a topic of ongoing research in various fields. We evaluate the performance of skin lesion classification using ResNet-based CNNs, focusing on patient sex variations in training data and three different learning strategies. We present a linear programming method for generating datasets with varying patient sex and class label… ▽ More

    Submitted 9 December, 2024; v1 submitted 24 July, 2024; originally announced July 2024.

    Comments: Published in the FAIMI EPIMI 2024 Workshop

    Journal ref: Ethics and Fairness in Medical Imaging. FAIMI EPIMI 2024 2024. Lecture Notes in Computer Science, vol 15198

  5. arXiv:2303.13151  [pdf, other

    cs.AI cs.SE

    Defining Quality Requirements for a Trustworthy AI Wildflower Monitoring Platform

    Authors: Petra Heck, Gerard Schouten

    Abstract: For an AI solution to evolve from a trained machine learning model into a production-ready AI system, many more things need to be considered than just the performance of the machine learning model. A production-ready AI system needs to be trustworthy, i.e. of high quality. But how to determine this in practice? For traditional software, ISO25000 and its predecessors have since long time been used… ▽ More

    Submitted 23 March, 2023; originally announced March 2023.

    Comments: Preprint - Paper accepted for CAIN23 - 2nd international conference on AI Engineering

  6. arXiv:2107.12734  [pdf, ps, other

    cs.CV cs.HC cs.LG

    ENHANCE (ENriching Health data by ANnotations of Crowd and Experts): A case study for skin lesion classification

    Authors: Ralf Raumanns, Gerard Schouten, Max Joosten, Josien P. W. Pluim, Veronika Cheplygina

    Abstract: We present ENHANCE, an open dataset with multiple annotations to complement the existing ISIC and PH2 skin lesion classification datasets. This dataset contains annotations of visual ABC (asymmetry, border, colour) features from non-expert annotation sources: undergraduate students, crowd workers from Amazon MTurk and classic image processing algorithms. In this paper we first analyse the correlat… ▽ More

    Submitted 24 December, 2021; v1 submitted 27 July, 2021; originally announced July 2021.

    Comments: Accepted for publication at the Journal of Machine Learning for Biomedical Imaging (MELBA) https://www.melba-journal.org

  7. arXiv:2103.10703  [pdf, other

    cs.AI cs.SE

    Lessons Learned from Educating AI Engineers

    Authors: Petra Heck, Gerard Schouten

    Abstract: Over the past three years we have built a practice-oriented, bachelor level, educational programme for software engineers to specialize as AI engineers. The experience with this programme and the practical assignments our students execute in industry has given us valuable insights on the profession of AI engineer. In this paper we discuss our programme and the lessons learned for industry and rese… ▽ More

    Submitted 19 March, 2021; originally announced March 2021.

    Comments: Acccepted for the 1st International Workshop on AI Engineering (WAIN21)

  8. arXiv:2011.01590  [pdf, other

    cs.SE cs.AI

    Turning Software Engineers into AI Engineers

    Authors: Petra Heck, Gerard Schouten

    Abstract: In industry as well as education as well as academics we see a growing need for knowledge on how to apply machine learning in software applications. With the educational programme ICT & AI at Fontys UAS we had to find an answer to the question: "How should we educate software engineers to become AI engineers?" This paper describes our educational programme, the open source tools we use, and the li… ▽ More

    Submitted 4 January, 2021; v1 submitted 3 November, 2020; originally announced November 2020.

  9. arXiv:2005.10050  [pdf, other

    cs.LG cs.AI stat.ML

    Risk of Training Diagnostic Algorithms on Data with Demographic Bias

    Authors: Samaneh Abbasi-Sureshjani, Ralf Raumanns, Britt E. J. Michels, Gerard Schouten, Veronika Cheplygina

    Abstract: One of the critical challenges in machine learning applications is to have fair predictions. There are numerous recent examples in various domains that convincingly show that algorithms trained with biased datasets can easily lead to erroneous or discriminatory conclusions. This is even more crucial in clinical applications where the predictive algorithms are designed mainly based on a limited or… ▽ More

    Submitted 17 June, 2020; v1 submitted 20 May, 2020; originally announced May 2020.

  10. arXiv:2004.14745  [pdf, other

    cs.HC cs.CV cs.LG eess.IV

    Multi-task Ensembles with Crowdsourced Features Improve Skin Lesion Diagnosis

    Authors: Ralf Raumanns, Elif K Contar, Gerard Schouten, Veronika Cheplygina

    Abstract: Machine learning has a recognised need for large amounts of annotated data. Due to the high cost of expert annotations, crowdsourcing, where non-experts are asked to label or outline images, has been proposed as an alternative. Although many promising results are reported, the quality of diagnostic crowdsourced labels is still unclear. We propose to address this by instead asking the crowd about v… ▽ More

    Submitted 6 July, 2020; v1 submitted 28 April, 2020; originally announced April 2020.