-
APICURON: a reactive infrastructure for credit attribution across distributed research data ecosystems
Authors:
Adel Bouhraoua,
Mehdi Zoubiri,
Gavin Farrell,
Maria Cristina Aspromonte,
Alex Bateman,
Henning Hermjakob,
Maria Victoria Nugnes,
Daniela Raciti,
Nicholas Stiffler,
Geert van Geest,
Ulrike Wittig,
Karen Yook,
Federica Quaglia,
Silvio C. E. Tosatto
Abstract:
Data-driven biology relies on structured knowledge generated by expert biocurators, yet this work remains largely unrecognized in traditional academic assessments. To bridge this gap, we present the updated APICURON platform, a credit-attribution infrastructure that formally acknowledges these scientific contributions. Rather than relying on delayed batch reporting, the system captures curation ev…
▽ More
Data-driven biology relies on structured knowledge generated by expert biocurators, yet this work remains largely unrecognized in traditional academic assessments. To bridge this gap, we present the updated APICURON platform, a credit-attribution infrastructure that formally acknowledges these scientific contributions. Rather than relying on delayed batch reporting, the system captures curation events as they happen and transforms them into verifiable units of work. This design allows independent resources to define and update their own recognition models while preserving the historical record of each contribution. For researchers, APICURON highlights recent activity alongside lifetime achievements and connects verified activities to persistent academic profiles via ORCID. APICURON has been successfully integrated across biological knowledgebases and data resources, demonstrating its application to diverse workflows. Extending beyond biodata resources, it also supports recognition of non-traditional research artefacts, including training materials and research software, without imposing a rigid definition of contribution.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Traceable Trust for action-ready artificial intelligence in bioscience
Authors:
Huayu Xin,
Yizhi Cai,
Mukilan Deivarajan Suresh,
Gavin Michael Farrell,
Iwona Gajda,
Charlie Harrison,
Conor Houghton,
Mato Lagator,
Yang Lu,
Virginia Portillo,
Reyer Zwiggelaar,
Sebastian Lobentanzer
Abstract:
Artificial intelligence (AI) is becoming part of the working infrastructure of the biosciences. AI models can predict biomolecular structures, design proteins, rank variants, annotate images, recommend strains and optimise experimental conditions. We argue that the decision to use an AI output to guide laboratory action is a key juncture for trustworthy research and should follow a defined, review…
▽ More
Artificial intelligence (AI) is becoming part of the working infrastructure of the biosciences. AI models can predict biomolecular structures, design proteins, rank variants, annotate images, recommend strains and optimise experimental conditions. We argue that the decision to use an AI output to guide laboratory action is a key juncture for trustworthy research and should follow a defined, reviewable process. We propose Traceable Trust as a proportionate assessment-and-design framework for this output-to-action boundary. It asks what evidence supports the output, what capability is being claimed, what agency has been delegated, what threshold authorises action, who can override it and how outcomes inform later decisions. We illustrate the framework through three case studies spanning ecosystem resources, project design and laboratory action. Together, the cases show how trust can be documented where AI outputs begin to shape scientific work.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Little Red Dots on FIRE: Exploring the formation and observational signatures of ultra-compact early galaxies
Authors:
Niranjan Chandra Roy,
Daniel Anglés-Alcázar,
Rachel K. Cochrane,
Alexander J. Richings,
Jonathan Mercedes-Feliz,
Christopher C. Hayward,
Claude-André Faucher-Giguère,
Erini Lambrides,
Robert Feldmann,
Boon Kiat Oh,
Andrew Marszewski,
Guochao Sun,
Kelcey Davis,
Jed McKinney,
Caitlin M. Casey,
Tanio Díaz-Santos,
Madisyn Brooks,
Grace Farrell
Abstract:
Little Red Dots (LRDs) are compact sources with broad Balmer lines, Balmer breaks, anomalous UV emission, rising red continuum, and uncertain origin. We use FIRE cosmological simulations, 3D dust radiative transfer, and synthetic emission-line data cubes to test whether ultra-compact early galaxies can reproduce LRD-like observables without invoking AGN. In progenitors of present-day group halos (…
▽ More
Little Red Dots (LRDs) are compact sources with broad Balmer lines, Balmer breaks, anomalous UV emission, rising red continuum, and uncertain origin. We use FIRE cosmological simulations, 3D dust radiative transfer, and synthetic emission-line data cubes to test whether ultra-compact early galaxies can reproduce LRD-like observables without invoking AGN. In progenitors of present-day group halos ($M_{\rm halo} > 10^{13.5} M_{\odot}$), we identify transient phases at $z \approx 4-8$ lasting $\sim 150-400$ Myr in which strong dissipative inflows build massive ($M_{\star} \sim 10^{8.5}-10^{10.5} M_{\odot}$), UV-bright ($-23 \lesssim M_{\rm UV} \lesssim -20$), ultra-compact ($R_{\rm eff} < 300$ pc) stellar cores with extreme circular velocity ($V_{\rm circ} > 500$ km s$^{-1}$) and consistent with several LRD properties: strong Balmer breaks ($F_ν(4200{\rm Å})/F_ν(3500{\rm Å}) \sim 2$); blue UV beta slopes ($β_{\rm UV} \approx -1.25$); dust masses; ALMA non-detections; and Balmer-line widths up to $\sim 1500$ km s$^{-1}$ broadened by galaxy-scale dynamics. However, stellar emission and host-galaxy kinematics alone do not reproduce the red rest-optical continuum, more extreme Balmer breaks ($\gtrsim 2.5$) and line widths ($\gtrsim 2000$ km s$^{-1}$), or the broad-Balmer/narrow-forbidden-line signature of broad-line AGN. The same ultra-compact conditions efficiently fuel central BHs, suggesting a hybrid stellar+AGN scenario in which compact stars explain the UV continuum, Balmer break, and intermediate line widths while AGN supply the red optical continuum and more extreme line properties. With halo masses $M_{\rm halo} \sim 10^{11-12.5} M_\odot$ and comoving abundance $\sim 2 \times 10^{-5} {\rm cMpc}^{-3}$ (for $\sim 20\%$ duty-cycle at $z \approx 4-8$), ultra-compact galaxies can contribute to the massive, bright LRD population.
△ Less
Submitted 22 June, 2026;
originally announced June 2026.
-
Hybrid Cloud Architectures for Research Computing: Applications and Use Cases
Authors:
Xaver Stiensmeier,
Alexander Kanitz,
Jan Krüger,
Santiago Insua,
Adrián Rošinec,
Viktória Spišáková,
Lukáš Hejtmánek,
David Yuan,
Gavin Farrell,
Jonathan Tedds,
Juha Törnroos,
Harald Wagener,
Alex Sczyrba,
Nils Hoffmann,
Matej Antol
Abstract:
Scientific research increasingly depends on robust and scalable IT infrastructures to support complex computational workflows. With the proliferation of services provided by research infrastructures, NRENs, and commercial cloud providers, researchers must navigate a fragmented ecosystem of computing environments, balancing performance, cost, scalability, and accessibility. Hybrid cloud architectur…
▽ More
Scientific research increasingly depends on robust and scalable IT infrastructures to support complex computational workflows. With the proliferation of services provided by research infrastructures, NRENs, and commercial cloud providers, researchers must navigate a fragmented ecosystem of computing environments, balancing performance, cost, scalability, and accessibility. Hybrid cloud architectures offer a compelling solution by integrating multiple computing environments to enhance flexibility, resource efficiency, and access to specialised hardware.
This paper provides a comprehensive overview of hybrid cloud deployment models, focusing on grid and cloud platforms (OpenPBS, SLURM, OpenStack, Kubernetes) and workflow management tools (Nextflow, Snakemake, CWL). We explore strategies for federated computing, multi-cloud orchestration, and workload scheduling, addressing key challenges such as interoperability, data security, reproducibility, and network performance. Drawing on implementations from life sciences, as coordinated by the ELIXIR Compute Platform and their integration into a wider EOSC context, we propose a roadmap for accelerating hybrid cloud adoption in research computing, emphasising governance frameworks and technical solutions that can drive sustainable and scalable infrastructure development.
△ Less
Submitted 7 January, 2026;
originally announced January 2026.
-
AI Benchmark Democratization and Carpentry
Authors:
Gregor von Laszewski,
Wesley Brewer,
Jeyan Thiyagalingam,
Juri Papay,
Armstrong Foundjem,
Piotr Luszczek,
Murali Emani,
Shirley V. Moore,
Vijay Janapa Reddi,
Matthew D. Sinclair,
Sebastian Lobentanzer,
Sujata Goswami,
Benjamin Hawks,
Marco Colombo,
Nhan Tran,
Christine R. Kirkpatrick,
Abdulkareem Alsudais,
Gregg Barrett,
Tianhao Li,
Kirsten Morehouse,
Shivaram Venkataraman,
Rutwik Jain,
Kartik Mathur,
Victor Lu,
Tejinder Singh
, et al. (6 additional authors not shown)
Abstract:
Benchmarks are a cornerstone of modern machine learning, enabling reproducibility, comparison, and scientific progress. However, AI benchmarks are increasingly complex, requiring dynamic, AI-focused workflows. Rapid evolution in model architectures, scale, datasets, and deployment contexts makes evaluation a moving target. Large language models often memorize static benchmarks, causing a gap betwe…
▽ More
Benchmarks are a cornerstone of modern machine learning, enabling reproducibility, comparison, and scientific progress. However, AI benchmarks are increasingly complex, requiring dynamic, AI-focused workflows. Rapid evolution in model architectures, scale, datasets, and deployment contexts makes evaluation a moving target. Large language models often memorize static benchmarks, causing a gap between benchmark results and real-world performance.
Beyond traditional static benchmarks, continuous adaptive benchmarking frameworks are needed to align scientific assessment with deployment risks. This calls for skills and education in AI Benchmark Carpentry. From our experience with MLCommons, educational initiatives, and programs like the DOE's Trillion Parameter Consortium, key barriers include high resource demands, limited access to specialized hardware, lack of benchmark design expertise, and uncertainty in relating results to application domains. Current benchmarks often emphasize peak performance on top-tier hardware, offering limited guidance for diverse, real-world scenarios.
Benchmarking must become dynamic, incorporating evolving models, updated data, and heterogeneous platforms while maintaining transparency, reproducibility, and interpretability. Democratization requires both technical innovation and systematic education across levels, building sustained expertise in benchmark design and use. Benchmarks should support application-relevant comparisons, enabling informed, context-sensitive decisions. Dynamic, inclusive benchmarking will ensure evaluation keeps pace with AI evolution and supports responsible, reproducible, and accessible AI deployment. Community efforts can provide a foundation for AI Benchmark Carpentry.
△ Less
Submitted 12 December, 2025;
originally announced December 2025.
-
All-$k$-Isolation in Trees
Authors:
Geoffrey Boyer,
Garrett C. Farrell,
Wayne Goddard
Abstract:
We define an all-$k$-isolating set of a graph to be a set $S$ of vertices such that, if one removes $S$ and all its neighbors, then no component in what remains has order $k$ or more. The case $k=1$ corresponds to a dominating set and the case $k=2$ corresponds to what Caro and Hansberg called an isolating set. We show that every tree of order $n \neq k$ contains an all-$k$-isolating set $S$ of si…
▽ More
We define an all-$k$-isolating set of a graph to be a set $S$ of vertices such that, if one removes $S$ and all its neighbors, then no component in what remains has order $k$ or more. The case $k=1$ corresponds to a dominating set and the case $k=2$ corresponds to what Caro and Hansberg called an isolating set. We show that every tree of order $n \neq k$ contains an all-$k$-isolating set $S$ of size at most $n/(k+1)$, and moreover, the set $S$ can be chosen to be an independent set. This extends previous bounds on variations of isolation, while improving a result of Luttrell et al., who called the associated parameter the $k$-neighbor component order connectivity. We also characterize the trees where this bound is achieved. Further, we show that for~$k\le 5$, apart from one exception every tree with $n\neq k$ contains $k+1$ disjoint independent all-$k$-isolating sets.
△ Less
Submitted 15 September, 2025;
originally announced September 2025.
-
Open and Sustainable AI: challenges, opportunities and the road ahead in the life sciences (October 2025 -- Version 2)
Authors:
Gavin Farrell,
Eleni Adamidi,
Rafael Andrade Buono,
Mihail Anton,
Omar Abdelghani Attafi,
Salvador Capella Gutierrez,
Emidio Capriotti,
Leyla Jael Castro,
Davide Cirillo,
Lisa Crossman,
Christophe Dessimoz,
Alexandros Dimopoulos,
Raul Fernandez-Diaz,
Styliani-Christina Fragkouli,
Carole Goble,
Wei Gu,
John M. Hancock,
Alireza Khanteymoori,
Tom Lenaerts,
Fabio G. Liberante,
Peter Maccallum,
Alexander Miguel Monzon,
Magnus Palmblad,
Lucy Poveda,
Ovidiu Radulescu
, et al. (5 additional authors not shown)
Abstract:
Artificial intelligence (AI) has recently seen transformative breakthroughs in the life sciences, expanding possibilities for researchers to interpret biological information at an unprecedented capacity, with novel applications and advances being made almost daily. In order to maximise return on the growing investments in AI-based life science research and accelerate this progress, it has become u…
▽ More
Artificial intelligence (AI) has recently seen transformative breakthroughs in the life sciences, expanding possibilities for researchers to interpret biological information at an unprecedented capacity, with novel applications and advances being made almost daily. In order to maximise return on the growing investments in AI-based life science research and accelerate this progress, it has become urgent to address the exacerbation of long-standing research challenges arising from the rapid adoption of AI methods. We review the increased erosion of trust in AI research outputs, driven by the issues of poor reusability and reproducibility, and highlight their consequent impact on environmental sustainability. Furthermore, we discuss the fragmented components of the AI ecosystem and lack of guiding pathways to best support Open and Sustainable AI (OSAI) model development. In response, this perspective introduces a practical set of OSAI recommendations directly mapped to over 300 components of the AI ecosystem. Our work connects researchers with relevant AI resources, facilitating the implementation of sustainable, reusable and transparent AI. Built upon life science community consensus and aligned to existing efforts, the outputs of this perspective are designed to aid the future development of policy and structured pathways for guiding AI implementation.
△ Less
Submitted 14 October, 2025; v1 submitted 22 May, 2025;
originally announced May 2025.
-
DOME Registry: Implementing community-wide recommendations for reporting supervised machine learning in biology
Authors:
Omar Abdelghani Attafi,
Damiano Clementel,
Konstantinos Kyritsis,
Emidio Capriotti,
Gavin Farrell,
Styliani-Christina Fragkouli,
Leyla Jael Castro,
András Hatos,
Tom Lenaerts,
Stanislav Mazurenko,
Soroush Mozaffari,
Franco Pradelli,
Patrick Ruch,
Castrense Savojardo,
Paola Turina,
Federico Zambelli,
Damiano Piovesan,
Alexander Miguel Monzon,
Fotis Psomopoulos,
Silvio C. E. Tosatto
Abstract:
Supervised machine learning (ML) is used extensively in biology and deserves closer scrutiny. The DOME recommendations aim to enhance the validation and reproducibility of ML research by establishing standards for key aspects such as data handling and processing, optimization, evaluation, and model interpretability. The recommendations help to ensure that key details are reported transparently by…
▽ More
Supervised machine learning (ML) is used extensively in biology and deserves closer scrutiny. The DOME recommendations aim to enhance the validation and reproducibility of ML research by establishing standards for key aspects such as data handling and processing, optimization, evaluation, and model interpretability. The recommendations help to ensure that key details are reported transparently by providing a structured set of questions. Here, we introduce the DOME Registry (URL: registry.dome-ml.org), a database that allows scientists to manage and access comprehensive DOME-related information on published ML studies. The registry uses external resources like ORCID, APICURON and the Data Stewardship Wizard to streamline the annotation process and ensure comprehensive documentation. By assigning unique identifiers and DOME scores to publications, the registry fosters a standardized evaluation of ML methods. Future plans include continuing to grow the registry through community curation, improving the DOME score definition and encouraging publishers to adopt DOME standards, promoting transparency and reproducibility of ML in the life sciences.
△ Less
Submitted 16 August, 2024; v1 submitted 14 August, 2024;
originally announced August 2024.
-
Anomalously slow transport in single-file diffusion with slow binding kinetics
Authors:
Spencer G. Farrell,
Andrew D. Rutenberg
Abstract:
We computationally study the effects of binding kinetics to the channel wall, leading to transient immobility, on the diffusive transport of particles within narrow channels, that exhibit single-file diffusion (SFD). We find that slow binding kinetics leads to an anomalously slow diffusive transport. Remarkably, the scaled diffusivity $\hat{D}$ characterizing transport exhibits scaling collapse wi…
▽ More
We computationally study the effects of binding kinetics to the channel wall, leading to transient immobility, on the diffusive transport of particles within narrow channels, that exhibit single-file diffusion (SFD). We find that slow binding kinetics leads to an anomalously slow diffusive transport. Remarkably, the scaled diffusivity $\hat{D}$ characterizing transport exhibits scaling collapse with respect to the occupation fraction $p$ of sites along the channel. We present a simple "cage-physics" picture that captures the characteristic occupation fraction $p_{scale}$ and the asymptotic $1/p^2$ behavior for $p/p_{scale} \gtrsim 1$. We confirm that subdiffusive behavior of tracer particles is controlled by the same $\hat{D}$ as particle transport.
△ Less
Submitted 1 June, 2018;
originally announced June 2018.
-
Probing the network structure of health deficits in human aging
Authors:
Spencer G. Farrell,
Arnold B. Mitnitski,
Olga Theou,
Kenneth Rockwood,
Andrew D. Rutenberg
Abstract:
We confront a network model of human aging and mortality in which nodes represent health attributes that interact within a scale-free network topology, with observational data that uses both clinical and laboratory (pre-clinical) health deficits as network nodes. We find that individual health attributes exhibit a wide range of mutual information with mortality and that, with a re- construction of…
▽ More
We confront a network model of human aging and mortality in which nodes represent health attributes that interact within a scale-free network topology, with observational data that uses both clinical and laboratory (pre-clinical) health deficits as network nodes. We find that individual health attributes exhibit a wide range of mutual information with mortality and that, with a re- construction of their relative connectivity, higher-ranked nodes are more informative. Surprisingly, we find a broad and overlapping range of mutual information of laboratory measures as compared with clinical measures. We confirm similar behavior between most-connected and least-connected model nodes, controlled by the nearest-neighbor connectivity. Furthermore, in both model and observational data, we find that the least-connected (laboratory) nodes damage earlier than the most-connected (clinical) deficits. A mean-field theory of our network model captures and explains this phenomenon, which results from the connectivity of nodes and of their connected neighbors. We find that other network topologies, including random, small-world, and assortative scale-free net- works, exhibit qualitatively different behavior. Our disassortative scale-free network model behaves consistently with our expanded phenomenology observed in human aging, and so is a useful tool to explore mechanisms of and to develop new predictive measures for human aging and mortality.
△ Less
Submitted 6 June, 2018; v1 submitted 23 February, 2018;
originally announced February 2018.
-
Loose packings of frictional spheres
Authors:
Greg R. Farrell,
K. Michael Martini,
Narayanan Menon
Abstract:
We have produced loose packings of cohesionless, frictional spheres by sequential deposition of highly-spherical, monodisperse particles through a fluid. By varying the properties of the fluid and the particles, we have identified the Stokes number (St) - rather than the buoyancy of the particles in the fluid - as the parameter controlling the approach to the loose packing limit. The loose packing…
▽ More
We have produced loose packings of cohesionless, frictional spheres by sequential deposition of highly-spherical, monodisperse particles through a fluid. By varying the properties of the fluid and the particles, we have identified the Stokes number (St) - rather than the buoyancy of the particles in the fluid - as the parameter controlling the approach to the loose packing limit. The loose packing limit is attained at a threshold value of St at which the kinetic energy of a particle impinging on the packing is fully dissipated by the fluid. Thus, for cohesionless particles, the dynamics of the deposition process, rather than the stability of the static packing, defines the random loose packing limit. We have made direct measurements of the interparticle friction in the fluid, and present an experimental measurement of the loose packing volume fraction, φ_{RLP}, as a function of the friction coefficient μ_s.
△ Less
Submitted 5 May, 2010;
originally announced May 2010.
-
Critical Dynamics of Dimers: Implications for the Glass Transition
Authors:
Dibyendu Das,
Greg Farrell,
Jane' Kondev,
Bulbul Chakraborty
Abstract:
The Adam-Gibbs view of the glass transition relates the relaxation time to the configurational entropy, which goes continuously to zero at the so-called Kauzmann temperature. We examine this scenario in the context of a dimer model with an entropy vanishing phase transition, and stochastic loop dynamics. We propose a coarse-grained master equation for the order parameter dynamics which is used t…
▽ More
The Adam-Gibbs view of the glass transition relates the relaxation time to the configurational entropy, which goes continuously to zero at the so-called Kauzmann temperature. We examine this scenario in the context of a dimer model with an entropy vanishing phase transition, and stochastic loop dynamics. We propose a coarse-grained master equation for the order parameter dynamics which is used to compute the time-dependent autocorrelation function and the associated relaxation time. Using a combination of exact results, scaling arguments and numerical diagonalizations of the master equation, we find non-exponential relaxation and a Vogel-Fulcher divergence of the relaxation time in the vicinity of the phase transition. Since in the dimer model the entropy stays finite all the way to the phase transition point, and then jumps discontinuously to zero, we demonstrate a clear departure from the Adam-Gibbs scenario. Dimer coverings are the "inherent structures" of the canonical frustrated system, the triangular Ising antiferromagnet. Therefore, our results provide a new scenario for the glass transition in supercooled liquids in terms of inherent structure dynamics.
△ Less
Submitted 22 June, 2005;
originally announced June 2005.