-
Harnessing X-ray Absorption Spectroscopy Data through Multimodal Mining of Battery Literature
Authors:
Tanjin He,
Aikaterini Vriza,
Logan Ward,
Xu Huang,
Yiming Chen,
Anubhav Jain,
Gerbrand Ceder,
Rajeev S. Assary,
Ian T. Foster,
Maria K. Y. Chan
Abstract:
X-ray absorption spectroscopy (XAS) is central to understanding the local electronic and atomic structure of materials, yet most published spectra remain inaccessible to data-driven analysis because they are embedded in figures and described through fragmented textual context in the literature. Here, we use multimodal (image and text) literature mining to transform this dispersed knowledge into an…
▽ More
X-ray absorption spectroscopy (XAS) is central to understanding the local electronic and atomic structure of materials, yet most published spectra remain inaccessible to data-driven analysis because they are embedded in figures and described through fragmented textual context in the literature. Here, we use multimodal (image and text) literature mining to transform this dispersed knowledge into an AI-ready experimental data resource. We developed a scalable spectroscopy data digitization pipeline that identifies XAS figures in full-text articles, digitizes spectral curves, and links each spectrum to accompanying metadata on the measured edge and material. Applying this pipeline to the battery literature produced an open dataset of 13,740 XAS spectra, spanning 66 absorbing elements and diverse battery chemistries, with expert validation confirming accurate extraction of spectral and metadata information. By converting literature-embedded spectra into structured numerical data, this dataset provides a foundation for large-scale XAS analysis, cross-laboratory comparison, high-throughput characterization, and autonomous discovery of advanced materials.
△ Less
Submitted 30 July, 2026; v1 submitted 26 July, 2026;
originally announced July 2026.
-
Steering an Active Learning Workflow Towards Novel Materials Discovery via Queue Prioritization
Authors:
Marcus Schwarting,
Logan Ward,
Nathaniel Hudson,
Xiaoli Yan,
Ben Blaiszik,
Santanu Chaudhuri,
Eliu Huerta,
Ian Foster
Abstract:
Generative AI poses both opportunities and risks for solving inverse design problems in the sciences. Generative tools provide the ability to expand and refine a search space autonomously, but do so at the cost of exploring low-quality regions until sufficiently fine tuned. Here, we propose a queue prioritization algorithm that combines generative modeling and active learning in the context of a d…
▽ More
Generative AI poses both opportunities and risks for solving inverse design problems in the sciences. Generative tools provide the ability to expand and refine a search space autonomously, but do so at the cost of exploring low-quality regions until sufficiently fine tuned. Here, we propose a queue prioritization algorithm that combines generative modeling and active learning in the context of a distributed workflow for exploring complex design spaces. We find that incorporating an active learning model to prioritize top design candidates can prevent a generative AI workflow from expending resources on nonsensical candidates and halt potential generative model decay. For an existing generative AI workflow for discovering novel molecular structure candidates for carbon capture, our active learning approach significantly increases the number of high-quality candidates identified by the generative model. We find that, out of 1000 novel candidates, our workflow without active learning can generate an average of 281 high-performing candidates, while our proposed prioritization with active learning can generate an average 604 high-performing candidates.
△ Less
Submitted 29 September, 2025;
originally announced September 2025.
-
MOFA: Discovering Materials for Carbon Capture with a GenAI- and Simulation-Based Workflow
Authors:
Xiaoli Yan,
Nathaniel Hudson,
Hyun Park,
Daniel Grzenda,
J. Gregory Pauloski,
Marcus Schwarting,
Haochen Pan,
Hassan Harb,
Samuel Foreman,
Chris Knight,
Tom Gibbs,
Kyle Chard,
Santanu Chaudhuri,
Emad Tajkhorshid,
Ian Foster,
Mohamad Moosavi,
Logan Ward,
E. A. Huerta
Abstract:
We present MOFA, an open-source generative AI (GenAI) plus simulation workflow for high-throughput generation of metal-organic frameworks (MOFs) on large-scale high-performance computing (HPC) systems. MOFA addresses key challenges in integrating GPU-accelerated computing for GPU-intensive GenAI tasks, including distributed training and inference, alongside CPU- and GPU-optimized tasks for screeni…
▽ More
We present MOFA, an open-source generative AI (GenAI) plus simulation workflow for high-throughput generation of metal-organic frameworks (MOFs) on large-scale high-performance computing (HPC) systems. MOFA addresses key challenges in integrating GPU-accelerated computing for GPU-intensive GenAI tasks, including distributed training and inference, alongside CPU- and GPU-optimized tasks for screening and filtering AI-generated MOFs using molecular dynamics, density functional theory, and Monte Carlo simulations. These heterogeneous tasks are unified within an online learning framework that optimizes the utilization of available CPU and GPU resources across HPC systems. Performance metrics from a 450-node (14,400 AMD Zen 3 CPUs + 1800 NVIDIA A100 GPUs) supercomputer run demonstrate that MOFA achieves high-throughput generation of novel MOF structures, with CO$_2$ adsorption capacities ranking among the top 10 in the hypothetical MOF (hMOF) dataset. Furthermore, the production of high-quality MOFs exhibits a linear relationship with the number of nodes utilized. The modular architecture of MOFA will facilitate its integration into other scientific applications that dynamically combine GenAI with large-scale simulations.
△ Less
Submitted 17 January, 2025;
originally announced January 2025.
-
Dirac-Schwinger Quantization for Emergent Magnetic Monopoles?
Authors:
A. Farhan,
M. Saccone,
B. F. L. Ward
Abstract:
In Refs.[1-4] Dirac and Schwinger showed the existence of a magnetic monopole required a charge quantization condition which we write following Dirac as $\frac{eg}{4π\hbar}=\frac{n}{2},\; n=0,\pm 1,\; \pm 2, \ldots$. Here, $g$ is the magnetic monopole charge and $e$ is the electric charge of the positron. Recently, in Refs. [5,6], it has been shown experimentally that frustrated spin-ice systems e…
▽ More
In Refs.[1-4] Dirac and Schwinger showed the existence of a magnetic monopole required a charge quantization condition which we write following Dirac as $\frac{eg}{4π\hbar}=\frac{n}{2},\; n=0,\pm 1,\; \pm 2, \ldots$. Here, $g$ is the magnetic monopole charge and $e$ is the electric charge of the positron. Recently, in Refs. [5,6], it has been shown experimentally that frustrated spin-ice systems exhibit 'emergent' magnetic monopoles. We show that, within the experimental errors, the respective magnetic charges obey the Dirac-Schwinger quantization condition. Possible implications are discussed.
△ Less
Submitted 28 April, 2025; v1 submitted 23 December, 2024;
originally announced January 2025.
-
A case study of multi-modal, multi-institutional data management for the combinatorial materials science community
Authors:
Sarah I. Allec,
Eric S. Muckley,
Nathan S. Johnson,
Christopher K. H. Borg,
Dylan J. Kirsch,
Joshua Martin,
Rohit Pant,
Ichiro Takeuchi,
Andrew S. Lee,
James E. Saal,
Logan Ward,
Apurva Mehta
Abstract:
Although the convergence of high-performance computing, automation, and machine learning has significantly altered the materials design timeline, transformative advances in functional materials and acceleration of their design will require addressing the deficiencies that currently exist in materials informatics, particularly a lack of standardized experimental data management. The challenges asso…
▽ More
Although the convergence of high-performance computing, automation, and machine learning has significantly altered the materials design timeline, transformative advances in functional materials and acceleration of their design will require addressing the deficiencies that currently exist in materials informatics, particularly a lack of standardized experimental data management. The challenges associated with experimental data management are especially true for combinatorial materials science, where advancements in automation of experimental workflows have produced datasets that are often too large and too complex for human reasoning. The data management challenge is further compounded by the multi-modal and multi-institutional nature of these datasets, as they tend to be distributed across multiple institutions and can vary substantially in format, size, and content. To adequately map a materials design space from such datasets, an ideal materials data infrastructure would contain data and metadata describing i) synthesis and processing conditions, ii) characterization results, and iii) property and performance measurements. Here, we present a case study for the low-barrier development of such a dashboard that enables standardized organization, analysis, and visualization of a large data lake consisting of combinatorial datasets of synthesis and processing conditions, X-ray diffraction patterns, and materials property measurements generated at several different institutions. While this dashboard was developed specifically for data-driven thermoelectric materials discovery, we envision the adaptation of this prototype to other materials applications, and, more ambitiously, future integration into an all-encompassing materials data management infrastructure.
△ Less
Submitted 6 February, 2024; v1 submitted 16 November, 2023;
originally announced November 2023.
-
Accelerating Electronic Stopping Power Predictions by 10 Million Times with a Combination of Time-Dependent Density Functional Theory and Machine Learning
Authors:
Logan Ward,
Ben Blaiszik,
Cheng-Wei Lee,
Troy Martin,
Ian Foster,
André Schleife
Abstract:
Knowing the rate at which particle radiation releases energy in a material, the stopping power, is key to designing nuclear reactors, medical treatments, semiconductor and quantum materials, and many other technologies. While the nuclear contribution to stopping power, i.e., elastic scattering between atoms, is well understood in the literature, the route for gathering data on the electronic contr…
▽ More
Knowing the rate at which particle radiation releases energy in a material, the stopping power, is key to designing nuclear reactors, medical treatments, semiconductor and quantum materials, and many other technologies. While the nuclear contribution to stopping power, i.e., elastic scattering between atoms, is well understood in the literature, the route for gathering data on the electronic contribution has for decades remained costly and reliant on many simplifying assumptions, including that materials are isotropic. We establish a method that combines time-dependent density functional theory (TDDFT) and machine learning to reduce the time to assess new materials to mere hours on a supercomputer and provides valuable data on how atomic details influence electronic stopping. Our approach uses TDDFT to compute the electronic stopping contributions to stopping power from first principles in several directions and then machine learning to interpolate to other directions at a cost of 10 million times fewer core-hours. We demonstrate the combined approach in a study of proton irradiation in aluminum and employ it to predict how the depth of maximum energy deposition, the "Bragg Peak," varies depending on incident angle -- a quantity otherwise inaccessible to modelers. The lack of any experimental information requirement makes our method applicable to most materials, and its speed makes it a prime candidate for enabling quantum-to-continuum models of radiation damage. The prospect of reusing valuable TDDFT data for training the model make our approach appealing for applications in the age of materials data science.
△ Less
Submitted 25 June, 2024; v1 submitted 1 November, 2023;
originally announced November 2023.
-
Reproducibility in Computational Materials Science: Lessons from 'A General-Purpose Machine Learning Framework for Predicting Properties of Inorganic Materials'
Authors:
Daniel Persaud,
Logan Ward,
Jason Hattrick-Simpers
Abstract:
The integration of machine learning techniques in materials discovery has become prominent in materials science research and has been accompanied by an increasing trend towards open-source data and tools to propel the field. Despite the increasing usefulness and capabilities of these tools, developers neglecting to follow reproducible practices creates a significant barrier for researchers looking…
▽ More
The integration of machine learning techniques in materials discovery has become prominent in materials science research and has been accompanied by an increasing trend towards open-source data and tools to propel the field. Despite the increasing usefulness and capabilities of these tools, developers neglecting to follow reproducible practices creates a significant barrier for researchers looking to use or build upon their work. In this study, we investigate the challenges encountered while attempting to reproduce a section of the results presented in "A general-purpose machine learning framework for predicting properties of inorganic materials." Our analysis identifies four major categories of challenges: (1) reporting computational dependencies, (2) recording and sharing version logs, (3) sequential code organization, and (4) clarifying code references within the manuscript. The result is a proposed set of tangible action items for those aiming to make code accessible to, and useful for the community.
△ Less
Submitted 10 October, 2023;
originally announced October 2023.
-
14 Examples of How LLMs Can Transform Materials Science and Chemistry: A Reflection on a Large Language Model Hackathon
Authors:
Kevin Maik Jablonka,
Qianxiang Ai,
Alexander Al-Feghali,
Shruti Badhwar,
Joshua D. Bocarsly,
Andres M Bran,
Stefan Bringuier,
L. Catherine Brinson,
Kamal Choudhary,
Defne Circi,
Sam Cox,
Wibe A. de Jong,
Matthew L. Evans,
Nicolas Gastellu,
Jerome Genzling,
María Victoria Gil,
Ankur K. Gupta,
Zhi Hong,
Alishba Imran,
Sabine Kruschwitz,
Anne Labarre,
Jakub Lála,
Tao Liu,
Steven Ma,
Sauradeep Majumdar
, et al. (28 additional authors not shown)
Abstract:
Large-language models (LLMs) such as GPT-4 caught the interest of many scientists. Recent studies suggested that these models could be useful in chemistry and materials science. To explore these possibilities, we organized a hackathon.
This article chronicles the projects built as part of this hackathon. Participants employed LLMs for various applications, including predicting properties of mole…
▽ More
Large-language models (LLMs) such as GPT-4 caught the interest of many scientists. Recent studies suggested that these models could be useful in chemistry and materials science. To explore these possibilities, we organized a hackathon.
This article chronicles the projects built as part of this hackathon. Participants employed LLMs for various applications, including predicting properties of molecules and materials, designing novel interfaces for tools, extracting knowledge from unstructured data, and developing new educational applications.
The diverse topics and the fact that working prototypes could be generated in less than two days highlight that LLMs will profoundly impact the future of our fields. The rich collection of ideas and projects also indicates that the applications of LLMs are not limited to materials science and chemistry but offer potential benefits to a wide range of scientific disciplines.
△ Less
Submitted 14 July, 2023; v1 submitted 9 June, 2023;
originally announced June 2023.
-
Machine Learning Prediction of Critical Cooling Rate for Metallic Glasses From Expanded Datasets and Elemental Features
Authors:
Benjamin T. Afflerbach,
Carter Francis,
Lane E. Schultz,
Janine Spethson,
Vanessa Meschke,
Elliot Strand,
Logan Ward,
John H. Perepezko,
Dan Thoma,
Paul M. Voyles,
Izabela Szlufarska,
Dane Morgan
Abstract:
We use a random forest model to predict the critical cooling rate (RC) for glass formation of various alloys from features of their constituent elements. The random forest model was trained on a database that integrates multiple sources of direct and indirect RC data for metallic glasses to expand the directly measured RC database of less than 100 values to a training set of over 2,000 values. The…
▽ More
We use a random forest model to predict the critical cooling rate (RC) for glass formation of various alloys from features of their constituent elements. The random forest model was trained on a database that integrates multiple sources of direct and indirect RC data for metallic glasses to expand the directly measured RC database of less than 100 values to a training set of over 2,000 values. The model error on 5-fold cross validation is 0.66 orders of magnitude in K/s. The error on leave out one group cross validation on alloy system groups is 0.59 log units in K/s when the target alloy constituents appear more than 500 times in training data. Using this model, we make predictions for the set of compositions with melt-spun glasses in the database, and for the full set of quaternary alloys that have constituents which appear more than 500 times in training data. These predictions identify a number of potential new bulk metallic glass (BMG) systems for future study, but the model is most useful for identification of alloy systems likely to contain good glass formers, rather than detailed discovery of bulk glass composition regions within known glassy systems.
△ Less
Submitted 24 May, 2023;
originally announced May 2023.
-
Quantifying the performance of machine learning models in materials discovery
Authors:
Christopher K. H. Borg,
Eric S. Muckley,
Clara Nyby,
James E. Saal,
Logan Ward,
Apurva Mehta,
Bryce Meredig
Abstract:
The predictive capabilities of machine learning (ML) models used in materials discovery are typically measured using simple statistics such as the root-mean-square error (RMSE) or the coefficient of determination ($r^2$) between ML-predicted materials property values and their known values. A tempting assumption is that models with low error should be effective at guiding materials discovery, and…
▽ More
The predictive capabilities of machine learning (ML) models used in materials discovery are typically measured using simple statistics such as the root-mean-square error (RMSE) or the coefficient of determination ($r^2$) between ML-predicted materials property values and their known values. A tempting assumption is that models with low error should be effective at guiding materials discovery, and conversely, models with high error should give poor discovery performance. However, we observe that no clear connection exists between a "static" quantity averaged across an entire training set, such as RMSE, and an ML property model's ability to dynamically guide the iterative (and often extrapolative) discovery of novel materials with targeted properties. In this work, we simulate a sequential learning (SL)-guided materials discovery process and demonstrate a decoupling between traditional model error metrics and model performance in guiding materials discoveries. We show that model performance in materials discovery depends strongly on (1) the target range within the property distribution (e.g., whether a 1st or 10th decile material is desired); (2) the incorporation of uncertainty estimates in the SL acquisition function; (3) whether the scientist is interested in one discovery or many targets; and (4) how many SL iterations are allowed. To overcome the limitations of static metrics and robustly capture SL performance, we recommend metrics such as Discovery Yield ($DY$), a measure of how many high-performing materials were discovered during SL, and Discovery Probability ($DP$), a measure of likelihood of discovering high-performing materials at any point in the SL process.
△ Less
Submitted 24 October, 2022;
originally announced October 2022.
-
Rapid Production of Accurate Embedded-Atom Method Potentials for Metal Alloys
Authors:
Elan J. Weiss,
Logan Ward,
Christian Oberdorfer,
Travis Withrow,
David C. Riegner,
Anupriya Agrawal,
Wolfgang Windl
Abstract:
A critical limitation to the wide-scale use of classical molecular dynamics for alloy design is the limited availability of suitable interatomic potentials. Here, we introduce the Rapid Alloy Method for Producing Accurate General Empirical Potentials or RAMPAGE, a computationally economical procedure to generate binary embedded-atom model potentials from already-existing single-element potentials…
▽ More
A critical limitation to the wide-scale use of classical molecular dynamics for alloy design is the limited availability of suitable interatomic potentials. Here, we introduce the Rapid Alloy Method for Producing Accurate General Empirical Potentials or RAMPAGE, a computationally economical procedure to generate binary embedded-atom model potentials from already-existing single-element potentials that can be further combined into multi-component alloy potentials. We present the quality of RAMPAGE calibrated Finnis-Sinclair type EAM potentials using binary Ag-Al and ternary Ag-Au-Cu as case studies. We demonstrate that RAMPAGE potentials can reproduce bulk properties and forces with greater accuracy than that of other alloy potentials. In some simulations, it is observed the quality of the optimized cross interactions can exceed that of the original off-the-shelf elemental potential inputs.
△ Less
Submitted 3 August, 2022;
originally announced August 2022.
-
Mapping Thermoelectric Transport in a Multicomponent Alloy Space
Authors:
Ramya Gurunathan,
Suchismita Sarker,
Christopher K. H. Borg,
James Saal,
Logan Ward,
Apurva Mehta,
G. Jeffrey Snyder
Abstract:
Interest in high entropy alloy thermoelectric materials is predicated on achieving ultralow lattice thermal conductivity $κ\sub{L}$ through large compositional disorder. However, here we show that for a given mechanism, such as mass contrast phonon scattering, $κ\sub{L}$ will be minimized along the binary alloy with the highest mass contrast, such that adding an intermediate-mass atom to increase…
▽ More
Interest in high entropy alloy thermoelectric materials is predicated on achieving ultralow lattice thermal conductivity $κ\sub{L}$ through large compositional disorder. However, here we show that for a given mechanism, such as mass contrast phonon scattering, $κ\sub{L}$ will be minimized along the binary alloy with the highest mass contrast, such that adding an intermediate-mass atom to increase atomic disorder can increase thermal conductivity. Only when each component adds an independent scattering mechanism (such as adding strain fluctuation to an existing mass fluctuation) is there a benefit. In addition, both charge carriers and heat-carrying phonons are known to experience scattering due to alloying effects, leading to a trade-off in thermoelectric performance. We apply analytic transport models, based on perturbation and effective medium theories, to predict how alloy scattering will affect the thermal and electronic transport across the full compositional range of several pseudo-ternary and pseudo-quaternary alloy systems. To do so, we demonstrate a multicomponent extension to both thermal and electronic binary alloy scattering models based on the virtual crystal approximation. Finally, we show that common functional forms used in computational thermodynamics can be applied to this problem to further generalize the scattering behavior that is modeled.
△ Less
Submitted 3 May, 2022;
originally announced May 2022.
-
Colmena: Scalable Machine-Learning-Based Steering of Ensemble Simulations for High Performance Computing
Authors:
Logan Ward,
Ganesh Sivaraman,
J. Gregory Pauloski,
Yadu Babuji,
Ryan Chard,
Naveen Dandu,
Paul C. Redfern,
Rajeev S. Assary,
Kyle Chard,
Larry A. Curtiss,
Rajeev Thakur,
Ian Foster
Abstract:
Scientific applications that involve simulation ensembles can be accelerated greatly by using experiment design methods to select the best simulations to perform. Methods that use machine learning (ML) to create proxy models of simulations show particular promise for guiding ensembles but are challenging to deploy because of the need to coordinate dynamic mixes of simulation and learning tasks. We…
▽ More
Scientific applications that involve simulation ensembles can be accelerated greatly by using experiment design methods to select the best simulations to perform. Methods that use machine learning (ML) to create proxy models of simulations show particular promise for guiding ensembles but are challenging to deploy because of the need to coordinate dynamic mixes of simulation and learning tasks. We present Colmena, an open-source Python framework that allows users to steer campaigns by providing just the implementations of individual tasks plus the logic used to choose which tasks to execute when. Colmena handles task dispatch, results collation, ML model invocation, and ML model (re)training, using Parsl to execute tasks on HPC systems. We describe the design of Colmena and illustrate its capabilities by applying it to electrolyte design, where it both scales to 65536 CPUs and accelerates the discovery rate for high-performance molecules by a factor of 100 over unguided searches.
△ Less
Submitted 6 October, 2021;
originally announced October 2021.
-
A high-throughput structural and electrochemical study of metallic glass formation in Ni-Ti-Al
Authors:
Howie Joress,
Brian L. DeCost,
Suchismita Sarker,
Trevor M. Braun,
Sidra Jilani,
Ryan Smith,
Logan Ward,
Kevin J. Laws,
Apurva Mehta,
Jason Hattrick-Simpers
Abstract:
Based on a set of machine learning predictions of glass formation in the Ni-Ti-Al system, we have undertaken a high-throughput experimental study of that system. We utilized rapid synthesis followed by high-throughput structural and electrochemical characterization. Using this dual-modality approach, we are able to better classify the amorphous portion of the library, which we found to be the port…
▽ More
Based on a set of machine learning predictions of glass formation in the Ni-Ti-Al system, we have undertaken a high-throughput experimental study of that system. We utilized rapid synthesis followed by high-throughput structural and electrochemical characterization. Using this dual-modality approach, we are able to better classify the amorphous portion of the library, which we found to be the portion with a full-width-half-maximum (FWHM) of 0.42 A$^{-1}$ for the first sharp x-ray diffraction peak. We demonstrate that the FWHM and corrosion resistance are correlated but that, while chemistry still plays a role, a large FWHM is necessary for the best corrosion resistance.
△ Less
Submitted 19 December, 2019;
originally announced December 2019.
-
Machine Learning Prediction of Accurate Atomization Energies of Organic Molecules from Low-Fidelity Quantum Chemical Calculations
Authors:
Logan Ward,
Ben Blaiszik,
Ian Foster,
Rajeev S. Assary,
Badri Narayanan,
Larry Curtiss
Abstract:
Recent studies illustrate how machine learning (ML) can be used to bypass a core challenge of molecular modeling: the tradeoff between accuracy and computational cost. Here, we assess multiple ML approaches for predicting the atomization energy of organic molecules. Our resulting models learn the difference between low-fidelity, B3LYP, and high-accuracy, G4MP2, atomization energies, and predict th…
▽ More
Recent studies illustrate how machine learning (ML) can be used to bypass a core challenge of molecular modeling: the tradeoff between accuracy and computational cost. Here, we assess multiple ML approaches for predicting the atomization energy of organic molecules. Our resulting models learn the difference between low-fidelity, B3LYP, and high-accuracy, G4MP2, atomization energies, and predict the G4MP2 atomization energy to 0.005 eV (mean absolute error) for molecules with less than 9 heavy atoms and 0.012 eV for a small set of molecules with between 10 and 14 heavy atoms. Our two best models, which have different accuracy/speed tradeoffs, enable the efficient prediction of G4MP2-level energies for large molecules and are available through a simple web interface.
△ Less
Submitted 7 June, 2019;
originally announced June 2019.
-
A Data Ecosystem to Support Machine Learning in Materials Science
Authors:
Ben Blaiszik,
Logan Ward,
Marcus Schwarting,
Jonathon Gaff,
Ryan Chard,
Daniel Pike,
Kyle Chard,
Ian Foster
Abstract:
Facilitating the application of machine learning to materials science problems will require enhancing the data ecosystem to enable discovery and collection of data from many sources, automated dissemination of new data across the ecosystem, and the connecting of data with materials-specific machine learning models. Here, we present two projects, the Materials Data Facility (MDF) and the Data and L…
▽ More
Facilitating the application of machine learning to materials science problems will require enhancing the data ecosystem to enable discovery and collection of data from many sources, automated dissemination of new data across the ecosystem, and the connecting of data with materials-specific machine learning models. Here, we present two projects, the Materials Data Facility (MDF) and the Data and Learning Hub for Science (DLHub), that address these needs. We use examples to show how MDF and DLHub capabilities can be leveraged to link data with machine learning models and how users can access those capabilities through web and programmatic interfaces.
△ Less
Submitted 20 July, 2019; v1 submitted 23 April, 2019;
originally announced April 2019.
-
Ternary mixed-anion semiconductors with tunable band gaps from machine-learning and crystal structure prediction
Authors:
Maximilian Amsler,
Logan Ward,
Vinay I. Hegde,
Maarten G. Goesten,
Xia Yi,
Chris Wolverton
Abstract:
We report the computational investigation of a series of ternary X$_4$Y$_2$Z and X$_5$Y$_2$Z$_2$ compounds with X={Mg, Ca, Sr, Ba}, Y={P, As, Sb, Bi}, and Z={S, Se, Te}. The compositions for these materials were predicted through a search guided by machine learning, while the structures were resolved using the minima hopping crystal structure prediction method. Based on $\textit{ab initio}$ calcul…
▽ More
We report the computational investigation of a series of ternary X$_4$Y$_2$Z and X$_5$Y$_2$Z$_2$ compounds with X={Mg, Ca, Sr, Ba}, Y={P, As, Sb, Bi}, and Z={S, Se, Te}. The compositions for these materials were predicted through a search guided by machine learning, while the structures were resolved using the minima hopping crystal structure prediction method. Based on $\textit{ab initio}$ calculations, we predict that many of these compounds are thermodynamically stable. In particular, 21 of the X$_4$Y$_2$Z compounds crystallize in a tetragonal structure with $\textit{I-42d}$ symmetry, and exhibit band gaps in the range of 0.3 and 1.8 eV, well suited for various energy applications. We show that several candidate compounds (in particular X$_4$Y$_2$Te and X$_4$Sb$_2$Se) exhibit good photo absorption in the visible range, while others (e.g., Ba$_4$Sb$_2$Se) show excellent thermoelectric performance due to a high power factor and extremely low lattice thermal conductivities.
△ Less
Submitted 6 December, 2018;
originally announced December 2018.
-
A General-Purpose Machine Learning Framework for Predicting Properties of Inorganic Materials
Authors:
Logan Ward,
Ankit Agrawal,
Alok Choudhary,
Christopher Wolverton
Abstract:
A very active area of materials research is to devise methods that use machine learning to automatically extract predictive models from existing materials data. While prior examples have demonstrated successful models for some applications, many more applications exist where machine learning can make a strong impact. To enable faster development of machine-learning-based models for such applicatio…
▽ More
A very active area of materials research is to devise methods that use machine learning to automatically extract predictive models from existing materials data. While prior examples have demonstrated successful models for some applications, many more applications exist where machine learning can make a strong impact. To enable faster development of machine-learning-based models for such applications, we have created a framework capable of being applied to a broad range of materials data. Our method works by using a chemically diverse list of attributes, which we demonstrate are suitable for describing a wide variety of properties, and a novel method for partitioning the data set into groups of similar materials in order to boost the predictive accuracy. In this manuscript, we demonstrate how this new method can be used to predict diverse properties of crystalline and amorphous materials, such as band gap energy and glass-forming ability.
△ Less
Submitted 1 July, 2016; v1 submitted 30 June, 2016;
originally announced June 2016.
-
Structural evolution and kinetics in Cu-Zr Metallic Liquids
Authors:
Logan Ward,
Dan Miracle,
Wolfgang Windl,
Oleg Senkov,
Katharine Flores
Abstract:
The atomic structure of the supercooled liquid has often been discussed as a key source of glass formation in metals. The presence of icosahedrally-coordinated clusters and their tendency to form networks have been identified as one possible structural trait leading to glass forming ability in the Cu-Zr binary system. In this work, we show that this theory is insufficient to explain glass formatio…
▽ More
The atomic structure of the supercooled liquid has often been discussed as a key source of glass formation in metals. The presence of icosahedrally-coordinated clusters and their tendency to form networks have been identified as one possible structural trait leading to glass forming ability in the Cu-Zr binary system. In this work, we show that this theory is insufficient to explain glass formation at all compositions in that binary system. Instead, we propose that the formation of ideally-packed clusters at the expense of atomic arrangements with excess or deficient free volume can explain glass-forming by a similar mechanism. We show that this behavior is reflected in the structural relaxation of a metallic glass during constant pressure cooling and the time evolution of structure at a constant volume. We then demonstrate that this theory is sufficient to explain slowed diffusivity in compositions across the range of Cu-Zr metallic glasses.
△ Less
Submitted 9 May, 2013;
originally announced May 2013.
-
Rapid Production of Accurate Embedded-Atom Method Potentials for Metal Alloys
Authors:
Logan Ward,
Anupriya Agrawal,
Katharine M. Flores,
Wolfgang Windl
Abstract:
The most critical limitation to the wide-scale use of classical molecular dynamics for alloy design is the availability of suitable interatomic potentials. In this work, we demonstrate a simple procedure to generate a library of accurate binary potentials using already-existing single-element potentials that can be easily combined to form multi-component alloy potentials. For the Al-Ni, Cu-Au, and…
▽ More
The most critical limitation to the wide-scale use of classical molecular dynamics for alloy design is the availability of suitable interatomic potentials. In this work, we demonstrate a simple procedure to generate a library of accurate binary potentials using already-existing single-element potentials that can be easily combined to form multi-component alloy potentials. For the Al-Ni, Cu-Au, and Cu-Al-Zr systems, we show that this method produces results comparable in accuracy to alloy potentials where all parts have been fitted simultaneously, without the additional computational expense. Furthermore, we demonstrate applicability to both crystalline and amorphous phases.
△ Less
Submitted 4 September, 2012;
originally announced September 2012.