-
Omics-scale polymer computational database transferable to real-world artificial intelligence applications
Authors:
Ryo Yoshida,
Yoshihiro Hayashi,
Hidemine Furuya,
Ryohei Hosoya,
Kazuyoshi Kaneko,
Hiroki Sugisawa,
Yu Kaneko,
Aiko Takahashi,
Yoh Noguchi,
Shun Nanjo,
Keiko Shinoda,
Tomu Hamakawa,
Mitsuru Ohno,
Takuya Kitamura,
Misaki Yonekawa,
Stephen Wu,
Masato Ohnishi,
Chang Liu,
Teruki Tsurimoto,
Arifin,
Araki Wakiuchi,
Kohei Noda,
Junko Morikawa,
Teruaki Hayakawa,
Junichiro Shiomi
, et al. (81 additional authors not shown)
Abstract:
Developing large-scale foundational datasets is a critical milestone in advancing artificial intelligence (AI)-driven scientific innovation. However, unlike AI-mature fields such as natural language processing, materials science, particularly polymer research, has significantly lagged in developing extensive open datasets. This lag is primarily due to the high costs of polymer synthesis and proper…
▽ More
Developing large-scale foundational datasets is a critical milestone in advancing artificial intelligence (AI)-driven scientific innovation. However, unlike AI-mature fields such as natural language processing, materials science, particularly polymer research, has significantly lagged in developing extensive open datasets. This lag is primarily due to the high costs of polymer synthesis and property measurements, along with the vastness and complexity of the chemical space. This study presents PolyOmics, an omics-scale computational database generated through fully automated molecular dynamics simulation pipelines that provide diverse physical properties for over $10^5$ polymeric materials. The PolyOmics database is collaboratively developed by approximately 260 researchers from 48 institutions to bridge the gap between academia and industry. Machine learning models pretrained on PolyOmics can be efficiently fine-tuned for a wide range of real-world downstream tasks, even when only limited experimental data are available. Notably, the generalisation capability of these simulation-to-real transfer models improve significantly as the size of the PolyOmics database increases, exhibiting power-law scaling. The emergence of scaling laws supports the "more is better" principle, highlighting the significance of ultralarge-scale computational materials data for improving real-world prediction performance. This unprecedented omics-scale database reveals vast unexplored regions of polymer materials, providing a foundation for AI-driven polymer science.
△ Less
Submitted 7 November, 2025;
originally announced November 2025.
-
Scaling Law of Sim2Real Transfer Learning in Expanding Computational Materials Databases for Real-World Predictions
Authors:
Shunya Minami,
Yoshihiro Hayashi,
Stephen Wu,
Kenji Fukumizu,
Hiroki Sugisawa,
Masashi Ishii,
Isao Kuwajima,
Kazuya Shiratori,
Ryo Yoshida
Abstract:
To address the challenge of limited experimental materials data, extensive physical property databases are being developed based on high-throughput computational experiments, such as molecular dynamics simulations. Previous studies have shown that fine-tuning a predictor pretrained on a computational database to a real system can result in models with outstanding generalization capabilities compar…
▽ More
To address the challenge of limited experimental materials data, extensive physical property databases are being developed based on high-throughput computational experiments, such as molecular dynamics simulations. Previous studies have shown that fine-tuning a predictor pretrained on a computational database to a real system can result in models with outstanding generalization capabilities compared to learning from scratch. This study demonstrates the scaling law of simulation-to-real (Sim2Real) transfer learning for several machine learning tasks in materials science. Case studies of three prediction tasks for polymers and inorganic materials reveal that the prediction error on real systems decreases according to a power-law as the size of the computational data increases. Observing the scaling behavior offers various insights for database development, such as determining the sample size necessary to achieve a desired performance, identifying equivalent sample sizes for physical and computational experiments, and guiding the design of data production protocols for downstream real-world tasks.
△ Less
Submitted 7 August, 2024;
originally announced August 2024.
-
Semi-automatic staging area for high-quality structured data extraction from scientific literature
Authors:
Luca Foppiano,
Tomoya Mato,
Kensei Terashima,
Pedro Ortiz Suarez,
Taku Tou,
Chikako Sakai,
Wei-Sheng Wang,
Toshiyuki Amagasa,
Yoshihiko Takano,
Masashi Ishii
Abstract:
We propose a semi-automatic staging area for efficiently building an accurate database of experimental physical properties of superconductors from literature, called SuperCon2, to enrich the existing manually-built superconductor database SuperCon. Here we report our curation interface (SuperCon2 Interface) and a workflow managing the state transitions of each examined record, to validate the data…
▽ More
We propose a semi-automatic staging area for efficiently building an accurate database of experimental physical properties of superconductors from literature, called SuperCon2, to enrich the existing manually-built superconductor database SuperCon. Here we report our curation interface (SuperCon2 Interface) and a workflow managing the state transitions of each examined record, to validate the dataset of superconductors from PDF documents collected using Grobid-superconductors in a previous work. This curation workflow allows both automatic and manual operations, the former contains ``anomaly detection'' that scans new data identifying outliers, and a ``training data collector'' mechanism that collects training data examples based on manual corrections. Such training data collection policy is effective in improving the machine-learning models with a reduced number of examples. For manual operations, the interface (SuperCon2 interface) is developed to increase efficiency during manual correction by providing a smart interface and an enhanced PDF document viewer. We show that our interface significantly improves the curation quality by boosting precision and recall as compared with the traditional ``manual correction''. Our semi-automatic approach would provide a solution for achieving a reliable database with text-data mining of scientific documents.
△ Less
Submitted 16 November, 2023; v1 submitted 19 September, 2023;
originally announced September 2023.
-
Automatic extraction of materials and properties from superconductors scientific literature
Authors:
Luca Foppiano,
Pedro Baptista de Castro,
Pedro Ortiz Suarez,
Kensei Terashima,
Yoshihiko Takano,
Masashi Ishii
Abstract:
The automatic extraction of materials and related properties from the scientific literature is gaining attention in data-driven materials science (Materials Informatics). In this paper, we discuss Grobid-superconductors, our solution for automatically extracting superconductor material names and respective properties from text. Built as a Grobid module, it combines machine learning and heuristic a…
▽ More
The automatic extraction of materials and related properties from the scientific literature is gaining attention in data-driven materials science (Materials Informatics). In this paper, we discuss Grobid-superconductors, our solution for automatically extracting superconductor material names and respective properties from text. Built as a Grobid module, it combines machine learning and heuristic approaches in a multi-step architecture that supports input data as raw text or PDF documents. Using Grobid-superconductors, we built SuperCon2, a database of 40324 materials and properties records from 37700 papers. The material (or sample) information is represented by name, chemical formula, and material class, and is characterized by shape, doping, substitution variables for components, and substrate as adjoined information. The properties include the Tc superconducting critical temperature and, when available, applied pressure with the Tc measurement method.
△ Less
Submitted 22 November, 2022; v1 submitted 25 October, 2022;
originally announced October 2022.
-
Third-order Electrical Conductivity of the Charge-ordered Organic Salt $α$-(BEDT-TTF)$_2$I$_3$
Authors:
Mayu Ishii,
Ryuji Okazaki,
Masafumi Tamura
Abstract:
We performed third-order electrical conductivity measurements on the organic conductor $α$-(BEDT-TTF)$_2$I$_3$ using an ac bridge technique sensitive to nonlinear signals. Third-order conductance $G_3$ is clearly observed even at low electric fields, and interestingly, $G_3$ is critically enhanced above the charge-order transition temperature $T_{\rm CO}=136$~K. The observed frequency dependence o…
▽ More
We performed third-order electrical conductivity measurements on the organic conductor $α$-(BEDT-TTF)$_2$I$_3$ using an ac bridge technique sensitive to nonlinear signals. Third-order conductance $G_3$ is clearly observed even at low electric fields, and interestingly, $G_3$ is critically enhanced above the charge-order transition temperature $T_{\rm CO}=136$~K. The observed frequency dependence of $G_3$ is incompatible with a percolation model, in which a Joule heating in a random resistor network is relevant to the nonlinear conduction. We instead argue the nonlinearity of the relaxation time according to a phenomenological model on the mobility in materials with large dielectric constants, and find that the third-order conductance $G_3$ corresponds to the third-order electric susceptibility $χ_3$. Since the nonlinear susceptibility is known as a probe for higher-order multipole ordering, the present observation of the divergent behavior of $G_3$ above $T_{\rm CO}$ reveals an underlying quadrupole instability at the charge-order transition of the organic system.
△ Less
Submitted 27 January, 2022;
originally announced January 2022.
-
SuperMat: Construction of a linked annotated dataset from superconductors-related publications
Authors:
Luca Foppiano,
Sae Dieb,
Akira Suzuki,
Pedro Baptista de Castro,
Suguru Iwasaki,
Azusa Uzuki,
Miren Garbine Esparza Echevarria,
Yan Meng,
Kensei Terashima,
Laurent Romary,
Yoshihiko Takano,
Masashi Ishii
Abstract:
A growing number of papers are published in the area of superconducting materials science. However, novel text and data mining (TDM) processes are still needed to efficiently access and exploit this accumulated knowledge, paving the way towards data-driven materials design. Herein, we present SuperMat (Superconductor Materials), an annotated corpus of linked data derived from scientific publicatio…
▽ More
A growing number of papers are published in the area of superconducting materials science. However, novel text and data mining (TDM) processes are still needed to efficiently access and exploit this accumulated knowledge, paving the way towards data-driven materials design. Herein, we present SuperMat (Superconductor Materials), an annotated corpus of linked data derived from scientific publications on superconductors, which comprises 142 articles, 16052 entities, and 1398 links that are characterised into six categories: the names, classes, and properties of materials; links to their respective superconducting critical temperature (Tc); and parametric conditions such as applied pressure or measurement methods. The construction of SuperMat resulted from a fruitful collaboration between computer scientists and material scientists, and its high quality is ensured through validation by domain experts. The quality of the annotation guidelines was ensured by satisfactory Inter Annotator Agreement (IAA) between the annotators and the domain experts. SuperMat includes the dataset, annotation guidelines, and annotation support tools that use automatic suggestions to help minimise human errors.
△ Less
Submitted 15 April, 2021; v1 submitted 7 January, 2021;
originally announced January 2021.
-
Ultra-low power on-chip learning of speech commands with phase-change memories
Authors:
Venkata Pavan Kumar Miriyala,
Masatoshi Ishii
Abstract:
Embedding artificial intelligence at the edge (edge-AI) is an elegant solution to tackle the power and latency issues in the rapidly expanding Internet of Things. As edge devices typically spend most of their time in sleep mode and only wake-up infrequently to collect and process sensor data, non-volatile in-memory computing (NVIMC) is a promising approach to design the next generation of edge-AI…
▽ More
Embedding artificial intelligence at the edge (edge-AI) is an elegant solution to tackle the power and latency issues in the rapidly expanding Internet of Things. As edge devices typically spend most of their time in sleep mode and only wake-up infrequently to collect and process sensor data, non-volatile in-memory computing (NVIMC) is a promising approach to design the next generation of edge-AI devices. Recently, we proposed an NVIMC-based neuromorphic accelerator using the phase change memories (PCMs), which we call as Raven. In this work, we demonstrate the ultra-low-power on-chip training and inference of speech commands using Raven. We showed that Raven can be trained on-chip with power consumption as low as 30~uW, which is suitable for edge applications. Furthermore, we showed that at iso-accuracies, Raven needs 70.36x and 269.23x less number of computations to be performed than a deep neural network (DNN) during inference and training, respectively. Owing to such low power and computational requirements, Raven provides a promising pathway towards ultra-low-power training and inference at the edge.
△ Less
Submitted 21 October, 2020;
originally announced October 2020.
-
Enhanced Seebeck coefficient by a filling-induced Lifshitz transition in KxRhO2
Authors:
Naoko Ito,
Mayu Ishii,
Ryuji Okazaki
Abstract:
We have systematically measured the transport properties in the layered rhodium oxide K$_{x}$RhO$_{2}$ single crystals ($0.5\lesssim x \lesssim 0.67$), which is isostructural to the thermoelectric oxide Na$_{x}$CoO$_{2}$. We find that below $x = 0.64$ the Seebeck coefficient is anomalously enhanced at low temperatures with increasing $x$, while it is proportional to the temperature like a conventi…
▽ More
We have systematically measured the transport properties in the layered rhodium oxide K$_{x}$RhO$_{2}$ single crystals ($0.5\lesssim x \lesssim 0.67$), which is isostructural to the thermoelectric oxide Na$_{x}$CoO$_{2}$. We find that below $x = 0.64$ the Seebeck coefficient is anomalously enhanced at low temperatures with increasing $x$, while it is proportional to the temperature like a conventional metal above $x=0.65$, suggesting an existence of a critical content $x^{*} \simeq 0.65$. For the origin of this anomalous behavior, we discuss a filling-induced Lifshitz transition, which is characterized by a sudden topological change in the cylindrical hole Fermi surfaces at the critical content $x^*$.
△ Less
Submitted 7 January, 2019;
originally announced January 2019.
-
XANES study of rare-earth valency in LRu4P12 (L = Ce and Pr)
Authors:
C. H. Lee,
H. Oyanagi,
C. Sekine,
I. Shirotani,
M. Ishii
Abstract:
Valency of Ce and Pr in LRu4P12 (L = Ce and Pr) was studied by L2,3-edge x-ray absorption near-edge structure (XANES) spectroscopy. The Ce-L3 XANES spectrum suggests that Ce is mainly trivalent, but the 4f state strongly hybridizes with ligand orbitals. The band gap of CeRu4P12 seems to be formed by strong hybridization of 4f electrons. Pr-L2 XANES spectra indicate that Pr exists in trivalent st…
▽ More
Valency of Ce and Pr in LRu4P12 (L = Ce and Pr) was studied by L2,3-edge x-ray absorption near-edge structure (XANES) spectroscopy. The Ce-L3 XANES spectrum suggests that Ce is mainly trivalent, but the 4f state strongly hybridizes with ligand orbitals. The band gap of CeRu4P12 seems to be formed by strong hybridization of 4f electrons. Pr-L2 XANES spectra indicate that Pr exists in trivalent state over a wide range in temperature, 20 < T < 300 K. We find that the metal-insulator (MI) transition at TMI = 60 K in PrRu4P12 does not originate from Pr valence fluctuation.
△ Less
Submitted 20 January, 2000;
originally announced January 2000.