Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5858_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
30.08.2026
Размер:
43 Мб
Скачать
72 R. Yoshida
https://t.me/med1917
trained using data from the target domain, where Y
f
target domain and the NN acquires the synthetic feature
is an arbitrary model. From the learning of the source domain,
target
φ
(X )
source
is the output variable of the
target
suitable for predicting Y
source
.If a common mechanism exists between the source and target domains, this feature can be utilized to predict sufficiently reduced and if
. In particular, if the dimension of φ
target
f
can be represented by a simple model, such as a
target
source
(X )
can be
Y
linear model, a high-performance model may be constructed with considerably less data than direct learning of the target domain. This type of transfer learning is known as feature extraction. There are various other methods of transfer learning, such as fine-tuning and affine transfer; however, owing to space limitations, we will not go into further detail.
4.3.2 Application of Transfer Learning: Prediction of Lattice
Thermal Conductivity of Inorganic Compounds
Ju et al. (2021) [13] constructed a statistical model to predict the physical proper­ties of crystal structures using transfer learning and successfully identified inorganic compounds with high thermal conductivity. In this study, only 45 samples were used to predict the lattice thermal conductivity (LTC), and the required level of predic­tion accuracy could not be achieved using conventional direct supervised learning. Compared to LTC, the cost of the first-principles calculations for SPS is considerably lower. We trained an NN using the 320 instances of the SPS that we produced and transferred to the target domain of the LTC model using 45 training instances. Using the transferred model, we predicted the LTC of approximately 60,000 compounds and identified 14 compounds that were expected to have high thermal conductivity. The LTC values of the 14 identified compounds were verified by first-principles calculations, and their LTCs reached the highest level, exceeding 3000 W/(mK). On the other hand, the LTCs of the 45 compounds used for training were distributed in the range of less than 400 W/(mK). In machine learning, the range of applicability of predictions is generally limited to the neighboring region of the training dataset distribution. In fact, the NN constructed using only 45 training instances could not predict the LTC of the 14 compounds. In contrast, the model transferred via the prediction of SPS was able to predict the LTCs for 14 compounds with some accu­racy. Cases where extrapolation is provided for prediction using transfer learning, as
11
in this case, are often observed [ domain contained some information that contributed to the acquisition of general­purpose features. By reusing the feature extractors, we successfully constructed a statistical model with predictive power even in regions that deviated greatly from the range of the training data.
]. The 320 observed instances of SPS in the source
4 Materials Informatics with Limited Data 73
https://t.me/med1917
4.3.3 Application of Transfer Learning: Prediction
and Discovery of High-Thermal-Conductivity Amorphous Polymers
Wu et al. (2019) [6] constructed a prediction model for the thermal conductivity of polymeric materials using transfer learning, as described in Sect. recording the thermal conductivities of 28 amorphous polymers from the PoLyInfo polymer properties database was used. The glass transition temperature, melting point, specific heat capacity at constant volume, and viscosity of the polymers were used as the source domains. A trained model was constructed in each domain and transferred to a thermal conductivity prediction model using the 28 instances of thermal conductivity. The model with the smallest mean absolute error was selected by cross-validation. The inverse problem was solved by designing 1000 hypothetical polymers predicted to have high thermal conductivities, as described in Sect. Finally, three aromatic polyamides were selected as candidates, synthesized, and had their properties measured. One of the synthesized polymers was found to have a thermal conductivity of 0.41 W/(mK). This corresponds to a performance improve­ment of approximately 80% compared with that of typical unoriented polyamide polymers. Notably, synthesized polymers with similar structures were rarely included in the training data. As in the previous case study, this case study illustrates the extrapolative nature of transfer learning.
4.2.1. A dataset
4.2.1.
4.3.4 Multitask Learning to Predict Polymer–Solvent
Miscibility
Predicting and understanding the miscibility of polymers with solvents is of great importance in polymer chemistry because the dissolution of polymers and solvents occurs in a variety of processes during material development, including plastic recy­cling, polymer blend synthesis, refining, painting, and coating. However, it is diffi­cult to accurately predict the phase behaviors of various polymer–solvent systems using current computational chemistry techniques. According to the Flory–Huggins theory [ given temperature, volume fraction, and molecular length, the change in the mixing free energy of a polymer solution is determined by a quantity called the χ param­eter, which represents the molecular interaction between the polymer and solvent. However, it is technically difficult to measure the χ parameter quickly and precisely with current experimental techniques, and the training dataset is known to be quanti­tatively insufficient and severely biased due to the nature of the experimental system. Aoki et al. (2023) [ applicable to a wide range of polymer–solvent pairs by integrating a large amount of quantum chemical calculation data and a limited amount of experimental data using a machine learning method called multi-task learning [
2426] that describes the thermodynamic properties of polymer solutions,
27] successfully constructed a highly accurate prediction model
28
].
74 R. Yoshida
https://t.me/med1917
They aimed to use machine learning to predict the χ parameter, which describes
the polymer–solvent interaction. According to the Flory–Huggins theory, the change
ΔG
in Gibbs free energy
upon mixing polymer p with solvent s is expressed by
mix
the following equation:
mix
=
Φ
p
log Φp +
N
p
Φ
s
log Φs + ΦpΦsχ (4.8)
N
s
ΔG
kT
where k, T , and Φi denote the Boltzmann’s constant, absolute temperature, and volume fraction of component
i {p, s}, respectively. Further, Np and Ns denote the
lengths of the molecular chains of p and s, respectively. In this study, the reference volume was set to the molecular volume of the solvent,
Ns = 1. The first two terms
represent the changes in the combinatorial entropy of p and s, respectively, which are related to the number of possible conformational states in the mixture. The third term includes the Flory–Huggins interaction parameter χ. This parameter is a critical dimensionless metric that represents the difference in the noncombinatorial entropy and strength of the pairwise interaction energies between p and s in the mixture. The two combinatorial entropy terms can be calculated based on the volume fractions p and s and the length of the molecular chain. Therefore, to analyze the solution phase behaviors of p and s, only the temperature-dependent value of the χ parameter has to be observed or estimated: the smaller the value of the χ parameter, the smaller the
ΔG
, and the more likely the mixing of p and s.
mix
The accurate prediction of the phase behaviors of various polymer–solvent systems is difficult using current computational chemistry techniques. Currently, empirical prediction methods based on the distance between the solubility param­eters of the polymers and solvents are widely used. For example, the Hansen solu­bility parameter (HSP) represents the potential solubility of a molecule as a three­dimensional vector consisting of dispersive force, polar, and hydrogen bonding
2931]. The compatibility of the polymer with the solvent is estimated based
terms [ on the distance between the HSP vectors. Although the solubility parameters of various molecules have been measured experimentally for molecules with undeter­mined solubility parameters, empirical models such as the atomic group contribu-
31
tion method [
] have been applied to estimate solubility parameters. However, such
empirical models have very low prediction accuracy, except for certain molecular
], a molecular simulation based on quantum
species. The COSMO-RS method [
32, 33
chemical calculations, can also estimate χ parameters. However, quantum chemical calculations are computationally expensive, making it difficult to apply them to the large-scale screening of candidate solvent molecules. Additionally, the prediction accuracy did not reach a sufficient level.
Aoki et al. (2023) [27
] used the experimental values of the χ parameter for 1190 polymer–solvent pairs, consisting of 46 different polymers and 140 different solvent molecules, to train the model. The dataset also included measurements of the χ parameter for different temperatures and polymer–solvent compositions. The molecular species of the polymers/solvents in the dataset were distributed over a very
4 Materials Informatics with Limited Data 75
https://t.me/med1917
Fig. 4.6 Bias of polymer and solvent species in the experimental dataset. This figure is a reprint from Aoki et al. (2023) [ data set D PoLyInfo. b Polymer–solvent pairs in the experimental dataset of the χ parameters were mapped to the soluble/insoluble labels in PoLyInfo, and the groupwise histograms are displayed for the two groups as a function of χ parameter values. The total number of polymer–solvent pairs for each group is included in the legend, which shows a significant bias in obtaining χ parameters for soluble polymer–solvent pairs
χ
27]. a UMAP projections of polymers’ chemical structures in the experimental
and computational dataset Dc of the χ parameters, and the soluble/insoluble dataset of
limited region of the entire chemical space (Fig. 4.6(a)). In addition, in certain exper­imental systems, it is difficult to measure the χ parameters of the polymer–solvent system in an immiscible state, resulting in a significant bias in the distribution of the data (Fig.
4.6(b)). Therefore, models trained using only this dataset generally
have narrow predictive applicability and are unable to predict the χ parameter in immiscible states.
To address this issue, we used quantum chemical calculations (COSMO-RS method) to generate a dataset of χ parameters for 9129 polymers and solvents. We also generated a dataset of 429 pairs of polymers and solvents with binary class labels indicating whether they were good or poor solvents. Using these three datasets, we trained a deep NN to predict from the chemical structures of an input polymer and solvent (1) the experimental χ parameter, the computational χ parameter of the quantum chemical calculation, and the binary class labels representing the solubility of the polymer/solvent pair (Fig.
4.7). This method is known as multi-task learning.
In multi-task learning, different tasks with a common mechanism are learned simul­taneously using a unified model. The experimental data on χ parameters for the main task were limited in quantity and included systematic biases caused by the nature of the experimental system. Therefore, by using two auxiliary tasks with data encom­passing a wide variety of molecular species, we could expand the applicability of the prediction model.
The model was experimentally confirmed to have significantly high predictive
performance for all three tasks (Fig.
4.8(a)). Additionally, it exhibited much better
predictive power than quantum chemical calculations based on COSMO-RS and the empirical method based on HSP (Fig.
4.8(b)). The architecture of the model was
76 R. Yoshida
https://t.me/med1917
Fig. 4.7 Architecture of a multitasking neural network (NN) for predicting experimental and computational χ parameters and the binary classification task to discriminate between good and poor polymer–solvent pairs. This figure is a reprint from Aoki et al. (2023) [
27]
designed to extend the concept of HSP, which assumes that the potential solubility of a molecule is determined by its dispersion force, polarity, and hydrogen bond strength. However, the low accuracy of the HSP-based prediction suggests that these three factors alone cannot explain the real system. In contrast, the machine learning algorithm, through learning the data, presented 34 different factors to describe the solubility of the molecule. Several of these factors correspond to three HSP factors. This suggests that there are unknown factors behind the determination of polymer– solvent compatibility that have been ignored in HSP.
4.3.5 Functional Output Regression for Limited Data
The output variable in the ordinary case of MI is often given as a scalar variable, such as the physical properties of materials. However, the variables to be predicted in materials research are often presented in the form of functions. For example, when predicting the optical absorption spectrum of a molecule, the input variable is the chemical structure, and the output variable is a spectral function defined over a wavelength domain. Several physical properties are determined by the tempera­ture, pressure, and frequency of the external electric field. The dielectric property of a material, that is, the dielectric constant or dielectric loss tangent, is represented as a function of frequency and temperature. For example, in the analysis of the microstructure of a composite material, the composition and processing conditions are treated as input variables, and the output variable is an intensity matrix repre­senting a grayscale image of the microstructure measured using a scanning electron microscope (SEM). In other words, it is a regression analysis in which the output variable is given as a matrix or function in two-dimensional coordinates.
4 Materials Informatics with Limited Data 77
https://t.me/med1917
Fig. 4.8 Results of the prediction of polymer–solvent miscibility. This figure is a reprint from Aoki et al. (2023) [ tasks. b Prediction performance of the quantum chemical calculations using the COSMO-RS method and the empirical predictor based on HSP
27]. a Prediction performance of the multi-task machine learning for the three different
Here, we present the kernel regression for the functional output (KRFO) method
], along with its application to microstructure
d
i=1
34
k(t, s
4.9):
β
+ μ(t) + ε (4.9)
)
(X )
i
i
d
}
{
k(t, s
.The
)
i
i=1
proposed by Iwayama et al. (2022) [ prediction. The KRFO was modeled as follows (see Fig.
Y (X , t) =
The first term is the weighted sum of d kernel basis functions
kernel centers s
are equally spaced in the two-dimensional image coordinate space.
i
For example, a Gaussian radial basis function was used as the kernel basis function.
β
The regression coefficient the weight of the kernel placed at each location. The intercept term μ only on the coordinate vector The model represents a system in which each kernel function located on or deactivated depending on the value of the input variable
is a function of the input variable X and determines
(X )
i
depends
(t)
t. The last term represents the measurement noise ε.
t is activated
X .
Here, we introduce an example application of microstructure prediction. The material is a thin film of Al coated on a Cr-based metal plate; the Cr and Al metal plates are placed facing each other, and Ar gas is sprayed onto the Al metal plate at high velocity via magnetron sputtering to adsorb ejected Al atoms onto the Cr metal plate. The model input
X is a six-dimensional descriptor vector representing
78 R. Yoshida
https://t.me/med1917
Fig. 4.9 Model architecture of the kernel regression for functional outputs. This figure is a reprint from Iwayama et al. (2022) [
34]
the composition CraAl
N and the following process conditions: (1) Cr and Al
1-aOb
content denoted by a; (2) O content denoted by b; (3) temperature at which Al is adsorbed; (4) pressure at which Al is adsorbed; (5) average Ar ion energy when incident on the Al metal plate; and (6) ionization degree of Ar gas. The output variable was an SEM image of the microstructure. In total, 123 images were used. From the overall dataset, 90%, 5%, and 5% of the randomly selected images were used as the training dataset, validation set for hyperparameter adjustment, and test set for generalization performance evaluation, respectively.
Figure 4.10 shows the KRFO prediction results for the seven selected test cases and SEM images of the experiment. Despite the small amount of training data (109), the predicted images adequately captured the morphological features of the microstruc­ture, such as the particle size and shape. This observation raises the simple question of why a highly accurate prediction model was obtained with only 109 training instances, even though the output variable was ultra-high dimensional (a function of the two-dimensional image coordinate space). To answer this question, we inves­tigated the prediction performance of NNs that learned to predict scalar variables without functionalization, that is, image luminance, independently for each pixel in the image. As shown in Fig.
4.10, we confirmed that NNs trained independently
for each pixel could not predict the microstructures. The prediction of the func­tional variables has a learning mechanism similar to that of multi-task learning. The tasks that predict each value of the function are not independent but related to each other. By learning multiple related tasks simultaneously, multi-task learning expects the model to capture latent features and task-specific components that are common
4 Materials Informatics with Limited Data 79
https://t.me/med1917
Fig. 4.10 Results of KRBO, conditional GAN, and pixel-by-pixel independent prediction using NN for the seven test images (top row). This figure is a reprint from Iwayama et al. (2022) [
across tasks. Additionally, the simultaneous use of data from multiple tasks compen­sates f or the lack of data for individual tasks and prevents the model from overfitting to task-specific observational noise. Functional output regression is a special case of multi-task learning in which data from different tasks (different function values) are observed simultaneously. This mechanism may allow functional output regression to achieve high predictive performance despite having only a small amount of data.
These observations provide an opportunity to reconsider machine-learning approaches in MI. For example, the prediction of temperature-dependent proper­ties often involves modeling scalar output variables with the measured temperature limited to room temperature. Such an approach would make the problem more diffi­cult and may lead to lower prediction accuracy. In fact, the model obtained from independent pixel-by-pixel training did not predict the microstructures at all, but simultaneous training of the entire image with KRFO successfully predicted the microstructures. If functional data are available, the prediction accuracy may be significantly improved by jointly predicting the entire function.
34]
4.4 Toward Creating a Large Physical Property Database
of Polymer Materials
In MI, a significant amount of data from computer simulations is often integrated and analyzed to compensate for the lack of data for a target task. Currently, large­scale computational property databases for various material systems are being devel­oped. The development of first-principles computational databases (e.g., Materials
] and QM9 [35]), which include tens of thousands to millions of materials
Project [ with their computational properties, has dramatically advanced the development and spread of MI, particularly for inorganic materials and low-molecular-weight
22
80 R. Yoshida
https://t.me/med1917
compounds. However, the development of databases for polymeric materials has made little progress owing to technical difficulties in automating the calculation of material properties and enormous computational costs. The existing databases are summarized in Table
4.1. The amount of data in the existing polymer property
databases is rather small, and tools for the automatic extraction of digital data are not yet available. In addition, information such as sample preparation conditions and higher-order structures has rarely been recorded. It is no exaggeration to state that there are practically no systematic open data sources that contribute to data-driven polymer materials research.
RadonPy is open-source software that fully automates polymer property calcu-
9, 36
lations based on all-atom classical MD simulations [
]. Once the chemical structure, degree of polymerization, and temperature of the polymer repeating unit are entered, the entire MD simulation process, including polymer chain forma­tion, charge calculation, force field parameter assignment, equilibrium and non­equilibrium MD calculations, determination of equilibrium completion, and various property calculations, is completely automated (Fig.
4.11). Currently, 17 proper-
ties, including thermal, mechanical, and optical, can be automatically calculated. Computable polymer systems include amorphous linear polymers, oriented polymer systems, and polymer solution systems.
Table 4.1 Representative polymer physical property open-source data. Compared to other applied fields of data science, there is very little data, and most of them do not have tools for automatic data extraction
Database Overview
PoLyInfo (polymer.nims.go.jp)
Polymer Genome—Khazana (khazana.gatech.edu)
Polymer Property Predictor and Database (pppdb.uchicago.edu)
NanoMine (materialsmine.org)
CROW (polymerdatabase.com)
Polymers: A Property Database (poly.chemnetbase.com)
The world’s largest physical property database of polymers that summarizes data extracted from the academic literature. Contains data on approximately 100 physical properties of polymers that are polymerized from about 18,000 types of monomers
A platform that provides experimental data extracted from 24 publications and physical property values calculated by ab initio calculations. Contains property data of approximately 1400 types of polymer/organic materials and approximately 2600 types of inorganic materials
It contains about 260 Flory–Huggins χ parameters and about 210 glass transition temperature data extracted from the academic literature
Includes composition, process, electron microscopy data, and physical properties of the microstructure of polymer composites
A repository that contains polymer thermophysical data. Includes experimental data extracted from the academic literature and computed physical property data computed from quantitative structure–activity relationship analysis
Polymer property data provided as an appendix to the Wiley book “Polymers: A Property Database.”
4 Materials Informatics with Limited Data 81
https://t.me/med1917
Fig. 4.11 Overview of RadonPy, software for automatically calculating polymer properties
The input variables for RadonPy are SMILES strings, which represent the repeating units of the polymer, degree of polymerization, and other properties. Hayashi et al. ( 2022) [
9] adjusted the degree of polymerization such that the number
of atoms in the polymer chain was approximately 1000. The repeating units were linked using the random-walk method to produce a randomly coiled polymer chain. The initial structure of an amorphous polymer was generated by replicating 10 of these polymer chains and randomly arranging and rotating them so that they did not overlap with each other, and the force field assignments and other settings necessary for MD calculations were fully automated. Equilibration MD for the structural relax­ation of the system was performed using LAMMPS. The end of the equilibration calculation was determined by the convergence status of the energy, density, and mean-square displacement of the atomic coordinates. From this equilibration MD, physical properties such as density and specific heat were calculated. In addition, heat-conduction MD was performed on the amorphous structure after equilibration. The thermal conductivity and thermal diffusivity of the amorphous structure were calculated from the results of thermal conduction MD, and the physical properties such as Young’s modulus and Poisson’s r atio were calculated from the results of uniaxial stretching MD. These calculation conditions were determined based on a systematic comparative verification of the physical properties of 1070 polymers in the polymer properties database PoLyInfo, as shown in Fig.
4.12(a).
In MD calculations of polymers, the calculation conditions, such as the degree of polymerization, have a significant impact on the calculation results. Therefore, to determine the conditions for high-throughput calculations, we systematically inves­tigated the effects of simulation conditions and molecular species on the prediction accuracy based on PoLyInfo experimental data. Furthermore, the biases and varia­tions in the calculated properties were corrected using transfer learning (Fig.
4.12(b)).