Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5338_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
30.08.2026
Размер:
43 Мб
Скачать
82 R. Yoshida
https://t.me/med1917
Fig. 4.12 Calibration of MD-calculated physical properties (specific heat capacity at constant pres­sure (C
), linear expansion coefficient, volume expansion coefficient). This figure is a reprint from
P
Hayashi et al. (2022) [ (top) were significantly improved by calibration using transfer learning (bottom)
9]. Bias and variations between the MD-calculated and experimental values
In the examples shown in Fig. specific heat capacity, linear expansion coefficient, and volume expansion coeffi­cient. As shown in Fig. between the experimental and MD-calculated values, which stemmed from the pres­ence or absence of quantum effects. The latter two had significantly large variations even within the same polymer in both experimental and calculated properties. For each property, the source task of transfer learning was defined to predict the MD­calculated properties, and the target task was to predict the experimental properties in PoLyInfo. A prediction model defines a mapping from the fingerprinted chemical structure of a given polymer repeating unit to the experimental or MD-calculated properties. As shown in Fig. showed significant improvements in predicting the experimental data compared with the direct predictions from the MD calculations. The systematic bias in the specific heat capacity has almost disappeared. Interestingly, for the linear and volume expan­sion coefficients, the transferred model not only corrected the systematic bias of the MD-calculated properties but also significantly reduced the variability of the experimental values.
No calculation conditions can be applied to a wide variety of polymers. There­fore, biases and variations always occur in the MD properties obtained from fully
4.12(b), the target properties to be predicted were the
4.12(a), the specific heat capacity exhibited a significant bias
4.12(b), for all three properties, the transferred models
4 Materials Informatics with Limited Data 83
https://t.me/med1917
automated calculations. Biases and variations also occur in the experimental values owing to the experimental and sample preparation conditions and the nature of the measurement equipment. Transfer learning bridges the gap between complex real-world systems and imperfect computer models.
Using RadonPy, we aim to create one of the world’s largest polymer prop­erty databases containing more than 100,000 molecular skeletons. Furthermore, in October 2022, an industry-academia consortium was established for the joint devel­opment of RadonPy and the computational polymer property database. To date, approximately 150 members from one national institute, three universities, and 29 companies have participated in the consortium. This project is supported by the “Program for Promoting Research on the Supercomputer Fugaku” of the Ministry of Education, Culture, Sports, Science and Technology (MEXT) in Japan and is making maximum use of the computational resources of one of the fastest supercomputers, “Fugaku,” to produce and accumulate a vast amount of data on a daily basis. The main computational targets were virtual polymers. We classified the polymer skele­tons into 20 classes, including polyester, polyimide, and polyacrylate, and built a molecular generator for each polymer class using the chemical language model of Ikebata et al. (2017) [ models with the chemical structures of existing polymers and built structure gener­ators that mimic the frequent patterns (e.g., fragmentation and bonding rules) that appear in existing polymers. A comprehensive virtual library consisting of 1,778,039 candidate molecules with diverse molecular skeletons was created.
Currently, the calculations of the physical properties of 47,500 amorphous poly­mers have been completed. The joint distribution of multiple properties of a signif­icant number of polymers has been clearly and comprehensively observed by conducting computer experiments at this scale, as shown in Fig. systematic knowledge of the location of the Pareto frontier formed by the trade­offs of multiple properties and the structural features of the polymer groups consti­tuting it was obtained. With the current experimental techniques for measurement and synthesis, the physical properties and material space could not be comprehen­sively observed on such a scale. Furthermore, high-throughput thermal conductivity calculations have identified novel polymers beyond the Pareto frontier. The thermal conductivity of ordinary amorphous polymers i s approximately 0.2–0.3 W/(mK) at best; however, some of the calculated polymers had thermal conductivities exceeding
0.5 W/(mK) (Fig. revealed that the rigidity of the molecular backbone, presence of a high density of hydrogen-bondable units, and mechanisms involving hydrogen bonding and dipole– dipole interactions are responsible for the high thermal conductivity of amorphous
9
polymers [
This project aims to create a large map of polymer material properties. In partic­ular, the project aims to create a systematic dataset of biodegradable plastics, highly thermally conductive polymers, and thermosetting resins and to create new materials that contribute to a decarbonized, recycling-oriented society and thermal management.
].
5] as described in Sect. 4.2.1. Here, we trained machine-learning
4.13 (a). As a result,
4.13(b)). Analysis of the structural features of these polymers
84 R. Yoshida
https://t.me/med1917
Fig. 4.13 Polymer world map constructed by the high-throughput MD simulation using RadonPy. This figure is a reprint from Hayashi et al. (2022) [ multiple physical properties (thermal conductivity, density, specific heat capacity at constant pres­sure (CP), volume expansion coefficient, linear expansion coefficient, refractive index) of polymer materials. b Eight types of polymers that exhibited a high thermal conductivity of more than 0.4 W/(mK) in the amorphous state
9]. a Joint distribution and Pareto frontier of
4.5 Concluding Remarks
This section provides an overview of MI in terms of forward and inverse problems. Input/output variables in materials research can take various forms. Owing to this diversity, methodologies and tools must be developed for each problem. The conflu­ence of academic advances in data, computational, and experimental sciences with this conventional workflow has produced new scientific methods and discoveries. In recent years, the time lag between the confluence of cutting-edge technologies in data science and applied fields has rapidly decreased.
The most important aspect of data-driven research is the data. Compared with other applied fields of data science, the amount of data in materials research is far less. Because a data-driven approach has not yet been fully introduced, the devel­opment of databases is still in its infancy. Additionally, because scientific outcomes and industrial applications are closely linked, researchers are highly conscious of information confidentiality, and in some areas, data sharing may not progress in the future. In these areas, the barrier of limited data remains a problem. However, these data are an endless source of knowledge. The volume and diversity of the data never diminish, but rather increase monotonically. Simultaneously, a gap exists between those who have data and those who do not. Power games are the essence of data­driven research. It is also important to examine the state of MI from a bird’s eye view in an area where big data and small amounts of data coexist.
4 Materials Informatics with Limited Data 85
https://t.me/med1917
References
1. Bohacek RS, McMartin C, Guida WC (1996) The Art and Practice of Structure-Based Drug Design: A Molecular Modeling Perspective. Med Res Rev 16(1):3–50. 10.1002/(SICI)1098­1128(199601)16:1<3::AID-MED1>3.0.CO;2-6
2. Kim S, Thiessen PA, Bolton EE, Chen J, Fu G, Gindulyte A, Han L, He J, He S, Shoemaker BA, Wang J, Yu B, Zhang J, Bryant SH (2016) PubChem Substance and Compound Databases. Nucleic Acids Res 44(D1):D1202–1213.
3. Miyao T, Arakawa M, Funatsu K (2010) Exhaustive Structure Generation for Inverse-QSPR/ QSAR. Mol Inform 29:111–125.
4. Miyao T, Kaneko H, Funatsu K (2016) Inverse QSPR/QSAR Analysis for Chemical Structure Generation (from y to x). J Chem Inf Model 56(2):286–299.
5b00628
5. Ikebata H, Hongo K, Isomura T, Maezono R, Yoshida R (2017) Bayesian Molecular Design with a Chemical Language Model. J Comput-Aided Mol Des 31:379–391.
1007/s10822-016-0008-z
6. Wu S, Kondo Y, Kakimoto M, Yang B, Yamada H, Kuwajima I, Lambard G, Hongo K, Xu Y, Shiomi J, Schick C, Morikawa J, Yoshida R (2019) Machine-Learning-Assisted Discovery of Polymers with High Thermal Conductivity Using a Molecular Design Algorithm. NPJ Comput Mater 5(1):66.
7. Liu C, Fujita E, Katsura Y, Inada Y, Ishikawa A, Tamura R, Kimura K, Yoshida R (2021) Machine Learning to Predict Quasicrystals from Chemical Compositions. Adv Mater 33(36):2102507.
8. LiuC,KitaharaK,IshikawaA,HirotoT,Singh A, FujitaE,Katsura Y, InadaY,TamuraR, Kimura K, Yoshida R (2023) Quasicrystals Predicted and Discovered by Machine Learning. Phys Rev Mater 7(9):093805.
9. Hayashi Y, Shiomi J, Morikawa J, Yoshida R (2022) RadonPy: Automated Physical Prop­erty Calculation Using All-Atom Classical Molecular Dynamics Simulations for Polymer Informatics. NPJ Comput Mater 8(1):222.
10. Otsuka S, Kuwajima I, Hosoya J, Xu Y, Yamazaki M (2011) PoLyInfo: Polymer Database for Polymeric Materials Design. In 2011 International Conference on Emerging Intelligent Data and Web Technologies, 22–29.
11. Yamada H, Liu C, Wu S, Koyama Y, Ju S, Shiomi J, Morikawa J, Yoshida R (2019) Predicting Materials Properties with Little Data Using Shotgun Transfer Learning. ACS CentL Sci 5(10):1717–1730.
12. Minami S, Liu S, Wu S, Fukumizu K, Yoshida R (2021) A General Class of Transfer Learning Regression Without Implementation Cost. In Proceedings of the AAAI Conference on Artificial Intelligence 35(10):8992–8999.
13. Ju S, Yoshida R, Liu C, Wu S, Hongo K, Tadano T, Shiomi J (2021) Exploring Diamondlike Lattice Thermal Conductivity Crystals Via Feature-Based Transfer Learning. Phys Rev Mater 5(5):053801.
14. Torres P, Wu S, Ju S, Liu C, Tadano T, Yoshida R, Shiomi J (2022) Descriptors of Intrinsic Hydrodynamic Thermal Transport: Screening a Phonon Database in a Machine Learning Approach. J Phys: Condens Matter 34(13):135702.
15. Minami S, Fukumizu K, Hayashi Y, Yoshida R (2023) Transfer learning with Affine Model Transformation. Adv Neural Inf Process Syst 36, in press.
16. Weininger D (1988) SMILES, A Chemical Language and Information System. 1. Introduction to Methodology and Encoding Rules. J Chem Inf Comput Sci 28:31–36.
1021/ci00057a005
17. Lowe DM (2012) Extraction of Chemical Structures and Reactions from the Literature. Ph.D. Thesis. University of Cambridge,
18. Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, Polosukhin I (2017) Attention Is All You Need. Adv Neural Inf Process Syst 30:5998–6008.
https://doi.org/10.1038/s41524-019-0203-2
https://doi.org/10.1002/adma.202102507
https://doi.org/10.1103/PhysRevMaterials.7.093805
https://doi.org/10.1109/EIDWT.2011.13
https://doi.org/10.1021/acscentsci.9b00804
https://doi.org/10.1609/aaai.v35i10.17087
https://doi.org/10.1103/PhysRevMaterials.5.053801
https://doi.org/10.1093/nar/gkv951
https://doi.org/10.1002/minf.200900038
https://doi.org/10.1021/acs.jcim.
https://doi.org/10.
https://doi.org/10.1038/s41524-022-00906-4
https://doi.org/10.1088/1361-648X/ac49c9
https://doi.org/10.
https://doi.org/10.17863/CAM.16293.
86 R. Yoshida
https://t.me/med1917
19. Schwaller P, Laino T, Gaudin T, Bolgar P, Hunter CA, Bekas C, Lee AA (2019) Molecular Transformer: A Model for Uncertainty-Calibrated Chemical Reaction Prediction. ACS CentL Sci 5:1572–1583.
20. Guo Z, Wu S, Ohno M, Yoshida R (2020) Bayesian Algorithm for Retrosynthesis. J Chem Inf Model 60:4474–4486.
21. Zhang Q, Liu C, Wu S, Hayashi Y, Yoshida R (2023) A Bayesian Method for Concurrently Designing Molecules and Synthetic Reaction Networks. Sci Technol Adv Mater: Methods 3(1):2204994.
22. Jain A, Ong SP, Hautier G, Chen W, Richards WD, Dacek S, Cholia S, Gunter D, Skinner D, Ceder G, and others (2013) Commentary: The Materials Project: A Materials Genome Approach to Accelerating Materials Innovation. APL Mater 1:011002.
1.4812323
23. Wu S, Lambard G, Liu C, Yamada H, Yoshida R (2020) iQSPR in XenonPy: A Bayesian Molecular Design Algorithm. Mol Inform 39(1):1900107.
900107
24. Flory PJ (1942) Thermodynamics of High Polymer Solutions. J Chem Phys 10(1):51–61.
https://doi.org/10.1063/1.1723621
25. Huggins ML (1942) Some Properties of Solutions of Long-Chain Compounds. J Phys Chem 46(1):151–158.
26. Coleman MM, Painter PC (1995) Hydrogen Bonded Polymer Blends. Prog Polym Sci 20(1):1–
59.
https://doi.org/10.1016/0079-6700(94)00038-4
27. Aoki Y, Wu S, Tsurimoto T, Hayashi Y, Minami S, Okubo T, Shiratori K, Yoshida R (2023) Multitask Machine Learning to Predict Polymer–Solvent Miscibility Using Flory–Huggins Interaction Parameters. Macromolecules 56(14):5446–5456.
romol.2c02600
28. Zhang Y, Yang Q (2018) An Overview of Multi-Task Learning. Natl Sci Rev 5(1):30–43.
https://doi.org/10.1093/nsr/nwx105
29. Hansen C (1967) The Three Dimensional Solubility Parameter and Solvent Diffusion Coefficient. Danish Technical Press, Copenhagen.
30. Hansen C (2007) Hansen Solubility Parameters: A User’s Handbook, 2nd ed.; CRC Press: Boca Raton, FL.
31. Stefanis E, Panayiotou C (2008) Prediction of Hansen Solubility Parameters with a New Group-Contribution Method. Int J Thermophys 29:568–585.
008-0415-z
32. Klamt A (2005) COSMO-RS. From Quantum Chemistry to Fluid Phase Thermodynamics and Drug Design. Elsevier: Amsterdam.
33. Loschen C, Klamt A (2014) Prediction of Solubilities and Partition Coefficients in Poly­mers Using COSMO-RS. Ind Eng Chem Res 53(28):11478–11487.
ie501669z
34. Iwayama M, Wu S, Liu C, Yoshida R (2022) Functional Output Regression for Machine Learning in Materials Science. J Chem Inf Model 62(20):4837–4851.
acs.jcim.2c00626
35. Blum LC, Reymond J-L (2009) 970 Million Druglike Small Molecules for Virtual Screening in the Chemical Universe Database GDB-13. J Am Chem Soc 131(25):8732–8733.
org/10.1021/ja902302h
36. RadonPy. https://github.com/RadonPy/RadonPy Accessed 24 Feb 2024
37. Shechtman D, Blech I, Gratias D, Cahn JW (1984) Metallic Phase with Long-Range Orienta­tional Order and No Translational Symmetry. Phys Rev Lett 53(20):1951.
1103/PhysRevLett.53.1951
38. Steurer W. Deloudi S (2009) Crystallography of Quasicrystals, Springer Series in Materials Science, Vol. 126, Springer, Berlin, Heidelberg.
https://doi.org/10.1021/acscentsci.9b00576
https://doi.org/10.1021/acs.jcim.0c00320
https://doi.org/10.1080/27660400.2023.2204994
https://doi.org/10.1063/
https://doi.org/10.1002/minf.201
https://doi.org/10.1021/j150415a018
https://doi.org/10.1021/acs.mac
https://doi.org/10.1007/s10765-
https://doi.org/10.1021/
https://doi.org/10.1021/
https://doi.
https://doi.org/10.
Chapter 5
https://t.me/med1917
Primer on Graph Machine Learning
Masatsugu Yamada and Mahito Sugiyama
5.1 Introduction
Designing de novo molecules for drugs and materials with desired properties is a highly challenging task due to the massive number of candidates, that is, the search space of molecules is exponentially huge with respect to the type of atoms and bonds. To discover drug candidates, one of the most fundamental approaches is using the quantitative structure-activity relationship (QSAR), a mathematical model to explain relationships between biological activities and the structural properties of molecules [ molecules, which is helpful to search datasets. However, to design new molecules, one needs to solve an inverse problem of QSAR models, and it is challenging to solve the optimization problem because a QSAR model loses graph topological structures.
Since molecules can be essentially represented as graphs with node and edge attributes, graph mining methods and graph machine learning methods have been studied to address this task [ in Sect. mental approaches to treat graphs in machine learning and data mining. Then we review graph neural networks (GNNs) in Sect. approaches in graph machine learning. We also introduce some of the recent tech­niques of reinforcement learning that have been used for efficient search of molecules in Sect.
1]. A QSAR model is often used for high-throughput screening of
2]. In this chapter, after reviewing basics of graph theory
5.2, we introduce graph kernels in Sect. 5.3, which is one of the funda-
5.4, which are now the most popular
5.5.
M. Yamada (B) · M. Sugiyama National Institute of Informatics, Chiyoda-ku, Tokyo 101-8430, Japan e-mail:
masatsugu-yamada@nii.ac.jp
M. Sugiyama e-mail:
mahito@nii.ac.jp
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2024 H. Satoh et al. (eds.), Drug Development Supported by Informatics,
https://doi.org/10.1007/978-981-97-4828-0_5
87
88 M. Yamada and M. Sugiyama
https://t.me/med1917
5.2 Graph Theory
First, we define terminology and notation of graphs. A graph is a tuple G = (V , E), where V and E denote the set of nodes (or vertices) E ={e1, e2,..., em}, respectively. Each edge is a pair of nodes (vi, vj) V × V . When one considers an undirected graph, an edge does not have a direction and it is treated as a set. The number of nodes and edges of G, the cardinality of V and E,are represented as via label functions label domain is a triple
lV and lE, respectively. For two graphs G = (V , E) and G'= (V', E'),wesay
from
G'is a subgraph of G, denoted by G'⊑ G,if V'⊆ V and E'⊆ (V V') E.
that
|V | and |E|, respectively. Nodes and edges can have labels (attributes)
lV : V V for nodes and lE : E E for edges with some
∑V and ∑E, which can be any set such as Z and Rd . A labeled graph
G = (V , E, L), where L is the set of node labels and edge labels obtained
Two nodes vi and vj in a graph G = (V , E) are said to be adjacent if an edge
e
= (vi, vj) exists in E. The neighboring information of a graph is represented by
ij
an adjacency matrix. The adjacency matrix
A
=
ij
A R
, vj) E,
1if (v
i
0 otherwise.
The set of neighborhood nodes with respect to a node vi ∈ V is denoted as
V ={v1, v2,..., vn} and edges
|V |×|V |
of a graph is defined as
(5.1)
N (vi) ={vj ∈ V | (vi, vj) E}.
The degree or valency of a node vi inagraph G is the number of neighboring nodes, that is,
deg(vi) =|N (vi)|. (5.2)
A graph can be generalized as a hypergraph H = (X , E), where X is a s et of vertices and E is hyper edges that have multiple edges rather than single pair of nodes. Examples of a graph and a hypergraph are shown in Fig. in the figure shows a simple graph
G = (V , E), where V ={v1, v2, v3, v4} and
5.1. The left graph
E ={e1, e2, e3, e4} with e1 = (v1, v2), e2 = (v1, v3), e3 = (v2, v3), and e4 =
, v4). The right graph in the figure shows a hypergraph H = (X , E) defined as
(v
3
X ={v1, v2, v3, v4} and E ={e1, e2, e3} with e1 ={v1, v2, v3}, e2 ={v1, v3}, and e3 ={v3, v4}. This notation is convenient when considering graph generation by
combining two subgraphs.
A graph can be decomposed into a junction tree composed of the set of subgraphs or cliques C. For a graph
G = (V , E),atree T with subsets C1, ... , Cn of V is called
a junction tree for G if it satisfies the following properties:
5 Primer on Graph Machine Learning 89
https://t.me/med1917
Fig. 5.1 An undirected graph and an undirected hypergraph. Left: An undirected graph that has 4 vertices and 4 edges. Right: An undirected hypergraph that has 4 vertices and 3 edges. An edge e1 is overlapped with half-toned. Edges e1 and e2 are duplication of vertices (v1, v3), but can have different edge labels
1. The union of all sets C1, ... , Cn equals to V; that is,
i
Ci = V .
2. For every edge (u, v) E, there exists Ci ∈ V such that u Ci and v Ci.
3. If Ck is on a path from Ci to Cj in T , Vi ∩ Vj ⊆ Vk.
An example of a junction tree is shown in Fig. 5.2. A graph G is decomposed into a junction tree
, v3, v4}, C3 ={v3, v5}, and C4 ={v5, v6}. The intersection of nodes becomes
{v
2
T . Cliques in G include four cliques: C1 ={v1, v2}, C2 =
edges of the original graph G in reconstruction from the junction tree. This notation is helpful when computing marginalization in general graphs in Bayesian networks and generating graphs being retained with cycles.
Fig. 5.2 An undirected graph and its corresponding junction tree. Left: An undirected graph with a cycle that has 6 vertices and 6 edges. Right: A junction tree that has 4 cliques encircled with blue. Intersection in the junction tree is colored orange. Graphs and trees can be converted interchangeably
90 M. Yamada and M. Sugiyama
https://t.me/med1917
5.3 Graph Kernels
The graph kernel is a kernel function that computes the similarity between graphs by the inner product of features obtained from graphs, which can be plugged into any kernel-based machine learning methods like SVM and kernel PCA [ essential idea for considering structured objects like graphs has been established as R-convolutional kernels in the seminal paper by Haussler [
4], where we decompose
each object into sub-parts and count the common sub-parts to measure the similarity between them. The most basic and natural instance of R-convolutional kernels for graphs is known to be all-subgraph kernels, which is defined as
k
subgraph
(G, G') =
SGS'⊑G
k
isomorphism
'
(S, S'), (5.3)
where
k
isomorphism
(S, S') =
1if S and S 0 otherwise.
'
are isomorphism,
Although it gives a canonical way of computing the similarity between graphs, Gärtner et al. show that this computation is NP-hard [
5], hence it is not practical.
To date, a number of graph kernels have been proposed to efficiently and effec-
], and libraries for computing graph
tively compute the similarity between graphs [ kernels are widely available [7]. A kernel function k(x, x larity between x and features, that is,
x'of features. Every kernel function must be symmetric about
k(x, x') = k(x', x) and semi-definite [8]. The feature f computed
6
'
) is a measure of simi-
from a graph is mainly composed of graph specific structures. In the following, we introduce two representative graph kernels, the vertex histogram kernel, which is one of the simplest graph kernel, and the Weisfeiler–Lehman graph kernel, which is the most popular kernel.
3]. The
5.3.1 The Vertex Histogram Kernel
Given a pair of graphs G = (V , E) and G'= (V', E'), the vertex histogram kernel counts the joint occurrences of node labels in them [ of node labels, that is, generality. The node label histogram is given as
The vertex histogram kernel kV between G and G'is defined as
9]. Let d be the number of types
d =|∑v|, and assume that ∑v ={1, 2,..., d } without loss of
f = (f1,..., fd ) Nd of a graph G = (V , E)
fi =|{v V | lv = i}|. (5.4)
5 Primer on Graph Machine Learning 91
https://t.me/med1917
d
kV (G, G') =⟨f, f'⟩=
'
fif
. (5.5)
i
i=1
The time complexity of computing the vertex histogram kernel is O(|V |+|V'|).
5.3.2 The Weisfeiler–Lehman Graph Kernel
Although the vertex histogram kernel is efficient, it does not take any graph topolog­ical structure into account. To consider the subgraph structures, a number of graph kernels have been proposed. Among them, one of the most powerful graph kernels is the Weisfeiler–Lehman (WL) subtree kernel [ Lehman test of isomorphism [
11]. Fig. 5.3 shows the procedure of the Weisfeiler–
Lehman algorithm at the second iteration with respect to v simplicity, the compression for relabeling is not shown in Fig. label is determined after finishing iteration. The WL procedure is to aggregate node labels in the neighboring nodes and count their occurrence. For example, let us take the node label of
v0 and v
'
. In the first iteration, the neighboring node of v0 is v1, and
0
they are aggregated as a new label [0, 3]. In order to compress the representation, the new node label is replaced as with that of labels in G and
v0 at first iteration. After completing iteration over entire vertices, new
G'are shown in Fig. 5.4.
5 := l([0, 3]) in this case. The label of v
A new label obtained by the WL procedure represents the subtree structure of a
graph as shown in Fig.
5.4. At the end of the iteration, feature vector representation
is obtained as a series of counts of original and compressed node labels, resulting in the WL kernel
10], which is based on the Weisfeiler–
'
and v
0
. Note that, for
0
5.3. Each compressed
'
is the same
0
where φ is the vertex histogram of node labels after iterations.
5.3.3 Extend Connectivity Fingerprints
Extend connectivity fingerprints (ECFPs) are a class of topological fingerprints for molecular characterization [ related to the WL graph kernel procedure.
The algorithm of ECFPs is basically the same process of aggregating neighboring node labels, such as atom labels. After obtaining a new node label, it is relabeled by a hash function that converts it to unique integer values to represent the subgraph
kWL(G, G') =⟨φ(G), φ (G'), (5.6)
12
]. Although it is not a rigorous graph kernel, it is highly