Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_117_библиотеки_им_акад_М_И_Перельмана

.pdf
Скачиваний:
2
Добавлен:
15.09.2026
Размер:
15 Мб
Скачать
☆
116
expensive studies are carried out. By allowing for informed action while also taking into account model logic, medicinal chemistry expertise, and knowledge of the limi­tations of the system. The goal of XAI-assisted medication development is to aid in resolving some of these issues [43].
Data analysts, chemoinformaticians, and medicinal chemistry researchers will be able to collaborate together more effectively, thanks to XAI [44, 45]. Actuality, XAI already makes it possible to mechanistically evaluate how drugs work [46, 47] and it helps to improve medication safety and plan organic synthesis [48]. If long­term success is achieved, XAI will offer crucial assistance in the interpretation and analysis of ever-more complicated chemical data as well as in the development of fresh pharmaceutical ideas, all while avoiding human bias. Intense challenges with drug discovery, such as the coronavirus pandemic, may accelerate the development of application-tailored XAI algorithms to swiftly address particular scientic prob­lems pertaining to human biology and pathophysiology [49, 50].
Although XAI is still in its early stages, it continues to grow quickly, and I antici­pate that its importance will rise over the next years. I aim to present an in-depth overview of the current XAI research in this book chapter, emphasising its advan­tages, drawbacks, and potential for future drug development. After giving a brief overview of the most pertinent XAI approaches organised into conceptual catego­ries, the next section presents some current and future applications to drug develop­ment. I conclude by summarising the limitations of present XAI and suggesting potential advances in methodology that might help make these techniques more usefully useful in pharmaceutical research.
A. V. Geevarghese
2 Applications ofArticial Intelligence inDrug Discovery
2.1 Relation ofQSAR/QSPR withStructure-Based Modelling
withArticial Intelligence
Throughout more than 50years since its inception, QSAR/QSPR modelling has progressed tremendously [51]. The efcacy of these type of computational models for predicting biological activities and pharmacokinetic properties, which include absorption, distribution, metabolism, excretion, and toxicity (ADMET) [52–55], is unambiguous evidence of their impact on the development of drugs. Those so-called molecular descriptors are frequently employed to transform the structural character­istics of molecules (such as pharmacophore distribution, physicochemical proper­ties, and functional groups) into machine-readable values for ligand-based QSAR/ QSPR modelling [56]. There are many different types of manually created molecu­lar descriptors, each of which aims to communicate a different feature of the under­lying molecular structure. Support vector machines (SVM) and gradient boosting methods (GBM) are two widely used machine learning techniques that have typi­cally replaced fundamental models like regression models like linear and k-nearest
Explainable Articial Intelligence inDrug Discovery
117
relatives in QSAR/QSPR approaches, often at the expense of interpretability, to address more complex and likely non-linear connections between the structure of a compound and its physicochemical/biological properties [57]. Deep neural network applications are not new [58]. The 1990s saw the introduction of the bulk of the most recent advancements in chemoinformatics, including deep and adaptive net­work topologies, autonomous maps, recurrent systems for sequencing and time­series analysis, and autoencoders [59–61]. Deep networks, however, went the next step after winning the Merck Molecular Activity Competition in 2012 [61]. While there is much debate about whether this specic class of models outperforms com­peting tactics (such as gradient boosting machines) [62].
Methods for deep learning have multiple advantages when employing the same set of variables [63]. The potential of deep neural networks to autonomously gather features throughout training seems possibly the most important. Particularly recur­ring neural networks [64] and neural networks with graphs (which are additionally referred to as message-passing techniques) have the ability of producing intrinsic context-specic representation for chemical structures. This may be done in the specic case of graph neural networks by learning latent atom and bond representa­tions during the training stage. As a result, modelling activities for which traditional descriptors were not initially designed is possible using deep learning methodolo­gies. Examples include macrocycles [65], proteolysis-targeting chimaeras (PROTACs), and modelling of peptides [65, 66].
Deep architectures may also benet from multitask learning, which seeks to identify a common internal representation useful for a collection of connected end­points. It differs from multi-output learning in that it doesn’t explicitly take use of relationships between the tasks that need to be learnt [67–69], perform a variety of activities. Since drug development is a multi-parameter optimisation issue, learning may be able to better take advantage of data correlation without the necessity for previous imputation in cases when a chemical library has not been thoroughly eval­uated on all relevant outcomes. Prior to the adoption of deep learning approaches, the concept of multi-output QSAR simulation, which tried to connect a collection of identied chemical descriptors to measurements, was researched [70–75]. Despite the potential of multitask learning, it hasn’t yet been demonstrated that it can out­perform single-task models [76–79]. Deep learning’s poor performance when there is little to no data is a well-known issue [80]. By using additional genetic or biologi­cal interactome data sources, certain chemogenomic-based techniques could be able to provide further light on these scenarios [81]. Additionally, there have been recent developments in “few-shot” learning [82] and meta-learning [83] (a family of approaches that aims to provide a set of learnable parameters that can quickly adapt to new, unknown jobs). Therefore, in contrast to approaches that are totally or par­tially based on physics, the capacity of completely data-driven methodology for molecular characteristic projections to extrapolate and generate trustworthy predic­tions for unknown chemical classes is highly constrained. The introduction of extra active learning approaches (strategies where the model participates in requesting particular training data for enhanced generalisation) and physics-inspired machine learning algorithms can overcome these constraints [84, 85]. Given that sufcient
118
A. V. Geevarghese
sources that would enable good data imputation are usually hard to come by, how effectively each of these techniques’ individual implementations handle data spar­sity will also have a signicant impact on how effective they are [86]. The “black­box” nature of deep learning models and their often challenging debugging have also attracted a lot of criticism [87]. To manually include background information in a way that is easier to grasp, domain-specic features [88, 89] (i.e. descriptors explicitly constructed with a specic objective) are still an option. By offering understandable interpretations of the decision-making process used by deep learn­ing systems, explainable AI techniques may be able to provide some partial reme­dies to these issues [90]. A gap between deep learning and drug discovery knowledge will be able to close with the ongoing development of feature attribution method­ologies [89]. Examples of instance-based explanations include counterfactuals, model-generated instances that are conditioned on user-dened queries, and atten­tion-based networks [89, 90]. The high expense of deep learning techniques is another drawback that is frequently mentioned. Deep learning usually requires lengthier training and evaluation times than many other machine learning tech­niques because it requires specialised hardware, such as tensor processing units or consumer-grade graphics processing. Although the aforementioned supposition is usually accurate, deep learning models can be capable of learning in an Internet environment by automatically utilising its most well-liked training approach, sto­chastic gradient-descent optimisation [90].
The benet of this is that it grows exponentially in proportion to the size of the training dataset, preventing the latter from needing to use the whole system’s mem­ory. Since deep learning models may be stochastically trained on sequential, ran­dom batches of data, researchers contend that they may perform better than competing solutions in big data environments [91]. In a similar vein, predicting deep learning typically requires far more human expertise in many real-world cir­cumstances compared to other, more well tested approaches. Despite the ease with which a high-performing random forest model may be trained for hyperparameter adjustment, our understanding of current deep learning approaches is still insuf­cient to provide trustworthy defaults [92]. However, recent theory suggests that this may change soon.
Furthermore, even when the predictions are obviously incorrect, neural networks show a propensity to give correct responses for deceptive reasons (such as the infa­mous Clever Hans effect [93]). The dilemma is made signicantly worse by the possibility that comparable trial circumstances might provide results that are notice­ably different when used to forecast characteristics in drug development. The wide­spread application of uncertainty estimation approaches, whether through deep learning systems that explicitly include uncertainty into their design, like Bayesian neural networks [94], or post hoc techniques, such ensemble learning [95], should minimise this problem in the following years. In contrast to classical QSAR, which needs a co-crystal or a docking equilibrium to generalise over numerous targets, incredible progress has also been made in the structure-based modelling of protein­ligand activity. To accurately account for the impact from individual descriptors (such as physical and chemical properties) on a target property, many conventional
Explainable Articial Intelligence inDrug Discovery
119
methods used partial least squares or multiple linear regression models to model an explicit, predened mathematical connection of the protein-ligand complex [96–98]. The use of methods that combine various descriptors, such as protein-ligand atom pair counts [99], property-encoded shape distributions [100], or basic atomic inter­actions, with sophisticated and exible non-linear models, such as random forests or support vector machines, increased in the early 2010s.
This unique subject has lately observed the emergence of deep learning and used it, much as its strictly ligand-based cousin. Early methods for predicting bioactivity were impacted by the advancement of computer vision and picture recognition, which was primarily driven by convolutional neural networks [101]. To achieve the same result, further study [102] combined graph-based techniques with feature enhancements based on distance and angle. In structure-based virtual evaluation and lead optimisation competitions, several of them were purportedly shown to pro­vide marginal performance advantages over existing methodologies [103, 104]. However, [105–109] it is debatable if certain reputable benchmarks favour ML-based grading systems over traditional ones. One conceptual limitation of approaches based on three-dimensional convolutional neural network models is the lack of rota­tional invariance with reference to the input, a quality crucial for representing atomic systems. Thanks to recently created neural network architectures like the Euclidean Neural Networks [110–112] and SchNet [113], which directly incorpo­rate equivariance with respect to the special Euclidean group in three dimensions (SE(3)) (i.e. rotations and translations) into their design, how to approach this prob­lem has recently become a very active area of study. In the past, these structures have been used for a number of molecular activities, including the study of mole­cules’ electrical characteristics [114].
It is predicted that further study will be conducted in this area in the future, expanding the modelling possibilities. Because deep learning applications in drug development are expanding quickly and require sizable training sets, thorough data curation and appropriate benchmarking of newly developed models are crucial. Chemical substance libraries have grown in size and accessibility over the past sev­eral years, with tools like ZINC [115] and ChEMBL [116] acting as standard entry points for ligand-based programmes. The same pattern was observed for structure­based modelling, for which databases like PDBbind [117] and BindingDB [118] provide incredibly precise structural information on protein-ligand complexes together with information on the biological activity associated with such complexes. The prospect of soon having access to structural data for a large number of potential therapeutic targets is encouraged by recent developments in protein structure pre­dicting and determination [119].
A lot of money has already been spent on open, standardised assessments of machine learning techniques in the eld of chemoinformatics. A quick evaluation of numerous important deep learning methods for drug-related property predictions in well-curated datasets from disciplines including biophysics, physical chemistry, and physiology is provided by the MoleculeNet benchmarking suite, in particular. These improvements to AI based drug developments helps pharmaceutical rms, publishers, and commercial research bodies continue to produce most structural
120
activity/property connection data [120–122]. Despite the fact that we have main­tained that the amount of public data continues to grow rapidly, who typically see the information collected as a differentiating advantage that should be kept private. Recent work suggests that molecular descriptors are routinely used to partly rebuild molecular structures, which may make it more challenging to communicate data even at the latent feature level [123]. There have been several attempts to circum­vent these limitations, such as the creation of federated and IP-preserving learning systems [124]. Since then, it has become clear that using sets pulled in a pseudo­random manner from a database to test a model’s performance might result in too optimistic results. Alternatives like scaffold-based [125] or time-based splits [126], which aim to approximate the development of a lead optimisation project, may be more illuminating.
Although there is no “one-size-ts-all” method, it is important to remember that each evaluation shows how well a model performs in a certain application area. Potential applications ought to be thought of as the ideal situation for model bench­marking, but we can show that they are not always objective and are not without bias [127]. Machine learning scoring algorithms [128] have been shown to be rea­sonably predictive in a number of virtual screening initiatives [129], even if there isn’t a general benchmarking agreement. The limitations of employing proper per­formance measures for regression and classication models have also received a lot of attention.
A. V. Geevarghese
2.2 Articial Intelligence-Based Approaches inDe Novo
Drug Design
De novo design, which entails the creation of novel molecular structures with desired pharmacological characteristics from scratch, can be regarded as one of the most challenging automated technology tasks in drug discovery due to the cardinal­ity of the chemical eld of drug-like molecules, which is thought to range in the order of 1060–10,100 [130, 131]. De novo molecule synthesis is complicated by the combinatorial issue even if there may be a vast array of potential atomic forms and molecular structures to examine [132]. Depending on the data used to guide the de novo design, similar methodologies may be ligand-based, structure-based, or any mix of the two [133].
Another approaches which lead to drug discovery in de novo-based approaches is ligand-based methodologies. Ligand-based methods are important area of research drug discovery and optimisation process, this approach doesn’t need the isotopic labelling of targeted proteins. Ligand-based methodologies in de novo drug design approaches can be broadly categorised into two main categories: (i) rule­based methods, which use a set of constructing rules for molecules to be built from a variety of “building blocks” (such as the reagents or molecular fragments), and (ii) rule-free methods, which do not use clear construction regulations. One of the
Explainable Articial Intelligence inDrug Discovery
121
forerunners of contemporary rule-driven de novo design is the Topliss technique [134], for the serial production of analogues of a strong lead molecules with the maximum potency. Modern methods entail employing a specied set of chemical transformations for optimisation, such as matching up molecules in correlation [135] or using molecular structure and functional group change rules-of-thumb [136]. Building block assembling and ligands creation are specically included in synthesis rules in synthesis-oriented techniques. These techniques can be applied, for illustration, to the development of electronically accessible libraries, like BI CLAIM [112] and CHIPMUNK [135]. During the late 1990s, hybrid techniques have been to guide the creation of novel compounds by jointly maximising their similarity to recognised bioactive ligands and the chemical synthesisability of the designs, such as TOPAS [136], DOGS [137], and DINGOS [138]. They were devel­oped to regulate the synthesis of novel molecules by optimising both of the design’s resemblance to already- known active interactions and their potential for chemical synthesis.
Overcoming molecular developing standards, rule-free strategies aim to generate compounds with specied characteristics. Contemporary techniques frequently depend on generating models based on deep learning [139], which take samples new atoms from a hidden chemical description that has been learnt. The concept of choosing a molecule from a numerical model for de novo synthesis is related to the “inverse QSAR” problem discussed in Skvortsova and Zerov’s landmark work in the early 1990s [140–142]. The use of these approaches is looks increasing in recent years. Reverse QSAR employs an earlier QSAR model to pinpoint the description values that t a desired attribute while making molecules.
With the reason of producing molecules, inverse QSAR utilises a current QSAR modelling to determine descriptor variables that correspond to an ideal trait. The latter approaches offer an array of disadvantages, notably the challenge in reverse­decoding the descriptors of molecules into suitable structures and the existence of multiple options for each specic characteristic. Creative machine learning handles some of these problems by simulating the underlying structure of a certain group of chemicals and then creating new molecules by choosing the obtained distribution [143]. Most frequently used generative models combine Simplied Molecular Input Line Entry Systems (SMILES) with natural language processing techniques [144]. The models in question undergo training to acquire the SMILES “syntax” (which describes about the capacity to generate a scientically acceptable string) on selected “semantics” (i.e. its similar appealing structural characteristics or bioactiv­ity). Recurring articial neural networks [145, 146] and transfers or learning by reinforcement [147–149] were the primary underpinnings of these systems. Several well-known deep learning-derived generated models for learning, such as varia­tional autoencoders [150], generative networks of adversarial networks [151, 152], as well as others that utilise graph the convolutions, have additionally been widely published [153]. Recently, instances of conditioned productive approaches are being offered. These methods employ more data to direct the design process, includ­ing molecular descriptor values [154], expression patterns [155], drug-likeness syn­thesisability, shape in three dimensions, and similarity to drugs. In this context, the
122
development of harmonious objectives that permit complex and constrained multi­parameter optimisations, such as those employed in Pareto [156] or in desirability­based techniques [157], which are often required in the discovery of pharmaceuticals, will provide a considerable challenge in the future.
Most of studies on deep learning-driven de novo synthesis thus far have concen­trated on ligand-based approaches. With the reason of concentrating on orphaned receptor and formerly unexplored macromolecules. These structure-based design using generative algorithms offers an exciting additional study area [158]. For the greatest extent of our knowledge, machine learning is still not substantially inte­grated into these methods, which typically utilise knowledge about the site where the ligand binds (e.g. via fragment linkage or growing). The makeup and features of the binding site were nevertheless taken into consideration in the early stages of ligand design [159–161].
A. V. Geevarghese
3 Articial Intelligence-Based Design forAutomated
Drug Synthesising
The bulk of all known chemicals can be synthesised using a select few reliable tech­niques [162]. Chemistry still stands in the way of reliable, fully automatic synthesis planning [163]. One of the reasons is the extensive chemical understanding required for efcient forward and retrosynthetic planning [156]. Synthesis planning using AI has a lengthy history in the eld of computer-aided retrosynthetic prediction, dating back to the 1970s. The use of articial intelligence (AI) for organic synthesis has experienced a comeback as a consequence of enhanced processing power, the emer­gence of big data, and the creation of novel deep neural networks and optimisation techniques. Retrosynthesis, where the main objective is to repeatedly create ef­cient synthetic pathways for the target molecule, unquestionably benets from rule­based techniques. They attempt to identify retrosynthetic paths by storing mechanisms for reaction and building skeletal structure. Their dependence on direct chemical modications or reaction constitutes one of their key disadvantages. Usually, these require human developing and curation. In recent years, methods employed in the processing of natural languages, such sequence-to-sequence struc­tures and transformers designs, have acted as motivation for this eld of study [164]. The reality that the order of arrangement of fragment in molecular biology matches that of phrases in the English language provides an impetus for this area of research [165]. Rule-free strategies often take into consideration outcomes in written repre­sentations (such as SMILES) and analyse them using an architecture consisting of encoders and decoders in order to foresee the related synthetic precursor at an a step response distance [166]. An improvement over this architecture is provided by tiered articial neural networks [167], which divide the retrosynthesis predictions issue into response type categorisation and response rule selection processes. A molecular similarity approach that had previously been disclosed [168] and which
Explainable Articial Intelligence inDrug Discovery
123
has been demonstrated to give improved performance over earlier baselines for comparison served as the impetus for the creation of this division. The bulk of the previously discussed solutions concentrate on the linear one-step retrosynthesis problem, however there is also a combinatorial opponent that is gaining ground.
One of the most signicant developments in the last 10years has been the effec­tive exploration of chemically reactive spaces using advanced search techniques like Monte Carlo Tree Search [163]. This development was sparked by improve­ments in reinforcement learning. One-step precursor predictions and the construc­tion of hypergraphs, or directed acyclic graphs with edges that can link several nodes at simultaneously, were employed in a recent study [165] to represent fake pathways in an effort to better understand the reactants and reagents. While the majority of the remedies described previously focus on the linear one-step retrosyn­thesis issue, an alternative scenario includes a combinatorial opposition that is rap­idly expanding.
Despite the fact that such problems may be handled by using reaction data already available, forward synthesis requires knowledge from reactions that pro­duce no products at all. The databases now in use for chemical reactions are heavily biased in favour of data on successful reactions [169]. There is a critical need for further data, such as details on byproducts or experimental conditions (such solvent and temperature). Some steps have been done to extend known reaction databases with unfavourable reaction outcomes in an effort to get around some of these restric­tions [170]. By doing this, new tailored data compilations for automated synthesis planning have also been produced [171]. Earlier methods used are data-derived reaction templates and ranking machine learning proof-of-concept response tem­plates to rate candidate compounds [170, 171], once the information on reactants and reagent had been given [172]. The objective of newer techniques is to rate com­pounds immediately by approaching the problem of chemical response predictions as a graphical conversion job [173]. A different set of techniques opted to employ rst-principle computations to evaluate the energy obstacles of a specic procedure, prompted by advances in the eld of quantum mechanics.
For medium-to-large structures, this approach is physically impractical. This dis­crepancy might soon be bridged by quantum-mechanical articial intelligence’s accurate estimations of energy and force [171]. The use of natural language process­ing techniques that utilise the transformers [172] or recurring neural networks archi­tecture [173] are additionally gaining popularity with regard to template-free forward synthesis predictions. A top-1 reactant precision over 90% was observed documented for them. Some novel alternative deep learning algorithms [174, 175] chose to represent reaction predictions as an electron rearrangement exercise in addition to using message-passing neural networks to learn. The latter method, however, lters out many pertinent organic processes because they cannot be clearly identied as electron ows.
124
A. V. Geevarghese
4 Discussion
Given the variety of explanations and methods that may be used to complete a task [176], current XAI also faces technological difculties. The majority of approaches need to be customised for every application instead of being offered as “out-of-the­box” solutions which can be used instantly. In addition, in order to gure out which model decisions require additional justications, what kinds of responds are impor­tant to the consumer and which ones are just simple or expected [177], an in-depth comprehension of the issue area is essential. Human decision-making justications produced by explainable articial intelligence must be complex, plausible, and use­fully informative for the relevant scientic community. It is going to be essential to explore further the benets and drawbacks when utilising traditional chemical lan­guage for expressing the range of choices of these models. Drawing on comprehen­sible “low level” chemical representations that are appropriate for algorithmic learning and have direct meaning for chemists (such as SMILES strings [145, 178], sequences of amino acids [179, 180], and different three-dimensional voxelised rep­resentations is an advance in the correct direction. Several recent studies utilise established biochemical descriptors, that capture structural characteristics that are predened a priori, such as hashed binary ngerprints [181, 182], topochemical and geometric descriptors [183, 184]. Because they may be more readily articulated using the well-known language of chemistry, molecular representations have a clear inclination to be used when attempting to implement XAI.The interpretability of the model is inuenced by both the specied machine learning approach and the chemical description. In light of this, creating unique, understandable molecular representations for machine learning will be an important area of research in the years to come. The next step will also involve the creation of simple methods that circumvent the problems posed by non-interpretable yet dense data descriptions by making sufciently accurate forecasts and providing explanations that are compa­rable to those provided by people. Because there are currently no techniques that incorporate all of the outlined desirable XAI features (transparency, justication, informativeness, and uncertainty estimation), consensus (jury) techniques that com­bine the benets of various (X)AI approaches and boost model dependability will be crucial in the short- to medium-term. In the long run, juror XAI techniques will represent a means to offer numerous perspectives on the simulated biochemical process by depending on various algorithms and chemical representations. The bulk of machine learning drug discovery techniques now in use [185, 186] ignore appli­cability domain limits, or the region of the chemical space where statistical learning assumptions are met. According to the author, these constraints ought to be consid­ered an essential part of XAI as an accurate assessment of modelling accuracy has been shown to be of greater signicance in making choices compared to the model­ling method itself [187]. Knowing whether to employ a particular algorithm would undoubtedly assist address the issue of deep learning models’ high condence in incorrect predictions and prevent needless extrapolations simultaneously. The bulk of machine learning drug discovery techniques now in use [185, 186] ignore
Explainable Articial Intelligence inDrug Discovery
applicability domain limits, or the region of the chemical space where statistical learning assumptions are met. Knowing whether to employ a particular algorithm would undoubtedly assist address the issue of deep learning models’ high con­dence in incorrect predictions and prevent needless extrapolations simultaneously. In light of this, machine learning practitioners who work in time- and money­sensitive scenarios, such as drug discovery, have a responsibility to carefully exam­ine and analyse the predictions resulting from their modelling decisions. There is presently no open-community platform for XAI in the creation of pharmaceuticals, with the aim of sharing and enhancing software, modelling interpretations, and related training data through cooperative efforts of academics with varied scientic backgrounds. The initial step in the right direction is being implemented by initia­tives such MELLODDY (Machine Learning Ledger Orchestration for Drug Discovery), that aims at creating federated, decentralised modelling for secure han­dling of information among pharmaceutical companies. Such collaborations ought to promote the creation, verication, and adoption of XAI and the rationales these tools offer.
125
5 Conclusion
Complete comprehension of models based on deep learning may be difcult in the environment of therapeutic research, although the provided forecasts might still be helpful to the researcher. It is going to be needed to meticulously organise an array control examines to assess the machine-driven ideas and improve their reliability and neutrality whereas aiming for interpretation that closely match the senses of humans. Additionally, there is proof suggesting applications of articial intelligence are starting to be utilised extensively in the eld of drug research and design cycle. Considering signicant recent advances in the QSAR simulation, de novo molecu­lar design etc., these techniques are now gradually coming to many of the commu­nity’s aspirations. Yet, it is yet to be seen how these approaches are going to be successful in helping scientists develop and synthesise “more effective drugs quickly.” Within the setting of predicting characteristics associated with ligands that techniques depending upon relatively “raw” biochemical depictions, like neural net­works using graphs and SMILES-based neural networks with recurrent neurons, are expected to perform a minimum as well as descriptor-based models. These tech­niques also allow for better use of information, such as via multitask and Internet­based learning, and are readily applicable to a bigger category of chemical substances and modelling tasks. On the contrary, conformation-aware machine learning is still in its early stages, particularly when taking into consideration methods that include three-dimensional symmetry into the design of the system. However, it is realistic to expect swift progress in the application of them for the discovery of drugs as well as associated elds like quantum physics or material research, especially as an alterna­tive for rst-principle computations, that are substantially more challenging. Over the last couple of decades, rules-based and rule-free approaches to de novo drug