Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5884_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
02.09.2026
Размер:
21 Мб
Скачать
388 J. E. Gonçalves
Table 13.1 Physicochemical, biological properties, and parameters used in the establishment of in silico methods with examples of the main platforms available for predictions
Examples of available
Properties Description
software
Level 1
Water Solubility (Log
)
S
aq
Ability of a compound to dissolve in an aqueous solvent, important for absorption
ChemAxon [7] ACD/ Percepta (ACD Labs) [8] SwissADME [9] EPI SUITE [10]
Log P - PgK
/ LogDOctanol/water Partition Coefcient. Indicative of the
ow
ability of a molecule to cross cell membranes
ADMETlab [11] admetSAR [12] SwissADME [9] EPI SUITE [10]
H
don/Hacc
Ability to donate or receive hydrogen bonds, an important component of Lipinskis rule determining the potential to be absorbed orally
TPSA Topological polar surface area indicates the polarity
of a molecule, the higher the value of this parameter, the lower the permeability through cell membranes
admetSAR [12] Corina Sym­phony [13]
Molinspiration [14] Corina Sym­phony [13] SwissADME [9]
Level 2
Violation of Lipinskis rule
Violation of 2 or more of the characteristics (Log P > 5; >5 H bond; >10 H bond acceptors; M
SwissADME [9]
wt > 500 Da) associated with low oral absorption
Vebers rule for good oral bioavailability
Effective permeability in humans (Peff)
a drug may be achieved if it possesses 10 or fewer rotatable bonds (RTB) and a polar surface area (TPSA) that does not exceed 140 Å
2
Determines the ability of a molecule to permeate through the human intestine
ADMETlab [11]
SwissADME [9] Dallphin-AtoM [15]
Apparent permeabil­ity (Papp)
F%Absolute oral bioavailability
Indicates the ability of a compound to cross mem-
SwissADME [9]
branes in in vitro or animal models Fraction effectively absorbed orally SwissADME [9]
Dallphin-AtoM [15]
PPB Plasma Protein binding, an important determinant of
distribution
ADMETlab [11] admetSAR [12] SwissADME [9]
V
d
P-gp substrate, inducer or inhibitor
Volume of distribution, indicative of the molecules ability to migrate from the blood to different peripheral tissues
The potential to be a substrate, inhibitor, or inducer of P-glycoprotein (P-gp), a key efux transporter that signicantly impacts bioavailability
ADMETlab [11] admetSAR [12] SwissADME [9]
ADMETlab [11] admetSAR [12] SwissADME [9] Simcyp [16]
(continued)
13 Challenges Faced in the Development of Computational Methods... 389
Table 13.1 (continued)
Examples of
Properties Description Blood-brain barrier
(BBB) partitioning
CYP substrate or inhibitor
Clearance The volume of blood effectively cleared of a mole-
T
½
C
max
T
max
AUC Area under the concentration– time curve (AUC),
Level 3
Target tissue exposure The ability of a drug to reach different tissues PK-SIM®[18]
Physiological absorp­tion model
Enzyme inhibitor binding process
Adapted from Madden and Thompson [6]
Blood-brain barrier partitioning refers to the ability of a substance to cross the blood-brain barrier
Predicts the predominant metabolic pathways and potential sites of metabolism in the liver Interactions with CYP1A2, CYP3A4, CYP2C9, CYP2C19, and CYP2D6 most commonly predicted as these enzymes are responsible for the metabolism of the majority of drugs. Particularly relevant for predicting drug–drug interactions
cule per unit time (per unit body weight); total clearance comprises contributions from renal excre­tion (Clr), hepatic clearance (Clh), and clearance via other routes such as metabolism, sweat, secretion into breast milk, and exhalation
Estimates the time needed for half of the drug to be eliminated from the plasma
Maximum concentration attained in blood, tissue, or an organ
Time to reach the maximum concentration in blood or tissue/organ
indicating overall internal exposure to the drug, in blood or a specic tissue/organ of interest
Absolute bioavailability and absorption in humans, animal models, and in vitro
Competitive and mechanism-based inhibition Simcyp [16]
available software
ADMETlab [11] admetSAR [12] SwissADME [9] Simcyp [16]
ADMETlab [11] admetSAR [12] SwissADME [9] Simcyp [16]
ADMETlab [11] admetSAR [12] SwissADME [9] Simcyp [16]
pkCSM [17]
PK-SIM®[18] PLETHEM [19] GastroPlus [20]
PK-SIM®[18] PLETHEM [19] GastroPlus [20]
PK-SIM®[18] PLETHEM [19] GastroPlus [20]
GastroPlus [20] GastroPlus [20]
3.1 Data Collection
This step is considered pivotal as it enables a comprehensi ve understanding of the systems behavior under investigation. The data that will inform the model can be obtained through various methods. Primarily, it is sourced from experimental data, gathered from scientic literature, chemical properties databases, and nonclinical studies. This dataset may include information on solubility, topology, ionization
390 J. E. Gonçalves
constant (pKa) permeability, metabolism, protein binding, and other relevant prop­erties. This diverse range of data is essential for accurately modeling and predicting the pharmacokinetic and pharmacodynamic behaviors of compounds, thereby enhancing the reliability and robustness of the in silico models [22].
Public databases such as PubChem [23], ChEMBL [24], and DrugBank [25] (please, see Chap. 2) are invaluable resources, providing comprehensive information on the chemical properties, biological activities, and pharmacokinetic data of various compounds [21]. These databases serve as crucial repositories of data that support computational modeling and drug discovery processes.
In specic scenarios, it becomes necessary to generate experimental data through in vitro and in vivo assays. This approach allows for the incorporation of empirical data into models, enabling a more accurate characterization of the pharmacokinetics of particular compounds. By integrating data from both in vitro assays and in vivo studies, researchers can enhance the predictive power and reliability of pharmaco­kinetic models, ultimately improving the drug development process [4].
3.2 Data Preprocessing
Given the available data, it is often necessary to undertake a process known as data cleaning. This involves rigorously analyzing the dataset to identify and eliminate inconsistencies, incorrect entries, missing values, and outliers [6]. During this stage, variable coding is also performed, wherein categorical variables are transformed into numerical formats. This process includes standardizing formats, magnitudes, and units to ensure consistency across the dataset. This step is crucial for ensuring that the data are suitable for subsequent modeling and analysis, thereby enhancing the accuracy and reliability of the computational models.
3.3 Identication of Relevant Variables
During this phase, the relevant variables for the conceptualized model are identied. Using statistical techniques or machine learning algorithms, the most critical phys­icochemical properties for predicting pharmacokinetics are determined. In certain cases, dimensionality reduction may be necessary. This involves applying methods such as principal component analysis (PCA) to address multicollinearity and improve the modelsefficiency [6 ].
13 Challenges Faced in the Development of Computational Methods... 391
3.4 Model Choice
The next step involves selecting the most suitable machine learning algorithm. The choice of model depends on the specic context, the size and nature of the data, and the research objectives [6]. Often, a trial-and-error approach, combined with cross­validation, is used to determine which model best ts the data [22]. Each model has its own set of advantages and limitations, and the selection should be guided by the specic requirements of the pharmacokinetic prediction problem. Commonly employed models include:
(a) Linear regression: This is a straightforward and interpretable method that pre-
sumes a linear relationship between the independent variables and the response [26]. It proves to be useful when a clear linear relationship exists between the characteristics and pharmacokinetic parameters.
(b) Support vector machines (SVM): SVM proves to be effective in both classica-
tion and regression problems. It is capable of handling complex and nonlinear datasets, making it useful when the relationships between variables are not strictly linear [27 ].
(c) Articial neural networks (ANN): Neural networks are potent models capable of
learning intricate patterns in data. They are particularly benecial in scenarios where the relationship between characteristics and outcomes is nonlinear and highly complex [28].
(d) Decision Trees: These are easy-to-interpret models that segregate data into
distinct sets based on decision rules . They are applicable in both classication and regression problems [29].
(e) Random forest: An enhancement of decision trees, random forest constructs
multiple trees, and amalgamates their results to bolster accuracy and mitigate overtting [26].
(f) Gradient boosting: Boosting methods like gradient boosting generate a sequence
of weak models that are combined to form a more robust model. They are effective in enhancing model accuracy [30 ].
(g) Bayesia n models: Bayesian models integrate uncertainties in modeling, proving
useful when it is crucial to consider uncertainty in model parameters [31].
(h) K-nearest neighbors (KNN): KNN is an instance-based learning model that
predicts the class or value of a data point based on the classes or values of its nearest neighbors [32 ].
(i) Nonlinear regression models: Specic nonlinear regression models, such as
polynomial regression, can be benecial when the relationship between vari­ables is more intricate than simple linearity [33].
(j) Physiologically based pharmacokinetics models (PBPK): This model is
constructed based on a vast number of drug physicochemical and absorption, distribution, metabolism, and elimination (ADME) attributes. These include lipophilicity, solubility, pKa, molecular weight, and plasma unbound fraction, along with physiological parameters like blood ow, tissue volume, vessel surface area, transporters, and enzyme expression level [34].
392 J. E. Gonçalves
3.5 Model Training
The model training phase involves using the chosen algorithm on the training data, allowing the model to discern and internalize patterns and relationships within these data [35]. Prior to commencing model training, the data are usual ly partitioned into three distinct subsets: training, validation, and testing. The proportions of these subsets can vary based on the datasets size, but a typical split is 50% for training, 30% for validation, and 20% for testing [36].
During the training phase, the algorithm is provided with data from the training set. The model adjusts its parameters based on these data, aiming to minimize the discrepancy between the models predictions and the actual values in the training set. This process allows the model to learn specic patterns and relationships present in the data [21].
Parameter Tuning: As the model undergoes training, certain algorithms have parameters that can be ne-tuned to optimize performance. Two common methods for analyzing the results obtained by the model are residual analysis and goodness­of-t[35].
In residual analysis, the difference between the values predicted by the model and those observed in experimental studies is calculated. An adequate model should exhibit an average residual value close to zero and should not show correlated residuals when plotting the dispersion diagram with predicted values on the ordinate and residuals on the abscissa, where no discernible trend in the points should be observed [35].
Goodness of t evaluates the correlation between predicted and observed values using metrics such as the coefcient of determination (R or coefcient of agreement. Parameter tuning is often performed using the validation set to avoid overtting, ensuring that decisions are not based solely on performance on the train ing data.
Training is considered complete when the model attains acceptable performance on the validation set. The nalized model is then evaluated on the test set, providing adefinitive measure of its ability to generalize unseen data [21, 22, 35].
2
), correlation coefcient (r),
3.6 Model Assessment
Once the model is prepared, it will be employed to evaluate a sample or a set of test samples to predict the desired information. Using these results, performance metrics such as mean absolute error, mean squared error, and the coefcient of determination
2
(R
) can be obtained to assess the models accuracy [35].
Upon analyzing results derived from external validation, it may be necessary to optimize the model by rening selected features, adjusting hyperparameters, or opting for a different algorithm [35].
13 Challenges Faced in the Development of Computational Methods... 393
3.7 External Validation
External validation of a pharmacokinetic computational prediction method involves evaluating the models performance on independent datasets that were not used during model training [21, 22]. The objective is to con rm that the model can effectively generalize to new, unseen data, providing accurate predictions in various contexts beyond those used for model development. This validation is essential to ensure the robustness and applicability of the method in real-world scenarios.
3.8 Implementation and Availability
The modeling process concludes with comprehensive documentation that provides a detailed description of all stages. This includes the criteria for data collection and the conditions under which the model should be used [35]. Typically, research groups submit this documentation for publication in compendia or scientic journals.
Following the publication of the model, the next step is to integrate it with drug development tools. This integration is accomplished using platforms and software, whether open-source or commercial [35].
3.9 Continuous Update
As the volume of information about a specic parameter increases, it becomes essential to update the model to improve its predictive accuracy and robustness . To achieve this, it is necessary to feed the model with the new data and restart all the previously outlined steps [35].
In a generalized and simplied manner, these are the steps involved in constructing a computational model for predicting pharmacokinetics. Each step is crucial for achieving the desired predictive power. Successful execution of the project relies on the collaboration between IT specialists and professionals from the pharmaceutical, chemistry, and biology elds. For more detailed information on the choice, as well as the implementation and subsequent evaluation of the model, it is recommended to consult Chap. 4.
394 J. E. Gonçalves
4 Challenges in Achieving Computational Models
with Enhanced Predictive Capability
Numerous challenges are encountered when developing a computational model to predict the pharmacokinetic behavior of candidate drug molecules. These challenges can be associated with each of the stages of pharmacokinetics processes and modeling development.
The initial challenge pertains to the quality of pharmacokinetic data used in constructing the model. Much of the data utilized is available in publi c databases, from companies, or published in scientic journal articles. Presently, there is a signicant surge in the availability of databases that can assist in predicting ADMET, such as the ADME database, SuperToxic, PKKB, and DSSTox [37
40]. However, the vast amount of information contained in these databases does
not guarantee the necessary quality to enable the models to exhibit the desired accuracy.
Another potential approach involves generating information through in-house experimentation. However, it is not always possible to have greater control over the quality of the data to be generated to feed the model. A limitation of this initiative is the use of a reduced number and variety of molecules to determine the parameters that will feed the model [21, 41]. Attempting to conduct a greater number of experiments to obtain in vivo data can be complex and costly.
Often, in the context of nonclinical studies, predictive models are also utilized, which can result in computational methods with less predictive power than those generated with data from in vivo studies. The literature presents discussions related to the quality of data obtained experimentally in different laboratories. Such incon­sistencies may be linked to errors in result acquisition, processing, and manipulation, as well as the use of inappropriate experimental methods [21] which, in certain circumstances, is associated with the difculty due to the inexperience of the programmer or modeler in critically analyzing the experimental data obtained from databases or literature.
Among the in vitro models typically used for computational modeling, methods for evaluating intestinal permeability stand out. These methods use synthetic mem­brane models, such as PAMPA (Parallel Articial Membrane Permeability), and cellular permeability studies, such as Caco-2 cell monolayers [42, 43]. Other studies include permeability studies across the blood-brain barrier, metabolism studies in hepatic microsomal systems, and plasma protein binding studies, among others [44
46]. The heterogeneity of results obtained by different laboratories for these in-house
analyses signicantly complicates the comparability of results, as well as subsequent analyses and predictions.
Recently, a measure has been adopted to address challenges related to data acquisition by aggregating sources from elds such as biol ogy, chemistry, pha rma­cology, and clinical trials to create big datasets for medicine research and development. However, this strategy still faces signicant obstacles, including
13 Challenges Faced in the Development of Computational Methods... 395
Table 13.2 Challenges and strategies to mitigate in silico methods applied on pharmacokinetics predictions
Challenges Strategies Reduced programming skills for building
model structures
The disparity between existing models and physicochemical and physiological processes
Difculties and limitations in obtaining exper­imental data
Restriction in predictive tools to estimate desired parameters due to the relative scarcity of data from available in vivo studies
Challenges in sampling in studies, covering intra- and inter-subject variability, as well as physiological differences between species, normal individuals and special populations
Adapted from Wang and Ouyang [51]
Multidisciplinary training create qualied and user-friendly platforms for Physiology-Based Physiological Modeling (PBPK)
Collect or estimate a more comprehensive range of physiological data and integrate a greater variety of physiological processes and tools to build more mechanistic models
Improve the design of in vitro experiments, carry out preliminary in vivo studies and establish connections between them
Create additional, user-friendly and qualied in silico tools for property prediction, and incor­porate data from these models
Perform more rened pharmacokinetic studies, adopt conservative conclusions, and consider model simplication or exploration in animal models
missing data, dimensional inaccuracies, and bias control challenges, which add complexity to big data analysis [48].
It is usually possible to obtain information in the literature about molecules that have been promising, have advanced in their evaluation or have become drugs. On the other hand, data on molecules that did not prove to be promising is not publicly accessible, constituting a private database for pharmaceutical companies [49]. This fact can harm the robustness of a computational method because it is only fed with a universe of results from molecules with selected characteristics. Therefore, to resolve these difculties, the development of security policies and data and infor­mation sharing are essential for building robust big data that takes into account greater expected variability. In this sense, organizations have worked to promote guidance by proposing recommendations for formatting standardized data that can be interchangeable [50].
An added challenge in establishing an in silico method is the necessity of selecting the most suitable machine learn ing algorithm. This choice calls for an enhanced understanding of the model to be implemented, necessitating the developer to possess profound knowledge of machine learning and deep learn ing models, which are crucial for predicting pharmacokinetic properties [51].
The ceaseless evolution and escalating complexity of AI-based models require researchers to swiftly comprehend new techniques capable of accurately predicting outcomes with diverse and large datasets. Table 13.2 presents the main challenges and possible ways to minimize their occurrence when developing an in silico method for predicting pharmacokinetics.
396 J. E. Gonçalves

5 Conclusions and Perspectives

The development of a new drug or medicine requires decision-making, which often entails discontinuing studies with a particular molecule or investing time and resources to progress through successive stages until an effective and safe medicine is achieved. These decisions frequently involve substantial uncertainty, posing a signicant challenge. The absence of adequately predictive met hods for pharmaco­kinetic characterization, target validation, and the identication and optimization of therapeutic candidates is currently viewed as the primary technical bottleneck in drug discovery [52].
Despite the advancements made, the evaluation of the predictive accuracy of existing in silico models continues to face considerable challenges. Numerous comparative studies have been conducted, as evidenced by the statistical data available in the scientic literature. However, the disparity in datasets used in individual research represents a recurring concern, and efforts to minimize it should be increasingly encouraged through the adoption of Good Laboratory Practices [53]. This disparity, as presented in this chapter and in various literatures, is attributed to the diverse origin of the data, which includes information generated internally, data published in the literature, and information extracted from public datasets [54]. This difculty could be minimized by incorporating a broader range of data sources, including clinical trials, electronic health records, and real-world evidence [55].
Future research is expected to request time and effort from multidisciplinary researchers to standardize sets of tests and evaluation criteria. This will facilitate more robust comparisons between in silico models, signicantly contributing to consistent advancements in the predictability of ADME proles and promoting the reliability and applicability of these computational approaches in pharmaceutical practice.

References

1. Storpirtis, S., Gai, M. N., Campos, D. R., & Gonçalves, J. E. (2011). Chapter 1: Farmacocinética: Conceitos, Denições e Relação com a Farmacodinâmica e a Biofarmácia (Biofarmacotécnica). In S. Storpirtis, M. N. Gai, D. R. Campos, & J. E. Gonçalves (Eds.), Farmacocinética Básica e Aplicada. Guanabara Koogan.
2. Pantaleão, S. Q., Fernandes, P. O., Gonçalves, J. E., Maltarollo, V. G., & Honorio, K. M. (2022). Recent advances in the prediction of pharmacokinetics properties in drug design studies: A review. ChemMedChem, 17, e202100542. PMID: 34655454.
3. Takebe, T., Imai, R., & Ono, S. (2018). The current status of drug discovery and development as originated in United States academia: The inuence of industrial and academic collaboration on drug discovery and development. Clinical and Translational Science, 11(6), 597–606. Epub 2018 Jul 30. PMID: 29940695; PMCID: PMC6226120.
13 Challenges Faced in the Development of Computational Methods... 397
4. Komura, H., Watanabe, R., & Mizuguchi, K. (2023). The trends and future prospective of in silico models from the viewpoint of ADME evaluation in drug discovery. Pharmaceutics, 15,
2619.
5. Khanna, I. (2012). Drug discovery in pharmaceutical industry: Productivity challenges and trends. Drug Discovery Today, 17(19–20), 1088–1102. Epub 2012 May 22. PMID: 22627006.
6. Madden, J. C., & Thompson, C. V. (2022). Chapter 3: Pharmacokinetic tools and applications. In E. Benfenati (Ed.), In silico methods for predicting drug toxicity (Methods in molecular biology 2425). Humana Press, Springer Protocols.
7. https://chemaxon.com/calculators-and-predictors
8. https://www.acdlabs.com/products/percepta/
9. http://www.swissadme.ch/
10. https://www.epa.gov/tsca-screeningtools/epi-suitetmestimationprogram-interface
11. http://admet.scbdd.com/calcpre/index/
12. http://lmmd.ecust.edu.cn/admetsar2/
13. https://www.mn-am.com/
14. http://www.molinspiration.com/
15. http://pipet.or.kr/board/apps_list.asp
16. https://www.certara.com/software/simcyp-pbpk/
17. http://biosig.unimelb.edu.au/pkcsm/
18. https://github.com/Open-Systems-Pharmacology/PK-Sim.wiki.git
19. https://scitovation.com/plethem/
20. https://www.simulations-plus.com/software/gastroplus/
21. Tran, T. T. V., Tayara, H., & Chong, K. T. (2023). Recent studies of articial intelligence on in silico drug distribution prediction. International Journal of Molecular Sciences, 24, 1815.
22. Grime Barton, P., & McGinnity, D. F. (2013). Application of in silico, in vitro and preclinical pharmacokinetic data for the effective and efcient prediction of human pharmacokinetics. Molecular Pharmaceutics, 10(4), 1191–1206.
23. Wishart, D. S., Feunang, Y. D., Guo, A. C., Lo, E. J., Marcu, A., Grant, J. R., et al. (2018). DrugBank 5.0: A major update to the drugbank database for 2018. Nucleic Acids Research, 46, D1074–D1082.
24. Kim, S., Chen, J., Cheng, T., Gindulyte, A., He, J., He, S., et al. (2019). PubChem 2019 update: Improved access to chemical data. Nucleic Acids Research, 47, D1102–D1109.
25. Gaulton, A., Hersey, A., Nowotka, M., Bento, A. P., Chambers, J., Mendez, D., et al. (2017). The ChEMBL database in 2017. Nucleic Acids Research, 45, D945–D954.
26. Lombardo, F., & Jing, Y. (2016). In silico prediction of volume of distribution in humans. Extensive data set and the exploration of linear and nonlinear methods coupled with molecular interaction elds descriptors. Journal of Chemical Information and Modeling, 56(10), 2042–2052. Epub 2016 Sep 28. PMID: 27602694.
27. Xue, Y., Yap, C. W., Sun, L. Z., Cao, Z. W., Wang, J. F., & Chen, Y. Z. (2004). Prediction of P-glycoprotein substrates by a support vector machine approach. Journal of Chemical Infor- mation and Computer Sciences, 44(4), 1497–1505.
28. Keutzer, L., You, H., Farnoud, A., Nyberg, J., Wicha, S. G., Maher-Edwards, G., Vlasakakis, G., Moghaddam, G. K., Svensson, E. M., Menden, M. P., Simonsson, U. S. H., & On Behalf of The UNITE4TB Consortium. (2022). Machine learning and pharmacometrics for prediction of pharmacokinetic data: Differences, similarities and challenges illustrated with rifampicin. Pharmaceutics, 14(8), 1530. PMID: 35893785; PMCID: PMC9330804.
29. Yun, Y. E., Cotton, C. A., & Edginton, A. N. (2014). Development of a decision tree to classify the most accurate tissue-specic tissue to plasma partition coefcient algorithm for a given compound. Journal of Pharmacokinetics and Pharmacodynamics, 41(1), 1–14. Epub 2013 Nov
21. PMID: 24258064.
30. Liu, H., Zhang, W., Nie, L., Ding, X., Luo, J., & Zou, L. (2019). Predicting effective drug combinations using gradient tree boosting based on features extracted from drug-protein heterogeneous network. BMC Bioinformatics, 20(1), 645.