Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5586_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
35 Мб
Скачать
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
long piece of DNA has been sequenced, the rst objective is to regulate if it contains any genes. This objective is comparatively simple in the case of prokaryotes, but the search for eukaryotic genes is quite difcult [39].
9.4.1 Identication of genes
With the increase in genomic and functional genomics information, approaches for disease gene identication are rapidly being developed [40]. Databases are now indispensable to the process of selecting candidate disease genes. Relating positional information with disease characteristics and functional information is the normal approach by which candidate disease genes are selected. Enrichment for candidate disease genes, however, depends on the skills of the operating researcher [41].
Prior to genome sequencing, gene identication used approaches based on tran­script mapping. The genomic clone of interest was exposed to zoo blot hybridization, i.e., hybridization with the whole genomic DNA of a range of species. It has been discovered that coding sequences are strongly conserved during development. Thus, a clone that tested positive in a zoo blot is possible to characterize as a gene. The clone of interest can be hybridized with cDNA libraries or employed for northern blot or reverse northern blot assays. A constructive assay identies a gene as only genes are transcribed. CpGProD is an application for identifying mammalian promoter regions associated with CpG islands in large genomic sequences [24]. Detection of CpG islands in the clone suggests it to be a gene as about 50% of human genes have related CpG islands. CpG islands are short stretches of G. C-rich DNA often found in association with vertebrate genes. Other methods of gene identication included cDNA selection, cDNA capture and exon trapping. These methods provide research that is appropriate for individual gene identication but unsuitable for genome annotation. Consequently, the correct reading frame is identied by carrying out a six-frame translation of the DNA sequence. The accurate reading frame is predicted to be the longest frame uninterupted by a stop codon (TGA, TAA or TAG); this type of reading frame is known as an open-reading frame (ORF). The longer this ORF, the better is the probability that it represents a gene. Locating the 3-end of an ORF is relatively easier than locating its 5-end, although the 5-end may be indicated by an ATG. However, more tools are required to locate the 5-end of an ORF, such as the presence of a Kozak sequence (CCGCCA UGGG) anking the AUG codon. Investigation of codon usage may also be supportive, and the presence of CpG islands may designate the 5-end of numerous vertebrate genes. However, determi­nation of ORFs may be hindered by sequencing errors. The most signicant single progress in genome annotation is the use of PCs to predict genes from DNA sequences. Identification of genes may achieve using one of the following two approaches. In the rst technique, sequences of known genes, cDNAs, ESTs and proteins present in databases are equated with the genome sequence; this is achieved by homology search techniques, such as BLAST. In the second method, specic software is employed to identify genes [42].
In prokaryotes, programs such as GenMark (a modied GeneScan algorithm) and Glimmercanidentify all the genes, including overlapping genes, present in the genome.
9-13
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Numerous sophisticated software programs have been developed for gene prediction in eukaryotes, e.g. Genie, GeneScan, Grail, GeneFinder, HMM Gene, etc. These programs required training with information or need to be provided with a set of rules. Grail can be employed with human, mouse, Arabidopsis, etc, genome sequences, but Genie has only been trained on human and Drosophila sequences. However, no single program is 100% perfect at identifying genes. An algorithm trained with the DNA sequence of one organism, Caenorhabditis elegans, will not execute suitably with that of another organism, e.g. a plant, without being retrained [43].
The gene prediction programs search for gene-specic characteristics, e.g. promoters, splice sites and polyadenylation sites, or for pertinent gene content such as ORFs. The currently available gene search programs are associated with different search criteria and their sensitivities vary widely. The identication of ORFs, usually exceeding 300 nucleotides, is sufcient to nd most genes in prokaryotic genomes. However, such a simple search criterion will miss smaller genes and overlapping genes. These problems are resolved by using algorithms that consider differences in base composition between genes and noncoding DNA, e.g. in GenMark. The gene prediction programs used in eukaryotes use the output of several algorithms to generate a whole gene model. In this model, a gene is dened as a series of exons that are coordinately transcribed. The various features of eukaryotic genes recognized during gene detection include transcriptional and transla­tional controls, e.g. the TATA box, cap site, Kozak consensus and polyadenylation sites. But problems arise as the TATA box is missing in 70% of human genes, and polyadenylation signals can differ considerably from the consensus sequence AATAAA. Moreover, these sequences recognize only the rstandthelastexonsofa gene. Thus, additional features have been encompassed in modern genesearchtools;these features include 5-and3′-splice sites, differences in base composition between coding and noncoding DNA (typically, comparison of hexamer base composition), etc.
9.4.2 Identication of the function of a new gene
A convenient approach to detect the function of a new gene is as follows. The gene sequence is translated into the amino acid sequence and thus the protein it is anticipated to encode. This protein sequence is then compared with a protein database. A program such as tBLASTx will execute both these functions. If the encoded protein is homologous to a protein in the databank, it allows identication of the new gene and also suggests the function of the new gene [44]. The FASTA and HMMER programs are slower than BLAST, but they are more sensitive.
9.4.3 Identication of functional domains
Several bioinformatics techniques for the detection of protein motifs and protein domains have been explored. Among these tools are PRINTS, PROSITE, SMART, BLOCKS, etc.
9.4.4 Detection of noncoding RNA
A number of RNAs are noncoding, such as rRNA, tRNA, small RNAs, etc. Of these, rRNAs are the easiest to nd; this is done by sequence similarity search. The program tRNAScanSE searches for aRNA sequences [44].
9-14
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
9.4.5 Genome annotation
After the detection of a gene and prediction of its functions, suitable experiments have to be devised to authenticate these ndings. This information is then employed for genome annotation. Standard genome annotation languages have become extensively accepted [45]. GAME is a program for describing investigational conrmation to support annotation. Likewise DAS (the distributed annotation system) is mainly useful for indexing and visualization. Several software programs such as BioPerl 2001, BioJava 2001, etc, are employed for storing, manipulating and imagining the genome annotations [46].
9.4.6 Molecular phylogenetics
DNA and protein sequence statistics can be employed to examine the develop­ment of genes and their protein products; this is called molecular phylogeny. Bioinformatics techniques are employed to dene phylogenetic relationships, which may be accessible in the form of either a phylogenetic tree or a dendo­gram. While the two depictions look different, they depict exactly the same relationship [47].

9.5 Computational approaches in bioinformatics

9.5.1 Algorithm development
Sequence alignment is a crucial task in bioinformatics, which provides the groundwork for several applications, such as functional annotation, structure prediction, and phylogenetics. The main goal of sequence alignment is to arrange DNA, RNA, or protein sequences to nd comparable sections that may result from functional, structural, or evolutionary links between the sequences. Sequence alignment may be roughly classied into multiple sequence alignment (MSA) and pairwise alignment [48]. Two sequences align in pairwise alignment, while three or more sequences align in multiple-sequence alignment [49]. The Nee dl eman – Wunsch technique [50] for global alignment and the Smith–Waterman algorithm [51] for local alignment are two popular pairwise alignment algorithms. The goal of global alignment, which works best for equal-length sequences, is to align every residue in every sequence. Local alignment, on the other hand, nds comparable sections within lengthy sequences that are often distinct overall. Algorithms like MUSCLE [52] and ClustalW [53] are widely used in multiple sequence alignment. MSA is crucial for phylogenetic analyses and identifying conserved domains across species. It is more complex than pairwise alignment due to the increased computa­tional demands and the complexity of optimally aligning multiple sequences. A fundamental aspect of sequence alignment algorithms is the scoring system, which uses substitution matrices to score alignments based on the substitution, insertion, or deletion of residues. Gap penalties are a crucial part of the scoring system, imposing penalties for opening and extending gaps, which is essential for manag­ing indels (insertions or deletions) [54]. Various algorithmic strategies are employed in sequence alignment. Dynamic programming is a common approach
9-15
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
used in both the Needleman–Wunsch and Smith–Waterman algorithms to ensure the optimality of the alignment. However, due to the computational intensity of dynamic programming, empirical methods are often employed to accelerate the alignment p rocess at the cost of optimality, with BLAST (Basic Local Alignment Search Tool) [55] and FASTA being prime examples [56]. In the case of multiple sequence alignment, a progressive alignment strategy is often employed, where pairwise alignments are extended to multiple alignments b y aligning the most similar sequences rst. Advancements in computing have also impacted the eld of sequence alignment. Parallel computing, for instance, has signicantly reduced the computational time required f or alignment tasks, making real-time analysis of large datasets feasible. Additionally, the integration of machine learning techni­ques has emerged as a promising avenue to improve alignment accuracy and predict novel alignments, signifying a blend of traditional bioinformatics approaches with modern computational methods [57].
9.5.2 Phylogenetic tree construction algorithms
Phylogenetic tree construction is fundamental for understanding evolutionary relationships among species or sequences. Various algorithms have been developed to implement this method, and they fall into different categories based on their approach [58]. A schematic diagram representing the structure and ow of the phylogenetic tree is shown in gure 9.4.
Figure 9.4. Phylogenetic tree representation: mapping evolutionary relationships and ancestral lineages of present-day species.
9-16
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
9.5.2.1 Distance-based algorithms
Neighbor-Joining (NJ) [59] and UPGMA (Unweighted Pair Group Method with Arithmetic Mean) are distance-based algorithms [60]. These algorithms begin by computing a matrix of pairwise distances between sequences. NJ is a widely used method due to its efciency. It iteratively joins the nearest neighboring taxa into a tree, updating the distance matrix in each iteration to reect the newly formed clusters. The UPGMA assumes a constant rate of evolution (molecular clock hypothesis) across different lineages, which may not always hold, thus limiting its accuracy in some scenarios [61].
9.5.2.2 Character-based algorithms
The character-based algorithms include maximum parsimony [62] and maximum likelihood. The maximum parsimony method seeks to nd the tree that explains the data with the fewest evolutionary changes. Maximum likelihood estimates tree topologies and branch lengths based on a probabilistic model of sequence evolution, aiming to nd the tree that maximizes the likelihood of observing the given data [63].
9.5.2.3 Bayesian phylogenetics
The Bayesian phylogenetics approach employs Bayesian assumption to estimate the posterior probabilities of different tree topologies given in the data, incorporating a priori knowledge through prior distributions [64].
9.5.3 Machine learning algorithms in bioinformatics
Machine learning (ML) encompasses various algorithms applied to various bio­informatics tasks. Support Vector Machines (SVM) are used for classication tasks like distinguishing between disease and healthy samples based on gene expression data. Random Forests are ensemble learning methods used for classication and regression tasks, providing feature importance scores that can be insightful in bioinformatics. Unsupervised learning includes K-means clustering [65]. It is employed for grouping data into clusters based on feature similarity, like clustering genes based on expression proles. Hierarchical clustering is useful for generating a dendrogram of clustered data, often employed in gene expression analysis. Convolutional neural networks (CNNs) have shown to be very benecial in image analysis applications, such as histopathology image analysis, within the eld of deep learning. Since recurrent neural networks (RNNs) are good at processing sequential data, they are often used for sequence analysis tasks [ 66]. Drug discovery process optimization may benet from reinforcement learning, a subeld of ML in which a method learns how to act in a given environment to maximize some concept of cumulative reward [67].
9.5.3.1 Structural bioinformatics algorithms
Structural bioinformatics focuses on studying and predicting biological macro­moleculesthree-dimensional (3D) structures. Homology modeling, sometimes called comparison modeling is a technique used to predict the 3D structures of
9-17
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
proteins and nucleic acids based on the known structures of similar molecules. It is based on the idea that even when protein sequences differ dramatically, the tertiary structures of the proteins remain constant throughout evolution [68]. There are four primary phases in the process:
1. Template identication is used for searching databases for known structures that share sequence similarity with the target molecule.
2. The target sequence is aligned with the template structures to identify corresponding residues in the sequence alignment phase.
3. In the model-building phase, a 3D model is generated for the target based on the template structures, often using spatial restraints derived from the templates.
4. Finally, the model renement and validation of the model to improve its stereochemical quality and validate it using various structural and statistical checks.
The 3D protein structure homology modeling program MODELLER is frequently used. It creates models by comparing the modeled sequence to similar sequences [69]. Molecular docking aims to determine which orientation of two molecules will result in the most stable combination. Predicting how compounds will react with their intended targets is essential in the drug design process [70]. Necessary procedures in molecular docking include:
1. Preparing the receptor and ligand structures, including adding hydrogen atoms, dening rotatable bonds, and assigning charges.
2. Dening scoring functions to evaluate different binding orientations and conformations.
3. Employing algorithms to explore possible binding modes, including stochas­tic or deterministic search methods.
4. Ranking the predicted binding modes based on scoring functions and analyzing the results to understand the binding interactions.
AutoDock is a suite of automated docking tools designed to predict how small molecules, such as substrates or drug candidates, bind to a receptor of known 3D structure [71]. Ab initio modeling, or de novo modeling, predicts protein structures solely from their amino acid sequences without using any template structure. It is benecial when no suitable template is available for homology modeling [72, 73]. The de novo modeling process often requires:
1. Building protein conformations by assembling short fragments from a library of known structures.
2. Employing computational methods to nd low-energy conformations that are likely close to the native structure.
3. Using statistical mechanics techniques to explore conformational space and nd the global minimum energy conformation.
Rosetta is a software suite for predicting and designing protein structures, folding, and protein –protein interactions from amino acid sequences, using a fragment-based approach when no homologous templates are available [74].
9-18
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Alpha-helices, beta-sheets, and twists are all examples of local protein structures that can be predicted using just the amino acid sequence [75]. The main procedures in secondary structure method are:
1. Statistical techniques use the known structures statistics to predict secondary structure elements.
2. Machine learning techniques like neural networks are used to predict secondary structures based on patterns in training data.
3. The prole-based methods are used in multiple sequence alignments to generate proles and predict secondary structures based on conserved patterns.
PSIPRED is a technique that uses neural networks to accurately predict secondary structure components like alpha-helices and beta-sheets from amino acid sequences [76].
9.5.4 High-performance computing (HPC) in bioinformatics
High-performance computing (HPC) has become vital in bioinformatics, enabling large-scale biological and genomic data processing and analysis [77]. The advent of next-generation sequencing (NGS) technologies and the surge in publicly available biological data have made modern biological experiments both data and computa­tionally intensive [77]. HPC provides the computational power necessary to handle the Big Datachallenges posed by bioinformatics, facilitating the extraction of knowledge from raw data through larger computational platforms known as supercomputers [78]. Parallel computing is a type of computation in which many calculations or processes are carried out simultaneously. Bioinformatics accelerates the analysis of large datasets and complex computational tasks [79]. The HPC in bioinformatics is particularly used in the following areas:
Sequence alignment: Parallel algorithms can drastically reduce the time required to align large or multiple sequences [80].
Phylogenetic analysis: Constructing phylogenetic trees from large datasets is computationally demanding, and parallel computing can signicantly accel­erate this process [81].
Structural modeling: Parallel computing enables faster processing in structural bioinformatics, such as molecular docking and homology modeling.
The background in the high-performance computing area, including parallel computing, opens up signicant opportunities for simulating relevant biological systems and applications in bioinformatics, computational biology, and computa­tional chemistry [82]. A schematic workow of HPC and how it works is shown in gure 9.5.
9.5.5 Cloud computing in genomics
Cloud computing provides a exible and scalable environment for storing, managing, and analyzing a gigantic amount of genomic data. It has emerged as
9-19
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Figure 9.5. HPC framework: data processing from big data storage through hpc nodes to user interface.
a vital resource for bioinformatics and genomics, offering several advantages. Cloud computing reduces the costs associated with data storage and c omputa­tional resources, which is especially benecial for smaller research groups or institutions. It allows for the easy scaling of resources as the data volume grows. Cloud platforms provide easy access to data and computational resources from anywhere, anytime. The technology of cloud computing and its applications in biology and bioinformatics have enormously increased over recent years, provid­ing a platform for the analysis of biological data that would not be possible without signicant computational resources [83]. Resources such as NSF XSEDE, Google Cloud, and Amazon AWS have become more available, and a growing community of academicians are working on teaching the utility of HPC resources in genomics and big data analyses [ 84].
9.5.6 GPGPU (general-purpose computing on graphics processing units)
GPGPU stands for General-Purpose computing on Graphics Processing Units. It represents an innovative approach in which the powerful processing capabilities of graphics processing units (GPUs) are employed for general computing tasks beyond just rendering graphics. In bioinformatics, GPGPU has shown signicant promise in accelerating various computationally intensive tasks. The utilization of GPGPU in bioinformatics is primarily facilitated through libraries such as Nvidias CUDA (Compute Unied Device Architecture). CUDA is a widely used library for developing GPU-based tools in bioinformatics, computational biology, and systems biology. It is specically tailored to use Nvidia GPUs, although alternative solutions like Microsoft DirectCompute also exist. The adoption of GPGPU in bioinfor­matics has been driven by the need to reduce the running time required by standard CPU-based software, allowing for more intensive investigations of biological systems. GPGPU offers computational power comparable to a small computer cluster but at a signicantly lower cost, around $400. This advantage, coupled with
9-20
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
the developing of specic algorithms tailored for GPU computing, has led to a growing interest in GPGPU within the scientic community [85].
A collection of GPU tools has been developed to perform computational analyses in life science disciplines, emphasizing the advantages and drawbacks of these parallel architectures. GPUs can considerably reduce the running time, enabling more efcient data processing and analysis in bioinformatics [86].
9.5.7 Big data analytics in bioinformatics
Big data analytics in bioinformatics refers to applying advanced analytic techniques on extensive and complex biological datasets. The emergence of high-throughput technologies, like NGS, has led to the generation of massive amounts of biological and genomic data. Big data analytics is crucial for extracting meaningful insights from these data, including discovering new knowledge, identifying patterns, and making informed decisions. Several primary approaches are used in data analysis to derive meaning from large datasets. Data mining is a fundamental approach, utilizing algorithms to lter through large datasets and reveal hidden patterns and relationships. ML takes center stage, as complex algorithms are applied to develop predictive models that can estimate future events based on the existing data. Statistical analysis is crucial to drawing meaningful conclusions from data, which uses statistical techniques to analyze the data, identify patterns, and rigorously test hypotheses. To better grasp a dataset, visualization methods are often used to construct visual representations of the data, providing a more transparent picture of the underlying patterns and connections. Several tools and frameworks have been built expressly for big data analytics in bioinformatics to manage the absolute volume and variety of biological data. Accelerating research in the biological sciences, these tools improve data processing, visualization, and interpretation [87].
9.5.8 Systems biology modelling
Systems biology modeling uses computer models to assess biological data and predict system behavior to comprehend the complicated relationships within bio­logical systems. Network analysis and metabolic pathway analysis are essential analytical techniques in system biology. In systems biology, network analysis often involves the examination of biological networks to clarify the intricate relationships between diverse biological entities, including proteins, metabolites, and genes. This method helps to comprehend the fundamental design and operation of biological systems. Building and analyzing networks using molecular data is known as molecular networking. Vital biological insights can be made by studying the relationships between molecules, such as nding crucial regulatory molecules or comprehending the processes behind the disease. Networks are inferred from experimental data using a variety of instruments and techniques, and these networks are then analyzed to determine their topological characteristics, dynamics, and functional consequences [88]. Metabolic pathway analysis includes the study of metabolic pathways to understand the movement of materials and information inside biological systems. Flux-based analysis is carried out to learn more about how
9-21
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
metabolic pathways are controlled and how they contribute to cellular function, which entails the use of computer models to evaluate the ow of metabolites via metabolic networks [89]. To offer a thorough knowledge of biological systems, systems biology analysis is a more complete approach that includes metabolic pathway analysis, among other techniques. It analyzes and interprets the behavior of complex biological systems, such as metabolic networks, using computer models and experimental data [88].
9.5.9 Systems pharmacology
An interdisciplinary eld called systems pharmacology combines computational and experimental methods to study how medications interact with biological networks and disease states. The goal of system pharmacology is to create mechanistic, quantitative knowledge that can estimate medication reactions and direct the development of novel treatment approaches. Quantitative systems pharmacology (QSP) is a mechanistic platform in pharmacology that models the interplay between medications, biological networks, and disease states to foretell optimal therapy responses. Small molecules, nucleic acids, proteins, pathways, cells, organs, and disease processes are all explored regarding medication interactions. Step-by-step enhancement of the modeling process to make it more effective and repeatable is essential to creating and validating QSP models. These models are exible and ever­evolving, so they can easily absorb newly available data and insights [90]. Multiscale modeling in systems pharmacology, especially in disease situations, links the cellular responses to protein or drug interactions in the context of animal or human physiology. When it comes to converting preclinical scientic discoveries into useful clinical applications, these models are essential [91]. In addition, discovering combination therapy treatments, particularly those targeting the immune system, might benet from multiscale systems pharmacology modeling techniques [92].
9.5.10 Multiscale modeling
The goal of multiscale modeling is to combine data from different scales and different physics in order to determine the processes behind the formation of function in biological systems. It includes analyses of biological and medical phenomena at several levels of organization, from the molecular to the organismal. ML methods are rapidly being included in multiscale modeling because of the valuable tools they provide for the robust management of challenging issues and the handling of sparse and noisy data. When multiscale modeling is combined with ML, the resulting models can better capture the nuanced details of biological systems [93]. Recently, ML has found its way into the multiscale modeling of hierarchical engineering materials and the solution of high-dimensional partial differential equations, unveiling its interdisciplinary relevance across various elds [94]. Constructing computational and mathematical models based on experimental data is a common approach to comprehending and predicting the behavior of complex biological systems. Simple modeling techniques like differential equations or Boolean networks benet from incorporating multiscale modeling, enabling a
9-22