Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5864_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
35 Мб
Скачать
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
alignment-free trees based on statistical properties of the genome;
gene content trees based on the presence and absence of genes;
phylogenomics-based genome trees;
trees based on average sequence similarity; and
trees based on chromosomal gene order.
Despite their current advancement, genome tree approaches have already had some inuence on the phylogenetic arrangement of bacterial species. However, their main impact so far has been on our understanding of the nature of genome evolution and the role of horizontal gene transfer therein [57]. Relative genome analyses disclose that most functional domains of human genes have homologs in commonly divergent species. These shared functional domains, however, are differentially shufed among evolutionary lineages to create an increasing number of domain architectures. Combined with duplication and adaptive evolution, domain shufing is accountable for the great phenotypic complexity of higher eukaryotes [8385]. Mobile elements within genomes have determined genome evolution by different means. Mainly in plants and mammals, retrotransposons have been shown to establish a large fraction of the genome and have shaped both genes and the entire genome. Although the host can frequently govern their numbers, massive expan­sions of retrotransposons have been accepted during evolution. Currently, mobile elements are becoming valuable tools for learning more about genome evolution and gene function.
Based on paleological investigation, it was demonstrated that approximately
3.5 billion years ago, cells identical to bacteria existed on our planet. Based on this observation we can only accept that the genomes of these organisms or the organisms that evolved soon afterward contain double-stranded DNA molecules. According to previous ndings (dated to about 1.4 billion years ago) the rst eukaryotic fossils were similar to single-celled algae, which means that both prokaryotes and eukaryotes experienced variations in size, shape and complexity, thus their genomes also transformed vigorously. These variations were motivated by mutation, transposition, gene transfer and recombination, along with gene deletion and duplication. Through investigating and equating genomes of organisms that exist currently, we can gain understandings into how these mechanisms have shaped genomes.
7.8.1 Microbial genome reduction in bacteria
Once bacterial pathogens make the evolution from free-living life cycles to a permanent relationship with a host, they experience a signicant loss of genes and DNA [86]. Complete genome sequences offer information on how signicant genome reduction affects the developmental directions and metabolic capabilities of obligate pathogens and symbionts [58]. Reconstruction of the gene content and order of the last common ancestor of human pathogens can be achieved, for example in the case of Mycobacterium leprae and Mycobacterium tuberculosis .An understanding into which genes are the most indispensable for existence comes from
7-28
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
equating the genomes of two related species of Mycobacterium: M. tuberculosis, which is responsible for tuberculosis and M. leprae, which is the causative agent of leprosy. As per reports, M. tuberculosis has 3959 protein-coding genes and a
4.41 Mb genome. In contrast, M. leprae has 1604 protein-coding genes and a
3.26 Mb genome. Therefore, during development or evolution, M. leprae has misplaced or lost approximately 2000 genes (more than 50% of its ancestral gene set). Several reasons can account for this, such as codon elimination or deletion and deterioration, through mutation and incorporation of transposable elements, which have disabled many genes. Consequently, metabolic functions vital for cell develop­ment have been abolished, and the bacterium develops very slowly; it multiplies once every 14 days. Moreover, genetic mutations for recombination and DNA repair are incapable of averting additional loss to the M. leprae genome. This blend of mutations may have comdemmed M. leprae to extinction, as epidemiologists have projected that the organism is at the limit of being able to sustain itself by infecting new individuals [59].
7.8.2 Role of duplications in the origin and evolution of the eukaryotic genome
Assessments of prokaryotic and eukaryotic genomes are providing indications of the lineages of the eukaryotic genome and by what means eukaryotes arose from prokaryotes. Several lines of inquiry, comprising amino acid analysis of proteins, gene sequences and metabolic pathways, determined that the eukaryotic genomes are actually mosaics that have notable contributions from both the Archaea and the Eubacteria. For example, the eukaryotic nuclear genome has few, if any, operons, and its genes encompass introns, both of which are features of the Archaea.
The mitochondrial genome of a eukaryotic looks very much like that of alpha­proteobacteria. To explain these observations, it has been suggested that eukaryotes arose as a result of a symbiotic relationship, or fusion, between an anaerobic archaebacterial host and an alpha proteobacterium (such as Rickettsia), which developed into the mitochondrion [90]. The monophyletic character of all eukar­yotes suggests that this incident occurred effectively only once in the history of the Earth.
The duplication–divergence concept was suggested many decades ago, i.e., that new genes evolve from pre-existing ones via gene duplication and subsequent divergence of the extra copy to acquire a new function (gure 7.12). Based on this concept Susumu Ohno wrote the book Evolution by Gene Duplication. Since then, comparative genomics, genetics and biochemistry have clearly demonstrated that duplication–divergence mechanisms are a key contributor to the evolution of new genes. In contrast, the de novo origination model of new gene evolution is more recent and less supported. De novo origination of protein-coding genes takes place when genes arise from a nonfunctional DNA sequence that was formerly not a gene. It would seem highly questionable whether functional proteins could arise sponta­neously from non-coding DNA, as the DNA must both be transcriptionally active and comprise a translatable open reading frame. Also, translation of any random ORF devoid of genes is anticipated to offer unimportant polypeptides rather than
7-29
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Figure 7.12. Role of duplications in the origin and evolution of the eukaryotic genome.
proteins with specic functions. Certainly it has been claimed that de novo origination of new genes is tremendously unlikely, however, irrespective of these reports, the arrival of large-scale sequencing and comparative genomics has delivered increasing proof that new genes have developed and continue to originate from non-coding sequences [60].
Gene duplication plays a signicant role in the development of eukaryotic genomes, particularly their size and complexity. Susumo Ohno has predicted that complete genome duplications are a powerful development mechanism [91]. Examination of nucleotide sequence information from genome projects encourages the impression that a considerable portion of the difference in gene number that differentiates prokaryotes from eukaryotes arose from genome expansions. Currently, it is understood that a signicant expansion in eukaryote genome size arose from a genome duplication event that accompanied the arrival of vertebrates in the fossil record. Genome duplications have also occurred at other times during eukaryotic development. A current investigation of the yeast genome displays traces of its earliest development by genomic duplication. The yeast genome has been determined to hold at least 55 duplicated regions encompassing 376 genes, which covers 5% of the genome. Of these 55 regions, 50 are in a similar comparative location on diverse chromosomes. For example, chromosomes XI and XIII comprise duplicated blocks of genes in which gene order and positioning relative to the centromere has been well-maintained. Additional chromosomes cover internal duplications, such as the region positioned on either side of the centromere on chromosome XII. Phylogenetic examination designates that this genome duplication incident took place around 100 million years ago. Examination of the human genome also displays proof of an early large-scale duplication, tracked by reorgan­izations and gene damage in certain duplicated regions. These duplications range
7-30
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Figure 7.13. Functional genomics investigates the coordination among the genome, transcripts (genes), proteins, and metabolites to generate specic phenotypes.
from minor sections covering only few genes to great stretches that cover nearly a complete chromosome. The larger duplications date to the origin of the vertebrates, around 500 million years ago. Overall, there are 1077 blocks of duplicated regions in the human genome, covering 10 000 genes (around one-third of the genome). Such duplication includes chromosomes 18 and 20. Different potential areas covered under proteomics are presented in gure 7.13.
7.8.3 Gene duplications increase genetic diversity and complexity
Genome duplication is both a historic and continuing process in yeast, plants and animals. It was recently established that the yeast genus Saccharomyces experienced an ancient whole-genome duplication, and genome duplication is pervasive in plants. Numerous incidents of genome duplication have occurred during the divergence of angiosperms, including ancient polyploidization processes and the assembly of self-governing duplications in different ancestries. Although not as common as in plants, many genome duplications have also happened in independent lineages in metazoans. During chordate development, whole-genome duplications have occurred coincident with the origin of vertebrates gnathostomes, and teleosts.
Gene and genome duplications offer a basis of genetic material for mutation, drift and selection to act upon, making new evolutionary changes more likely. Thus, several researchers have claimed that genome duplication is a dominant factor in the
7-31
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
evolution of complexity and diversity. However, a strong connection between a genome duplication incident and increased complexity and diversity is not con­clusive, and there are variations in the patterns of diversity raised to support this claim. Remarkably, many studies of genome duplication processes in vertebrates show they are preceeded by numerous extinct ancestries, resulting in preduplication gaps in extant taxa.
Genomic nucleotide sequence information indicates that multigene families are found in the human genome, if not all genomes. In addition to genome-wide duplications, small chunks of genes and single genes can be duplicated by numerous mechanisms, including:
unequal crossing over and
replication errors.
Throughout replication of a template molecule, a slippage can cause the introduc­tion of a short segment into the freshly produced strand. When produced, participants of multigene families may remain associated on a single chromosome or can scatter to other parts of the genome. Numerous mechanisms determine this procedure, such as inversions, translocations and transposition by mobile elements.
The ancestry and relationships among members of gene families have been recently established by several groups of molecular phylogeneticists. Among the most studied examples is the globin gene superfamily. Some 800 million years ago, duplication of a family gene encoding an oxygen transport protein occurred in this family. This has formed two sister genes, one of which developed into the modern day myoglobin gene. Myoglobin is an oxygen-carrying muscular protein, i.e., it is found in muscles. Around 500 million years ago, the ancestral globin gene duplicated to produce the prototypes of the α- and β-globin subfamilies. The α­and β-globin genes encode for the proteins that are present in hemoglobin, the oxygen-carrying molecule in red blood cells. Further duplications within the α- and β-globin genes occurred in the last 200 million years. Subsequent events dispersed members of this superfamily, and each is now on a separate chromosome. Similar patterns of development are reported in other gene families, including the trypsin­chymotrypsin family of proteases, the homeotic selector genes of animals and the rhodopsin family of visual pigments.

7.9 Algae bioinformatics

An alga is considered a unique source of various bioactive compounds and having signicant biological activities. Algae bioinformatics, as the name suggests is the application of information technology to decipher more algae with the aid of computational tools and software. This emerging eld, just like the other streams of bioinformatics employs computational or dry-lab techniques for applications in various areas such as gene prediction, comparative genomics, genome analysis, and functional genomics and so on. Algae bioinformatics cannot be performed without computer science, biology and genetics with a good-sized dollop of mathematics, statistics and other medical specialties thrown into the mix. There are many tools
7-32
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
available for algae bioinformatics research, which are briey discussed at the very end of this article.
7.9.1 Scope of algae bioinformatics
Algae bioinformatics is a valuable resource for geneticists, phycologists and other scientists who sequence algal genomes as a part of their wet-lab research. Therefore, there is an immense requirement for more reliable, sophisticated, computerized methods for examining this information bringing in the requirement of an algae bioinformaticist.
7.9.2 What is involved in algae bioinformatics
Algae bioinformatics involves the development of novel algorithms and statistics with which relationships among the different algal species can be measured. Examination and elucidation of different types of data and their related information is done. This data study involves DNA, RNA and protein sequences and structures. Development of various techniques allows the efcient access and management of different types of algal samples.
7.9.3 Role of algae bioinformatics
As discussed earlier, algae bioinformatics will only involve dry-lab research and it requires the use of RNA, DNA, protein sequence data. A series of wet-lab work is done so as to obtain the sequencing data and is subjected to bioinformatics analysis.
7.9.4 Steps involved in obtaining the data for analysis using bioinformatics
Design of primers;
Extraction of DNA, RNA;
PCR amplication;
Denaturizing gradient gel electrophoresis;
Sequencing.
Nucleic acids are usually derived from the algal samples which are then utilized to synthesize and design primers and then PCR amplication is done. As the products of the PCR reactions are equivalent in their size, a denaturing gradient gel electrophoresis is done to identify variation between the sequences of the products. After this step algal bioinformatics work is performed. This can be achieved by making phylogenetic trees, BLAST searches for similar sequences, and annotation of the new sequences.

7.10 Functional genomics

Functioning genomics is an academic discipline that explores the relationship among a living things genetic composition, often termed the genome, and its discernible characteristics, known as an organisms phenotype. Recognizing the genetic processes that determine the operational dynamics and interaction of genes inside
7-33
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
an individual is the fundamental goal. The study of gene expression, gene function, and research into the gene structure are all methods included in genomics with function [61]. An overview of functional genomics is presented in gure 7.13.
7.10.1 Introduction to functional genomics
The branch of research known as operational genomics seeks to understand the complex relationship that links the genome and its phenotype on a broad scale that encompasses the whole genome. The strategies used in this work include utilizing high throughput techniques to investigate genesactivation and interacting patterns. This facilitates the examination and comprehension of the actions performed by these genes and the cellular processes that govern their roles inside an organism.
7.10.2 Transcriptomics: studying the RNA molecules
Transcriptomics relates to the science of the eld as it examines a transcriptome, which covers the totality of RNA molecules created by the genetic code during a specic time. The procedure involves the examination of the expression of gene designs, the recognition of transcription factors, and the exploration of functional elements inside the genome. The sequencing of RNA is the technique that is often used in the transcriptomics study. It gives us helpful information about how genes are expressed and how the genome is structured [62].
7.10.3 Proteomics: understanding the world of proteins
Proteomics is an extensive examination of peptides on a broad scale, focusing mainly on elucidating their functions and structural traits. The eld of research being discussed has signicance in understanding changes in metabolism under different conditions and assumes a pivotal role in the early detection of illnesses, forecasting their results, designing medications, and monitoring disease progression. Proteomics methods allow for in-depth analysis and characterization of its proteome at various stages of development and physiology by measuring protein expression, structure, function, interaction between molecules, and changes following trans­lation [63].
7.10.4 Metabolomics: exploring cellular metabolites
Metabolomics is a contemporary eld of study within the omicssciences, focusing on the comprehensive analysis of metabolites, including qualitative and quantitative evaluations. These metabolites encompass essential intermediates and nal products that arise from metabolic processes. The integration of high-throughput analytical techniques and bioinformatics enables the systematic identication and quantica­tion of several metabolites, hence facilitating the comprehension of metabolic disruptions and the elucidation of underlying mechanisms associated with diverse diseases [64].
7-34
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
7.10.5 Interactomics investigating protein–protein interactions
Interactomics, as a sub-eld within the discipline of systems biology, is primarily concerned with investigating protein–protein interactions and their consequential impact on phenotypic traits. Various methodologies are utilized to investigate and analyze the interactome, facilitating the comprehension of protein functionality and regulation. The utilization of interactomics encompasses biochemical and clinical domains, signicantly contributing to understanding cellular signaling, disease pathogenesis and advancing therapeutic approaches [65].

7.11 Structural genomics

The primary objective of structural genomics is to get comprehensive insights into the spatial arrangement of all proteins encoded by genetic material. This discipline extends beyond conventional molecular genetics and biochemistry. To offer a more comprehensive comprehension of the genome. The main aims of this study are comprehensive genetic and physical mapping and sequencing of the entire genome. A crucial objective is elucidating the experimental structures of all potential protein folds. Structural genomics holds promise in offering a comprehensive comprehen­sion of several facets of cellular existence, encompassing metabolic activities, DNA replication, transcriptional processes, protein synthesis, and protein folding [66]. The ow of structural genomics is represented in gure 7.14.
Figure 7.14. Structural genomics, facilitated by the National Institutes of Health (NIH) Protein Structure Initiative (PSI), established a collaborative network of research and resource centers. The primary objective was to comprehensively explore the structural aspects of proteins, particularly in the uncharted territories of the protein sequence space. A key challenge addressed by the PSI was elucidating the intricate connections between protein sequences and their respective structures. The ultimate aim was to enhance accessibility to structural data for the majority of proteins based on their gene sequences.
7-35
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
7.11.1 Introduction to structural genomics
Structural genetics may be seen as an extension of genomics that encompasses the domain that includes structural inquiry, aiming to uncover the features of genome architecture. High-throughput methods are used to determine the structural features of proteins across the whole genome. The insights gained by structural genomics are valuable in many fields, including medication creation and a molecular-level under- standing of illnesses, since they allow for modifying genes and DNA segments. One of the main goals is to assign functional annotations to proteins whose activities have not yet been established, mainly by analyzing their structural properties [67]. The aim may be achieved using several methodologies, such as inferring distant homology relation­ships, identifying ligands, and characterizing electrostatic bands and cavities among a protein’s structures. Genomic structural variants (SVs) are also studied in this eld; chromosomal rearrangements affect at least 50 bp. Understanding structural varia­tions is of the highest relevance owing to their ability to give valuable insights into mechanisms of evolution and the deep molecular foundations of illnesses such as varied forms of carcinoma and neurodevelopmental conditions [68].
7.11.2 The approaches used in the domain of structural genomics
Deciphering a three-dimensional (3D) structure that composes biological macro­molecules, having a signicant focus on certain peptide regions, is the goal of structural genomics. New methods have been created to nd and analyze these combinations, which helps us learn more about the molecular and functional settings where these giant molecules work. The development of protein structure determination pipelines is mainly attributable to structural genomics, which has led to significant advances in identifying hitherto uncharacterized proteins. Notable advancements in protein production and the rapid identification of new protein structures have resulted from investigating proteins obtained from microbial dark matter and human diseases [68]. Crystallography using x-rays, NMR (nuclear magnetic resonance) spectral analysis, and cryo-electron microscopy (cryo-EM) are just a few examples of the many methods included in structural genomics and are essential to the eldsoverall goals and objectives. These techniques are often supplemented by other method­ologies, including connecting mass spectroscopy, small angle scattering of x-rays (SAXS), and neutron diffraction. This collective approach offers a complete means of investigating biological macromolecules. The challenges encountered within structural genomics are manifold. The large volumes of data generated necessitate effective collection, display, and analysis protocols to maximize the value derived from such substantial investments [69]. Moreover, the quality of structures determined is a concern, particularly when resolution data are articially limited, which might undermine the benets of structural genomics endeavors.
7.11.3 Importance of structural genomics in drug design
Structural genomics holds signicant promise for drug discovery. The growing repository of high-resolution structures of known and potential drug target proteins
7-36
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
is poised to benet future drug discovery programs substantially. When the target protein of a drug is known, structural genomics facilitates the rapid and efcient acquisition of ligand-bound structures, leveraging high-throughput x-ray crystallog­raphy and NMR [70]. In cases where the drug target is unknown, structural genomics provides puried proteins for numerous potential drug targets, thus supporting drug discovery efforts. The impact is already being felt in the early stages of drug discovery and target validation, with the contribution of new structures, complexes with ligands, and supportive protocols and reagents for additional structural work within drug discovery programs. The role of structural genomics in modern structure-based drug design is underscored by its ability to analyze many target proteins concurrently. While several structural genomics initiatives have been launched, a relative few have focused on integral membrane proteins, which represent a critical area for future exploration in drug design and discovery [71].

7.12 Epigenomics and epigenetics

Epigenomics is a sub-eld of genomics that explores the comprehensive character­ization and analysis of all epigenetic modications across the genome. Epigenetics is the study of reversible, heritable changes in gene function that occur without altering the DNA sequence itself. DNA methylation is a hallmark epigenetic regulation mechanism [72]. Epigenomics is concerned with understanding the broader epigenetic landscape across the genome, focusing on modications like DNA methylation, histone modications, and chromatin remodeling that do not alter the DNA sequence but signicantly impact gene expression, cellular function and phenotype. The conventional concept of epigenetics, articulated initially by Conrad Waddington during the 1950s, pertains to the investigation of enduringly inheritable phenotypic traits that arise from modications to a chromosome without concomitant alterations to the DNA sequence [73]. Epigenetic alterations play a critical role in various biological processes, such as development, differentiation, and the ability to respond to environmental stimuli. The rapid acceleration of epigenetic research in the 21st century has generated excitement and hope, linking genetics to environmental factors and diseases [74]. DNA methylation, a pivotal epigenetic mechanism, involves adding a methyl group to the cytosine base of DNA, predominantly at CpG dinucleotides. DNA methylation patterns are established early in mammalian development and are maintained during somatic cell division, playing crucial roles in gene silencing, x-chromosome inactivation, genomic imprinting, and suppressing transposable ele­ment activity. These dynamic patterns can be altered in response to environmental stimuli and during different developmental stages. For instance, the review highlights the dynamic erasure and re-establishment of DNA methylation in embryonic, germ­line, and somatic cell development, emphasizing its signicance in mice and humans. DNA methylation and demethylation processes also occur in the nervous system, associating with other epigenetic mechanisms like histone modications and non­coding RNAs, showcasing the interplay between different epigenetic modications [75]. The structure and distribution of CpG islands and the role of methylation in gene expression regulation, embryogenesis, ageing, and cancer development are also critical
7-37