Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5344_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
35 Мб
Скачать
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
storage, examination, analysis and utilization of the information about biological systems. For example, it involves processes such as accumulating genome sequences, documentation of genes, assigning functions to the identied genes, planning of databases, etc. Different genetic compilers are available today. A software compiler is a way to use genetic information and provides a design platform allowing the geneticist to manipulate and design everything from single genes to whole genomes. So as to guarantee that the nucleotide sequence of a genetic material is complete and error-free, the genetic material is sequenced more than once, such as by using the shotgun method. The bacterial genome (Pseudornonas aeruginosa) was sequenced seven times to make the sequence accurate and free from faults. Stover and co­workers (2010) reported the complete genome sequence of P. aeruginosa PAO1. They suggest that the size and complexity of the P. aeruginosa genome reveals an evolutionary adaptation allowing it to ourish in various environments and limit the effects of a diversity of antimicrobial compounds [12]. However, the assembler software (a sequence assembler is innovative bioinformatics software for sponta­neous DNA sequence assembly, contig editing, le format alteration, DNA sequence analysis and mutation recognition) identied 1604 regions that need further justication. These areas were reinvestigated and re-sequenced to success­fully complete the genome sequence. The efciency/accuracy of the shotgun technique was equated with the sequence derived from the clone-by-clone procedure of two extensively separated genomic regions of P. aeruginosa. These two considered areas together were 81 843 nucleotides long. The sequences derived by the two procedures were in perfect agreement. This assessment exposed the precision potential of the shotgun method of genome sequencing. This also demonstrates the safeguards to be considered during genome sequencing projects. This level of maintenance is not uncommon. Similar safety measures are employed in all genome projects. The Human Genome Project, an international scientic research project, sequenced the 3.2 billion bp of the human genome a total of 12 times. Celera Genomics (a private biotech company), using shotgun-cloning, used an approach of sequencing from both ends of DNA fragments. This project sequenced the human genome 35.6 times. While a rough draft of the human genome sequence is complete, some other objectives are yet to be accomplished. These include obtaining the remaining sequence and rectifying miscalculations (called proofreading the genome), lling in gaps and then sequencing the 7%–10% of the genome that comprise heterochromatin. This is present in ample amounts inside the cells that are less active or not active, whereas euchromatin is predominant inside cells that are vigorous or active in the transcription of several of their genes. Since they comprise low priority sections of repetitive DNA sequences, heterochromatic regions of the genome were not initially sequenced. Moreover, intially it was thought that the heterochromatin does not contain genes. However, the Drosophila genome sequence revealed that their heterochromatic regions contain a very small number of genes (around 50). Based on this discovery, the heterochromatic regions of the human genome must be sequenced to conrm that all genes in the human genome have been recognized. As soon as the genome of an organism is sequenced, compiled and proofread, the subsequent process of genomics, namely annotation, begins.
7-8
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)

7.5 Understanding bioinformatics and sequencing

Advance computers and high-speed networks are compulsory to evaluate genome projects. Genetic data signify a potential source for investigators and organizations interested in how genes contribute to our health and wellbeing [11]. Nearly half of the genes identied by the Human Genome Project have no known function. Scientists, by means of bioinformatics, can identify genes, establish their roles and develop gene-based strategies for preventing, diagnosing and treating disease [12]. Simplied presentation of bioinformatics and sequencing is presented in gure 7.3.
Genome sequencing and investigation is an area that has developed very quickly over the last 10–20 years, in particular after publication of the nal draft of the human genome sequence in 2003. From procuring a complete genome sequence from a characteristic individual or strain of a small number of species, the genomics community has moved on to recording and authenticating genetic assortment within species, with a special focus on human beings, and by sequencing the genomes of a quickly growing number of species. The introduction of economical, very high throughput sequencing tools has also made it promising to sample transcriptomes (the summation of all mRNA fragments expressed from the genes of an organism), genomic areas bound by proteins, bacterial populations, or sometimes even whole ecosystems, by means of sequencing methods. This has germinated an innovative generation of software techniques to manage the very large numbers of short sequences, often called reads, which are produced by the new machines, and has also moved the area of computational genomics into the territory of big data, requiring
Figure 7.3. Bioinformatics and sequencing.
7-9
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
large and sophisticated computers for the organization and investigation of these sequences.
The entities that are selected for genome projects are those typically used in genetic and other systematic studies. Thus they are usually called model organisms. A model organism can be a prokaryotic or eukaryotic micro-organism, or an animal or plant, which is extensively investigated as it can be easily studied in the laboratory to understand biological processes. The Generic Model Organism Database project offers genetic researchers tools of open-source software machinery for imaging, interpreting, managing and procuring biological information. Model organisms include organisms such as E. coli, Archaeoglobus fulgidus, yeast (Saccharorny cerevisiae), Bacillus subtilis, A. thaliana, the fruit y(Droso melanogaster) and nematide worm (C. elegans). The Human Genome Project focused on sequencing the whole human genome.
E. coli is, by far, the most extensively studied micro-organism. Various tools that are available currently were established for the E. coli genome project. The sequencing of the E. coli genome was accomplished in 1997 [13]. The E. coli K-12 genome contains 4408 genes (4288 protein-coding genes annotated) with a size of around 4.64 × 10
6
bp. B. subtislis is a gram-positive bacterium that settles over leaf surfaces and is signicant for both production of enzyme and food supply fermentation. GRASis an acronym for the phrase Generally Recognized as Safe. GRAS is well suited for enzymes, given the general availability of scientic data supporting enzyme safety, and the generally recognized (peer-reviewed) methodology and decision trees for evaluating the safety of microbial enzymes used in food processing and in animal feed, respectively. It is usually named as safe (GRAS) with genome of 4.21 × 10
6
bp (4212 genes). The reference database SubtiList was exclusively created for the genome of B. subtilis 168, the paradigm of gram-positive endospore-forming bacteria [14]. Another example of a model organism is Aracheoglobus fulgidus, a strictly anaerobic archaebacterium with a genome size of 2.17 × 106 bp (2493 genes) [15] (1 563 423 bp, 1858 protein-coding, 52 RNA genes, according to the Genomic Encyclopedia of Bacteria and Archaea project) [16]. The yeast organism S. cerevisiae, so called because of its ability to ferment saccharose (sugar), is the most key fungal species used in biotechnological developments. In 1996 the S. cerevisiae genome was the rst completely sequenced from a eukaryote. The current version, called S288C 2010,was determined from a single yeast colony by means of advance sequencing tools and functions as the anchor for further innovations in yeast genomic science [16]. In 1989, genome sequencing of S. cerevisiae was accomplished. The yeast genome is 12.8 × 10
6
bp in size and contains 6548 genes. PlantGDB (http://www.plantgdb.org/) is a databank of molecular sequence data for all plant species with important sequencing efforts. The database arranges EST sequences into contigs that represent tentative unique genes. PlantGDB offers genome browsing abilities that integrate all accessible EST and cDNA reported for existing gene models (for A. thaliana, see the AtGDB site at
http://www.plantgdb.org/AtGDB/)[16].
The rst plant for which the complete genome was sequenced and published was
A. thaliana. In 1907, it was rst considered as a good organism for genetic
7-10
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
investigations and has become a model plant for genetic investigations as it has the smallest genome of the higher plants. Furthermore, its genome contains only a small amount of repetitive DNA [17]. The genomic sequencing project of the A. thaliana genome began in 1990 and was completed in 2000. According to the sequencing results, the size of the A. thaliana rabidopsis genome is 130 × 10
6
bp and a predicted 26 000 genes [18]. The MIPS A. thaliana Database (MAtDB; http://mips.gsf.de/proj/
thal/db) started out as a repository for genome sequence data in the European
Scientists Sequencing Arabidopsis (ESSA) project and the Arabidopsis Genome Initiative. Another model organism, C. elegans (a free-living nematode) was completely sequenced in 1999 and was the rst multicellular animal whose genome was sequenced. The C. elegans genome has 97 × 10
6
bp and has over 20 000 estimated genes. WormBase (http://www.wormbase.org) is a web-based resource for the C. elegans genome and its biology. It is based on the existing ACeDB databank of the C. elegans genome and offers data curation services, and a considerably extended scope [19].
D. melanogaster (the fruit y) is often called the Queen of Genetics. In 2000, the
Drosophila genome was completely sequenced. It has 180 × 10
6
bp (16 000 genes). This breakthrough has offered a platform for further genomic research, as the number of genes present in the Drosophila genome is less than four times that of the bacterium E. coli. Since the human genome draft sequence was published in 2001, a number of observations have been made. Several signicant features of the human genome are as follows:
A minimum 50% of the genome is obtained from transposable elements.
The genome sequences of individuals vary by less than 0.2% of their base
pairs.
It has approximately 35 000 genes.
It encompasses over 3.2 billion base pairs.
Only 5% of the genome encodes proteins.
The genome contains gene-rich regions separated by gene poor regions, often
called gene deserts.
The largest gene is the gene encoding dystrophin (2.5 × 10
6
bp).
The majority of the variations in the human genome are present in the form of single-base differences in the sequence. A single-base difference is known as a single­nucleotide polymorphism (SNP), frequently called snips. One SNP is present in roughly every 1000 bp of human genome. Approximately 85% of all variations in human DNA is due to SNPs. Advantages of genome projects include [20]:
Better knowledge of human genetic diseases and their correlations should allow their management/cure.
A number of methods have been established for genome sequencing projects.
Information related to SNPs has become accessible; this may be benecial in
different ways.
The complete genetic information existing in the genomes of various organisms can be determined.
7-11
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
They have unlocked exciting opportunities for future research, e.g. functional genomics.
They offer knowledge on why the individuals react differently to the same drugs (pharmacogenomics).
They offer understanding of genome association and its evolution, and the most important mechanisms involved within that.
Genome sequences allow the study of the many molecular interactions resulting in the normal development of organisms.
Micro-organism pathogenicity can be better understood. This can offer protection from such diseases.
The associations between genes can be studied with condence.

7.6 Comparative genomics as a technique to understand evolution

The whole-genome sequence of an individual can be considered to be their nal genetic map, in the sense that the heritable physical characteristics are encoded within the DNA and that the order of all the nucleotides along each chromosome is known. However, information on the DNA sequence does not tell us directly how these genetic data result in the observable traits and behaviors (phenotypes) that we want to understand. Identifying all the functional parts of genome sequences and using these data to improve the wellbeing of individuals and society are the focus of the next phase of the Human Genome Project [21]. Relative studies of genome sequences will be a major part of this effort. The investigation of variation and similarity in genomic structures and, in particular, their organization in various organisms is known as comparative genomics. The aims of comparative genomics are as follows:
to recognize the route of evolution and
to translate the DNA sequence data into proteins of known functions.
In order to equate the genomes of various organisms it is essential to understand the meaning of orthologues and paralogues. Orthologues are homologous genes present in different organisms and only encode proteins that have a similar function. Orthologues arise by direct vertical descent, and have deviated only by storing mutations. Comparatively, paralogues are homologous genes within similar organ­isms and always encode proteins that have interrelated but not similar functions. Paralogues have developed by gene duplication, followed by mutation accumula­tion, e.g. genes of the globin family.
7.6.1 The role of exon shufing
Exon shufing is a molecular tool for the development of new genes. It is a method by which two or more exons from diverse genes can be taken together ectopically (atypical gene expression in a cell, tissue, or at any developmental stage in which the gene is not typically expressed), or the same exon can be duplicated, to create a new exon–intron structure. Exon shufing has been considered as one of the main development powers determining both the genome and the proteome of eukaryotes.
7-12
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Figure 7.4. Exon shufing. In this illustration the end gene is produced by shufing exons from more than one gene.
This is vital in the formation of multidomain proteins throughout animal evolution, obtaining a number of functional genetic novelties [22]. A simple presentation of exon shufing is depicted in gure 7.4.
Various proteins contain distinct domains. Such proteins are known as mosaic proteins, e.g. serine proteases responsible for blood coagulation. The majority of mosaic proteins are extracellular and are mainly present in metazoa. In addition, these types of proteins are also present in unicellular organisms.
The study of the genes that encode mosaic proteins explores the strong association between domain establishment and intron–exon structure. Each domain is inclined to be encoded by one or a combination of exons, which are synthesized by recombination within the intervening sequences; this is known as exon shufing. This process synthesizes new genes that encode proteins with improved function. As introns are much longer than exons, the chances of crossovers in introns are greater than in exons.
7.6.2 Horizontal or lateral gene transfer
Horizontal gene transfer was rst reported in 1928, in an experiment by Frederick Grifth. During this experiment, Grifth demonstrated that virulence was able to transfer from virulent to nonvirulent strains of Streptococcus pneumoniae, establish­ing that genetic information can be horizontally transferred between bacteria via a mechanism called transformation [23]. A schematic representation of horizontal and vertical gene transfer is presented in gure 7.5.
Horizontal or lateral gene transfer involves genetic transfer/exchange between organisms of different evolutionary origin/lineages (sequences of cells, genes, organisms or populations, linked by a continuous line of origin from ancestor to descendent) [24]. It is usually assumed that such gene transfers/exchanges have happened many times throughout the course of evolution. The following two approaches are employed for the discovery of those genes that are likely to develop by horizontal gene transfer:
7-13
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Figure 7.5. Horizontal and vertical gene transfer representation.
Detection of genes having unusual base composition and bias in codon usage can also be used for the detection.
Failure to nd a similar gene in closely related species.
The accessibility of entire genome sequences has allowed the detection of such genes. It has been determined that hundreds (between 113 and 223) of human genes have been transferred directly from bacteria into the human genome rather than having developed from bacterial genes.
7.6.3 Genome similarity or homology
The sequence similarity searching tool is the most extensively used, and most reliable, approach for describing newly determined sequences. This tool can identify homologous proteins or genes by identifying a statistically signicant similarity that reveals common ancestry [25].
Extensively used similarity searching databases are:
BLAST [27].
PSI-BLAST [27].
SSEARCH [30].
FASTA [28].
HMMER3 [29].
These programs offer precise statistical evaluations, guaranteeing protein sequences that share considerable similarity also have similar structures. One of the amazing discoveries from the investigation of genome sequences of various organisms is as follows: organisms whose genomes are quite similar in appearance may look very different, for instance humans and mice.
7-14
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
These two organisms in fact have 97.5% of their DNA sequences in common that achieve the same genetic functions. It was also observed that humans and mice had a mutual ancestor some 100 million years ago. Since then, functional DNA has deviated only to a small extent, whereas the non-coding DNAs of the two species have deviated to a much greater extent. Likewise, it is estimated that human and chimpanzee genomes differ by only 1%–3% of their DNA sequence. It may be noticed that the organisms above are more closely related to each other than organisms such as yeast (a fungus) and the nematode worm (an animal). The nematode worm is estimated to have 19 000 genes, of which over 2000 encode proteins that have functional equivalents as in yeast.
7.6.4 SNPs
For the last few decades, the utilization of molecular markers, highlighting poly­morphism at the DNA level, has contributed a great deal in the development of animal genetics [32]. Among all the techniques, microsatellite DNA markers have been the most extensively utilized. A new marker tool, the SNP (see above) [32], is now in the picture and has recently gained high recognition. It is just a single-base change in a DNA sequence, with a usual alternative of two likely nucleotides at a certain position [26].
SNPs, the most abundant type of genetic variation, are now the main raw material essential for most genetic investigations and databases. Variation compris­ing indels, copy number variants, microsatellites and epigenetic markers remain parameters to consider and can inuence disease [27]. There are some reported 10 million SNPs in the human genome. The basic concept and types of SNPs are depicted in gures 7.6 and 7.7, respectively.
As mentioned above, SNPs are the most common type of genetic variant among people. Each and every SNP signies a variance in a single DNA building block, known as a nucleotide. For example, an SNP can substitute the nucleotide cytosine with the nucleotide thymine in a specic stretch of DNA. SNPs usually take place throughout a persons DNA. They occur once in every 300 nucleotides on average, i.e., there are approximately 10 million SNPs present in the human genome. As a rule, these deviations are present in the DNA between genes. They can behave as biological markers, facilitating researchers to locate genes that are related to disease. Once SNPs occur inside a gene or in a promoter region closer to a gene, they might show an additional direct role in disease by inuencing the genes function. The majority of SNPs have no inuence on health or development. Certain genetic differences, however, have been shown to affect human health. Scientists have discovered SNPs that may help to predict an individuals reaction to certain drugs, vulnerability against external environmental factors such as toxins and the threat of developing specic illnesses. SNPs may also be utilized to track the causes of disease, in particular the inheritance of disease genes inside families. Upcoming investiga­tions will work to recognize SNPs linked to complex diseases, e.g. heart disease, diabetes and cancer. Various academic institutions and industries are currently involved in the development of large-scale SNP detection projects, and these projects introduced data on thousands of publicly available SNPs. The most common
7-15
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Figure 7.6. Representation of SNPs.
Figure 7.7. Types of SNPs.
example of these is the HGBASE database. It has been projected that 90% of sequence variation in humans is because of SNPs. Human genome is projected to contain 3–17 million SNPs. Out of these, 5% SNPs are likely to occur in genes. Thus each individual gene is likely to contain six SNPs. Therefore, SNPs offer a molecular marker that is present in the genome at a very high density. The SNPs can thus be utilized to map genes involved in human diseases. These genes are exceptionally difcult to map using other molecular markers as these latter markers are not distributed at an adequate density in the genome. Therefore, by using SNPs as markers, every gene in the human genome can be mapped. Additionally, it might be possible to identify all the genes involved in human diseases. An SNP might be present within the relevant gene or very near the gene which is tightly linked with the SNP. SNP1 and SNP2 are the two alleles of an SNP.
Some SNPs are situated in the recognition sequence of an endonuclease. They alter the recognition sequence (base sequence) and produce restriction fragment
7-16
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
length polymorphism (RFLP). Some other SNPs are situated in the coding regions of genes. They are linked with denite phenotypes. The differences in the genetic sequences of individuals are understood to be responsible for the following:
Disease susceptibility.
Drug response.
Normal development and aging.
Reaction against environmental factors.
The field of study that is more concerned with the effect of genetic variation disease susceptibility and drug response is known as pharmacogenomics. Pharmacogenomics is the investigation of how genetic make-up regulates the response to a therapeutic intervention. Pharmacogenomics is likely to advance individualized treatments that will be safer and more effective [28]. Genetic variations may affect the drug reaction of an individual in the following ways:
They may inuence the metabolism of the drug itself.
They may affect the action of the drug on its target molecule.
Pharmacogenomics aims to raise drug efcacy and safety by comparing drug proper­ties to the genetic make-up of an individual. The response of the drug is expected to be inuenced by various genes associated with drug metabolism, drug transport and creation of drug targets. Thus, various inputs have been made to employ SNPs to further map hundreds or thousands of genes that have an inuence on the safety and efficiency of different drug treatments. These SNPs can be positioned on gene chips that can be employed to explore the genotype of those genes that are involved in responses to specific drugs. SNPs have the potential to describe the danger of an individuals susceptibility against numerous illnesses and reaction against various drugs. Biologically significant SNPs that are specically related with the threat of disease need to be identified [29]. The identication and understanding of great numbers of these SNPs are essential before they can be used widely as genetic tools. Depending upon these data, more effective and safe treatment can be considered for these individuals, such as variants in the gene encoding apolipoprotein E (apoE), which were found to be associated with variation in the reaction to a drug used for treatment. apoE supports lipid transport and binds to cell surface receptors to facilitate lipoprotein uptake. It is the main apolipoprotein produced in the brain [37]. The apoE protein is present in three different isoforms (E2, E3 and E4) that are the outcome of two non-synonymous SNPs (rs429358 and rs7412) present in exon 4 of the APOE gene. Gene apoE is involved in vulnerability to Alzheimer’s disease, and a cholines- terase inhibitor is employed to mitigate the symptoms of illness [30].
7.6.5 Inferences from comparative genomics
In the case of prokaryotes, the following genomes and characteristics have been considered:
Bacillus megaterium is important in the Bacillus phylogeny, as an evolutio- narily important species and, in particular, in understanding genome
7-17