Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5586_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
35 Мб
Скачать
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
evolution, dynamics and plasticity in the bacilli. B. megaterium is a non­pathogenic commercially existing host for the biotechnological production of numerous substances such as vitamin B12, penicillin acylase and amylases. The genome sizes of bacilli vary signicantly, such as from 0.58 × 106 bp (Mycoplasma genitalium) to 30 × 106 bp (Bacillus megaterium)[39, 41]. The gene density is, however, comparatively constant at one gene per kilo base pairs.
Vibrio cholerae O1 is reported as an important causative agent for infectious diseases [41]. A number of bacteria like V. cholerae have two or more circular chromosomes in their genomes. As a minimum, one of these chromosomes is obtained from a plasmid.
In the case of larger genomes such as E. coli, almost 50% of the ORFs (open reading frames, the portion of a reading frame that has the prospect of being translated) are of unidentied biological function [31].
Around 25% of all ORFs are exclusive and do not have substantial sequence resemblance to any other available protein sequence. Therefore, there are some new protein families yet to be revealed. The incidence of paralogues (pairs of genes that develop from similar ancestral genes) ranges from 26% for M. genitalium (the smallest self-replicating organism and an effective human pathogen related to a variety of genitourinary diseases) to 75% for Pseudomonas aeruginosa. The incidence of paralogues increases with genome mass [32].
In close relation with bacteria, chromosomal transposals are more likely to arise nearby the origin or terminus of replication. Several bacterial genomes present in the database comprise phage DNA incorporated into the bacterial chromosome. It is not uncommon for bacteria to hold several prophages in their chromosomes, which then establish a substantial part of the total bacterial DNA. Several prophages and prophage remnants litter many genomes [33]
Since severe complete genomes have been sequenced, gene order conservation between diverse organisms is evolving as a revealing feature of genome [33]. Gene order conservation has been employed for forecasting the function and functional interactions of proteins, as well as for investigating the devel­opmental relationships between genomes [34]. The explanations for gene order maintenance are still not well understood, since the association of the prokaryote genome into operons and adjacent gene transfer cannot perhaps be considered for all the examples of conservation found [45]. In fact, there is very small preservation of gene order between phylogenetically distant genomes. A statistically important preservation of gene order is most likely indicative of the related genes being organized into operons. This character­istic has been employed to explore a number of new operons and to allocate functions to the genes based on the prediction of operon function.
Genes that are out of use for an organism are often called pseudogenes. In general, inactivated gene copies are characterized by interference to their reading frames due to frameshifts and premature stop codons. Pseudogenes take place in prokaryotic genomes [35].
7-18
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
In eukaryotes, the following features of genomes have been observed:
The detection of orthologous groups is benecial for genome annotation, investigations on gene/protein evolution, comparative genomics and the determination of taxonomically restricted sequences. The procedures effec­tively utilized for prokaryotic genome examination have, however, proved difcult to apply to eukaryotes, since larger genomes can hold multiple paralogous genes and sequence information is often inadequate. A great number of the proteins encoded by the human genome have orthologues in other eukaryotic genomes. For example, 60% of human proteins have sequence resemblances to yeast, y, nematode or plant proteins.
Ancient transposon copies are found more in the human genome, whereas the genomes of Arabidopsis, Drosophila and Caenorhabditis have transposons of more recent origin [51, 52]. Bennett et al reported on comparisons between Caenorhabditis (100 Mb) and Drosophila ( 175 Mb) using ow cytometry. This work showed the genome size in Arabidopsis to be 157 Mb and therefore 25% larger than the Arabidopsis Genome Initiative Estimate of 125 Mb [36].
Transcription factors control the expression of genes at the transcriptional level. Arabidopsis has not more than 20% transcription factors. These factors are zinc-coordinating proteins [53]. In contrast, in the fruit y, yeast and nematode case, 51%–64% of the transcription factors are of this type.
Arabidopsis transcription factors are represented by several genes and by the assortment of gene families matched with those of D. melanogaster or C. elegans. Around 50% of the genes recognized so far have unknown function [37].
SINEs are retrotransposons that have appreciably signicant reproductive success throughout the course of mammalian development, and have played an important role in shaping mammalian genomes [55]. Around 75% of the repetitive DNA in human genome is because of LINEs (long interspersed nuclear elements), these are sets of non-LTR retrotransposons [56] which are extensive in the genome of several eukaryotes and in SINE (short interspersed nuclear element) sequences. SINEs are sequences of non-coding DNA existing at high incidences in numerous eukaryotic genomes.
As the genome size increases, gene density declines. It ranges from one gene per 2 kb in yeast to one gene per 10 kb in Drosophila, and one gene in 100 kb in humans.
The beginning of the genomic era has unlocked the doors to the study of complete genome organization, not only in eukaryotes but also in prokar­yotes, especially bacteria. Sufcient information on operon organization in E. coli, together with the completed chromosomal sequence of this bacterium, allowed the examination of the distances between genes and of functional interactions of adjacent genes in the same operon, as opposed to nearby genes in diverse transcription units. Bacterial genomes have distinguishing operons (a section of DNA that a repressor binds to); the E. coli genome has 600 operons [38].
7-19
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Approximately 15% of the 20 000 C. elegans genes are present in operons, multigene clusters governed by a single promoter [39]. The C. elegans genome contains operons similar to those of bacteria, and they contain 25% of its genes.
Eukaryotic genomes have relatively large amounts of repetitive DNA; this is a key reason for the variances in genome size.
Transposable elements are found in the genomes of nearly all eukaryotes. The human genome has a much higher density of transposable elements (44.4% of the genome) than A. thaliana (10.5%), D. melanogaster (3.1%) or C. elegans (6.5%) [40].
In plants such as maize and barley, genes are assembled in sections of DNA that can further be divided by long stretches of intergenic DNA. Earlier studies on the nuclear genomes of angiosperms showed that they are distinguished by a compositional compartmentalization and that the vast majority of genes derived from maize, rice, and barley are clustered in long DNA stretches that are separated by vast expanses of gene-empty DNA.
Several human ailment networks are highly preserved in the fruit y. Therefore, this insect can function as a model organism for the investigation of even complex human diseases, such as neurological disorders. D. mela- nogaster is a well-studied and highly tractable genetic model organism to explore the molecular mechanisms of human diseases. Various biological, physiological and neurological characteristics are maintained between mam­mals and D. melanogaster, and approximately 75% of human disease-causing genes are assumed to have an efcient homolog in the y.
The incidence of sole genes (having no paralogues in the genome) is highest (>70%) in yeast and fruit ies, and the smallest in Arabidopsis (35%).
Variances in intron and exon structure are recognized in eukaryotic genomes; the Arabidopsis genome varies considerably. It has been reported that there are three basic patterns of exon–intron variation in eukaryotic genomes, and these patterns were consistently present in almost all 13 genomes analyzed [41]. There is a better degree of alternative splicingin humans than in the other three eukaryotes. This facilitates more proteins to be encoded by each gene.
1. Various eukaryotic proteins are mosaic proteins, i.e., they are made of a number of different domains. Around 90% of the domains known in human proteins are present in Drosophila and C. elegans proteins. Therefore, vertebrate development has been central for the creation of the few new domains [
61]. Looking at both eukaryotes and prokaryotes,
we have learned the following: certain bacterial species have more genes than lower eukaryotes.
2. Certain genes are present in a wide variety of organisms. For instance, some of yeast genes have homologs among the human genes of unknown function.
3. Unexpectedly, gene number does not essentially link with the compli­cation of the organism or its evolutionary ladder.
Archaea is one of the three domains of life (alongside bacteria and eukaryotes). To date, 16 complete archaeal genomes have been sequenced, with a conserved core of
7-20
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
313 genes that are present in all sequenced archaeal genomes. The genomes of Archaea have similarities to eukaryotic genomes. For example, certain genes look like those of eukaryotes as they have histone genes and their DNA appears to be prearranged into chromatin [42].
7.6.6 Gene order comparisons (for phylogenetic inference)
Comprehensive information on gene maps or even complete nucleotide sequences for small genomes leads to the possibility of evolutionary interpretation based on the macrostructure of complete genomes. The mathematical modeling of development at the genomic level and the related inferential tools are, however, qualitatively different from the normal sequence comparison theory established to study evolution [43]. Equating gene order in dissimilar organisms is one of the techniques for developing molecular phylogeny. Once the gene instruction in a given region of the genome of two organisms is similar, they are called syntenic and the phenom­enon is known as synteny. Gene order comparison in S. cerevisiae and C. albicans discloses numerous cases of inversions, around half of which are single gene inversions, in these two species separated by 140–330 million years [44]. Comparatively, inversions in prokaryotes tend to include much larger regions of genomes and are focused on the origin and terminus of replication. Additionally, the degree of inversion is much higher in eukaryotes than in prokaryotes. Furthermore, the path of transcription appears to have little effect on gene location. A number of regions of the human genome are highly conserved in dogs, cattle and sheep. One of the main benets of synteny is that information on gene location from a highly mapped organism can be employed to locate the analogous gene in a poorly mapped relative.
7.6.7 Phylogenetic footprinting (computational method)
Phylogenetic footprinting is a process for the detection of regulatory elements in a set of orthologous regulatory regions from various species [45]. In other words, phylogenetic footprinting is a method employed to detect transcription factor binding sites (TFBS) inside a non-coding region of DNA of interest. This is done by equating it to orthologous sequences in different species. Scientists have discovered that non-coding pieces of DNA hold binding sites for regulatory proteins that direct the (spatiotemporal) expression of genes. A general representation of phylogenetic footprinting is shown in gure 7.8.
These TFBS or regulatory motifs are difcult to detect, mainly because they are very small in length, and can display sequence variation. Two main concepts on which phylogenetic footprinting rely are:
Transcription factor functions and their DNA binding preferences are well­conserved between diverse species.
Signicant non-coding DNA sequences that are important for directing gene expression will display differential selective pressure. A more gradual rate of change occurs in TFBS than in others.
7-21
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Figure 7.8. Phylogenetic footprinting.
7.6.8 Origins, evolution and phenotypic impact of new genes
The genetic code is shared, and the organization of the codons in the standard codon table is highly non-random [46]. The three main ideas on the origin and evolution of the code are the following:
Stereochemical theory: In this theory codon assignments are dictated by physico-chemical afnity between amino acids and the cognate codons (anticodons);
Coevolution theory: This postulates that the code structure coevolved with amino acid biosynthesis pathways.
Error minimization theory: According to this theory selection to reduce the adverse effect of point mutations and translation errors was the principal factor of the code s evolution.
Thousands of genes present in eukaryotic genomes do not have complements in prokaryotes. Several newgenes are mostly regular prokaryotic genes that have been altered beyond recognition, such as domains involved in protein–protein interactions. Several neweukaryotic domains are α-helical; these could have developed from the condensed coiled structures present in prokaryotes, which are now particularly abundant in eukaryotes. Certain other genes have arisen in completely unpredicted ways. For example, the hedgehog gene seems to have developed by gene fusion; during development an intein domain has fused with an extensively modied metalloprotease domain. However, the basis of several other genes remains unknown. A number of new genes might have risen from transposable elements. Among 13 799 human genes studied, 533 (4%) protein-coding regions comprised of transposable elements or sections of such elements were found. Such a transposable element could have incorporated into an exon. Otherwise, it may have initially rst incorporated into an intron and then enrolled as an exon. This can become possible once the transposable element has possible splice sites. The incorporation of transposable elements into eukaryotic genomes could be one
7-22
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
reason for the high incidence of alternative joining in human protein-coding genes. Therefore, incorporation of transposable elements into genes can produce new genes. For example, in their protein-coding regions, two mouse genes have one or numerous transposable elements, these genes have no human or rat orthologues.
7.6.9 The concept of minimum genome size
The minimal genomemethod aims to estimate the smallest number of genetic elements adequate to create a modern-type free-living cellular organism. Or, in other words, the minimum genome size can be dened as the least number of genes essential to sustain life, and can at best only be predicted. A simplied scheme of the concept of minimum genome size is shown in gure 7.9.
For the following reasons, mycoplasma is considered a model organism for
minimum genome size:
It is a member of the mollicutes.
It evolved by massive genome reduction.
It lacks genomic redundancy.
It is an obligate parasite.
It has the smallest identied genome of any free-living organism capable of
growing in axenic culture.
It is wall-less.
Essential genes are considered important for its survival, e.g. genes that are encoding for protein that maintains the central metabolism, replicate DNA, translate genes into proteins, maintain a basic cellular structure, and mediate transport processes into and out of the cell. Some essential genes can tolerate mutations that are deleterious, but not wholly lethal, since they do not completely abolish the genes function. Some of the essential genes, however, appear to perform non-essential functions, such as aging and cell death, while many of the non-essential genes play critical roles in cell survival. Essential genes can be difcult to study. If an essential
Figure 7.9. Concept of minimum genome size.
7-23
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
gene is disrupted in the macronucleus, it is unlikely that complete replacement of the wild-type copies with the selectable marker will ever be achieved. There are a number of different transposon-based approaches available for dening essential genes. There are two ways to identify essential genes or regions of the bacterial chromosome which can be based on the location of the transposon insertion: the negativeand the positiveapproaches. A problem in trying to identify essential genes is that a knockout of an essential gene is lethal. Therefore, the use of the negative approach will help with the identication of many regions that are not essential and enable one to say with some certainty that regions in which transposon insertions are not observed are likely to be essential.
Whole-genome sequences are now available for a great number of assorted species [47]. Gene content quantication, expansion of the gene family, orthologous gene conservation and gene displacement are promising for the assessment of the minimal set of proteins adequate for cellular life.
The genes found in different organisms with the smallest genomes and their roles are matched carefully. Moreover, investigational ndings are collected by deactivat­ing specic genes of these organisms and measuring the outcome on the organisms survival. Based on these investigations, it has been estimated that living creatures need at least 250–350 genes. This minimum gene number is essential for the organisms to exist as autonomous, self-replicating organisms. The unicellular eukaryote minimum genome size was found to be 2.9 × 10
6
bp (the parasite Encephalitozoan cuniculi). E. cuniculi infects different mammals, including humans, and can be a source of digestive and nervous clinical syndromes in HIV-infected or cyclosporine-treated people [48]. The genome of E. cuniculi has very little repetitive DNA other than that for rDNA, and is expected to contain approximately 2000 genes [49]. Therefore, transformation from prokaryotes to eukaryotes has aug­mented the gene number by a factor of 7–8, however, the genome size has not grown remarkably. The minimum gene number in multicellular eukaryotes, e.g. Drosophila, Arabidopsis, etc, appears to be 16 000; this signicant increase would be essential for development and growth, and for sufciently reacting to environ­mental or other external factors. Despite its toxicity, Fugu rubripes, the puffersh is widely consumed in Japan and considered the most delicious of all sh. Interest in F. rubripes is not only restricted to taste and toxicity, however. The smallest genome size for a vertebrate is that for the F. rubripes; it has only 4 × 10
8
bp DNA, but this genome has 35 000 genes, which is similar to that of the human genome [50]. The puffersh genome has tightly packed genes, lacks repetitive factors and has single short intergenic and intronic sequences. This drop in genome size is attained by decreasing the sequences not involved in genetic functions to the minimum; it is known as genome compaction. Otherwise, genome size may be decreased by gene loss; this has happened in several parasites.
7.6.10 Comparative genomics analysis of mitochondria and chloroplasts
Comparative genomics investigation has offered new understanding of the origin of organelles by endosymbiosis and exposed the vast evolutionary dynamics of
7-24
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
organellar genomes. Moreover, they have signicantly assisted in elucidating phylogenetic relationships, particularly in algae and primary land plants with inadequate morphological and anatomical diversity [51]. The genomes of mitochon­dria are often circular but linear molecules also occur. A number of observations on mtDNA are as follows:
Animal and fungal mitochondrial DNA (mtDNA) is very small (15–20 kb), compared to plant mtDNA (200–2000 kb) which is more complex than that of either animals or fungi.
The greater genome size of plant mitochondria is mainly due to a greater amount of spacer DNA. For example, Arabidopsis mtDNA is almost 20 times larger than human mtDNA, but has less than twice the number of genes.
The mtDNA of numerous species can be divided into two basic types: ancestral and derived DNA. Plant and animal mtDNAs are derived genomes. In the case of angiosperms, there has been extensive gene loss, however, the mtDNA has grown in size because of the duplication and capture of DNA from cpDNA and the nuclear genome.
Mitochondria are thought to be derived from Rickettsia prowazekii, the agent responsible for epidemic, louse-borne typhus in humans [75]. The Rickettsia (alpha-proteobacteria) multiply in eukaryotic cells only. The organization, structure and gene content of this bacterium look exactly like the mtDNA of Reclinomonas americana [52].
Seed plant mitochondrial genomes are exceptionally uid in size, structure and sequence content, with the accumulation and activity of repetitive sequences underlying much of this variation. mtDNA genes have been transferred into the nucleus, as this transfer has stopped in animals, however, it continues to occur in the case of plants and protists [76]. For example, a gene from mitochondria, cox2, is still under the development of transfer in the case of legumes. In some legumes, cox2 is present in the mitochondria, whereas in other organisms it is present inside the nuclear genome, and in some it is present in both the nuclear and mitochondrial genomes [53].
The three plant genomes gene types display differences in their evolutionary rates. Nuclear genes develop the fastest, followed by chloroplast genes and nally mitochondrial genes, in spite of the fact that the mitochondrial genome displays extensive reorganization of its structural organization. This gradual rate of evolution has made mitochondrial gene sequences unattractive to plant phylogenetic studies at the suborder and subfamily levels, but has proven to be valuable in inferring ancient phylogenetic relationships and in approximating the time of early diversication events in seed plants. There has been cross-species acquisition of DNA or horizontal gene transfer of genes by plant mitochondrial genomes such as the homing group I intron, introns that encode site-specic endonucleases to further catalyze their effective distribution from introns including alleles to intron-lacking alleles of similar gene in genetic crosses. This intron is present in the cox gene of some angiosperm species [54].
7-25
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)

7.7 Gene estimation and counting

Annonation is a procedure which detects genes, and explores their regulatory sequences and promising functions. Annonation designates the non-protein-coding genes (those that do not encode protein sequences), e.g. coding genes for rRNA, tRNA and nuclear RNAs, mobile genetic elements (DNA segments that encode enzymes and other proteins that control the DNA movement inside genomes, called intracellular mobility or, between bacterial cells, intercellular mobility) and repet­itive sequence families existing in the genome. The role of annonation begins only after the genome sequence has been attained. Gene estimation is a signicant issue for computational science. There are numerous logarithms that can be employed for gene estimation by utilizing identied genes as a training data set which can be further utilized to form a model or to test (or validate) the built model. Gene counting is problematic until the actual locations of genes in a genome are determined. The occurrence of overlapping genes and splice variations makes the job more challenging as it becomes more difcult to forecast which region of DNA should be considered as the same or as many different genes.
Numerous genes in eukaryotes have an arrangement of exons (coding regions) trailed by introns (non-coding regions). Because of this, genes are not systematized as continuous ORFs (open reading frames). An ORF has a series of codons that require an amino acid sequence.
Also, eukaryotic genes are often widely spaced, thus there is a greater probability of nding false genes. However, novel methods of ORF scanning software for eukaryotic genes allow more effective scanning. It has been observed that approx­imately 99.8% of the 3.2 billion base pairs of two humans are identical and merely
0.2% are different. Of every 500 nucleotides, only one nucleotide varies between two individuals. It means that a difference in a few sites in the DNA sequence can result in severe disorders and dissimilar features in human beings.
7.7.1 Genome similarity, SNPs and comparative genomics
As discussed above, approximately 99.8% of every human genome is identical to any other human genome and 0.2% different, which indicates that two individuals vary only in 6 million locations out of 3.2 billion locations (gures 7.10 and 7.11). Two human beings will be similar if these locations have little inuence. As already discussed, the human genome is closely related to that of chimpanzees (more than 98%). Therefore, variance in some locations in DNA results in a unique individual. One nucleotide among every 500–1000 nucleotides always differs between two individuals. The most considerable variations in individual genomes are SNPs, which have been detected in both coding and non-coding regions of genome [55]. SNPs as discussed above are fundamentally varied in a single base, i.e., A, C, G or T resulting in DNA variations which further result in the presence of different bases at such locations. With varying individuals the type of nucleotide or base existing at a given position on a chromosome can also vary. It has been observed that 90% of sequence variation among humans is because of the presence of SNPs. Thus SNPs offer a molecular marker that is present in high density. During DNA nger-printing
7-26
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Figure 7.10. An SNP map.
Figure 7.11. An SNP map to determine how patients are likely to respond to a particular drug.
in mainly non-coding parts of the genome, such genetic deviations among diverse human beings are used. These genetic variations are also accountable for the severity of disease and the reaction of the body to treatment.

7.8 Genomes: genome evolution

Genome trees are a way to contain the astounding amount of phylogenetic data that is found in genomes. Diverse formalisms have been presented to construct genome trees on the basis of various features of the genome [ 56]. On the basis of these characteristics, we divide genome trees into ve classes:
7-27