Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5586_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Acknowledgement
- •Author biographies
- •Professor Ahmed Al-Harrasi
- •Dr Saurabh Bhatia
- •Dr Ajmal Khan
- •1.1 Introduction
- •1.2 Properties of enzymes
- •1.3 Catalysis
- •1.4 The structure of enzymes
- •1.5 Structural features: primary and secondary structures
- •1.6 Nomenclature and classification
- •1.6.1 Class 1—oxidoreductase
- •1.6.2 Class 2—transferase
- •1.6.3 Class 3—hydrolases
- •1.6.4 Class 4—lyases
- •1.6.5 Class 5—isomerases
- •1.6.6 Class 6—ligases
- •1.7 The mechanism of action of enzymes
- •1.7.3 Covalent catalysis
- •1.8 Catalysis via chymotrypsin
- •1.8.1 Intermediary stages of chymotrypsin
- •1.8.2 Kinetic behavior of α-chymotrypsin
- •1.8.3 Selective proteolysis in creation of the catalytic sites of enzymes
- •1.8.4 Kinetic models for enzymes
- •1.8.5 Enzyme mediated acid–base (general) catalysis
- •1.8.6 Metallozymes
- •1.9 Enzyme inhibition
- •1.10 Pharmaceutical applications
- •1.10.1 Diagnostic applications of enzymes
- •1.10.2 Enzymes in therapeutics
- •1.11 Plants and algae enzyme systems
- •1.12 Enzyme safety
- •1.13 Enzyme structure determination
- •1.13.1 X-ray crystallography
- •1.13.2 NMR spectroscopy
- •1.13.3 Cryo-electron microscopy
- •1.14 Enzyme engineering and design
- •1.14.1 Directed evolution of enzymes
- •1.14.2 Rational design of enzymes
- •1.14.3 Applications of engineered enzymes
- •1.15 Enzymes in medicine and healthcare
- •1.15.1 Enzyme-targeted drug delivery
- •1.15.2 Enzymes as drug targets
- •1.15.3 Challenges and opportunities in enzyme drug discovery
- •1.15.4 Enzymes in gene therapy
- •1.15.5 Enzymes in personalized medicine
- •1.15.6 Enzyme biomarkers in disease diagnosis
- •1.15.7 Pharmacogenomics and enzyme variability
- •1.15.8 Enzyme-based therapies for personalized treatment
- •1.16 Enzymes in bioremediation
- •1.17 Enzymes in agriculture and crop production
- •1.18 Enzymes in waste management
- •References
- •2.1 Introduction
- •2.1.1 Sources of enzymes
- •2.2 Enzyme production technology
- •2.2.1 Selection of microorganisms
- •2.2.2 Medium selection
- •2.2.3 Production process
- •2.2.5 Cell debris removal
- •2.2.6 Nucleic acid removal
- •2.2.7 Precipitation of enzymes
- •2.2.8 Liquid–liquid partition
- •2.2.9 Chromatographic separation
- •2.2.10 Drying and packing
- •2.2.11 Regulation of microbial enzyme production
- •2.2.12 Induction
- •2.2.13 Feedback repression
- •2.2.14 Nutrient repression
- •2.3 Procedures involved in enzyme production
- •2.3.1 Source and location of enzymes
- •2.3.2 The variety of microorganisms
- •2.3.3 Media for fermentation
- •2.3.4 Fermentation
- •2.3.5 Enzyme extraction
- •2.3.7 Finishing operations
- •2.4 Recombinant proteins from algae
- •2.5 Enzyme immobilization techniques
- •2.5.1 Advantages and applications of enzyme immobilization
- •2.5.2 Methods of enzyme immobilization
- •2.6 Enzyme engineering for enhanced stability and activity
- •2.6.1 Protein engineering strategies
- •2.6.2 Improving enzyme thermostability
- •2.7 Upstream process intensification
- •2.7.1 High cell density fermentation
- •2.7.2 Solid-state fermentation
- •2.7.3 Continuous fermentation
- •2.7.4 Microbial consortia for enzyme production
- •2.7.5 In situ product removal strategies
- •2.8 Enzyme production from extreme environments
- •2.8.1 Psychrophiles (cold-loving)
- •2.9.4 Automation and robotics in downstream processing
- •References
- •2.8.2 Thermophiles (heat-loving)
- •2.8.3 Acidophiles (acid-loving)
- •2.8.4 Alkaliphiles (alkaline-loving)
- •2.8.5 Halophiles (salt-loving)
- •2.8.6 Applications of extremozymes in biotechnology
- •2.9 Downstream process intensification
- •2.9.1 Continuous chromatography
- •2.9.2 Process integration and optimization
- •3.1 Industrial enzymes
- •3.2 Bacterial α-amylases
- •3.3 Fungal α-amylases
- •3.4 Bacterial proteases
- •3.5 Fungal proteases
- •3.6 Glucose isomerase (d-xylose ketol-isomerase; EC. 5.3.1.5)
- •3.7 Penicillinase
- •3.8 Chloramphenicol acetyltransferase
- •3.9 Aminoglycoside antibiotic inactivating enzymes
- •3.10 Fibrinolytic enzymes
- •3.10.1 Streptokinase
- •3.10.2 Urokinase
- •3.10.3 Tissue plasminogen activator (t-PA)
- •3.11 Biotechnological applications of enzymes
- •3.11.1 Algae and plant research
- •3.11.2 Immobilization
- •3.12 Industrial enzymes
- •3.12.1 Glucoamylase
- •3.12.2 Cellulases
- •3.13 The role of enzymes in the synthesis of functional foods
- •3.13.1 Lipases
- •3.13.2 Proteases
- •3.13.3 Carbohydrate-modifying enzyme
- •3.13.4 Tannase
- •3.13.5 Asparaginase
- •3.13.6 The phytases
- •3.14 Enzymes used as additives to food
- •3.14.1 The enzymatic synthesis of dietary antioxidants
- •3.14.2 The use of ascorbyl esters
- •3.14.3 Polyphenolic esters
- •3.14.4 Synthesis of sugars esters surfactants by enzymes
- •References
- •4.1 Introduction
- •4.2 Types of immobilization
- •4.2.1 Surface immobilization by covalent coupling
- •4.2.2 Adsorption
- •4.2.3 Complexation and chelation
- •4.2.4 Within-support immobilization
- •4.2.5 Cell immobilization
- •4.2.6 Commercial production of enzymes
- •4.3 Genetic engineering for microbial enzyme production
- •4.3.1 Cloning methods
- •4.4 Protein studies for modification of commercial enzymes
- •4.5 Enzyme and cell immobilization
- •4.6 Immobilization methods
- •4.6.1 Adsorption methods
- •4.6.3 Ionic binding
- •4.6.4 Hydrophobic adsorption
- •4.6.6 Entrapment method
- •4.6.7 Covalent binding
- •4.6.8 Cross-linking
- •4.7 Choice of immobilization technique
- •4.7.1 Immobilization of l-amino acid acylase
- •4.7.2 Stabilization of soluble enzymes
- •4.8 Immobilization of cells
- •4.8.1 Immobilization of viable cells
- •4.8.2 Immobilized non-viable cells
- •4.8.3 Drawbacks of immobilizing eukaryotic cells
- •4.8.4 The effect of immobilization on enzyme properties
- •4.8.5 Immobilized enzyme reactors
- •4.8.6 Applications of immobilized enzymes and cells
- •4.9 Manufacture of commercial products
- •4.9.1 Production of l-amino acids
- •4.9.2 Production of high-fructose syrup
- •4.9.3 Immobilized enzyme and cell analytical applications
- •4.10 Immobilized enzymes for biomedical applications
- •4.11.1 Bioluminescence
- •4.11.2 The measurement of biomass using bioluminescence-based techniques
- •4.11.4 Biosensors relying on bioluminescence
- •4.12 Bioluminescence-based microbial biosensors
- •4.12.1 The microencapsulation process involves the utilization of polymers and cells
- •4.12.2 Microcapsule evaluation
- •4.12.4 Modern developments in cell encapsulation
- •4.13 Immobilization of microalgae
- •4.13.1 Techniques for immobilization
- •4.13.2 Use of cryopreserved algae
- •4.13.3 Removal of nitrogen and phosphorous
- •4.13.4 Disposal of metals
- •4.13.5 Biosensor development
- •References
- •5.1 Introduction
- •5.2 Principles of a biosensor
- •5.3 Different types of biosensors
- •5.3.1 Electrochemical biosensors
- •5.3.2 Thermometric biosensors
- •5.3.3 Optical biosensors
- •5.3.4 Piezoelectric biosensors
- •5.3.5 Whole-cell biosensors
- •5.3.6 Immunobiosensors
- •5.4 Applications of biosensors
- •5.4.1 Applications in medicine and health
- •5.4.2 Applications in industry
- •5.4.3 Applications in pollution control
- •5.4.4 Applications in the military
- •5.4.5 Immobilized enzymes and cell therapeutic applications
- •5.5 Recent advancements in biosensor technology
- •5.5.1 Electrochemical biosensors
- •5.5.2 Optical/visual biosensors
- •5.5.3 Silica, quartz/crystal, and glass biosensors
- •5.5.4 Nanomaterials-based biosensors
- •5.5.5 Fluorescent biosensors that are either genetically encoded or synthetic
- •5.7 Technological comparison of biosensors
- •5.9 Grand challenges in biosensors and biomolecular electronics
- •5.9.1 Sensitivity
- •5.9.2 Multiplex capability
- •5.9.3 Continuous monitoring in vivo
- •5.10.1 Sustainability to the ecosystem
- •References
- •6.1 Introduction
- •6.2 Types of biotransformation reactions
- •6.3 Sources of biocatalysts and techniques for biotransformation
- •6.3.1 Growing cells
- •6.3.2 Non-growing cells
- •6.3.3 Immobilized cells
- •6.3.4 Immobilized enzymes
- •6.4 Product recovery in biotransformations
- •6.5 Application of biotransformation in the production of pharmaceutical products
- •6.5.1 Biotransformation of steroids
- •6.5.2 Biotransformation of antibiotics
- •6.5.3 Biotransformation of arachidonic acid to prostaglandins
- •6.5.4 Biotransformation for the production of ascorbic acid
- •6.5.5 Biotransformation of glycerol to dihydroxyacetone
- •6.5.6 Biotransformation for the production of indigo
- •6.6 Mechanisms of enzyme action in biotransformation
- •6.6.1 Enzyme kinetics and biotransformation
- •6.6.2 Cofactors and coenzymes in biotransformation
- •6.6.3 Enzyme inhibition and activation
- •6.7 Biotransformation in environmental applications
- •6.7.1 Degradation of pollutants
- •6.7.2 Enzymatic breakdown of pesticides
- •6.8 Emerging technologies in biotransformation
- •6.8.1 Enzyme engineering and directed evolution
- •6.8.3 Biotransformation of lipids for healthy oils
- •6.9 Biotransformation challenges and future perspectives
- •6.9.1 Scalability issues in industrial applications
- •6.9.2 Regulatory and safety concerns
- •6.9.3 Challenges in enzyme storage and stability
- •6.9.4 Future trends and emerging areas of research
- •6.9.5 Biotransformation in biofuel production
- •6.9.6 Biotransformation in the cosmetic industry
- •6.9.7 Specialized enzyme systems: lignin-modifying enzymes in biotransformation
- •References
- •7.1 Introduction
- •7.2 Characterizations in genomics
- •7.3 Historical background
- •7.4 Genome sequencing
- •7.4.1 Clone-by-clone sequencing
- •7.4.2 Human whole-genome shotgun sequencing
- •7.4.3 Compilation of genome resources
- •7.5 Understanding bioinformatics and sequencing
- •7.6 Comparative genomics as a technique to understand evolution
- •7.6.2 Horizontal or lateral gene transfer
- •7.6.3 Genome similarity or homology
- •7.6.4 SNPs
- •7.6.5 Inferences from comparative genomics
- •7.6.6 Gene order comparisons (for phylogenetic inference)
- •7.6.7 Phylogenetic footprinting (computational method)
- •7.6.8 Origins, evolution and phenotypic impact of new genes
- •7.6.9 The concept of minimum genome size
- •7.6.10 Comparative genomics analysis of mitochondria and chloroplasts
- •7.7 Gene estimation and counting
- •7.7.1 Genome similarity, SNPs and comparative genomics
- •7.8 Genomes: genome evolution
- •7.8.1 Microbial genome reduction in bacteria
- •7.8.2 Role of duplications in the origin and evolution of the eukaryotic genome
- •7.8.3 Gene duplications increase genetic diversity and complexity
- •7.9 Algae bioinformatics
- •7.9.1 Scope of algae bioinformatics
- •7.9.2 What is involved in algae bioinformatics
- •7.9.3 Role of algae bioinformatics
- •7.9.4 Steps involved in obtaining the data for analysis using bioinformatics
- •7.10 Functional genomics
- •7.10.1 Introduction to functional genomics
- •7.10.2 Transcriptomics: studying the RNA molecules
- •7.10.3 Proteomics: understanding the world of proteins
- •7.10.4 Metabolomics: exploring cellular metabolites
- •7.10.5 Interactomics investigating protein–protein interactions
- •7.11 Structural genomics
- •7.11.1 Introduction to structural genomics
- •7.11.2 The approaches used in the domain of structural genomics
- •7.11.3 Importance of structural genomics in drug design
- •7.12 Epigenomics and epigenetics
- •7.12.1 Epigenetic inheritance and diseases
- •7.13 Pharmacogenomics
- •7.13.1 The importance of personalized medicine
- •7.13.2 The impact of genetic variations on drug response
- •7.13.3 Additional insights on pharmacogenomics
- •7.13.4 Pharmacogenomic tests in the market
- •7.13.5 Challenges in implementing pharmacogenomics
- •7.14 Population genomics
- •7.14.1 Studying genetic variation across populations
- •7.14.2 Population genomics techniques
- •7.14.3 Understanding human migration and evolution through population genomics
- •7.14.4 Conservation genomics in endangered species
- •7.15 Microbiome genomics
- •7.15.1 Introduction to the human microbiome
- •7.15.2 Techniques in studying microbial communities
- •7.15.3 Role of microbiome in human health and disease
- •7.15.4 Environmental microbiomes and their importance
- •7.16 Synthetic biology and genome editing
- •7.16.1 Techniques like CRISPR/Cas9 in genome editing
- •7.17 Systems biology and genomics
- •7.17.1 Integrative approaches in genomics
- •7.17.2 Modeling biological systems and networks
- •7.17.3 Challenges and opportunities in systems biology
- •7.18 Genome-wide association studies (GWAS)
- •7.18.1 Introduction to GWAS
- •7.18.2 Techniques and platforms for GWAS
- •7.18.3 Challenges in interpreting GWAS results
- •7.19 Future of genomics
- •7.19.1 Next-generation sequencing technologies
- •7.19.2 Ethical considerations in genomics research
- •7.19.3 The role of AI and machine learning in genomics
- •7.19.4 Personalized medicine and its potential impact
- •8.1 Introduction
- •8.2 Types of proteomics
- •8.2.1 Structural proteomics
- •8.2.2 Functional proteomics (strategy)
- •8.2.3 Expression proteomics
- •8.3 Basic techniques involved in proteomics
- •8.3.1 Sequence alignment (algorithms)
- •8.3.2 Protein structure (annotation resources)
- •8.3.3 Protein structural investigation
- •8.3.4 Two-dimensional gel electrophoresis in proteomics
- •8.3.5 Domain fusion method (or rosetta stone method)
- •8.4 Complete proteome of Mycoplasma genitalium
- •8.5 Architecture and design of the nuclear pore complex
- •8.6 Functional genomics and systems biology
- •8.6.2 Transcriptome, proteome and genomes
- •8.6.3 DNA arrays: a potential genomic tool
- •8.6.4 Gene function determination from sequence information
- •8.6.5 Protein interactions
- •8.7 Synthetic genomics
- •8.8 Advanced techniques in proteomics
- •8.8.1 Mass spectrometry in proteomics
- •8.8.2 Tandem mass spectrometry
- •8.8.3 Quantitative proteomics using mass spectrometry
- •8.8.4 Other advanced techniques in proteomics
- •8.8.5 Chromatography in proteomics
- •8.9 Proteogenomics
- •8.9.1 Proteogenomics role in precision medicine
- •8.10 Single-cell proteomics
- •8.10.1 Technologies enabling single-cell proteomics
- •8.11 Clinical and diagnostic proteomics
- •8.12 Metaproteomics
- •8.13 Emerging topics in proteomics
- •8.13.1 Data-independent acquisition (DIA)
- •8.13.2 Top-down proteomics
- •8.13.3 Targeted proteomics and selected reaction monitoring (SRM)
- •8.13.4 Proteomics in plant research
- •8.14 Ethical and data management issues in proteomics
- •8.14.1 Open-source platforms for proteomic analysis
- •8.15 Cellular and molecular dynamics
- •8.15.1 Molecular mechanisms of protein function
- •8.15.2 Protein degradation pathways
- •8.15.4 Cellular signaling pathways
- •8.15.5 Proteomic analysis of signaling networks
- •8.15.6 Signaling pathway dysregulation in disease
- •8.15.7 Targeting signaling pathways in drug discovery
- •8.15.8 Crosstalk between signaling pathways
- •8.16 Membrane proteomics
- •8.16.1 Techniques for membrane protein analysis
- •8.16.2 Membrane protein structure and function
- •8.16.3 Membrane proteins in disease
- •8.16.4 Drug targeting of membrane proteins
- •8.17 Subcellular proteomics
- •8.17.3 Proteomics of cellular compartments
- •8.17.4 Techniques for subcellular proteomic analysis
- •References
- •9.1 Introduction
- •9.2 History of bioinformatics
- •9.3 Sequences and nomenclature
- •9.3.1 DNA sequences
- •9.3.2 Amino acid sequences of proteins
- •9.3.3 Types of sequences in nucleotide sequence databases
- •9.3.4 Databases
- •9.3.5 Search engines and analysis tools
- •9.3.6 Various indian databases
- •9.4 Investigation by means of bioinformatics tools
- •9.4.4 Detection of noncoding RNA
- •9.4.5 Genome annotation
- •9.4.6 Molecular phylogenetics
- •9.5 Computational approaches in bioinformatics
- •9.5.1 Algorithm development
- •9.5.2 Phylogenetic tree construction algorithms
- •9.5.3 Machine learning algorithms in bioinformatics
- •9.5.4 High-performance computing (HPC) in bioinformatics
- •9.5.5 Cloud computing in genomics
- •9.5.6 GPGPU (general-purpose computing on graphics processing units)
- •9.5.7 Big data analytics in bioinformatics
- •9.5.8 Systems biology modelling
- •9.5.9 Systems pharmacology
- •9.5.10 Multiscale modeling
- •9.5.11 Computational genomics
- •9.5.12 Functional genomics
- •9.5.13 Comparative genomics
- •9.5.14 Epigenomics
- •9.5.15 Metagenomics
- •9.6 Bioinformatics in precision medicine
- •9.7 Translational bioinformatics
- •9.8 Bioinformatics in drug discovery and development
- •9.8.2 AI-driven drug discovery
- •9.9 CRISPR and genome editing in bioinformatics
- •9.10 Integrative and multi-omics analysis
- •References
- •10.1 Protein and enzyme engineering
- •10.2 Designing macromolecules
- •10.3 Protein engineering versus enzyme engineering
- •10.4 Protein engineering
- •10.5 Foundation of protein (enzyme) engineering
- •10.6 Basic assumptions for protein engineering
- •10.7 Steps involved in protein engineering
- •10.7.1 Studying three-dimensional protein structure
- •10.7.2 Protein modeling
- •10.7.3 Perturbation theory
- •10.8 Methods of protein engineering
- •10.9 Mutagenesis and selection of mutant enzymes
- •10.10 Gene modifications or gene synthesis for protein engineering
- •10.11 Multi-enzyme systems
- •10.12 Chemical modification of enzyme
- •10.13 Some early achievements of protein engineering
- •10.14 Computational approaches in protein engineering
- •10.14.1 Molecular dynamics simulations
- •10.14.2 Quantum mechanical calculations
- •10.14.3 Docking and ligand optimization
- •10.14.4 Machine learning algorithms in protein design
- •10.15 Directed evolution techniques
- •10.15.1 Error-prone PCR
- •10.15.3 Saturation mutagenesis
- •10.15.4 Phage display
- •10.16 Post-translational modifications
- •10.16.1 Glycosylation engineering
- •10.16.2 Phosphorylation engineering
- •10.16.3 Methylation and acetylation
- •10.16.4 PEGylation for enzyme stability
- •10.17 Structural flexibility and allosteric regulation
- •10.17.1 Intraprotein communication pathways
- •10.17.3 Modulator design
- •10.17.4 Coupling allosteric regulation with catalytic function
- •10.18 Protein–protein and protein–ligand interactions
- •10.18.1 Characterizing binding sites
- •10.18.3 Interaction networks
- •10.18.4 Biophysical methods for interaction studies
- •10.19 Applications in synthetic biology
- •10.19.1 Metabolic pathway engineering
- •10.19.2 Genetically encoded sensors
- •10.19.3 Protein-based logic gates
- •10.19.4 Gene circuits for dynamic control
- •10.20 Engineering multi-functional proteins
- •10.20.1 Fusion proteins
- •10.20.2 Protein scaffolds
- •10.20.3 Modular protein design
- •10.20.4 Dual-enzyme systems
- •10.21 Ethical and safety considerations
- •10.21.1 Bioethics in protein engineering
- •10.21.2 Biosafety and environmental concerns
- •10.21.3 Intellectual property rights
- •10.21.4 Regulatory frameworks
- •10.22 Studies in protein engineering
- •10.22.1 Therapeutic proteins
- •10.22.2 Industrial enzymes
- •10.22.3 Diagnostic proteins
- •10.23 Single-molecule techniques in protein engineering
- •10.23.1 Atomic force microscopy
- •10.23.2 Single-molecule FRET
- •10.23.3 Optical tweezers
- •10.23.4 Patch-clamp technique
- •10.24 High throughput screening methods
- •10.24.1 Fluorescence-activated cell sorting (FACS)
- •10.24.3 Yeast surface display
- •10.24.4 Mass spectrometry-based methods
- •10.25 Protein engineering for nanotechnology
- •10.25.1 Protein-based nanocarriers
- •10.25.2 Biosensors
- •10.25.3 Protein nanowires and nanotubes
- •10.25.4 DNA–protein hybrid structures

Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
storage, examination, analysis and utilization of the information about biological
systems. For example, it involves processes such as accumulating genome sequences,
documentation of genes, assigning functions to the identified genes, planning of
databases, etc. Different genetic compilers are available today. A software compiler
is a way to use genetic information and provides a design platform allowing the
geneticist to manipulate and design everything from single genes to whole genomes.
So as to guarantee that the nucleotide sequence of a genetic material is complete and
error-free, the genetic material is sequenced more than once, such as by using the
shotgun method. The bacterial genome (Pseudornonas aeruginosa) was sequenced
seven times to make the sequence accurate and free from faults. Stover and coworkers (2010) reported the complete genome sequence of P. aeruginosa PAO1.
They suggest that the size and complexity of the P. aeruginosa genome reveals an
evolutionary adaptation allowing it to flourish in various environments and limit the
effects of a diversity of antimicrobial compounds [12]. However, the assembler
software (a sequence assembler is innovative bioinformatics software for spontaneous DNA sequence assembly, contig editing, file format alteration, DNA
sequence analysis and mutation recognition) identified 1604 regions that need
further justification. These areas were reinvestigated and re-sequenced to successfully complete the genome sequence. The efficiency/accuracy of the shotgun
technique was equated with the sequence derived from the clone-by-clone procedure
of two extensively separated genomic regions of P. aeruginosa. These two considered
areas together were 81 843 nucleotides long. The sequences derived by the two
procedures were in perfect agreement. This assessment exposed the precision
potential of the shotgun method of genome sequencing. This also demonstrates
the safeguards to be considered during genome sequencing projects. This level of
maintenance is not uncommon. Similar safety measures are employed in all genome
projects. The Human Genome Project, an international scientific research project,
sequenced the 3.2 billion bp of the human genome a total of 12 times. Celera
Genomics (a private biotech company), using shotgun-cloning, used an approach of
sequencing from both ends of DNA fragments. This project sequenced the human
genome 35.6 times. While a rough draft of the human genome sequence is complete,
some other objectives are yet to be accomplished. These include obtaining the
remaining sequence and rectifying miscalculations (called proofreading the genome),
filling in gaps and then sequencing the 7%–10% of the genome that comprise
heterochromatin. This is present in ample amounts inside the cells that are less active
or not active, whereas euchromatin is predominant inside cells that are vigorous or
active in the transcription of several of their genes. Since they comprise low priority
sections of repetitive DNA sequences, heterochromatic regions of the genome were
not initially sequenced. Moreover, intially it was thought that the heterochromatin
does not contain genes. However, the Drosophila genome sequence revealed that
their heterochromatic regions contain a very small number of genes (around 50).
Based on this discovery, the heterochromatic regions of the human genome must be
sequenced to confirm that all genes in the human genome have been recognized. As
soon as the genome of an organism is sequenced, compiled and proofread, the
subsequent process of genomics, namely annotation, begins.
7-8

Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
7.5 Understanding bioinformatics and sequencing
Advance computers and high-speed networks are compulsory to evaluate genome
projects. Genetic data signify a potential source for investigators and organizations
interested in how genes contribute to our health and wellbeing [11]. Nearly half of
the genes identified by the Human Genome Project have no known function.
Scientists, by means of bioinformatics, can identify genes, establish their roles and
develop gene-based strategies for preventing, diagnosing and treating disease [12].
Simplified presentation of bioinformatics and sequencing is presented in figure 7.3.
Genome sequencing and investigation is an area that has developed very quickly
over the last 10–20 years, in particular after publication of the final draft of the
human genome sequence in 2003. From procuring a complete genome sequence
from a characteristic individual or strain of a small number of species, the genomics
community has moved on to recording and authenticating genetic assortment within
species, with a special focus on human beings, and by sequencing the genomes of a
quickly growing number of species. The introduction of economical, very high
throughput sequencing tools has also made it promising to sample transcriptomes
(the summation of all mRNA fragments expressed from the genes of an organism),
genomic areas bound by proteins, bacterial populations, or sometimes even whole
ecosystems, by means of sequencing methods. This has germinated an innovative
generation of software techniques to manage the very large numbers of short
sequences, often called reads, which are produced by the new machines, and has also
moved the area of computational genomics into the territory of ‘big data’, requiring
Figure 7.3. Bioinformatics and sequencing.
7-9

Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
large and sophisticated computers for the organization and investigation of these
sequences.
The entities that are selected for genome projects are those typically used in
genetic and other systematic studies. Thus they are usually called model organisms.
A model organism can be a prokaryotic or eukaryotic micro-organism, or an animal
or plant, which is extensively investigated as it can be easily studied in the laboratory
to understand biological processes. The Generic Model Organism Database project
offers genetic researchers tools of open-source software machinery for imaging,
interpreting, managing and procuring biological information. Model organisms
include organisms such as E. coli, Archaeoglobus fulgidus, yeast (Saccharorny
cerevisiae), Bacillus subtilis, A. thaliana, the fruit fly(Droso melanogaster) and
nematide worm (C. elegans). The Human Genome Project focused on sequencing
the whole human genome.
E. coli is, by far, the most extensively studied micro-organism. Various tools that
are available currently were established for the E. coli genome project. The
sequencing of the E. coli genome was accomplished in 1997 [13]. The E. coli K-12
genome contains 4408 genes (4288 protein-coding genes annotated) with a size of
around 4.64 × 10
6
bp. B. subtislis is a gram-positive bacterium that settles over leaf
surfaces and is significant for both production of enzyme and food supply
fermentation. ‘GRAS’ is an acronym for the phrase Generally Recognized as
Safe. GRAS is well suited for enzymes, given the general availability of scientific
data supporting enzyme safety, and the generally recognized (peer-reviewed)
methodology and decision trees for evaluating the safety of microbial enzymes
used in food processing and in animal feed, respectively. It is usually named as safe
(GRAS) with genome of 4.21 × 10
6
bp (4212 genes). The reference database
SubtiList was exclusively created for the genome of B. subtilis 168, the paradigm of
gram-positive endospore-forming bacteria [14]. Another example of a model
organism is Aracheoglobus fulgidus, a strictly anaerobic archaebacterium with a
genome size of 2.17 × 106 bp (2493 genes) [15] (1 563 423 bp, 1858 protein-coding,
52 RNA genes, according to the Genomic Encyclopedia of Bacteria and Archaea
project) [16]. The yeast organism S. cerevisiae, so called because of its ability to
ferment saccharose (sugar), is the most key fungal species used in biotechnological
developments. In 1996 the S. cerevisiae genome was the first completely sequenced
from a eukaryote. The current version, called ‘S288C 2010,’ was determined from a
single yeast colony by means of advance sequencing tools and functions as the
anchor for further innovations in yeast genomic science [16]. In 1989, genome
sequencing of S. cerevisiae was accomplished. The yeast genome is 12.8 × 10
6
bp in
size and contains 6548 genes. PlantGDB (http://www.plantgdb.org/) is a databank
of molecular sequence data for all plant species with important sequencing efforts.
The database arranges EST sequences into contigs that represent tentative unique
genes. PlantGDB offers genome browsing abilities that integrate all accessible EST
and cDNA reported for existing gene models (for A. thaliana, see the AtGDB site at
http://www.plantgdb.org/AtGDB/)[16].
The first plant for which the complete genome was sequenced and published was
A. thaliana. In 1907, it was first considered as a good organism for genetic
7-10

Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
investigations and has become a model plant for genetic investigations as it has the
smallest genome of the higher plants. Furthermore, its genome contains only a small
amount of repetitive DNA [17]. The genomic sequencing project of the A. thaliana
genome began in 1990 and was completed in 2000. According to the sequencing
results, the size of the A. thaliana rabidopsis genome is 130 × 10
6
bp and a predicted
26 000 genes [18]. The MIPS A. thaliana Database (MAtDB; http://mips.gsf.de/proj/
thal/db) started out as a repository for genome sequence data in the European
Scientists Sequencing Arabidopsis (ESSA) project and the Arabidopsis Genome
Initiative. Another model organism, C. elegans (a free-living nematode) was
completely sequenced in 1999 and was the first multicellular animal whose genome
was sequenced. The C. elegans genome has 97 × 10
6
bp and has over 20 000
estimated genes. WormBase (http://www.wormbase.org) is a web-based resource for
the C. elegans genome and its biology. It is based on the existing ACeDB databank
of the C. elegans genome and offers data curation services, and a considerably
extended scope [19].
D. melanogaster (the fruit fly) is often called the ’Queen of Genetics’. In 2000, the
Drosophila genome was completely sequenced. It has 180 × 10
6
bp (16 000 genes).
This breakthrough has offered a platform for further genomic research, as the
number of genes present in the Drosophila genome is less than four times that of the
bacterium E. coli. Since the human genome draft sequence was published in 2001, a
number of observations have been made. Several significant features of the human
genome are as follows:
• A minimum 50% of the genome is obtained from transposable elements.
• The genome sequences of individuals vary by less than 0.2% of their base
pairs.
• It has approximately 35 000 genes.
• It encompasses over 3.2 billion base pairs.
• Only 5% of the genome encodes proteins.
• The genome contains gene-rich regions separated by gene poor regions, often
called gene deserts.
• The largest gene is the gene encoding dystrophin (2.5 × 10
6
bp).
The majority of the variations in the human genome are present in the form of
single-base differences in the sequence. A single-base difference is known as a singlenucleotide polymorphism (SNP), frequently called ‘snips’. One SNP is present in
roughly every 1000 bp of human genome. Approximately 85% of all variations in
human DNA is due to SNPs. Advantages of genome projects include [20]:
• Better knowledge of human genetic diseases and their correlations should
allow their management/cure.
• A number of methods have been established for genome sequencing projects.
• Information related to SNPs has become accessible; this may be beneficial in
different ways.
• The complete genetic information existing in the genomes of various
organisms can be determined.
7-11

Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
• They have unlocked exciting opportunities for future research, e.g. functional
genomics.
• They offer knowledge on why the individuals react differently to the same
drugs (pharmacogenomics).
• They offer understanding of genome association and its evolution, and the
most important mechanisms involved within that.
• Genome sequences allow the study of the many molecular interactions
resulting in the normal development of organisms.
• Micro-organism pathogenicity can be better understood. This can offer
protection from such diseases.
• The associations between genes can be studied with confidence.
7.6 Comparative genomics as a technique to understand evolution
The whole-genome sequence of an individual can be considered to be their final
genetic map, in the sense that the heritable physical characteristics are encoded
within the DNA and that the order of all the nucleotides along each chromosome is
known. However, information on the DNA sequence does not tell us directly how
these genetic data result in the observable traits and behaviors (phenotypes) that we
want to understand. Identifying all the functional parts of genome sequences and
using these data to improve the wellbeing of individuals and society are the focus of
the next phase of the Human Genome Project [21]. Relative studies of genome
sequences will be a major part of this effort. The investigation of variation and
similarity in genomic structures and, in particular, their organization in various
organisms is known as comparative genomics. The aims of comparative genomics
are as follows:
• to recognize the route of evolution and
• to translate the DNA sequence data into proteins of known functions.
In order to equate the genomes of various organisms it is essential to understand the
meaning of orthologues and paralogues. Orthologues are homologous genes present
in different organisms and only encode proteins that have a similar function.
Orthologues arise by direct vertical descent, and have deviated only by storing
mutations. Comparatively, paralogues are homologous genes within similar organisms and always encode proteins that have interrelated but not similar functions.
Paralogues have developed by gene duplication, followed by mutation accumulation, e.g. genes of the globin family.
7.6.1 The role of exon shuffling
Exon shuffling is a molecular tool for the development of new genes. It is a method
by which two or more exons from diverse genes can be taken together ectopically
(atypical gene expression in a cell, tissue, or at any developmental stage in which the
gene is not typically expressed), or the same exon can be duplicated, to create a new
exon–intron structure. Exon shuffling has been considered as one of the main
development powers determining both the genome and the proteome of eukaryotes.
7-12

Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Figure 7.4. Exon shuffling. In this illustration the end gene is produced by shuffling exons from more than one
gene.
This is vital in the formation of multidomain proteins throughout animal evolution,
obtaining a number of functional genetic novelties [22]. A simple presentation of
exon shuffling is depicted in figure 7.4.
Various proteins contain distinct domains. Such proteins are known as mosaic
proteins, e.g. serine proteases responsible for blood coagulation. The majority of
mosaic proteins are extracellular and are mainly present in metazoa. In addition,
these types of proteins are also present in unicellular organisms.
The study of the genes that encode mosaic proteins explores the strong
association between domain establishment and intron–exon structure. Each domain
is inclined to be encoded by one or a combination of exons, which are synthesized by
recombination within the intervening sequences; this is known as exon shuffling.
This process synthesizes new genes that encode proteins with improved function. As
introns are much longer than exons, the chances of crossovers in introns are greater
than in exons.
7.6.2 Horizontal or lateral gene transfer
Horizontal gene transfer was first reported in 1928, in an experiment by Frederick
Griffith. During this experiment, Griffith demonstrated that virulence was able to
transfer from virulent to nonvirulent strains of Streptococcus pneumoniae, establishing that genetic information can be horizontally transferred between bacteria via a
mechanism called transformation [23]. A schematic representation of horizontal and
vertical gene transfer is presented in figure 7.5.
Horizontal or lateral gene transfer involves genetic transfer/exchange between
organisms of different evolutionary origin/lineages (sequences of cells, genes,
organisms or populations, linked by a continuous line of origin from ancestor to
descendent) [24]. It is usually assumed that such gene transfers/exchanges have
happened many times throughout the course of evolution. The following two
approaches are employed for the discovery of those genes that are likely to develop
by horizontal gene transfer:
7-13

Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Figure 7.5. Horizontal and vertical gene transfer representation.
• Detection of genes having unusual base composition and bias in codon usage
can also be used for the detection.
• Failure to find a similar gene in closely related species.
The accessibility of entire genome sequences has allowed the detection of such genes.
It has been determined that hundreds (between 113 and 223) of human genes have
been transferred directly from bacteria into the human genome rather than having
developed from bacterial genes.
7.6.3 Genome similarity or homology
The sequence similarity searching tool is the most extensively used, and most
reliable, approach for describing newly determined sequences. This tool can identify
homologous proteins or genes by identifying a statistically significant similarity that
reveals common ancestry [25].
Extensively used similarity searching databases are:
• BLAST [27].
• PSI-BLAST [27].
• SSEARCH [30].
• FASTA [28].
• HMMER3 [29].
These programs offer precise statistical evaluations, guaranteeing protein sequences
that share considerable similarity also have similar structures. One of the amazing
discoveries from the investigation of genome sequences of various organisms is as
follows: organisms whose genomes are quite similar in appearance may look very
different, for instance humans and mice.
7-14

Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
These two organisms in fact have 97.5% of their DNA sequences in common that
achieve the same genetic functions. It was also observed that humans and mice had a
mutual ancestor some 100 million years ago. Since then, functional DNA has
deviated only to a small extent, whereas the non-coding DNAs of the two species
have deviated to a much greater extent. Likewise, it is estimated that human and
chimpanzee genomes differ by only 1%–3% of their DNA sequence. It may be
noticed that the organisms above are more closely related to each other than
organisms such as yeast (a fungus) and the nematode worm (an animal). The
nematode worm is estimated to have 19 000 genes, of which over 2000 encode
proteins that have functional equivalents as in yeast.
7.6.4 SNPs
For the last few decades, the utilization of molecular markers, highlighting polymorphism at the DNA level, has contributed a great deal in the development of animal
genetics [32]. Among all the techniques, microsatellite DNA markers have been the most
extensively utilized. A new marker tool, the SNP (see above) [32], is now in the picture
and has recently gained high recognition. It is just a single-base change in a DNA
sequence, with a usual alternative of two likely nucleotides at a certain position [26].
SNPs, the most abundant type of genetic variation, are now the main raw
material essential for most genetic investigations and databases. Variation comprising indels, copy number variants, microsatellites and epigenetic markers remain
parameters to consider and can influence disease [27]. There are some reported
10 million SNPs in the human genome. The basic concept and types of SNPs are
depicted in figures 7.6 and 7.7, respectively.
As mentioned above, SNPs are the most common type of genetic variant among
people. Each and every SNP signifies a variance in a single DNA building block,
known as a nucleotide. For example, an SNP can substitute the nucleotide cytosine
with the nucleotide thymine in a specific stretch of DNA. SNPs usually take place
throughout a person’s DNA. They occur once in every 300 nucleotides on average,
i.e., there are approximately 10 million SNPs present in the human genome. As a
rule, these deviations are present in the DNA between genes. They can behave as
biological markers, facilitating researchers to locate genes that are related to disease.
Once SNPs occur inside a gene or in a promoter region closer to a gene, they might
show an additional direct role in disease by influencing the gene’s function. The
majority of SNPs have no influence on health or development. Certain genetic
differences, however, have been shown to affect human health. Scientists have
discovered SNPs that may help to predict an individual’s reaction to certain drugs,
vulnerability against external environmental factors such as toxins and the threat of
developing specific illnesses. SNPs may also be utilized to track the causes of disease,
in particular the inheritance of disease genes inside families. Upcoming investigations will work to recognize SNPs linked to complex diseases, e.g. heart disease,
diabetes and cancer. Various academic institutions and industries are currently
involved in the development of large-scale SNP detection projects, and these projects
introduced data on thousands of publicly available SNPs. The most common
7-15

Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Figure 7.6. Representation of SNPs.
Figure 7.7. Types of SNPs.
example of these is the HGBASE database. It has been projected that 90% of
sequence variation in humans is because of SNPs. Human genome is projected to
contain 3–17 million SNPs. Out of these, 5% SNPs are likely to occur in genes. Thus
each individual gene is likely to contain six SNPs. Therefore, SNPs offer a molecular
marker that is present in the genome at a very high density. The SNPs can thus be
utilized to map genes involved in human diseases. These genes are exceptionally
difficult to map using other molecular markers as these latter markers are not
distributed at an adequate density in the genome. Therefore, by using SNPs as
markers, every gene in the human genome can be mapped. Additionally, it might be
possible to identify all the genes involved in human diseases. An SNP might be
present within the relevant gene or very near the gene which is tightly linked with the
SNP. SNP1 and SNP2 are the two alleles of an SNP.
Some SNPs are situated in the recognition sequence of an endonuclease. They
alter the recognition sequence (base sequence) and produce restriction fragment
7-16

Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
length polymorphism (RFLP). Some other SNPs are situated in the coding regions
of genes. They are linked with definite phenotypes. The differences in the genetic
sequences of individuals are understood to be responsible for the following:
• Disease susceptibility.
• Drug response.
• Normal development and aging.
• Reaction against environmental factors.
The field of study that is more concerned with the effect of genetic variation disease
susceptibility and drug response is known as pharmacogenomics. Pharmacogenomics
is the investigation of how genetic make-up regulates the response to a therapeutic
intervention. Pharmacogenomics is likely to advance individualized treatments that
will be safer and more effective [28]. Genetic variations may affect the drug reaction of
an individual in the following ways:
• They may influence the metabolism of the drug itself.
• They may affect the action of the drug on its target molecule.
Pharmacogenomics aims to raise drug efficacy and safety by comparing drug properties to the genetic make-up of an individual. The response of the drug is expected to be
influenced by various genes associated with drug metabolism, drug transport and
creation of drug targets. Thus, various inputs have been made to employ SNPs to
further map hundreds or thousands of genes that have an influence on the safety and
efficiency of different drug treatments. These SNPs can be positioned on gene chips
that can be employed to explore the genotype of those genes that are involved in
responses to specific drugs. SNPs have the potential to describe the danger of an
individual’s susceptibility against numerous illnesses and reaction against various
drugs. Biologically significant SNPs that are specifically related with the threat of
disease need to be identified [29]. The identification and understanding of great
numbers of these SNPs are essential before they can be used widely as genetic tools.
Depending upon these data, more effective and safe treatment can be considered for
these individuals, such as variants in the gene encoding apolipoprotein E (apoE),
which were found to be associated with variation in the reaction to a drug used for
treatment. apoE supports lipid transport and binds to cell surface receptors to facilitate
lipoprotein uptake. It is the main apolipoprotein produced in the brain [37]. The apoE
protein is present in three different isoforms (E2, E3 and E4) that are the outcome of
two non-synonymous SNPs (rs429358 and rs7412) present in exon 4 of the APOE
gene. Gene apoE is involved in vulnerability to Alzheimer’s disease, and a cholines-
terase inhibitor is employed to mitigate the symptoms of illness [30].
7.6.5 Inferences from comparative genomics
In the case of prokaryotes, the following genomes and characteristics have been
considered:
• Bacillus megaterium is important in the Bacillus phylogeny, as an evolutio-
narily important species and, in particular, in understanding genome
7-17
Соседние файлы в папке Библиотека им академика М.И. Перельмана
