Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5586_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
35 Мб
Скачать
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Figure 8.1. Basic protocol for proteomics.

8.2 Types of proteomics

8.2.1 Structural proteomics
In the postgenomic era, structural proteomics is one of the promising areas of investigation, and explains the structure–function relationships of uncharacterized gene products based on the three-dimensional protein structure (gure 8.4)[4]. It suggests the biochemical and cellular roles of unannotated proteins and thus classies possible drug design and protein engineering targets [8]. Recently, several innovative groups in structural proteomics research have attained proof of structural proteomic theory by forecasting the three-dimensional structures of theoretical proteins that correctly recognized the biological functions of those proteins.
Structural proteomics can show the three-dimensional structure and nature of protein complexes present in a specic cell/organelle [5]. The nal goal of structural proteomics is to form an organization of structural data that will assist in predicting
8-2
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Target identification and selection
Target isolation and purification
Structure determination
Analyze structure for potential ligand binding
Docking of small molecules using
Biochemical assays and further testing
Lead optimization to improve potency
Cytotoxicity tests, pharmacokinetics studies &
toxicological investigations
sites
computational methods
Drug candidate
Figure 8.2. Basic procedure involved in proteomics.
Figure 8.3. Different areas in proteomics.
8-3
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Figure 8.4. Basic procedure involved in structural proteomics.
the possible structure and possible function of virtually any protein from informa­tion of its coding sequence. This proteomics branch can also help in accumulating data about protein–protein interactions and about the construction of cells, to describe how the expression of certain proteins results in a cell s unique features. Recent mass spectrometry tools have offered a useful platform that can be combined with several procedures to examine protein structure and dynamics. Moreover nuclear magnetic resonance (NMR)-based structural proteomics coupled with x-ray crystallography can provide a comprehensive structural database to predict the basic biological functions of hypothetical proteins identied by genome projects.
8.2.2 Functional proteomics (strategy)
As the number of genome sequencing projects is increasing, there is a parallel exponential growth in the number of protein sequences whose function is still unidentied (gure 8.5)[6]. Functional proteomics is a developing research area in the eld of proteomics whose methods are geared towards two main objectives: the interpretation of the biological function of unidentied proteins and the description of cellular mechanisms at the molecular level [7].
Functional proteomics uses proteomics tools to examine the features of the molecular protein networks involved in a living cell. Documentation and inves­tigation of molecular protein networks involved in the nuclear pore complex in yeast is one of the recent accomplishments of functional proteomics. This accomplishment allows understanding of the translocation of molecules from the nucleus to the cytoplasm and vice versa [8].
8.2.3 Expression proteomics
Expression proteomics is the investigation of dysregulated proteins as a function of stimulation or condition (disease, time, drug, etc). In other words it is the quantitative investigation of protein expression between samples differing in a number of variables. It can also be dened as the examination of protein expression at a larger scale. The arrangement of expression of the entire proteome or a portion of it (a subproteome) between samples can be ascertained with the help of this
8-4
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Figure 8.5. Functional proteomic analysis.
method. Expression proteomics is quite benecial in nding disease-specic proteins. For example, over-expression or under-expression of proteins in cancerous cells and normal cells derived from a cancer patient and a normal individual, respectively, can be examined using several tools such as 2DE, mass spectrometry, microarrays, etc. This approach can recognize the growth of cancers and allows improvement of drugs in the management of cancer.

8.3 Basic techniques involved in proteomics

8.3.1 Sequence alignment (algorithms)
Fast emerging new sequencing tools offer information on an unparalleled scale. A principal challenge to the examination of these data is sequence alignment, wherein sequence reads must be matched to a reference [9]. An extensive variety of alignment algorithms and software have been developed over the past few years [9]. These techniques involve the study of the sequences of the DNA, RNA or protein to detect sites of resemblance that may be a result of functional, structural or evolutionary relationships between the sequences. Nucleotide or amino acid residues which are aligned sequences are characteristically presented as rows within a matrix. The residues gaps are incorporated between these, so that identical or similar features are associated in successive columns. The fastest way to allocate a new gene or cDNA is to perform pairwise assessment (a procedure of relating objects in pairs to judge which object is favored) with cDNA and EST sequences, and recognize gene sequences procured in databanks using alignment techniques such as BLAST (Basic Local Alignment Search Tool) or FASTA (a DNA and protein sequence alignment software package). However, this method is really only benecial when there is >30% sequence uniqueness between the new gene and the databank sequence. If a similar gene/protein is found, it proposes the possible role of the
8-5
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
product of new gene. However, additional experiments will be essential to determine the real in vivo function of the new gene product. To nd resemblances below 30% sequence uniqueness, numerous sequence assessments are made by means of an algorithm such as PSI-BLAST (Position-Specic Iterated Blast). However, such a search may provide false positive ndings. PSI-BLAST develops a position-specic scoring matrix or outline from the various sequence alignments detected above a given score threshold by protein–protein BLAST [10].
8.3.2 Protein structure (annotation resources)
An important reason three-dimensional protein structures are interpreted with supportive or derived data is to determine the molecular background of the protein function. In this endeavor, protein structure annotation databanks curate important evidence and explanations, based on community-accepted standards, for the 100 000 three-dimensional investigational protein structures that allow further understanding of the structure–function relationship [11].
A method to control the function of an orphan gene is to inspect the three­dimensional structural arrangements of the protein, that the gene responsible for this protein is expected to encode; such a protein is called a theoretical protein. This three­dimensional structural arrangement is matched with the protein structures stored in databases. One of the bases for this method is the fact that the three-dimensional structures of proteins are better maintained than are their primary structures. For this reason protein function depends on the three-dimensional structure instead of the primary structure of the protein. It has been projected that 20%–30% of orphan genes could be allocated function by dening the three-dimensional structures of the proteins encoded by them, and equating these to protein structures in databanks. For example, the three-dimensional structural congurations of hemoglobin and myoglobin are very comparable, considering that they both are oxygen carriers. However, a BLAST search will fail to detect resemblance between myoglobin and α­or β-globin (both are hemoglobin polypeptides). Thus it has been projected that if a theoretical protein has a similar structure to a known protein, there is a 66% chance that it has a role related to that of the known protein. The structure of a theoretical protein can be evaluated in the following three ways:
If this target sequence displays >25% resemblance to a sequence whose structure is already known (called a template sequence), relative modeling can be employed to forecast the structure of the target protein sequences.
When the uniqueness is <25%, the sequence can be matched to known structures to examine the level to which the experimental data match the values anticipated by theory.
The target sequence structure can also be examined from the primary data without reference to known structures. Protein structures are experimentally evaluated using x-ray crystallography and NMR spectroscopy. Both of these methods are slow, time consuming and require greater quantities (milligrams) of the puried protein. However, current developments in both procedures will allow their alteration for high throughput analyses.
8-6
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Figure 8.6. Procedure involved in protein structural investigation.
8.3.3 Protein structural investigation
Figure 8.6 provides a schematic representation of the process involved in protein structural investigation.
Structural investigation can often disclose the general function of a protein when sequence and structural assessments fail to suggest function. For example, protein surface scanning may disclose splits that may signify ligand-binding sites. A protein with a large split or cleft could be cautiously assigned the name enzyme. A greater resolution in the description can be conceivable by comparing the shape of the split/ cleft to a library of small molecular shapes using drug design software, e.g. DOCK and HOOK. These programs are capable of identifying possible ligands with binding sites present on protein surfaces. Scanning of the protein surface also facilitates identification of the domains most likely involved in interaction with other proteins. Knowing gene function involves more than recognizing the gene products. After they are produced, several gene products are altered by the breakdown of end groups (e.g. signal sequences, propeptides or initiator methionine residues), by the accumulation of chemical groups (such as methyl-, acetyl-, phosphoryl) or sometimes by adding linkages to sugars and lipids. In addition to the wide variety produced by the alternate splicing of mRNA, nearly a hundred mechanisms of post-translational modication are recognized. Therefore, the human genome, which may have 35 000–40 000 protein-coding genes, may offer over 350 000 different gene products. The main aim of proteomics is to provide data on each protein encoded in a genome, on its role, structure, cellular localization, post-translational modications, associations (shared domains, evolutionary history) with other proteins and variants.
8.3.4 Two-dimensional gel electrophoresis in proteomics
Gel-based proteomics is one of the most multipurpose approaches for fractionating protein complexes [12]. Among these approaches, two-dimensional polyacrylamide
8-7
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Figure 8.7. Two-dimensional gel electrophoresis.
gel electrophoresis (2DE) provides a key orthogonal method. It is commonly used to simultaneously fractionate, detect, and quantify proteins when coupled with mass spectrometric identication or other immunological tests.
The basic tools in proteomics include separation and identication of proteins that are isolated from cells. To achieve this, 2DE can be utilized. 2DE involves placement of the protein extract on a polyacrylamide gel after which an electric charge is applied across the gel to separate proteins based on their molecular weight (gure 8.7). Once this is complete, the gel is turned 90°, and through a second phase of electrophoresis, the proteins are separated in a second dimension based on their molecular mass. Once the gels are stained, the proteins are exposed as spots; different gels display 200–10 000 spots. To detect different proteins, spots are cut from the gel and are allowed to digest with enzymes, e.g. trypsin, to offer a characteristic set of fragments. These fragments are further studied by mass spectrometry, which is known as peptide mass ngerprinting. To recognize the proteins, the peptides mass is matched with the masses predicted from data in genetic or protein databanks.
8.3.5 Domain fusion method (or rosetta stone method)
The examination of amino acid sequences from different individuals often provides results in which two or more proteins encoded for discretely in a genome also act as fusions, either in a similar genome or that of some other organism [13]. These fusion proteins, called Rosetta stone sequences, support the linkage of dissimilar proteins, and suggest the probability of functional interactions between the linked objects, creating local and global relationships within the proteome [14]. The domain fusion method searches for functionally active proteins that are distinct in a number of organisms, but are fused into a single protein, a Rosetta stone protein, in some other organisms. The basic principle of the Rosetta stone technique is as follows. The protein A sequence from an organism is utilized to search for a homolog in a different organism. This examination can recognize a single domain protein such as
8-8
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
A. However, in a number of organisms, it may recognize a Rosetta stone protein, such as A–B. As domains A and B are portions of a single protein, they are likely to be functionally related. Now the sequence of domain B is used to search for its homolog in the rst organism. The homolog of the domain identied in this organism will be functionally related to protein A. Thus if the function of protein A were known, this strategy will reveal the function of protein/gene B, which was formerly an orphan gene.

8.4 Complete proteome of Mycoplasma genitalium

M. genitalium is the smallest member of the Mollicutes, has a genome size of 580 kb and the potential to express 480 gene products, and is thus an outstanding model to assess [15]:
The minimum metabolism necessary for a free-living cell.
Proteomic tools and the information derived by proteome analysis.
Wasinger and co-workers utilized proteomics to offer a portrait of what type of genes are expressed in the bacterium M. genitalium during exponential and sta­tionary growth periods. M. genitalium is one of the simplest independent bacteria with a reduced genome. Using 2DE, the researchers explored 427 protein spots in exponentially growing cells. Of these, 201 were examined and recognized. The rest of the spots present in fragments derived from larger proteins, variants of similar proteins (isoforms) and post-translationally modified forms. The recognized proteins included enzymes involved in DNA replication, transcription, translation, delivery of materials across the cell membrane and energy metabolism. The study, however, could not cover 158 known proteins (33% of the proteome) and 17 unknown proteins. It was analyzed that there was a 42% reduction in the number of proteins synthesized during the transition from the exponential growth phase to the sta­tionary phase. Moreover, a number of new proteins appeared, and additional proteins experienced intense variations because of nutrient exhaustion, enhanced acidity of the growth medium and other adaptations against environmental changes. In M. genitalium, a mere 33% of the proteome is expressed during maximal growth, and the remaining 67% of the proteome are the proteins to be expressed under different environmental conditions. Wasingers examination of the M. genitalium proteome assisted in establishing the least number of expressed genes essential for independent existence and the modications in gene expression that accompany the changeover from the exponential growth phase to the stationary phase. It was also noticed that the extensive range of information offered by proteome analysis cannot be derived only by genome sequencing.

8.5 Architecture and design of the nuclear pore complex

Three-dimensional examination of the nuclear pore complex discloses the funda­mental, highly symmetric framework of this supramolecular assembly, how it is attached in the nuclear membrane and how it is constructed from many distinct, interconnected subunits [16]. The organization of the subunits within the membrane
8-9
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
pore makes a large central channel, by which active nucleocytoplasmic transport is known to occur, and eight smaller peripheral channels that are probable routes for passive diffusion of ions and small molecules. The nuclear pore complex (NPC) is made up of proteins embedded in the nuclear membrane. This complex joins the cytoplasmic and nuclear compartments of the cell and permits transport of materials between the nucleus and the cytoplasm. As this transportation contains mRNA intended for the cytoplasm, nuclear pores are vital control points for regulating gene expression. Their structural and functional sophistication has been an obstacle in understanding their molecular organization. Genomics in association with proteo­mics has offered scientists techniques to investigate the three-dimensional construc­tion of nuclear pores and to recognize their proteins. Depending on its size, the yeast nuclear pore complex contains around 200 different proteins. By using a rat genome databank as a resource, scientists were able to recognize potential nuclear pore complex proteins, also called nucleoporins, by means of sequence resemblance to already identied nuclear pore complex proteins from other organisms. In another example, proteomics was employed to outline the molecular construction of a nuclear pore complex. Accordingly, pore complexes were separated and puried, and then nucleoporins were detected by means of mass spectrometry on peptide digests of proteins puried from the NPCs. Yeasts NPC only includes around 30 different proteins, but with each protein present in multiple copies. The nucleoporins are prearranged into 16 subunits (eight on the pores nuclear side and eight on its cytoplasmic side). Moreover, there are lament-base structures on both sides of the pore, formed into a bag on the nuclear side. Using multiple tools such as mutant analysis, microscopy and proteomics, the actual regions/sites of NPC proteins in the three-dimensional arrangement of the pore have been mapped. The majority of nucleoporins are proportionally distributed on the cytoplasmic and nuclear sides of the pore. Five proteins are present only on one side or the other, and seven are distributed more on one side than on the other. The function of nucleoporins in nucleocytoplasmic transportation or the regulation of transport can be studied by examining protein–protein interactions within the pore, as well as interactions of pore proteins with transport proteins, signal molecules and other cellular components.

8.6 Functional genomics and systems biology

Functional genomics is an efcient method for determining the roles of the novel genes revealed by complete genome sequences (gure 8.8)[17]. Such a method should take a hierarchical approach, as this will both bound the number of trials to be done and allow a closer and closer estimate of the function of any individual gene to be attained. Furthermore, hierarchical studies have, in their early stages, remarkable integrative power and functional genomics aims at a comprehensive and integrative view of the workings of living cells [18]. Functional genomics is an area where the function of a gene product is determined. This includes answers to the questions:
8-10
Introduction to Pharmaceutical Biotechnology, Volume 2 (Second Edition)
Figure 8.8. Schematic representation of functional genomics.
How are genes expressed?
How are the gene sequence and structure related to the end product, and its
relationship with gene versus an individual from the same generation?
Exactly how is a product associated with a sequence and structure, and how is it associated with the end products of other genes of similar organisms?
What are the implications of the environment over a gene and how does this interaction help in understanding the rate of genetic variation? Gene– environment interaction helps in understanding genetic inuences modeled as latent [19].
These questions can be answered by studying the following:
Determining at what time and exactly at which place specic genes are expressed on a genomic scale (called expression proling). This permits the identication of those particular genes that are over-expressed or under­expressed. This concept can be used to understand different biological processes in health and disease. This concept can also be utilized to study the expression in a particular tissue/organ at the molecular level. This type of gene-based proling, and more specically recent RNA sequence based proling, can be used as an important tool for drawing correlations between gene activity and various physiological or developmental states. It also helps in creating collections of gene expression data across various cell types, development times, varied species and against different stimuli [20].
The replacement of a particular natural gene with mutated gene under in vitro conditions to access its in vivo function (the knockout method) [21].
The genetic relationship between proteins and other molecules (as rst suggested by Archibald Garrod in 1902, who predicted that genes control the function of proteins during his study on patients suffering from alkaptonuria).
Functional genomics aims towards answering such questions thoroughly for most of the genes present in a genome, in comparison to classical strategies that report one
8-11