Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5319_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
29.08.2026
Размер:
92 Мб
Скачать
Gene Technology in Drug
Research
Contents
12.1 The History and Basics of Gene Technology – 170
12.2 Gene Technology: AKey Technology in Drug Design – 171
12.3 Genome Projects Decipher Biological Constructions – 173
12.4 What Is Contained in the Biological Space of the Human Proteome? – 174
12.5 Knock in, Knock out: Validation of Therapeutic Concepts – 176


12.6 Recombinant Proteins for Molecular Test Systems – 177
12.7 Silencing Genes by RNA Interference – 178
12.8 PROTAC: How to force therapeutically untargetable proteins into targeted degradation – 179
12.9 Proteomics and Metabolomics – 180
12.10 Expression Patterns on aChip: Microarray Technology – 182
12.11 SNPs and Polymorphism: What Makes Us Dierent – 183
12.12 The Personal Genome: Access to an Individualized Therapy? – 184
12.13 When Genetic Dierences Turn into Disease – 184
12.14 Epigenetics: Lifestyle and Environment Inuence Gene Activity Like aPen Leaves aMark in the Book of Life – 185
12.15 The Scope and Limitations of Gene Therapy – 187
12.16 Synopsis – 189
Bibliography and Further Reading – 190
© The Author(s), under exclusive license to Springer-Verlag GmbH, DE, part of Springer Nature 2024 G. Klebe, Drug Design, https://doi.org/10.1007/978-3-662-68998-1_12
Chapter  • Gene Technology in Drug Research
12
Engineers and writers have predicted many developments in science and technology. Among other sophisticated machines, Leonardo da Vinci described the principle of the helicopter. In the early 1820s, Charles Babbage designed an automatic calculating machine that was far ahead of its time. More than 160 years later, the mechan­ical precursor of aprogrammable computer was actually built, and it worked! Jules Verne described submarines and atrip to the moon, and Hans Dominik described creating energy by splitting the atom. All these visions became reality. Only one application of gene technology, the most groundbreaking invention of our time, was fore­seen: the cloning of two genetically identical individuals in Aldous Huxley’s Brave New World in 1932. Mamma- lian cloning has already been carried out. When Dolly, the rst cloned sheep, was born in Scotland on July5, 1996, it was hailed as abreakthrough in genetic research. Cloned mice, cattle, pigs, horses and, most recently, mon­keys have followed. It is to be hoped that researchers will respect ethical boundaries and not make use of Huxley’s idea of human cloning, despite its tangible feasibility.
Genetic engineering makes it possible to introduce new genes into acell, multiply them, and exchange or remove them. When removed or altered, the cell can no longer produce the original protein derived from that gene. With the introduction of anew gene and aclever choice of method, the cell produces aforeign product, either an intentionally modied protein or acompletely new one. For many diseases, the molecular cause is known to be the absence of aprotein or agenetically caused mutation in aprotein. These are only afew of the more common examples:
Diabetes as aresult of insulin deciency,
-
Particular, hereditary cancer forms (e.g., familial co-
-
lon cancer, malignant melanoma),
Chorea Huntington, achronic form of brain atrophy,
-
Sickle cell anemia, agenetic disease producing mal-
-
formed red blood cells (Sect.12.13), and
Bleeding disorders that are caused by the absence of
-
particular coagulation factors (see Sect.12.13).
The possibility of purposefully producing arbitrary pro­teins has yielded the following main applications of gene
technology:
The identication of genes and proteins that could
-
play arole in the treatment of adisease,
The development of animal models to test athera-
-
peutic principle,
The production of proteins for therapies in which
-
aparticular protein is missing,
The manufacture of monoclonal antibodies and vac-
-
cines,
The manufacture of proteins for molecular test sys-
-
tems, and the determination of the 3D structures of
enzymes and receptor proteins,
The generation of proteins of which atargeted muta-
-
genesis has been undertaken to exchange one or more amino acids for the elucidation of the mode of action of enzymes and for the characterization of receptor binding sites,
Somatic individual gene therapy for specic patients,
-
and
Immunotherapy using genetically engineered T-cells
-
with antigen-specic receptors.
Other application possibilities, for example, the manipu­lation of the human germline, or genetic changes in crops to achieve herbicide resistance, or to prolong the shelf life of fruits, are only briey mentioned here.
12.1 The History and Basics of Gene
Technology
The foundations of gene technology were rst established in the middle of the twentieth century. It all started in 1953 after James Watson and Francis Crick in Cambridge (England) became aware of Rosalind Franklin’s X-ray data; they correctly interpreted the available knowledge about deoxyribonucleic acid (DNA) and proposed the structure of adouble helix as the three-dimensional geo­metry for the genetic material of all living beings. Im­mediate indications were obtained from the structure about the mechanism of our hereditary transfer and about the genetic code for the biosynthesis of proteins. Afew years later, Werner Arber discovered enzymes that attack specic sites on the double helix and cleave DNA in asequence-specic manner. These enzymes are called restriction enzymes. What was initially seen as acuriosity proved to be an exceedingly important discovery for gene technology. It is possible to selectively cleave DNA with these enzymes and to introduce new fragments. Next, the merging of new information with the original DNA, the recombination of the genetic constitution, is accom­plished with ligases from special viruses called bacterio­phages. The techniques for DNA sequencing have also made decisive progress. Soon afterwards, the amino acid sequence of aprotein was no longer directly determined, but rather deduced from the analysis of the correspond­ing DNA. Today, sequencing is mainly done via cDNA, which is complementary to the RNA (Sect.12.6).
In 1973, Stanley Cohen and Herbert Boyer managed to recombine the genome of abacterium for the rst time (. Fig.12.1). Then, things happened one after the other: two years later the bacterial strain Escherichia coli K12, which is still used today, was developed. Part of its genetic constitution is missing, making it viable only under laboratory conditions. This bacterium can be ge­netically manipulated at will without fear of harm. The British scientists H.Williams-Smith and E.S. Anderson
. • Gene Technology: AKey Technology in Drug Design


. Fig. 12.1 The principle of gene technological recombination of
hereditary information. Bacteria often contain additional genetic ma­terial in addition to their “chromosome” in the form of ring-shaped plasmids; these are used in gene technology as vectors to introduce foreign genes. Plasmids are removed from the cell and sequence-spe­cically cut with so-called restriction enzymes, which are isolated from bacteria. The target DNA that carries the desired gene, which was typically also treated with the same restriction enzyme, is bound to
carried out independent experiments in which they took Escherichia coli K12 orally. They demonstrated that these bacteria survive only for ashort time in the gastrointesti­nal tract and that the K12 gene, which confers antibiotic resistance for selection of the transformed cells, cannot be transferred to normal Escherichia coli found in the intestinal ora. At aconference in Asilomar, California, experts discussed the potential dangers of genetic engi­neering and dened various risk and safety classes. Ge­nentech was founded in 1976. Its founder, Herbert Boyer, had to borrow US$ 500 as start-up capital! When the company went public in 1980, the value of his shares made him amillionaire within minutes. In 1982, Genen­tech introduced the rst drug to be produced using gene technology: human insulin (Humulin®).
In 1983, Kary Mullis made aseminal contribution to gene technology when he developed the polymerase chain reaction (PCR) while working at Cetus, aCalifor­nian company founded in 1971. Double-stranded DNA is melted into its single strands by heating, then two short pieces of single-stranded DNA complementary to the regions at the beginning of the DNA, called primers, are added along with the four DNA nucleotides. A poly­merase can now be used to synthesize new DNA in atest tube. This means that anew double strand is formed by starting with the primers (. Fig.12.2). A heat-stable DNA polymerase (originally from the bacterium Ther- mus aquaticus, endemic to the hot springs of Yellowstone National Park) is used for DNA synthesis. Each repeti­tion of this step doubles the amount of DNA. Within afew hours, billions and trillions of DNA molecules can be produced from asingle starting molecule. This
the overlapping single-stranded DNA ends in vitro. The DNA ends are coupled with the enzyme DNA ligase, and the modied, recom­binant plasmid is brought into the bacterial cell. In addition to the DNA segment that is necessary for replication, plasmid vectors that are used in gene technology carry additional information that allows for the recognition and selection of the transformed cells (usually an antibiotic-resistance gene). In the presence of the selecting agent, only plasmid-containing cells grow
amount is sufcient to sequence the DNA segment of interest.
PCR methods are applied diversely. The entire genetic information of an individual can be derived from asin­gle DNA molecule. In medical diagnostics, this serves to provide evidence regarding genetic disorders, cancer, infectious diseases, and risk factors. PCR methods are also used to establish agenetic ngerprint in paternity tests and in forensic science.
New genetic information can be introduced not only into bacterial cells, but also into yeast, virus-infected in­sect cells, and even mammalian cells. However, as arst approximation, the more complex the organism, from bacteria to mammalian cells, the more difcult it is to produce proteins in these cells. On the other hand, insect and mammalian cells have the advantage of producing not only ordinary proteins, but also more complex ones (e.g., glycosylated proteins) in afunctional form. Thus, in many cases, we are dependent on such organisms for the production of proteins.
12.2 Gene Technology: AKey Technology
in Drug Design
The 1970s and 1980s were the golden age of receptor binding assays with membrane preparations. Radioac­tively labeled ligands were used to determine the specic binding of new drug candidates. The major receptors for hormones and neurotransmitters were known, and in some cases the difference between pre- and postsyn­aptic receptors. However, the different subtypes and
Chapter  • Gene Technology in Drug Research
12
. Fig. 12.2 The polymerase chain reaction (PCR) can make an un-
limited number of identical copies of aDNA molecule. The DNA is heated to split the double-stranded DNA into complementary single strands. Synthetic oligonucleotides of about 20bases, called primers, that are complementary to these DNA strands hybridize to the cor­responding strand. Each primer must bind to one end of each strand of DNA. The primers dene the boundaries of the amplied DNA. In addition, an excess of primers must be used because one pair of primers is required for each double strand of DNA in each cycle. The primers are needed to synthesize the new DNA in the presence of the DNA polymerase and an excess of the four different nucleotides. This
their amino acid sequences were not known. As aconse­quence, the results of these studies were quite imprecise.
Gene technology methods allow the production of homogeneous recombinant proteins in virtually unlimited quantities. They play an important role in the very rst step of drug design: the identication of atarget protein. Advances in methodology have led to the discovery of new receptors, some of which had unknown function or specicity. The next steps are to test the therapeutic con- cept in genetically modied animals. Another important contribution is the preparation of proteins for molecular test systems and the isolation of adequate material for the elucidation of the 3D protein structure (Chap.13). With, perhaps, the exception of afew proteins that can be iso­lated from blood or other natural sources, the production of large quantities of proteins is dependent on gene tech­nology. Today, the purication of proteins from animal or human blood is done rather reluctantly. The risk of transmitting viruses or infections is deemed to be too high.
Gene technology offers the possibility to selectively produce structural variants of proteins. The generation of point mutations (site-directed mutagenesis) allows partic-
occurs in the reverse direction (dashed arrow) due to the opposite ori- entation of the DNA strands and the specicity of the polymerase. The newly synthesized DNA segment can be several hundred to sever­al thousand base pairs long. The result is two identical double-strand­ed DNA molecules. After heating, single strands are obtained and the above procedure is repeated. Because DNA polymerase is heat stable, it does not need to be added repeatedly. Each repetition of the above steps results in adoubling of the DNA molecule. Its number increases exponentially. Ten cycles lead to about 1000 DNA molecules, 20to amillion, and 30to abillion. In this way, asingle DNA molecule can be multiplied into abiochemically analyzable quantity
ular properties in proteins to be improved, and the bind­ing and catalytic properties of enzymes to be purposefully changed. Membrane-bound receptors can be probed posi­tion by position to establish which amino acids are respon­sible for the maintenance and stability of the 3D structure, the adoption of aparticular conformation, or are of crit­ical importance with respect to the binding of aligand. Three-dimensional structural models of receptors can be generated in this way, or their relevance can be appraised.
In many cases, it has also proven worthwhile to in­troduce point mutations that change the surface proper­ties of proteins and help to elucidate the 3D structure of proteins. Sometimes the charge on individual amino acids or highly exible amino acids that tend to cause disorder must be exchanged for the sake of protein crystallization. For proteins in which part of the sequence is membrane anchored, the membrane anchor, which would interfere with crystallization, is removed prior to the crystallization experiment. For soluble receptors, it has been found useful to remove individual domains, crystallize them, and de­termine their structure. Of course, such modied proteins must still perform their specic functions, such as ligand
. • Genome Projects Decipher Biological Constructions


binding or DNA docking. Once the difcult crystalliza­tion step has been accomplished, the actual structure de­termination nowadays only takes amatter of afew weeks in most cases (Chap.13). Membrane-bound proteins can be stabilized in their structure by point mutations so that they can be detached from the membrane. They can then be crystallized in special media, such as cubic lipid phases.
When considering the benets of all this progress for humanity, we cannot avoid asking the question: where do the fears of broad sections of society about gene technol-
ogy come from? It is not difcult to understand these reser­vations: with the use of gene technology, almost everything that is theoretically conceivable in the eld of genetics be­comes possible. However, people’s trust in science is not as unshakable as it used to be before the atomic bomb. Now that the opportunities outweigh the risks, the sins of our forefathers have come back to haunt us. Too often in the past, scientists have underestimated potential risks and put their ethical concerns on hold. Scientists have still not managed to allay the public’s fears. We must take these fears seriously and rebuild trust by acting responsibly.
12.3 Genome Projects Decipher Biological
Constructions
The entire human genome is organized on 23 chro­mosomes. In 1990, the Human Genome Organization (HUGO), with abudget of US$ 3 billion, began the then-ambitious task of sequencing the entire human genetic code from DNA within 15years. By the end of 1993, the rst annotated genome maps were available, and these were later rened. By 2001, the project had progressed to the point where the entire genome was pub­lished in Science and Nature by two parallel consortia.
The two competing consortia followed different strat-
egies. The publicly funded international consortium chose the approach of setting progressively narrower parameters, the stepwise digestion of the genome, and the systematic elucidation of sequences for the complete genome analyses. In humans, this means that in addition to the 5% of DNA that corresponds to genes, the other 95% of sequenced DNA, the function of which was unknown, was classi­ed with the somewhat derogatory term “junk DNA.” It is now established that these regions play a pivotal role in the regulation of gene expression and host to a class of small non-coding RNA molecules, collectively termed microRNA. In 2024, Victor Ambros (UMass Chen Med­ical School, Worcester, USA) and Gary Ruvkun (Harvard Medical School, Boston, USA) were awarded the Nobel Prize in Medicine for their discovery of microRNA and the elucidation of their function in post-transcriptional gene regulation, particularly in the context of organismal development and function (Sect.12.7). The second strat­egy, pursued by the privately funded consortium, used the shotgun approach. This was done by amplifying alonger
strand of DNA and then cutting it into many small seg ments. After these segments were sequenced, the sequences were reconstructed into the original long strand of DNA using apowerful computer program. Of course, this can only work if the sequences of the cleaved segments overlap sufciently. This technique proved to be much faster than the usual systematic sequencing methods. In particular, it beneted from the development of ever faster sequencing machines and powerful bioinformatics programs. In the end, it was not adisadvantage that the shotgun method required multiple sequencing of the genome due to the high redundancy of the method. Interestingly, the shot­gun method was also used in the end by the international consortium that followed the systematic approach to elu­cidate local sequence regions. Since the initial intention of the private company was to patent the sequenced genome, the competition between the two initiatives was great. In March 2000, American President Bill Clinton declared that the human genome was not patentable and advocated its use by all for the common good.
How did it come that acompeting private initiative started to sequence the genome? In spring 1995, Craig Ven­ter and his group identied the entire genome for the bac­terium Haemophilus inuenzae using the shotgun method. The enormous number of 1,830,121 base pairs that code for 1749 genes was sequenced. The complete genomes of individual viruses were already known, but this was the
. Table 12.1 Examples for the sequenced genomes of differ-
ent organisms
Organism
HI virus
HI-9.2 virus, Phageλ
Intestinal bacteria Escherichia coli 4.6 × 10
Baker’s yeast, Saccharomyces
cerevisiae
Pin worm, Caenorhabditis elegans 8 × 10
Wallcress, Arabidopsis thaliana 1 × 10
Fruit y, Drosophila melanogaster 2 × 10
Green blow sh, Tetraodon
nigroviridis
Human, Homo sapiens 3.2 × 10
Common newt, Triturus vulgaris 2.5 × 10
Ethiopian lung sh, Protopterus
aethiopicus
Amoeba, Amoeba dubia 6.70 × 10
a
b
c
c
Number of base pairs Single-stranded RNA
human immunodeciency virus
Genome sizeaGenes
9.2 × 103 b 9
4.85 × 104 70
6
4800
7
2 × 10
3.85 × 10
1.3 × 10
6275
7
19,000
8
25,500
8
13,600
8
9
~21,500
10
10
10
-
Chapter  • Gene Technology in Drug Research
12
decoding of the genetic information of a self-contained creature. The subsequent decoding of the sequence of 580,067 base pairs of the Mycoplasma genitalium genome by Venter’s wife, Claire Fraser, took only four months.
Venter and his group worked with the shotgun method on the entire genome, the so-called “whole-ge­nome shotgun sequencing.” The statistical approach that was followed by Venter initially seemed so unusual and utopian that his application for aresearch grant from the American National Institutes of Health (NIH) was re­jected. This brought about the founding of The Institute for Genomic Research (TIGR) and the company Celera Genomics. There, Venter could pursue his research ac­cording to his ideas and plans. Finally, the success proved the feasibility of the proposed strategy.
Whose genome was actually sequenced? In both ini­tiatives, the DNA of multiple individuals were mixed and the individual differences were purposefully calcu­lated out. In this way the “consensus sequence” of the human genome was determined. But it did not stop with the human genome. The complete elucidation of baker’s yeast Saccharomyces cerevisiae, and the common thale cress Arabidopsis thaliana, the rice plant Oryza sativa, the pinworm Caenorhabditis elegans, the fruit y Drosophila
melanogaster, the chimpanzee Pan troglodytes, the mouse Mus musculus, and many other organisms (. Table12.1)
has been accomplished. In the meantime, new ones emerge weekly. This raises new questions: how should this pleth­ora of information be managed? How can the genetic in­formation be translated into useful knowledge? The eld of bioinformatics has been challenged. Computer pro­grams for the intelligent comparison of sequences and the analysis of metabolic pathways and signaling cascades already existed. New initiatives were launched with the goal of determining the spatial structure of all or at least many sequences. The spatial structures of all real, natu­rally occurring proteins were slowly being elucidated. The crystal structures of all members of some protein families of the human genome have now been determined. In the meantime, structure predictions derived from amino acid sequences have become much more reliable (Sect.20.6). Thus, it is only amatter of time until we will be able to place the sequence catalogs of very many genomes next to those with all the spatial structure maps.
12.4 What Is Contained in the Biological
Space of the Human Proteome?
After the human genome was sequenced, the exciting question arose as to what gene products all these DNA se­quences code for. The rst thing to note is that the genome is not static but constantly changing. Only this way can the genetic variations that make up the diversity of all living creatures occur. In the course of evolution, the genetic constitution has expanded. Simple unicellular organisms
without cell nuclei (prokaryotes) have acircular genome containing only coding genes. Single-celled organisms with anucleus (eukaryotes), such as yeast, have alarger genome, of which about 20% consists of coding genes. Multicellular organisms, such as humans, have genomes that are several 100 times larger than that of yeast (. Ta­ble12.1). However, the number of coding genes is not greater. In fact, some organisms, such as amoebae, have genomes that are 200 times larger than that of humans. Even the tiny water ea, with its 31,000 genes, dwarfs us numerically. Thus, the alleged masterpiece of creation does not necessarily have the largest genome. Obviously, only asmall number of additional DNA sequences, which actually code for additional gene products, have accrued during the course of evolution. Many genes in higher or­ganisms are similar to those in simpler species. If the num­ber of coding genes has barely increased from unicellular organisms to humans, and even the gene products encoded are similar, what explains the massive increase in genome complexity in higher organisms? The answer does not lie in the diversity of required gene products, but rather in the nely tuned regulation of gene expression (Sect.12.13).
In higher organisms, it is critical where and when spe­cic gene products are synthesized. The 95% of human DNA that does not code for proteins contains numerous sequences and signals that control gene expression. There­fore, the total number of genes does not seem to increase in higher organisms, but rather gene density decreases. On average, there are 12genes per million base pairs in the human genome, compared to 118 in the fruit y, 197 in the nematode, and 221 in the common thale cress (Arabi- dopsis thaliana). Moreover, the human genome is highly scattered. It seems that it is not the number of genes, but rather how they are used and how their activation is regulated that determines the developmental state of an organism. It must also be considered that multicellular organisms also require agreat deal of cell differentiation in the different organs. These processes must be reliably regulated and controlled. In addition, higher organisms achieve amuch greater diversity in their protein composi­tion through alternative splicing. Posttranslational modi- cation after biosynthesis also plays arole. This is observed to amuch lesser extent in prokaryotes, for example. After transcription from DNA to RNA, the splicing process cuts out parts of the RNA that do not code for proteins. In alternative splicing, adecision is made during splicing as to what is cut out and what is used for translation. Thus, one DNA sequence can code for several different proteins.
To date, one of the largest genomes of a prokaryote that has been found belongs to the pathogenic protozoa Trichomonas vaginalis. It consists of 160 million base pairs. In humans, this pathogen is usually transmitted through sexual intercourse and causes urinary tract infections. Its huge genome is disproportionately large in the cell. This could be an advantage for the pathogen because its large surface area makes it easier to adhere to the vaginal mu-
. • What Is Contained in the Biological Space of the Human Proteome?


cosa. In addition, the immune system has atough time attacking and destroying such an oversized parasite. The genome of the soil bacterium Sorangium cellosum with 13million bases and 10,000 genes is four times as large as the average genome of other bacteria. This may have some­thing to do with the fact that this soil bacterium is able to perform special tasks that make its therapeutic use interest­ing. It is aversatile producer of complex natural products such as epothilones, which are potent chemotherapeutics that have great potential in the treatment of cancer.
According to estimates in 2016, the human genome consists of more than 3.088 billion bases. It contains ap­proximately 21,500 protein-coding genes and several thou- sand RNA genes. The former textbook knowledge that there is agene product behind each DNA sequence needs to be expanded. It should not be overlooked that our ge­nome contains many thousands of genes for noncoding RNA segments. The resulting RNA molecules perform important functions in our bodies. Particularly notewor­thy are the large groups of tRNAs that serve as adapter molecules for reading and translating base-pair triplets in the genome into the correct amino acid sequence. In ad­dition, it has been shown that the ribosome itself, the mo­lecular machinery for protein synthesis, consists largely of RNA (Sect.32.7). The spliceosome, the complex ma­chinery for removing noncoding segments of the genome, contains RNA molecules called snRNAs. There are even more small RNA molecules (snoRNAs) that are respon­sible for processing and modifying other RNA molecules.
The number of protein-coding genes is, as mentioned, about 21,500, but we still do not know what functions all these proteins perform. Bioinformatics has contributed greatly to the classication of their biochemical function, that is, whether the protein is an enzyme (e.g., protease, kinase, or oxidoreductase) or areceptor, ion channel, or transporter. The function or class of protein to which anew sequence belongs can be determined by comparing it to previously annotated proteins. Multiple sequence comparisons within aprotein family often reveal signi­cant similarity. Information about spatial architecture and folding (Sect.14.2) can be analyzed using relation­ships, because the spatial geometry of proteins is much more conserved than the sequential composition of the folded protein chain. Often, individual motifs or charac­teristic sequence segments reveal aparticular biochemi­cal function of aprotein. Another tool in this detective tour de force of functional annotation has been the com­parison of protein sequences across species.
Assigning abiochemical function to a protein se­quence provides arst idea about its molecular function. It shows, for example, whether it acts as acatalyst to cleave apeptide sequence, carries out ametabolic reduc­tion or, as areceptor, transmits asignal to the cell. What this regulation and control means for the organism re­mains to be resolved. It is also not known whether apar­ticular protein causes adisease because of its defective
function or its dysregulation. Correcting such adefect could lead to asuccessful pharmaceutical therapy.
In the Science publication from the Venter group in 2001, it was assumed that the genome coded for more than 26,500 proteins. At that time, adenitive function could not be assigned to 40% of the sequences. In the remaining part, about 10% were detected to be enzymes. Another 12% proved to be involved in signal transduction, and
13.5% are nucleic acid-binding proteins. The large remain­ing group was scattered across many different functions such as proteins of the cytoskeleton, surface receptors, ion channels, transporters, extracellular matrix proteins, immune system proteins, or chaperones. Seven years later, this picture could be rened. The largest protein family with more than 7000 members contains the zinc nger do­main (Sect.28.2). These proteins assume an important role in transcribing sequence segments of the DNA into RNA. Most zinc nger proteins belong to the group of transcrip­tion factors. Another large protein family contains the immunoglobulins. These domains (Sect.32.1), which are
. Table 12.2
man genome (number of family members at the time of the study)
Protein superfamily Number
Zinc nger (C2H2 and C2HC) 7707
Protein kinase-like 876
G-Protein-coupled receptor-like 784
α/β-Hydrolases
Cysteine proteases 164
Trypsin-like serine proteases 155
Metalloprotease (“Zincins”), catalytic domains 132
FAD/NAD(P)-binding domains 79
Cytochrome P450 79
Integrinα, N-terminal domains
Cytokines 52
Cycl. Nucleotide-phosphodiesterase, catalytic
domains
Caspase-like 39
Carbonic anhydrases 23
Aquaporin-like 20
Integrin domains 18
Aspartic proteases 16
ClC-chloride channel 16
Subtilisin-like 14
a
Based on: http://hodgkin.mbu.iisc.ernet.in/~human/
For updated data see: https://www.proteinatlas.org/search/
(The Human Protein Atlas)
Selected examples of protein families in the hu-
a
151
51
50
Chapter  • Gene Technology in Drug Research
. Fig. 12.3 The composition of protein families
that are particularly often associated with human diseases (GPCR G-protein-coupled receptor; Fibronectin extracellular glycoproteins in tissue construction; homeobox proteins that inuence the morphogenetic development; spectrin cytoskeletal proteins; MHCI major histocompatibility complex proteins that are involved in immune-recognition processes; myosin motor protein in muscle control;
RRM RNA-recognition motif transcriptions factor; trypsin-like serine proteases; laminin EGF agrowth
factor in the extracellular matrix; Ras oncoprotein in tumorigenesis; SH2 protein domains in the phos­phorylation signal cascade)
12
constructed from β-pleated sheets, are present in antibod­ies. Afew protein families are listed in . Table12.2 and are presented in more detail in Chaps.23–32 of this book. It is interesting to note which protein family is frequently associated with which disease (. Fig.12.3). This list is headed by protein kinases (Chap.26). It is, therefore, not surprising that drug research in the pharmaceutical indus­try has intensely focused for years on the modulation and inhibition of protein kinases. Next on the list are cadher­ins. These proteins are important for stabilizing cell–cell contacts. They play arole in embryonic morphogenesis, in signal transduction, and in the assembly of the cyto­skeleton in cells. G-protein-coupled receptors, ion chan­nels, trypsin-like serine proteases or RAS proteins are also prominent on the list of proteins potentially associated with disease, especially when genetically altered.
Finally, consider how the human genome differs from other eukaryotes. Of the more than 2200 protein families discovered in organisms with anucleus, more than 1000 are missing in the human genome. Most of these families have specic functions in their respective organisms or can be explained phylogenetically. For example, venoms are found in snakes, scorpions, and insects. In plants, proteins are found that have avery specic function for the plant, such as nutrient storage in seeds or defense against the attack of pests. The proteins that are absent in humans usually perform biochemical functions that are irrelevant to our organism, or they perform avery specic task in lower eukaryotes.
12.5 Knock in, Knock out: Validation of
Therapeutic Concepts
Molecular biology provides a wealth of information about how diseases develop and how their course can be inuenced. This is the basis of the long road from the dis­covery to the development of anew drug. At the end of
the process, it may be found that the result, although well planned, does not lead to the desired clinical success. It is therefore important to have an animal model that can be used to validate the therapeutic concept at an early stage. Classical test models are often not available because the disease in question does not occur in animals.
Since the 1980s, transgenic animals have been increas- ingly used in pharmacological research. Transgenic animals are animals in which aspecic gene has been completely or partially turned off or replaced by ahuman gene. An animal in which the gene is completely knocked out corre­sponds to an animal in which the corresponding protein is absent or nonfunctional. Aheterozygous animal in which the gene of only one parent is present corresponds to an animal in which the corresponding protein is only partially blocked. When the gene for an enzyme or receptor is af­fected, the effect of an inhibitor or antagonist can be sim­ulated. The onset and progression of adisease, or the inu­ence of protein inhibition on adisease, can be observed in such an animal. In this way, the relevance of atherapeutic concept can be established before an extremely long re­search and development process is launched. The increased production of aparticular protein can be induced by the amplication of a gene. If the absence of agene causes the overexpression of another gene, this will also become transparent. The gene product that is then produced in in­creased amounts can take over the missing function of the silenced product. In such acase, the planned therapeutic principle would only work if the function of the other gene product is also blocked. This question plays an important role in the inhibition of kinases (Sect.26.2).
The knock-out method involves switching off avery specic gene. This technique was developed by Mario Capecchi at the University of Utah, USA, in 1987. The sequence of the gene to be knocked out must be known. Astructurally homologous gene is created that is not functional, for example, because of the insertion of astop signal. The gene is introduced into an animal and
. • Recombinant Proteins for Molecular Test Systems


the intact gene is replaced at exactly the same location. This process is called homologous recombination or gene targeting. Mice are particularly well suited because the technology for manipulating their embryonic stem cells is particularly well established. Aforeign gene, such as ahuman gene, can also be introduced. Mice are also well suited for this because their genome is surprisingly simi­lar to the human genome.
To create atransgenic mouse, female mice are treated to produce alarge number of egg cells. After fertilization, stem cells are extracted from the embryos at avery early stage, the blastocyst stage. They are cultured in vitro and the desired gene is injected into the cell. This procedure results in alow yield. Atechnique has been developed to differentiate transfected from nontransfected cells. The gene to be transferred is coupled to agene that confers resistance to the cytotoxin neomycin. When cells are treated with neomycin, only transformed cells survive. The blastocytes are combined with blastocytes from other mice, and the altered embryos are carried to birth by mice. The offspring of the surrogate mothers are chi­meric, meaning that they carry the genetic information of both the donor and acceptor mice. Here, mice with differently colored fur are selected so that the trans­formed mice are easily recognizable by their spotted fur.
Another method is to directly inject foreign DNA during an early embryonic stage. Adisadvantage of ran­dom insertion of agene is the possibility of destroying another gene, alack of expression of the new gene, or multiple insertions. Animals from the rst litter are bred to produce both genetically mixed, heterozygous animals and genetically homogeneous, homozygous animals. So­phisticated techniques can even selectively turn the new genes on and off.
In this way, transgenic animals are generated in which hereditary diseases such as cystic brosis, Crohn’s dis­ease, phenylketonuria, and others can be studied. Today, animal models also exist for diseases that have different or multiple causes, such as cancer, diabetes, rheumatoid arthritis, and Alzheimer’s disease. Since 1988, when the U.S. Patent Ofce granted the rst patent for atransgenic mouse, there has been controversy over whether aliving creature can be patented at all. European patent law will prohibit patents if genetic modications can lead to ani­mal suffering. Exceptions are allowed only if asignicant medical benet is to be expected. In 1992, the rst patent in Europe was approved for agenetically modied “can­cer mouse.” Since then, many similar patents have been led, mostly on laboratory animals. But now farm ani­mals such as cattle and pigs are also included. Because of their genetic similarity to humans, the European Patent Ofce declared patents on genetically modied primates invalid on ethical grounds in July 2020. Currently, pat­ents on laboratory animals are not completely banned, but should be limited to afew exceptional cases.
12.6 Recombinant Proteins for Molecular
Test Systems
Early on, pure or enriched enzymes were available for in vitro assays, but only in cases where the material was read-
ily available, for example, human thrombin from blood. In other cases, animal material had to be used, with all the risks this entails, given the relevance to rational design (see Sect.19.11). There are many proteins that cannot be isolated in sufcient quantities or in ahomogeneous form. The sequence determination and production of such pro­teins is nowadays easy. The incredibly small amount of afew picomoles (1 pmol = 10 termine the primary structure of ashort protein segment. From the amino acid sequence determined in this way, the genetic code can be reconstructed into agene. It should be noted that several base triplets can stand for apartic­ular amino acid (so-called degenerate code, Sect.32.7). Aset of single-stranded oligonucleotides is synthesized that could theoretically cover the entire original peptide segment. These molecules can be used to nd acomple­mentary sequence in acDNA library. cDNA (comple- mentary DNA) is the DNA complementary to mRNA (messenger RNA). It is obtained from the mRNA, which contains only the sequence needed for protein biosyn­thesis, by reverse transcription with areverse transcrip­tase (Sect.32.5). Finally, the gene is produced in larger quantities using the PCR technique, and the amino acid sequence is determined from its base sequence, simply because polynucleotides are much easier to sequence.
Next, the gene is introduced into cells that are allowed to reproduce. In afew cases, there may be difculties with this step. In bacteria, such as the intestinal bacterium Escherichia coli, or in yeast cells, only soluble proteins can be produced. Some proteins accumulate in inclusion bodies. They must be extracted, solubilized, and refolded under specic conditions. The gene segment for asmall protein is often fused with the gene for another protein, and the fusion protein is then expressed. Often, the large protein conjugate formed in the cell is more soluble and/ or better protected from metabolic degradation than small proteins. During purication, the nonessential part of the protein conjugate is cleaved off. Problems may arise if the protein is not folded correctly or if several chains (such as insulin) have to be linked by disulde bridges. Larger proteins that require sugar groups to per­form their function (glycosylation) must be produced in cells of higher organisms, such as mammalian cells. The production of complex proteins in insect cells has become particularly attractive. These cells are infected with abac- ulovirus that has incorporated the desired information into its genome. The virus encodes the target protein, and the insect cells provide for its production and subsequent glycosylation. Not only enzymes, but also receptors, ion channels, and entire signaling cascades can be produced in cells in this way.
−12
mol) is sufcient to de-
Chapter  • Gene Technology in Drug Research
12

12.7 Silencing Genes by RNA Interference

How genetically modied species can be created by inter­vening in the germline of an organism was described in Sect.12.5. For example, such species may lack aparticular gene and, therefore, agene product, or anew gene may have been introduced. In this way, the function of certain genes can be studied in aliving organism. This makes the consequences of blocking the corresponding gene product transparent before developing an effective drug. At the end of the 1990s, another technique that made it possible to silence genes without using mutagenesis to interfere with the genes of the organism was developed. This work was done by Andrew Fire and Craig Mello, who were awarded the Nobel Prize for their achievements in 2006.
Genes are stored on DNA. For gene expression, the coding parts of the genome are rst transcribed onto mRNA. The ribosome then uses this transcribed informa­tion to convert the base sequence into apeptide sequence (Sect.32.7). The idea of trapping the transcribed informa­tion on the single-stranded mRNA by adding acomple­mentary single-stranded RNA was rst proposed in the early 1980s. The two strands can be joined by “hybridiza­tion,” which means that the single nucleotide strand is sup­plemented by the complementary strand to form adouble strand. The resulting double-stranded RNA is no longer suitable as atemplate for protein biosynthesis. In practice, however, this antisense principle (Sect.32.4) did not pro­duce the expected breakthrough results. In some cases, genes were only weakly suppressed, and even the addition of the normal RNA strand could achieve asuppression. Fire and Mello suspected that neither the normal nor the antisense strand caused gene blockage, but rather the double-stranded form, which was present as an impurity. Further experiments conrmed their suspicions. Interest­ingly, even small amounts of double-stranded RNA were sufcient to render alarge number of mRNA molecules unusable. In contrast, the use of antisense strands would require stoichiometric amounts. It was further shown that even short double-stranded RNA fragments of about 20nucleotides were sufcient to silence entire mRNA gene sequences. Fire and Mello named this phenomenon RNA interference. What had happened? An enzyme called dicer chopped up the double-stranded RNA into pieces with alength of 21–23 nucleotides, which then resulted in block­ade. The double-stranded RNA pieces are incorporated into the enzyme complex RISC (RNA-induced silencing complex) and separated into individual strands. One strand is released from the complex, while the other remains as atemplate for capturing other mRNA molecules.
The sequence of the captured strand allows the RISC complex to recognize and sequentially cleave all mRNAs with acomplementary base sequence. They are then di­gested by enzymes in the cytoplasm. The cell selectively eliminates only those mRNAs that contain the sequence pattern complementary to the short RNA strands in
the RISC complex. In practice, this gene blockade has proven to be simpler and more reliable than the anti­sense technique. RNA interference can, thus, be used to systematically silence discovered genes in order to draw conclusions about the resulting consequences for the or­ganism. RNA interference is not only used for analytical purposes. There are already biotech companies that use small RNA fragments to silence disease-causing genes.
Another major problem in the development of RNAi therapeutics is how to get a22-base RNA molecule into the cell where it is supposed to work. Highly charged molecules cannot cross the cell membrane. Therefore, aspecial delivery system is needed. Intensive research is underway to develop such systems, but the problem is far from being solved. Areliable and highly efcient system that can selectively deliver such polar and nuclease-sen­sitive molecules into the interior of the cell is likely to open up acompletely new and currently unforeseeable perspective for disease therapy.
The goal is to construct delivery systems that can package the fragile and polar freight of RNA molecules and dock onto the cell. Once there, the coat of these carriers must fuse with the cell membrane or selectively penetrate to reach the interior of the cell. One concept is to package and compartmentalize RNA in polymers such as polyethyleneimine. The positive charges on the polymer backbone can bind and encapsulate anegatively charged polymer molecule such as RNA or DNA build­ing blocks. Other systems attempt to make the RNA or DNA molecules bioavailable to the cell by encapsulating them in amembrane-like coating. This packaging in lipo­somes results in selective adhesion of the articial cell to the membrane of the target cell, and then the liposome fuses with the target cell in an endocytosis-like process.
Another problem is the risk that small, silencing RNA molecules (siRNAs) could trigger an immune response. One solution to this dilemma is the chemical modica­tion of siRNAs. The RNA molecules are modied in such away that they still hybridize optimally to the targeted segment in the mRNA, but have improved properties in terms of transport, immunogenicity, and stability. The OH groups of the ribose building block of the nucleotides have been replaced by uorine, methoxy, or hydrogen.
Certainly, siRNA research is still in its infancy. The potential of the methods seems impressive, since they use the principles of gene regulation that are applied in Na­ture. As described earlier, we have genes in our genome that encode microRNAs and that have long stretches of sequence complementarity. Structurally, they exist as double strands. They are cut by the dicer protein and can be used to interfere with RNA, resulting in an alternative means of gene regulation. For abroad therapeutic appli­cation of externally delivered RNA snips, these problems need to be overcome. In 2018, the company Alnylam launched the rst RNAi therapeutic, patisiran, for the treatment of polyneuropathy in hereditary transthyre-