Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5319_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Preface and Acknowledgement
- •Chemical Structures of Amino Acids,Molecular Graphics and Introduction
- •Introduction
- •Literature
- •Chapter Abstract Videos
- •Contents
- •About the author
- •1.10 Synopsis
- •1.3 The Battle Against Infectious Disease
- •1.4 Biological Concepts in Drug Research
- •Bibliography and Further Reading
- •2.8 A Long List of Accidents
- •2.10 Synopsis
- •Bibliography and Further Reading
- •3. Classical Drug Research
- •3.2 Malaria: Success and Failure
- •3.6 Synopsis
- •Bibliography and Further Reading
- •4.1 The Lock-and-Key Principle
- •4.2 The Essential Role of the Membrane
- •4.6 Blame It All on Water!
- •4.11 Lessons for Drug Design
- •4.12 Synopsis
- •Bibliography and Further Reading
- •5.1 Louis Pasteur Sorts Crystals
- •5.2 Structural Basis of Optical Activity
- •5.4 Lipases Separate Racemates
- •5.8 Synopsis
- •Bibliography and Further Reading
- •6.2 Lead Structures from Plants
- •6.9 Synopsis
- •Bibliography and Further Reading
- •7.2 Color Change Demonstrates Activity
- •7.7 Biophysics Supports Screening
- •7.11 Synopsis
- •Bibliography and Further Reading
- •8.1 Strategies for Drug Optimization
- •8.5 From Agonists to Antagonists
- •8.9 Synopsis
- •Bibliography and Further Reading
- •9. Designing Prodrugs
- •9.1 Foundations of Drug Metabolism
- •9.2 Esters Are Ideal Prodrugs
- •9.6 Synopsis
- •Bibliography and Further Reading
- •10. Peptidomimetics
- •10.1 Therapeutic Relevance of Peptides
- •10.2 Designing Peptidomimetics
- •Bibliography and Further Reading
- •11.4 What Is Contained in Chemical Space?
- •Bibliography and Further Reading
- •12.7 Silencing Genes by RNA Interference
- •12.9 Proteomics and Metabolomics
- •Bibliography and Further Reading
- •13.3 Crystal Lattices Diffract X-Rays
- •Bibliography and Further Reading
- •Bibliography and further reading
- •15. Molecular Modeling
- •15.2 Strategies in Molecular Modeling
- •15.3 Knowledge-Based Approaches
- •15.4 Force Field Methods
- •15.5 Quantum Chemical Methods
- •Bibliography and further reading
- •16. Conformational Analysis
- •16.8 Synopsis
- •Bibliography and Further Reading
- •Bibliography and Further Reading
- •18.4 Lipophilicity and Biological Activity
- •Bibliography and Further Reading
- •19.3 The Role of Hydrogen Bonds
- •19.5 Absorption Profiles of Acids and Bases
- •19.8 From In Vitro to In Vivo Activity
- •Bibliography and Further Reading
- •Bibliography and Further Reading
- •21.5 LUDI Discovers the First Leads
- •Bibliography and Original Papers
- •22.1 The Druggable Genome
- •22.4 Enzymes and Their Inhibitors
- •22.9 Resistance and Its Origin
- •Bibliography and Further Reading
- •23.1 Serine-Dependent Hydrolases
- •23.10 Synopsis
- •Bibliography and Further Reading
- •24. Aspartic Protease Inhibitors
- •24.2 Design of Renin Inhibitors
- •24.8 Synopsis
- •Bibliography and Further Reading
- •25.1 Structure of Zinc Metalloproteases
- •25.9 What Zinc Can Do, Iron Can Too
- •25.11 Synopsis
- •Bibliography and Further Reading
- •26. Transferase Inhibitors
- •26.1 The Kinase “Gold Rush”
- •Bibliography and Further Reading
- •27. Oxidoreductase Inhibitors

Gene Technology in Drug
Research
Contents
12.1 The History and Basics of Gene Technology – 170
12.2 Gene Technology: AKey Technology in Drug Design – 171
12.3 Genome Projects Decipher Biological Constructions – 173
12.4 What Is Contained in the Biological Space
of the Human Proteome? – 174
12.5 Knock in, Knock out: Validation of
Therapeutic Concepts – 176
12.6 Recombinant Proteins for Molecular Test Systems – 177
12.7 Silencing Genes by RNA Interference – 178
12.8 PROTAC: How to force therapeutically untargetable proteins
into targeted degradation – 179
12.9 Proteomics and Metabolomics – 180
12.10 Expression Patterns on aChip: Microarray Technology – 182
12.11 SNPs and Polymorphism: What Makes Us Dierent – 183
12.12 The Personal Genome: Access to an
Individualized Therapy? – 184
12.13 When Genetic Dierences Turn into Disease – 184
12.14 Epigenetics: Lifestyle and Environment Inuence Gene
Activity Like aPen Leaves aMark in the Book of Life – 185
12.15 The Scope and Limitations of Gene Therapy – 187
12.16 Synopsis – 189
Bibliography and Further Reading – 190
© The Author(s), under exclusive license to Springer-Verlag GmbH, DE, part of Springer Nature 2024
G. Klebe, Drug Design, https://doi.org/10.1007/978-3-662-68998-1_12

Chapter • Gene Technology in Drug Research
12
Engineers and writers have predicted many developments
in science and technology. Among other sophisticated
machines, Leonardo da Vinci described the principle
of the helicopter. In the early 1820s, Charles Babbage
designed an automatic calculating machine that was far
ahead of its time. More than 160 years later, the mechanical precursor of aprogrammable computer was actually
built, and it worked! Jules Verne described submarines
and atrip to the moon, and Hans Dominik described
creating energy by splitting the atom. All these visions
became reality. Only one application of gene technology,
the most groundbreaking invention of our time, was foreseen: the cloning of two genetically identical individuals
in Aldous Huxley’s Brave New World in 1932. Mamma-
lian cloning has already been carried out. When Dolly,
the rst cloned sheep, was born in Scotland on July5,
1996, it was hailed as abreakthrough in genetic research.
Cloned mice, cattle, pigs, horses and, most recently, monkeys have followed. It is to be hoped that researchers will
respect ethical boundaries and not make use of Huxley’s
idea of human cloning, despite its tangible feasibility.
Genetic engineering makes it possible to introduce
new genes into acell, multiply them, and exchange or
remove them. When removed or altered, the cell can no
longer produce the original protein derived from that
gene. With the introduction of anew gene and aclever
choice of method, the cell produces aforeign product,
either an intentionally modied protein or acompletely
new one. For many diseases, the molecular cause is
known to be the absence of aprotein or agenetically
caused mutation in aprotein. These are only afew of the
more common examples:
Diabetes as aresult of insulin deciency,
-
Particular, hereditary cancer forms (e.g., familial co-
-
lon cancer, malignant melanoma),
Chorea Huntington, achronic form of brain atrophy,
-
Sickle cell anemia, agenetic disease producing mal-
-
formed red blood cells (Sect.12.13), and
Bleeding disorders that are caused by the absence of
-
particular coagulation factors (see Sect.12.13).
The possibility of purposefully producing arbitrary proteins has yielded the following main applications of gene
technology:
The identication of genes and proteins that could
-
play arole in the treatment of adisease,
The development of animal models to test athera-
-
peutic principle,
The production of proteins for therapies in which
-
aparticular protein is missing,
The manufacture of monoclonal antibodies and vac-
-
cines,
The manufacture of proteins for molecular test sys-
-
tems, and the determination of the 3D structures of
enzymes and receptor proteins,
The generation of proteins of which atargeted muta-
-
genesis has been undertaken to exchange one or more
amino acids for the elucidation of the mode of action
of enzymes and for the characterization of receptor
binding sites,
Somatic individual gene therapy for specic patients,
-
and
Immunotherapy using genetically engineered T-cells
-
with antigen-specic receptors.
Other application possibilities, for example, the manipulation of the human germline, or genetic changes in crops
to achieve herbicide resistance, or to prolong the shelf life
of fruits, are only briey mentioned here.
12.1 The History and Basics of Gene
Technology
The foundations of gene technology were rst established
in the middle of the twentieth century. It all started in
1953 after James Watson and Francis Crick in Cambridge
(England) became aware of Rosalind Franklin’s X-ray
data; they correctly interpreted the available knowledge
about deoxyribonucleic acid (DNA) and proposed the
structure of adouble helix as the three-dimensional geometry for the genetic material of all living beings. Immediate indications were obtained from the structure
about the mechanism of our hereditary transfer and
about the genetic code for the biosynthesis of proteins.
Afew years later, Werner Arber discovered enzymes that
attack specic sites on the double helix and cleave DNA
in asequence-specic manner. These enzymes are called
restriction enzymes. What was initially seen as acuriosity
proved to be an exceedingly important discovery for gene
technology. It is possible to selectively cleave DNA with
these enzymes and to introduce new fragments. Next,
the merging of new information with the original DNA,
the recombination of the genetic constitution, is accomplished with ligases from special viruses called bacteriophages. The techniques for DNA sequencing have also
made decisive progress. Soon afterwards, the amino acid
sequence of aprotein was no longer directly determined,
but rather deduced from the analysis of the corresponding DNA. Today, sequencing is mainly done via cDNA,
which is complementary to the RNA (Sect.12.6).
In 1973, Stanley Cohen and Herbert Boyer managed
to recombine the genome of abacterium for the rst
time (. Fig.12.1). Then, things happened one after the
other: two years later the bacterial strain Escherichia coli
K12, which is still used today, was developed. Part of
its genetic constitution is missing, making it viable only
under laboratory conditions. This bacterium can be genetically manipulated at will without fear of harm. The
British scientists H.Williams-Smith and E.S. Anderson

. • Gene Technology: AKey Technology in Drug Design
. Fig. 12.1 The principle of gene technological recombination of
hereditary information. Bacteria often contain additional genetic material in addition to their “chromosome” in the form of ring-shaped
plasmids; these are used in gene technology as vectors to introduce
foreign genes. Plasmids are removed from the cell and sequence-specically cut with so-called restriction enzymes, which are isolated from
bacteria. The target DNA that carries the desired gene, which was
typically also treated with the same restriction enzyme, is bound to
carried out independent experiments in which they took
Escherichia coli K12 orally. They demonstrated that these
bacteria survive only for ashort time in the gastrointestinal tract and that the K12 gene, which confers antibiotic
resistance for selection of the transformed cells, cannot
be transferred to normal Escherichia coli found in the
intestinal ora. At aconference in Asilomar, California,
experts discussed the potential dangers of genetic engineering and dened various risk and safety classes. Genentech was founded in 1976. Its founder, Herbert Boyer,
had to borrow US$ 500 as start-up capital! When the
company went public in 1980, the value of his shares
made him amillionaire within minutes. In 1982, Genentech introduced the rst drug to be produced using gene
technology: human insulin (Humulin®).
In 1983, Kary Mullis made aseminal contribution
to gene technology when he developed the polymerase
chain reaction (PCR) while working at Cetus, aCalifornian company founded in 1971. Double-stranded DNA
is melted into its single strands by heating, then two short
pieces of single-stranded DNA complementary to the
regions at the beginning of the DNA, called primers, are
added along with the four DNA nucleotides. A polymerase can now be used to synthesize new DNA in atest
tube. This means that anew double strand is formed by
starting with the primers (. Fig.12.2). A heat-stable
DNA polymerase (originally from the bacterium Ther-
mus aquaticus, endemic to the hot springs of Yellowstone
National Park) is used for DNA synthesis. Each repetition of this step doubles the amount of DNA. Within
afew hours, billions and trillions of DNA molecules
can be produced from asingle starting molecule. This
the overlapping single-stranded DNA ends in vitro. The DNA ends
are coupled with the enzyme DNA ligase, and the modied, recombinant plasmid is brought into the bacterial cell. In addition to the
DNA segment that is necessary for replication, plasmid vectors that
are used in gene technology carry additional information that allows
for the recognition and selection of the transformed cells (usually an
antibiotic-resistance gene). In the presence of the selecting agent, only
plasmid-containing cells grow
amount is sufcient to sequence the DNA segment of
interest.
PCR methods are applied diversely. The entire genetic
information of an individual can be derived from asingle DNA molecule. In medical diagnostics, this serves
to provide evidence regarding genetic disorders, cancer,
infectious diseases, and risk factors. PCR methods are
also used to establish agenetic ngerprint in paternity
tests and in forensic science.
New genetic information can be introduced not only
into bacterial cells, but also into yeast, virus-infected insect cells, and even mammalian cells. However, as arst
approximation, the more complex the organism, from
bacteria to mammalian cells, the more difcult it is to
produce proteins in these cells. On the other hand, insect
and mammalian cells have the advantage of producing
not only ordinary proteins, but also more complex ones
(e.g., glycosylated proteins) in afunctional form. Thus,
in many cases, we are dependent on such organisms for
the production of proteins.
12.2 Gene Technology: AKey Technology
in Drug Design
The 1970s and 1980s were the golden age of receptor
binding assays with membrane preparations. Radioactively labeled ligands were used to determine the specic
binding of new drug candidates. The major receptors
for hormones and neurotransmitters were known, and
in some cases the difference between pre- and postsynaptic receptors. However, the different subtypes and

Chapter • Gene Technology in Drug Research
12
. Fig. 12.2 The polymerase chain reaction (PCR) can make an un-
limited number of identical copies of aDNA molecule. The DNA is
heated to split the double-stranded DNA into complementary single
strands. Synthetic oligonucleotides of about 20bases, called primers,
that are complementary to these DNA strands hybridize to the corresponding strand. Each primer must bind to one end of each strand
of DNA. The primers dene the boundaries of the amplied DNA.
In addition, an excess of primers must be used because one pair of
primers is required for each double strand of DNA in each cycle. The
primers are needed to synthesize the new DNA in the presence of the
DNA polymerase and an excess of the four different nucleotides. This
their amino acid sequences were not known. As aconsequence, the results of these studies were quite imprecise.
Gene technology methods allow the production of
homogeneous recombinant proteins in virtually unlimited
quantities. They play an important role in the very rst
step of drug design: the identication of atarget protein.
Advances in methodology have led to the discovery of
new receptors, some of which had unknown function or
specicity. The next steps are to test the therapeutic con-
cept in genetically modied animals. Another important
contribution is the preparation of proteins for molecular
test systems and the isolation of adequate material for the
elucidation of the 3D protein structure (Chap.13). With,
perhaps, the exception of afew proteins that can be isolated from blood or other natural sources, the production
of large quantities of proteins is dependent on gene technology. Today, the purication of proteins from animal
or human blood is done rather reluctantly. The risk of
transmitting viruses or infections is deemed to be too high.
Gene technology offers the possibility to selectively
produce structural variants of proteins. The generation of
point mutations (site-directed mutagenesis) allows partic-
occurs in the reverse direction (dashed arrow) due to the opposite ori-
entation of the DNA strands and the specicity of the polymerase.
The newly synthesized DNA segment can be several hundred to several thousand base pairs long. The result is two identical double-stranded DNA molecules. After heating, single strands are obtained and the
above procedure is repeated. Because DNA polymerase is heat stable,
it does not need to be added repeatedly. Each repetition of the above
steps results in adoubling of the DNA molecule. Its number increases
exponentially. Ten cycles lead to about 1000 DNA molecules, 20to
amillion, and 30to abillion. In this way, asingle DNA molecule can
be multiplied into abiochemically analyzable quantity
ular properties in proteins to be improved, and the binding and catalytic properties of enzymes to be purposefully
changed. Membrane-bound receptors can be probed position by position to establish which amino acids are responsible for the maintenance and stability of the 3D structure,
the adoption of aparticular conformation, or are of critical importance with respect to the binding of aligand.
Three-dimensional structural models of receptors can be
generated in this way, or their relevance can be appraised.
In many cases, it has also proven worthwhile to introduce point mutations that change the surface properties of proteins and help to elucidate the 3D structure of
proteins. Sometimes the charge on individual amino acids
or highly exible amino acids that tend to cause disorder
must be exchanged for the sake of protein crystallization.
For proteins in which part of the sequence is membrane
anchored, the membrane anchor, which would interfere
with crystallization, is removed prior to the crystallization
experiment. For soluble receptors, it has been found useful
to remove individual domains, crystallize them, and determine their structure. Of course, such modied proteins
must still perform their specic functions, such as ligand

. • Genome Projects Decipher Biological Constructions
binding or DNA docking. Once the difcult crystallization step has been accomplished, the actual structure determination nowadays only takes amatter of afew weeks
in most cases (Chap.13). Membrane-bound proteins can
be stabilized in their structure by point mutations so that
they can be detached from the membrane. They can then
be crystallized in special media, such as cubic lipid phases.
When considering the benets of all this progress for
humanity, we cannot avoid asking the question: where do
the fears of broad sections of society about gene technol-
ogy come from? It is not difcult to understand these reservations: with the use of gene technology, almost everything
that is theoretically conceivable in the eld of genetics becomes possible. However, people’s trust in science is not
as unshakable as it used to be before the atomic bomb.
Now that the opportunities outweigh the risks, the sins
of our forefathers have come back to haunt us. Too often
in the past, scientists have underestimated potential risks
and put their ethical concerns on hold. Scientists have still
not managed to allay the public’s fears. We must take these
fears seriously and rebuild trust by acting responsibly.
12.3 Genome Projects Decipher Biological
Constructions
The entire human genome is organized on 23 chromosomes. In 1990, the Human Genome Organization
(HUGO), with abudget of US$ 3 billion, began the
then-ambitious task of sequencing the entire human
genetic code from DNA within 15years. By the end of
1993, the rst annotated genome maps were available,
and these were later rened. By 2001, the project had
progressed to the point where the entire genome was published in Science and Nature by two parallel consortia.
The two competing consortia followed different strat-
egies. The publicly funded international consortium chose
the approach of setting progressively narrower parameters,
the stepwise digestion of the genome, and the systematic
elucidation of sequences for the complete genome analyses.
In humans, this means that in addition to the 5% of DNA
that corresponds to genes, the other 95% of sequenced
DNA, the function of which was unknown, was classied with the somewhat derogatory term “junk DNA.” It
is now established that these regions play a pivotal role
in the regulation of gene expression and host to a class
of small non-coding RNA molecules, collectively termed
microRNA. In 2024, Victor Ambros (UMass Chen Medical School, Worcester, USA) and Gary Ruvkun (Harvard
Medical School, Boston, USA) were awarded the Nobel
Prize in Medicine for their discovery of microRNA and
the elucidation of their function in post-transcriptional
gene regulation, particularly in the context of organismal
development and function (Sect.12.7). The second strategy, pursued by the privately funded consortium, used the
shotgun approach. This was done by amplifying alonger
strand of DNA and then cutting it into many small seg
ments. After these segments were sequenced, the sequences
were reconstructed into the original long strand of DNA
using apowerful computer program. Of course, this can
only work if the sequences of the cleaved segments overlap
sufciently. This technique proved to be much faster than
the usual systematic sequencing methods. In particular, it
beneted from the development of ever faster sequencing
machines and powerful bioinformatics programs. In the
end, it was not adisadvantage that the shotgun method
required multiple sequencing of the genome due to the
high redundancy of the method. Interestingly, the shotgun method was also used in the end by the international
consortium that followed the systematic approach to elucidate local sequence regions. Since the initial intention of
the private company was to patent the sequenced genome,
the competition between the two initiatives was great. In
March 2000, American President Bill Clinton declared
that the human genome was not patentable and advocated
its use by all for the common good.
How did it come that acompeting private initiative
started to sequence the genome? In spring 1995, Craig Venter and his group identied the entire genome for the bacterium Haemophilus inuenzae using the shotgun method.
The enormous number of 1,830,121 base pairs that code
for 1749 genes was sequenced. The complete genomes of
individual viruses were already known, but this was the
. Table 12.1 Examples for the sequenced genomes of differ-
ent organisms
Organism
HI virus
HI-9.2 virus, Phageλ
Intestinal bacteria Escherichia coli 4.6 × 10
Baker’s yeast, Saccharomyces
cerevisiae
Pin worm, Caenorhabditis elegans 8 × 10
Wallcress, Arabidopsis thaliana 1 × 10
Fruit y, Drosophila melanogaster 2 × 10
Green blow sh, Tetraodon
nigroviridis
Human, Homo sapiens 3.2 × 10
Common newt, Triturus vulgaris 2.5 × 10
Ethiopian lung sh, Protopterus
aethiopicus
Amoeba, Amoeba dubia 6.70 × 10
a
b
c
c
Number of base pairs
Single-stranded RNA
human immunodeciency virus
Genome sizeaGenes
9.2 × 103 b 9
4.85 × 104 70
6
4800
7
2 × 10
3.85 × 10
1.3 × 10
6275
7
19,000
8
25,500
8
13,600
8
9
~21,500
10
10
10
-

Chapter • Gene Technology in Drug Research
12
decoding of the genetic information of a self-contained
creature. The subsequent decoding of the sequence of
580,067 base pairs of the Mycoplasma genitalium genome
by Venter’s wife, Claire Fraser, took only four months.
Venter and his group worked with the shotgun
method on the entire genome, the so-called “whole-genome shotgun sequencing.” The statistical approach that
was followed by Venter initially seemed so unusual and
utopian that his application for aresearch grant from the
American National Institutes of Health (NIH) was rejected. This brought about the founding of The Institute
for Genomic Research (TIGR) and the company Celera
Genomics. There, Venter could pursue his research according to his ideas and plans. Finally, the success proved
the feasibility of the proposed strategy.
Whose genome was actually sequenced? In both initiatives, the DNA of multiple individuals were mixed
and the individual differences were purposefully calculated out. In this way the “consensus sequence” of the
human genome was determined. But it did not stop with
the human genome. The complete elucidation of baker’s
yeast Saccharomyces cerevisiae, and the common thale
cress Arabidopsis thaliana, the rice plant Oryza sativa, the
pinworm Caenorhabditis elegans, the fruit y Drosophila
melanogaster, the chimpanzee Pan troglodytes, the mouse
Mus musculus, and many other organisms (. Table12.1)
has been accomplished. In the meantime, new ones emerge
weekly. This raises new questions: how should this plethora of information be managed? How can the genetic information be translated into useful knowledge? The eld
of bioinformatics has been challenged. Computer programs for the intelligent comparison of sequences and the
analysis of metabolic pathways and signaling cascades
already existed. New initiatives were launched with the
goal of determining the spatial structure of all or at least
many sequences. The spatial structures of all real, naturally occurring proteins were slowly being elucidated. The
crystal structures of all members of some protein families
of the human genome have now been determined. In the
meantime, structure predictions derived from amino acid
sequences have become much more reliable (Sect.20.6).
Thus, it is only amatter of time until we will be able to
place the sequence catalogs of very many genomes next
to those with all the spatial structure maps.
12.4 What Is Contained in the Biological
Space of the Human Proteome?
After the human genome was sequenced, the exciting
question arose as to what gene products all these DNA sequences code for. The rst thing to note is that the genome
is not static but constantly changing. Only this way can the
genetic variations that make up the diversity of all living
creatures occur. In the course of evolution, the genetic
constitution has expanded. Simple unicellular organisms
without cell nuclei (prokaryotes) have acircular genome
containing only coding genes. Single-celled organisms
with anucleus (eukaryotes), such as yeast, have alarger
genome, of which about 20% consists of coding genes.
Multicellular organisms, such as humans, have genomes
that are several 100 times larger than that of yeast (. Table12.1). However, the number of coding genes is not
greater. In fact, some organisms, such as amoebae, have
genomes that are 200 times larger than that of humans.
Even the tiny water ea, with its 31,000 genes, dwarfs us
numerically. Thus, the alleged masterpiece of creation
does not necessarily have the largest genome. Obviously,
only asmall number of additional DNA sequences, which
actually code for additional gene products, have accrued
during the course of evolution. Many genes in higher organisms are similar to those in simpler species. If the number of coding genes has barely increased from unicellular
organisms to humans, and even the gene products encoded
are similar, what explains the massive increase in genome
complexity in higher organisms? The answer does not lie
in the diversity of required gene products, but rather in
the nely tuned regulation of gene expression (Sect.12.13).
In higher organisms, it is critical where and when specic gene products are synthesized. The 95% of human
DNA that does not code for proteins contains numerous
sequences and signals that control gene expression. Therefore, the total number of genes does not seem to increase
in higher organisms, but rather gene density decreases. On
average, there are 12genes per million base pairs in the
human genome, compared to 118 in the fruit y, 197 in
the nematode, and 221 in the common thale cress (Arabi-
dopsis thaliana). Moreover, the human genome is highly
scattered. It seems that it is not the number of genes,
but rather how they are used and how their activation is
regulated that determines the developmental state of an
organism. It must also be considered that multicellular
organisms also require agreat deal of cell differentiation
in the different organs. These processes must be reliably
regulated and controlled. In addition, higher organisms
achieve amuch greater diversity in their protein composition through alternative splicing. Posttranslational modi-
cation after biosynthesis also plays arole. This is observed
to amuch lesser extent in prokaryotes, for example. After
transcription from DNA to RNA, the splicing process
cuts out parts of the RNA that do not code for proteins.
In alternative splicing, adecision is made during splicing
as to what is cut out and what is used for translation. Thus,
one DNA sequence can code for several different proteins.
To date, one of the largest genomes of a prokaryote
that has been found belongs to the pathogenic protozoa
Trichomonas vaginalis. It consists of 160 million base pairs.
In humans, this pathogen is usually transmitted through
sexual intercourse and causes urinary tract infections. Its
huge genome is disproportionately large in the cell. This
could be an advantage for the pathogen because its large
surface area makes it easier to adhere to the vaginal mu-

. • What Is Contained in the Biological Space of the Human Proteome?
cosa. In addition, the immune system has atough time
attacking and destroying such an oversized parasite. The
genome of the soil bacterium Sorangium cellosum with
13million bases and 10,000 genes is four times as large as
the average genome of other bacteria. This may have something to do with the fact that this soil bacterium is able to
perform special tasks that make its therapeutic use interesting. It is aversatile producer of complex natural products
such as epothilones, which are potent chemotherapeutics
that have great potential in the treatment of cancer.
According to estimates in 2016, the human genome
consists of more than 3.088 billion bases. It contains approximately 21,500 protein-coding genes and several thou-
sand RNA genes. The former textbook knowledge that
there is agene product behind each DNA sequence needs
to be expanded. It should not be overlooked that our genome contains many thousands of genes for noncoding
RNA segments. The resulting RNA molecules perform
important functions in our bodies. Particularly noteworthy are the large groups of tRNAs that serve as adapter
molecules for reading and translating base-pair triplets in
the genome into the correct amino acid sequence. In addition, it has been shown that the ribosome itself, the molecular machinery for protein synthesis, consists largely
of RNA (Sect.32.7). The spliceosome, the complex machinery for removing noncoding segments of the genome,
contains RNA molecules called snRNAs. There are even
more small RNA molecules (snoRNAs) that are responsible for processing and modifying other RNA molecules.
The number of protein-coding genes is, as mentioned,
about 21,500, but we still do not know what functions all
these proteins perform. Bioinformatics has contributed
greatly to the classication of their biochemical function,
that is, whether the protein is an enzyme (e.g., protease,
kinase, or oxidoreductase) or areceptor, ion channel, or
transporter. The function or class of protein to which
anew sequence belongs can be determined by comparing
it to previously annotated proteins. Multiple sequence
comparisons within aprotein family often reveal signicant similarity. Information about spatial architecture
and folding (Sect.14.2) can be analyzed using relationships, because the spatial geometry of proteins is much
more conserved than the sequential composition of the
folded protein chain. Often, individual motifs or characteristic sequence segments reveal aparticular biochemical function of aprotein. Another tool in this detective
tour de force of functional annotation has been the comparison of protein sequences across species.
Assigning abiochemical function to a protein sequence provides arst idea about its molecular function.
It shows, for example, whether it acts as acatalyst to
cleave apeptide sequence, carries out ametabolic reduction or, as areceptor, transmits asignal to the cell. What
this regulation and control means for the organism remains to be resolved. It is also not known whether aparticular protein causes adisease because of its defective
function or its dysregulation. Correcting such adefect
could lead to asuccessful pharmaceutical therapy.
In the Science publication from the Venter group in
2001, it was assumed that the genome coded for more than
26,500 proteins. At that time, adenitive function could
not be assigned to 40% of the sequences. In the remaining
part, about 10% were detected to be enzymes. Another
12% proved to be involved in signal transduction, and
13.5% are nucleic acid-binding proteins. The large remaining group was scattered across many different functions
such as proteins of the cytoskeleton, surface receptors,
ion channels, transporters, extracellular matrix proteins,
immune system proteins, or chaperones. Seven years later,
this picture could be rened. The largest protein family
with more than 7000 members contains the zinc nger domain (Sect.28.2). These proteins assume an important role
in transcribing sequence segments of the DNA into RNA.
Most zinc nger proteins belong to the group of transcription factors. Another large protein family contains the
immunoglobulins. These domains (Sect.32.1), which are
. Table 12.2
man genome (number of family members at the time of the study)
Protein superfamily Number
Zinc nger (C2H2 and C2HC) 7707
Protein kinase-like 876
G-Protein-coupled receptor-like 784
α/β-Hydrolases
Cysteine proteases 164
Trypsin-like serine proteases 155
Metalloprotease (“Zincins”), catalytic domains 132
FAD/NAD(P)-binding domains 79
Cytochrome P450 79
Integrinα, N-terminal domains
Cytokines 52
Cycl. Nucleotide-phosphodiesterase, catalytic
domains
Caspase-like 39
Carbonic anhydrases 23
Aquaporin-like 20
Integrin domains 18
Aspartic proteases 16
ClC-chloride channel 16
Subtilisin-like 14
a
Based on: http://hodgkin.mbu.iisc.ernet.in/~human/
For updated data see: https://www.proteinatlas.org/search/
(The Human Protein Atlas)
Selected examples of protein families in the hu-
a
151
51
50

Chapter • Gene Technology in Drug Research
. Fig. 12.3 The composition of protein families
that are particularly often associated with human
diseases (GPCR G-protein-coupled receptor;
Fibronectin extracellular glycoproteins in tissue
construction; homeobox proteins that inuence the
morphogenetic development; spectrin cytoskeletal
proteins; MHCI major histocompatibility complex
proteins that are involved in immune-recognition
processes; myosin motor protein in muscle control;
RRM RNA-recognition motif transcriptions factor;
trypsin-like serine proteases; laminin EGF agrowth
factor in the extracellular matrix; Ras oncoprotein
in tumorigenesis; SH2 protein domains in the phosphorylation signal cascade)
12
constructed from β-pleated sheets, are present in antibodies. Afew protein families are listed in . Table12.2 and
are presented in more detail in Chaps.23–32 of this book.
It is interesting to note which protein family is frequently
associated with which disease (. Fig.12.3). This list is
headed by protein kinases (Chap.26). It is, therefore, not
surprising that drug research in the pharmaceutical industry has intensely focused for years on the modulation and
inhibition of protein kinases. Next on the list are cadherins. These proteins are important for stabilizing cell–cell
contacts. They play arole in embryonic morphogenesis,
in signal transduction, and in the assembly of the cytoskeleton in cells. G-protein-coupled receptors, ion channels, trypsin-like serine proteases or RAS proteins are also
prominent on the list of proteins potentially associated
with disease, especially when genetically altered.
Finally, consider how the human genome differs from
other eukaryotes. Of the more than 2200 protein families
discovered in organisms with anucleus, more than 1000
are missing in the human genome. Most of these families
have specic functions in their respective organisms or
can be explained phylogenetically. For example, venoms
are found in snakes, scorpions, and insects. In plants,
proteins are found that have avery specic function for
the plant, such as nutrient storage in seeds or defense
against the attack of pests. The proteins that are absent
in humans usually perform biochemical functions that
are irrelevant to our organism, or they perform avery
specic task in lower eukaryotes.
12.5 Knock in, Knock out: Validation of
Therapeutic Concepts
Molecular biology provides a wealth of information
about how diseases develop and how their course can be
inuenced. This is the basis of the long road from the discovery to the development of anew drug. At the end of
the process, it may be found that the result, although well
planned, does not lead to the desired clinical success. It is
therefore important to have an animal model that can be
used to validate the therapeutic concept at an early stage.
Classical test models are often not available because the
disease in question does not occur in animals.
Since the 1980s, transgenic animals have been increas-
ingly used in pharmacological research. Transgenic animals
are animals in which aspecic gene has been completely
or partially turned off or replaced by ahuman gene. An
animal in which the gene is completely knocked out corresponds to an animal in which the corresponding protein is
absent or nonfunctional. Aheterozygous animal in which
the gene of only one parent is present corresponds to an
animal in which the corresponding protein is only partially
blocked. When the gene for an enzyme or receptor is affected, the effect of an inhibitor or antagonist can be simulated. The onset and progression of adisease, or the inuence of protein inhibition on adisease, can be observed in
such an animal. In this way, the relevance of atherapeutic
concept can be established before an extremely long research and development process is launched. The increased
production of aparticular protein can be induced by the
amplication of a gene. If the absence of agene causes
the overexpression of another gene, this will also become
transparent. The gene product that is then produced in increased amounts can take over the missing function of the
silenced product. In such acase, the planned therapeutic
principle would only work if the function of the other gene
product is also blocked. This question plays an important
role in the inhibition of kinases (Sect.26.2).
The knock-out method involves switching off avery
specic gene. This technique was developed by Mario
Capecchi at the University of Utah, USA, in 1987. The
sequence of the gene to be knocked out must be known.
Astructurally homologous gene is created that is not
functional, for example, because of the insertion of
astop signal. The gene is introduced into an animal and

. • Recombinant Proteins for Molecular Test Systems
the intact gene is replaced at exactly the same location.
This process is called homologous recombination or gene
targeting. Mice are particularly well suited because the
technology for manipulating their embryonic stem cells
is particularly well established. Aforeign gene, such as
ahuman gene, can also be introduced. Mice are also well
suited for this because their genome is surprisingly similar to the human genome.
To create atransgenic mouse, female mice are treated
to produce alarge number of egg cells. After fertilization,
stem cells are extracted from the embryos at avery early
stage, the blastocyst stage. They are cultured in vitro and
the desired gene is injected into the cell. This procedure
results in alow yield. Atechnique has been developed to
differentiate transfected from nontransfected cells. The
gene to be transferred is coupled to agene that confers
resistance to the cytotoxin neomycin. When cells are
treated with neomycin, only transformed cells survive.
The blastocytes are combined with blastocytes from
other mice, and the altered embryos are carried to birth
by mice. The offspring of the surrogate mothers are chimeric, meaning that they carry the genetic information
of both the donor and acceptor mice. Here, mice with
differently colored fur are selected so that the transformed mice are easily recognizable by their spotted fur.
Another method is to directly inject foreign DNA
during an early embryonic stage. Adisadvantage of random insertion of agene is the possibility of destroying
another gene, alack of expression of the new gene, or
multiple insertions. Animals from the rst litter are bred
to produce both genetically mixed, heterozygous animals
and genetically homogeneous, homozygous animals. Sophisticated techniques can even selectively turn the new
genes on and off.
In this way, transgenic animals are generated in which
hereditary diseases such as cystic brosis, Crohn’s disease, phenylketonuria, and others can be studied. Today,
animal models also exist for diseases that have different
or multiple causes, such as cancer, diabetes, rheumatoid
arthritis, and Alzheimer’s disease. Since 1988, when the
U.S. Patent Ofce granted the rst patent for atransgenic
mouse, there has been controversy over whether aliving
creature can be patented at all. European patent law will
prohibit patents if genetic modications can lead to animal suffering. Exceptions are allowed only if asignicant
medical benet is to be expected. In 1992, the rst patent
in Europe was approved for agenetically modied “cancer mouse.” Since then, many similar patents have been
led, mostly on laboratory animals. But now farm animals such as cattle and pigs are also included. Because of
their genetic similarity to humans, the European Patent
Ofce declared patents on genetically modied primates
invalid on ethical grounds in July 2020. Currently, patents on laboratory animals are not completely banned,
but should be limited to afew exceptional cases.
12.6 Recombinant Proteins for Molecular
Test Systems
Early on, pure or enriched enzymes were available for in
vitro assays, but only in cases where the material was read-
ily available, for example, human thrombin from blood.
In other cases, animal material had to be used, with all
the risks this entails, given the relevance to rational design
(see Sect.19.11). There are many proteins that cannot be
isolated in sufcient quantities or in ahomogeneous form.
The sequence determination and production of such proteins is nowadays easy. The incredibly small amount of
afew picomoles (1 pmol = 10
termine the primary structure of ashort protein segment.
From the amino acid sequence determined in this way, the
genetic code can be reconstructed into agene. It should
be noted that several base triplets can stand for aparticular amino acid (so-called degenerate code, Sect.32.7).
Aset of single-stranded oligonucleotides is synthesized
that could theoretically cover the entire original peptide
segment. These molecules can be used to nd acomplementary sequence in acDNA library. cDNA (comple-
mentary DNA) is the DNA complementary to mRNA
(messenger RNA). It is obtained from the mRNA, which
contains only the sequence needed for protein biosynthesis, by reverse transcription with areverse transcriptase (Sect.32.5). Finally, the gene is produced in larger
quantities using the PCR technique, and the amino acid
sequence is determined from its base sequence, simply
because polynucleotides are much easier to sequence.
Next, the gene is introduced into cells that are allowed
to reproduce. In afew cases, there may be difculties with
this step. In bacteria, such as the intestinal bacterium
Escherichia coli, or in yeast cells, only soluble proteins
can be produced. Some proteins accumulate in inclusion
bodies. They must be extracted, solubilized, and refolded
under specic conditions. The gene segment for asmall
protein is often fused with the gene for another protein,
and the fusion protein is then expressed. Often, the large
protein conjugate formed in the cell is more soluble and/
or better protected from metabolic degradation than
small proteins. During purication, the nonessential
part of the protein conjugate is cleaved off. Problems may
arise if the protein is not folded correctly or if several
chains (such as insulin) have to be linked by disulde
bridges. Larger proteins that require sugar groups to perform their function (glycosylation) must be produced in
cells of higher organisms, such as mammalian cells. The
production of complex proteins in insect cells has become
particularly attractive. These cells are infected with abac-
ulovirus that has incorporated the desired information
into its genome. The virus encodes the target protein, and
the insect cells provide for its production and subsequent
glycosylation. Not only enzymes, but also receptors, ion
channels, and entire signaling cascades can be produced
in cells in this way.
−12
mol) is sufcient to de-

Chapter • Gene Technology in Drug Research
12
12.7 Silencing Genes by RNA Interference
How genetically modied species can be created by intervening in the germline of an organism was described in
Sect.12.5. For example, such species may lack aparticular
gene and, therefore, agene product, or anew gene may
have been introduced. In this way, the function of certain
genes can be studied in aliving organism. This makes the
consequences of blocking the corresponding gene product
transparent before developing an effective drug. At the
end of the 1990s, another technique that made it possible
to silence genes without using mutagenesis to interfere
with the genes of the organism was developed. This work
was done by Andrew Fire and Craig Mello, who were
awarded the Nobel Prize for their achievements in 2006.
Genes are stored on DNA. For gene expression, the
coding parts of the genome are rst transcribed onto
mRNA. The ribosome then uses this transcribed information to convert the base sequence into apeptide sequence
(Sect.32.7). The idea of trapping the transcribed information on the single-stranded mRNA by adding acomplementary single-stranded RNA was rst proposed in the
early 1980s. The two strands can be joined by “hybridization,” which means that the single nucleotide strand is supplemented by the complementary strand to form adouble
strand. The resulting double-stranded RNA is no longer
suitable as atemplate for protein biosynthesis. In practice,
however, this antisense principle (Sect.32.4) did not produce the expected breakthrough results. In some cases,
genes were only weakly suppressed, and even the addition
of the normal RNA strand could achieve asuppression.
Fire and Mello suspected that neither the normal nor
the antisense strand caused gene blockage, but rather the
double-stranded form, which was present as an impurity.
Further experiments conrmed their suspicions. Interestingly, even small amounts of double-stranded RNA were
sufcient to render alarge number of mRNA molecules
unusable. In contrast, the use of antisense strands would
require stoichiometric amounts. It was further shown that
even short double-stranded RNA fragments of about
20nucleotides were sufcient to silence entire mRNA gene
sequences. Fire and Mello named this phenomenon RNA
interference. What had happened? An enzyme called dicer
chopped up the double-stranded RNA into pieces with
alength of 21–23 nucleotides, which then resulted in blockade. The double-stranded RNA pieces are incorporated
into the enzyme complex RISC (RNA-induced silencing
complex) and separated into individual strands. One strand
is released from the complex, while the other remains as
atemplate for capturing other mRNA molecules.
The sequence of the captured strand allows the RISC
complex to recognize and sequentially cleave all mRNAs
with acomplementary base sequence. They are then digested by enzymes in the cytoplasm. The cell selectively
eliminates only those mRNAs that contain the sequence
pattern complementary to the short RNA strands in
the RISC complex. In practice, this gene blockade has
proven to be simpler and more reliable than the antisense technique. RNA interference can, thus, be used to
systematically silence discovered genes in order to draw
conclusions about the resulting consequences for the organism. RNA interference is not only used for analytical
purposes. There are already biotech companies that use
small RNA fragments to silence disease-causing genes.
Another major problem in the development of RNAi
therapeutics is how to get a22-base RNA molecule into
the cell where it is supposed to work. Highly charged
molecules cannot cross the cell membrane. Therefore,
aspecial delivery system is needed. Intensive research is
underway to develop such systems, but the problem is far
from being solved. Areliable and highly efcient system
that can selectively deliver such polar and nuclease-sensitive molecules into the interior of the cell is likely to
open up acompletely new and currently unforeseeable
perspective for disease therapy.
The goal is to construct delivery systems that can
package the fragile and polar freight of RNA molecules
and dock onto the cell. Once there, the coat of these
carriers must fuse with the cell membrane or selectively
penetrate to reach the interior of the cell. One concept
is to package and compartmentalize RNA in polymers
such as polyethyleneimine. The positive charges on the
polymer backbone can bind and encapsulate anegatively
charged polymer molecule such as RNA or DNA building blocks. Other systems attempt to make the RNA or
DNA molecules bioavailable to the cell by encapsulating
them in amembrane-like coating. This packaging in liposomes results in selective adhesion of the articial cell to
the membrane of the target cell, and then the liposome
fuses with the target cell in an endocytosis-like process.
Another problem is the risk that small, silencing RNA
molecules (siRNAs) could trigger an immune response.
One solution to this dilemma is the chemical modication of siRNAs. The RNA molecules are modied in such
away that they still hybridize optimally to the targeted
segment in the mRNA, but have improved properties in
terms of transport, immunogenicity, and stability. The
OH groups of the ribose building block of the nucleotides
have been replaced by uorine, methoxy, or hydrogen.
Certainly, siRNA research is still in its infancy. The
potential of the methods seems impressive, since they use
the principles of gene regulation that are applied in Nature. As described earlier, we have genes in our genome
that encode microRNAs and that have long stretches of
sequence complementarity. Structurally, they exist as
double strands. They are cut by the dicer protein and can
be used to interfere with RNA, resulting in an alternative
means of gene regulation. For abroad therapeutic application of externally delivered RNA snips, these problems
need to be overcome. In 2018, the company Alnylam
launched the rst RNAi therapeutic, patisiran, for the
treatment of polyneuropathy in hereditary transthyre-
Соседние файлы в папке Библиотека им академика М.И. Перельмана
