Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5942_Библиотеки_им_академика_М_И_Перельмана
.pdf
184 6 The SWISS-MODEL Repository of 3D Protein Structures and Models
https://t.me/medicina_free
(a)
(b)
(c)
Figure 6.3 Structure Assessment tool results for PDB entry 3MLA. (a) A Ramachandran
Plot shows the distribution of Φ/Ψ dihedral angles. (b) Molprobity results highlight residues
of low quality. (c) Interactive view of the 3D model, highlighting selected residues of low
quality found by MolProbity. (d) Bar chart displaying the QMEANDisCo local quality
estimation above the sequence for every polypeptide chain. One-letter codes of the
residues selected in Figure 6.3b are highlighted in red. The same analysis can be performed
on a predicted model rather than a PDB structure.
that cross-reference to other databases (Figure 6.5b (1)). Figure 6.5b (2) displays
the annotation tracks, which summarize which parts of the input sequence are
covered by models or experimental structures, as well as the regions corresponding
to multiple annotations (e.g. natural variants, protein domains, catalytic sites, or
user-provided annotations). All regions are interactive and can be selected by clicking, and zoomed in or out by scrolling while pressing the control key. Figure 6.5b (3)
provides detailed information about the template used for modeling (in the case
of homology models), the version of the AlphaFold model, or the experimental
(d)

(a)
https://t.me/medicina_free
6.3 Protein Feature Annotation and Cross-References to Computational Resources 185
(b)
Figure 6.4 Assessment of conformational changes. (a) The Structure Comparison tool of
SWISS-MODEL highlights conformational changes of apo (PDB entry 3TPL, model 02,
yellow) and holo (PDB entry 3TPR, model 01, turquoise) states of Human Beta-secretase 1
(P56817, BACE1_HUMAN). This can be seen both in the 3D view (curved grey arrow) and in a
local drop in the “Consistency with Ensemble” plot (straight grey arrow). (b) Template
Search results for D. discoideum Adenylate kinase (Q54QJ9, KAD2_DICDI) show the presence
of templates in both open (2AK2) and closed (2AKY) conformations depending on the
ligands bound. The curved grey arrows show the main conformational change.
structure. It also includes information about global and local (per-residue) model
condence (for computational models) and on biologically relevant ligands as well
as protein–ligand interactions (based on PLIP).
Figure 6.5b (4) provides the interactive 3D view of the structure where annotated
regions are highlighted. If a ligand is present in the model, the protein–ligand
interaction annotations can be displayed by clicking on the “Contains 2 ligands”

186 6 The SWISS-MODEL Repository of 3D Protein Structures and Models
https://t.me/medicina_free
(a)
(b)
Figure 6.5 The SMR web interface (https://swissmodel.expasy.org/repository/). (a) A
snapshot of the main landing page, highlighting its search field. (b) Example of a typical
UniProtKB entry page is demonstrated here for the Mycothiol ligase from Mycobacterium
tuberculosis (strain ATCC 25618/H37Rv) (UniProtKB AC: P9WJM9). Zoom into the ligandbinding site can be activated when a specific ligand is selected from the “Contains n ligands”
button at the top of the protein view (here: n = 2). Circled numbers, from 1 to 7, indicate the
different sections of the entry page: (1) the entry identifier and cross-reference to other
databases, (2) the annotation tracks (including user-provided annotations), (3) information
about the selected structure (here: template used and model confidence), (4) interactive
view of the 3D model, (5) zoom in to an annotated ligand binding site, (6) tools to change
figure representation and color options, (7) list of computational and experimental models,
(8) sequence alignments as relevant for selected structure (here: target–template
alignment), and (9) links to the Structure Assessment and Structure Comparison tools.

6.3 Protein Feature Annotation and Cross-References to Computational Resources 187
https://t.me/medicina_free
button (Figure 6.5b (3)), and selecting the ligand of interest with the down arrow.
This will highlight the interacting residues in the model (Figure 6.5b (5)). The
appearance of the protein model can be changed with the options in Figure 6.5b (6),
which includes representation (e.g. cartoon or surface) and coloring mode (e.g.
amino acid properties or local quality estimates, accessible with the cogwheel
button).
Figure 6.5b (7) lists all models and experimental structures available for the target
entry, and controls what is displayed in the viewer. The initial selection is based
on several criteria: type of model (experimental or computational), oligomeric
state, coverage range, experimental method, resolution, and model quality. Finally,
Figure 6.5b (8) displays alignments between the target and the template (for
homology models) or the alignment between the experimental structure and the
UniProtKB sequence, or between the model and the UniProtKB sequence for other
models. Here, the residues are colored according to the display in the 3D viewer
(Figure 6.5b (4)) and their sequence numbering can be shown by hovering the
alignment.
For users who require structural information for a large number of proteins, the
repository provides two downloadable les for each of the 13 core species covered by
SWISS-MODEL. They are available from the landing page of the repository (https://
swissmodel.expasy.org/repository) and are updated shortly after every UniProtKB
release. The rst corresponds to the metadata, which contains an index of all the
homology-based models and experimental structures available for that organism. It
(i) lists information on the part of the protein that each model covers, (ii) provides
a link for the download of the structure’s atomic coordinates le, and (iii) provides
template and model quality information for homology-based models. The same data
are also available in tab-separated and JSON formats. The second le contains the
atomic coordinates of all the homology-based models for the organism in PDB format. It does not contain experimental structures or models built by AlphaFold, but
models that can be built with templates mapped via SIFTS are included. Those models can be useful even if the protein is covered in the PDB: (i) template selection aims
to pick a high-quality structure if multiple exist; (ii) gaps will be lled; (iii) target
sequence will be aligned with the UniProtKB sequence; and (iv) sidechains may be
adjusted when sequence identity is below 100%. Users who require a large number
of homology models for an organism that is not part of the 13 core species are free
to contact the SWISS-MODEL help desk, who can generate a bulk download with
recently updated homology models.
Programmatic access to the repository via the SMR REST API (https://swissmodel
.expasy.org/docs/repository_help#smr_api) provides most exibility, including
detailed access to per-residue information on homology models and experimental
structures. It follows the OpenAPI specications (https://www.openapis.org/),
and its full documentation is available on the Repository Help page. In addition,
the 3D-Beacons (https://3d-beacons.org) API provides summary information per
UniProtKB entry for several model providers, including those in SMR.

188 6 The SWISS-MODEL Repository of 3D Protein Structures and Models
https://t.me/medicina_free
6.4 Quality Estimates and Benchmarking
SMR aims to only include high-quality models in its database. But what does high
quality mean, especially in the context of SBDD? To answer such questions, the
CAMEO benchmarking eort has been providing blind, weekly, automated benchmarking services for structure prediction methods since 2012 [63–66]. Based on the
PDB pre-release data, which is part of the PDB release cycle, CAMEO participants
are provided with a selection of sequences of proteins whose structures are going to
be published in the following PDB release. The predictions are then compared to the
newly released experimental structures, which are considered the gold standard.
To perform this comparison, several dierent metrics are used. Calculating the
root-mean-square deviation (RMSD) of atomic positions between the reference
structure and the model is a straightforward approach, but is strongly inuenced
by outliers in poorly predicted regions, is insensitive to missing parts of the model,
and is strongly dependent on the superposition of the model with the reference
structure [46]. Another approach that has been developed in the context of CASP
is the Global Distance Test (GDT) score [67, 68], which quanties the number of
corresponding atoms in the model that can be superposed within a set of predened tolerance thresholds to the reference structure. Although it is not sensitive
to outliers and accounts for missing parts of the model, it is still dependent on
superposition. In the case of exible proteins composed of several domains, which
can naturally change their relative orientation with respect to each other, a global
rigid-body superposition is typically dominated by the largest domain, and as a
consequence, the smaller domains may not correctly match, which can result
in articially unfavorable scores. In CASP, the eects of domain movement are
mitigated by splitting the target into the so-called assessment units (AUs), which
are evaluated separately. However, this process is based on visual inspection and
remains mostly manual and often based on subjective criteria [46].
Superposition-free scores such as the local Distance Dierence Test (lDDT) [46],
which is a measure for the conservation of interatomic distances between the prediction and the reference structure, or the Contact Area Dierence (CAD) score [69],
based on dierences in residue-residue contact area using Voronoitessellation, have
solved these problems. Both methods calculate a “local” (per-residue) score, which is
often averaged into a “global” score representing the overall condence in the model.
The lDDT score and several eld-specic variants have been developed. Of particular note in the context of SBDD is the lDDT-BS score, which assesses the accuracy
of binding sites of biologically relevant ligands [64].
However, at the time of modeling, structure prediction methods do not know the
correct answer, which prevents them from providing a direct metric of model quality.
Thus, they typically report an estimation instead, which represents their expectation
of the accuracy of the model. Of course, such estimations should be accurate, which
is also assessed by CAMEO. Even for very good prediction methods, the quality of
predictions can vary from one model to another, and even between regions within
a single model. Therefore, CAMEO participants are tasked to provide reliable estimates of the quality of their predictions at a per-residue level, also known as “Model

6.5 Binding Site Conformational States 189
https://t.me/medicina_free
Condence” scores. The accuracy of the model condence scores is evaluated by
assessing the ability of those scores to distinguish between high-quality (lDDT ≥ 0.6)
and low-quality (lDDT < 0.6) residues using an area under the receiver operating
characteristic curve (ROC AUC) metric, which shows how sensitivity and specicity change as the decision threshold of local quality changes. Higher ROC AUC
indicates more accurate predictions. Over the years, benchmarking of model quality estimation methods resulted in signicant improvements in the performance of
those methods [65], and has impacted internal model quality estimates too. Prediction of local lDDT scores has become a standard way to report internal model quality
predictions, with both AlphaFold (pLDDT, which predicts lDDT-Cα,avariationof
the lDDT score restricted to Cα atoms) and SWISS-MODEL (with QMEANDisCo
which predicts the all-atom lDDT score) using comparable scores.
All computational models in SMR provide such a local condence measure, with
a method that has been critically benchmarked in CAMEO [65] and CASP [70].
Although all computational models are of high global quality (with criteria depending on the model provider), users should carefully review local quality information
before attempting to use the model for SBDD. Docking a compound in a low-quality
region is likely to produce wrong or misleading results. While there is no exact cuto, users should be particularly cautious with any region where residues show local
lDDT scores below 0.6.
6.5 Binding Site Conformational States
Most experimental structures and computational models are a snapshot in time and
do not account for conformational changes and exibility. However, proteins are
dynamic in nature and their conformation can change, e.g. upon substrate binding. The aspartyl proteases [71], adenylate kinases [72], and cytochrome P450 [73]
are three examples. In other cases, allosteric eects due to binding of a molecule
outside the active site trigger structural changes in the binding site [74]. Thus, for
computer-aided SBDD applications, it is essential to make sure that the predicted
model or experimental structure represents the conformational state that is adequate
for the specic downstream application.
With experimental structures from the PDB, such assessment is a relatively easy
task and can be directly derived from the description and presence or absence
of ligands in the structure. This can be illustrated in the example of the Human
Beta-secretase 1 (P56817, BACE1_HUMAN). A wealth of experimental structures
for this member of the aspartyl protease family is available for the bound (e.g.
3TPL)andunbound(e.g.3TPR)statesinboththePDBandSMR.TheStructure
Comparison tool (Figure 6.4a), available by clicking on the “compare” icon in the
list of structures (Figure 6.5b 9), carries out a superposition, which highlights the
structural dierences to better inform the selection of a structure in the adequate
conformation.
On the other hand, the homology models from SWISS-MODEL listed in SMR are
built from automatically selected templates. They undergo minimal postprocessing

190 6 The SWISS-MODEL Repository of 3D Protein Structures and Models
https://t.me/medicina_free
and will adopt the binding site conformational state of the template regardless of
the presence of a ligand in the nal model. Because of its strict ligand transfer rules,
SWISS-MODEL may produce a model of a bound state without the relevant ligand,
even when the model is based on a template with relatively high sequence identity.
Therefore, the absence of a ligand in a predicted structure is no guarantee that a
conformation in the unbound state was built. Instead, users need to check whether
ligands were present in the template used for modeling. They can do so by following
the link to the template used for modeling (Figure 6.5b 3).
In cases where ligand binding causes signicant structural changes, SWISSMODEL can produce models for several of the conformations of the protein if corresponding template structures are available. The adenylate kinase from P. aerug inosa
(Q9HXV4, KAD_PSEAE) is an illustrative example: since the conformational
changes on ligand binding are rather large (more than 10Å after a superposition
based on the target–template alignment), SWISS-MODEL built two models of the
protein, one in each state (open and closed).
However, in other cases, the conformational changes may also be more subtle
and not picked up automatically by the algorithm. Since there is no preferential
conformation, models may represent either open (as in Dictyostelium discoideum,
Q54QJ9, KAD2_DICDI) or closed (as in S. cerevisiae, P07170, KAD2_YEAST) states.
Alternative conformational states of the target protein can be remodeled directly
from within SMR by using the “Interactive Modeling” button above the 3D view
(Figure 6.5b 3) and clicking on the “Search For Templates” button in the next window. In the resulting page (Figure 6.4b, which shows interactive template search
results for D. discoideum, Q54QJ9, KAD2_DICDI), a selection of templates is presented which attempts to balance structural diversity with the predicted quality of
the resulting model. Using the template descriptions (Figure 6.4b, left), bound ligands and 3D viewer (Figure 6.4b, right), users can explore the diversity of the template conformations and choose an appropriate selection of templates for modeling.
Once that is done, clicking on the “Build Models” button (Figure 6.4b, top) will start
the modeling process. If needed, models based on other templates can always be
built at a later time point.
With AlphaFoldmodels, the exact conformational state of a protein cannot be controlled. It has been shown that AlphaFold predicts the holo (bound) state of a protein
in 70% of the cases [75] and users should exercise extra caution when using these
models for docking studies.
6.6 SMR and Computer-Aided Structure-based Drug
Design
Despite the recent astonishing results in protein structure prediction by the application of deep-learning-based methods [76, 77], homology-based modeling remains
the most commonly used technique in SBDD due to its short response time, its
direct relation to experimentally determined structures and well-established track
record. One of the most recent applications of the SWISS-MODEL infrastructure

6.7 Conclusion and Outlook 191
https://t.me/medicina_free
was the generation of models for proteins from SARS-CoV-2, such as the spike
protein for which a model was produced to accompany the publication of the
sequenced genome of the virus [78]. To respond to the demand for such models,
SMR introduced a dedicated page (https://swissmodel.expasy.org/repository/
species/2697049) to make models for all single proteins and relevant protein
complexes of SARS-CoV-2 publicly available. These models are generated and
annotated following the same procedures and summary formats as regular SMR
entries and are frequently updated. In addition, annotations for most “Variants of
Interest” and “Variants of Concern” are available.
Since the outbreak of COVID-19, the SWISS-MODEL infrastructure and these
models have been explored by multiple teams worldwide in the quest for a therapy, including the mapping of druggable cavities in diverse, therapeutically relevant
SARS-CoV-2 proteins [79, 80], the search for possible drugs by virtual screening
[81, 82] or drug repurposing [83–85], the design of vaccines by antibody [86, 87],
and the analysis of variant eects [88–90]. In these studies, researchers used SMR
as a source of protein structure models and their quality estimations and also utilized the SWISS-MODEL pipeline for the identication of ligand-binding sites or
the modeling of specic targets or antibody scaolds.
But the use of SMR and SWISS-MODEL predictions for SBDD goes beyond
COVID-19. For example, in a study about the toxic eect of curcumin, a avonoid
derived from the traditional medicinal plant Curcuma longa L, on female reproduction and embryo development, high-quality structure predictions of several
key curcuma-binding proteins of unknown structure were collected from SMR
and used for molecular docking in order to analyze the curcuma-binding mode
and estimate its anity [91]. Watanabe et al. followed a similar approach in a
recent study on the possible anticancer eect of rabdosianone I, a bitter diterpene
extracted from the oriental herb Isodon japonicus Hara, but in this case, the authors
opted to predict 3D structural models using their own selection of templates [92]. In
another study, Kwarteng et al. used SWISS-MODEL and its quality estimation tools
to obtain a single high-quality structural model of 5
from the Wolbachia bacteria (wALAS) and used it for virtual screening over 3200
FDA-approved drugs in the quest of treatment of neglected tropical diseases caused
by these endosymbionts and endemic in African and Latin American countries [93].
Finally, in a study about the unexpected secondary eects of two azole antifungals
by the overseen, potential inhibition of human 11β-hydroxysteroid dehydrogenase
2(11β-HSD2), Inderbinen et al. predicted structures for the human and mouse
proteins with SWISS-MODEL using multiple user-selected templates and used
dierent predicted models to inspect the eects of species-specic variations in the
binding of the two ligands [94].
′
-aminolevulinic acid synthase
6.7 Conclusion and Outlook
SMR is a source of high-quality, experimental, and computational structural models of proteins from UniProtKB. These models come from various providers: the

192 6 The SWISS-MODEL Repository of 3D Protein Structures and Models
https://t.me/medicina_free
PDB, ModelArchive, the AlphaFold DB, and the homology modeling pipeline implemented in SWISS-MODEL. The repository is updated frequently in order to assure
a comprehensive, high-quality snapshot of the current experimental and computational structural knowledge is provided. The cross-references to protein sequence
and structural feature annotation databases, the mapping of transmembrane segments, and the inclusion of ligand information transferred from templates of known
structure allow for an integrative analysis of the models, which may aid and inform
dierent approaches in SBDD.
Deep-learning-based methods, such as AlphaFold [18] and RoseTTAFold [17], are
revolutionizing the eld of protein structure prediction [48, 95], impacting various
elds of research, and could have ripple eects on many others, including SBDD
[96]. The high accuracy of such models for both proteins with and without homologs
of known structure, accompanied by their increasing speed, allows researchers and
companies to not only produce a model of their favorite protein in a timely manner
but also the generation of large sets across entire proteomes. The availabilityof these
models in online repositories such as ModelArchive or the AlphaFold DB makes it
possible for a fast bridging of the protein sequence–structure gap, which was out of
the reach of biophysical methods and homology modeling.
However, when it comes to SBDD, the impact of deep-learning-based methods is
still unclear. While the accuracy of these methods is high, it is higher for the protein
backbone than for the side chains [49], which may result in low-quality binding sites.
Downstream methods that require a highly accurate binding site and do not account
for conformational changes will fail [97]. It is thus imperative to inspect local quality estimates in binding sites and, especially for full-length protein models from the
AlphaFold DB, watch out for predicted disordered regions and uncertain domain
orientations. When homologs of known structure exist, homology modeling users
can hand-pick their templates, which allows for some control over the orientation
of side chains and overall conformational changes. Hence, our recommendation is
to always check rst if a template-based method can provide a high-quality model
and possibly rene the model with a manually selected template, and then extend
the structural coverage with models from AlphaFold if this is needed. For SBDD to
make full use of the extended structural coverage provided by AlphaFold, docking
tools able to better account for inaccurately modeled and exible binding sites are
needed.
Additionally, users must bear in mind that current structure prediction methods
are not able to include and consider post-translational modications (PTMs), such
as phosphorylations and glycosylations, which may aect conformational states
or surface accessibility. There are ongoing eorts to use deep neural networks
for PTMs [98], as well as for reliably and accurately predicting and modeling
protein–protein interaction pairs [55], and for the prediction and annotation of
functional regions [99], ligand-binding sites [100] and variant eects [101]. We
expect that the fast-paced development of structural and annotation methods will
continue in the coming years, which will drive a rapid and continuous adaptation
of model repositories and model quality estimators.

References 193
https://t.me/medicina_free
When using any model for downstream SBDD applications, such as docking,
the most critical parameter is model accuracy [102], especially at the all-atom and
side-chain conformation levels. SMR displays all PDB structures and AlphaFold
DB models available for a given UniProtKB AC, and only high-quality homology
models or manually curated models from ModelArchive. Quality is measured based
on quality estimation metrics, such as QMEANDisCo, which predict the average
conservation of interatomic distances in the neighborhood of a residue (lDDT)
when compared to the “correct” structure. The integration of protein assemblies
and the modeling of complexes with ligands and other macromolecules will likely
require the development of quality estimation metrics beyond those currently
employed to better account for protein interfaces. Future scores should also be
able to distinguish regions with disorder and exibility from modeling inaccuracies
since currently both would be represented by low lDDT values. In addition, the
development of clearer and comparable quality estimates is a must. We expect
that the CASP experiment [48] and CAMEO [63] will act as major forces in that
direction.
The eld is developing rapidly, and signicant developments will occur in the
coming years. We expect that computational model repositories and automated
method benchmarking will help researchers along this path and, by providing
state-of-the-art annotated models to everyone, remain useful tools for those who
are not directly involved in structure prediction method development.
References
1 Anderson, A.C. (2003). The process of structure-based drug design. Chemistry &
Biology 10 (9): 787–797.
2 Sliwoski, G. et al. (2014). Computational methods in drug discovery. Pharmaco-
logical Reviews 66 (1): 334–395.
3 wwPDB consortium (2019). Protein Data Bank: the single global archive for 3D
macromolecular structure data. Nucleic Acids Research 47 (D1): D520–D528.
4 Anderson, S. and Chiplin, J. (2002) ‘Structural genomics: shaping the future of
drug design?’, Drug Discovery Today, 105–107. https://doi.org/10.1016/s1359-
6446(01)02125-0.
5 Jones, M.M. et al. (2014). The structural genomics consortium: a knowledge
platform for drug discovery: a summary. Rand Health Quarterly 4 (3): 19.
6 Schapira, M. (2010). Structural genomics, its application in chemistry, biology,
and drug discovery. Burger’s Medicinal Chemistry and Drug Discovery [Preprint].
https://doi.org/10.1002/0471266949.bmc135.
7 Schwede, T. (2013). Protein modeling: what happened to the “protein structure
gap”? Structure 21 (9): 1531–1540.
8 Blundell, T.L. et al. (2006). Structural biology and bioinformatics in drug
design: opportunities and challenges for target identication and lead discovery.
Philosophical transactions of the Royal Society of London Series B, Biological
Sciences 361 (1467): 413–423.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
