Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5664_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
11 Мб
Скачать
☆
184 6 The SWISS-MODEL Repository of 3D Protein Structures and Models
https://t.me/medicina_free
(a)
(b)
(c)
Figure 6.3 Structure Assessment tool results for PDB entry 3MLA. (a) A Ramachandran Plot shows the distribution of Φ/Ψ dihedral angles. (b) Molprobity results highlight residues of low quality. (c) Interactive view of the 3D model, highlighting selected residues of low quality found by MolProbity. (d) Bar chart displaying the QMEANDisCo local quality estimation above the sequence for every polypeptide chain. One-letter codes of the residues selected in Figure 6.3b are highlighted in red. The same analysis can be performed on a predicted model rather than a PDB structure.
that cross-reference to other databases (Figure 6.5b (1)). Figure 6.5b (2) displays the annotation tracks, which summarize which parts of the input sequence are covered by models or experimental structures, as well as the regions corresponding to multiple annotations (e.g. natural variants, protein domains, catalytic sites, or user-provided annotations). All regions are interactive and can be selected by click­ing, and zoomed in or out by scrolling while pressing the control key. Figure 6.5b (3) provides detailed information about the template used for modeling (in the case of homology models), the version of the AlphaFold model, or the experimental
(d)
(a)
https://t.me/medicina_free
6.3 Protein Feature Annotation and Cross-References to Computational Resources 185
(b)
Figure 6.4 Assessment of conformational changes. (a) The Structure Comparison tool of SWISS-MODEL highlights conformational changes of apo (PDB entry 3TPL, model 02, yellow) and holo (PDB entry 3TPR, model 01, turquoise) states of Human Beta-secretase 1 (P56817, BACE1_HUMAN). This can be seen both in the 3D view (curved grey arrow) and in a local drop in the “Consistency with Ensemble” plot (straight grey arrow). (b) Template Search results for D. discoideum Adenylate kinase (Q54QJ9, KAD2_DICDI) show the presence of templates in both open (2AK2) and closed (2AKY) conformations depending on the ligands bound. The curved grey arrows show the main conformational change.
structure. It also includes information about global and local (per-residue) model condence (for computational models) and on biologically relevant ligands as well as protein–ligand interactions (based on PLIP).
Figure 6.5b (4) provides the interactive 3D view of the structure where annotated regions are highlighted. If a ligand is present in the model, the protein–ligand interaction annotations can be displayed by clicking on the “Contains 2 ligands”
186 6 The SWISS-MODEL Repository of 3D Protein Structures and Models
https://t.me/medicina_free
(a)
(b)
Figure 6.5 The SMR web interface (https://swissmodel.expasy.org/repository/). (a) A snapshot of the main landing page, highlighting its search field. (b) Example of a typical UniProtKB entry page is demonstrated here for the Mycothiol ligase from Mycobacterium tuberculosis (strain ATCC 25618/H37Rv) (UniProtKB AC: P9WJM9). Zoom into the ligand­binding site can be activated when a specific ligand is selected from the “Contains n ligands” button at the top of the protein view (here: n = 2). Circled numbers, from 1 to 7, indicate the different sections of the entry page: (1) the entry identifier and cross-reference to other databases, (2) the annotation tracks (including user-provided annotations), (3) information about the selected structure (here: template used and model confidence), (4) interactive view of the 3D model, (5) zoom in to an annotated ligand binding site, (6) tools to change figure representation and color options, (7) list of computational and experimental models, (8) sequence alignments as relevant for selected structure (here: target–template alignment), and (9) links to the Structure Assessment and Structure Comparison tools.
6.3 Protein Feature Annotation and Cross-References to Computational Resources 187
https://t.me/medicina_free
button (Figure 6.5b (3)), and selecting the ligand of interest with the down arrow. This will highlight the interacting residues in the model (Figure 6.5b (5)). The appearance of the protein model can be changed with the options in Figure 6.5b (6), which includes representation (e.g. cartoon or surface) and coloring mode (e.g. amino acid properties or local quality estimates, accessible with the cogwheel button).
Figure 6.5b (7) lists all models and experimental structures available for the target entry, and controls what is displayed in the viewer. The initial selection is based on several criteria: type of model (experimental or computational), oligomeric state, coverage range, experimental method, resolution, and model quality. Finally, Figure 6.5b (8) displays alignments between the target and the template (for homology models) or the alignment between the experimental structure and the UniProtKB sequence, or between the model and the UniProtKB sequence for other models. Here, the residues are colored according to the display in the 3D viewer (Figure 6.5b (4)) and their sequence numbering can be shown by hovering the alignment.
For users who require structural information for a large number of proteins, the repository provides two downloadable les for each of the 13 core species covered by SWISS-MODEL. They are available from the landing page of the repository (https:// swissmodel.expasy.org/repository) and are updated shortly after every UniProtKB release. The rst corresponds to the metadata, which contains an index of all the homology-based models and experimental structures available for that organism. It (i) lists information on the part of the protein that each model covers, (ii) provides a link for the download of the structure’s atomic coordinates le, and (iii) provides template and model quality information for homology-based models. The same data are also available in tab-separated and JSON formats. The second le contains the atomic coordinates of all the homology-based models for the organism in PDB for­mat. It does not contain experimental structures or models built by AlphaFold, but models that can be built with templates mapped via SIFTS are included. Those mod­els can be useful even if the protein is covered in the PDB: (i) template selection aims to pick a high-quality structure if multiple exist; (ii) gaps will be lled; (iii) target sequence will be aligned with the UniProtKB sequence; and (iv) sidechains may be adjusted when sequence identity is below 100%. Users who require a large number of homology models for an organism that is not part of the 13 core species are free to contact the SWISS-MODEL help desk, who can generate a bulk download with recently updated homology models.
Programmatic access to the repository via the SMR REST API (https://swissmodel .expasy.org/docs/repository_help#smr_api) provides most exibility, including detailed access to per-residue information on homology models and experimental structures. It follows the OpenAPI specications (https://www.openapis.org/), and its full documentation is available on the Repository Help page. In addition, the 3D-Beacons (https://3d-beacons.org) API provides summary information per UniProtKB entry for several model providers, including those in SMR.
188 6 The SWISS-MODEL Repository of 3D Protein Structures and Models
https://t.me/medicina_free
6.4 Quality Estimates and Benchmarking
SMR aims to only include high-quality models in its database. But what does high quality mean, especially in the context of SBDD? To answer such questions, the CAMEO benchmarking eort has been providing blind, weekly, automated bench­marking services for structure prediction methods since 2012 [63–66]. Based on the PDB pre-release data, which is part of the PDB release cycle, CAMEO participants are provided with a selection of sequences of proteins whose structures are going to be published in the following PDB release. The predictions are then compared to the newly released experimental structures, which are considered the gold standard.
To perform this comparison, several dierent metrics are used. Calculating the root-mean-square deviation (RMSD) of atomic positions between the reference structure and the model is a straightforward approach, but is strongly inuenced by outliers in poorly predicted regions, is insensitive to missing parts of the model, and is strongly dependent on the superposition of the model with the reference structure [46]. Another approach that has been developed in the context of CASP is the Global Distance Test (GDT) score [67, 68], which quanties the number of corresponding atoms in the model that can be superposed within a set of prede­ned tolerance thresholds to the reference structure. Although it is not sensitive to outliers and accounts for missing parts of the model, it is still dependent on superposition. In the case of exible proteins composed of several domains, which can naturally change their relative orientation with respect to each other, a global rigid-body superposition is typically dominated by the largest domain, and as a consequence, the smaller domains may not correctly match, which can result in articially unfavorable scores. In CASP, the eects of domain movement are mitigated by splitting the target into the so-called assessment units (AUs), which are evaluated separately. However, this process is based on visual inspection and remains mostly manual and often based on subjective criteria [46].
Superposition-free scores such as the local Distance Dierence Test (lDDT) [46], which is a measure for the conservation of interatomic distances between the predic­tion and the reference structure, or the Contact Area Dierence (CAD) score [69], based on dierences in residue-residue contact area using Voronoitessellation, have solved these problems. Both methods calculate a “local” (per-residue) score, which is often averaged into a “global” score representing the overall condence in the model. The lDDT score and several eld-specic variants have been developed. Of particu­lar note in the context of SBDD is the lDDT-BS score, which assesses the accuracy of binding sites of biologically relevant ligands [64].
However, at the time of modeling, structure prediction methods do not know the correct answer, which prevents them from providing a direct metric of model quality. Thus, they typically report an estimation instead, which represents their expectation of the accuracy of the model. Of course, such estimations should be accurate, which is also assessed by CAMEO. Even for very good prediction methods, the quality of predictions can vary from one model to another, and even between regions within a single model. Therefore, CAMEO participants are tasked to provide reliable esti­mates of the quality of their predictions at a per-residue level, also known as “Model
6.5 Binding Site Conformational States 189
https://t.me/medicina_free
Condence” scores. The accuracy of the model condence scores is evaluated by assessing the ability of those scores to distinguish between high-quality (lDDT ≥ 0.6) and low-quality (lDDT < 0.6) residues using an area under the receiver operating characteristic curve (ROC AUC) metric, which shows how sensitivity and speci­city change as the decision threshold of local quality changes. Higher ROC AUC indicates more accurate predictions. Over the years, benchmarking of model qual­ity estimation methods resulted in signicant improvements in the performance of those methods [65], and has impacted internal model quality estimates too. Predic­tion of local lDDT scores has become a standard way to report internal model quality predictions, with both AlphaFold (pLDDT, which predicts lDDT-Cα,avariationof the lDDT score restricted to Cα atoms) and SWISS-MODEL (with QMEANDisCo which predicts the all-atom lDDT score) using comparable scores.
All computational models in SMR provide such a local condence measure, with a method that has been critically benchmarked in CAMEO [65] and CASP [70]. Although all computational models are of high global quality (with criteria depend­ing on the model provider), users should carefully review local quality information before attempting to use the model for SBDD. Docking a compound in a low-quality region is likely to produce wrong or misleading results. While there is no exact cut­o, users should be particularly cautious with any region where residues show local lDDT scores below 0.6.
6.5 Binding Site Conformational States
Most experimental structures and computational models are a snapshot in time and do not account for conformational changes and exibility. However, proteins are dynamic in nature and their conformation can change, e.g. upon substrate bind­ing. The aspartyl proteases [71], adenylate kinases [72], and cytochrome P450 [73] are three examples. In other cases, allosteric eects due to binding of a molecule outside the active site trigger structural changes in the binding site [74]. Thus, for computer-aided SBDD applications, it is essential to make sure that the predicted model or experimental structure represents the conformational state that is adequate for the specic downstream application.
With experimental structures from the PDB, such assessment is a relatively easy task and can be directly derived from the description and presence or absence of ligands in the structure. This can be illustrated in the example of the Human Beta-secretase 1 (P56817, BACE1_HUMAN). A wealth of experimental structures for this member of the aspartyl protease family is available for the bound (e.g. 3TPL)andunbound(e.g.3TPR)statesinboththePDBandSMR.TheStructure Comparison tool (Figure 6.4a), available by clicking on the “compare” icon in the list of structures (Figure 6.5b 9), carries out a superposition, which highlights the structural dierences to better inform the selection of a structure in the adequate conformation.
On the other hand, the homology models from SWISS-MODEL listed in SMR are built from automatically selected templates. They undergo minimal postprocessing
190 6 The SWISS-MODEL Repository of 3D Protein Structures and Models
https://t.me/medicina_free
and will adopt the binding site conformational state of the template regardless of the presence of a ligand in the nal model. Because of its strict ligand transfer rules, SWISS-MODEL may produce a model of a bound state without the relevant ligand, even when the model is based on a template with relatively high sequence identity. Therefore, the absence of a ligand in a predicted structure is no guarantee that a conformation in the unbound state was built. Instead, users need to check whether ligands were present in the template used for modeling. They can do so by following the link to the template used for modeling (Figure 6.5b 3).
In cases where ligand binding causes signicant structural changes, SWISS­MODEL can produce models for several of the conformations of the protein if corre­sponding template structures are available. The adenylate kinase from P. aerug inosa (Q9HXV4, KAD_PSEAE) is an illustrative example: since the conformational changes on ligand binding are rather large (more than 10Å after a superposition based on the target–template alignment), SWISS-MODEL built two models of the protein, one in each state (open and closed).
However, in other cases, the conformational changes may also be more subtle and not picked up automatically by the algorithm. Since there is no preferential conformation, models may represent either open (as in Dictyostelium discoideum, Q54QJ9, KAD2_DICDI) or closed (as in S. cerevisiae, P07170, KAD2_YEAST) states. Alternative conformational states of the target protein can be remodeled directly from within SMR by using the “Interactive Modeling” button above the 3D view (Figure 6.5b 3) and clicking on the “Search For Templates” button in the next win­dow. In the resulting page (Figure 6.4b, which shows interactive template search results for D. discoideum, Q54QJ9, KAD2_DICDI), a selection of templates is pre­sented which attempts to balance structural diversity with the predicted quality of the resulting model. Using the template descriptions (Figure 6.4b, left), bound lig­ands and 3D viewer (Figure 6.4b, right), users can explore the diversity of the tem­plate conformations and choose an appropriate selection of templates for modeling. Once that is done, clicking on the “Build Models” button (Figure 6.4b, top) will start the modeling process. If needed, models based on other templates can always be built at a later time point.
With AlphaFoldmodels, the exact conformational state of a protein cannot be con­trolled. It has been shown that AlphaFold predicts the holo (bound) state of a protein in 70% of the cases [75] and users should exercise extra caution when using these models for docking studies.
6.6 SMR and Computer-Aided Structure-based Drug Design
Despite the recent astonishing results in protein structure prediction by the appli­cation of deep-learning-based methods [76, 77], homology-based modeling remains the most commonly used technique in SBDD due to its short response time, its direct relation to experimentally determined structures and well-established track record. One of the most recent applications of the SWISS-MODEL infrastructure
6.7 Conclusion and Outlook 191
https://t.me/medicina_free
was the generation of models for proteins from SARS-CoV-2, such as the spike protein for which a model was produced to accompany the publication of the sequenced genome of the virus [78]. To respond to the demand for such models, SMR introduced a dedicated page (https://swissmodel.expasy.org/repository/ species/2697049) to make models for all single proteins and relevant protein complexes of SARS-CoV-2 publicly available. These models are generated and annotated following the same procedures and summary formats as regular SMR entries and are frequently updated. In addition, annotations for most “Variants of Interest” and “Variants of Concern” are available.
Since the outbreak of COVID-19, the SWISS-MODEL infrastructure and these models have been explored by multiple teams worldwide in the quest for a ther­apy, including the mapping of druggable cavities in diverse, therapeutically relevant SARS-CoV-2 proteins [79, 80], the search for possible drugs by virtual screening [81, 82] or drug repurposing [83–85], the design of vaccines by antibody [86, 87], and the analysis of variant eects [88–90]. In these studies, researchers used SMR as a source of protein structure models and their quality estimations and also uti­lized the SWISS-MODEL pipeline for the identication of ligand-binding sites or the modeling of specic targets or antibody scaolds.
But the use of SMR and SWISS-MODEL predictions for SBDD goes beyond COVID-19. For example, in a study about the toxic eect of curcumin, a avonoid derived from the traditional medicinal plant Curcuma longa L, on female repro­duction and embryo development, high-quality structure predictions of several key curcuma-binding proteins of unknown structure were collected from SMR and used for molecular docking in order to analyze the curcuma-binding mode and estimate its anity [91]. Watanabe et al. followed a similar approach in a recent study on the possible anticancer eect of rabdosianone I, a bitter diterpene extracted from the oriental herb Isodon japonicus Hara, but in this case, the authors opted to predict 3D structural models using their own selection of templates [92]. In another study, Kwarteng et al. used SWISS-MODEL and its quality estimation tools to obtain a single high-quality structural model of 5 from the Wolbachia bacteria (wALAS) and used it for virtual screening over 3200 FDA-approved drugs in the quest of treatment of neglected tropical diseases caused by these endosymbionts and endemic in African and Latin American countries [93]. Finally, in a study about the unexpected secondary eects of two azole antifungals by the overseen, potential inhibition of human 11β-hydroxysteroid dehydrogenase 2(11β-HSD2), Inderbinen et al. predicted structures for the human and mouse proteins with SWISS-MODEL using multiple user-selected templates and used dierent predicted models to inspect the eects of species-specic variations in the binding of the two ligands [94].
′
-aminolevulinic acid synthase
6.7 Conclusion and Outlook
SMR is a source of high-quality, experimental, and computational structural mod­els of proteins from UniProtKB. These models come from various providers: the
192 6 The SWISS-MODEL Repository of 3D Protein Structures and Models
https://t.me/medicina_free
PDB, ModelArchive, the AlphaFold DB, and the homology modeling pipeline imple­mented in SWISS-MODEL. The repository is updated frequently in order to assure a comprehensive, high-quality snapshot of the current experimental and computa­tional structural knowledge is provided. The cross-references to protein sequence and structural feature annotation databases, the mapping of transmembrane seg­ments, and the inclusion of ligand information transferred from templates of known structure allow for an integrative analysis of the models, which may aid and inform dierent approaches in SBDD.
Deep-learning-based methods, such as AlphaFold [18] and RoseTTAFold [17], are revolutionizing the eld of protein structure prediction [48, 95], impacting various elds of research, and could have ripple eects on many others, including SBDD [96]. The high accuracy of such models for both proteins with and without homologs of known structure, accompanied by their increasing speed, allows researchers and companies to not only produce a model of their favorite protein in a timely manner but also the generation of large sets across entire proteomes. The availabilityof these models in online repositories such as ModelArchive or the AlphaFold DB makes it possible for a fast bridging of the protein sequence–structure gap, which was out of the reach of biophysical methods and homology modeling.
However, when it comes to SBDD, the impact of deep-learning-based methods is still unclear. While the accuracy of these methods is high, it is higher for the protein backbone than for the side chains [49], which may result in low-quality binding sites. Downstream methods that require a highly accurate binding site and do not account for conformational changes will fail [97]. It is thus imperative to inspect local qual­ity estimates in binding sites and, especially for full-length protein models from the AlphaFold DB, watch out for predicted disordered regions and uncertain domain orientations. When homologs of known structure exist, homology modeling users can hand-pick their templates, which allows for some control over the orientation of side chains and overall conformational changes. Hence, our recommendation is to always check rst if a template-based method can provide a high-quality model and possibly rene the model with a manually selected template, and then extend the structural coverage with models from AlphaFold if this is needed. For SBDD to make full use of the extended structural coverage provided by AlphaFold, docking tools able to better account for inaccurately modeled and exible binding sites are needed.
Additionally, users must bear in mind that current structure prediction methods are not able to include and consider post-translational modications (PTMs), such as phosphorylations and glycosylations, which may aect conformational states or surface accessibility. There are ongoing eorts to use deep neural networks for PTMs [98], as well as for reliably and accurately predicting and modeling protein–protein interaction pairs [55], and for the prediction and annotation of functional regions [99], ligand-binding sites [100] and variant eects [101]. We expect that the fast-paced development of structural and annotation methods will continue in the coming years, which will drive a rapid and continuous adaptation of model repositories and model quality estimators.
References 193
https://t.me/medicina_free
When using any model for downstream SBDD applications, such as docking, the most critical parameter is model accuracy [102], especially at the all-atom and side-chain conformation levels. SMR displays all PDB structures and AlphaFold DB models available for a given UniProtKB AC, and only high-quality homology models or manually curated models from ModelArchive. Quality is measured based on quality estimation metrics, such as QMEANDisCo, which predict the average conservation of interatomic distances in the neighborhood of a residue (lDDT) when compared to the “correct” structure. The integration of protein assemblies and the modeling of complexes with ligands and other macromolecules will likely require the development of quality estimation metrics beyond those currently employed to better account for protein interfaces. Future scores should also be able to distinguish regions with disorder and exibility from modeling inaccuracies since currently both would be represented by low lDDT values. In addition, the development of clearer and comparable quality estimates is a must. We expect that the CASP experiment [48] and CAMEO [63] will act as major forces in that direction.
The eld is developing rapidly, and signicant developments will occur in the coming years. We expect that computational model repositories and automated method benchmarking will help researchers along this path and, by providing state-of-the-art annotated models to everyone, remain useful tools for those who are not directly involved in structure prediction method development.
References
1 Anderson, A.C. (2003). The process of structure-based drug design. Chemistry &
Biology 10 (9): 787–797.
2 Sliwoski, G. et al. (2014). Computational methods in drug discovery. Pharmaco-
logical Reviews 66 (1): 334–395.
3 wwPDB consortium (2019). Protein Data Bank: the single global archive for 3D
macromolecular structure data. Nucleic Acids Research 47 (D1): D520–D528.
4 Anderson, S. and Chiplin, J. (2002) ‘Structural genomics: shaping the future of
drug design?’, Drug Discovery Today, 105–107. https://doi.org/10.1016/s1359- 6446(01)02125-0.
5 Jones, M.M. et al. (2014). The structural genomics consortium: a knowledge
platform for drug discovery: a summary. Rand Health Quarterly 4 (3): 19.
6 Schapira, M. (2010). Structural genomics, its application in chemistry, biology,
and drug discovery. Burger’s Medicinal Chemistry and Drug Discovery [Preprint]. https://doi.org/10.1002/0471266949.bmc135.
7 Schwede, T. (2013). Protein modeling: what happened to the “protein structure
gap”? Structure 21 (9): 1531–1540.
8 Blundell, T.L. et al. (2006). Structural biology and bioinformatics in drug
design: opportunities and challenges for target identication and lead discovery.
Philosophical transactions of the Royal Society of London Series B, Biological Sciences 361 (1467): 413–423.