Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5418_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
11 Мб
Скачать
☆
250 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
Figure 8.12 Filter value enrichment.
● Top left: In a target list, the filters show the counts of targets that have each annotation.
In this example, the counts of targets associated with each disease are shown for the complete unfiltered target list. Some diseases are associated with a very large number of targets.
● Top right: In a filtered list, these value counts are often affected by this bias in the data
toward heavily documented filter values. In this example, this list is filtered to targets that interact with ATP1B4. Many of the top diseases in this filtered list are the same diseases as top diseases in the unfiltered list. The “calculate enrichment” button here will calculate the probability that the measured counts are explainable by random chance.
● Bottom: Results of an enrichment calculation for targets that interact with ATP1B4. Chronic
heart failure is associated with 11% of the targets in the list (Count: 8 targets, Observed Frequency: 0.11). The Expected Frequency, based on the full dataset, is 0.23%. Based on Fisher’s Exact Test, the probability of finding that many targets by random chance is 4e-12. p-adjust is calculated based on an adjustment for performing a large number of tests, controlling the False Discovery Rate to 0.05.
8.3 Use Cases 251
https://t.me/medicina_free
concept helps validate the method and lends credence to the idea of discovering relevant annotations for less well-studied targets.
8.3.1.3 Downloading Data
Easy data extraction is a key requirement for many researchers. Researchers study­ing ATP1B4 may build a download query and select the “Associated Disease” elds to download all the associated disease data for targets in the list. Users wanting to compile a list of approved drugs or active ligands for targets in the list would select the “Drugs and Ligands” elds. Heat maps and the Sequence Alignment component also contain shortcuts for initiating a data download with the appropriate elds.
8.3.1.4 Variations on this Use Case
Other Ways to Expand the Dataset
The path chosen in this use case was to expand the search for documentation from one dark target to a list of related targets by compiling a list of interacting targets via the “Explore Interacting Targets” button in the Protein–Protein Interactions compo­nent of the target details page.
Another way to expand the dataset would be to compile a list of targets that share one or more of the annotations that exist on this sparsely documented target. Beyond simply clicking what annotations are there (i.e. PANTHER Class for cation trans­porter) to generate a list of targets, users can also compile lists of targets based on having any overlap in documentation through the “Find Similar Targets” buttons (Section 8.2.4.3) that are available on a number of components including GO Terms and Pathways.
Alternatively, the Protein Sequence and Structure component allows users to easily initiate a search for targets with a homologous protein sequence (Section 8.2.4.2). A Sequence search can also be initiated from a target list page via a button high­lighted in Figure 8.2 panel J. Figure 8.13 shows the resulting Sequence Alignment component on the resulting List Analysis page that illustrates, which regions of the sequence are matching in each entry in the target list.
It should also be noted that these expansions can be compounded, for example by nding a list of targets with a homologous sequence, and then ltering that list by PANTHER Class or GO term, thereby nding targets with multiple points of similarity.
Other Filters to use for Enrichment Calculation
In addition to looking for Associated Diseases that are enriched in the target list, other interesting lters to try might be IDG Family, PANTHER Class, DTO Class, Pathways, GO Functions, or GO Processes. There are no restrictions on which l­ters are available to calculate enrichment, beyond the requirement that the lter is categorical (i.e. not numerical).
8.3.2 Characterizing a Novel Chemical Compound
This use case is for a researcher who would like to explore the potential eects of a novel chemical compound. For the sake of the example, we will use a “novel”
252 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
Figure 8.13 Sequence search.
● Sequence Search Results in the List Analysis tab of the resulting target list. The top plot is
a density plot of aligned residues for all the aligning sequences. Below is a representation of each matching target’s region of alignment, color-coded by the percent identity of the alignment.
compound represented by the following SMILES: CC1CC(O)(CCN1CCCC(=O)C1= CC=C(F)C=C1)C1=CC=C(Cl)C=C1. Note that this is a slightly modied version of haloperidol that is just for the purposes of this example, to help illustrate the con­cepts. Our researcher has two main questions: What targets might be aected by this compound? What eects might it have on the human body?
8.3.2.1 Finding Predicted Targets
Since this is a novel chemical compound, there would be no primary documenta­tion to review, as was the case with the dark target in Section 8.3.1. The researcher would begin on the ligand list page, and follow the link for “Structure Search” or navigate to https://pharos.nih.gov/structure. Starting on the Structure Search page (Figure 8.14), the user can paste in the SMILES to attempt to resolve the structure to a known compound and render the compound in the MarvinJS widget.
Following the search link in the “Find Predicted Targets” card will fetch the list of targets from NCATS Predictor [30] as described in Section 8.2.4.4. The ensuing target list (Figure 8.15) is a compilation of proteins predicted as targets of the query structure, and the targets that have a known activity to the compound in the database if a matching compound was found. Since our example is a novel compound, only predicted data are available. Predicted data are highlighted in the Target Card, and in the lter panel with a shaded background.
Details about the prediction are shown in the Table View of the list pages, which includes the predicted activity, the activity for the nearest compound from the model training set, and a measure of the Applicability Domain. Prediction applicability is quantied by the Tanimoto similarity between the query structure and the nearest
8.3 Use Cases 253
https://t.me/medicina_free
Figure 8.14 Searching by chemical structure.
● Pharos’ Structure Search page allows users to initiate a search based on any chemical
structure. Many types of inputs are used to resolve into a SMILES string to load into the Marvin JS Sketcher. The Sketcher can then be used to edit the query structure, upload a file, or draw a compound from scratch.
● Options are available to find similar ligands to the query structure or find targets predicted
to interact with the query structure.
compound from the training set. Users can lter the list based on the numeric lter for prediction applicability, or based on the numeric lter for the predicted activity.
For our example compound, NCATS Predictor found 25 targets with predicted activity. There are several target lters to consider when the list consists of targets with activity against a given compound. Calculating lter value enrichment, as described in more detail in section “Filter Value Enrichment”, yields some interest­ing results. Figure 8.16 shows the top ve GO Functions, and the top ve Associated Diseases, for this list of predicted targets. The presence of serotonin-related GO Terms, and depression-related diseases, are what we should expect from this “novel” compound that is so closely related to haloperidol.
254 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
Figure 8.15 Predicted targets.
● Sample results from a target list consisting of targets with predicted activity to the query
structure.
● The Target Prediction Details panel shows the predicted activity the query structure will
have for the target. The nearest activity for this compound from the training set determines the prediction applicability, which is quantified by the Tanimoto coefficient shown over the ≈ between the query structure and the nearest structure from the training set.
● The list can be sorted by a number of different fields. This list has been sorted according
to the potency of the predicted activity to the query structure.
● Experimentally determined activity will also be shown here for targets with a known activ-
ity to the query structure.
8.3.2.2 Analyzing Similar Ligands
Another method of investigating the potential targets or target families for a novel compound is in the analysis of the eects of a set of similar compounds. As in the pre­vious search for predicted targets, a Similarity Search is initiated from the Structure Search page (Figure 8.14 bottom). Section 8.2.4.1 describes how this functionality works.
Pharos will display a list of all compounds from TCRD that are structurally similar to the query structure (Figure 8.17). In the TableView of the ligand list, the 2D struc­tures are presented in order of decreasing similarity for users to browse and choose the best similarity cuto. A Tanimoto coecient of 0.6 is by default used to cast a wide net for matching compounds. It is usually prudent to restrict the similarity fur­ther using the numeric lter for the similarity score. Figure 8.17 shows the resulting ligand list page after searching for similar compounds to our query structure, and subsequently ltering the result list to those ligands with a Structure Similarity (i.e. Tanimoto coecient on molecular ngerprints) greater than 0.8. This results in a list of 66 ligands from TCRD.
8.3 Use Cases 255
https://t.me/medicina_free
Figure 8.16 Enriched filter values in a list of predicted targets.
● Top : The top five GO Functions that are overrepresented in a list of predicted targets.
● Bottom: The top five Associated Diseases that are overrepresented in a list of predicted
targets.
Figure 8.17 Similarity search results.
● Similarity search results for a sample query structure that was the methylated analog of
haloperidol.
● The list shown here is filtered to ligands with a Structure Similarity (i.e. Tanimoto coeffi-
cient
on molecular fingerprints) greater than 0.8.
256 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
Figure 8.18 Enriched filter values in a list of similar compounds.
● In a list of compounds found through a Similarity Search to a methylated analog of
haloperidol, filter value enrichment is calculated for the PANTHER Class filter and the Target filter.
● The PANTHER Class filter in a ligand list represents the count of ligands in the list that
have activity for a target in each class. The enrichment scores show that G-protein coupled receptors are highly overrepresented in the list, suggesting that the query structure may also have activity for some G-protein coupled receptors.
● The Target filter in a ligand list represents the simple count of ligands in the list that have
an activity to each target. The enrichment scores show that several dopamine receptors are overrepresented in the list, suggesting that the query structure may also have activity for dopamine receptors.
There are several lters in this ligand list that can help understand the poten­tial eect of the query compound. Similar to the workow in Section 8.3.2.1, lter value enrichment for the PANTHER Class lter, and for the Target lter, show that G-protein coupled receptors, and specically, dopamine receptors, are overrepre­sented in the list of similar compounds to our query structure (Figure 8.18).
8.3.2.3 Ligand Details Pages
As always for the listing pages, clicking entries in the list will take the user to the corresponding details page. Ligand details pages (Figure 8.19) show all data com­piled for each compound and primarily consist of a rendering of the 2D structure of the compound, a listing of known synonyms or other identiers for the compound, and a listing of associated targets and all measured activities against those targets. The ligand details page can also serve as a jumping-o point for performing either a Substructure Search, a Similarity Search for other compounds in TCRD, or Exploring the targets in a target list page.
A table of target activities is displayed (Figure 8.19 bottom), along with a button to translate the list of targets into a more interactive target list page. That list page will also include predicted activities, based on functionality described in Section 8.2.4.4.
8.3 Use Cases 257
https://t.me/medicina_free
Figure 8.19 Ligand details.
● Top , Ligand Summary – This component shows the 2D structure and equivalent IDs. A
button allows users to initiate either a Similarity Search or a Substructure Search, of the database based on the current ligand as a starting point.
● Bottom, Target Activities – All experimentally determined target activities are shown here.
A button allows users to translate this view into a target list page, which will include targets this compound that is predicted to have activity for.
8.3.2.4 Variations on this Use Case
Other Ligand List filters
Ligand lists have lters that can help users nd compounds that meet their requirements. Some of those lters were introduced in Section 8.3.2.2. Additionally, Figure 8.17 includes the Type lter, which reports whether or not the compounds have achieved FDA approval. In this example, three compounds are approved drugs, while 63 are other active ligands.
The number of targets each compound has an activity for, a measure of target specicity, is reected in the Target Count lter (Figure 8.20). This can be used to lter the list based on how selective the compounds are.
258 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
Figure 8.20 Target selectivity filters in a ligand list.
● Top: The Target Count filter in a ligand list represents the number of ligands in the list that
have documented activity to different numbers of targets. Compounds with a high number are more promiscuous, while compounds with only one or two known target activities may be selective or less well-studied. The sliders at the bottom of the histogram can be used to filter the ligand list to only compounds that meet the requirements the user sets.
● Bottom: An UpSet Chart is constructed for the Target filter on a ligand list. The counts of
compounds in the list with each combination of active targets are represented. Clicking on the bars or circles will filter the list to only those compounds that have the required combination of active targets.
Using the UpSet Plots
Another useful method to dive into the specicity is to build an UpSet plot [56] for the Target lter. On the List Analysis tab of the list pages, the Filter Visualizations component shows Donut Charts and UpSet Charts for the dierent lters, depend­ing on how many values an entry in the list can have for that lter. If an entry can have multiple values (e.g. one compound can have activity for multiple targets), an
8.3 Use Cases 259
https://t.me/medicina_free
UpSet plot is shown to help the user understand the overlap of dierent lter values present in the list.
Figure 8.20 bottom shows an UpSet plot for the Target lter in this ligand list. This was constructed by selecting the Target lter from the list of buttons along the top, then selecting the targets that correspond to dopamine receptors using the “Se­lect Filter Values” button, appearing below the rst generated graph (for all targets). Interpreting the UpSet chart tells us how many compounds in the list have activity for each combination of lter values. For example, Figure 8.20 tells us there are four­teen compounds that have activity for the D2, D3, and D4 dopamine receptors, and not D1A, D1B, or the dopamine transporter. Furthermore, clicking the appropriate bar on the graph can lter the list to those fourteen compounds for follow-up.
Using the Heat Maps
The List Analysis tab will also display buttons to create heat maps (Figure 8.21). Target list pages can construct heat maps of targets in the list vs. their associated diseases, active ligands, or protein–protein interacting partners. Disease and ligand list pages can construct heat maps of the elements in the list vs. their associated targets. Heat maps are interactive in their sorting and can be used to drill down into the details of each cell in the heat map. Figure 8.21 shows the heat map of the average reported activities between compounds in the list and the associated targets for the current use case. A Details View is shown as would appear when the user clicks a cell of the heat map. Links within the Details View will navigate to the appropriate details pages for the corresponding clicked cell of the heat map.
Generating a Ligand List Based on a Target
The target details pages have Drugs and Ligands components (See Sections 8.2.3.2 and “Behavioral Data”) that display all the active compounds that TCRD has for a given target. Those components have a link to explore the list of compounds on a lig­and list page (See Figure 8.7). Ligand lists based on a target are an obvious use case for nding selective compounds, using the Target Count lter (see section “Other Ligand List lters”), or using the UpSet plot (see section “Using the UpSet Plots”). A special feature of ligand lists based on an associated target is the potency lter that becomes available. Figure 8.22 shows how the Target Count lter and the Potency lter can be used together to lter the list to only compounds with the required speci­city and potency for the given target.
Uploading a Custom Ligand List
The functionality provided by a ligand list page can be used to explore common­alities among a set of ligands identied to be active against a target in a biological assay. From a ligand list page, the Upload button (Figure 8.2) can be used to resolve compounds by a number of chemical identiers (see Section 8.2.4.1), and the List Analysis features such as lter value enrichment, and heat maps are useful for this task as well.
Given a chemist who has screened thousands of compounds in a cell-based phe­notypic assay, the resulting hit list would be loaded into Pharos. The workows