Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5664_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
11 Мб
Скачать
☆
246 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
are 70 targets with documented physical interaction or text-mined interaction. These 70 targets will be the basis of the target list analysis presented in Section 8.3.1.2.
For the Pathways component, data are aggregated from ve data sources: Reac­tome [45], KEGG [47], PathwayCommons [48], UniProt [19], and WikiPathways [49]. The Reactome Pathway tab includes an interactive widget showing a graphi­cal representation of each annotated pathway, which highlights other targets in the pathways according to their TDL. All pathway annotations can be used as a start­ing point to generate a target list page that includes all the targets in each pathway. With pathway data users can generate a target list page of similar targets, a list of all targets that share any pathway annotation with the chosen target, sorted by the degree of overlap in their pathway annotations. Details on that calculation are in Section 8.2.4.3.
Phenotypic Data
Gene Ontology (GO) [50, 51] denes terms for Molecular Functions, Biological Processes, and Cellular Components that represent the normal function of gene products. From the Gene Ontology Terms component, Pharos shows each annotation that has been assigned to the target and can generate ltered lists of targets for any annotation.
The Disease Associations component (Figure 8.9 – Example GPR68) includes data from eight sources, including CTD [14], DisGeNET [15], DrugCentral [9], eRam [16], Expression Atlas [17], JensenLab [7], Monarch [18], and UniProt [19]. Users can pivot from a target details page showing a list of associated diseases to a disease list page where there are more analysis capabilities and visualizations for the list of diseases. Our example target has no entries for direct disease association annotations.
The Disease Novelty component (Figure 8.10 – Example DRD2) highlights data from Tin-X [52] which discovers target–disease relationships through natural language processing of PubMed abstracts. Tin-X quanties two metrics represent­ing the Importance, the strength of the association between the target and the disease, and the Disease Novelty, the relative scarcity of publications about a disease. Target–disease associations that have high Importance, and high Target Novelty, are promising areas for following up. Figure 8.10 shows the Tin–X associations found for DRD2, with highlights on diseases related to substance-related disorders. Note the tendency for target–disease associations for DRD2 and substance-related disorders are high-scoring associations in both Importance and Novelty.
The GWAS traits component (Figure 8.11 – Example DRD2) shows traits identi- ed through genome-wide association studies (GWAS) [53]. GWAS traits are sorted according to an Evidence Score (Mean Rank Score), which is a ranking of the traits based on the number and power of the GWAS studies that found the association. These rankings of GWAS traits are calculated by Target Illumination GWAS Analyt­ics (TIGA) [54] before being ingested into TCRD.
Similar to the Pathways data, a list of Similar Targets can be generated based on GO Terms, Associated Diseases, or GWAS Traits, as well as many other data points from this section. For example, a user could generate a target list that shares any of the associated diseases as the target of interest. Presenting the list in order of the
8.3 Use Cases 247
https://t.me/medicina_free
Figure 8.9 Associated diseases.
● A pageable table of associated diseases for the current target. Entries are expandable to
show details of the association from each data source.
● Users can navigate to the disease details page through the Explore Disease button.
● The Explore Associated Diseases button will transfer the diseases in this table to a disease
list page where further analysis can be done.
● The Find Similar Targets button will construct a target list that shares any associated dis-
eases with the target of interest, and sort them by the degree of overlap in the sets of diseases, as described in Section 8.2.4.3.
degree of overlap allows them to easily nd other targets that have the most similar sets of those associated diseases, or whichever data eld the similarity calculation was based on.
Resources and Publications
Researchers can learn which typical experimental models (17 species supported) have orthologous versions of the target. Additionally, resources, such as genetic constructs, cells, mice, chemical tools, or data resources, generated by IDG grant awardees may be possible to acquire for research. The last section of the target details pages contains temporal plots of bibliometric scores [7, 55] and patent counts that inform users about the popularity of research into a particular target. Text-mined references from JensenLab, and manually curated GeneRIFs (https:// www.ncbi.nlm.nih.gov/gene/about-generif) are available to review.
8.3.1.2 List Analysis
So far in this use case, as a researcher investigating the role of the dark tar­get ATP1B4, we learned that there is some data for the target, but notably no known associated diseases. Next, we will continue our search for possible disease associations by expanding the search to related targets. One promising idea is to
248 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
Figure 8.10 Disease novelty.
● This component highlights data from Tin-X [52] which discovers target–disease relation-
ships through natural language processing of PubMed abstracts.
● Left shows a scatterplot of the Importance, the strength of the association between the
target and the disease, versus the Disease Novelty, and the relative scarcity of publications about a disease.
● Right shows a circle-pack plot of the Importance metric for each disease association.
Selecting different regions within the disease hierarchy highlights points on the scatter­plot, and vice-versa.
● Target–disease associations that have high Importance, and high Target Novelty, are
promising areas for following up.
look for consistent documentation for the 70 documented protein–protein inter­actions mentioned in section “Behavioral Data”. The Protein–Protein Interactions component (Figure 8.8) presents a link to “Explore Interacting Targets,” which takes the user to a list of 71 targets, including the 70 interacting targets plus ATP1B4 (the target of interest).
Filter Value Enrichment
As introduced in Section 8.2.3.1, the lters in the lter panel (Figure 8.2 panel E) will tell us how many of the targets in the list have each lter value. Since ATP1B4 did not have any direct disease associations, we used the Associated Disease lter to discover, which diseases may be associated with the set of interacting targets.
Figure 8.12 top-right shows the counts of targets in this list that are associated with each disease. At rst glance, it might seem signicant that 35 of the 71 (49%) targets in the list are associated with ovarian cancer, but knowing that ovarian cancer and many other cancers are associated with changes in the function and expression of many targets should lead us to be skeptical. In fact, in the unltered target list (Figure 8.12 top-left), 8589 of the 20,412 (42%) targets in TCRD have a documented association with ovarian cancer, meaning that even a randomly chosen
Figure 8.11 GWAS traits.
https://t.me/medicina_free
8.3 Use Cases 249
● A pageable table of GWAS traits for a given target. The listing is sorted based on the
Evidence Score, as calculated by the TIGA algorithm described in Section 8.3.1.4.
● The lower panel shows a plot of the Mean Rank Score vs. Beta Count. The most reliable
traits would be found in the upper right of this plot.
list of proteins would include a large number of targets associated with ovarian cancer.
Pharos allows users to calculate enrichment scores using Fisher’s Exact Test [31] (Section 8.2.4.5), to determine the signicance of nding a given count of lter val­ues in a ltered list. In our example, Fisher’s Exact Test calculates a p-value of 0.13, which can be interpreted as the probability of nding 35 (or more) targets associated with ovarian cancer in a randomly selected list of 71 targets. The implication is that nding so many targets associated with ovarian cancer in this list is not very surpris­ing. Calculating the enrichment scoresfor the entirelist (Figure 8.12 bottom) reveals, which diseases are signicantly overrepresented in the ltered list. The top values in this sorted list would be more promising for further investigation, especially since three out of the top four overrepresented diseases are forms of heart failure.
A similar example is given on Pharos in the “Filter Value Enrichment” tutorial. In that example, enrichment scores are calculated for the list of interacting targets for a very well-understood target, the D(2) dopamine receptor (DRD2). In that example, the enrichment scores reveal many neurological and substance-related diseases that are commonly known to be associated with DRD2. That proof of
250 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
Figure 8.12 Filter value enrichment.
● Top left: In a target list, the filters show the counts of targets that have each annotation.
In this example, the counts of targets associated with each disease are shown for the complete unfiltered target list. Some diseases are associated with a very large number of targets.
● Top right: In a filtered list, these value counts are often affected by this bias in the data
toward heavily documented filter values. In this example, this list is filtered to targets that interact with ATP1B4. Many of the top diseases in this filtered list are the same diseases as top diseases in the unfiltered list. The “calculate enrichment” button here will calculate the probability that the measured counts are explainable by random chance.
● Bottom: Results of an enrichment calculation for targets that interact with ATP1B4. Chronic
heart failure is associated with 11% of the targets in the list (Count: 8 targets, Observed Frequency: 0.11). The Expected Frequency, based on the full dataset, is 0.23%. Based on Fisher’s Exact Test, the probability of finding that many targets by random chance is 4e-12. p-adjust is calculated based on an adjustment for performing a large number of tests, controlling the False Discovery Rate to 0.05.
8.3 Use Cases 251
https://t.me/medicina_free
concept helps validate the method and lends credence to the idea of discovering relevant annotations for less well-studied targets.
8.3.1.3 Downloading Data
Easy data extraction is a key requirement for many researchers. Researchers study­ing ATP1B4 may build a download query and select the “Associated Disease” elds to download all the associated disease data for targets in the list. Users wanting to compile a list of approved drugs or active ligands for targets in the list would select the “Drugs and Ligands” elds. Heat maps and the Sequence Alignment component also contain shortcuts for initiating a data download with the appropriate elds.
8.3.1.4 Variations on this Use Case
Other Ways to Expand the Dataset
The path chosen in this use case was to expand the search for documentation from one dark target to a list of related targets by compiling a list of interacting targets via the “Explore Interacting Targets” button in the Protein–Protein Interactions compo­nent of the target details page.
Another way to expand the dataset would be to compile a list of targets that share one or more of the annotations that exist on this sparsely documented target. Beyond simply clicking what annotations are there (i.e. PANTHER Class for cation trans­porter) to generate a list of targets, users can also compile lists of targets based on having any overlap in documentation through the “Find Similar Targets” buttons (Section 8.2.4.3) that are available on a number of components including GO Terms and Pathways.
Alternatively, the Protein Sequence and Structure component allows users to easily initiate a search for targets with a homologous protein sequence (Section 8.2.4.2). A Sequence search can also be initiated from a target list page via a button high­lighted in Figure 8.2 panel J. Figure 8.13 shows the resulting Sequence Alignment component on the resulting List Analysis page that illustrates, which regions of the sequence are matching in each entry in the target list.
It should also be noted that these expansions can be compounded, for example by nding a list of targets with a homologous sequence, and then ltering that list by PANTHER Class or GO term, thereby nding targets with multiple points of similarity.
Other Filters to use for Enrichment Calculation
In addition to looking for Associated Diseases that are enriched in the target list, other interesting lters to try might be IDG Family, PANTHER Class, DTO Class, Pathways, GO Functions, or GO Processes. There are no restrictions on which l­ters are available to calculate enrichment, beyond the requirement that the lter is categorical (i.e. not numerical).
8.3.2 Characterizing a Novel Chemical Compound
This use case is for a researcher who would like to explore the potential eects of a novel chemical compound. For the sake of the example, we will use a “novel”
252 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
Figure 8.13 Sequence search.
● Sequence Search Results in the List Analysis tab of the resulting target list. The top plot is
a density plot of aligned residues for all the aligning sequences. Below is a representation of each matching target’s region of alignment, color-coded by the percent identity of the alignment.
compound represented by the following SMILES: CC1CC(O)(CCN1CCCC(=O)C1= CC=C(F)C=C1)C1=CC=C(Cl)C=C1. Note that this is a slightly modied version of haloperidol that is just for the purposes of this example, to help illustrate the con­cepts. Our researcher has two main questions: What targets might be aected by this compound? What eects might it have on the human body?
8.3.2.1 Finding Predicted Targets
Since this is a novel chemical compound, there would be no primary documenta­tion to review, as was the case with the dark target in Section 8.3.1. The researcher would begin on the ligand list page, and follow the link for “Structure Search” or navigate to https://pharos.nih.gov/structure. Starting on the Structure Search page (Figure 8.14), the user can paste in the SMILES to attempt to resolve the structure to a known compound and render the compound in the MarvinJS widget.
Following the search link in the “Find Predicted Targets” card will fetch the list of targets from NCATS Predictor [30] as described in Section 8.2.4.4. The ensuing target list (Figure 8.15) is a compilation of proteins predicted as targets of the query structure, and the targets that have a known activity to the compound in the database if a matching compound was found. Since our example is a novel compound, only predicted data are available. Predicted data are highlighted in the Target Card, and in the lter panel with a shaded background.
Details about the prediction are shown in the Table View of the list pages, which includes the predicted activity, the activity for the nearest compound from the model training set, and a measure of the Applicability Domain. Prediction applicability is quantied by the Tanimoto similarity between the query structure and the nearest
8.3 Use Cases 253
https://t.me/medicina_free
Figure 8.14 Searching by chemical structure.
● Pharos’ Structure Search page allows users to initiate a search based on any chemical
structure. Many types of inputs are used to resolve into a SMILES string to l oad into the Marvin JS Sketcher. The Sketcher can then be used to edit the query structure, upload a file, or draw a compound from scratch.
● Options are available to find similar ligands to the query structure or find targets predicted
to interact with the query structure.
compound from the training set. Users can lter the list based on the numeric lter for prediction applicability, or based on the numeric lter for the predicted activity.
For our example compound, NCATS Predictor found 25 targets with predicted activity. There are several target lters to consider when the list consists of targets with activity against a given compound. Calculating lter value enrichment, as described in more detail in section “Filter Value Enrichment”, yields some interest­ing results. Figure 8.16 shows the top ve GO Functions, and the top ve Associated Diseases, for this list of predicted targets. The presence of serotonin-related GO Terms, and depression-related diseases, are what we should expect from this “novel” compound that is so closely related to haloperidol.
254 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
Figure 8.15 Predicted targets.
● Sample results from a target list consisting of targets with predicted activity to the query
structure.
● The Target Prediction Details panel shows the predicted activity the query structure will
have for the target. The nearest activity for this compound from the training set determines the prediction applicability, which is quantified by the Tanimoto coefficient shown over the ≈ between the query structure and the nearest structure from the training set.
● The list can be sorted by a number of different fields. This list has been sorted according
to the potency of the predicted activity to the query structure.
● Experimentally determined activity will also be shown here for targets with a known activ-
ity to the query structure.
8.3.2.2 Analyzing Similar Ligands
Another method of investigating the potential targets or target families for a novel compound is in the analysis of the eects of a set of similar compounds. As in the pre­vious search for predicted targets, a Similarity Search is initiated from the Structure Search page (Figure 8.14 bottom). Section 8.2.4.1 describes how this functionality works.
Pharos will display a list of all compounds from TCRD that are structurally similar to the query structure (Figure 8.17). In the TableView of the ligand list, the 2D struc­tures are presented in order of decreasing similarity for users to browse and choose the best similarity cuto. A Tanimoto coecient of 0.6 is by default used to cast a wide net for matching compounds. It is usually prudent to restrict the similarity fur­ther using the numeric lter for the similarity score. Figure 8.17 shows the resulting ligand list page after searching for similar compounds to our query structure, and subsequently ltering the result list to those ligands with a Structure Similarity (i.e. Tanimoto coecient on molecular ngerprints) greater than 0.8. This results in a list of 66 ligands from TCRD.
8.3 Use Cases 255
https://t.me/medicina_free
Figure 8.16 Enriched filter values in a list of predicted targets.
● Top: The top five GO Functions that are overrepresented in a list of predicted targets.
● Bottom: The top five Associated Diseases that are overrepresented in a list of predicted
targets.
Figure 8.17 Similarity search results.
● Similarity search results for a sample query structure that was the methylated analog of
haloperidol.
● The list shown here is filtered to ligands with a Structure Similarity (i.e. Tanimoto coeffi-
cient on molecular fingerprints) greater than 0.8.