Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5422_Библиотеки_им_академика_М_И_Перельмана
.pdf
246 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
are 70 targets with documented physical interaction or text-mined interaction. These
70 targets will be the basis of the target list analysis presented in Section 8.3.1.2.
For the Pathways component, data are aggregated from ve data sources: Reactome [45], KEGG [47], PathwayCommons [48], UniProt [19], and WikiPathways
[49]. The Reactome Pathway tab includes an interactive widget showing a graphical representation of each annotated pathway, which highlights other targets in the
pathways according to their TDL. All pathway annotations can be used as a starting point to generate a target list page that includes all the targets in each pathway.
With pathway data users can generate a target list page of similar targets, a list of
all targets that share any pathway annotation with the chosen target, sorted by the
degree of overlap in their pathway annotations. Details on that calculation are in
Section 8.2.4.3.
Phenotypic Data
Gene Ontology (GO) [50, 51] denes terms for Molecular Functions, Biological
Processes, and Cellular Components that represent the normal function of gene
products. From the Gene Ontology Terms component, Pharos shows each annotation
that has been assigned to the target and can generate ltered lists of targets for any
annotation.
The Disease Associations component (Figure 8.9 – Example GPR68) includes
data from eight sources, including CTD [14], DisGeNET [15], DrugCentral [9],
eRam [16], Expression Atlas [17], JensenLab [7], Monarch [18], and UniProt [19].
Users can pivot from a target details page showing a list of associated diseases to a
disease list page where there are more analysis capabilities and visualizations for
the list of diseases. Our example target has no entries for direct disease association
annotations.
The Disease Novelty component (Figure 8.10 – Example DRD2) highlights data
from Tin-X [52] which discovers target–disease relationships through natural
language processing of PubMed abstracts. Tin-X quanties two metrics representing the Importance, the strength of the association between the target and the
disease, and the Disease Novelty, the relative scarcity of publications about a disease.
Target–disease associations that have high Importance, and high Target Novelty,
are promising areas for following up. Figure 8.10 shows the Tin–X associations
found for DRD2, with highlights on diseases related to substance-related disorders.
Note the tendency for target–disease associations for DRD2 and substance-related
disorders are high-scoring associations in both Importance and Novelty.
The GWAS traits component (Figure 8.11 – Example DRD2) shows traits identi-
ed through genome-wide association studies (GWAS) [53]. GWAS traits are sorted
according to an Evidence Score (Mean Rank Score), which is a ranking of the traits
based on the number and power of the GWAS studies that found the association.
These rankings of GWAS traits are calculated by Target Illumination GWAS Analytics (TIGA) [54] before being ingested into TCRD.
Similar to the Pathways data, a list of Similar Targets can be generated based on
GO Terms, Associated Diseases, or GWAS Traits, as well as many other data points
from this section. For example, a user could generate a target list that shares any
of the associated diseases as the target of interest. Presenting the list in order of the

8.3 Use Cases 247
https://t.me/medicina_free
Figure 8.9 Associated diseases.
● A pageable table of associated diseases for the current target. Entries are expandable to
show details of the association from each data source.
● Users can navigate to the disease details page through the Explore Disease button.
● The Explore Associated Diseases button will transfer the diseases in this table to a disease
list page where further analysis can be done.
● The Find Similar Targets button will construct a target list that shares any associated dis-
eases with the target of interest, and sort them by the degree of overlap in the sets of
diseases, as described in Section 8.2.4.3.
degree of overlap allows them to easily nd other targets that have the most similar
sets of those associated diseases, or whichever data eld the similarity calculation
was based on.
Resources and Publications
Researchers can learn which typical experimental models (17 species supported)
have orthologous versions of the target. Additionally, resources, such as genetic
constructs, cells, mice, chemical tools, or data resources, generated by IDG grant
awardees may be possible to acquire for research. The last section of the target
details pages contains temporal plots of bibliometric scores [7, 55] and patent
counts that inform users about the popularity of research into a particular target.
Text-mined references from JensenLab, and manually curated GeneRIFs (https://
www.ncbi.nlm.nih.gov/gene/about-generif) are available to review.
8.3.1.2 List Analysis
So far in this use case, as a researcher investigating the role of the dark target ATP1B4, we learned that there is some data for the target, but notably no
known associated diseases. Next, we will continue our search for possible disease
associations by expanding the search to related targets. One promising idea is to

248 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
Figure 8.10 Disease novelty.
● This component highlights data from Tin-X [52] which discovers target–disease relation-
ships through natural language processing of PubMed abstracts.
● Left shows a scatterplot of the Importance, the strength of the association between the
target and the disease, versus the Disease Novelty, and the relative scarcity of publications
about a disease.
● Right shows a circle-pack plot of the Importance metric for each disease association.
Selecting different regions within the disease hierarchy highlights points on the scatterplot, and vice-versa.
● Target–disease associations that have high Importance, and high Target Novelty, are
promising areas for following up.
look for consistent documentation for the 70 documented protein–protein interactions mentioned in section “Behavioral Data”. The Protein–Protein Interactions
component (Figure 8.8) presents a link to “Explore Interacting Targets,” which
takes the user to a list of 71 targets, including the 70 interacting targets plus ATP1B4
(the target of interest).
Filter Value Enrichment
As introduced in Section 8.2.3.1, the lters in the lter panel (Figure 8.2 panel E)
will tell us how many of the targets in the list have each lter value. Since ATP1B4
did not have any direct disease associations, we used the Associated Disease lter to
discover, which diseases may be associated with the set of interacting targets.
Figure 8.12 top-right shows the counts of targets in this list that are associated
with each disease. At rst glance, it might seem signicant that 35 of the 71 (49%)
targets in the list are associated with ovarian cancer, but knowing that ovarian
cancer and many other cancers are associated with changes in the function and
expression of many targets should lead us to be skeptical. In fact, in the unltered
target list (Figure 8.12 top-left), 8589 of the 20,412 (42%) targets in TCRD have a
documented association with ovarian cancer, meaning that even a randomly chosen

Figure 8.11 GWAS traits.
https://t.me/medicina_free
8.3 Use Cases 249
● A pageable table of GWAS traits for a given target. The listing is sorted based on the
Evidence Score, as calculated by the TIGA algorithm described in Section 8.3.1.4.
● The lower panel shows a plot of the Mean Rank Score vs. Beta Count. The most reliable
traits would be found in the upper right of this plot.
list of proteins would include a large number of targets associated with ovarian
cancer.
Pharos allows users to calculate enrichment scores using Fisher’s Exact Test [31]
(Section 8.2.4.5), to determine the signicance of nding a given count of lter values in a ltered list. In our example, Fisher’s Exact Test calculates a p-value of 0.13,
which can be interpreted as the probability of nding 35 (or more) targets associated
with ovarian cancer in a randomly selected list of 71 targets. The implication is that
nding so many targets associated with ovarian cancer in this list is not very surprising. Calculating the enrichment scoresfor the entirelist (Figure 8.12 bottom) reveals,
which diseases are signicantly overrepresented in the ltered list. The top values
in this sorted list would be more promising for further investigation, especially since
three out of the top four overrepresented diseases are forms of heart failure.
A similar example is given on Pharos in the “Filter Value Enrichment” tutorial.
In that example, enrichment scores are calculated for the list of interacting targets
for a very well-understood target, the D(2) dopamine receptor (DRD2). In that
example, the enrichment scores reveal many neurological and substance-related
diseases that are commonly known to be associated with DRD2. That proof of

250 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
Figure 8.12 Filter value enrichment.
● Top left: In a target list, the filters show the counts of targets that have each annotation.
In this example, the counts of targets associated with each disease are shown for the
complete unfiltered target list. Some diseases are associated with a very large number of
targets.
● Top right: In a filtered list, these value counts are often affected by this bias in the data
toward heavily documented filter values. In this example, this list is filtered to targets that
interact with ATP1B4. Many of the top diseases in this filtered list are the same diseases
as top diseases in the unfiltered list. The “calculate enrichment” button here will calculate
the probability that the measured counts are explainable by random chance.
● Bottom: Results of an enrichment calculation for targets that interact with ATP1B4. Chronic
heart failure is associated with 11% of the targets in the list (Count: 8 targets, Observed
Frequency: 0.11). The Expected Frequency, based on the full dataset, is 0.23%. Based on
Fisher’s Exact Test, the probability of finding that many targets by random chance is 4e-12.
p-adjust is calculated based on an adjustment for performing a large number of tests,
controlling the False Discovery Rate to 0.05.

8.3 Use Cases 251
https://t.me/medicina_free
concept helps validate the method and lends credence to the idea of discovering
relevant annotations for less well-studied targets.
8.3.1.3 Downloading Data
Easy data extraction is a key requirement for many researchers. Researchers studying ATP1B4 may build a download query and select the “Associated Disease” elds
to download all the associated disease data for targets in the list. Users wanting to
compile a list of approved drugs or active ligands for targets in the list would select
the “Drugs and Ligands” elds. Heat maps and the Sequence Alignment component
also contain shortcuts for initiating a data download with the appropriate elds.
8.3.1.4 Variations on this Use Case
Other Ways to Expand the Dataset
The path chosen in this use case was to expand the search for documentation from
one dark target to a list of related targets by compiling a list of interacting targets via
the “Explore Interacting Targets” button in the Protein–Protein Interactions component of the target details page.
Another way to expand the dataset would be to compile a list of targets that share
one or more of the annotations that exist on this sparsely documented target. Beyond
simply clicking what annotations are there (i.e. PANTHER Class for cation transporter) to generate a list of targets, users can also compile lists of targets based on
having any overlap in documentation through the “Find Similar Targets” buttons
(Section 8.2.4.3) that are available on a number of components including GO Terms
and Pathways.
Alternatively, the Protein Sequence and Structure component allows users to easily
initiate a search for targets with a homologous protein sequence (Section 8.2.4.2).
A Sequence search can also be initiated from a target list page via a button highlighted in Figure 8.2 panel J. Figure 8.13 shows the resulting Sequence Alignment
component on the resulting List Analysis page that illustrates, which regions of the
sequence are matching in each entry in the target list.
It should also be noted that these expansions can be compounded, for example
by nding a list of targets with a homologous sequence, and then ltering that list
by PANTHER Class or GO term, thereby nding targets with multiple points of
similarity.
Other Filters to use for Enrichment Calculation
In addition to looking for Associated Diseases that are enriched in the target list,
other interesting lters to try might be IDG Family, PANTHER Class, DTO Class,
Pathways, GO Functions, or GO Processes. There are no restrictions on which lters are available to calculate enrichment, beyond the requirement that the lter is
categorical (i.e. not numerical).
8.3.2 Characterizing a Novel Chemical Compound
This use case is for a researcher who would like to explore the potential eects of
a novel chemical compound. For the sake of the example, we will use a “novel”

252 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
Figure 8.13 Sequence search.
● Sequence Search Results in the List Analysis tab of the resulting target list. The top plot is
a density plot of aligned residues for all the aligning sequences. Below is a representation
of each matching target’s region of alignment, color-coded by the percent identity of the
alignment.
compound represented by the following SMILES: CC1CC(O)(CCN1CCCC(=O)C1=
CC=C(F)C=C1)C1=CC=C(Cl)C=C1. Note that this is a slightly modied version of
haloperidol that is just for the purposes of this example, to help illustrate the concepts. Our researcher has two main questions: What targets might be aected by this
compound? What eects might it have on the human body?
8.3.2.1 Finding Predicted Targets
Since this is a novel chemical compound, there would be no primary documentation to review, as was the case with the dark target in Section 8.3.1. The researcher
would begin on the ligand list page, and follow the link for “Structure Search” or
navigate to https://pharos.nih.gov/structure. Starting on the Structure Search page
(Figure 8.14), the user can paste in the SMILES to attempt to resolve the structure to
a known compound and render the compound in the MarvinJS widget.
Following the search link in the “Find Predicted Targets” card will fetch the list
of targets from NCATS Predictor [30] as described in Section 8.2.4.4. The ensuing
target list (Figure 8.15) is a compilation of proteins predicted as targets of the query
structure, and the targets that have a known activity to the compound in the database
if a matching compound was found. Since our example is a novel compound, only
predicted data are available. Predicted data are highlighted in the Target Card, and
in the lter panel with a shaded background.
Details about the prediction are shown in the Table View of the list pages, which
includes the predicted activity, the activity for the nearest compound from the model
training set, and a measure of the Applicability Domain. Prediction applicability is
quantied by the Tanimoto similarity between the query structure and the nearest

8.3 Use Cases 253
https://t.me/medicina_free
Figure 8.14 Searching by chemical structure.
● Pharos’ Structure Search page allows users to initiate a search based on any chemical
structure. Many types of inputs are used to resolve into a SMILES string to l oad into the
Marvin JS Sketcher. The Sketcher can then be used to edit the query structure, upload a
file, or draw a compound from scratch.
● Options are available to find similar ligands to the query structure or find targets predicted
to interact with the query structure.
compound from the training set. Users can lter the list based on the numeric lter
for prediction applicability, or based on the numeric lter for the predicted activity.
For our example compound, NCATS Predictor found 25 targets with predicted
activity. There are several target lters to consider when the list consists of targets
with activity against a given compound. Calculating lter value enrichment, as
described in more detail in section “Filter Value Enrichment”, yields some interesting results. Figure 8.16 shows the top ve GO Functions, and the top ve Associated
Diseases, for this list of predicted targets. The presence of serotonin-related GO
Terms, and depression-related diseases, are what we should expect from this
“novel” compound that is so closely related to haloperidol.

254 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
Figure 8.15 Predicted targets.
● Sample results from a target list consisting of targets with predicted activity to the query
structure.
● The Target Prediction Details panel shows the predicted activity the query structure will
have for the target. The nearest activity for this compound from the training set determines
the prediction applicability, which is quantified by the Tanimoto coefficient shown over the
≈ between the query structure and the nearest structure from the training set.
● The list can be sorted by a number of different fields. This list has been sorted according
to the potency of the predicted activity to the query structure.
● Experimentally determined activity will also be shown here for targets with a known activ-
ity to the query structure.
8.3.2.2 Analyzing Similar Ligands
Another method of investigating the potential targets or target families for a novel
compound is in the analysis of the eects of a set of similar compounds. As in the previous search for predicted targets, a Similarity Search is initiated from the Structure
Search page (Figure 8.14 bottom). Section 8.2.4.1 describes how this functionality
works.
Pharos will display a list of all compounds from TCRD that are structurally similar
to the query structure (Figure 8.17). In the TableView of the ligand list, the 2D structures are presented in order of decreasing similarity for users to browse and choose
the best similarity cuto. A Tanimoto coecient of 0.6 is by default used to cast a
wide net for matching compounds. It is usually prudent to restrict the similarity further using the numeric lter for the similarity score. Figure 8.17 shows the resulting
ligand list page after searching for similar compounds to our query structure, and
subsequently ltering the result list to those ligands with a Structure Similarity (i.e.
Tanimoto coecient on molecular ngerprints) greater than 0.8. This results in a
list of 66 ligands from TCRD.

8.3 Use Cases 255
https://t.me/medicina_free
Figure 8.16 Enriched filter values in a list of predicted targets.
● Top: The top five GO Functions that are overrepresented in a list of predicted targets.
● Bottom: The top five Associated Diseases that are overrepresented in a list of predicted
targets.
Figure 8.17 Similarity search results.
● Similarity search results for a sample query structure that was the methylated analog of
haloperidol.
● The list shown here is filtered to ligands with a Structure Similarity (i.e. Tanimoto coeffi-
cient on molecular fingerprints) greater than 0.8.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
