Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5418_Библиотеки_им_академика_М_И_Перельмана
.pdf
https://t.me/medicina_free

8
https://t.me/medicina_free
Pharos and TCRD: Informatics Tools for Illuminating Dark
Targets
Keith J. Kelleher1, Timothy K. Sheils1, Stephen L. Mathias2, Dac-Trung
Nguyen1, Vishal Siramshetty1, Ajay Pillai1, Jeremy J. Yang2, Cristian G.
Bologa2, Jeremy S. Edwards2, Tudor I. Oprea
1
National Center for Advancing Translational Science, 9800 Medical Center Drive, Rockville, MD 20850, USA
2
University of New Mexico Health Sciences Center, Department of Internal Medicine, Translational
Informatics Division, MSC09-5025, 1 University of New Mexico, Albuquerque, NM 87131, USA
3
Expert Systems Inc., 12760 High Bluff Dr St e 370, San Diego, CA 92130, USA
2,3
, and Ewy Mathé
1
8.1 Introduction
The current focus of translational and biomedical research tends to be dominated
by a relatively small number of well-studied proteins. According to one estimate,
around 10% of human proteins receive 75% of the focus of research [1]. Another
study reported that at least one-third of the human proteome is understudied, based
on mined information from PubMed or the granted patent corpus, antibody count,
NIH-funded R01 grants, and other criteria [2].
A recent analysis found that much of the reason for this bias can be attributed
to the publication history of a target and the availability of chemical probes for targets [3]. As such, there is a need to incentivize more diversity in the range of targets
under investigation, such that more novel targets are featured in publications, and
to develop molecular probes (e.g. chemicals or antibodies) and genetic constructs
for them. This process of illumination can have a ripple eect – as more targets are
studied and data are generated, similar understudied targets may receive additional
focus. To address this bias and expand our knowledge of understudied proteins,
the National Institutes of Health (NIH) initiated the Illuminating the Druggable
Genome (IDG) Consortium [4] in 2013 to shed light on these understudied targets,
oftentimes referred to as the “dark genome.”
In this chapter, we use the term target to generically refer to both the characteristics of the gene and the protein in a drug discovery context: drug targets are
macromolecules to which the drug (or its bioactive derivative) binds to exert the
intended therapeutic eect [5].
As a primary step in guiding research toward understudied targets, the IDG Consortium dened a categorical metric called the TargetDevelopment Level (TDL) [6].
Each target protein is classied into one of four categories, which describe the degree
231
Open Access Databases and Datasets for Drug Discovery, First Edition.
Edited by Antoine Daina, Michael Przewosny, and Vincent Zoete.
© 2024 WILEY-VCH GmbH. Published 2024 by WILEY-VCH GmbH.

232 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
of progress and available knowledge for these targets in terms of their chemistry,
biology, and clinical activity. The four categories and their requirements are:
● Tdark proteins satisfy two or more of the following requirements:
○ A PubMed score [7] less than ve
○ Three or fewer documented GeneRIFs (Reference Into Functions)
○ No more than 50 commercially available antibodies[2]
● Tbio proteins satisfy one of the following requirements:
○ Fail to satisfy two or more of the criteria for Tdark
○ Have a documented Gene Ontology Molecular Function or Biological Process
with an experimental evidence code
● Tchem proteins must have:
○ At least one active small molecule documented by ChEMBL [8] or DrugCentral
[9] that must be one of the following measures: IC50,Ki,Kb,Kd,EC50,AC50,
XC50,andKm. Depending on the target type, the activities must also satisfy the
following potency cutos:
◾ GPCRs and nuclear hormone receptors: 100nM
◾ Kinases: 30nM
◾ Ion Channels: 10μM
◾ All others: 1μM
● Tclin proteins must have:
○ A documented activity for an approved drug, with a known mechanism of
action (MoA).
Figure 8.1 shows the yearly distribution of the number of targets at each TDL, and
how the generation of new knowledge has changed the distribution of the TDLs over
time. Specically, the number of Tdark proteins dramatically decreased from 9199
(December 2013) to 5932 (July 2022), indicating an increased awareness of the dark
genome and the need to further study associated proteins (see also Figure 8.1).
As part of the IDG program, the Target Central Resource Database (TCRD) and
Pharos were created to provide free, public access to information on targets and associated annotations [6, 10]. TCRD, accessible at http://juniper.health.unm.edu/tcrd/
download/, is an aggregating database of 79 data sources and Pharos (https://pharos
.nih.gov/) is the interactive web-based frontend that allows users to browse, search,
lter, and analyze the rich information contained within TCRD. Users can log in
via pre-established social media credentials, such as Google, Twitter, and GitHub,
to save lists of targets, diseases, or ligands that are available upon return. Other
than the ability to save custom lists, all functionalities are available without registration. Since its rst publication, TCRD has had 27 new releases, and Pharos has a
bimonthly release cycle. TCRD has had approximately 22,000 database downloads
as of 31 March 2022. Pharos receives more than 2500 new visitors every month.
TCRD and Pharos are tailored to a broad range of user types, from those that need
full access to the data available to others that may only be interested in a specic
dataset for a select group of targets. Other types of users may generate tools or other
web components that query Pharos’ GraphQL API.

8.2 Methods 233
https://t.me/medicina_free
9199
9214
443
117 1
601
Figure 8.1 Changes in the TDL distribution over time.
● This Sankey plot represents TDL transition for proteins, as assigned in different versions
of TCRD between December 2013 and December 2021.
● The counts of targets in each group at the endpoints are shown on the left and right sides
of the plot.
● The last two-digit number indicates the year. An additional category, “Tvoid,” was intro-
duced for this plot to trace target transition for protein identifiers that have been obsoleted
(e.g. pseudogenes) or more recently introduced in UniProt.
5997
118 4 5
1878
216
692
While many primary resources within TCRD already have a web portal where the
data can be accessed, there is a clear benet in database aggregators such as OpenTargets
[11], GeneCards [12], or Pharos. These tools display detailed information
about a target on a single page, allowing a user to learn about dierent aspects of target biology at once, a role that is also lled by Pharos’ details pages (Section 8.2.3.2).
In addition, the visualizations and analysis functionality oered by Pharos’ list pages
(Section 8.2.3.1) opens up a range of new possibilities for research questions that capitalize on the wide breadth of knowledge in the database. The use cases (Section 8.3)
illustrate how generating lists in dierent ways, performing calculations, and constructing visualizations can help address dierent scientic questions, and reveal
patterns in the data.
This chapter provides details on the methods (data organization, analysis, and programmatic access), gives an overview of the content and navigation, and exemplies
the utility and usage of Pharos through use cases. Our goal is to highlight and demonstrate how Pharos empowers users to gain chemical, biological, and clinical insight
on targets, with the ability to analyze and visualize lists of targets, diseases, or ligands, and the ability to generate those lists in interesting and relevant ways.
8.2 Methods
8.2.1 Data Organization
TCRD is a relational database compiled from 79 datasets from various knowledge
domains. Further details about each resource can be found in previous publications

234 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
[6, 10] and on the Pharos About page (https://pharos.nih.gov/about). Those include
eight sources for target–disease associations, three for target–ligand activities,
three for protein–protein interactions, ve for pathway annotations, ve for protein
expression data, GeneRIFs and publications from NCBI, GO Terms from Gene
Ontology, and more.
This section summarizes our approach to aligning information on targets, diseases, and ligands across these complementary primary resources. In general, these
primary resources use a variety of ontologies describing targets, diseases, ligands,
and associations between them, and so require a systematic approach to align data
from the dierent sources.
8.2.1.1 Target Alignment
The primary framework for data related to targets comes from the “reviewed” (manually curated) human proteins described by UniProt [13]. TCRD maps protein and
gene-related data to this dataset given the target identiers used by each dataset,
i.e. gene symbols, ENSG IDs, and UniProt IDs, and mapped to the synonyms that
UniProt provides for each target.
8.2.1.2 Disease Alignment
Eight data sources [7, 9, 14–19] provide target–disease associations for TCRD and
use seven dierent ontologies, including DO, UMLS, MESH, OMIM, Orphanet,
AmyCo, and NCBIGene [20–25]. Pharos relies on the MONDO Disease Ontology
(DO) [18] to align disparate input source types into a common disease annotation. For example, when one source reports an association between APP and
“Cerebral amyloid angiopathy, APP-related (OMIM:605714),” and another source
reports an association between APP and “APP-related cerebral amyloid angiopathy
(DOID:0070028),” Pharos correctly reports those two associations as the same
disease.
Overall, the eight data sources provide data for 26,418 unique disease IDs and
17,989 unique disease names among the set of documented target–disease associations. By aligning equivalent terms through the MONDO mappings, Pharos has
been able to consolidate data down to 13,704 unique disease entities.
8.2.1.3 Ligand Alignment
Three data sources for activity data on ligands are used, including ChEMBL [8],
DrugCentral [9], and Guide to Pharmacology [26]. Each source could represent a
given ligand using a dierent SMILES string. To standardize molecules, Pharos uses
Layered Chemical Identier (LyChI – https://github.com/ncats/lychi), a hierarchical hash key representation for chemical structures. LyChI is a lexicographically
meaningful four-layered hash key that is generated after standardization of the
chemical structure, which involves valency check, kekulization, tautomerization,
salt/solvent removal, protonation (or deprotonation), followed by perception of
mobile charges and stereochemistry. Briey, the hash keys are generated from
SMILES, and all activities associated with each ligand across dierent targets are
grouped based on that unique hash key.

8.2 Methods 235
https://t.me/medicina_free
8.2.1.4 Data and UI Updates
TCRD and the Pharos web portal are living resources that are constantly being
updated. TCRD releases occur approximately twice a year. The list of data sources
is constantly changing to meet the research community’s needs and to track new
trends and advances in methods and technologies. New data sources are chosen
based on the quality and utility of the data source, with the focus on data sources
that will help in the project’s main focus, which is to illuminate less well-studied
areas of the proteome. Pharos follows a bimonthly release cycle that incorporates
new changes from TCRD and adds new functionality and analysis methods.
8.2.2 Programmatic Access and Data Download
Pharos queries TCRD using a Node.js server (https://nodejs.org/) running a
GraphQL API (https://graphql.org/). GraphQL is a powerful tool for generating
easily extensible queries that are ideal for hierarchical data. For example, top-level
information about a ligand can be fetched alongside a list of targets for which
the ligand has an activity. The query can easily retrieve further information about
targets, such as aected pathways, protein classes, or associated diseases for those
targets. External developers can use GraphQL to fetch only data they are interested
in, making lighter payloads instead of a traditional REST API where the Pharos
developers would dictate what data are returned from a request.
Users can query our GraphQL API directly through the website (https://pharos
.nih.gov/api) or programmatically (https://pharos-api.ncats.io/graphql). Pharos’
API page has several Example Queries to help get started querying the API.
Beyond querying the API, users can also download data through the UI, or download complete dumps of the SQL database TCRD from http://juniper.health.unm
.edu/tcrd/download/.
The Pharos web application constructs GraphQL queries based on user input. The
web server component is currently written in Angular 13 (https://angular.io/), using
Server Side Rendering and pre-rendering for slow-loading pages. Each details page
and list page provides a link to download data for oine analysis. Users can build a
download query that includes any related data elds, such as Pathways for a target
list or GWAS data for a disease list. The download builder can also help SQL users
of TCRD to understand the table structure of the underlying database, as it presents
the SQL query that will be used to execute the download query. The downloaded zip
le will also contain a metadata le that reviews the version information at the time
of the download, a summary of the downloaded elds, and the SQL query used to
generate the download. Downloads are limited to 250,000 rows for this functionality.
Users, who need more data, should download a full version of TCRD as described
above.
8.2.3 UI Organization
This section provides a broad overview of the information layout in Pharos without
going into detail about the data and functionality of each component. Many of the

236 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
components and features will be introduced during the course of the use cases that
will follow in Section 8.3. Regardless, each component has a help panel that provides
those details to users as they need them (Figure 8.2 panel L and Figure 8.3 panel D).
There are two main types of pages that contain the bulk of biological data from
TCRD, those being list pages (Section 8.2.3.1), which provide pageable lists of entities for browsing and population analysis, and details pages (Section 8.2.3.2), which
present detailed data and source information for a single entity. List pages and details
pages are available for all targets, diseases, and ligands in the database.
All Pharos pages contain a top menu with links to the homepage, the main list
pages for targets, diseases, and ligands, as well as the About page, FAQ page, and Use
Case page (Figure 8.2 panel A). Also from the toolbar, users can search the database,
which will be discussed in more detail in Section 8.2.3.3. Rounding out the top menu
are buttons to submit feedback to the Pharos development team and to Sign In via
a number of social authentication providers (e.g. Google and GitHub). Signing into
Pharos is optional, but is useful for users who will be saving their own custom lists
(see section 8.3.2 “Uploading a Custom Ligand List”) since they will be able to return
to Pharos and access those lists later.
8.2.3.1 List Pages
List pages are browsable listings of targets, diseases, or ligands, that can be ltered
by a number of data elds. The example in Figure 8.2 shows a list ltered to include
only the 683 targets that have a TDL of Tclin. List pages can be toggled between
Table View, where results are shown in a table, and List Analysis View, which oers
functionality to show visualizations and perform calculations at the population
level. More details on the functionality provided in the List Analysis tab are shown
in context in Section 8.3.
The lter panel (Figure 8.2 panel E), shows the counts of entries in the list that
have each lter value. Each list page provides the ability to upload a custom list and
download data for all entries currently on the list.
The lower right quadrant, labeled Figure 8.2 panel K, is where the entries of the
list are shown. Target list pages show each entry in a card view, which highlights
some summary data for each target in the list. Disease list pages show a more simplied table of data, while ligand list pages show a card view including the chemical
structure of each compound. The lists are pageable, and each entry will link to a
details page for the entry. A popup information panel will explain the data behind
each data point on the cards.
8.2.3.2 Details Pages
As mentioned above, each entry on the list pages will link to a corresponding details
page. Figure 8.3 shows an example of a target details page. This example page, as well
as all disease and ligand details pages, will have a header that includes the name of
the entry, and a link to download data for the given entry. The left panel is a clickable
table of contents that highlights, which sections have data and which do not, while
the right panel shows the primary data in distinct components. All headers and data
labels will show a tooltip on hover that helps the user understand the data and its
context. A popup information panel can also be accessed for this information. There

8.2 Methods 237
https://t.me/medicina_free
Figure 8.2 Overview of a list page.
A. To p Menu Bar. All pages contain links to the list pages for targets, diseases, and ligands,
as well as a link to the API page, the About pages, and Tutorials.
B. Search bar. Search the database from the top menu bar of all pages. Suggestions will
appear as users type based on common searches.
C. Feedback button. Provide feedback or report bugs to the Pharos team.
D. Sign In button. Signing in via social media credentials is available, but is optional. Users
who create their own custom lists will likely want to sign in so that they can access their
lists on subsequent visits.
E. Filter Panel. This panel is available on all list pages and displays the counts of entries in
the list that have each filter value. In this example, the Target Development Lev el filter
shows counts of targets at each TDL. This list is currently filtered to targets with a TDL of
Tc l i n .
F. Selected Filters Panel. A complete listing of all filters applied to the list.
G. Filter Description. An expandable section that shows a description of the filter and where
the data came from.
H. Table Header. Shows the type of list, and count of entries in the list.
I. View Toggle Button. Toggles the panel to show the Table View (the current view) or List
Analysis View, which displays v isualizations and functionality for analyzing a population
J. Other Buttons. All list pages contain buttons in this section to allow the user to Upload
their own list of targets, diseases, or ligands, as well as to Download data for entries in the
list. Target lists (shown here) will have a button to initiate targets in the database based
on an amino acid sequence. Ligand lists (not pictured) will have a button for initiating a
search for compounds in the database based on a chemical structure.
K. Data Table. A pageable listing of data for each entry in the list. Note: Disease lists have a
more simplified Table View of the list, and ligand lists have a card view with the rendered
chemical structure.
L. Info Panel Button. A button that triggers a panel of information describing each piece of
data on the card.

238 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
Figure 8.3 Overview of a details page.
A. Details Page Header. Details pages show the name of the main entity of the details page,
as well as a download button, where users can download any piece of information shown
on the page.
B. Table of Contents. Target details pages (shown here) and disease details pages (not pic-
tured) will show this navigable menu for the different components on the details page.
Bolded entries are components that have data. Ligand details pages do not have a table
of contents due to the small number of components on the pages.
C. Main Panel is the scrollable region that holds the page’s main content.
● The Protein Summary component shows descriptions and aliases. The radar graph (Illu-
mination Graph) from Harmonizome, quantifies the amount of data available for a target
on a large number of dimensions.
● Protein Classes as annotated by PANTHER and Drug Target Ontology appear in this
section when they are available.
D. Info Panel Button. A button that triggers a panel of information describing each piece of
data on the component, and some details about the data source. Some components may
also display a button to launch a tutorial for understanding data in the component.
are many types of data on the details pages, especially the target details components,
and many of them will be explained in more detail in the context of the use cases in
Section 8.3.
8.2.3.3 Search
The search bar from either the top menu or the home page (Figure 8.4), will
auto-complete partial queries with common searches. Users can navigate directly
to matching details pages when the autocompletion match is to that of a name or
common symbol for a target, disease, or ligand. When the autocompletion match
is to that of a common lter value, such as a GO Term or a pathway, an option
will be presented to navigate directly to a listing page ltered to entries that match
that lter value. The bottom of Figure 8.4 shows the result of a general text search,
which occurs when the user does not select an autocompletion match, or selects

8.2 Methods 239
https://t.me/medicina_free
Figure 8.4 Search functionality.
● Top : Autocomplete suggestions based on the user’s input will take the user directly to a
details page, or a relevant list page. The user input, in this case, is “acetylcholine” and the
suggestions include the ligand “acetylcholine” and the target “Acetylcholinesterase.”
● Bottom: By selecting the “Search Pharos…” option, the search will be done for the term in
many locations of the database. Matching targets, diseases, and ligands are profiled, as well
as matching text within a variety of filters, such as the PANTHER Class for “acetylcholine
receptor.”
Соседние файлы в папке Библиотека им академика М.И. Перельмана
