Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5938_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
11 Мб
Скачать
☆
https://t.me/medicina_free
8
https://t.me/medicina_free
Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
Keith J. Kelleher1, Timothy K. Sheils1, Stephen L. Mathias2, Dac-Trung Nguyen1, Vishal Siramshetty1, Ajay Pillai1, Jeremy J. Yang2, Cristian G. Bologa2, Jeremy S. Edwards2, Tudor I. Oprea
1
National Center for Advancing Translational Science, 9800 Medical Center Drive, Rockville, MD 20850, USA
2
University of New Mexico Health Sciences Center, Department of Internal Medicine, Translational
Informatics Division, MSC09-5025, 1 University of New Mexico, Albuquerque, NM 87131, USA
3
Expert Systems Inc., 12760 High Bluff Dr St e 370, San Diego, CA 92130, USA
2,3
, and Ewy Mathé
1
8.1 Introduction
The current focus of translational and biomedical research tends to be dominated by a relatively small number of well-studied proteins. According to one estimate, around 10% of human proteins receive 75% of the focus of research [1]. Another study reported that at least one-third of the human proteome is understudied, based on mined information from PubMed or the granted patent corpus, antibody count, NIH-funded R01 grants, and other criteria [2].
A recent analysis found that much of the reason for this bias can be attributed to the publication history of a target and the availability of chemical probes for tar­gets [3]. As such, there is a need to incentivize more diversity in the range of targets under investigation, such that more novel targets are featured in publications, and to develop molecular probes (e.g. chemicals or antibodies) and genetic constructs for them. This process of illumination can have a ripple eect – as more targets are studied and data are generated, similar understudied targets may receive additional focus. To address this bias and expand our knowledge of understudied proteins, the National Institutes of Health (NIH) initiated the Illuminating the Druggable Genome (IDG) Consortium [4] in 2013 to shed light on these understudied targets, oftentimes referred to as the “dark genome.”
In this chapter, we use the term target to generically refer to both the charac­teristics of the gene and the protein in a drug discovery context: drug targets are macromolecules to which the drug (or its bioactive derivative) binds to exert the intended therapeutic eect [5].
As a primary step in guiding research toward understudied targets, the IDG Con­sortium dened a categorical metric called the TargetDevelopment Level (TDL) [6]. Each target protein is classied into one of four categories, which describe the degree
231
Open Access Databases and Datasets for Drug Discovery, First Edition. Edited by Antoine Daina, Michael Przewosny, and Vincent Zoete. © 2024 WILEY-VCH GmbH. Published 2024 by WILEY-VCH GmbH.
232 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
of progress and available knowledge for these targets in terms of their chemistry, biology, and clinical activity. The four categories and their requirements are:
● Tdark proteins satisfy two or more of the following requirements:
○ A PubMed score [7] less than ve ○ Three or fewer documented GeneRIFs (Reference Into Functions) ○ No more than 50 commercially available antibodies[2]
● Tbio proteins satisfy one of the following requirements:
○ Fail to satisfy two or more of the criteria for Tdark ○ Have a documented Gene Ontology Molecular Function or Biological Process
with an experimental evidence code
● Tchem proteins must have:
○ At least one active small molecule documented by ChEMBL [8] or DrugCentral
[9] that must be one of the following measures: IC50,Ki,Kb,Kd,EC50,AC50, XC50,andKm. Depending on the target type, the activities must also satisfy the following potency cutos:
◾ GPCRs and nuclear hormone receptors: 100nM ◾ Kinases: 30nM ◾ Ion Channels: 10μM ◾ All others: 1μM
● Tclin proteins must have:
○ A documented activity for an approved drug, with a known mechanism of
action (MoA).
Figure 8.1 shows the yearly distribution of the number of targets at each TDL, and how the generation of new knowledge has changed the distribution of the TDLs over time. Specically, the number of Tdark proteins dramatically decreased from 9199 (December 2013) to 5932 (July 2022), indicating an increased awareness of the dark genome and the need to further study associated proteins (see also Figure 8.1).
As part of the IDG program, the Target Central Resource Database (TCRD) and Pharos were created to provide free, public access to information on targets and asso­ciated annotations [6, 10]. TCRD, accessible at http://juniper.health.unm.edu/tcrd/ download/, is an aggregating database of 79 data sources and Pharos (https://pharos .nih.gov/) is the interactive web-based frontend that allows users to browse, search, lter, and analyze the rich information contained within TCRD. Users can log in via pre-established social media credentials, such as Google, Twitter, and GitHub, to save lists of targets, diseases, or ligands that are available upon return. Other than the ability to save custom lists, all functionalities are available without regis­tration. Since its rst publication, TCRD has had 27 new releases, and Pharos has a bimonthly release cycle. TCRD has had approximately 22,000 database downloads as of 31 March 2022. Pharos receives more than 2500 new visitors every month.
TCRD and Pharos are tailored to a broad range of user types, from those that need full access to the data available to others that may only be interested in a specic dataset for a select group of targets. Other types of users may generate tools or other web components that query Pharos’ GraphQL API.
8.2 Methods 233
https://t.me/medicina_free
9199
9214
443
117 1
601
Figure 8.1 Changes in the TDL distribution over time.
● This Sankey plot represents TDL transition for proteins, as assigned in different versions
of TCRD between December 2013 and December 2021.
● The counts of targets in each group at the endpoints are shown on the left and right sides
of the plot.
● The last two-digit number indicates the year. An additional category, “Tvoid,” was intro-
duced for this plot to trace target transition for protein identifiers that have been obsoleted (e.g. pseudogenes) or more recently introduced in UniProt.
5997
118 4 5
1878 216
692
While many primary resources within TCRD already have a web portal where the data can be accessed, there is a clear benet in database aggregators such as Open­Targets
[11], GeneCards [12], or Pharos. These tools display detailed information about a target on a single page, allowing a user to learn about dierent aspects of tar­get biology at once, a role that is also lled by Pharos’ details pages (Section 8.2.3.2). In addition, the visualizations and analysis functionality oered by Pharos’ list pages (Section 8.2.3.1) opens up a range of new possibilities for research questions that cap­italize on the wide breadth of knowledge in the database. The use cases (Section 8.3) illustrate how generating lists in dierent ways, performing calculations, and con­structing visualizations can help address dierent scientic questions, and reveal patterns in the data.
This chapter provides details on the methods (data organization, analysis, and pro­grammatic access), gives an overview of the content and navigation, and exemplies the utility and usage of Pharos through use cases. Our goal is to highlight and demon­strate how Pharos empowers users to gain chemical, biological, and clinical insight on targets, with the ability to analyze and visualize lists of targets, diseases, or lig­ands, and the ability to generate those lists in interesting and relevant ways.
8.2 Methods
8.2.1 Data Organization
TCRD is a relational database compiled from 79 datasets from various knowledge domains. Further details about each resource can be found in previous publications
234 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
[6, 10] and on the Pharos About page (https://pharos.nih.gov/about). Those include eight sources for target–disease associations, three for target–ligand activities, three for protein–protein interactions, ve for pathway annotations, ve for protein expression data, GeneRIFs and publications from NCBI, GO Terms from Gene Ontology, and more.
This section summarizes our approach to aligning information on targets, dis­eases, and ligands across these complementary primary resources. In general, these primary resources use a variety of ontologies describing targets, diseases, ligands, and associations between them, and so require a systematic approach to align data from the dierent sources.
8.2.1.1 Target Alignment
The primary framework for data related to targets comes from the “reviewed” (man­ually curated) human proteins described by UniProt [13]. TCRD maps protein and gene-related data to this dataset given the target identiers used by each dataset, i.e. gene symbols, ENSG IDs, and UniProt IDs, and mapped to the synonyms that UniProt provides for each target.
8.2.1.2 Disease Alignment
Eight data sources [7, 9, 14–19] provide target–disease associations for TCRD and use seven dierent ontologies, including DO, UMLS, MESH, OMIM, Orphanet, AmyCo, and NCBIGene [20–25]. Pharos relies on the MONDO Disease Ontology (DO) [18] to align disparate input source types into a common disease anno­tation. For example, when one source reports an association between APP and “Cerebral amyloid angiopathy, APP-related (OMIM:605714),” and another source reports an association between APP and “APP-related cerebral amyloid angiopathy (DOID:0070028),” Pharos correctly reports those two associations as the same disease.
Overall, the eight data sources provide data for 26,418 unique disease IDs and 17,989 unique disease names among the set of documented target–disease associ­ations. By aligning equivalent terms through the MONDO mappings, Pharos has been able to consolidate data down to 13,704 unique disease entities.
8.2.1.3 Ligand Alignment
Three data sources for activity data on ligands are used, including ChEMBL [8], DrugCentral [9], and Guide to Pharmacology [26]. Each source could represent a given ligand using a dierent SMILES string. To standardize molecules, Pharos uses Layered Chemical Identier (LyChI – https://github.com/ncats/lychi), a hierarchi­cal hash key representation for chemical structures. LyChI is a lexicographically meaningful four-layered hash key that is generated after standardization of the chemical structure, which involves valency check, kekulization, tautomerization, salt/solvent removal, protonation (or deprotonation), followed by perception of mobile charges and stereochemistry. Briey, the hash keys are generated from SMILES, and all activities associated with each ligand across dierent targets are grouped based on that unique hash key.
8.2 Methods 235
https://t.me/medicina_free
8.2.1.4 Data and UI Updates
TCRD and the Pharos web portal are living resources that are constantly being updated. TCRD releases occur approximately twice a year. The list of data sources is constantly changing to meet the research community’s needs and to track new trends and advances in methods and technologies. New data sources are chosen based on the quality and utility of the data source, with the focus on data sources that will help in the project’s main focus, which is to illuminate less well-studied areas of the proteome. Pharos follows a bimonthly release cycle that incorporates new changes from TCRD and adds new functionality and analysis methods.
8.2.2 Programmatic Access and Data Download
Pharos queries TCRD using a Node.js server (https://nodejs.org/) running a GraphQL API (https://graphql.org/). GraphQL is a powerful tool for generating easily extensible queries that are ideal for hierarchical data. For example, top-level information about a ligand can be fetched alongside a list of targets for which the ligand has an activity. The query can easily retrieve further information about targets, such as aected pathways, protein classes, or associated diseases for those targets. External developers can use GraphQL to fetch only data they are interested in, making lighter payloads instead of a traditional REST API where the Pharos developers would dictate what data are returned from a request.
Users can query our GraphQL API directly through the website (https://pharos .nih.gov/api) or programmatically (https://pharos-api.ncats.io/graphql). Pharos’ API page has several Example Queries to help get started querying the API.
Beyond querying the API, users can also download data through the UI, or down­load complete dumps of the SQL database TCRD from http://juniper.health.unm .edu/tcrd/download/.
The Pharos web application constructs GraphQL queries based on user input. The web server component is currently written in Angular 13 (https://angular.io/), using Server Side Rendering and pre-rendering for slow-loading pages. Each details page and list page provides a link to download data for oine analysis. Users can build a download query that includes any related data elds, such as Pathways for a target list or GWAS data for a disease list. The download builder can also help SQL users of TCRD to understand the table structure of the underlying database, as it presents the SQL query that will be used to execute the download query. The downloaded zip le will also contain a metadata le that reviews the version information at the time of the download, a summary of the downloaded elds, and the SQL query used to generate the download. Downloads are limited to 250,000 rows for this functionality. Users, who need more data, should download a full version of TCRD as described above.
8.2.3 UI Organization
This section provides a broad overview of the information layout in Pharos without going into detail about the data and functionality of each component. Many of the
236 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
components and features will be introduced during the course of the use cases that will follow in Section 8.3. Regardless, each component has a help panel that provides those details to users as they need them (Figure 8.2 panel L and Figure 8.3 panel D).
There are two main types of pages that contain the bulk of biological data from TCRD, those being list pages (Section 8.2.3.1), which provide pageable lists of enti­ties for browsing and population analysis, and details pages (Section 8.2.3.2), which present detailed data and source information for a single entity. List pages and details pages are available for all targets, diseases, and ligands in the database.
All Pharos pages contain a top menu with links to the homepage, the main list pages for targets, diseases, and ligands, as well as the About page, FAQ page, and Use Case page (Figure 8.2 panel A). Also from the toolbar, users can search the database, which will be discussed in more detail in Section 8.2.3.3. Rounding out the top menu are buttons to submit feedback to the Pharos development team and to Sign In via a number of social authentication providers (e.g. Google and GitHub). Signing into Pharos is optional, but is useful for users who will be saving their own custom lists (see section 8.3.2 “Uploading a Custom Ligand List”) since they will be able to return to Pharos and access those lists later.
8.2.3.1 List Pages
List pages are browsable listings of targets, diseases, or ligands, that can be ltered by a number of data elds. The example in Figure 8.2 shows a list ltered to include only the 683 targets that have a TDL of Tclin. List pages can be toggled between Table View, where results are shown in a table, and List Analysis View, which oers functionality to show visualizations and perform calculations at the population level. More details on the functionality provided in the List Analysis tab are shown in context in Section 8.3.
The lter panel (Figure 8.2 panel E), shows the counts of entries in the list that have each lter value. Each list page provides the ability to upload a custom list and download data for all entries currently on the list.
The lower right quadrant, labeled Figure 8.2 panel K, is where the entries of the list are shown. Target list pages show each entry in a card view, which highlights some summary data for each target in the list. Disease list pages show a more sim­plied table of data, while ligand list pages show a card view including the chemical structure of each compound. The lists are pageable, and each entry will link to a details page for the entry. A popup information panel will explain the data behind each data point on the cards.
8.2.3.2 Details Pages
As mentioned above, each entry on the list pages will link to a corresponding details page. Figure 8.3 shows an example of a target details page. This example page, as well as all disease and ligand details pages, will have a header that includes the name of the entry, and a link to download data for the given entry. The left panel is a clickable table of contents that highlights, which sections have data and which do not, while the right panel shows the primary data in distinct components. All headers and data labels will show a tooltip on hover that helps the user understand the data and its context. A popup information panel can also be accessed for this information. There
8.2 Methods 237
https://t.me/medicina_free
Figure 8.2 Overview of a list page.
A. To p Menu Bar. All pages contain links to the list pages for targets, diseases, and ligands,
as well as a link to the API page, the About pages, and Tutorials.
B. Search bar. Search the database from the top menu bar of all pages. Suggestions will
appear as users type based on common searches. C. Feedback button. Provide feedback or report bugs to the Pharos team. D. Sign In button. Signing in via social media credentials is available, but is optional. Users
who create their own custom lists will likely want to sign in so that they can access their
lists on subsequent visits. E. Filter Panel. This panel is available on all list pages and displays the counts of entries in
the list that have each filter value. In this example, the Target Development Lev el filter
shows counts of targets at each TDL. This list is currently filtered to targets with a TDL of
Tc l i n .
F. Selected Filters Panel. A complete listing of all filters applied to the list.
G. Filter Description. An expandable section that shows a description of the filter and where
the data came from. H. Table Header. Shows the type of list, and count of entries in the list.
I. View Toggle Button. Toggles the panel to show the Table View (the current view) or List
Analysis View, which displays v isualizations and functionality for analyzing a population
J. Other Buttons. All list pages contain buttons in this section to allow the user to Upload
their own list of targets, diseases, or ligands, as well as to Download data for entries in the
list. Target lists (shown here) will have a button to initiate targets in the database based
on an amino acid sequence. Ligand lists (not pictured) will have a button for initiating a
search for compounds in the database based on a chemical structure. K. Data Table. A pageable listing of data for each entry in the list. Note: Disease lists have a
more simplified Table View of the list, and ligand lists have a card view with the rendered
chemical structure.
L. Info Panel Button. A button that triggers a panel of information describing each piece of
data on the card.
238 8 Pharos and TCRD: Informatics Tools for Illuminating Dark Targets
https://t.me/medicina_free
Figure 8.3 Overview of a details page.
A. Details Page Header. Details pages show the name of the main entity of the details page,
as well as a download button, where users can download any piece of information shown on the page.
B. Table of Contents. Target details pages (shown here) and disease details pages (not pic-
tured) will show this navigable menu for the different components on the details page. Bolded entries are components that have data. Ligand details pages do not have a table of contents due to the small number of components on the pages.
C. Main Panel is the scrollable region that holds the page’s main content.
● The Protein Summary component shows descriptions and aliases. The radar graph (Illu-
mination Graph) from Harmonizome, quantifies the amount of data available for a target on a large number of dimensions.
● Protein Classes as annotated by PANTHER and Drug Target Ontology appear in this
section when they are available.
D. Info Panel Button. A button that triggers a panel of information describing each piece of
data on the component, and some details about the data source. Some components may also display a button to launch a tutorial for understanding data in the component.
are many types of data on the details pages, especially the target details components, and many of them will be explained in more detail in the context of the use cases in Section 8.3.
8.2.3.3 Search
The search bar from either the top menu or the home page (Figure 8.4), will auto-complete partial queries with common searches. Users can navigate directly to matching details pages when the autocompletion match is to that of a name or common symbol for a target, disease, or ligand. When the autocompletion match is to that of a common lter value, such as a GO Term or a pathway, an option will be presented to navigate directly to a listing page ltered to entries that match that lter value. The bottom of Figure 8.4 shows the result of a general text search, which occurs when the user does not select an autocompletion match, or selects
8.2 Methods 239
https://t.me/medicina_free
Figure 8.4 Search functionality.
● Top : Autocomplete suggestions based on the user’s input will take the user directly to a
details page, or a relevant list page. The user input, in this case, is “acetylcholine” and the suggestions include the ligand “acetylcholine” and the target “Acetylcholinesterase.”
● Bottom: By selecting the “Search Pharos…” option, the search will be done for the term in
many locations of the database. Matching targets, diseases, and ligands are profiled, as well as matching text within a variety of filters, such as the PANTHER Class for “acetylcholine receptor.”