Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5942_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
11 Мб
Скачать
☆
(a)
https://t.me/medicina_free
3.3 Protocols 81
(b)
Figure 3.6 Using DrugBank Online’s advanced search functionality. This figure shows the process of conducting an advanced search. (a) The completed advanced query includes multiple search conditions and display fields. (b) Search results are displayed in list format, with each result containing the fields used in the search conditions and requested in the display fields.
and management of cardiovascular diseases. As per World Health Organization (WHO) guidance around International Nonproprietary Names (INN), inhibitors of ACEare given the sux “-pril” [23]. For the rst search condition, the eld and predicate can remain in their default state (“Name” and “matches”) – this will search through drug names in DrugBank and return any results that match the query entered in the “search drug name” textbox. To complete the rst search con­dition, type *pril into the textbox – because the asterisk can be used as a wildcard matching any number of characters, this search term will nd any drug names that end with the string “pril.”
82 3 DrugBank Online: A How-to Guide
https://t.me/medicina_free
a. The rst dropdown box for a given search condition species the eld in
which the user wishes to search. The advanced search function supports a number of search elds, including drug identiers and chemical properties (e.g. brands/products, CAS number, InChI, chemical formula, and predicted logP) as well as drug type and availability (e.g. small molecule, approved, and withdrawn; see Section 3.2.2.2), among others.
b. The second dropdown box for a given search condition species the predicate,
which simply tells the search how to query the chosen eld for the inputted text. Predicates provide additional search exibility by allowing users to build more complex queries – for example, the predicate in the above exercise may be set to “does not match” to generate a list of drugs that do not contain the string “pril.” Supported predicates are dependent on the selected search eld, and in general include functions like “does not equal,” “starts with,” and “is present,” among others.
i. Manipulating the predicate in the above example allows us to run a similar
search without using the wildcard (*) character. With the search eld set to “Name,” we can set the predicate to “ends with” and type “pril” in the textbox, this anchors the search to the end of the string and will nd any instances in which a drug name ends with the sux “-pril.”
ii. Wildcard searching is supported when the predicate is set to either “match-
es” or “does not match.” Users can input an asterisk (*) to match any num­ber of characters or a question mark (?) to match a single character.
3. After completing the rst search condition, select Add Search Condition to
include another. For this exercise, we want to search only for approved ACE inhibitor drugs, so click the search eld dropdown and select “Approved.” a. When the selected search eld can only evaluate to true or false (i.e. the drug
is either approved or is not), the available predicates will also change to reect this. After selecting “Approved,” leave the predicate set to its default state, “is true.”
4. With our search conditions set, we next need to set display elds. These are addi-
tional elds that will appear alongside our search results, which are not part of the actual query. Click the Add Display Field button to create our rst display eld. a. Display elds can be selected from the same list available for search elds (e.g.
Name, CAS number, and InChI). They do not require a predicate or the input of a query, as they are simply additional pieces of data that we would like to display alongside our search results.
5. We will set three display elds: one each for CAS number, UNII, and average
mass. First, create two more display elds by clicking the Add Display Field button two more times. In the rst display eld, select “CAS Number.” In the second, select “UNII,” and in the third select “Average mass.” a. Similar to creating search conditions, users can add as many display elds as
necessary to achieve the desired results.
6. Once the appropriate search conditions and display elds are set (Figure 3.6a),
clicking the Search button will generate the search results below the search
3.3 Protocols 83
https://t.me/medicina_free
widget. Each result will rst display the data found by the search conditions – in this case, “Name” and whether the drug is “Approved” – followed by the specied display elds. By default, each returned drug will also populate with its approval status (e.g. approved, withdrawn, and investigational) in its top-right corner (Figure 3.6b).
7. Users with a free DrugBank Online account can export the results of an advanced search as a CSV le. a. To create a new DrugBank Online account, click the Sign Up button near the
top of the advanced search page and follow the instructions provided. If you have an existing DrugBank Online account, click Login and enter your user­name and password.
b. Once signed in, users can export their search results in CSV format by clicking
the Export button found at the top-right of the search results.
8. The advanced search function also supports searches of DrugBank’s drug tar­get data. To search through targets instead of drugs, click the Target Advanced Search button in the top-right of the advanced search page. a. The functionality of the target advanced search is essentially identical to that
of the drug search. Users can add one or more search conditions, specifying a search eld and, if necessary, a predicate and text query for each. Display elds can also be added to target searches, and search results can be exported using the method described in Step 7.
b. Rather than returning a list of drugs, target searches return a list of
biomolecules that may interact with drugs (e.g. receptors and enzymes). The available search elds for this dataset are dierent than those available for the advanced drug search, with more focus on target-specic data like UniProt ID and taxonomy.
DrugBank Online’s advanced search functionality serves to illustrate the power and potential of DrugBank data. The exibility aorded by this advanced search, as well as the ability to export its results, means that users can generate highly focused and specic datasets for use in a variety of applications, such as ML (see Section
3.3.3). A less focused exploration of drugs and compounds with similar traits can be achieved by browsing through DrugBank Online’s drug categories as described in the following section.
3.3.1.3 Browsing Drugs Using DrugBank Online’s Drug Categories
As described in Section 3.2.2.4, drugs in DrugBank are assigned categories that serve to group similar drugs together based on shared characteristics. Drugs may be grouped into categories based on mechanistic similarities (e.g. “Proton Pump Inhibitors”), pharmacokinetic properties (e.g. “CYP3A4 Substrates”), structural similarities (e.g. “Catecholamines”), or clinical use (e.g. “Antifungal Agents”). Grouping like drugs together within drug categories can help to elucidate common­alities between member drugs, for example, a common target that might represent an MoA or a common metabolic pathway through which member drugs may be metabolized.
84 3 DrugBank Online: A How-to Guide
https://t.me/medicina_free
1. From the DrugBank Online landing page, navigate to the drug category browser (https://go.drugbank.com/categories) by clicking the Browse tab in the naviga­tion bar at the top of the page and selecting Categories. a. Individual categories can also be accessed directly from a drug card by clicking
the hyperlinked title of the category of interest from the “Categories” section of the drug card (see Section 3.2.2.4).
2. Drug categories are presented as a searchable table that can be ltered by the approval status and/or market availability of the drugs within them. a. Each category in the table contains the name of the category, a truncated
description of the category, the number of drugs within the category, and the total number of targets associated with those drugs.
i. The category table can be additionally ltered via these columns. Users can
search through category names and descriptions by inputting their search term(s) in the text boxes at the top of each column. Inputting a value into the text box at the top of the “# of drugs” or “# of targets” column will lter the table to show only the categories, which contain a number of drugs or targets greater than or equal to the value input at the top of the column.
b. Clicking the hyperlinked category name will direct the user to a category-
specic page with additional information about the selected category.
3. For this exercise, we will navigate to the “ACE inhibitors” category in order to view the same drugs returned in the advanced search query outlined in Section
3.3.1.2. In the textbox at the top of the category column, type “enzyme inhibitors” and click the magnifying glass or hit “Enter” to lter the list down to a handful of categories (Figure 3.7a). Navigate to the category page for “ACE inhibitors” by clicking the hyperlinked category name in the leftmost column.
4. Every category in DrugBank has a number of data elds that can be viewed from the page of that specic category. Information about the category as a whole includes its name, accession number (a 6-digit number prexed with “DBCAT”), a description of the category, and its equivalent ATC classication [12] (Figure 3.7a). a. Some categories in DrugBank are associated with multiple accession numbers,
which will be indicated by additional bracketed accession numbers following the rst. This means that two (or more) categories were, at some point, deemed synonymous and merged together.
b. Within each category, users can browse through its member drugs and their
associated targets, or search through drugs and targets using the search bar in the top right of the respective sections. In the “Drugs” section of the page, each drug contained within the category will populate with its name (hyperlinked to the relevant drug card) and a brief description of the drug. In the “Drug & Drug Targets” section, the targets for each drug will display alongside the name of the drug and the type of relationship between the drug and its target (see Section 3.2.2.6 for more information on types of drug–target interactions). Both the drug name and the target names in this section are hyperlinked to their respective pages on DrugBank Online.
(a)
https://t.me/medicina_free
3.3 Protocols 85
(b)
Figure 3.7 Browsing drug categories using DrugBank Online.Thisfigureshowsan example of DrugBank’s browsing feature for drug categories. (a) The drug category browser with search results narrowed to show only categories with “enzyme inhibitor” in the title. Categories can be broadly filtered by group or market availability (of the drugs within them), or more specifically searched via the text boxes at the top of each column. (b) DrugBank’s drug category page for ACE inhibitors. Note that both the “Drugs” and “Drugs and Drug Targets” lists can be searched using the search box to the upper-right of each list, and can be reordered using the up-down arrow icons at the top of each column.
Organizing drugs into categories allows users to examine groups of similar drugs at a higher level of abstraction. Previously hidden relationships might become apparent when browsing drugs in this way – for example, we may notice a target or enzyme common to several members of a given drug category that can provide clues or additional context in the process of drug discovery or in the evaluation of a newly synthesized molecule.
86 3 DrugBank Online: A How-to Guide
https://t.me/medicina_free
3.3.2 Identifying Chemicals and Relevant Sequences
Text searching, including both basic (Section 3.3.1.1) and advanced (Section 3.3.1.2) searching, was covered in Section 3.3. It is also possible to query DrugBank using richer data structures including chemical structures and nucleic acid or protein sequences.
3.3.2.1 Searching Using Chemical Structure Search
DrugBank Online’s chemical structure search, powered by ChemAxon (https:// chemaxon.com/), allows users to search for drugs based on their similarity to a specied chemical structure. This type of search functionality is particularly useful for chemists who are interested in nding similar molecules to newly synthesized or identied compounds. It is also useful for searching for compounds that have the same parent molecule or belong to the same drug class.
1. From the DrugBank Online landing page, navigate to the chemical structure search (https://go.drugbank.com/structures/search/small_molecule_drugs/ structure) by clicking on the Search tab in the navigation bar and selecting Chemical Structure (Figure 3.8a). a. By default, the search parameters are set to nd drugs based on their “Similar-
ity” with a similarity threshold of 0.7 and will return a maximum of up to 100 results. These parameters can be adjusted as described in the next step.
b. Structures can be manually drawn into the MarvinJS drawing applet using the
provided tools in the drawing box. If a SMILES, InChI, or similar identier is known, it can also be pasted into the canvas. For a complete explanation of the MarvinJS drawing applet, refer to its ocial documentation [24] or click the MarvinJS Tutorials button at the lower-right of the window.
c. To view an example of a pre-drawn structure, click the “Load example” button
located on the lower right side of the window.
2. To modify or rene a structure similarity search, users can edit the search options located to the right of the drawing canvas. a. Users can use radio buttons to specify that a query structure be searched based
on its “Similarity” to other molecules, whether it is a “Substructure” contained within other drugs, or to specify that it must be an “Exact” match to other drugs in DrugBank.
b. Additionally, users can adjust query parameters to specify a similarity thresh-
old, a minimum and maximum molecular weight, the maximum number of displayed results, and the types of drugs returned.
c. The similarity threshold allows users to set a minimum similarity score for the
results of a chemical structure similarity search. A similarity score is a value between 0 and 1 that represents the degree of similarity between the queried structure and each returned structure, with a greater value indicating greater similarity. These scores are generated by rst creating a chemical hashed ngerprint – a bit string encoding structural features – of the structure being queried, which by default is a 1024-bit ngerprint with a maximum pattern length of seven. Using this ngerprint, the Tanimoto similarity
3.3 Protocols 87
https://t.me/medicina_free
(a)
(c) (d)
Figure 3.8 Using DrugBank Online’s chemical structure similarity search.Thisfigure illustrates the process of conducting a chemical structure similarity search. (a) The Marvin JS drawing canvas for drawing and inputting chemical structures to query. All search options shown are in their default state. (b) The chemical structure of testosterone is drawn on the canvas. Note the indexed atoms, which can aid in drawing and communicating more complex structures – atoms indices are not displayed by default but can be turned on in the settings menu indicated by the cogwheel icon at the top of the canvas. (c) Chemical structure similarity search results using testosterone (b) as the queried structure. Results are displayed in descending order of similarity to the queried structure, evident here by the inclusion of testosterone itself as the first result. (d) A screenshot of DrugBank’s drug card for testosterone, with the Similar Structures button below the structure image, highlighted.
(b)
metric between the queried structure and other structures in the database is calculated. A more technical explanation of similarity scores is available via ChemAxon’s documentation [25]. Note that the minimum allowable similarity threshold is 0.3 – attempting to set it any lower will instead run the search using the default value of 0.7.
3. For this exercise, we will assume the role of a researcher interested in developing a novel anabolic steroid. We will draw the structure of testosterone, a simple anabolic steroid, directly in the MarvinJS canvas in order to examine previously synthesized testosterone derivatives and identify potential novel derivatives that have yet to be tested. We will leave the stereochemistry of our molecule unspecied – when the queried structure does not contain stereo information, the search results will include molecules both with and without stereo information.
88 3 DrugBank Online: A How-to Guide
https://t.me/medicina_free
a. To start, we will draw the four-ring steroid nucleus common to all steroid
compounds, which comprises three cyclohexane rings and one cyclopentane ring. Select the cyclohexane ring from the bottom of the canvas and attach two along their vertical axis, with the third attached to the top-right face of the rightmost ring. Select the cyclopentane ring and attach it to the right side of the third cyclohexane ring.
i. At this stage, it is useful to index (i.e. number) the atoms for ease of refer-
ence. Click the “View settings” button at the top of the canvas, represented by a cogwheel icon, check the “Index atoms” checkbox, then hit “Ok.”
b. Next, we need to add some functional groups. Select the bond tool from the
left side of the canvas and add a single bond to carbons 3, 6, 12, and 17 by clicking on each carbon. Note that, by default, the addition of a single bond to an atom will attach a methyl group to the other end of that bond. Carbons 6 and 12 require methyl groups, but carbons 3 and 17 require a ketone and hydroxyl group, respectively.
c. Select the oxygen atom from the right side of the canvas, and click on the
methyl group attached to carbons 3 and 17. This action will substitute the carbon atom at these positions with oxygen and results in a hydroxyl group attached to both carbons 3 and 17.
d. Finally, we will add a double bond between carbons 4 and 5, and to the
hydroxyl group at carbon 3 to create a ketone. Select the bond tool again and click on the existing bond between carbons 4 and 5 to make it into a double bond. Similarly, click on the single bond between carbon 3 and its hydroxyl group to convert it into a double bond and the hydroxyl group into a ketone. i) The complete structure should look identical to the one shown in
Figure 3.8b.
4. Prior to executing the search, click the “Approved” checkbox to limit the results to only compounds, which have been approved for use in humans. The remain­der of the search options can stay in their default setting. After conrming your structure and search options, click the “Search” button to run the search.
5. The list of search results will appear below the MarvinJS structure editor and will be organized in descending order of similarity to the queried structure (Figure 3.8c). Each result will display along with a number of data points, including the DrugBank ID, a similarity score (with higher scores indicating better matches), a vector image of the matched structure, the name and CAS number of the matched drug, its approval status, and its formula and molecular weight. a. Clicking the DrugBank ID will direct you to the DrugBank drug card entry for
that compound.
b. Clicking the vector image of the returned structure will open a new window
with a larger image.
6. The abovementioned protocol has described the steps involved in drawing a struc­ture for which to search in MarvinJS. As mentioned previously, the structure search can also generate structures to query based on certain chemical notation formats like SMILES or InChI. If a compound already has a known SMILES or
3.3 Protocols 89
https://t.me/medicina_free
InChI string, copying and pasting this string directly into the canvas is generally much easier and faster than drawing a structure from scratch.
7. A structure similarity search can also be performed from directly within a DrugBank drug card. In the “Identication” section of the drug card, next to the “Structure” heading, is an image of the drug’s structure (Figure 3.8d). Clicking the button labeled “Similar Structures” directly below will immediately run a structure similarity search using all of the default search options and return a list of similar structures and their similarity scores.
This kind of chemical structure-based searching has a number of potential applica­tionsinregardtodrugdiscovery.Intheexampleabove,oursearchreturnedalistof approved drug molecules with a structure similar to testosterone. One potential next step might be to examine structural dierences in these testosterone derivatives as compared to their relative potencies in order to determine the importance of certain functional groups and their position within the molecule. Even this relatively simple approach can provide the context required to guide further research and narrow the focus of future drug discovery eorts.
When the structure is determined for a newly discovered or synthesized bioactive compound, it can often provide clues as to the compound’s potential actions – in other words, structural similarity to an existing compound might imply a similar mechanism. Taking our admittedly simplied example from above, suppose we were unaware that our starting compound was testosterone. By running a structure sim­ilarity search and looking at the results, we could immediately identify our mystery compound as some type of steroid, and could then make inferences about things like its MoA and pharmacokinetics based on known properties of similar compounds. This search can also be used to identify potential protein targets (viewable by clicking the hyperlinked DrugBank ID), predict side eects, and predict unexpected interac­tions with unintended protein targets.
3.3.2.2 Using Sequence Search to Find Similar Targets
It is possible to search DrugBank for similar protein sequences to a known sequence, including targets, enzymes, carriers, and transporters. This can be useful to under­stand the types of molecules that are known to interact with your sequence (in cases of an exact match) or sequences similar to your sequence. The similarity search is powered by BLAST [26].
As an example, assume you have identied a putative target sequence based on in silico or in vitro means, and wish to understand the chemical nature of drugs that may bind to it. Navigate to the search page (https://go.drugbank.com/ structures/search/bonds/sequence), either by selecting Search -> Target Sequences from the main navigation bar or manually entering the URL; you should see an input form (Figure 3.9a). Next:
1. Enter one or more DNA or protein sequences to search in the main text input
box at the top of the form. These sequences should be in FASTA format (hovering over the small question mark icon at the top right provides a full explanation of this input format).
90 3 DrugBank Online: A How-to Guide
https://t.me/medicina_free
(a)
(b)
Figure 3.9 Using the DrugBank Online sequence search. This figure provides an overview of using the DrugBank Online sequence searching tool. (a) The sequence input form, which accepts one or more FASTA-formatted nucleic acid or protein sequences. Users can adjust a subset of BLAST parameters using the entry fields and radio buttons. The filters allow users to restrict the search to sequences associated with subsets of drugs based on approval status and to specific types of sequences (target, enzyme, carrier, or transporter; see Section
3.2.2.6). (b) The first result displayed after searching using the human C-X-C chemokine receptor type 5. Note the hit metrics in the top right and the BLAST output alignment present below; exact matches are denoted by the one-letter code between sequences, while similar residues are denoted with a plus symbol (“+”). The bottom table lists the drugs with which the identified sequence has known interactions in DrugBank.
2. Adjust the BLAST parameters, if desired. Note that not all parameters available in BLAST, such as the choice of substitution matrix, are available to change. The “Expectation value” controls the cuto for returning hits; increasing this value will result in more hits but many more will be only slightly similar to the target sequence. The default gap opening cost is set at one (as opposed to the normal BLAST default of 11); this may result in hits with more gaps than otherwise