Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5440_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
10.10.2026
Размер:
9 Мб
Скачать
☆
4.2 Structure-Based Drug Discovery Concept 73
sequence length and prediction mode. It returns a unique job ID and URL for monitoring the
progress of submitted jobs and provides a web interface for job submissions. The server can be
utilized for genome-scale structure property prediction and is designed to respond promptly to
users who only require structure property prediction. Additional research could enhance predic-
tion accuracy and decrease execution time during sequence profile construction. Furthermore,
the accuracy and utility of the server in protein analysis could be significantly improved by aug-
menting the quality of training data utilized for secondary structure and disorder prediction [14].
HHpred: A bioinformatics application known as HHpred detects protein homology and predicts
its structure. The tool’s rapidity, user-friendliness, and sensitivity in detecting remote homology
are emphasized. HHpred is a rapid option for structure prediction, as it can search the PDB in
approximately one minute using a 300-amino acid sequence. It facilitates the exploration of an
extensive assortment of structure and protein family databases that are consistently updated.
Such databases comprise the PDB, SCOP, Pfam, SMART, COG, and CDD. HHpred is suited for
predicting domain boundaries because it has been designed to function effectively with query
sequences. HHpred employs profile hidden Markov models (HMMs) to detect protein homology
and compares the query HMM with every HMM in the chosen database. Precalculated secondary
structure information, either predicted by PSIPRED or assigned by DSSP from the 3D structure,
is contained in the database HMMs. The HHsearch software is utilized to conduct the database
search in order to compare HMM–HMM. This software enhances sensitivity through the imple-
mentation of position-specific gap penalties. In addition, HHpred can compute secondary struc-
ture similarity scores, thereby significantly improving the sensitivity toward homologous proteins.
Nevertheless, structurally analogous proteins that are not homologous may yield scores that are
only marginally significant. HHsearch, which utilizes profile HMMs and facilitates HMM–HMM
comparison, enables HHpred to conduct a thorough and sensitive search for homologous pro-
teins, thereby augmenting its capability to identify protein homology. Investigating the homolo-
gous relationship among distinct protein families is facilitated by the useful bioinformatics utility
Hhpred [15]. Robetta: The Robetta protein structure prediction server predicts and analyses struc-
tures using multiple computational approaches. Through RosettaNMR, the Robetta server derives
protein structures from experimental NMR limitations. If available, we generate fragment librar-
ies that match chemical shifts, NOE constraint data, and residual dipolar couplings. After using
these fragment libraries, the RosettaNMR de novo fragment insertion approach generates decoys
using constraints data in its scoring function. In Robetta, computational interface alanine scan-
ning detects energetically relevant side chains in protein–protein complexes to anticipate muta-
tion effects on protein–protein interactions. Lab-developed protein design and protein–protein
docking methods will be added to improve the Robetta server. To expand the Robetta server, the
server will deliver high-quality structural information to enhance research, infer function, and
aid medication creation [16]. AlphaFold: AlphaFold accurately predicts protein structures. It
describes the AlphaFold network’s training, construction, performance, and bioinformatics appli-
cations. AlphaFold improves structure prediction accuracy by training and inferring with raw
MSAs and attention methods. AlphaFold uses a new architecture to include MSAs and paired
features to learn from sequence data evolutionary constraints and patterns in the training process.
Bioinformatics inter-residue contact prediction requires linked mutations and evolutionary
sequence variation, which the system captures. AlphaFold uses the predicted confidence score to
select the best model per target to improve structure prediction and runs the network trunk
numerous times with different random MSA cluster center choices. Structured bioinformatics
and biological applications benefit from the AlphaFold system’s atomic-accuracy protein struc-
ture prediction. This innovation could transform protein structure and function research and
          74
speed up drug discovery and protein engineering. Accurate protein structure predictions can illu-
minate protein function pathways and shed light on molecular biology. Identifying pharmacologi-
cal targets and designing more effective therapies can result. Large-scale protein structural studies
could benefit structural bioinformatics if protein structures can be accurately predicted. This
helps us comprehend biological systems by identifying novel protein folds, characterizing
protein–protein interactions, and studying protein dynamics. Accurate protein structure predic-
tion can also help develop novel proteins with specific functionalities, such as enzymes with
improved catalytic activity or proteins with increased stability, which have biotechnology and
industrial uses. In structural bioinformatics and biological applications, the AlphaFold system’s
atomic-accuracy protein structure predictions could affect drug discovery, protein engineering,
and our understanding of biological systems [17]. SNPWeb: The SNPWEB web server simulates
the effects of a single amino acid residue substitution. It inputs the wild-type protein’s structure
and substitution. The server predicts the mutant’s function and justifies its influence based on
wild-type and mutant structures. Without locating the wild-type structure in PDB or MODBASE,
modeling with MODWEB is attempted. The server computes characteristics for wild-type and
mutant proteins based on their sequences and structures, such as accessible surface area, rigidity,
variations in residue volume, charge, hydrophobicity, evolutionary conservation, and replace-
ment likelihood. In exceptional circumstances, supplementary characteristics, such as the struc-
tural significance and location of substituting established functional sites, may be incorporated.
A decision tree categorizes the mutation as either harmful or neutral. According to the protocol,
a mutation is detrimental in two ways: it may significantly alter functional sites’ composition
when exposed to the solvent or impede the changes in the native fold when submerged in the
core [18]. QUARK: QUARK predicts protein structure using energy concepts and algorithms.
QUARK protein structure prediction uses energy concepts and algorithms in many ways. Energy
terms and conformational change motions are easier to calculate and apply. Predicting structural
properties with neural networks and assembling small fragments with REMC simulations aid the
development of force fields and search engines. A composite knowledge-based force field leads
REMC simulations, improving the force field and search engine with unique energy terms and
Monte Carlo movements. QUARK uses multiple feature predictions, fragment synthesis, REMC
simulation using the semi-reduced protein model, decoy structure clustering, and full-atomic
refinement. QUARK improves ab initio protein structure prediction accuracy and efficiency by
tackling force field design and conformational search with energy concepts and algorithms.
QUARK outperforms Rosetta in protein structure prediction in numerous instances. In a study of
145 test sequences, QUARK models outperformed Rosetta models in RMSD (root mean square
deviation), TM (template modeling) score, and HB score for 96 targets. Small- to medium-sized
globular protein structures are also well-predicted by QUARK. QUARK beat Rosetta in both sets of
proteins, with a higher TM and HB score, indicating better folding of tiny proteins. Due to its more
properly generated potentials for low-resolution simulations, QUARK has been shown to generate
structural models for protein targets more accurately than Rosetta. Force field design and confor-
mational search issues complicate ab initio protein structure prediction. Protein structure predic-
tion requires accurate force fields and efficient conformational space exploration [19]. I-TASSER:
The sequence-to-structure-to-function paradigm used by I-TASSER predicts protein structure and
function. I-TASSER generates 3D atomic models from amino acid sequences utilizing multiple
threading alignments and iterative structure assembly simulations. For forecasts, the I-TASSER
server utilizes the C-score, a confidence score. The C-score depends on threading template align-
ments and structure assembly simulation convergence settings. Based on its substantial C-score
correlation, the first I-TASSER model estimates quality in large-scale benchmark testing. The
4.2 Structure-Based Drug Discovery Concept 75
C-score and quality of lower-ranked models are less associated, making it impossible to determine
their quality. Despite numerous benchmark tests, automated structure and function prediction
quality estimates can be unpredictable and error-prone. Thus, user-collected experimental data is
necessary to validate predictions. The I-TASSER service offers Threading, Ab Initio Modeling,
C-score, Sequence Submission, External Restraints, and result output for protein structure predic-
tion [20]. MODBASE: The relational database MODBASE contains annotated comparative protein
structural models. Applying MODPIPE to all SWISS-PROT protein sequences yielded datasets for
all protein sequences matched to at least one known protein structure. MODBASE models domains
in 415-937 of 733-239 distinct SWISS-PROT protein sequences. MySQL permits direct SQL que-
ries [21]. MOULDER: MODWEB’s optional MOULDER protocol optimizes alignment and implied
model via genetic algorithms. The protocol optimizes model evaluation scores by aligning, realign-
ing, constructing, and assessing models. New alignments, spatial restraint-based comparative mod-
els, and composite criterion assessments are made in iteration. The iterative method blurs
comparative modeling and threading. MOULDER in a Linux cluster takes a day on 100 CPUs to
compute a 150-residue target sequence [21]. MODLOOP: MODLOOP is a web server for precise
protein structure loop modeling using coordinate files and residue positions. Users can optimize
several loops simultaneously, especially coupled loops. Starting with random beginning conforma-
tions and optimizing nonhydrogen atom locations, the service generates 300 loop predictions. The
lowest objective function score determines the final loop prediction. The server submits calcula-
tions to a Linux cluster to reply faster and can only calculate 300 loop conformations. Loops not
longer than 20 residues are allowed [21]. MODWEB: MODWEB, an automatic comparison struc-
ture of proteins for modeling purposes, accepts FASTA sequences and produces models using the
best PDB template structures. This method helps structural genomics examine how a newly estab-
lished structure affects sequence space modeling. MODPIPE, a software pipeline uses template
structures and sequence–structure alignments to construct protein sequence comparison models.
By matching the PSI-BLAST sequence profile to each PDB template sequence and scanning the
target sequence against IMPALA’s template profile database, sequence–structure matches are pro-
duced. We evaluate sequence–structure relationships after building and evaluating the model [21].

4.2.2 Active Binding Site Within the Target

Moreover, based on the final target-generated structure before the docking analysis, it is essential
that the ligand molecule will strongly bind with the target residue there, for the active residue
within the target is most promising for the docking purpose. The binding site of a protein is a
region where a molecule binds to produce the desirable product. Understanding the structure of
ligands with proteins can help in structure-based drug design. A lack of structural information on
binding pockets can lead to possible issues. Binding pockets anticipated using in silico approaches.
While these methods are useful for predicting binding sites, their accuracy is affected by features,
including template similarity and pocket size [6]. There are various tools with different algorithms
and specifications, including CASTp, Prankweb, BSpred, DoGSitesScorer, Consurf, GRaSP,
Metapocket, COACH, Pockdrug, SiteMap, FPocket, PocketDepth, and Caver. The list of available
tools is illustrated in Table 4.2.
4.2.2.1 The Detailed Description of Each Tool
CastP: The computed atlas of surface topography of proteins lets you find, draw, and meas-
ure the geometric and topology characteristics of protein forms online. Using the alpha
shape approach, the server finds physical characteristics, counts area and volume,
          76
determines impose, and retrieves secondary structure data from UniProt and SIFTS. The
method gives a full and accurate numerical summary of the topographic aspects of proteins
and has been demonstrated to be useful for various research [22]. Prankweb: PrankWeb is a
web-based application that helps you use P2Rank, regarded as the most current way to pre-
dict ligand sites’ binding. PrankWeb employs a template-free machine learning approach
that predicts the capacity of near compounds to attach to peptides according to places on
their surface that agents can access. PrankWeb helps viewers estimate and observe which
ligands and proteins will bind and match their predictions to real binding locations and
areas that stay the same [23]. DoGSitesScorer: Identifying probable interaction sites and
subpockets of a particular protein of interest is possible with DoGSitesScorer. After that, it
looks at these pockets’ geometric and physicochemical characteristics while employing a
support vector machine (SVM) to determine how druggable they are [24]. ConSurf: Per the
PDB, ConSurf-DB is a location where the evolutionary conservation investigation into pep-
tides via established structures can be viewed. A grouping of conserved results is currently
calculated based on MSAs taken from the HSSP database. It uses Bayesian analysis in order
to figure out conservation scores, which are then broken down into levels of conservation
using 1–9 codes of color [25]. GRaSP: The GRaSP method is a new, scalable way to predict
the residues that will bind to ligands. It works better than other methods, can be used on a
large scale, and can predict the binding site for a protein complex in 10–20 seconds on aver-
age. It is a way to learn with supervision that shows a residue and its structure neighbors as
a graph, storing them as a feature vector. The atoms and interactions of each residue are
used to determine the relative solvent accessibility, physicochemical traits, and interac-
tions [26]. PockDrug: PockDrug is a strong model for studying druggability in pockets that
can handle unknown pocket boundaries. The PockDrug-Server consistently gives druggabil-
ity results when using various pocket measurement methods. Because it is robust against
pocket boundary and estimation uncertainties, it can be used effectively with apo pockets
that are hard to predict. It used various estimation methods to separate druggable and less
druggable pockets and did better than recent druggability models for apo pockets [27].
Fpocket: This is an open-source package for finding pockets. It uses Voronoi tessellation and
alpha spheres, also available to the public. Fpocket is a solid starting point for quicker, more
reliable, and freely available protein pocket identification, pocket descriptor of the extrac-
tion process, and druggability prediction [28].
Table 4.2 List of associated tools for the active site prediction.
Sl. no. Name URLs References
1 CASTp http://sts.bioe.uic.edu/castp/index.html?1ycs [22]
2 Prankweb https://prankweb.cz/ [23]
3 DoGSitesScorer https://proteins.plus/ [24]
4 ConSurf https://consurfdb.tau.ac.il/ [25]
5 GRaSP https://grasp.ufv.br [26]
6 PockDrug https://pockdrug.rpbs.univ-paris-diderot.fr/cgi-bin/index.
py?page=home
[27]
7 Fpocket https://mobyle2.rpbs.univ-paris-diderot.fr/cgi-bin/portal.
py#forms::fpocket
[28]
4.2 Structure-Based Drug Discovery Concept 77
4.2.2.2 Molecular Docking Analysis
The approach to designing drugs based on shape discusses the docking of molecules. The
molecular docking process finds how the ligand and target molecule will mix. Determining
the preferred position of the low free binding energy shows how well the ligand will bind to the
protein and form a stable complex. Several noncovalent relationships occur in this reaction,
including hydrogen bonds, ionic bonds, hydrophobic interactions, and interactions mediated
by van der Waals. Investigations into molecular docking can be done among proteins, between
proteins and ligands, and between proteins and nucleotides [6]. A ligand is identified with high
binding affinity when the structure and binding site are determined during the drug design
process. The molecular docking process can predict the ligand’s perfect orientation with the
protein’s binding site in this ligand identification method. The docking analysis is anticipated
to be of three types: rigid, flexible, and semiflexible docking. Semiflexible docking is mainly
used for multiple protein structures, and rigid docking can only show the target’s and ligand’s
geometry. Additional flexible docking is used for the refinement [1]. Docking can be manual or
automated, and it is found to fit the ligand best. There are two ways of docking protocol, i.e.,
one ligand with different orientations and the binding affinity score of the ligand. Various tools
and web servers are available for analysis. These tools (mentioned in Table 4.3) are different in
their algorithm, scoring, and strategy.
Table 4.3 List of available tools and databases for the docking analysis.
Sl. no. Name URLs References
1 Cluspro https://cluspro.bu.edu/login.php [29]
2 PatchDock https://bioinfo3d.cs.tau.ac.il/PatchDock/ [30]
3 SymmDock https://bioinfo3d.cs.tau.ac.il/SymmDock/php.php [30]
4 GRAMM-X https://gramm.compbio.ku.edu/request [31]
5 RosettaDock https://rosie.graylab.jhu.edu/ [32]
6 FoXSDock https://modbase.compbio.ucsf.edu/foxsdock/ [33]
7 HDOCK http://hdock.phys.hust.edu.cn/ [34]
8 HADDOCK https://wenmr.science.uu.nl/ [35]
9 ZDOCK https://zdock.umassmed.edu/ [36]
10 FRODOCK https://chaconlab.org/modeling/frodock/frodock-donwload [37]
11 MEGADOCK https://www.bi.cs.titech.ac.jp/megadock/ [38]
12 CombDock http://bioinfo3d.cs.tau.ac.il/CombDock/download/ [39]
13 FibreDock http://bioinfo3d.cs.tau.ac.il/FiberDock/php.php [40]
14 F2Dock https://www.cs.utexas.edu/~bajaj/cvc/software/f2dock.shtml [41]
15 GOLD https://www.ccdc.cam.ac.uk/solutions/software/gold/ [42]
16 pyDock https://life.bsc.es/pid/pydock/ [43]
17 MOE https://www.chemcomp.com/Products.htm [44]
18 AutoDock Vina https://autodock.scripps.edu/ [45]
19 SurflexDock https://www.biopharmics.com/ [46]
20 GEMDOCK http://gemdock.life.nctu.edu.tw/dock/ [47]
          78
4.2.2.3 The Detailed Description of Each Tool
ClusPro: ClusPro is an online docking tool that works autonomously and employs the clustering
technique to find the bound complexes of proteins with the best breakdown and electrostatic free
energy properties. You may modify your search criteria using several advanced methods with
ClusPro. With the help of desolvation and electrostatic energies (calculated using a Coulombic
potential), it quickly filters the output of the Fourier correlation method. This method lets many
structures very close to being native using the filter while eliminating many erroneous results [29].
PatchDock and SymmDock: PatchDock and SymmDock represent two online docking applica-
tions that use shape matching and geometry-based docking algorithms to determine the structure
of amino acid compounds. These two tools are used to figure out the molecular makeup of PPI
complexes and a homo multimer with cyclic symmetry. Both use the Shape Similarity Principles
and the symmetrical limitations method for docking [30]. GRAMM-X: GRAMM-X is a docking
web-based user interface, which employs the Fast Fourier transformation (FFT) method. FFT
incorporates based on knowledge evaluation, smooth potential, and improvement. The global
search FFT stage uses a smooth Lennard-Jones possible on an extremely fine grid. This is followed
by refined optimization in continuous dimensions and rescoring with several possible phrases
based on knowledge [31]. RosettaDock: RosettaDock is an online docking assistance that sorts
bound protein molecules by rigid-body alignment and side chain geometry. It uses a Monte Carlo-
based algorithm for local protein–protein interaction [32]. FoXSDock: FoXSDock is an internet-
based coupling technique employing rigid-body docking and the SAXS (small angle X-ray
scattering) models to generate better complexes of proteins docked with a small amount of vitality.
It has rigid global modeling and flexible docking methodology for rigid body docking to show the
geometry [33]. HDOCK: HDOCK is a web server that uses an alternate docking method to make
good docked complexes. It does both modeling and docking with provided PDB frameworks [34].
HADDOCK: HADDOCK is an online tool using an information-driven coupling method to guess
how protein complexes are assembled. Biomolecular interaction with high inconsistency is what it
does. The flexible docking method is used to show the binding strength between the object being
studied and its receptor [35]. ZDOCK: ZDOCK is an easy-to-use, rigid docking-based protein dock-
ing program that guesses how complexes of proteins and uniform complexes are put together. The
FFT method is used to predict PP complexes and symmetric multimers [36]. FRODOCK:
FRODOCK serves as a docking framework, which assists in the two proteins binding collectively
through its complimentary based on knowledge possibilities. This platform does rapid circular
docking by using added knowledge-based capability [37]. MEGADOCK: MEGADOCK is a speedy
protein–protein docking application built on FFT. Various types of supercomputers are used to
speed up the docking procedure. The system uses various types of supercomputers to accomplish
outstanding efficiency [38]. CombDock: Randomized techniques are employed by CombDock to
do protein–protein docking. The function numerous peptides make in how they communicate
alongside each other is put jointly and anticipated. Sequential development was done by docking
more than one item at a time [39]. FibreDock: FiberDock is the initial web server for dock biomol-
ecules that takes into consideration the adaptability of both sides of chains and the core to enhance
the docked molecules’ strength. In addition to performing rigid-body refining and rescoring, this
application performs adaptable induced-fit backbone improvement [40]. F2Dock: F2Dock is a
quick Fourier transform-based tool for docking biomolecules to one another. It boosts up the pro-
cedure by using numerous threads and Lennard-Jones’ perspective to rescore the docked com-
plexes of proteins depending on their desolvation activity [41]. GOLD: To predict how easily
malleable compounds will bind to protein molecules, GOLD is a docking equipment. Protein–
ligand docking in GOLD is carried out using an algorithm based on genetics technique, and the
corresponding ligand and the protein can be either entirely or partially flexible [42]. pyDock: The
4.2 Structure-Based Drug Discovery Concept 79
pyDockWEB website host forecasts the rigid-body docking of intricate protein structures, employing
an updated version of the pyDock scoring method. The pyDockWEB website serves the pyDock
rigid-body docking and assessment technique while making it easy for scholars to use through a
web-based user interface [43]. MOE: The MOE drug discovery application platform combines visu-
alization, simulation, modeling, and method design. The MOE command, scripting, and applica-
tion programming syntax is SVL. MOE uses SVL, a flexible, high-performance programming
language [44]. AutoDock Vina: AutoDock Vina is a quick and popular open-source docking tech-
nology. The computerized docking apparatus uses a simple scoring process and quick gradient-
optimization conformation exploration. These universal computational docking tools accept
receptor and ligand coordinates data and suggest optimum docked compliance [45]. SurflexDock:
Surflex is a completely automatic and adaptable molecular docking method that uses the
Hammerhead docking system’s scoring mechanism and a surface-based molecular homology
search tool to create molecular segment poses quickly. Surflex docking efficiency was equal to the
best [46]. GEMDOCK: GEMDOCK, a protein–ligand docking software, uses the differential evolu-
tion algorithm to compute elegantly physiologically. A docking program like GEMDOCK uses a
search algorithm and a scoring system to identify target’s binding site. Differential evolutionary
algorithm piecewise potential energy function is used in the GEMDOCK scoring function [47].

4.2.3 Molecular Dynamic Simulations

Molecular dynamics (MD) simulations have been adopted as a substitute for computationally forecast-
ing protein binding pockets and ligands. Protein structural characteristics and stability of protein–ligand
complexes contribute to the drug design. It can facilitate the development of more viable drugs by iden-
tifying additional druggable binding sites and the virtual screening of chemical compounds. Typically,
simulations are performed to validate promising complexes [6]. Various simulation applications are uti-
lized to simulate the dynamics of protein and ligand topologies generated with AMBER CharmGUI
force fields, including GROMOS, GROMACS, YASARA, LAMMPS, GENESIS, NAMD, Gaussian soft-
ware, Discovery Studio, Chimera, TINKER, and OpenMM as shown in Table 4.4 [59]. MD simulations
require a more exact anatomical force field, which may result in a greater computing overhead.
Table 4.4 List of available tools and databases for the MD simulation analysis.
Sl. no. Name URLs References
1 AMBER https://ambermd.org/ [48]
2 CHARMM https://www.charmm.org/ [49]
3 GROMACS https://www.gromacs.org/ [50]
4 YASARA http://www.yasara.org/mdanalysis.htm [51]
5 LAMMPS https://www.lammps.org/#gsc.tab=0 [52]
6 GENESIS https://www.r-ccs.riken.jp/labs/cbrt/
[53]
7 NAMD https://www.ks.uiuc.edu/Research/namd/ [54]
8 Gaussian software N/A [55]
9 Discovery Studio https://www.3ds.com/products/biovia/
discovery-studio/simulations
[56]
10 TINKER https://dasher.wustl.edu/tinker/ [57]
11 OpenMM https://openmm.org/ [58]
          80
4.2.3.1 The Detailed Description of Each Tool
Amber: Amber is a collection of codes that work together for molecular simulations. It consists of
separate programs for system preparation, simulation, and trajectory analysis, allowing for flexibil-
ity, modularity, and compatibility with different coding practices. While this code separation has
advantages, such as easy upgrades and portability, it also has disadvantages, such as a lack of con-
sistent user interface and limited interaction between different program components. The main MD
program in the Amber suite is called Sander. It is a parallel program written in Fortran 90 that uses
the MPI (Message Passing Interface) programming interface for communication among processors.
The program divides force-field tasks among processors and performs MD simulations by comput-
ing potential energy and gradients, communicating force vectors, and updating positions. An opti-
mized version called pmemd has also been developed to improve performance by efficiently sharing
only the necessary coordinate information for energy calculations. The suite also includes the
nmode code for analyzing nonperiodic simulations and computing thermodynamic quantities, pri-
marily simulations [48]. CHARMM: CHARMM-GUI is a web application that helps users assemble
complicated simulation systems without software or modeling expertise. CHARMM input scripts
can be downloaded and reused because it generates and executes them using the CHARMM binary.
CHARMM-GUI uses the model-view-controller (MVC) architecture pattern to divide the program
into three interrelated parts: a model for data and logic, a view for rendering information, and a
controller for user input and model or view control. This modular approach lets new modules be
developed quickly without extensive module knowledge. CHARMM-GUI also uses Python scripts
and third-party applications for difficult instances that CHARMM cannot handle. Users can inspect
their system at each step using JSmol to visualize three-dimensional molecule structures. The view-
ing modes include protein-only complexes containing band illustrations, protein/membrane com-
plexes with lipids and proteins, orientation view of the bilayer hydrophobic core, packing picture
view for lipid-like pseudo atoms, and electrostatic potential view. Users can see and analyze struc-
tures in CHARMM-GUI without installing additional applications, improving workflow [49].
GROMACS: GROMACS is a popular free and open-source chemistry program for biomolecule
dynamical simulations. GROMACS is a popular MD simulation tool that provides spatial and tem-
poral resolution that is impossible in experiments. It runs on supercomputers, embedded systems,
and laptops and is optimized for performance and efficiency. GROMACS is a nearly two-million-
line software project. GROMACS 5 uses C++ and improves code modularity, memory handling,
and error handling. Balanced hardware works well with GROMACS 5; however, unbalanced hard-
ware may slow performance. GROMACS 5 and the Copernicus ensemble framework support
ensemble-level parallelism. GROMACS evaluates short-ranged nonbonded interactions using
domain-level parallel decomposition and a new algorithm. It supports OpenMP-based multithread-
ing and uses SIMD and GPU acceleration [50]. YASARA: YASARA represents atomic visual mode-
ling, and YASARA uses PVL, a revolutionary programming environment that beats conventional
software. PVL uses GPUs’ 1993 molecular simulation application for Windows, Linux, MacOS, and
Android to view even the biggest molecules and do full active real-time calculations using extremely
accurate force fields on regular PCs. VR-enabled chemical modeling program YASARA Model
blends efficiency and enjoyment. It interacts with VR headsets via OpenVR and broadcasts front-
camera footage for viewing a mouse and keyboard. The alpha gradient on the billboard material
combines the image when looking or moving aside [51]. LAMMPS: MD’s ability to model huge
systems over extended timeframes has made it popular. Advanced computer hardware has made
MD simulations faster and more efficient, enabling complex phenomenon analysis. Complex
potentials like many-body and machine learning have also enhanced material property predictions.
A popular open-source MD algorithm, LAMMPS allows parallel computations and includes many
     81
material models and customization possibilities. A parallel MD code for materials and biomolecular
modeling, LAMMPS, was developed in the mid-1990s. It offers 200 pair styles, hundreds of fix and
compute styles, and hybrid models from various models. Model simulation may determine the best
GPU exploit approach. Spatial decomposition can divide the simulation box into subdomains, with
each processor handling a subset of atoms and their interactions [52]. GENESIS: Genesis is a novel
software for general-purpose supercomputers that efficiently simulates huge biomolecular systems
with all-atom MD. GENESIS includes ATDYN and SPDYN MD simulators plus analysis and setup
tools. This version of GENESIS supports CHARMM force-fields and FFTE for 3D FFT (Fast Fourier
Transform). Basic structural parameters and advanced analysis functions like PCA (Principal
Component Analysis) are also available [53]. NAMD: MD simulations on all-atom models have
improved, but sampling uncommon occurrences remains an issue. Charm++ manages NAMD’s
hybrid spatial/force decomposition, which is highly scalable. Charm++ supports many NAMD
instances, and its versatile Tcl interface makes scripting easy. Another set of MPI communicators
allows independent NAMD instances to communicate. Charm++ can be broken into numerous
local subcommunicators [54]. Gaussian software: Gaussian accelerated molecular dynamics
(GaMD) computes biomolecule-free energy and unconstrained increased sampling. It speeds up
and accurately reconstructs biomolecule-free energy landscapes. This enhanced sampling method
uses Gaussian boost potentials to calculate biomolecule-free energy landscapes accurately. GaMD
smooths the potential energy surface using monotonicity, minor potential difference, and threshold
energy range. The dual-boost GaMD accelerates more than the other two simulations. Simulation
settings include harmonic force constants and threshold energy values [55]. Discovery Studio:
Discovery Studio is an easy-to-use, single-graphical interface for drug design and protein modeling
research. Additionally, Discovery Studio Standalone provides a complete molecular modeling envi-
ronment for independent modelers. DS CHARMm can reduce receptor atoms and analyze vast
numbers of ligands in Discovery Studio. It optimizes docked poses and calculates molecular system
entropic energy [56]. TINKER: Tinker offers the latest molecular mechanics and dynamic programs
and tools for modeling molecules and biopolymers that attempt to make Tinker user-friendly with-
out a GUI [57]. OpenMM: OpenMM, a multilayer computer package, mimics molecules on robust
computing systems. It is expandable to accommodate new hardware architectures and add func-
tionality. The molecular simulation code OpenMM solves previous systems’ reusability, extension,
and availability issues. Intended for easy integration into programs, it allows adding new function-
ality without altering the OpenMM library. With its open-source license, OpenMM can be used in
any application. Covalent and noncovalent interactions, implicit solvent models, integrators, ther-
mostats, and barostats are OpenMM’s main features. OpenMM is expandable so that users can add
functionality. OpenMM plugins add new functionalities. Overall, OpenMM’s extensibility simpli-
fies simulation technique prototyping and testing [58].

4.3 Ligand-Based Drug Discovery Concept

When the target’s three-dimensional structure is not known, LBDD is a handy technique during
the drug discovery process. When a target structure is not comparable, the known ligand molecule
is employed to identify the target by leveraging its physiochemical, structural, and pharmacologi-
cal characteristics. The correlation between the target compound’s biological activity and the com-
pounds’ chemical information is established. LBDD approaches predict novel drug compounds
with comparable biological effects using prior knowledge of active drugs’ structural, physical, and
chemical characteristics, as shown in Figure 4.2. The high structural and physicochemical
          82
similarity between chemical compounds signifies a greater biological similarity upon which these
predictions are founded. Several techniques are used in the drug discovery process, one of which
is similarity search. This approach identifies compounds that share a ligand molecule through
their structural and functional similarities. Pharmacophore and quantitative structure–activity
relationship (QSAR) methodologies are implemented after comparable datasets have been identi-
fied. These are the two most prevalent methodologies employed in the LBDD [1, 6, 45]. The tech-
nique known as QSARs is utilized to determine the relationship between a compound’s chemical
structure and its function. In order to determine the biological activity of compounds, the QSAR
model is utilized to optimize them. In contrast, pharmacophore is also a valuable method for
describing the essential characteristics that enable a compound to exert its biological activity. It
enhances knowledge regarding the interactions between ligands and proteins.
By utilizing the structural data of the ligand molecule, it is possible to construct it. However,
within AI and machine learning, several databases and tools (DrugBank, Chembl, Zinc,
ChemBank, and TCM) are accessible to facilitate LBDD-related procedures. Furthermore, because
the LBDD is predominately founded upon a collection of ligand libraries, certain databases are
accessible. These databases encompass a vast quantity of compound data, spewing specific infor-
mation regarding a natural product, FDA-approved drug, phytochemical, small molecular, and
numerous others. The worldwide researcher is utilizing several impactful tools and databases,
which are detailed in Table 4.5.

4.3.1.1 The Detailed Description of Each Tool

LigandScout: LigandScout is a comprehensive, integrated platform that employs 3D chemistry fea-
ture pharmacophore models for exact virtual screening. The program supports streamlined
Ligand-based drug discovery
Pharmacophore modelling
Quantitative structure–
activity relationships
Similarity search
Virtual screening
ADME analysis
MD simulation
Experimental evaluation
Figure 4.2 Illustration of basic concepts that are involved in the LBDD.