Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5908_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Contents
- •Foreword
- •Preface
- •About the Editors
- •Contributors
- •References
- •2.3.4 Barriers to Automation Adoption
- •2.4 Core Ingredients for Successful Digital Transformation
- •2.1 Introduction
- •2.3.1 Operational Challenges
- •2.3.2 Cultural Challenges
- •2.4.2 Cloud Computing
- •2.5 Case Studies of Successful Digital Transformation
- •2.6 Conclusion
- •References
- •3. Computational Protein Design Strategies for Optimization of Antigen Generation to Drive Antibody Discovery
- •3.1 Introduction
- •3.3 Antigen Generation Strategies
- •3.4 Computational Methods
- •3.4.2 Computational Protein Structure Prediction
- •References
- •4. Bioinformatic Analyses of Antibody Repertoires and Their Roles in Modern Antibody Drug Discovery
- •4.1 Introduction
- •4.6 Summary and Future Directions
- •Acknowledgments
- •References
- •5.1 Introduction
- •5.2 Databases
- •5.2.1 Databases in Machine Learning Approaches
- •5.2.2 Database Types
- •5.3 Applications of Machine Learning in Antibody Discovery and Development
- •5.3.1 Structure Prediction with Deep Learning
- •5.3.3 Developability
- •5.4 Antibody Generation and Design by Language Models
- •5.4.1 Antibody Representations
- •5.4.2 Representation Learning
- •5.4.3 Language Models
- •References
- •6.1 Introduction
- •6.2 Antibody Generation through Deep Generative Models
- •6.3.1 Sampling and Scoring
- •6.5 Conclusions and Perspectives
- •Acknowledgments
- •References
- •7.1 Introduction
- •7.2.3 Computational Approaches to Predict Antibody–Antigen Interaction
- •7.3 Conclusion
- •Competing Interests
- •Acknowledgments
- •References
- •8.2 Common Types of Molecular Simulations for Biomolecules
- •8.2.1 Molecular Dynamics (MD) Simulations
- •8.2.2 Monte Carlo (MC) Simulations
- •8.2.3 Challenges of Molecular Simulations
- •8.3.1 Periodic Boundary Conditions
- •8.4 Uses of Molecular Simulation in Antibody Drug Development
- •8.5 Conclusion
- •References
- •9. Considerations of Developability During the Early Stages of Antibody Drug Discovery and Design
- •9.1 Introduction
- •9.2 Historical Perspective
- •9.3 Clinical Antibody Data Set
- •9.5 Control Antibodies
- •9.7 Assessment of Chemical Liabilities
- •9.8 Conclusions and Future Perspectives
- •Acknowledgments
- •References
- •Abbreviations
- •10.1 Introduction
- •10.4.1 Conclusions and Outlook
- •Acknowledgments
- •References
- •11.8 Conclusions and Future Directions
- •References
- •12.1 Introduction to PK/PD and QSP Modeling
- •12.1.1 PK/PD Modeling
- •12.1.2 QSP Modeling
- •12.2.1 Monoclonal Antibodies (mAbs)
- •12.2.3 Cell Therapies
- •12.2.4 Gene Therapies
- •12.2.5 Vaccines
- •12.2.6 mRNA/siRNA/Oligonucleotide Therapeutics
- •12.4 Case Studies
- •12.5 Conclusions and Future Perspectives
- •References
- •13.1 Introduction
- •13.2 AI/ML: A Game Changer for Antibody Design
- •13.3 Multispecific Antibody Design
- •13.4 Adapting AI to the Design of Multispecific Antibodies
- •13.4.1 Structure Prediction and Modeling
- •13.4.2 Developability Prediction and Optimization
- •13.4.4 In Silico Modeling and Simulation
- •13.5 The Future: Beyond Optimization
- •13.5.1 Market Trends and Commercialization
- •13.5.2 Logic Gates, Biosensors, and De Novo Design
- •13.5.3 Challenges and Opportunities
- •13.6 Conclusion
- •Acknowledgments
- •References
- •Index

2 • Digital Transformation 21
tools and machinery become more sophisticated, higher levels of automation are attain‑
able, consequently reducing human involvement. Common tools like pipettes are stan‑
dard, alongside instruments automating single tasks (e.g. thermocyclers, centrifuges,
plate readers). These instruments are integral to developing automated workows, with
the coordination of these diverse instruments and automated modules posing a signi‑
cant challenge.
Beyond unit process automation equipment (level 1 automation), there is a plethora
of tools for automating multi‑step processes and entire experimental workows. Liquid
handling robots and microuidic devices, for instance, have been designed in various
formats to automate critical tasks and multiple steps in lengthy pipelines. Liquid han‑
dling robots enable more exibility and throughput, whereas microuidic devices enable
precision and miniaturisation, allowing researchers to control liquids on a microscopic
scale. Adaptable automation platforms are now being developed, notably by HighRes
Biosolutions, offering highly versatile modular platforms with increased system exibil‑
ity, adaptive scheduling software, and improved protocol design. These technologies and
an overview of their constraints and benets are summarised in Figure2.2. Combining
robust physical automation with computational tools then enables closed‑loop iterative
experimentation, enhancing discovery speed and productivity.
Recent trends emphasise systems approaches that combine hardware and soft‑
ware automation, especially crucial in cell‑based work for generating large, high‑
quality datasets. This is particularly evident in drug discovery’s lead optimisation stage,
where complex assay readouts are essential and processing of data is more complex.
The DBTL framework becomes increasingly automated at level 3 and above. Physical
automation, in tandem with software tools, can then enhance efciency, foster reproduc‑
ibility through standardisation, and minimise experimental error and variability (e.g. by
using data normalisation). These improvements are essential for advancing the DBTL
cycle. To effectively screen multiple designs concurrently and expedite DBTL cycles,
comprehensive automation is indispensable at each stage of the DBTL process.
Design tools in synthetic biology have evolved to include functionalities for DNA
part design and assembly, utilising established methods like Golden Gate assembly,
Gibson assembly, SLIC, and digestion‑ligation methods (Chao etal., 2015; Chao etal.,
2017). Recent years have seen the integration of computational tools and ML, with
innovations like generative design (Ingraham etal., 2023) and protein language models
(Fenoy etal., 2022; Alley etal., 2019; Brandes etal., 2022; Rives etal., 2021; Wu etal.,
2021; Choi, 2022; Dounas etal., 2024; You etal., 2021; Heinzinger etal., 2019; Lu etal.,
2020) augmenting traditional computational biology approaches.
In the build phase, DNA assembly stands out as a particularly time‑intensive and
error‑prone process, making it an ideal candidate for automation (Walsh etal., 2019;
Storch etal., 2020; Linshiz etal., 2016; Kang etal., 2022). Transitioning manual assem‑
bly to high‑throughput automated systems is complex, but using standardised parts and
modular designs can facilitate the large‑scale prototyping of combinatorial DNA librar‑
ies (Casini etal., 2015) as well as faster assembly and screening of combinatorial design
spaces. Methods like Golden Gate assembly, which are executed as one‑pot reactions
with limited steps, such as liquid handling, incubation, and thermocycling, can be auto‑
mated to signicantly cut down on time and manual errors.

FIGURE2.2 Overview of automation technologies and characteristics that need to be considered when incorporating automation. Adapted
from Stephenson etal. (2023).
22 Biopharmaceutical Informatics

2 • Digital Transformation 23
Steps like cell transformation, transfection, or transduction are also challenging to
automate due to specic experimental conditions and high labour intensity, but many
successful implementations exist (Si etal., 2017).
The testing phase presents another automation challenge. Protocols vary widely,
with some steps being more conducive to automation than others. Cell‑based work,
in particular, introduces unique challenges for automation, and current solutions often
struggle to upscale or simply shift bottlenecks along the pipeline (Doulgkeroglou etal.,
2020). Killing assays, for example, using the IncuCyte® by Sartorius, are notoriously
hard to automate. Recent systems like the Tecan Freedom EVO or Tecan Cellerity allow
for better integration of asks (so‑called RoboFlask) into the process and thereby sim‑
plify cell‑based work. Automated analysis involving techniques like chromatography
and mass spectrometry also presents distinct challenges. The complexity of the data and
inefcient resolution renders it difcult for software to accurately make an assessment,
such as correctly distinguishing between peaks and often requires human intervention
(Cappadona etal., 2012). However, the integration of liquid handling technologies with
component instruments that automate critical or labour‑intensive steps, combined with
advanced software, can signicantly alleviate the workload on researchers and enhance
experimental results.
A notable application of this approach is the modular DBTL platform developed
by LabGenius, which focusses on the design, synthesis, and testing of multi‑specic
antibodies using cell‑based assays. LabGenius has developed a exible robotic liquid
handling platform that can perform disease‑relevant, functional, cell‑based assays to
produce machine learning‑grade data at high throughput, minimising variability and
noise and rendering it ideally suited for active learning. Computational models enable
the design and learning steps, while the automated liquid handling system covers the
automated build and test steps. By combining high‑throughput experimentation with
modelling, LabGenius can today evaluate the performance of thousands of antibody
designs in disease‑relevant cell‑based assays in just a few weeks. This is possible
through an active learning‑based approach that generates data in an optimal way to
train ML models (LabGenius, 2023).
While automation has made signicant strides in elds like antibody engineering,
manual intervention is often still necessary for monitoring and connecting different ele‑
ments of the pipeline.
The connectivity among each step in the discovery process and integration of com‑
puting throughout makes it possible to capture a wealth of information: explicit repre‑
sentations of goals, hypotheses, experimental results, and ndings, as well as metadata
that can be critical to tracking why and how experiments were carried out in order to
understand and repeat them or diagnose errors (e.g. environmental conditions, instru‑
ment settings, protocols, runtime logs). For this reason, many systems that are ideally
capable of level 4 automation (e.g. King etal., 2009; Williams etal., 2015; Sparkes etal.,
2010), often function as level 3 systems due to the regular need for human intervention
in hardware and software issues. Nevertheless, they highlight the potential of combined
human‑robot efforts in accelerating research through strategic automation.
It is also important to recognise that fully autonomous systems often require a high
degree of integration, which often coincides with less exibility or modularity of the

24 Biopharmaceutical Informatics
constituent parts. Such systems are not possible in many organisations due to the need
for equipment to work across many different projects with a multitude of workows.
In scientic research, pipetting assistant robots offer a basic level of automation,
emulating manual pipettes with a small footprint and lower cost than more complex plat‑
forms. Their typical applications include reagent dispensing, serial dilutions, and basic
sample preparations. While having lower throughput due to limited channels and deck
size, they alleviate manual pipetting errors and save time. Manufacturers like Integra,
Hudson Robotics, and Gilson offer various models, from dedicated microplate dispens‑
ers to versatile workstations that can integrate with other devices for expanded functions.
Some manufacturers provide pre‑congured workstations for specic research
areas such as NGS, polymerase chain reaction (PCR), and proteomics. These are larger
and costlier, starting from $100 to 200k, but offer high throughput, exibility, and reli‑
ability through advanced hardware and software. Companies like PerkinElmer, Agilent,
and Beckman Coulter offer various models tailored to specic applications.
Contrastingly, microuidic devices, primarily developed in‑house by research
labs, face a commercialisation gap, with notable exceptions in sequencing and analysis
sample preparation. Commercial systems like the Agilent 2100 Bioanalyser and the
Beckman Coulter Echo 650 series have succeeded due to their innovative approaches in
reducing sample consumption, analysis time, and contamination. However, challenges
remain in the commercial viability and reliability of microuidic devices, as seen in the
discontinuation of certain products. Efforts continue in commercialising microuidics,
with emerging low‑cost or open‑source platforms.
In synthetic biology, as automation and digitalisation advance, the need for stan‑
dard data management and communication protocols is crucial for creating an inte‑
grated laboratory environment. While liquid handling robots and automated analytical
units are widespread, they typically operate as standalone units with proprietary stan‑
dards, command sets, device driver software, and data formats. The result of this is poor
connectivity among devices, data silos, and difculty in integrating multiple devices
to form continuous, exible automation pipelines. In automated laboratories, multiple
devices are set up for steps, such as sample preparation and analysis, and connected to
software systems that orchestrate the ow of samples through experimental pipelines.
Piecing together diverse instruments needed for a given workow without standardised
hardware interfaces complicates the connection of equipment to a centralised control
system. Variations in software among devices make integration very challenging, where
even devices with similar functionality have vastly different command sets, behav‑
iours, and capabilities for handling device states and errors. Furthermore, data collected
from various instruments in a workow are typically stored in proprietary formats
that complicate visualisation, analysis, and sharing. Arguably, having a manufacturer‑
independent interface denition that applies to all devices and stores results in a com‑
mon format would greatly simplify the process of setting up exible and modular
laboratory automation systems that can then be sustainably tailored to application and
researcher‑specic needs. There is hence a big need for standardisation of data formats
across analytical devices and interfaces (communication standards)–both are needed
for true end‑to‑end integration. Such standards (e.g. as introduced by the Standardisation
in Laboratory Automation consortium or via the Analytical Information Markup

2 • Digital Transformation 25
Language) would be able to facilitate easier data storage, visualisation, and sharing in
a unied format and can enhance reproducibility, traceability, and remote monitoring
of experiments. They can hence enable the automatic upload to global repositories via
cloud computing, simplifying data management, enabling remote process monitoring,
and making it possible to efciently share experimental results with all stakeholders.
We note that large language models (LLMs) like Generative Pre‑trained
Transformer3 (GPT‑3) and ChatGPT are emerging as promising tools for laboratory
automation. These models, trained on extensive datasets, show adaptability across vari‑
ous applications, notably in intuitive robot control via natural language inputs. However,
it is too early to tell how they will actually impact the state of laboratory automation.
2.4.2 Cloud Computing
Cloud computing, a vital component in drug discovery, involves the on‑demand delivery
of congurable computing resources over a network. It relies on virtualisation technol‑
ogy, where a hypervisor software abstracts and shares hardware resources as virtual
machines (VMs). This environment allows for rapid scaling of resources with minimal
interaction with the service provider (Spjuth etal., 2021).
The growing reliance on ML and statistical modelling underscores the need for
effective data management solutions that ensure rapid and efcient data access. The
data collection phase in the ML life cycle involves various tasks, including conducting
experiments to generate data, extracting existing data from databases, and selecting spe‑
cic data for modelling. Cloud computing plays a crucial role in this context, offering
scalable and on‑demand infrastructure for storage, databases, and middleware without
initial costs. Using cloud services for data hosting provides easy and quick global access
to data, particularly when modelling is conducted within cloud environments (Dalpé
and Joly, 2014). This accessibility not only streamlines the modelling process but also
enhances collaborative efforts and simultaneous access.
Before the advent of cloud computing, drug discovery organisations requiring signi‑
cant computational power were constrained to on‑premise servers, entailing substantial
overhead costs and maintenance responsibilities. This setup posed insurmountable chal‑
lenges for smaller entities and formidable obstacles even for larger corporations, render‑
ing it an impractical endeavour for most. The emergence of cloud computing, particularly
through infrastructure or platform as a service, has revolutionised this landscape, democra‑
tising access to computing resources previously unattainable to biotechnology companies.
Cloud based platforms such as Microsoft’s Azure or Amazon’s AWS for example,
have meanwhile entered majority of the corporate pharmaceutical world with a more or
less complete set of governance, storage and computing tools on demand. In the absence
of such solutions, companies like LabGenius would have required large amounts of
additional human resources and may have deemed the endeavour nancially unviable.
Many small teams and companies nd themselves in similar circumstances, where the
cost of leveraging ML technologies would have otherwise been prohibitive, barring
those organisations for which such technologies constitute the core business operations.
Cloud computing has signicantly lowered this barrier, enabling even the smallest of

26 Biopharmaceutical Informatics
teams to engage in exploratory ML initiatives by leveraging appropriate infrastructure
and platforms to assess the viability of integrating ML into their workows.
From a data acquisition and analysis standpoint, cloud computing has also facilitated
the adoption of distributed, asynchronous, simultaneous, and secure data handling prac‑
tices. Notably, Laboratory Information Management Systems (LIMS) and Electronic
Laboratory Notebooks (ELN) now benet from automated backups, universal acces‑
sibility across connected devices, and instantaneous sharing and analysis capabilities
across disparate locations and users. The cloud environment affords streamlined system
management, allowing for immediate implementation of software modications with
unprecedented agility and efciency. Increasingly, alongside traditional off‑the‑shelf
LIMS, custom bespoke applications are being developed and integrated due to their
low creation and deployment barriers. These tailored applications can be ne‑tuned to
meet the specic requirements of either a company as a whole or the distinct needs of
individual scientists. This trend highlights a shift towards greater efciency and adapt‑
ability in scientic research settings.
Cloud computing has three primary service levels: infrastructure as a service
(IaaS), platform as a service (PaaS), and software as a service (SaaS), each providing
an increasing level of abstraction from the hardware. There are four main cloud deploy‑
ment models: public cloud (accessible to the general public), private cloud (deployed on‑
premise for a single organisation), community cloud (shared by multiple organisations
with similar concerns), and hybrid cloud (a combination of any two models, often public
and private, with a software layer for seamless application portability). OpenStack, a
modular open‑source software stack, is widely used for operating cloud environments
in academia and by local cloud providers (www.openstack.org/).
Major public cloud providers offer extensive catalogues of virtual machine images
(VMIs) with pre‑installed software environments, enhancing tool distribution in drug
discovery and improving reproducibility.
Containerisation offers a lightweight method for deploying VMs by packaging an
application’s code, runtime, dependencies, and settings into a portable and repeatable
unit. It allows applications to run consistently across various platforms and is more
resource‑efcient than VMs. Kubernetes (K8s), developed by Google, is the leading
container orchestration platform, enabling easy management of containerised applica‑
tions. It’s commonly used in private cloud platforms and major cloud providers for scal‑
able access to computing resources.
While software as a service requires minimal IT skills, efciently working with
infrastructure as a service and platform as a service necessitates basic knowledge of
Linux or programming languages like Python.
Private clouds and on‑premises Kubernetes (K8s) clusters serve as valuable com‑
plements in this ecosystem (Emami etal., 2019). They provide scientists with the ability
to swiftly repurpose existing hardware resources to meet the diverse application needs
at different stages of the ML life cycle. This approach enables a more efcient and
cost‑effective way to manage computational resources, catering to the dynamic and
sometimes unpredictable requirements of ML projects. The integration of private clouds
and K8s clusters into the ML workow thus represents a strategic approach to balancing
resource demands while maintaining operational exibility and control.

2 • Digital Transformation 27
Using infrastructure as a service (IaaS) resources from cloud providers presents
a viable alternative, offering access to extensive resources on a pay‑per‑minute basis,
eliminating upfront and maintenance hardware costs. However, the costs for advanced
cloud infrastructure, such as graphics processing units (GPUs) or large VMs with ample
random‑access memory (RAM), should not be underestimated. Teams with consistently
high GPU demands are increasingly establishing private Kubernetes (K8s) infrastruc‑
tures to support GPU‑powered, containerised workows.
2.4.3 Machine Learning and MLOps to Support
the Machine Learning Life Cycle
The integration of machine learning (ML) in drug discovery involves comprehensive
management of data and model life cycles. This process encompasses stages like data
gathering, pre‑processing, training, validating, and deploying models, along with pre‑
dictions and inference. ML Operation (MLOps), a blend of practices and software, sup‑
ports these stages, enhancing reproducibility and robustness in iterative drug research.
Since machine learning in drug discovery is covered amply in other resources (e.g.
Vijayan etal., 2022; Sarkar et al., 2023; Wossnig et al., 2024), we mainly focus on
MLOps in the following section.
MLOps includes a range of functions primarily concerned with facilitating deploy‑
ment, audit logging, and reproducibility of ML systems, vis à vis version control and
automation for code, data, pipelines, and models. In drug discovery, one frequently
encounters changes in the experimental data that suddenly changes model performance.
Often, there is also a lack of metadata which hinders the diagnosis of errors and trouble‑
shooting. MLOps, specically model and data versioning, can enable the rapid diagno‑
sis of errors, reverting to previous models, and generally better data‑model‑prediction
lineage that ensures that the right method is applied to the right problem at the right
time.
Cloud computing offers scalable, resilient computing resources crucial for MLOps
and ML operations. Containerisation, coupled with platforms like Kubernetes and scien‑
tic workows, enables sturdy, reproducible ML analysis (e.g. consistent use of environ‑
ment and library versions), and higher speed in drug discovery. Additionally, emerging
federated methods promise to facilitate collaborative ML across multiple organisations,
leveraging both private and cloud‑based infrastructure.
Most drug discovery projects, specically in the lead optimisation stage follow the
iterative DBTL process (Figure2.3). ML is used to enhance decision‑making, particu‑
larly in the Learn phase, using both project‑specic and global models. MLOps emerges
as a critical framework, integrating data engineering, science, and operations to manage
the ML life cycle in production. It aims to streamline processes from data preparation to
model deployment, ensuring up‑to‑date models are accessible. Cloud services support
MLOps stages, with tools like Vertex AI, KubeFlow, STACKn, and H2O.ai aiding in
public and private infrastructures. Continuous ML modelling requires the integration of
data handling, pre‑processing, quality control, and validation into a reproducible pipe‑
line for optimal decision‑making in drug discovery.

28 Biopharmaceutical Informatics
FIGURE2.3 The use of ML models in the DBTL cycle in drug discovery. Within a specic
project, models are trained using data from the Test phase and used to make predictions
that support the Learn phase, such as running an assay to optimise a molecule for binding
to a particular target. Global models can then be used for predictions, such as predicting
binding to off-targets. Global models are typically trained using data collected from many
projects or merged with publicly available data. By contrast, project models are developed
for a particular objective within a drug discovery project. Project models are typically manually trained, and may need updating. Adapted from Spjuth etal. (2021).
FIGURE2.4 The machine learning life cycle. Diagram adapted from Spjuth etal. (2021).
The ML life cycle (Figure2.4) in drug discovery involves multiple steps with spe‑
cic infrastructure and software needs. Cloud IaaS reduces the need for on‑premises
computing infrastructure, offering scalable, on‑demand resources. Data collection is
crucial, requiring efcient data management solutions. Pre‑processing involves han‑
dling artefacts and quality control, with cloud computing facilitating transparent,
reproducible workows. Model training and validation are resource‑intensive, whereas
cloud IaaS offers a cost‑effective alternative to on‑site infrastructure. Prominent tools
like TensorFlow and PyTorch are readily available in cloud environments and have sig‑
nicantly simplied the process of training models on complex infrastructure such as
GPUs–thereby reducing the barrier of entry. Kubernetes‑based scientic workows,
exemplied by Recursion Pharma and BenevolentAI, orchestrate model development
pipelines efciently. For example, AlphaFold (Jumper etal., 2021) was quickly inte‑
grated by a large number of companies due to its availability in a container (c.f. the
AlphaFold GitHub repository, https://www.github.com/google‑deepmind/alphafold).
The initial step is data collection. MLOps should enable careful versioning of data
(e.g. using DataVersion Control(DVC), Pachyderm, or Neptune). This is essential to
later on trace the lineage from a dataframe (a specic version of the data) to a model and
the model predictions (results). Without proper data versioning and clear pointers to the
most recent data version, prediction errors or delays can easily occur.

2 • Digital Transformation 29
After selecting and assembling data in drug discovery, a crucial step involves
pre‑processing, which addresses issues like duplicate data, missing values, normalisa‑
tion, augmentation, and quality control. Cloud computing signicantly aids this process
by simplifying the construction and execution of pipelines or workows. This enhance‑
ment contributes to increased transparency, reproducibility, and robustness in data
handling.
Many workow systems in the cloud allow for declarative specications of anal‑
ysis pipelines, with capabilities to execute these workows on both public and private
clouds. Nextow, a popular workow engine in life science research, can directly
execute workows on infrastructure as a service (IaaS) resources and Kubernetes
clusters. Additionally, Argo and Pachyderm are two Kubernetes‑native workow
systems.
Notebooks deployed on cloud resources are commonly used for specifying pre‑pro‑
cessing steps. This approach not only facilitates pre‑processing execution but also sup‑
ports visual interpretations, enhancing understanding and analysis. Implementing a
complete workow of all pre‑processing steps in the cloud ensures that the pre‑process‑
ing is reproducible, portable, and scalable. On the other hand, notebooks allow people
to write lower‑quality and potentially unsafe code. Here again, MLOps practices can be
applied to mitigate risks and follow good standards.
Moreover, numerous platform as a service (PaaS) services focussed on big data
pre‑processing and analysis are available as public cloud services. An example is
Databricks, which is based on Apache Spark, offering advanced capabilities for han‑
dling and analysing large datasets in cloud environments.
After dataset construction, the focus shifts to model development, including train‑
ing, validation, optimisation, and model hyperparameter tuning. This phase can be
resource‑intensive and time‑consuming. Cloud computing offers signicant advantages
here, providing exibility and scalability than on‑site workstations. These workstations,
while powerful, are not scalable, require substantial upfront investment, and necessitate
ongoing maintenance.
For AI modelling in drug discovery, software stacks like TensorFlow, PyTorch, and
SciKit‑Learn (the rst two are primarily for deep learning) are prominent. These tools
are readily available in cloud environments, with virtual machine images (VMIs) and
container images facilitating rapid setup for single‑node infrastructures. Kubernetes is
widely utilised for scaling across multiple nodes, along with specialised AI frameworks
for Kubernetes.
Scientic workows, orchestrated using engines like SciPipe and AirFlow, are
employed to manage complex model development processes, including tasks like nested
cross‑validations. These technologies and practices illustrate the evolving landscape of
AI modelling in drug discovery, where cloud computing and containerisation play criti‑
cal roles in enhancing efciency and scalability.
Model serving and inference in the ML modelling life cycle involve making AI
models accessible to end‑users in a production‑grade environment with appropri‑
ate governance. This step, often the most technically challenging, requires not only
model validation (Smiatek et. al., 2021) and versioning but also critical features like
privacy, access control, auditability, logging, monitoring, and a resilient infrastructure
capable of recovering from failures.

30 Biopharmaceutical Informatics
Modern software engineering practices, particularly DevOps, which merges devel‑
opment (Dev) and operations (Ops), offer valuable insights for this stage. Cloud com‑
puting and infrastructure‑as‑code have revolutionised DevOps, providing a exible,
scalable environment suited for production‑grade solutions. Kubernetes, popular in
DevOps, supports these requirements, offering features essential for serving ML mod‑
els, including resilience and scalability.
Public clouds typically provide services for serving and publishing ML models, and
numerous cloud‑native open‑source frameworks have emerged, leveraging Kubernetes.
These include TensorFlow Serving, Seldon, Kubeow, and STACKn, or ones directly
hosted by the cloud provider such as Google Kubernetes Engine (Google Cloud) or
Amazon Elastic Kubernetes Service (AWS) which facilitate model management and
serving.
An important aspect of model serving is making models available via an Application
Programming Interface (API), allowing for integration with other software components
and enabling cloud‑based inference, essentially predictions as a service (SaaS). In drug
discovery, platforms like OpenRiskNet and DeepCell Kiosk exemplify this approach.
OpenRiskNet, built on OpenShift, a Kubernetes distribution, allows for publishing AI
model services and adds layers of discoverability and interoperability. DeepCell Kiosk,
designed for microscopy image analysis using TensorFlow, also utilises Kubernetes for
scalable deployment and inference. These examples demonstrate the evolving integra‑
tion of cloud computing and AI in drug discovery, highlighting the critical role of model
serving and inference in the ML life cycle.
2.4.4 Company Culture That Drives
a Data‑Driven Approach
Senior management sponsorship and widespread cross‑functional support of a data
engineering culture, as opposed to just science, is integral to digital transformation and
a data‑driven approach (Saldanha, 2019). Take the example of MLOps, integrating data
engineering, science, and operations to manage the ML life cycle in production. Teams
need to work together with a shared understanding of the framework. However, chal‑
lenges often remain with respect to a company‑wide holistic adoption of digitisation.
Reinhardt etal. (2020) note that knowledge of Industry 4.0 is concentrated at the more
senior levels of organisations and highlights a disconnect based on seniority. To over‑
come this, data scientists should be hired early, and be given a senior position where
they can shape the company‑wide mindset from the get‑go. Senior team members should
articulate a compelling vision for becoming a data‑driven organisation. As an example,
GlaxoSmithKline has established the role of Senior Vice President Global Head of
Articial Intelligence and Machine Learning; Relay Therapeutics has established the
roles of a Chief Data Ofcer and Senior Vice President, Articial Intelligence; and
Recursion the role of Chief Technology Ofcer–to name just a few. Teams should be
trained and given the opportunity to learn from one another–for example, by encourag‑
ing biologists and data scientists to engage in conversation and to participate in seminars
together.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
