Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5387_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Contents
- •Foreword
- •Preface
- •About the Editors
- •Contributors
- •References
- •2.3.4 Barriers to Automation Adoption
- •2.4 Core Ingredients for Successful Digital Transformation
- •2.1 Introduction
- •2.3.1 Operational Challenges
- •2.3.2 Cultural Challenges
- •2.4.2 Cloud Computing
- •2.5 Case Studies of Successful Digital Transformation
- •2.6 Conclusion
- •References
- •3. Computational Protein Design Strategies for Optimization of Antigen Generation to Drive Antibody Discovery
- •3.1 Introduction
- •3.3 Antigen Generation Strategies
- •3.4 Computational Methods
- •3.4.2 Computational Protein Structure Prediction
- •References
- •4. Bioinformatic Analyses of Antibody Repertoires and Their Roles in Modern Antibody Drug Discovery
- •4.1 Introduction
- •4.6 Summary and Future Directions
- •Acknowledgments
- •References
- •5.1 Introduction
- •5.2 Databases
- •5.2.1 Databases in Machine Learning Approaches
- •5.2.2 Database Types
- •5.3 Applications of Machine Learning in Antibody Discovery and Development
- •5.3.1 Structure Prediction with Deep Learning
- •5.3.3 Developability
- •5.4 Antibody Generation and Design by Language Models
- •5.4.1 Antibody Representations
- •5.4.2 Representation Learning
- •5.4.3 Language Models
- •References
- •6.1 Introduction
- •6.2 Antibody Generation through Deep Generative Models
- •6.3.1 Sampling and Scoring
- •6.5 Conclusions and Perspectives
- •Acknowledgments
- •References
- •7.1 Introduction
- •7.2.3 Computational Approaches to Predict Antibody–Antigen Interaction
- •7.3 Conclusion
- •Competing Interests
- •Acknowledgments
- •References
- •8.2 Common Types of Molecular Simulations for Biomolecules
- •8.2.1 Molecular Dynamics (MD) Simulations
- •8.2.2 Monte Carlo (MC) Simulations
- •8.2.3 Challenges of Molecular Simulations
- •8.3.1 Periodic Boundary Conditions
- •8.4 Uses of Molecular Simulation in Antibody Drug Development
- •8.5 Conclusion
- •References
- •9. Considerations of Developability During the Early Stages of Antibody Drug Discovery and Design
- •9.1 Introduction
- •9.2 Historical Perspective
- •9.3 Clinical Antibody Data Set
- •9.5 Control Antibodies
- •9.7 Assessment of Chemical Liabilities
- •9.8 Conclusions and Future Perspectives
- •Acknowledgments
- •References
- •Abbreviations
- •10.1 Introduction
- •10.4.1 Conclusions and Outlook
- •Acknowledgments
- •References
- •11.8 Conclusions and Future Directions
- •References
- •12.1 Introduction to PK/PD and QSP Modeling
- •12.1.1 PK/PD Modeling
- •12.1.2 QSP Modeling
- •12.2.1 Monoclonal Antibodies (mAbs)
- •12.2.3 Cell Therapies
- •12.2.4 Gene Therapies
- •12.2.5 Vaccines
- •12.2.6 mRNA/siRNA/Oligonucleotide Therapeutics
- •12.4 Case Studies
- •12.5 Conclusions and Future Perspectives
- •References
- •13.1 Introduction
- •13.2 AI/ML: A Game Changer for Antibody Design
- •13.3 Multispecific Antibody Design
- •13.4 Adapting AI to the Design of Multispecific Antibodies
- •13.4.1 Structure Prediction and Modeling
- •13.4.2 Developability Prediction and Optimization
- •13.4.4 In Silico Modeling and Simulation
- •13.5 The Future: Beyond Optimization
- •13.5.1 Market Trends and Commercialization
- •13.5.2 Logic Gates, Biosensors, and De Novo Design
- •13.5.3 Challenges and Opportunities
- •13.6 Conclusion
- •Acknowledgments
- •References
- •Index

Digital
Transformation in
the Biopharmaceutical
Industry
Rebuilding the Way
We Discover Complex
Therapeutics
Tonya Frolov, Leonard Wossnig,
and Alexander Jung
2
2.1 INTRODUCTION
The biopharmaceutical sector, driven by the possibility of curing diseases, has made tre‑
mendous progress over the last decades: the rst biosynthetic human insulin in 1982; the
discovery of clustered regularly interspaced short palindromic repeats (CRISPR) and
the advent of synthetic biology; the Human Genome Project; the rst approved mono‑
clonal therapeutic antibody in 1996; the discovery of the Yamanaka factors in 2006; the
explosion of omics; the approval of the rst cell therapy in 2017; and the rapid growth of
11

12 Biopharmaceutical Informatics
biologics which in 2022 accounted for 40% of the 37 drugs approved by the Food and
Drug Administration (FDA) (de la Torre & Albericio, 2023).
Correspondingly, the pharmaceutical sector has experienced continued and signi‑
cant growth. Worldwide sales exceeded $1.13 trillion in 2017 and $1.48 trillion in 2022
(Mikulic, 2024). The total spending and global demand for medicines will increase to
approximately $1.9 trillion by 2027 (IQVIA, 2023).
Yet scientic and technological gains have not increased the efciency of research
and development (R&D). As noted by Scannell and Bosley, ination‑adjusted costs
per novel drug increased nearly 100‑fold between 1950 and 2010 and drugs are more
likely to fail in clinical development today than in the 1970s (Scannell & Bosley, 2016).
Recent estimates for the mean cost of developing a new small‑molecule drug range
from $314million to $2.8 billion (Wouters etal., 2020; Sarkar etal., 2023). Despite
a lot of effort to improve success rates, the majority of clinical trials continue to fail
and the probability of positive outcomes in phase I clinical trials remains low (~10%)
(Takebe etal., 2018). Most importantly, despite tremendous progress, too many diseases
remain incurable. Those that remain to cure are typically multi‑factorial and increas‑
ingly complex.
The scale of operations required to develop drugs, diagnostics, and devices is
unparalleled. The average time it takes to get a new drug approved is 13 years (Paul
etal., 2010), characterised by clinical trials and regulatory approval which requires
time. As an example, in 2019, Novartis alone was running 500 clinical trials across >70
countries, with over 80,000 patients participating. In parallel, scientists were design‑
ing and developing new therapeutics, products, and devices (Finelli & Narasimhan,
2020).
Digital transformation is the fundamental rewiring of how a company operates.
It is the process of adoption and implementation of digital technology with the aim of
being nimble, reducing costs, and improving efciency and productivity (Hole etal.,
2021). Digital transformation has already impacted every industry, although it is moving
much more rapidly in some than in others. Notably, aerospace, automotive industry, and
banking have been quick to adopt digital transformation that generates business value in
measurable ways. Surprisingly, the biopharmaceutical industry has lagged behind this
revolution. The last several years have seen a surge in digital transformation within this
sector, bringing optimisations across all facets of healthcare, including research, drug
discovery, clinical trials, hospital management, surgery, diagnostics, monitoring, and
wearables.
The discovery and development of new therapeutics comprises target selection, lead
discovery, lead optimisation, pre‑clinical testing, drug product development, and clini‑
cal testing. This is one of the most investment‑intensive, research‑intensive, and strictly
regulated sectors. The processes involve difcult, manually intensive, time‑consuming
tasks that are prone to human error. The complexity of human biology further renders
drug discovery and development very risky.
In parallel, the number of products and amount of data under review by the regula‑
tory authorities has been increasing signicantly over the last 20 years, accompanied
by a trend towards more comprehensive and cross‑validated evaluation processes and

2 • Digital Transformation 13
regulations (ICH, 2002). The regulatory authorities, notably the US Food and Drug
Administration (FDA), have also in recent years opted for improved drug data and
knowledge life cycle management (ICH, 2019). As this is a critical part of the bio‑
pharmaceutical value chain, it may trigger a much more focussed data and knowledge
transformation within the biopharmaceutical industry. In 2019, the FDA announced its
Technology Modernization Action Plan (FDA, 2019). In 2021, it further announced the
reorganisation of its information technology, data management, and cybersecurity func‑
tions, signalling continued commitment towards its technology and data modernisation
efforts (FDA, 2021), the aim of which is to move closer to the speed of industry by
improving timely delivery of services.
Beyond the need for companies to implement data transformation to keep
up with regulatory requirements, it offers a significant opportunity to leverage
cross‑product and cross‑project insights to avoid repeating mistakes and to improve
productivity and success rates. Publicly available data generated from healthcare
research and adjacent industries, such as from biobanks, biomarker studies, and
large‑population studies, is only growing. This results in increased potential for
predictive modelling in disease understanding, prevention, and cure. Last but not
least, it may already help to predict clinical success or failure in the very early
stages of testing a hypothesis.
Here, we review the principles of successful digital transformation that can be
applied specically to the earlier steps of the drug development pipeline after target
selection: lead discovery and lead optimisation. We discuss how automation, robotics,
and the systemic application of machine learning (ML) and predictive analytics could
generate novel insights and better molecules with lower non‑mechanism toxicity, faster.
The more complex the therapeutic and/or problem to solve for, the more we stand to
gain from the potential of a digital transformation. We propose a framework for wider
digital transformation, including a more strategic approach to experimentation, reduced
empiricism, and reshaping company culture, using the discovery of complex therapeutic
modalities, such as multi‑specic antibodies, as a case study.
As with digital transformation across any sector, the goal is to facilitate faster deci‑
sion‑making and operational efciency by generating novel data‑driven insights, or in
other words–better, faster, and more affordable. Within the biopharmaceutical sector,
this includes testing and falsifying hypotheses as early as possible and with less experi‑
mental effort and cost. By not starting clinical studies that are unlikely to be successful,
the biopharmaceutical sector would revolutionise the availability of resources needed to
generate new, better, disease hypotheses, ultimately beneting patients.
Although other sectors have witnessed a more rapid adoption of digital transforma‑
tion, it is evident that the biopharmaceutical sector is capable of drug discovery 4.0.
New analytical and technological developments, such as automation, targeted data
acquisition, and data from external knowledge bases, can now be integrated to gener‑
ate high‑quality data for ML methods. The continuous acceleration of data generation,
combined with unprecedented data storage availability and exponentially growing com‑
puting power are enabling data‑driven decision‑making. This, in turn, can redene the
iterative process of experimentation into one that is driven by insight.

14 Biopharmaceutical Informatics
2.2 CURRENT STATE OF DIGITALISATION
IN PRE‑CLINICAL R&D
Articial intelligence (AI) has vitalised computer‑aided drug design, enabled by
the increasing availability of large and well‑curated public and proprietary datasets,
enhanced computing capabilities, and access to cloud storage (Sarkar etal., 2023). AI
can be dened as the science and engineering of creating intelligent machines that
have the ability to learn, reason, perceive, interact, and solve problems autonomously
(Manning, 2020). Many specic methods underpin AI systems; deep neural networks
are at the heart of the current AI renaissance but other ML approaches may form com‑
ponents of AI systems; in general, they all seek to make a prediction or estimation based
on some input data via some learned mathematical function.
One application of AI and ML is through the use of trained prediction models from
a known dataset, to make inferences about as yet unseen data. This is known as interpo‑
lation and extrapolation depending on whether we are working inside or outside of the
known data distribution. This is particularly relevant for biology where our understand‑
ing of underlying molecular mechanisms is often limited, and for drug discovery where
information is missing for novel therapeutics or novel biology. A well‑known example
is using ML methods and statistical modelling for predicting protein structures from
amino acid sequences (Duch etal., 2007).
AI has already shown positive results in research and in speeding up clinical trials
as a result of the growing abundance of pharmacological and patient data. A well‑known
application of AI is the assessment of molecular characteristics by in silico screening
and the identication of molecules with desired properties (Klaus, 2023). For example,
a model capable of predicting the binding strength of a molecule may be developed
based on data from physically measuring binding afnities for a range of substrates
(Duch etal., 2007). Other applications include but are not limited to, peptide synthesis,
structure‑based virtual screening (Thompson etal., 2022; Varkaris etal., 2024), toxicity
prediction (Klambauer etal., 2023), and drug repositioning (Gupta etal., 2021).
Previously, there has been a tendency for high‑throughput systems with low biologi‑
cal complexity, as is the case with large‑scale screening for binding afnity. The chal‑
lenge is that the data rarely translates to the desired clinical, functional response because
existing proxy models are insufcient (Bender and Cortes‑Ciriano, 2021a; Bender and
Cortes‑Ciriano, 2021b). But AI can now be deployed to solve unique or complex problems,
where human intelligence is ill‑equipped, or where traditional methods would simply take
too long, or where earlier computational methods did not sufce, see e.g. (Lyu etal., 2024).
While it is indisputable that AI holds immense potential for accelerating drug dis‑
covery and development, the current focus needs to shift from building more models to
acquiring ‘good’ high‑quality data that will optimise the functionality of these models.
There is little to no data available for novel therapeutic candidates or novel targets, and
thus general models are unlikely to work for lead optimisation onwards. This is in con‑
trast to large language models (LLMs), where large, averaged, datasets are readily avail‑
able on the internet and sufce. For drug discovery, we need complex biological data

2 • Digital Transformation 15
with a higher predictive validity (Scannell etal., 2022), i.e. data from proxy models that
are more predictive of patient outcomes, even if it is challenging to generate. Data needs
to be both clinically relevant and of sufcient quality to train the models–i.e. machine
learning‑grade data. Key properties include data veracity, variety, volume, velocity, and
value, or the ve ‘Vs’ of drug discovery as we refer to them. For a more in‑depth review
of what we mean by high‑quality data, refer to Wossnig (2023).
The integrity of data collection, storage, protocols, and standardisation is critical. This
needs to be supported by a technology stack that maximises consistency and reproduc‑
ibility. If satised, this will result in superior predictive models and meaningful insights.
2.3 CHALLENGES IN DIGITAL
TRANSFORMATION OF PRE‑CLINICAL R&D
The requirement for automation and standardisation of reproducible, scalable, and clini‑
cally relevant data is now largely accepted, and many of the hardware and software tools
that facilitate the creation of automated experimental and data workows are now either
available or under development. Here, we discuss how companies can adopt these, and
the challenges that need to be overcome.
2.3.1 Operational Challenges
The integration of different technologies comes with its challenges. Teams typically
need to integrate software that ranges from the latest available technologies to those
that may be a couple of decades old, and as such there is no standardised approach for
achieving this. Reliance on web‑based software only is not possible when instruments
are connected to personal computers. A further challenge is that instrument processes
may be complex, or adapted for a specic and unique lab purpose. The development of
automation needs to be balanced with providing manual exibility for scientists to adapt
experimentation and workows.
Furthermore, tools are not typically suitable for immediate use by a scientist and
instead require the input from or execution by software engineers, which can limit
broader adoption. Given the complexity of workows and data, it is unlikely that scien‑
tists will acquire the skills of a software engineer, highlighting the ongoing requirement
for a mixed team of experts. While some people view the scientist of the future as a
hybrid of software developer and experimental biologist, we believe that the eld will
continue to require specialists for each area.
It is also difcult to maintain a consistent adoption of technologies across all teams,
which is particularly challenging for larger organisations. Modern software designed
for personal use is typically designed to be user‑friendly (e.g. the iPhone), but this is not
always the case in the biopharmaceutical sector. Teams from varied career backgrounds
show a range of technical competence, making it difcult to integrate technologies that

16 Biopharmaceutical Informatics
could be adopted by everyone. Scientists may desire to implement technologies they are
most familiar with, rather than those that would objectively benet the team at large.
Training teams is imperative but requires time investment, which may be difcult given
the pressure to deliver assets in biopharmaceutical organisations. Training is further
challenged by employee turnover. One of the trends that are forseen for the near future
would be, that the data generated by instruments and lab scientists feed directly into a
holistic data standard aka knowledge graph, reducing the need for dedicated data acqui‑
sition software and all their logistcal disadvantages.
2.3.2 Cultural Challenges
It is widely accepted that in early drug development and research organisations, cultural
resistance persists towards digital transformation. Perhaps a contradiction to the notion
that scientists should thrive for innovation, or a result of scientic mistrust of the benet
that new technologies could provide, possibly owing to the lack of success stories in the
long journey to bring a new therapeutic to the market.
There often appears to be a generational and diversity gap in current teams, and
one conclusion that could be drawn is that there is a reluctance to adopt digital transfor‑
mation triggered by external forces simply because it comes from elsewhere. However,
given that most scientists have adopted the benets of digital transformation in their
personal lives, one could suppose that they would be equally willing to adopt digital
transformation within their own eld of expertise.
Here, we will see a huge shift in the teams within the biopharmaceutical industry,
as the trend for agile and diverse product teams begins to take over the product devel‑
opment framework. Over time, this will lead to improved communication of scientists,
engineers, and data scientists.
Until this transformation materialises on a larger scale, communicating the value of,
and involving teams with, digital transformation is critical. For example, AI and genera‑
tive AI tend to be vastly misunderstood in the sense that either expectation rise too high,
or scepticism hinders the adoption. In order to become widely accepted, their properties,
benets, and use cases need to be explained during training, roadshows, and pilot proj‑
ects across the R&D sector. Perhaps more fundamentally, organisations need to think
differently about what constitutes a valuable data asset. Traditionally, data was merely a
stepping stone to nding candidates whereas now data is in itself a highly valuable asset.
With respect to the digital transformation of datasets, the benets are not always
immediate. Efforts need to be made towards maturing these datasets and towards
the FAIRication with an organisation that comprises 15 guiding principles out‑
lined by Wilkinson et al. (2016) aimed at enhancing the Findability, Accessibility,
Interoperability, and Reusability of data. A particular complexity is that experimental
data is typically expensive to produce and still relatively small scale. In order to create
large enough datasets for successful ML applications, datasets may need to be additive
and aggregated over many experiments and iterations. This gives rise to the difculty of
asserting consistent data models over time and ensuring the compatibility of datasets.
Other key considerations are that datasets should be representative, i.e. will want to

2 • Digital Transformation 17
include negative results as well as positive results, and will want to minimise bias such
that models trained on the data have an accurate and complete understanding of the
real‑world processes (Bender and Cortes‑Ciriano, 2021a; Bender and Cortes‑Ciriano,
2021b).
Data scientists in biotech require a rich understanding not only of data and engi‑
neering but also of complex science. This in turn may encompass biology, chemistry,
toxicology, etc. They need to be able to create the right datasets, process the data, and
make it timely available in the right format to the right model, and hence anticipate what
the scientists may want to infer from them (Benchling, 2023). Building a team with the
necessary expertise across AI, data science, data engineering, software engineering,
and science is critical. A strategy to enable this should be dened from the earliest
stages of company creation. Additionally, a company‑wide vision for the creation of the
right technology stack is also important.
The leadership’s challenge here is to link these activities directly to the company’s
business strategy and to the acceleration of bringing new therapeutics to patients. If
done successfully, teams will feel they are contributing to a meaningful outcome. As
more and more data roles and digitalisation roles are created within R&D organisations,
this will lead to an organic adoption and integration of digital tools in the daily work of
scientists.
2.3.3 Data Management, Data Analytics,
and Integrity Concerns
The establishment of a standardised digital framework for experimental data collection
that adheres to FAIR (Findable, Accessible, Interoperable, Reusable) principles, what
we referred to as FAIRication earlier, is critical to enhance the ability to nd, use, and
reuse the data. But challenges to achieving this persist, including data integration, which
requires managing and integrating large, diverse datasets; initial substantial investments
are needed to build and integrate advanced technologies such as new data management
systems (FTE time), and for clearing and organising existing data; incentivising sup‑
port, for example from shareholders; and ensuring the accuracy and reproducibility of
AI‑driven results, particularly in absence of benchmarks.
Experimental data in the early stages of drug discovery and research is as diverse as
the tasks and the purpose it is supposed to serve. There exist, however, critical inection
points throughout the value chain that result in the need for further data accumulation
and consolidation.
A robust data strategy can facilitate the entire R&D pipeline, and must take into
account three key attributes: data quality, data integrity, and data governance.
2.3.3.1 Data quality standards
While organisations need to ensure that data regulatory standards required for submis‑
sions are met (e.g. Standard for Exchange of Nonclinical Data (SEND), Identication
of Medicinal Products (IDMP), etc.), the need for structuring data and the need for

18 Biopharmaceutical Informatics
high‑quality data extends far beyond this. The value of the data comes not only from
insight‑driven decision‑making at critical inection points but from its ability to gener‑
ate new insights in earlier stages of R&D. This, however, requires that all data is col‑
lected to satisfy minimum quality criteria such as the capture of relevant metadata or
common identiers in order that it be usable.
2.3.3.2 Data integrity principles
ALCOA (Attributable, Legible, Contemporaneous, Original, and Accurate) and ALCOA
plus principles, introduced by the FDA, are often applied as data integrity measures to
data assets. Unfortunately, they impose high standards on data‑generation and automa‑
tion teams within R&D labs. However, if the data can be reused to generate high‑value
predictions going forward, then this apparent ‘burden’ turns into a competitive advan‑
tage. Furthermore, data integrity and organisation help statisticians, data scientists, and
data engineers to work collaboratively and integrate the data into the prediction frame‑
work with less domain knowledge.
2.3.3.3 Data governance
Researchers may encounter situations where information or data requests are not
directly denied but instead met with limitations based on perceived issues such as data
quality, researcher expertise, or condentiality concerns. This aligns with observations
in individual data‑sharing negotiations, where frameworks are established to address
condentiality, data protection regulations, and researcher qualications. A well‑estab‑
lished governance structure including roles such as data domain owners, and/or data
stewards is needed to maintain a high level of trust in the data‑sharing process. This
fosters data transparency and data usage along the value chain for non‑expert users to
generate new insights.
The landscape of available data is hugely diverse in terms of data standards, matu‑
rity, and ontological data models. Building data assets around domain knowledge
requires a huge initial effort to build the right taxonomies, ontologies, and data clas‑
sication rules within the sector. However, it is worthwhile and necessary to invest here,
mainly in the form of value‑added pilot use cases and projects, in order to create the data
platform to support AI and ML predictions in a sufcient way.
2.3.4 Barriers to Automation Adoption
Building out a high‑throughput, automated platform requires a high level of upfront
investment. But the appeal is that it can enable signicant scientic discovery, particu‑
larly in the context of complex and high‑impact problems that would otherwise take
decades, or longer to solve.
Given the signicant investment required, it is important to dene the concrete
problem that needs to be solved and to assess whether the development of the platform
is warranted. Parameters to consider that increase the value of the investment include:

2 • Digital Transformation 19
(i) the task is simple but highly repetitive and vulnerable to human error due to the sheer
repetitiveness of the task; (ii) the workow is highly complex and human interaction at
multiple levels along the workow would lead to an increased risk of error resulting in
catastrophic failure; (iii) the return on investment is such that a signicant saving in time
and in human resource can be realised with the ability to run automated tasks either
continuously or to completion with a high degree of precision and accuracy independent
of human shift patterns; (iv) the task is either amenable to automation or adapted to run
using automated processes, as automation is not always a like‑for‑like replacement for
tasks designed for manual handling.
2.4 CORE INGREDIENTS FOR SUCCESSFUL DIGITAL TRANSFORMATION
Four elements inuence digital transformation: (i) the steady acceleration of the rate at
which new data is generated; (ii) cloud computing that offers universally available data
storage and exponentially growing computing power; (iii) new ML methods/AI; and (iv)
a shift in culture to build teams equipped to develop and deploy a data‑driven approach
(Finelli & Narasimhan, 2020). Here, we review developments enabling the rapid digital
transformation of drug discovery.
2.4.1 New Data‑Generation Methods
and Laboratory Automation
In synthetic biology, considerable progress has been made towards achieving autono‑
mous and scalable experimentation, largely due to the combination of automation and
ML (Martin etal., 2023). While fully automating a laboratory is an ideal goal in some
cases, it is not always the most practical or benecial approach. In certain situations,
simpler and more modular methods may be more effective (Stephenson etal., 2023).
For antibody engineering, given the current technological advancements, a modular yet
fully integrated platform is in our view most appropriate.
A key application of automation is in improving the Design‑Build‑Test‑Learn
(DBTL) cycle, commonly employed in drug discovery programmes. This cycle involves
designing a series of candidates, constructing, and testing them, then using the gathered
data to inform the design of the next series of candidates, often employing ML or sta‑
tistical modelling.
Effective implementation of DBTL frameworks in laboratories hinges on several
factors. Consistent, reliable data are vital for ML applications, necessitating rigorous
tracking and monitoring of reagents, samples, protocols, and business rules as well
as control compounds that are used in programmes. Automation in data handling and
metadata management is crucial to reduce errors commonly associated with manual
data entry or processing.

20 Biopharmaceutical Informatics
In individual labs, various factors can impede clear and reproducible experimental
results. These include inconsistent sample tracking, varying lab conditions, manual pro‑
cess errors (including simple copy‑paste errors of data), and non‑uniform data collection
and storage practices. Additionally, differences in commercial and proprietary instru‑
ments and protocols can introduce signicant batch effects, leading to inconsistencies
across different research groups and labs. Therefore, standardisation and automation
across an organisation are essential.
The advent of advanced liquid handling technologies has revolutionised the move‑
ment of reagents and samples in automated experimental workows, boosting repro‑
ducibility and throughput. When paired with data management tools that consolidate
experimental parameters and data, these technologies improve sample tracking and
error detection, aiding decision‑making in drug discovery programmes. Technologies
like next‑generation sequencing (NGS) or new deep screening of sequence‑sequence
level interactions (Porebski etal., 2023), facilitate the generation of large datasets suit‑
able for training ML models. These datasets can be used to evaluate system perfor‑
mance and shape future experimental designs.
However, further improvements are needed in hardware and software systems
to enhance connectivity between automated steps and equipment. Additionally, stan‑
dardising data infrastructures, exchange formats, protocol descriptions, ontologies, and
data reporting will improve data interpretability and promote information sharing. In
the following sections, we will examine the current status of automation systems perti‑
nent to biologics engineering.
Automation technologies in synthetic biology exhibit a wide range in format and
degree, encompassing various levels of human interaction (Figure2.1). As automation
FIGURE2.1 Levels of automation in synthetic biology laboratories achieved through hardware and software automation technologies. Adapted from Stephenson etal. (2023) and
Martin etal. (2023).
Соседние файлы в папке Библиотека им академика М.И. Перельмана
