Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5908_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
16 Мб
Скачать
☆
Digital Transformation in the Biopharmaceutical Industry
Rebuilding the Way We Discover Complex Therapeutics
Tonya Frolov, Leonard Wossnig, and Alexander Jung
2

2.1 INTRODUCTION

The biopharmaceutical sector, driven by the possibility of curing diseases, has made tre‑ mendous progress over the last decades: the rst biosynthetic human insulin in 1982; the discovery of clustered regularly interspaced short palindromic repeats (CRISPR) and the advent of synthetic biology; the Human Genome Project; the rst approved mono‑ clonal therapeutic antibody in 1996; the discovery of the Yamanaka factors in 2006; the explosion of omics; the approval of the rst cell therapy in 2017; and the rapid growth of
11
12 Biopharmaceutical Informatics
biologics which in 2022 accounted for 40% of the 37 drugs approved by the Food and Drug Administration (FDA) (de la Torre & Albericio, 2023).
Correspondingly, the pharmaceutical sector has experienced continued and signi‑ cant growth. Worldwide sales exceeded $1.13 trillion in 2017 and $1.48 trillion in 2022 (Mikulic, 2024). The total spending and global demand for medicines will increase to approximately $1.9 trillion by 2027 (IQVIA, 2023).
Yet scientic and technological gains have not increased the efciency of research and development (R&D). As noted by Scannell and Bosley, ination‑adjusted costs per novel drug increased nearly 100‑fold between 1950 and 2010 and drugs are more likely to fail in clinical development today than in the 1970s (Scannell & Bosley, 2016). Recent estimates for the mean cost of developing a new small‑molecule drug range from $314million to $2.8 billion (Wouters etal., 2020; Sarkar etal., 2023). Despite a lot of effort to improve success rates, the majority of clinical trials continue to fail and the probability of positive outcomes in phase I clinical trials remains low (~10%) (Takebe etal., 2018). Most importantly, despite tremendous progress, too many diseases remain incurable. Those that remain to cure are typically multi‑factorial and increas‑ ingly complex.
The scale of operations required to develop drugs, diagnostics, and devices is unparalleled. The average time it takes to get a new drug approved is 13 years (Paul etal., 2010), characterised by clinical trials and regulatory approval which requires time. As an example, in 2019, Novartis alone was running 500 clinical trials across >70 countries, with over 80,000 patients participating. In parallel, scientists were design‑ ing and developing new therapeutics, products, and devices (Finelli & Narasimhan,
2020).
Digital transformation is the fundamental rewiring of how a company operates. It is the process of adoption and implementation of digital technology with the aim of being nimble, reducing costs, and improving efciency and productivity (Hole etal.,
2021). Digital transformation has already impacted every industry, although it is moving much more rapidly in some than in others. Notably, aerospace, automotive industry, and banking have been quick to adopt digital transformation that generates business value in measurable ways. Surprisingly, the biopharmaceutical industry has lagged behind this revolution. The last several years have seen a surge in digital transformation within this sector, bringing optimisations across all facets of healthcare, including research, drug discovery, clinical trials, hospital management, surgery, diagnostics, monitoring, and wearables.
The discovery and development of new therapeutics comprises target selection, lead discovery, lead optimisation, pre‑clinical testing, drug product development, and clini‑ cal testing. This is one of the most investment‑intensive, research‑intensive, and strictly regulated sectors. The processes involve difcult, manually intensive, time‑consuming tasks that are prone to human error. The complexity of human biology further renders drug discovery and development very risky.
In parallel, the number of products and amount of data under review by the regula‑ tory authorities has been increasing signicantly over the last 20 years, accompanied by a trend towards more comprehensive and cross‑validated evaluation processes and
2 • Digital Transformation 13
regulations (ICH, 2002). The regulatory authorities, notably the US Food and Drug Administration (FDA), have also in recent years opted for improved drug data and knowledge life cycle management (ICH, 2019). As this is a critical part of the bio‑ pharmaceutical value chain, it may trigger a much more focussed data and knowledge transformation within the biopharmaceutical industry. In 2019, the FDA announced its Technology Modernization Action Plan (FDA, 2019). In 2021, it further announced the reorganisation of its information technology, data management, and cybersecurity func‑ tions, signalling continued commitment towards its technology and data modernisation efforts (FDA, 2021), the aim of which is to move closer to the speed of industry by improving timely delivery of services.
Beyond the need for companies to implement data transformation to keep up with regulatory requirements, it offers a significant opportunity to leverage cross‑product and cross‑project insights to avoid repeating mistakes and to improve productivity and success rates. Publicly available data generated from healthcare research and adjacent industries, such as from biobanks, biomarker studies, and large‑population studies, is only growing. This results in increased potential for predictive modelling in disease understanding, prevention, and cure. Last but not least, it may already help to predict clinical success or failure in the very early stages of testing a hypothesis.
Here, we review the principles of successful digital transformation that can be applied specically to the earlier steps of the drug development pipeline after target selection: lead discovery and lead optimisation. We discuss how automation, robotics, and the systemic application of machine learning (ML) and predictive analytics could generate novel insights and better molecules with lower non‑mechanism toxicity, faster. The more complex the therapeutic and/or problem to solve for, the more we stand to gain from the potential of a digital transformation. We propose a framework for wider digital transformation, including a more strategic approach to experimentation, reduced empiricism, and reshaping company culture, using the discovery of complex therapeutic modalities, such as multi‑specic antibodies, as a case study.
As with digital transformation across any sector, the goal is to facilitate faster deci‑ sion‑making and operational efciency by generating novel data‑driven insights, or in other words–better, faster, and more affordable. Within the biopharmaceutical sector, this includes testing and falsifying hypotheses as early as possible and with less experi‑ mental effort and cost. By not starting clinical studies that are unlikely to be successful, the biopharmaceutical sector would revolutionise the availability of resources needed to generate new, better, disease hypotheses, ultimately beneting patients.
Although other sectors have witnessed a more rapid adoption of digital transforma‑ tion, it is evident that the biopharmaceutical sector is capable of drug discovery 4.0. New analytical and technological developments, such as automation, targeted data acquisition, and data from external knowledge bases, can now be integrated to gener‑ ate high‑quality data for ML methods. The continuous acceleration of data generation, combined with unprecedented data storage availability and exponentially growing com‑ puting power are enabling data‑driven decision‑making. This, in turn, can redene the iterative process of experimentation into one that is driven by insight.
14 Biopharmaceutical Informatics
2.2 CURRENT STATE OF DIGITALISATION IN PRE‑CLINICAL R&D
Articial intelligence (AI) has vitalised computer‑aided drug design, enabled by the increasing availability of large and well‑curated public and proprietary datasets, enhanced computing capabilities, and access to cloud storage (Sarkar etal., 2023). AI can be dened as the science and engineering of creating intelligent machines that have the ability to learn, reason, perceive, interact, and solve problems autonomously (Manning, 2020). Many specic methods underpin AI systems; deep neural networks are at the heart of the current AI renaissance but other ML approaches may form com‑ ponents of AI systems; in general, they all seek to make a prediction or estimation based on some input data via some learned mathematical function.
One application of AI and ML is through the use of trained prediction models from a known dataset, to make inferences about as yet unseen data. This is known as interpo‑ lation and extrapolation depending on whether we are working inside or outside of the known data distribution. This is particularly relevant for biology where our understand‑ ing of underlying molecular mechanisms is often limited, and for drug discovery where information is missing for novel therapeutics or novel biology. A well‑known example is using ML methods and statistical modelling for predicting protein structures from amino acid sequences (Duch etal., 2007).
AI has already shown positive results in research and in speeding up clinical trials as a result of the growing abundance of pharmacological and patient data. A well‑known application of AI is the assessment of molecular characteristics by in silico screening and the identication of molecules with desired properties (Klaus, 2023). For example, a model capable of predicting the binding strength of a molecule may be developed based on data from physically measuring binding afnities for a range of substrates (Duch etal., 2007). Other applications include but are not limited to, peptide synthesis, structure‑based virtual screening (Thompson etal., 2022; Varkaris etal., 2024), toxicity prediction (Klambauer etal., 2023), and drug repositioning (Gupta etal., 2021).
Previously, there has been a tendency for high‑throughput systems with low biologi‑ cal complexity, as is the case with large‑scale screening for binding afnity. The chal‑ lenge is that the data rarely translates to the desired clinical, functional response because existing proxy models are insufcient (Bender and Cortes‑Ciriano, 2021a; Bender and Cortes‑Ciriano, 2021b). But AI can now be deployed to solve unique or complex problems, where human intelligence is ill‑equipped, or where traditional methods would simply take too long, or where earlier computational methods did not sufce, see e.g. (Lyu etal., 2024).
While it is indisputable that AI holds immense potential for accelerating drug dis‑ covery and development, the current focus needs to shift from building more models to acquiring ‘good’ high‑quality data that will optimise the functionality of these models. There is little to no data available for novel therapeutic candidates or novel targets, and thus general models are unlikely to work for lead optimisation onwards. This is in con‑ trast to large language models (LLMs), where large, averaged, datasets are readily avail‑ able on the internet and sufce. For drug discovery, we need complex biological data
2 • Digital Transformation 15
with a higher predictive validity (Scannell etal., 2022), i.e. data from proxy models that are more predictive of patient outcomes, even if it is challenging to generate. Data needs to be both clinically relevant and of sufcient quality to train the models–i.e. machine learning‑grade data. Key properties include data veracity, variety, volume, velocity, and value, or the ve ‘Vs’ of drug discovery as we refer to them. For a more in‑depth review of what we mean by high‑quality data, refer to Wossnig (2023).
The integrity of data collection, storage, protocols, and standardisation is critical. This needs to be supported by a technology stack that maximises consistency and reproduc‑ ibility. If satised, this will result in superior predictive models and meaningful insights.
2.3 CHALLENGES IN DIGITAL
TRANSFORMATION OF PRE‑CLINICAL R&D
The requirement for automation and standardisation of reproducible, scalable, and clini‑ cally relevant data is now largely accepted, and many of the hardware and software tools that facilitate the creation of automated experimental and data workows are now either available or under development. Here, we discuss how companies can adopt these, and the challenges that need to be overcome.

2.3.1 Operational Challenges

The integration of different technologies comes with its challenges. Teams typically need to integrate software that ranges from the latest available technologies to those that may be a couple of decades old, and as such there is no standardised approach for achieving this. Reliance on web‑based software only is not possible when instruments are connected to personal computers. A further challenge is that instrument processes may be complex, or adapted for a specic and unique lab purpose. The development of automation needs to be balanced with providing manual exibility for scientists to adapt experimentation and workows.
Furthermore, tools are not typically suitable for immediate use by a scientist and instead require the input from or execution by software engineers, which can limit broader adoption. Given the complexity of workows and data, it is unlikely that scien‑ tists will acquire the skills of a software engineer, highlighting the ongoing requirement for a mixed team of experts. While some people view the scientist of the future as a hybrid of software developer and experimental biologist, we believe that the eld will continue to require specialists for each area.
It is also difcult to maintain a consistent adoption of technologies across all teams, which is particularly challenging for larger organisations. Modern software designed for personal use is typically designed to be user‑friendly (e.g. the iPhone), but this is not always the case in the biopharmaceutical sector. Teams from varied career backgrounds show a range of technical competence, making it difcult to integrate technologies that
16 Biopharmaceutical Informatics
could be adopted by everyone. Scientists may desire to implement technologies they are most familiar with, rather than those that would objectively benet the team at large. Training teams is imperative but requires time investment, which may be difcult given the pressure to deliver assets in biopharmaceutical organisations. Training is further challenged by employee turnover. One of the trends that are forseen for the near future would be, that the data generated by instruments and lab scientists feed directly into a holistic data standard aka knowledge graph, reducing the need for dedicated data acqui‑ sition software and all their logistcal disadvantages.

2.3.2 Cultural Challenges

It is widely accepted that in early drug development and research organisations, cultural resistance persists towards digital transformation. Perhaps a contradiction to the notion that scientists should thrive for innovation, or a result of scientic mistrust of the benet that new technologies could provide, possibly owing to the lack of success stories in the long journey to bring a new therapeutic to the market.
There often appears to be a generational and diversity gap in current teams, and one conclusion that could be drawn is that there is a reluctance to adopt digital transfor‑ mation triggered by external forces simply because it comes from elsewhere. However, given that most scientists have adopted the benets of digital transformation in their personal lives, one could suppose that they would be equally willing to adopt digital transformation within their own eld of expertise.
Here, we will see a huge shift in the teams within the biopharmaceutical industry, as the trend for agile and diverse product teams begins to take over the product devel‑ opment framework. Over time, this will lead to improved communication of scientists, engineers, and data scientists.
Until this transformation materialises on a larger scale, communicating the value of, and involving teams with, digital transformation is critical. For example, AI and genera‑ tive AI tend to be vastly misunderstood in the sense that either expectation rise too high, or scepticism hinders the adoption. In order to become widely accepted, their properties, benets, and use cases need to be explained during training, roadshows, and pilot proj‑ ects across the R&D sector. Perhaps more fundamentally, organisations need to think differently about what constitutes a valuable data asset. Traditionally, data was merely a stepping stone to nding candidates whereas now data is in itself a highly valuable asset.
With respect to the digital transformation of datasets, the benets are not always immediate. Efforts need to be made towards maturing these datasets and towards the FAIRication with an organisation that comprises 15 guiding principles out‑ lined by Wilkinson et al. (2016) aimed at enhancing the Findability, Accessibility, Interoperability, and Reusability of data. A particular complexity is that experimental data is typically expensive to produce and still relatively small scale. In order to create large enough datasets for successful ML applications, datasets may need to be additive and aggregated over many experiments and iterations. This gives rise to the difculty of asserting consistent data models over time and ensuring the compatibility of datasets. Other key considerations are that datasets should be representative, i.e. will want to
2 • Digital Transformation 17
include negative results as well as positive results, and will want to minimise bias such that models trained on the data have an accurate and complete understanding of the real‑world processes (Bender and Cortes‑Ciriano, 2021a; Bender and Cortes‑Ciriano, 2021b).
Data scientists in biotech require a rich understanding not only of data and engi‑ neering but also of complex science. This in turn may encompass biology, chemistry, toxicology, etc. They need to be able to create the right datasets, process the data, and make it timely available in the right format to the right model, and hence anticipate what the scientists may want to infer from them (Benchling, 2023). Building a team with the necessary expertise across AI, data science, data engineering, software engineering, and science is critical. A strategy to enable this should be dened from the earliest stages of company creation. Additionally, a company‑wide vision for the creation of the right technology stack is also important.
The leadership’s challenge here is to link these activities directly to the company’s business strategy and to the acceleration of bringing new therapeutics to patients. If done successfully, teams will feel they are contributing to a meaningful outcome. As more and more data roles and digitalisation roles are created within R&D organisations, this will lead to an organic adoption and integration of digital tools in the daily work of scientists.
2.3.3 Data Management, Data Analytics,
and Integrity Concerns
The establishment of a standardised digital framework for experimental data collection that adheres to FAIR (Findable, Accessible, Interoperable, Reusable) principles, what we referred to as FAIRication earlier, is critical to enhance the ability to nd, use, and reuse the data. But challenges to achieving this persist, including data integration, which requires managing and integrating large, diverse datasets; initial substantial investments are needed to build and integrate advanced technologies such as new data management systems (FTE time), and for clearing and organising existing data; incentivising sup‑ port, for example from shareholders; and ensuring the accuracy and reproducibility of AI‑driven results, particularly in absence of benchmarks.
Experimental data in the early stages of drug discovery and research is as diverse as the tasks and the purpose it is supposed to serve. There exist, however, critical inection points throughout the value chain that result in the need for further data accumulation and consolidation.
A robust data strategy can facilitate the entire R&D pipeline, and must take into account three key attributes: data quality, data integrity, and data governance.
2.3.3.1 Data quality standards
While organisations need to ensure that data regulatory standards required for submis‑ sions are met (e.g. Standard for Exchange of Nonclinical Data (SEND), Identication of Medicinal Products (IDMP), etc.), the need for structuring data and the need for
18 Biopharmaceutical Informatics
high‑quality data extends far beyond this. The value of the data comes not only from insight‑driven decision‑making at critical inection points but from its ability to gener‑ ate new insights in earlier stages of R&D. This, however, requires that all data is col‑ lected to satisfy minimum quality criteria such as the capture of relevant metadata or common identiers in order that it be usable.
2.3.3.2 Data integrity principles
ALCOA (Attributable, Legible, Contemporaneous, Original, and Accurate) and ALCOA plus principles, introduced by the FDA, are often applied as data integrity measures to data assets. Unfortunately, they impose high standards on data‑generation and automa‑ tion teams within R&D labs. However, if the data can be reused to generate high‑value predictions going forward, then this apparent ‘burden’ turns into a competitive advan‑ tage. Furthermore, data integrity and organisation help statisticians, data scientists, and data engineers to work collaboratively and integrate the data into the prediction frame‑ work with less domain knowledge.
2.3.3.3 Data governance
Researchers may encounter situations where information or data requests are not directly denied but instead met with limitations based on perceived issues such as data quality, researcher expertise, or condentiality concerns. This aligns with observations in individual data‑sharing negotiations, where frameworks are established to address condentiality, data protection regulations, and researcher qualications. A well‑estab‑ lished governance structure including roles such as data domain owners, and/or data stewards is needed to maintain a high level of trust in the data‑sharing process. This fosters data transparency and data usage along the value chain for non‑expert users to generate new insights.
The landscape of available data is hugely diverse in terms of data standards, matu‑ rity, and ontological data models. Building data assets around domain knowledge requires a huge initial effort to build the right taxonomies, ontologies, and data clas‑ sication rules within the sector. However, it is worthwhile and necessary to invest here, mainly in the form of value‑added pilot use cases and projects, in order to create the data platform to support AI and ML predictions in a sufcient way.

2.3.4 Barriers to Automation Adoption

Building out a high‑throughput, automated platform requires a high level of upfront investment. But the appeal is that it can enable signicant scientic discovery, particu‑ larly in the context of complex and high‑impact problems that would otherwise take decades, or longer to solve.
Given the signicant investment required, it is important to dene the concrete problem that needs to be solved and to assess whether the development of the platform is warranted. Parameters to consider that increase the value of the investment include:
2 • Digital Transformation 19
(i) the task is simple but highly repetitive and vulnerable to human error due to the sheer repetitiveness of the task; (ii) the workow is highly complex and human interaction at multiple levels along the workow would lead to an increased risk of error resulting in catastrophic failure; (iii) the return on investment is such that a signicant saving in time and in human resource can be realised with the ability to run automated tasks either continuously or to completion with a high degree of precision and accuracy independent of human shift patterns; (iv) the task is either amenable to automation or adapted to run using automated processes, as automation is not always a like‑for‑like replacement for tasks designed for manual handling.

2.4 CORE INGREDIENTS FOR SUCCESSFUL DIGITAL TRANSFORMATION

Four elements inuence digital transformation: (i) the steady acceleration of the rate at which new data is generated; (ii) cloud computing that offers universally available data storage and exponentially growing computing power; (iii) new ML methods/AI; and (iv) a shift in culture to build teams equipped to develop and deploy a data‑driven approach (Finelli & Narasimhan, 2020). Here, we review developments enabling the rapid digital transformation of drug discovery.
2.4.1 New Data‑Generation Methods and Laboratory Automation
In synthetic biology, considerable progress has been made towards achieving autono‑ mous and scalable experimentation, largely due to the combination of automation and ML (Martin etal., 2023). While fully automating a laboratory is an ideal goal in some cases, it is not always the most practical or benecial approach. In certain situations, simpler and more modular methods may be more effective (Stephenson etal., 2023). For antibody engineering, given the current technological advancements, a modular yet fully integrated platform is in our view most appropriate.
A key application of automation is in improving the Design‑Build‑Test‑Learn (DBTL) cycle, commonly employed in drug discovery programmes. This cycle involves designing a series of candidates, constructing, and testing them, then using the gathered data to inform the design of the next series of candidates, often employing ML or sta‑ tistical modelling.
Effective implementation of DBTL frameworks in laboratories hinges on several factors. Consistent, reliable data are vital for ML applications, necessitating rigorous tracking and monitoring of reagents, samples, protocols, and business rules as well as control compounds that are used in programmes. Automation in data handling and metadata management is crucial to reduce errors commonly associated with manual data entry or processing.
20 Biopharmaceutical Informatics
In individual labs, various factors can impede clear and reproducible experimental results. These include inconsistent sample tracking, varying lab conditions, manual pro‑ cess errors (including simple copy‑paste errors of data), and non‑uniform data collection and storage practices. Additionally, differences in commercial and proprietary instru‑ ments and protocols can introduce signicant batch effects, leading to inconsistencies across different research groups and labs. Therefore, standardisation and automation across an organisation are essential.
The advent of advanced liquid handling technologies has revolutionised the move‑ ment of reagents and samples in automated experimental workows, boosting repro‑ ducibility and throughput. When paired with data management tools that consolidate experimental parameters and data, these technologies improve sample tracking and error detection, aiding decision‑making in drug discovery programmes. Technologies like next‑generation sequencing (NGS) or new deep screening of sequence‑sequence level interactions (Porebski etal., 2023), facilitate the generation of large datasets suit‑ able for training ML models. These datasets can be used to evaluate system perfor‑ mance and shape future experimental designs.
However, further improvements are needed in hardware and software systems to enhance connectivity between automated steps and equipment. Additionally, stan‑ dardising data infrastructures, exchange formats, protocol descriptions, ontologies, and data reporting will improve data interpretability and promote information sharing. In the following sections, we will examine the current status of automation systems perti‑ nent to biologics engineering.
Automation technologies in synthetic biology exhibit a wide range in format and degree, encompassing various levels of human interaction (Figure2.1). As automation
FIGURE2.1 Levels of automation in synthetic biology laboratories achieved through hard­ware and software automation technologies. Adapted from Stephenson etal. (2023) and Martin etal. (2023).