Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5908_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
16 Мб
Скачать
☆
2 • Digital Transformation 21
tools and machinery become more sophisticated, higher levels of automation are attain‑ able, consequently reducing human involvement. Common tools like pipettes are stan‑ dard, alongside instruments automating single tasks (e.g. thermocyclers, centrifuges, plate readers). These instruments are integral to developing automated workows, with the coordination of these diverse instruments and automated modules posing a signi‑ cant challenge.
Beyond unit process automation equipment (level 1 automation), there is a plethora of tools for automating multi‑step processes and entire experimental workows. Liquid handling robots and microuidic devices, for instance, have been designed in various formats to automate critical tasks and multiple steps in lengthy pipelines. Liquid han‑ dling robots enable more exibility and throughput, whereas microuidic devices enable precision and miniaturisation, allowing researchers to control liquids on a microscopic scale. Adaptable automation platforms are now being developed, notably by HighRes Biosolutions, offering highly versatile modular platforms with increased system exibil‑ ity, adaptive scheduling software, and improved protocol design. These technologies and an overview of their constraints and benets are summarised in Figure2.2. Combining robust physical automation with computational tools then enables closed‑loop iterative experimentation, enhancing discovery speed and productivity.
Recent trends emphasise systems approaches that combine hardware and soft‑ ware automation, especially crucial in cell‑based work for generating large, high‑ quality datasets. This is particularly evident in drug discovery’s lead optimisation stage, where complex assay readouts are essential and processing of data is more complex. The DBTL framework becomes increasingly automated at level 3 and above. Physical automation, in tandem with software tools, can then enhance efciency, foster reproduc‑ ibility through standardisation, and minimise experimental error and variability (e.g. by using data normalisation). These improvements are essential for advancing the DBTL cycle. To effectively screen multiple designs concurrently and expedite DBTL cycles, comprehensive automation is indispensable at each stage of the DBTL process.
Design tools in synthetic biology have evolved to include functionalities for DNA part design and assembly, utilising established methods like Golden Gate assembly, Gibson assembly, SLIC, and digestion‑ligation methods (Chao etal., 2015; Chao etal.,
2017). Recent years have seen the integration of computational tools and ML, with innovations like generative design (Ingraham etal., 2023) and protein language models (Fenoy etal., 2022; Alley etal., 2019; Brandes etal., 2022; Rives etal., 2021; Wu etal., 2021; Choi, 2022; Dounas etal., 2024; You etal., 2021; Heinzinger etal., 2019; Lu etal.,
2020) augmenting traditional computational biology approaches.
In the build phase, DNA assembly stands out as a particularly time‑intensive and error‑prone process, making it an ideal candidate for automation (Walsh etal., 2019; Storch etal., 2020; Linshiz etal., 2016; Kang etal., 2022). Transitioning manual assem‑ bly to high‑throughput automated systems is complex, but using standardised parts and modular designs can facilitate the large‑scale prototyping of combinatorial DNA librar‑ ies (Casini etal., 2015) as well as faster assembly and screening of combinatorial design spaces. Methods like Golden Gate assembly, which are executed as one‑pot reactions with limited steps, such as liquid handling, incubation, and thermocycling, can be auto‑ mated to signicantly cut down on time and manual errors.
FIGURE2.2 Overview of automation technologies and characteristics that need to be considered when incorporating automation. Adapted from Stephenson etal. (2023).
22 Biopharmaceutical Informatics
2 • Digital Transformation 23
Steps like cell transformation, transfection, or transduction are also challenging to automate due to specic experimental conditions and high labour intensity, but many successful implementations exist (Si etal., 2017).
The testing phase presents another automation challenge. Protocols vary widely, with some steps being more conducive to automation than others. Cell‑based work, in particular, introduces unique challenges for automation, and current solutions often struggle to upscale or simply shift bottlenecks along the pipeline (Doulgkeroglou etal.,
2020). Killing assays, for example, using the IncuCyte® by Sartorius, are notoriously hard to automate. Recent systems like the Tecan Freedom EVO or Tecan Cellerity allow for better integration of asks (so‑called RoboFlask) into the process and thereby sim‑ plify cell‑based work. Automated analysis involving techniques like chromatography and mass spectrometry also presents distinct challenges. The complexity of the data and inefcient resolution renders it difcult for software to accurately make an assessment, such as correctly distinguishing between peaks and often requires human intervention (Cappadona etal., 2012). However, the integration of liquid handling technologies with component instruments that automate critical or labour‑intensive steps, combined with advanced software, can signicantly alleviate the workload on researchers and enhance experimental results.
A notable application of this approach is the modular DBTL platform developed by LabGenius, which focusses on the design, synthesis, and testing of multi‑specic antibodies using cell‑based assays. LabGenius has developed a exible robotic liquid handling platform that can perform disease‑relevant, functional, cell‑based assays to produce machine learning‑grade data at high throughput, minimising variability and noise and rendering it ideally suited for active learning. Computational models enable the design and learning steps, while the automated liquid handling system covers the automated build and test steps. By combining high‑throughput experimentation with modelling, LabGenius can today evaluate the performance of thousands of antibody designs in disease‑relevant cell‑based assays in just a few weeks. This is possible through an active learning‑based approach that generates data in an optimal way to train ML models (LabGenius, 2023).
While automation has made signicant strides in elds like antibody engineering, manual intervention is often still necessary for monitoring and connecting different ele‑ ments of the pipeline.
The connectivity among each step in the discovery process and integration of com‑ puting throughout makes it possible to capture a wealth of information: explicit repre‑ sentations of goals, hypotheses, experimental results, and ndings, as well as metadata that can be critical to tracking why and how experiments were carried out in order to understand and repeat them or diagnose errors (e.g. environmental conditions, instru‑ ment settings, protocols, runtime logs). For this reason, many systems that are ideally capable of level 4 automation (e.g. King etal., 2009; Williams etal., 2015; Sparkes etal.,
2010), often function as level 3 systems due to the regular need for human intervention in hardware and software issues. Nevertheless, they highlight the potential of combined human‑robot efforts in accelerating research through strategic automation.
It is also important to recognise that fully autonomous systems often require a high degree of integration, which often coincides with less exibility or modularity of the
24 Biopharmaceutical Informatics
constituent parts. Such systems are not possible in many organisations due to the need for equipment to work across many different projects with a multitude of workows.
In scientic research, pipetting assistant robots offer a basic level of automation, emulating manual pipettes with a small footprint and lower cost than more complex plat‑ forms. Their typical applications include reagent dispensing, serial dilutions, and basic sample preparations. While having lower throughput due to limited channels and deck size, they alleviate manual pipetting errors and save time. Manufacturers like Integra, Hudson Robotics, and Gilson offer various models, from dedicated microplate dispens‑ ers to versatile workstations that can integrate with other devices for expanded functions.
Some manufacturers provide pre‑congured workstations for specic research areas such as NGS, polymerase chain reaction (PCR), and proteomics. These are larger and costlier, starting from $100 to 200k, but offer high throughput, exibility, and reli‑ ability through advanced hardware and software. Companies like PerkinElmer, Agilent, and Beckman Coulter offer various models tailored to specic applications.
Contrastingly, microuidic devices, primarily developed in‑house by research labs, face a commercialisation gap, with notable exceptions in sequencing and analysis sample preparation. Commercial systems like the Agilent 2100 Bioanalyser and the Beckman Coulter Echo 650 series have succeeded due to their innovative approaches in reducing sample consumption, analysis time, and contamination. However, challenges remain in the commercial viability and reliability of microuidic devices, as seen in the discontinuation of certain products. Efforts continue in commercialising microuidics, with emerging low‑cost or open‑source platforms.
In synthetic biology, as automation and digitalisation advance, the need for stan‑ dard data management and communication protocols is crucial for creating an inte‑ grated laboratory environment. While liquid handling robots and automated analytical units are widespread, they typically operate as standalone units with proprietary stan‑ dards, command sets, device driver software, and data formats. The result of this is poor connectivity among devices, data silos, and difculty in integrating multiple devices to form continuous, exible automation pipelines. In automated laboratories, multiple devices are set up for steps, such as sample preparation and analysis, and connected to software systems that orchestrate the ow of samples through experimental pipelines. Piecing together diverse instruments needed for a given workow without standardised hardware interfaces complicates the connection of equipment to a centralised control system. Variations in software among devices make integration very challenging, where even devices with similar functionality have vastly different command sets, behav‑ iours, and capabilities for handling device states and errors. Furthermore, data collected from various instruments in a workow are typically stored in proprietary formats that complicate visualisation, analysis, and sharing. Arguably, having a manufacturer‑ independent interface denition that applies to all devices and stores results in a com‑ mon format would greatly simplify the process of setting up exible and modular laboratory automation systems that can then be sustainably tailored to application and researcher‑specic needs. There is hence a big need for standardisation of data formats across analytical devices and interfaces (communication standards)–both are needed for true end‑to‑end integration. Such standards (e.g. as introduced by the Standardisation in Laboratory Automation consortium or via the Analytical Information Markup
2 • Digital Transformation 25
Language) would be able to facilitate easier data storage, visualisation, and sharing in a unied format and can enhance reproducibility, traceability, and remote monitoring of experiments. They can hence enable the automatic upload to global repositories via cloud computing, simplifying data management, enabling remote process monitoring, and making it possible to efciently share experimental results with all stakeholders.
We note that large language models (LLMs) like Generative Pre‑trained Transformer3 (GPT‑3) and ChatGPT are emerging as promising tools for laboratory automation. These models, trained on extensive datasets, show adaptability across vari‑ ous applications, notably in intuitive robot control via natural language inputs. However, it is too early to tell how they will actually impact the state of laboratory automation.

2.4.2 Cloud Computing

Cloud computing, a vital component in drug discovery, involves the on‑demand delivery of congurable computing resources over a network. It relies on virtualisation technol‑ ogy, where a hypervisor software abstracts and shares hardware resources as virtual machines (VMs). This environment allows for rapid scaling of resources with minimal interaction with the service provider (Spjuth etal., 2021).
The growing reliance on ML and statistical modelling underscores the need for effective data management solutions that ensure rapid and efcient data access. The data collection phase in the ML life cycle involves various tasks, including conducting experiments to generate data, extracting existing data from databases, and selecting spe‑ cic data for modelling. Cloud computing plays a crucial role in this context, offering scalable and on‑demand infrastructure for storage, databases, and middleware without initial costs. Using cloud services for data hosting provides easy and quick global access to data, particularly when modelling is conducted within cloud environments (Dalpé and Joly, 2014). This accessibility not only streamlines the modelling process but also enhances collaborative efforts and simultaneous access.
Before the advent of cloud computing, drug discovery organisations requiring signi‑ cant computational power were constrained to on‑premise servers, entailing substantial overhead costs and maintenance responsibilities. This setup posed insurmountable chal‑ lenges for smaller entities and formidable obstacles even for larger corporations, render‑ ing it an impractical endeavour for most. The emergence of cloud computing, particularly through infrastructure or platform as a service, has revolutionised this landscape, democra‑ tising access to computing resources previously unattainable to biotechnology companies.
Cloud based platforms such as Microsoft’s Azure or Amazon’s AWS for example, have meanwhile entered majority of the corporate pharmaceutical world with a more or less complete set of governance, storage and computing tools on demand. In the absence of such solutions, companies like LabGenius would have required large amounts of additional human resources and may have deemed the endeavour nancially unviable. Many small teams and companies nd themselves in similar circumstances, where the cost of leveraging ML technologies would have otherwise been prohibitive, barring those organisations for which such technologies constitute the core business operations. Cloud computing has signicantly lowered this barrier, enabling even the smallest of
26 Biopharmaceutical Informatics
teams to engage in exploratory ML initiatives by leveraging appropriate infrastructure and platforms to assess the viability of integrating ML into their workows.
From a data acquisition and analysis standpoint, cloud computing has also facilitated the adoption of distributed, asynchronous, simultaneous, and secure data handling prac‑ tices. Notably, Laboratory Information Management Systems (LIMS) and Electronic Laboratory Notebooks (ELN) now benet from automated backups, universal acces‑ sibility across connected devices, and instantaneous sharing and analysis capabilities across disparate locations and users. The cloud environment affords streamlined system management, allowing for immediate implementation of software modications with unprecedented agility and efciency. Increasingly, alongside traditional off‑the‑shelf LIMS, custom bespoke applications are being developed and integrated due to their low creation and deployment barriers. These tailored applications can be ne‑tuned to meet the specic requirements of either a company as a whole or the distinct needs of individual scientists. This trend highlights a shift towards greater efciency and adapt‑ ability in scientic research settings.
Cloud computing has three primary service levels: infrastructure as a service (IaaS), platform as a service (PaaS), and software as a service (SaaS), each providing an increasing level of abstraction from the hardware. There are four main cloud deploy‑ ment models: public cloud (accessible to the general public), private cloud (deployed on‑ premise for a single organisation), community cloud (shared by multiple organisations with similar concerns), and hybrid cloud (a combination of any two models, often public and private, with a software layer for seamless application portability). OpenStack, a modular open‑source software stack, is widely used for operating cloud environments in academia and by local cloud providers (www.openstack.org/).
Major public cloud providers offer extensive catalogues of virtual machine images (VMIs) with pre‑installed software environments, enhancing tool distribution in drug discovery and improving reproducibility.
Containerisation offers a lightweight method for deploying VMs by packaging an application’s code, runtime, dependencies, and settings into a portable and repeatable unit. It allows applications to run consistently across various platforms and is more resource‑efcient than VMs. Kubernetes (K8s), developed by Google, is the leading container orchestration platform, enabling easy management of containerised applica‑ tions. It’s commonly used in private cloud platforms and major cloud providers for scal‑ able access to computing resources.
While software as a service requires minimal IT skills, efciently working with infrastructure as a service and platform as a service necessitates basic knowledge of Linux or programming languages like Python.
Private clouds and on‑premises Kubernetes (K8s) clusters serve as valuable com‑ plements in this ecosystem (Emami etal., 2019). They provide scientists with the ability to swiftly repurpose existing hardware resources to meet the diverse application needs at different stages of the ML life cycle. This approach enables a more efcient and cost‑effective way to manage computational resources, catering to the dynamic and sometimes unpredictable requirements of ML projects. The integration of private clouds and K8s clusters into the ML workow thus represents a strategic approach to balancing resource demands while maintaining operational exibility and control.
2 • Digital Transformation 27
Using infrastructure as a service (IaaS) resources from cloud providers presents a viable alternative, offering access to extensive resources on a pay‑per‑minute basis, eliminating upfront and maintenance hardware costs. However, the costs for advanced cloud infrastructure, such as graphics processing units (GPUs) or large VMs with ample random‑access memory (RAM), should not be underestimated. Teams with consistently high GPU demands are increasingly establishing private Kubernetes (K8s) infrastruc‑ tures to support GPU‑powered, containerised workows.
2.4.3 Machine Learning and MLOps to Support
the Machine Learning Life Cycle
The integration of machine learning (ML) in drug discovery involves comprehensive management of data and model life cycles. This process encompasses stages like data gathering, pre‑processing, training, validating, and deploying models, along with pre‑ dictions and inference. ML Operation (MLOps), a blend of practices and software, sup‑ ports these stages, enhancing reproducibility and robustness in iterative drug research. Since machine learning in drug discovery is covered amply in other resources (e.g. Vijayan etal., 2022; Sarkar et al., 2023; Wossnig et al., 2024), we mainly focus on MLOps in the following section.
MLOps includes a range of functions primarily concerned with facilitating deploy‑ ment, audit logging, and reproducibility of ML systems, vis à vis version control and automation for code, data, pipelines, and models. In drug discovery, one frequently encounters changes in the experimental data that suddenly changes model performance. Often, there is also a lack of metadata which hinders the diagnosis of errors and trouble‑ shooting. MLOps, specically model and data versioning, can enable the rapid diagno‑ sis of errors, reverting to previous models, and generally better data‑model‑prediction lineage that ensures that the right method is applied to the right problem at the right time.
Cloud computing offers scalable, resilient computing resources crucial for MLOps and ML operations. Containerisation, coupled with platforms like Kubernetes and scien‑ tic workows, enables sturdy, reproducible ML analysis (e.g. consistent use of environ‑ ment and library versions), and higher speed in drug discovery. Additionally, emerging federated methods promise to facilitate collaborative ML across multiple organisations, leveraging both private and cloud‑based infrastructure.
Most drug discovery projects, specically in the lead optimisation stage follow the iterative DBTL process (Figure2.3). ML is used to enhance decision‑making, particu‑ larly in the Learn phase, using both project‑specic and global models. MLOps emerges as a critical framework, integrating data engineering, science, and operations to manage the ML life cycle in production. It aims to streamline processes from data preparation to model deployment, ensuring up‑to‑date models are accessible. Cloud services support MLOps stages, with tools like Vertex AI, KubeFlow, STACKn, and H2O.ai aiding in public and private infrastructures. Continuous ML modelling requires the integration of data handling, pre‑processing, quality control, and validation into a reproducible pipe‑ line for optimal decision‑making in drug discovery.
28 Biopharmaceutical Informatics
FIGURE2.3 The use of ML models in the DBTL cycle in drug discovery. Within a specic project, models are trained using data from the Test phase and used to make predictions that support the Learn phase, such as running an assay to optimise a molecule for binding to a particular target. Global models can then be used for predictions, such as predicting binding to off-targets. Global models are typically trained using data collected from many projects or merged with publicly available data. By contrast, project models are developed for a particular objective within a drug discovery project. Project models are typically manu­ally trained, and may need updating. Adapted from Spjuth etal. (2021).
FIGURE2.4 The machine learning life cycle. Diagram adapted from Spjuth etal. (2021).
The ML life cycle (Figure2.4) in drug discovery involves multiple steps with spe‑ cic infrastructure and software needs. Cloud IaaS reduces the need for on‑premises computing infrastructure, offering scalable, on‑demand resources. Data collection is crucial, requiring efcient data management solutions. Pre‑processing involves han‑ dling artefacts and quality control, with cloud computing facilitating transparent, reproducible workows. Model training and validation are resource‑intensive, whereas cloud IaaS offers a cost‑effective alternative to on‑site infrastructure. Prominent tools like TensorFlow and PyTorch are readily available in cloud environments and have sig‑ nicantly simplied the process of training models on complex infrastructure such as GPUs–thereby reducing the barrier of entry. Kubernetes‑based scientic workows, exemplied by Recursion Pharma and BenevolentAI, orchestrate model development pipelines efciently. For example, AlphaFold (Jumper etal., 2021) was quickly inte‑ grated by a large number of companies due to its availability in a container (c.f. the AlphaFold GitHub repository, https://www.github.com/google‑deepmind/alphafold).
The initial step is data collection. MLOps should enable careful versioning of data (e.g. using DataVersion Control(DVC), Pachyderm, or Neptune). This is essential to later on trace the lineage from a dataframe (a specic version of the data) to a model and the model predictions (results). Without proper data versioning and clear pointers to the most recent data version, prediction errors or delays can easily occur.
2 • Digital Transformation 29
After selecting and assembling data in drug discovery, a crucial step involves pre‑processing, which addresses issues like duplicate data, missing values, normalisa‑ tion, augmentation, and quality control. Cloud computing signicantly aids this process by simplifying the construction and execution of pipelines or workows. This enhance‑ ment contributes to increased transparency, reproducibility, and robustness in data handling.
Many workow systems in the cloud allow for declarative specications of anal‑ ysis pipelines, with capabilities to execute these workows on both public and private clouds. Nextow, a popular workow engine in life science research, can directly execute workows on infrastructure as a service (IaaS) resources and Kubernetes clusters. Additionally, Argo and Pachyderm are two Kubernetes‑native workow systems.
Notebooks deployed on cloud resources are commonly used for specifying pre‑pro‑ cessing steps. This approach not only facilitates pre‑processing execution but also sup‑ ports visual interpretations, enhancing understanding and analysis. Implementing a complete workow of all pre‑processing steps in the cloud ensures that the pre‑process‑ ing is reproducible, portable, and scalable. On the other hand, notebooks allow people to write lower‑quality and potentially unsafe code. Here again, MLOps practices can be applied to mitigate risks and follow good standards.
Moreover, numerous platform as a service (PaaS) services focussed on big data pre‑processing and analysis are available as public cloud services. An example is Databricks, which is based on Apache Spark, offering advanced capabilities for han‑ dling and analysing large datasets in cloud environments.
After dataset construction, the focus shifts to model development, including train‑ ing, validation, optimisation, and model hyperparameter tuning. This phase can be resource‑intensive and time‑consuming. Cloud computing offers signicant advantages here, providing exibility and scalability than on‑site workstations. These workstations, while powerful, are not scalable, require substantial upfront investment, and necessitate ongoing maintenance.
For AI modelling in drug discovery, software stacks like TensorFlow, PyTorch, and SciKit‑Learn (the rst two are primarily for deep learning) are prominent. These tools are readily available in cloud environments, with virtual machine images (VMIs) and container images facilitating rapid setup for single‑node infrastructures. Kubernetes is widely utilised for scaling across multiple nodes, along with specialised AI frameworks for Kubernetes.
Scientic workows, orchestrated using engines like SciPipe and AirFlow, are employed to manage complex model development processes, including tasks like nested cross‑validations. These technologies and practices illustrate the evolving landscape of AI modelling in drug discovery, where cloud computing and containerisation play criti‑ cal roles in enhancing efciency and scalability.
Model serving and inference in the ML modelling life cycle involve making AI models accessible to end‑users in a production‑grade environment with appropri‑ ate governance. This step, often the most technically challenging, requires not only model validation (Smiatek et. al., 2021) and versioning but also critical features like privacy, access control, auditability, logging, monitoring, and a resilient infrastructure capable of recovering from failures.
30 Biopharmaceutical Informatics
Modern software engineering practices, particularly DevOps, which merges devel‑ opment (Dev) and operations (Ops), offer valuable insights for this stage. Cloud com‑ puting and infrastructure‑as‑code have revolutionised DevOps, providing a exible, scalable environment suited for production‑grade solutions. Kubernetes, popular in DevOps, supports these requirements, offering features essential for serving ML mod‑ els, including resilience and scalability.
Public clouds typically provide services for serving and publishing ML models, and numerous cloud‑native open‑source frameworks have emerged, leveraging Kubernetes. These include TensorFlow Serving, Seldon, Kubeow, and STACKn, or ones directly hosted by the cloud provider such as Google Kubernetes Engine (Google Cloud) or Amazon Elastic Kubernetes Service (AWS) which facilitate model management and serving.
An important aspect of model serving is making models available via an Application Programming Interface (API), allowing for integration with other software components and enabling cloud‑based inference, essentially predictions as a service (SaaS). In drug discovery, platforms like OpenRiskNet and DeepCell Kiosk exemplify this approach. OpenRiskNet, built on OpenShift, a Kubernetes distribution, allows for publishing AI model services and adds layers of discoverability and interoperability. DeepCell Kiosk, designed for microscopy image analysis using TensorFlow, also utilises Kubernetes for scalable deployment and inference. These examples demonstrate the evolving integra‑ tion of cloud computing and AI in drug discovery, highlighting the critical role of model serving and inference in the ML life cycle.
2.4.4 Company Culture That Drives
a Data‑Driven Approach
Senior management sponsorship and widespread cross‑functional support of a data engineering culture, as opposed to just science, is integral to digital transformation and a data‑driven approach (Saldanha, 2019). Take the example of MLOps, integrating data engineering, science, and operations to manage the ML life cycle in production. Teams need to work together with a shared understanding of the framework. However, chal‑ lenges often remain with respect to a company‑wide holistic adoption of digitisation. Reinhardt etal. (2020) note that knowledge of Industry 4.0 is concentrated at the more senior levels of organisations and highlights a disconnect based on seniority. To over‑ come this, data scientists should be hired early, and be given a senior position where they can shape the company‑wide mindset from the get‑go. Senior team members should articulate a compelling vision for becoming a data‑driven organisation. As an example, GlaxoSmithKline has established the role of Senior Vice President Global Head of Articial Intelligence and Machine Learning; Relay Therapeutics has established the roles of a Chief Data Ofcer and Senior Vice President, Articial Intelligence; and Recursion the role of Chief Technology Ofcer–to name just a few. Teams should be trained and given the opportunity to learn from one another–for example, by encourag‑ ing biologists and data scientists to engage in conversation and to participate in seminars together.