Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5525_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Editor biographies
- •Yi Wang
- •X. Sharon Qi
- •List of contributors
- •1.2.3 Feature engineering and representation
- •1.2.4 Linear separability
- •1.2.5 Classical models
- •1.1 A brief introduction to AI
- •1.2 Machine learning basics
- •1.2.1 Learning paradigms
- •1.3 Artificial neural networks
- •1.3.1 Feed-forward neural networks
- •1.3.2 Recurrent neural networks
- •1.3.3 Convolutional neural networks
- •1.3.4 Attention
- •1.3.5 Training neural networks
- •1.3.6 Applications and use cases of deep learning
- •1.4 Model training and evaluation
- •1.4.1 Hyperparameters
- •1.4.2 Data split
- •1.4.3 Evaluation metrics
- •1.5 Generative models
- •1.5.1 Generative adversarial networks
- •1.5.2 Diffusion models
- •1.5.3 Applications and use cases
- •1.6 Ethical consideration and bias
- •1.6.1 Transparency and explainability
- •1.6.2 Bias and fairness
- •1.6.3 Data privacy violation
- •1.6.4 Risk and misuse
- •1.7 Summary
- •References
- •2.1 Introduction
- •2.1.2 Staff roles in radiation therapy
- •2.2 Overview of AI in radiation therapy
- •2.2.1 Patient evaluation and dose prescription
- •2.2.2 Treatment simulation
- •2.2.3 Contouring
- •2.2.4 Treatment planning
- •2.2.5 Quality assurance
- •2.2.6 Treatment delivery
- •2.2.7 Response assessment and toxicity management
- •2.3 Summary
- •3.1 Introduction
- •3.1.1 Introduction of clinical decision making and AI
- •3.1.2 The role of AI in clinical decision making
- •3.2 AI algorithms for clinical decision making
- •3.2.2 Radiomics
- •3.2.3 Data integration by AI
- •3.2.4 Interpretability of AI models
- •3.3 Application of AI in clinical decision making
- •3.3.1 Diagnosis and disease phenotyping
- •3.3.2 Personalized treatment
- •3.3.3 Treatment outcome and prognosis prediction
- •3.4 Challenges and future directions of AI in clinical decision making
- •3.4.1 Challenges and concerns
- •3.4.2 Future directions
- •3.5 Summary
- •References
- •4.1 Introduction
- •4.2 Imaging for treatment planning
- •4.2.1 CT simulation
- •4.2.2 4D-CT
- •4.2.3 PET/CT
- •4.2.4 MRI
- •4.3 Imaging for treatment guidance
- •4.3.1 Portal imaging
- •4.3.2 CBCT
- •4.3.3 CT-on-rail and CT-linac
- •4.3.4 MR-linac
- •4.3.5 PET-linac
- •4.4 Imaging for motion management
- •4.4.1 ExacTrac
- •4.4.2 Varian triggered imaging
- •4.4.3 4D-CBCT
- •4.4.4 Cine MRI
- •4.4.5 4D-MRI
- •4.4.6 Surface imaging
- •4.5 Imaging for treatment assessment
- •4.5.1 Contrasted CT
- •4.5.2 PET/CT
- •4.5.3 Functional MRI
- •4.6 Summary
- •5.1 Introduction to big data in radiation oncology
- •5.1.1 Overview of big data
- •5.1.2 Sources of big data in radiation oncology
- •5.1.3 Big data and AI in radiation oncology
- •5.2 Big data lifecycle in radiation oncology
- •5.2.1 Data aggregation and storage
- •5.2.2 Data sharing and security
- •5.4 The application of big data in radiation oncology
- •5.4.1 Medical image segmentation
- •5.4.2 Automatic treatment planning
- •5.4.3 Treatment response prediction
- •5.4.4 Quality assurance and patient safety
- •5.4.5 Clinical decision support
- •5.2.3 Data visualization
- •5.2.4 Knowledge creation and implementation
- •5.2.5 Data archiving and deletion
- •5.3 Big data analytics with AI
- •5.3.1 Data processing and integration
- •5.3.2 AI modeling
- •5.5 Challenges and future perspectives
- •5.6 Summary
- •Reference
- •6.1 The road to ART
- •6.1.1 3D conformal radiotherapy (3DCRT)
- •6.1.2 Intensity modulated radiotherapy (IMRT)
- •6.1.3 Image-guided radiotherapy (IGRT)
- •6.1.4 Adaptive radiotherapy (ART)
- •6.2 ART workflow and implementation
- •6.2.2 Current practice
- •6.2.3 Clinical impact
- •6.3 Considerations for implementing online ART
- •6.3.1 Time as a limiting factor
- •6.3.2 Implications for fast and reliable re-planning
- •6.3.4 Clinical considerations
- •6.4 Summary
- •7.1 Components of ART workflow
- •7.1.1 Simulation
- •7.1.2 Pre-planning
- •7.1.3 Online imaging and daily re-planning
- •7.1.4 Quality assurance
- •7.2 AI-driven ART
- •7.2.1 Simulation
- •7.2.2 Pre-planning
- •7.2.3 AI for delivery
- •7.3 Outlook and future directions
- •7.3.1 Real-time ART with AI
- •7.3.2 Dose escalation and functional adaption with AI
- •7.4 Summary
- •References
- •8.1 Introduction
- •8.2 Synthetic CT: deep learning methods
- •8.2.1 Conventional methods
- •8.2.2 U-Net
- •8.2.3 Generative adversarial networks
- •8.2.4 Denoising diffusion probabilistic model
- •8.3 Synthetic CT from CBCT
- •8.3.1 Noise and artifact reduction
- •8.3.2 Online dose calculation
- •8.3.3 Online image segmentation
- •8.4 Synthetic CT from MRI
- •8.4.1 Synthetic image accuracy
- •8.4.2 Dose calculation in MR-only radiation therapy
- •8.4.3 PET attenuation correction
- •8.4.4 Image registration
- •8.5 Discussion and outlook
- •8.6 Summary
- •References
- •9.1 AI-based image registration and segmentation for ART
- •9.1.1 Adaptive radiation therapy
- •9.2 Artificial intelligence
- •9.2.1 What is machine learning?
- •9.2.2 What is deep learning?
- •9.3 Deep learning: the basic components
- •9.3.1 Convolutional neural networks: looking at the picture
- •9.3.2 Pooling layers: keeping what matters most
- •9.3.3 Fully connected (dense) layers: bringing it all together
- •9.3.4 Activations
- •9.3.5 Loss: driving the model
- •9.3.6 Auto-encoders: remove the noise
- •9.3.7 Supervised versus unsupervised learning
- •9.3.8 Pre-trained convolutional neural networks
- •9.4 Image registration: bringing two images together
- •9.4.1 Registration similarity metrics
- •9.4.2 Types of registrations
- •9.5 AI-based image registration
- •9.5.1 Supervised learning
- •9.5.2 Unsupervised learning
- •9.5.3 Registration in ART
- •9.5.4 Commonalities in architectures
- •9.6 Image segmentation
- •9.6.1 Introduction: coloring by the numbers
- •9.6.2 Segmentation networks
- •9.6.3 Best practices
- •9.7 Summary
- •10.1 Introduction
- •10.1.1 Overview of chapter content
- •10.2 The landscape of AI-assisted dose prediction
- •10.2.1 Traditional machine learning for dose prediction
- •10.2.2 Deep learning-based dose prediction
- •10.2.3 Challenges in AI-assisted dose prediction
- •10.3 Re-planning workflows powered by AI
- •10.3.1 Deep learning for re-planning pipelines
- •10.4 Future directions of AI-assisted dose prediction and re-planning
- •10.5 Summary
- •11.1 Introduction
- •11.2.1 Imaging-based motion monitoring
- •11.2.2 Delivery system actions
- •11.2.3 Challenges for real-time ART implementation
- •11.3 AI in real-time ART workflows
- •11.3.1 Improving intrafraction motion monitoring through AI
- •11.3.2 Mitigating system latency through AI
- •11.4 AI for ART delivery: future directions
- •11.4.1 Management of non-respiratory motion
- •11.4.2 Training AI models with small or unpaired datasets
- •11.4.4 Biology-guided ART delivery
- •11.5 Summary
- •References
- •12.1 Introduction
- •12.2 Patient QA
- •12.2.1 Pre-planning QA
- •12.2.2 Pre-treatment plan QA
- •12.2.3 On-treatment QA
- •12.3 Treatment delivery systems and instruments
- •12.3.1 Machine commissioning
- •12.3.2 Machine QA
- •12.3.3 Dosimetry tool QA
- •12.4 Summary
- •References
- •13.1 Data resources for response modeling in radiotherapy
- •13.1.1 Clinical data
- •13.1.2 Imaging (radiomics)
- •13.1.3 Treatment planning (dosiomics)
- •13.1.4 Multiomics
- •13.2 Radiotherapy treatment outcome modeling
- •13.2.1 TCP/NTCP in radiotherapy
- •13.2.2 Clinical outcomes versus PROs
- •13.2.3 Machine learning response prediction
- •13.2.4 Explainability of ML response models
- •13.2.5 Sample use cases
- •13.3 AI response-based adaptive radiotherapy
- •13.3.1 Requirements and challenges
- •13.3.2 Prediction versus treatment optimization
- •13.3.3 Sample use cases
- •13.4 Challenges and recommendations
- •13.5 Summary
- •Acknowledgments
- •References
- •14.1 Overview of challenges in AI-driven ART
- •14.2 Data challenges
- •14.2.1 Data availability
- •14.2.2 Data quality
- •14.2.3 Data privacy
- •14.3 Technical challenges
- •14.3.2 Model robustness and generalizability
- •14.3.3 Model explainability and interpretability
- •14.4 Challenges associated with online and real-time workflows
- •14.4.1 Image quality
- •14.4.2 Dose calculation
- •14.4.3 Real-time ART
- •14.5 Operational challenges
- •14.5.1 Clinical validation
- •14.5.3 Staff training
- •14.5.4 User experiences
- •14.5.5 Quality management program
- •14.5.6 Financial challenges
- •14.6 Ethical, regulatory, and legal challenges
- •14.6.1 Ethical issues
- •14.6.2 Regulatory and legal issues
- •14.7 Summary
- •References
- •15.1 Clinical considerations for CT-based offline ART
- •15.1.1 Patient and site selection
- •15.1.2 Re-simulation
- •15.1.3 Re-planning
- •15.1.4 Plan summation and evaluation
- •15.1.6 Limitations and future directions
- •15.2 Clinical considerations for CBCT/CT-based online ART
- •15.2.2 Patient and site selection
- •15.2.3 Simulation
- •15.2.4 Pre-planning review
- •15.2.5 Reference planning
- •15.2.9 Limitations and future directions
- •15.3 Summary
- •References
- •16.1 Introduction
- •16.2 Overview of MRI-guided ART systems
- •16.3 MRI-guided ART workflow
- •16.4 AI applications for MRI-guided ART
- •16.4.1 Synthetic CT generation
- •References
- •16.4.2 Auto-segmentation
- •16.4.3 Image registration
- •16.4.4 Others
- •16.4.5 Future AI development and implementation
- •16.5 Summary
- •17.1 Functional PET-guided ART
- •17.1.1 PET-based functional imaging overview
- •17.1.2 From anatomy to function: the power of PET in radiation therapy
- •17.1.5 Conclusions and future prospects
- •17.2 Functional MRI-guided ART
- •17.2.1 From anatomy to function: the power of functional MRI in radiation therapy
- •17.2.4 Conclusion and future prospects
- •17.3 Summary
- •References
- •18.1 Proton ART
- •18.1.1 Clinical context and necessity
- •18.1.2 Patient populations
- •18.1.4 Rationale for AI in proton ART
- •18.2 AI in proton ART
- •18.2.1 Imaging
- •18.2.2 Deformable and rigid registration
- •18.2.3 Contour propagation
- •18.2.4 Dose calculations
- •18.2.5 Plan optimization
- •18.2.6 Other developments
- •18.3 Implementation of adaptive proton therapy
- •18.4 Summary
- •References
- •19.1 Designing clinical trials with AI
- •19.1.1 The essential role of clinical trials
- •19.1.2 Trial protocols and methodologies
- •19.1.3 AI-driven clinical trial design and execution
- •19.1.4 Incorporation of digital twins (DTs) in clinical trials
- •19.2 Implementation of AI in ongoing clinical trials
- •19.2.1 Integration with existing clinical trial frameworks
- •19.2.2 Quality assurance, compliance, and standardization
- •19.3 Case studies of AI in adaptive radiotherapy trials
- •19.3.1 Overview of guidance for advanced radiotherapy in clinical trials
- •19.3.2 AI in the radiotherapy clinical trial quality assurance processes
- •19.4 Ethical and regulatory considerations
- •19.4.1 Patient consent and data privacy
- •19.4.2 Bias, fairness, and transparency
- •19.4.3 Regulatory guidelines and compliance
- •19.5 Future directions and challenges
- •19.5.1 Emerging technologies and techniques
- •19.5.2 Alternative strategies
- •19.6 Conclusion
- •19.7 Summary
- •References
- •20.1 Risk management
- •20.1.1 Prospective risk assessments
- •20.1.2 Root cause analysis

IOP Publishing
Artificial Intelligence in Adaptive Radiation Therapy
Yi Wang and X. Sharon Qi
Chapter 1
Fundamentals of artificial intelligence
Parsa Bagherzadeh, Laya Rafiee Sevyeri, Yujing Zou and Shirin Abbasinejad Enger
Artificial Intelligence (AI) is transforming healthcare, with adaptive radiotherapy
standing out as a key area where AI is reshaping cancer treatment. By leveraging
real-time data, AI enables more personalized and precise treatment plans, improving
patient outcomes. This chapter introduces the fundamentals of AI, focusing on
machine learning (ML) and deep learning (DL), which drive innovations in radiotherapy. It covers the evolution of AI from rule-based systems to advanced learning
algorithms, exploring their applications in clinical settings. It will discuss the ethical
considerations surrounding AI in healthcare, including transparency, bias, and data
privacy, setting the stage for a deeper dive into AI’s role in adaptive radiotherapy.
1.1 A brief introduction to AI
Artificial intelligence (AI), a subfield of computer science, focuses on developing
systems capable of performing tasks that typically require human intelligence. Such
tasks include problem-solving, learning, perception, and language understanding.
AI systems leverage algorithms, statistical models, and machine learning techniques
to analyse and interpret data, enabling them to make decisions, recognize patterns,
and improve performance over time. AI applications range from virtual assistants
and recommendation systems to complex tasks in various fields. In particular, AI
plays a significant role in healthcare, transforming medical decision-making and
treatment strategies by changing the existing landscape. In the context of adaptive
radiotherapy, an intricate interplay of AI methodologies is reshaping how we
approach cancer treatments. The term ‘artificial intelligence’ was coined at the
Dartmouth Conference in 1956, where researchers aimed to explore how machines
could simulate human intelligence. In the 1960s and 1970s, AI research mainly
focused on logic-based symbolic reasoning, which involves representing knowledge
using symbols and rules, allowing machines to manipulate these symbols to derive
logical conclusions and make decisions [1]. During this period, AI found its
doi:10.1088/978-0-7503-6119-4ch1 1-1 ª IOP Publishing Ltd 2025. All rights,
including for text and data mining (TDM), artificial intelligence (AI) training, and similar technologies, are reserved.

Artificial Intelligence in Adaptive Radiation Therapy
applications in medicine through systems such as MYCIN [2], a system for
diagnosing bacterial infections and recommending antibiotic treatments.
In parallel with symbolic reasoning, AI research also focused on optimization and
handling uncertainty, both important for developing intelligent systems. Fuzzy
logic, for instance, is a mathematical framework that deals with uncertainty as well
as imprecision, allowing for the representation of partial truths [3]. This flexibility
makes it particularly suitable for systems where variables may have ambiguous or
overlapping boundaries. Genetic algorithms are examples of evolutionary optimization techniques inspired by principles of natural selection and genetics [4]. They
employ a population of potential solutions, subjecting them to genetic operations
such as crossover, mutation, and selection to evolve toward an optimal solution over
successive generations.
The development of reasoning systems continued in the 1980s as expert systems.
Expert systems serve as digital counterparts to human domain experts, leveraging a
comprehensive knowledge base to guide decision-making processes. These systems
bring a wealth of rules and domain-specific knowledge into play, enhancing the
precision of decision-making [5].
ONCOCIN [6] is an example of a system designed to assist oncologists in the
treatment of cancer patients. Note that the rules in expert systems are predefined and
programmed into the system, outlining how the system needs to behave in different
situations. Expert systems, thus, lack the ability to adapt or learn from new data,
and they operate within the boundaries of the predetermined rules set by human
experts and the initial knowledge programmed into the system.
The limitations of expert systems later motivated the development of machine
learning (ML) models in the 1990s and 2000s. In contrast to expert systems, ML
models extrapolate from data. ML involves the development of algorithms that
enable machines to learn and improve from experience. In adaptive radiotherapy,
ML algorithms can leverage extensive patient data to discern patterns and forecast
customized treatment plans [7]. For example, these algorithms can be used in
segmentation tasks to help delineate volumes concerning the tumor and organs at
risk. This progresses to treatment management, addressing motion during the
treatment and treatment plan optimization. Ultimately, decisions are interconnected
through reviewing treatment specifics, enabling adaptive changes.
While classical ML models do not have predefined rules such as expert systems,
they still require human expert efforts to design a feature set. These features, specific
attributes or characteristics selected from input data, provide structured information
to help algorithms discern and differentiate between data inputs, enhancing pattern
recognition and prediction capabilities. Deep learning (DL), a subfield of ML,
employs neural networks with multiple layers to extract representations from raw
data [8] without explicit feature engineering. DL has dominated AI research since
2010. DL, with its capacity to autonomously acquire hierarchical features from
intricate data, has the potential to enhance precision and efficiency across all stages
of radiotherapy treatment and its automation.
The landscape of AI in healthcare has undergone a transformative shift with the
rise of ML and DL techniques. This chapter focuses on ML and DL, describing their
1-2

D
X
y
y
D
Artificial Intelligence in Adaptive Radiation Therapy
capacity to discern intricate patterns, adapt to dynamic patient conditions, and
optimize and automate radiotherapy treatments. The nuances of these data-driven
approaches within the context of adaptive radiotherapy are explored. A practical
understanding of these AI approaches is provided from classical ML to DL and
generative models. An overview of ethical considerations regarding transparency
and bias is also provided.
1.2 Machine learning basics
This section briefly introduces basic definitions of ML, including different learning
strategies, feature representation, feature selection, and classical ML models.
1.2.1 Learning paradigms
An ML algorithm refers to a type of algorithm that can learn from experience. The
learning process can occur in various ways. Considering the data and the type of
information it provides, three main learning paradigms can be defined: (i) supervised, (ii) unsupervised, and (iii) self-supervised learning.
Supervised learning is the most commonly used ML paradigm. It earns its name as
it resembles a teacher guiding a student’s progress. Here, the algorithm learns from a
training dataset, with known correct answers, by making predictions iteratively,
which are then corrected. This process continues until the algorithm achieves a
satisfactory performance level. Formally, supervised learning involves training a
model on a dataset
(
) is paired with its corresponding ground-truth label (yi∈ Y). In summary,
∈x
i
=…
xy xy xy,,,,,,
train 1
{( ) ( ) ( )}
2
1
2
the goal of supervised learning is to learn a function f that accurately maps input
an output
, where
=
provides a reliable prediction ofy[9].
fx()
Unsupervised learning assumes that access is limited to the input variables, and
their corresponding ground truth remains unavailable (
objective of unsupervised learning is to model the underlying structure or distribution of the data to gain insights. Unlike supervised learning, no supervision is
provided for the given training data. Unsupervised learning problems extend beyond
traditional clustering, which involves grouping similar samples into distinct clusters,
and association problems, which discover interesting relationships or patterns
among variables in large datasets. Unsupervised models include dimensionality
reduction models, such as principal component analysis (PCA) [10], as well as
generative models such as autoencoders [11] and generative adversarial networks
(GANs) [12], which are widely used to generate examples resembling the training
data.
Self-supervised learning is a fairly new ML paradigm where a model learns from
the data without relying on externally provided labeled annotations. Instead of using
labeled data created by human annotators, self-supervised learning leverages the
inherent structure or information within the data to generate supervisory signals.
The model is tasked with predicting certain parts or properties of the input data
based on other parts, creating a pretext task such as jigsaw puzzle [13] or rotations
detection [14]. This process enables the model to learn useful representations or
, where each sample
n
n
=…
xx x,,,
{}
). The
ntrain 1 2
x
to
1-3

Artificial Intelligence in Adaptive Radiation Therapy
features from the input data, which can later be utilized for downstream tasks such
as classification or regression. Self-supervised learning is particularly valuable in
scenarios where obtaining labeled data is challenging or expensive.
In addition to these three paradigms, one may also consider semi-supervised
learning and reinforcement learning paradigms. In semi-supervised learning, we
typically have access to a small set of labeled data along with a large set of unlabeled
data, while in reinforcement learning an agent learns a task by interacting with an
environment and updates its approach based on the positive or negative rewards it
receives.
1.2.2 Regression versus classification
Classification and regression are two fundamental types of supervised ML tasks.
Classification aims to assign input data points to predefined categories or classes, as
illustrated in figure 1.1(a). Classifying tissue samples as normal, benign, or
malignant based on histopathological images, identifying the presence or absence
of specific genetic mutations, distinguishing between different stages of cancer
progression based on imaging data, and detecting cancer recurrence or metastasis
from medical imaging scans, are examples of classification tasks in precision
oncology. In contrast, in regression the goal is to predict a continuous numeric
output based on input features, as shown in figure 1.1(b). A regression model may
use features such as patient age, sex, tumor size, Gleason score, stage, and grade to
predict radiation dose, which is a continuous numeric value.
1.2.3 Feature engineering and representation
Representing real-world samples using structures understandable by computers
(often vectors or matrices) is the fi rst step in developing an ML model.
Considering the task at hand, domain experts define a set of features that capture
the most important characteristics of a sample. Such features serve as the descriptors
that constitute the input for ML algorithms. This step in developing ML models is
often referred to as feature engineering and representation. A sample dataset for the
Figure 1.1. Classification versus regression. (a) A sample classification example of two classes with a clear
decision boundary. (b) A regression model fitted on a sample dataset.
1-4

Artificial Intelligence in Adaptive Radiation Therapy
Table 1.1. A sample dataset for predicting the therapeutic role of radiotherapy.
Time
since
ID Age Sex
1 45 Male 120 2.5 Chest Lung T2 Curative
2 62 Female 90 1.8 Abdomen Kidney T1 Palliative
3 55 Male 150 1.9 Neck Thyroid T3 Curative
4 70 Female 180 2.2 Pelvis Prostate T2 Palliative
5 38 Male 75 1.6 Extremities Arm T1 Curative
6 48 Female 110 2.3 Chest Breast T2 Curative
7 64 Male 200 2.8 Neck Tongue T3 Palliative
8 58 Male 130 2.1 Abdomen Pancreas T2 Palliative
9 50 Female 160 2.4 Pelvis Ovary T3 Curative
10 42 Male 85 1.9 Extremities Leg T1 Curative
1st RT
Dose/fraction
(Gy) Region Site
T
stage Output
task of predicting the therapeutic role of radiotherapy is provided in table 1.1. The
expert has designed/engineered a set of d
= 8 features (the first eight columns) as
in
predictive features for the output (the last column). Each entity in the world, in this
case a patient, is then characterized using the feature set.
1.2.3.1 Feature encoding
As table 1.1 shows, the features are inherently different in terms of the values they
can take. Features such as sex, region, site, etc, are nominal categorical and
comprise a finite set of discrete values with no inherent ranking or order. An ordinal
variable also comprises a finite set of discrete values, but it has a clear ranked
ordering between the values, such as ratings (e.g. low, medium, high) or stages (e.g.
T1, T2, T3). Another feature type is numerical features such as age, dose, time, etc,
that take continuous values from a domain. As the name suggests, such features are
characterized by numeric values and can represent a wide range of data types, such
as integers, real numbers, or ratios.
Identifying and understating different feature types is important since it determines which type of ML models might be applicable. Different ML models require
different encoding schema for their inputs. As introduced in section 1.2.5, decision
trees, for instance, assume that all features are categorical, while models such as knearest neighbors and support vector machines that operate on vector spaces require
the features to be numerical.
Ordinal encoding, for instance, is used when the categorical values have a clear
order or rank. It assigns a unique numerical value to each category based on its
position in the order as shown in the feature stage in table 1.2. On the other hand,
one-hot encoding is used for categorical variables that do not have a natural rank
ordering. Assuming m different categories, each category can be encoded by a onehot vector {0, 1}
m
, where each dimension corresponds to a category. For instance, in
table 1.2, the category ‘Male’ for sex is encoded as 〈1, 0〉.
1-5

Artificial Intelligence in Adaptive Radiation Therapy
Table 1.2. Different encoding schema.
Feature Categorical Numerical
Stage T1, T2, T3 1, 2, 3
Sex Male, female
Age
R [1, 20), [20, 40), [40, 60)
>< >1, 0 , 0, 1
Converting numerical features to categorical features, on the other hand, often
involves discretization of the domain into intervals, each of which is considered a
category (see ‘Age’ in table 1.2). When discretizing numerical features, an important
design decision is the number of intervals, which can affect the model’s performance.
For instance, if the number of intervals is too low, the categorical feature may
oversimplify the underlying numerical distribution, losing important information
and patterns. Discretizing numerical features can also be used for data deidentification [15]. This reduction in granularity serves as a privacy-enhancing
mechanism, especially in medical applications where the disclosure of fine-grained
details could lead to the re-identification of individuals.
1.2.3.2 Feature selection
Medical data are inherently complex, often characterized by many features, and
patient information extends far beyond basic demographics, comprising variables
such as genetic markers, vital signs, medical history, diagnostic results, etc. The
number of features in adaptive radiotherapy is even larger since it requires
monitoring changes over time. Patients’ anatomy, for instance, may evolve during
the course of treatment due to weight loss, tumor regression/progression, and each
time point adds new features. Moreover, radiotherapy involves a large set of
dosimetric features, including dose distributions and dose–volume histograms,
which leads to an increase in the number of features.
Many features, however, can lead to computational inefficiencies, resulting in
long training and prediction times. Moreover, it might increase the risk of overfitting (discussed in section 1.4), as the model may capture noise and intricacies
specific to the training data, hindering its ability to generalize to new instances. The
complexity imposed by numerous features impedes computational efficiency and
hampers model interpretability. Thus, pre-processing is required to select a subset of
features to mitigate these challenges and enhance the overall performance of the ML
model. Feature selection involves choosing a subset of the original features from the
dataset based on their relevance to the task at hand [16]. The goal is to retain the
most informative features while discarding irrelevant or redundant ones, reducing
the complexity of the model, and potentially improving its interpretability. As an
example, in some datasets, features such as patient ID often carry random values
without meaningful patterns for predictive purposes. Including such features may
introduce noise and deteriorate the predictive ability. Therefore, careful feature
selection is crucial to exclude random-valued features. Note that patient IDs may
1-6

)
Artificial Intelligence in Adaptive Radiation Therapy
contain some information. For instance, some databases might incrementally assign
IDs, meaning that the ID is correlated with the date of diagnosis. In other cases
where certain features are constant or nearly constant across a patient cohort,
including all instances of these features, they may not add value and could be
considered redundant. For instance, body site might be highly correlated with sex.
An example is the ovary, specific to the female reproductive system, rendering the
sex feature redundant. Finally, some features may bear some information for output
prediction. Their contribution as strong predictors needs to be evaluated to avoid
overemphasizing less impactful variables. Feature selection involves using various
measures to assess the relevance and importance of features in a dataset. Common
techniques include evaluating features based on feature correlation coefficients or
statistical measures such as mutual information.
Pearson correlation, for instance, is a statistical measure that quantifies the
strength and direction of a linear relationship between two variables x and y:
n
−−
xxyy
()()
i
∑
i
1
=
=
r
xy
n
−−
xx xy
()()
i
∑∑
i
1
==
i
n
2
i
i
1
,
2
()
1.1
wherexandyare the mean values forxandy. The correlation coefficient is a
dimensionless value that ranges from −1 to 1. A coefficient of 1 indicates a perfect
positive linear relationship, meaning that as one variable increases, the other also
increases proportionally. Conversely, a coefficient of −1 signifies a perfect negative
linear relationship, where one variable decreases as the other increases. A coefficient
of 0 implies no linear correlation between the variables.
Mutual information, on the other hand, can capture non-linear relationships
between features. Assuming
…,,
as possible values fory, mutual information
v
y1
∣∣
…
as possible values for featurexand
u,,
x1 ∣∣
is an information-
xy,(
theoretic measure which is defined based on individual and joint entropies of two
random variables:
=+−IxyHxHyHxy,,
() () () ()
x
∣
== =
Hx Px u P x u
() () ()
∑
=
j
1
y
∣
== =
Hy Py P y
() () ()
∑
=
k
1
log log ,
jj
log log .
vv
kk
1.2
()
This formulation of the mutual information is only applicable to categorical
features. For numerical continuous features, an estimate is used, which is implemented in most of the existing ML libraries such as scikit-learn [17].
Feature selection is a subset of the broader dimensionality reduction paradigm
[18]. In fact, any feature subset selection is a dimensionality reduction but not vice
versa. Some dimensionality reduction methods, such as PCA and autoencoders,
1-7

Artificial Intelligence in Adaptive Radiation Therapy
Figure 1.2. Two different geometric distributions of data: (a) linear separable and (b) linear non-separable.
transform the original feature space into a lower-dimensional representation by
creating new composite features. In medical applications, feature selection might be
preferable to such approaches to enhance interpretability. Feature selection retains a
subset of the original features, providing transparency in the model’s decisionmaking process. This transparency is important in medical contexts where understanding the relationship between input features and predictions is essential for
gaining trust from clinicians and ensuring the responsible adoption of the model in
clinical practice.
1.2.4 Linear separability
When designing an ML model, an important consideration is the geometric
distribution of data in terms of linear (non-)separability. Linear separability and
linear non-separability refer to the inherent structure of a dataset in relation to a
classification task. In a scenario of linear separability, classes within the data can be
effectively separated by a hyperplane—a geometric construct in the feature space as
presented in figure 1.2(a). This means that a straight line or plane can distinguish
between different classes, allowing the application of linear classification algorithms
such as linear support vector machines or logistic regression. On the other hand, in
cases of linear non-separability, a single hyperplane cannot accurately separate the
classes, requiring more complex decision boundaries, such as curves or non-linear
surfaces as illustrated in figure 1.2(b). In such cases, non-linear classification
techniques such as kernel methods or DL models may be more appropriate to
capture the intricate relationships within the data. Understanding whether a dataset
exhibits linear separability or not is important in selecting the most suitable ML
approach for a given classification problem.
1.2.5 Classical models
Classical ML models have been foundational in the evolution of AI, serving as
powerful tools for data analysis and decision-making across various domains,
1-8

Artificial Intelligence in Adaptive Radiation Therapy
including healthcare. These models rely on predefined features as described in
section 1.2.3 and statistical techniques to identify patterns and make predictions.
Understanding these models is important since it provides a knowledge of the
principles and methods widely used in the field, enabling practitioners to appreciate
the evolution of ML.
1.2.5.1 Decision trees
Decision trees are a fundamental concept in ML, widely employed for their intuitive
and transparent representation of decision-making processes. Serving as predictive
models, decision trees navigate through a series of (often binary) choices based on
input features to reach a final decision or prediction. Figure 1.3 shows a decision tree
learned based on the dataset provided in table 1.1. Decision trees are composed of
three types of nodes: root, internal nodes, and leaves. All root and internal nodes
correspond to features, and depending on the feature value, one of the sub-trees is
traversed. The sub-trees are iteratively traversed until a leaf node is reached,
corresponding to a classification decision.
The training process of a decision tree is often referred to as tree induction [20].
During the induction, at each root/internal node, the decision tree tries to partition
the training samples (based on feature value) such that the subsets have the least
impurity regarding class labels. The class impurity can be quantified using measures
such as entropy or Gini index. Assuming a binary classification problem, the
maximum entropy and Gini values are 1.0 and 0.5, respectively. The minimum value
for these measures often occurs for the leaf nodes—the least impurity and the most
certainty for the classification. In some cases, however, the designer might prefer to
Figure 1.3. A decision tree for predicting the therapeutic role of radiotherapy (inspired by [19]). Based on the
value of features, the branches are iteratively traversed until a leaf node is reached.
1-9

Artificial Intelligence in Adaptive Radiation Therapy
terminate the induction process to keep the tree below a certain height1(for instance,
to avoid over-fitting). In such cases, some leaf nodes might have a non-zero impurity
value. Note that the induction process may ignore some of the features. For
instance, in figure 1.3, features such as sex or age do not appear in the tree. This
means that the decision tree has implicitly performed a feature subset selection.
1.2.5.2 Random forests
Decision trees, while powerful, suffer from limitations. Their tendency to overfit
(discussed in section 1.4.4) and sensitivity to data variations can lead to unreliable
predictions. Moreover, decision trees have a greedy optimization scheme and they
have limited ability to capture complex relationships.
Random forests represent an ensemble machine learning method that leverages
the collective power of multiple decision trees. Each tree is constructed from a
random subset of features, introducing an element of stochasticity that enhances
generalization. By aggregating the predictions from this ensemble, random forests
achieve robustness to over-fitting and improve overall prediction accuracy compared
to individual decision trees. This technique has gained widespread adoption due to
its effectiveness in various classification and regression tasks.
1.2.5.3 Support vector machines
Assuming a linear separable distribution for data, the classes can be separated
d
using a line or hyperplan e characterized as
+=
wb0.
,where
in
∈wR
and b ∈
R
are the slope and intercept parameters of the hyperplane. There are, however,
infinite possible values for the parameters, resulting in many potential discriminators as shown in figure 1.4(a). While all t he hyperplanes effectively separate the
training samples, some of them may misclassify test samples, leading to poor
generalization.
Figure 1.4. Support vector machine. (a) Several possible separating lines and (b) the separating line with the
largest margin.
1
The height of a tree is defined as the maximum number of internal nodes on the longest path from the root to
any leaf node.
1-10
Соседние файлы в папке Библиотека им академика М.И. Перельмана
