Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5525_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Editor biographies
- •Yi Wang
- •X. Sharon Qi
- •List of contributors
- •1.2.3 Feature engineering and representation
- •1.2.4 Linear separability
- •1.2.5 Classical models
- •1.1 A brief introduction to AI
- •1.2 Machine learning basics
- •1.2.1 Learning paradigms
- •1.3 Artificial neural networks
- •1.3.1 Feed-forward neural networks
- •1.3.2 Recurrent neural networks
- •1.3.3 Convolutional neural networks
- •1.3.4 Attention
- •1.3.5 Training neural networks
- •1.3.6 Applications and use cases of deep learning
- •1.4 Model training and evaluation
- •1.4.1 Hyperparameters
- •1.4.2 Data split
- •1.4.3 Evaluation metrics
- •1.5 Generative models
- •1.5.1 Generative adversarial networks
- •1.5.2 Diffusion models
- •1.5.3 Applications and use cases
- •1.6 Ethical consideration and bias
- •1.6.1 Transparency and explainability
- •1.6.2 Bias and fairness
- •1.6.3 Data privacy violation
- •1.6.4 Risk and misuse
- •1.7 Summary
- •References
- •2.1 Introduction
- •2.1.2 Staff roles in radiation therapy
- •2.2 Overview of AI in radiation therapy
- •2.2.1 Patient evaluation and dose prescription
- •2.2.2 Treatment simulation
- •2.2.3 Contouring
- •2.2.4 Treatment planning
- •2.2.5 Quality assurance
- •2.2.6 Treatment delivery
- •2.2.7 Response assessment and toxicity management
- •2.3 Summary
- •3.1 Introduction
- •3.1.1 Introduction of clinical decision making and AI
- •3.1.2 The role of AI in clinical decision making
- •3.2 AI algorithms for clinical decision making
- •3.2.2 Radiomics
- •3.2.3 Data integration by AI
- •3.2.4 Interpretability of AI models
- •3.3 Application of AI in clinical decision making
- •3.3.1 Diagnosis and disease phenotyping
- •3.3.2 Personalized treatment
- •3.3.3 Treatment outcome and prognosis prediction
- •3.4 Challenges and future directions of AI in clinical decision making
- •3.4.1 Challenges and concerns
- •3.4.2 Future directions
- •3.5 Summary
- •References
- •4.1 Introduction
- •4.2 Imaging for treatment planning
- •4.2.1 CT simulation
- •4.2.2 4D-CT
- •4.2.3 PET/CT
- •4.2.4 MRI
- •4.3 Imaging for treatment guidance
- •4.3.1 Portal imaging
- •4.3.2 CBCT
- •4.3.3 CT-on-rail and CT-linac
- •4.3.4 MR-linac
- •4.3.5 PET-linac
- •4.4 Imaging for motion management
- •4.4.1 ExacTrac
- •4.4.2 Varian triggered imaging
- •4.4.3 4D-CBCT
- •4.4.4 Cine MRI
- •4.4.5 4D-MRI
- •4.4.6 Surface imaging
- •4.5 Imaging for treatment assessment
- •4.5.1 Contrasted CT
- •4.5.2 PET/CT
- •4.5.3 Functional MRI
- •4.6 Summary
- •5.1 Introduction to big data in radiation oncology
- •5.1.1 Overview of big data
- •5.1.2 Sources of big data in radiation oncology
- •5.1.3 Big data and AI in radiation oncology
- •5.2 Big data lifecycle in radiation oncology
- •5.2.1 Data aggregation and storage
- •5.2.2 Data sharing and security
- •5.4 The application of big data in radiation oncology
- •5.4.1 Medical image segmentation
- •5.4.2 Automatic treatment planning
- •5.4.3 Treatment response prediction
- •5.4.4 Quality assurance and patient safety
- •5.4.5 Clinical decision support
- •5.2.3 Data visualization
- •5.2.4 Knowledge creation and implementation
- •5.2.5 Data archiving and deletion
- •5.3 Big data analytics with AI
- •5.3.1 Data processing and integration
- •5.3.2 AI modeling
- •5.5 Challenges and future perspectives
- •5.6 Summary
- •Reference
- •6.1 The road to ART
- •6.1.1 3D conformal radiotherapy (3DCRT)
- •6.1.2 Intensity modulated radiotherapy (IMRT)
- •6.1.3 Image-guided radiotherapy (IGRT)
- •6.1.4 Adaptive radiotherapy (ART)
- •6.2 ART workflow and implementation
- •6.2.2 Current practice
- •6.2.3 Clinical impact
- •6.3 Considerations for implementing online ART
- •6.3.1 Time as a limiting factor
- •6.3.2 Implications for fast and reliable re-planning
- •6.3.4 Clinical considerations
- •6.4 Summary
- •7.1 Components of ART workflow
- •7.1.1 Simulation
- •7.1.2 Pre-planning
- •7.1.3 Online imaging and daily re-planning
- •7.1.4 Quality assurance
- •7.2 AI-driven ART
- •7.2.1 Simulation
- •7.2.2 Pre-planning
- •7.2.3 AI for delivery
- •7.3 Outlook and future directions
- •7.3.1 Real-time ART with AI
- •7.3.2 Dose escalation and functional adaption with AI
- •7.4 Summary
- •References
- •8.1 Introduction
- •8.2 Synthetic CT: deep learning methods
- •8.2.1 Conventional methods
- •8.2.2 U-Net
- •8.2.3 Generative adversarial networks
- •8.2.4 Denoising diffusion probabilistic model
- •8.3 Synthetic CT from CBCT
- •8.3.1 Noise and artifact reduction
- •8.3.2 Online dose calculation
- •8.3.3 Online image segmentation
- •8.4 Synthetic CT from MRI
- •8.4.1 Synthetic image accuracy
- •8.4.2 Dose calculation in MR-only radiation therapy
- •8.4.3 PET attenuation correction
- •8.4.4 Image registration
- •8.5 Discussion and outlook
- •8.6 Summary
- •References
- •9.1 AI-based image registration and segmentation for ART
- •9.1.1 Adaptive radiation therapy
- •9.2 Artificial intelligence
- •9.2.1 What is machine learning?
- •9.2.2 What is deep learning?
- •9.3 Deep learning: the basic components
- •9.3.1 Convolutional neural networks: looking at the picture
- •9.3.2 Pooling layers: keeping what matters most
- •9.3.3 Fully connected (dense) layers: bringing it all together
- •9.3.4 Activations
- •9.3.5 Loss: driving the model
- •9.3.6 Auto-encoders: remove the noise
- •9.3.7 Supervised versus unsupervised learning
- •9.3.8 Pre-trained convolutional neural networks
- •9.4 Image registration: bringing two images together
- •9.4.1 Registration similarity metrics
- •9.4.2 Types of registrations
- •9.5 AI-based image registration
- •9.5.1 Supervised learning
- •9.5.2 Unsupervised learning
- •9.5.3 Registration in ART
- •9.5.4 Commonalities in architectures
- •9.6 Image segmentation
- •9.6.1 Introduction: coloring by the numbers
- •9.6.2 Segmentation networks
- •9.6.3 Best practices
- •9.7 Summary
- •10.1 Introduction
- •10.1.1 Overview of chapter content
- •10.2 The landscape of AI-assisted dose prediction
- •10.2.1 Traditional machine learning for dose prediction
- •10.2.2 Deep learning-based dose prediction
- •10.2.3 Challenges in AI-assisted dose prediction
- •10.3 Re-planning workflows powered by AI
- •10.3.1 Deep learning for re-planning pipelines
- •10.4 Future directions of AI-assisted dose prediction and re-planning
- •10.5 Summary
- •11.1 Introduction
- •11.2.1 Imaging-based motion monitoring
- •11.2.2 Delivery system actions
- •11.2.3 Challenges for real-time ART implementation
- •11.3 AI in real-time ART workflows
- •11.3.1 Improving intrafraction motion monitoring through AI
- •11.3.2 Mitigating system latency through AI
- •11.4 AI for ART delivery: future directions
- •11.4.1 Management of non-respiratory motion
- •11.4.2 Training AI models with small or unpaired datasets
- •11.4.4 Biology-guided ART delivery
- •11.5 Summary
- •References
- •12.1 Introduction
- •12.2 Patient QA
- •12.2.1 Pre-planning QA
- •12.2.2 Pre-treatment plan QA
- •12.2.3 On-treatment QA
- •12.3 Treatment delivery systems and instruments
- •12.3.1 Machine commissioning
- •12.3.2 Machine QA
- •12.3.3 Dosimetry tool QA
- •12.4 Summary
- •References
- •13.1 Data resources for response modeling in radiotherapy
- •13.1.1 Clinical data
- •13.1.2 Imaging (radiomics)
- •13.1.3 Treatment planning (dosiomics)
- •13.1.4 Multiomics
- •13.2 Radiotherapy treatment outcome modeling
- •13.2.1 TCP/NTCP in radiotherapy
- •13.2.2 Clinical outcomes versus PROs
- •13.2.3 Machine learning response prediction
- •13.2.4 Explainability of ML response models
- •13.2.5 Sample use cases
- •13.3 AI response-based adaptive radiotherapy
- •13.3.1 Requirements and challenges
- •13.3.2 Prediction versus treatment optimization
- •13.3.3 Sample use cases
- •13.4 Challenges and recommendations
- •13.5 Summary
- •Acknowledgments
- •References
- •14.1 Overview of challenges in AI-driven ART
- •14.2 Data challenges
- •14.2.1 Data availability
- •14.2.2 Data quality
- •14.2.3 Data privacy
- •14.3 Technical challenges
- •14.3.2 Model robustness and generalizability
- •14.3.3 Model explainability and interpretability
- •14.4 Challenges associated with online and real-time workflows
- •14.4.1 Image quality
- •14.4.2 Dose calculation
- •14.4.3 Real-time ART
- •14.5 Operational challenges
- •14.5.1 Clinical validation
- •14.5.3 Staff training
- •14.5.4 User experiences
- •14.5.5 Quality management program
- •14.5.6 Financial challenges
- •14.6 Ethical, regulatory, and legal challenges
- •14.6.1 Ethical issues
- •14.6.2 Regulatory and legal issues
- •14.7 Summary
- •References
- •15.1 Clinical considerations for CT-based offline ART
- •15.1.1 Patient and site selection
- •15.1.2 Re-simulation
- •15.1.3 Re-planning
- •15.1.4 Plan summation and evaluation
- •15.1.6 Limitations and future directions
- •15.2 Clinical considerations for CBCT/CT-based online ART
- •15.2.2 Patient and site selection
- •15.2.3 Simulation
- •15.2.4 Pre-planning review
- •15.2.5 Reference planning
- •15.2.9 Limitations and future directions
- •15.3 Summary
- •References
- •16.1 Introduction
- •16.2 Overview of MRI-guided ART systems
- •16.3 MRI-guided ART workflow
- •16.4 AI applications for MRI-guided ART
- •16.4.1 Synthetic CT generation
- •References
- •16.4.2 Auto-segmentation
- •16.4.3 Image registration
- •16.4.4 Others
- •16.4.5 Future AI development and implementation
- •16.5 Summary
- •17.1 Functional PET-guided ART
- •17.1.1 PET-based functional imaging overview
- •17.1.2 From anatomy to function: the power of PET in radiation therapy
- •17.1.5 Conclusions and future prospects
- •17.2 Functional MRI-guided ART
- •17.2.1 From anatomy to function: the power of functional MRI in radiation therapy
- •17.2.4 Conclusion and future prospects
- •17.3 Summary
- •References
- •18.1 Proton ART
- •18.1.1 Clinical context and necessity
- •18.1.2 Patient populations
- •18.1.4 Rationale for AI in proton ART
- •18.2 AI in proton ART
- •18.2.1 Imaging
- •18.2.2 Deformable and rigid registration
- •18.2.3 Contour propagation
- •18.2.4 Dose calculations
- •18.2.5 Plan optimization
- •18.2.6 Other developments
- •18.3 Implementation of adaptive proton therapy
- •18.4 Summary
- •References
- •19.1 Designing clinical trials with AI
- •19.1.1 The essential role of clinical trials
- •19.1.2 Trial protocols and methodologies
- •19.1.3 AI-driven clinical trial design and execution
- •19.1.4 Incorporation of digital twins (DTs) in clinical trials
- •19.2 Implementation of AI in ongoing clinical trials
- •19.2.1 Integration with existing clinical trial frameworks
- •19.2.2 Quality assurance, compliance, and standardization
- •19.3 Case studies of AI in adaptive radiotherapy trials
- •19.3.1 Overview of guidance for advanced radiotherapy in clinical trials
- •19.3.2 AI in the radiotherapy clinical trial quality assurance processes
- •19.4 Ethical and regulatory considerations
- •19.4.1 Patient consent and data privacy
- •19.4.2 Bias, fairness, and transparency
- •19.4.3 Regulatory guidelines and compliance
- •19.5 Future directions and challenges
- •19.5.1 Emerging technologies and techniques
- •19.5.2 Alternative strategies
- •19.6 Conclusion
- •19.7 Summary
- •References
- •20.1 Risk management
- •20.1.1 Prospective risk assessments
- •20.1.2 Root cause analysis

Artificial Intelligence in Adaptive Radiation Therapy
Support vector machines (SVMs), on the other hand, try to identify a hyperplane
that not only accurately classifies training data points but also maintains a
significant distance, or margin, to the nearest data points of each class, which are
called support vectors as presented in figure 1.4(b). Acting as a buffer zone, this
margin provides robustness to variations in the data and improves the model’s
generalization ability. By maximizing this margin, an SVM aims to achieve a clear
separation between classes, reducing the risk of misclassification and enhancing the
model’s ability to generalize well to unseen data.
Support vectors are embedded within the model after the training in an SVM. In
the context of sensitive medical data, this can potentially violate patient privacy due
to the inclusion of support vectors in the model. In particular, if the support vectors
correspond to specific individuals in the dataset, it may inadvertently disclose
sensitive information about those patients, compromising their privacy. This means
that these models should not be trained on identifiable data.
1.2.5.4 k-nearest neighbors
k-nearest neighbors (kNN) is a simple ML algorithm that operates on the principle
of proximity. The intuition behind kNN is that similar data points in a feature space
tend to belong to the same class or category. Distance metrics often measure this
similarity in vector space. A common distance metric is Euclidean distance, which
measures the straight-line distance between points in a vector space. Other popular
distances include Manhattan distance, which calculates the sum of absolute differences along each dimension. Other metrics, such as Minkowski distance, allow a
tunable parameter to adjust the emphasis on different dimensions [20].
An illustration of the kNN classification of a test sample (the point with no color)
with different k values is provided in figure 1.5. The prediction is often made by a
majority voting of the labels that fall within the neighborhood. Note how the
decision changes with increasing k value. For a binary classification task, choosing
an even or odd value for k is important since it impacts the resolution of ties as
shown in figure 1.5(b). An odd k value is preferred to avoid ties (figure 1.6(c)),
ensuring a clear majority decision when voting among the nearest neighbors. Note
that in multi-label classification a tie situation may occur when multiple neighbors
have the same number of occurrences for different labels, and the resolution of the
tie does not depend on the specific value of k. In general, a smaller k value tends to
Figure 1.5. Different neighborhoods of an arbitrary test sample: (a) 1NN classifying as blue, (b) 2NN, no
classification due to tie, and (c) 3NN classifying as red.
1-11

t
Artificial Intelligence in Adaptive Radiation Therapy
Figure 1.6. Different activation functions: (a) sigmoid, (b) ReLu, and (c) hyperbolic tangent [22].
result in a more flexible and low-bias model, allowing it to adapt well to intricate
patterns in the data, but it might be sensitive to noise [21]. On the other hand, a
larger k value leads to a smoother decision boundary, reducing sensitivity to noise
but potentially introducing bias.
While simple, kNN can be a powerful and effective algorithm for certain types of
datasets, in particular when the underlying relationships are based on local patterns
and proximity. kNN can be categorized as a lazy learner [20], where it learns by
memorizing the training data and classifying new instances based on their proximity
to known examples. This means that the training data must always be available to
the model even when the model is deployed. An important consideration when using
such models is the privacy of patients.
1.3 Artificial neural networks
Artificial neural networks (ANNs), inspired by the human brain, consist of
interconnected layers of arti fi cial neurons. While traditional models such as decision
trees and SVMs require handcrafted features, neural networks can automatically
learn complex and hierarchical representations through multiple layers of interconnected neurons.
1.3.1 Feed-forward neural networks
Feed-forward neural networks are the most basic form of neural nets which express
a mapping:
=+yfxWb,1.3() ()
d
where f(·) is an activation function,
in
∈xR
is the input vector,
output. The mapping is characterized by two learnable parameters
d
(weight matrix) and
out
∈
(bias vector). If f(·) = I(·) (identity function), equation
R
(1.3) expresses a linear mapping. In this case, the neural net can only approximate
linear functions. On the other hand, if a non-linear activation is used, the network
can approximate more sophisticated functions. Examples of non-linear activations
such as sigmoid, rectified linear units (ReLU), and hyperbolic tangent (tanh) are
provided in figure 1.6.
1-12
∈yR
d
∈×R
ou
is the
dd
in out

Artificial Intelligence in Adaptive Radiation Therapy
A problem with sigmoid and hyperbolic tangent functions is that they can easily
saturate, i.e. large input values converge to 1.0, and small values converge to −1or0
for hyperbolic tangent and sigmoid, respectively. Once saturated, it becomes
challenging for the learning algorithm to continue to adapt the weights to improve
the model’s performance. The ReLU, on the other hand, is a piecewise linear
function and has greatly improved the performance of neural networks. Since
rectified linear units are nearly linear, they preserve many properties that make
linear models easy to optimize [8].
Another common activation function is softmax, which is often used to obtain
probability values from a vector h = [h
1,h2
, …,hd]:
softmax
()
⎡
⎢
=…h
⎢
⎢
⎣
h
exp
()
1
,,
h
exp
()
∑∑
j
j
exp
exp
j
h
()
d
h
()
j
⎤
⎥
.1.4
⎥
⎥
⎦
()
The softmax function is often applied to the output layer of a network and yields
class probabilities for a multi-class classification problem.
1.3.2 Recurrent neural networks
The primary assumption of feed-forward networks is that the features of a sample do
not change over time; thus, their input is often represented by a single feature vector.
Many phenomena, however, have temporal dynamics where the behavior changes
through time. A patient, for instance, might have new tests, radiation dose, changes
in organs, etc, which means that a sequence of feature vectors x
describes the patient where x
denotes a time-dependent feature. A predictive system
t
, x2, …, xTnow
1
should consider such changes to provide a personalized treatment recommendation.
Recurrent neural networks (RNNs) are a class of networks introduced for
processing sequences [23] where the current state h
current input x
t
d
in
∈ R
, and previous hidden state h
t−1
d
h
t
∈ R
is a function of the
. Note that the term state is
borrowed from the terminology used in control theory, where a system is often
described by its state, which represents a set of variables that describe the system at a
particular point in time [22]. Figure 1.7 shows a recurrent network. Note the
feedback loop, which allows a memory of different time steps/positions.
Figure 1.7. A recurrent network unfolded through time. The hidden representation at time stept(
function of
and
x
t
. Reproduced with permission from [22].
−ht 1
1-13
h
t
)isa

Artificial Intelligence in Adaptive Radiation Therapy
The network can be unfolded through time/positions, showing that a recurrent net is,
in fact, a feed-forward network that is applied over all positions.
Formally, an RNN can be characterized as
==++
hfxh fxWhWb,.
()( )
ttt tt11
−−
The function is characterized by two weight matrices
are learning parameters, as well as a bias vector
hidden state
sequence up to the current input. The
represents a history of observed inputs from the beginning of a
t
corresponding to the final position (T)is
t
∼
∼
dd
hin
∈×R
d
h
∈
R
,
. In a recurrent net, the
∈
1.5
()
×Rdd
hh
that
assumed to represent the whole sequence and can be followed by a simple classifier
to determine the label for the desired task. The weight matrix then needs to be
updated based on the error of the classifier’s prediction. In recurrent nets,
particularly with long sequences, the vanishing gradient problem can occur, leading
to difficulties in preserving relevant information across distant time steps/positions,
hindering the network’s ability to effectively capture long-term dependencies. This
issue can result in diminished learning capabilities for tasks requiring the retention of
context over extended temporal spans.
1.3.3 Convolutional neural networks
Convolutional neural networks (CNNs) were introduced for pattern recognition
tasks to process spatial data, such as images [24] and have been exponentially
applied in the medical computer vision domain. A major difference between a
neuron in a CNN from a regular ANN is that its neurons are designed to be threedimensional with width, height, and depth. A CNN comprises layers where each
layer transforms a 3D (i.e. RGB channels) input to a 3D output. The basic
components of a CNN are the convolutional, pooling, and fully connected layers.
Convolution layers use filters or kernels, which are spatially small 3D matrices, to
convolve or slide across through the width and height of the input volume. The
resulting output is then passed to an activation function (see section 1.3.1),
producing a 2D feature map or activation map for each filter that learns visual
patterns.
Pooling layers downsample an input volume to decrease the size of feature
representation progressively, hence reducing the number of parameters and the
computational time. It is commonly followed by a convolution layer, achieving
spatial invariance. Average pooling computes the average value in a sliding window
while preserving general information. Max pooling selects the maximum value in a
sliding window, effectively capturing the most prominent features and patterns
during down-sampling.
The fully connected (FC) layer flattens the output of the final convolutional or
pooling layers into a vector h. An output layer is responsible for producing the final
predictions or classifications (elaborated in section 1.3.5). In a CNN, the initial
layers often detect general patterns (image edges and colors) and deeper layers learn
task-specific representations. Figure 1.8 illustrates a simple example of a
1-14

Artificial Intelligence in Adaptive Radiation Therapy
Figure 1.8. A CNN architecture comprising five stacked layers: input layer of a histopathology whole slide
image, convolutional layer, pooling layer, fully connected layer, and output layer to predict a class.
classification task predicting cancer malignancy from an H&E histopathology whole
slide image (WSI) using a CNN-based architecture.
While CNN is often applied to classify entire images, a specific architecture of
CNNs designed for pixel-wise biomedical image segmentation is U-Net [25]. It has
become the basis of many auto-segmentation algorithms applied to medical images
[26]. In this design, the encoder network or the contracting path is a feature
extractor. It acquires an abstract representation of the input image through a series
of encoder blocks comprising convolutions while the spatial dimensions of the input
volume are gradually reduced. The decoder network or the expanding path increases
the spatial dimensions of the extracted features, localizing these features in the image
and producing a segmentation map.
1.3.4 Attention
CNNs effectively capture local spatial information through convolutional filters but
face challenges in capturing global context and long-range dependencies within data
[27]. Their limited receptive fields (due to filters with small window size) can hinder
understanding complex scenes where relationships between distant elements are
crucial. Adaptive attention to local and global context is required to address these
challenges. This has motivated the development of architecture such as transformers
[28] that use the attention mechanism.
1.3.4.1 Self-attention
The attention mechanism resembles a spotlight that an ML model can use to focus
on specific parts of input data when making predictions. Instead of treating all parts
of the input equally (equal weights), the model can assign different levels of
importance or ‘attention’ to different elements [29]. While different forms of
attention mechanism exist, one of its most well-known and widely used forms is
self-attention, defined as
…=
x x XW XW XW; ; softmax ,
[] (())
1
∼
X
T
qkT
A
v
1.6
()
1-15

]
]
Artificial Intelligence in Adaptive Radiation Therapy
whereq,k, and
v
are training parameters and
=…Xxx;;
[
is a matrix of
T1
input representations (positions/time steps 1 to T). In this formalism, the product
qkT
XW XW
()
in a square matrix
represents the similarity between inputs at different positions, resulting
TT
∈×R
, where
represents the amount of attention that the
tk,
tth input element pays to the kth element. A softmax activation is then applied to
(row-level) to obtain a probability distribution. The self-attention mechanism results
in a matrix of contextualized representations
obtained using an adaptive weighted average of
∼
=…
Xxx;;
[
, enabling the model to
…xx,,
T1
, where each row is
T1
focus on different input elements.
1.3.4.2 Positional encoding
Although a self-attention mechanism captures relationships between different
elements in a sequence, it does not account for positional information. Thus,
additional positional encoding is often introduced before applying self-attention to
ensure the model can discern the sequence’s order and the relative positions of
elements. This positional encoding matrix P (often a constant matrix) is added to
matrix X. A common form of positional encoding is sinusoidal, which is
represented as
2
=Ptsin 10000 1.7
tjjd,2
+
tjjd,2 1
/
() ()
=
tcos 10000 ,
()
where t is the position index and d is the number of input dimensions. An illustration
of a positional encoding matrix P for 15 positions (t = 0, …, 14) and d = 100
/
2
/
/
2
provided in figure 1.9(a). Note that the representations corresponding to adjacent
positions (rows) are similar while distance positions have dissimilar representations,
allowing the position of each element in a sequence to be captured explicitly.
is
1.3.4.3 Transformer encoder
Transformers are attention-based networks introduced to address the problem of
vanishing long-distance information that exists in RNNs and CNNs [28].
Figure 1.9(b) illustrates the transformer architecture, where its input is the matrix
X = [x
; …; xT]. In contrast to a recurrent net, where the input is processed one
1
position at a time, in transformers, the inputs at all positions are fed into the
network simultaneously. A self-attention layer obtains the new contextual representations. A residual connection is used for each layer to avoid issues such as
exploding gradient. The final output of the transformer is a matrix of hidden
representations h
(t = 1, …, T).
t
Modern neural network architectures are versatile, and advancements in different
domains within the broader field of DL are interconnected. CNNs, for instance,
were originally introduced for image data, and their success sparked interest in
exploring their adaptability to other domains such as natural language processing
2
The number of positions and the value of d are chosen arbitrarily.
1-16

Artificial Intelligence in Adaptive Radiation Therapy
Figure 1.9. Transformer architecture: (a) positional encoding and (b) transformer encoder.
Figure 1.10. A vision transformer.
(NLP) [30]. On the other hand, transformers, initially developed by the NLP
community, marked a significant paradigm shift in sequence modeling.
Surprisingly, their capabilities transcended their original NLP domain and were
later successfully applied to computer vision tasks. This adaptation is called the
vision transformer, where an image is seen as a sequence of patches, as shown in
figure 1.10 [31]. Benefiting from the attention mechanism, a vision transformer
allows adaptive attention to local and global contexts, often leading to superior
performances compared to CNNs.
1-17

)
)
c
Artificial Intelligence in Adaptive Radiation Therapy
1.3.5 Training neural networks Making predictions in neural nets involves using learned representations to generate
outputs. The neural layers discussed in the previous sections are mostly encoders
that try to learn a representation from an input. The learned representation
(usually a vector) is then used as input to a classifier/regressor to make predictions.
For classification, a simple projection
and
is the number of classes.
N
c
The projection results in a vector
=zhW
∈zR
cls
can be used, where
N
c
of logits where each dimension
∈×R
dNcls
hc
corresponds to a class. A softmax activation is then used to convert the logit values
to class probabilities
=pzsoftmax(
corresponds to the predicted class. For regression, a projection
used to map
to a single scalar value where
and the dimension with highest
=zhW
dreg 1
h
∈×R
.
probability
p
c
reg
is often
Loss functions are mathematical measures that quantify the difference between
predicted values and actual values in an ML model, serving as a guide for adjusting
model parameters. The goal is to minimize the loss function during the training process
to improve the accuracy of the model. For classification tasks, a commonly used loss
function is cross-entropy loss. The cross-entropy loss
two input arguments
∈y0,1
N
{}
, a one-hot vector representing the true class, andpas
yp y p,loglog
() (
=−
c
∑
cc
=
cN1
,has
the probability vector (softmax output). For regression tasks, squared error
=−
yz y z,2()( )
is a commonly used form.
Gradient descent is an iterative approach for training neural nets and the
minimization of their prediction error (over training data). The loss L is a function
of the network’s trainable parameters
zation involves finding an optimal
(see figure 1.11). Therefore, the minimi-
∗
that minimizes the error. Most neural
networks describe non-convex functions, meaning that several local minima exist.
In this case, a minimum is found using a gradient descent approach that relies on
Figure 1.11. An error/loss surface as a function of training parameters (here w1and w2).
1-18

Artificial Intelligence in Adaptive Radiation Therapy
∇L, the gradient of loss L with respect to the set of trainable parameters
The gradient represents the direction and the magnitude of the ascent of the
loss function. In each iteration, trainable parameters are updated with a step-size
η > 0 (also known as the learning rate):
new old
h=−∇WW L.
()
1.8
As illustrated in figure 1.11, −∇L directs the training process towards a descending
direction of the loss [22].
1.3.6 Applications and use cases of deep learning
Dosimetry utilizing the Monte Carlo (MC) method is widely regarded as the gold
standard and the most precise approach for calculating absorbed radiation dose in
heterogeneous materials such as human tissue. This method accurately models
fundamental physical processes within the context of patient-specific anatomy,
source and applicator geometry, and the presence of tissue and material variations
[32]. Monte Carlo simulations, however, involve the stochastic sampling of
numerous particle tracks, making them time-consuming. DL, with its capacity to
learn from large amounts of data, has the potential to approximate the complex
physics involved in radiation interactions. By training neural networks on large
datasets of simulated radiation scenarios, these models can potentially learn to
predict dose distributions with faster speed while significantly reducing the computational burden associated with Monte Carlo simulations. An example of such DL
systems is RapidBrachyDL [33], which is a 3D deep CNN. RapidBrachyDL
calculates dose distributions for high dose rate (HDR) brachytherapy, with patient
CT images and treatment plans used as inputs, while accelerating the simulation
process by 300 times.
In addition to dose prediction, deep models can also be used for predicting
possible side effects of a prescribed dose. For instance, early detection of toxicities in
radiotherapy is crucial for ensuring the safety and well-being of patients undergoing
cancer treatment. An example is the architecture proposed by Elhaminia et al which
is a CNN- and attention-based model for prediction of toxicity in pelvic radiotherapy [34], demonstrating an 80% accuracy. Investigation of attention weights in
this model provides insight on which anatomical regions are associated with high
risk of toxicity, and how dose maps impact the network’s prediction.
DL has shown significant promise in adaptive radiotherapy when applied to
various image modalities such as CT, MRI, and PET. One key application is in
image segmentation, where deep learning models can accurately delineate organs at
risk and target volumes. For example, in head and neck cancer treatment, precise
segmentation of critical structures such as the spinal cord or salivary glands is crucial
to avoid unnecessary radiation exposure and minimize side effects. Another
application is in image registration, where deep learning can align images from
different time points to track changes in tumor size and position. This is particularly
important for tumors that exhibit significant intrafractional motion, such as lung
.
1-19

Artificial Intelligence in Adaptive Radiation Therapy
tumors. By accurately registering images, clinicians can adapt the treatment plan to
ensure that the tumor receives the intended dose while sparing healthy tissue.
Despite the novel applications of deep learning models, the moderate-sized
datasets in the medical domain pose challenges for training robust ML models.
Transfer learning is a technique widely used in medical computer vision tasks with
limited datasets and computational resources. It involves utilizing a pre-trained
model, trained on a large dataset, as a starting point for a new task with a smaller
dataset. In this approach, a large number of parameters in the pre-trained model are
frozen (not updated by gradient descent) to retain learned representations, and finetuning allows updating a small subset of parameters during training to learn taskspecific features. This approach promotes faster model convergence and improved
performance by leveraging general knowledge gained from the original large dataset
[35]. VGGNet [36] is a well-known CNN, pre-trained on an extensive image
classification task, ImageNet [37], which comprises 1000 classes. This pre-training
facilitates the deployment of VGGNet in segmentation tasks, a common requirement in many medical imaging applications [38].
1.4 Model training and evaluation
1.4.1 Hyperparameters
Hyperparameters are external configuration settings that are not learned from the
data but are set prior to the training process (by an ML expert). These parameters
play a crucial role in determining the model’s performance. In kNNs, for instance,
the primary hyperparameter is the k itself, as well as the choice for the distance
metric. Other examples of a hyperparameter are the maximum depth of a decision
tree, the minimum number of samples required to split an internal node in a decision
tree, and for random forests, the number of estimators (decision trees).
For deep learning models, there are several key hyperparameters that play a
crucial role in determining their performance and behavior. The architecture-related
hyperparameters include the number of layers, the size of each layer (number of
neurons or units), the activation function used in each layer (e.g. ReLU, sigmoid,
tanh), and the type of layers used (e.g. dense, convolutional, recurrent). These
hyperparameters collectively define the model’s capacity to learn complex patterns
and representations from the data. Additionally, hyperparameters such as the
learning rate (
to the training process.
Hyperparameters are tuned during the training process, allowing the adjustment
of these external configuration settings to achieve optimal results. This iterative
tuning helps to find a middle ground between model complexity and generalization,
enhancing the overall performance of the ML model.
), batch size, the type of optimizer (e.g. Adam, SGD) are related
1.4.2 Data split
When humans learn, the process of acquiring knowledge and skills involves both
training and testing. Just as students need to undergo exams to demonstrate their
understanding and proficiency, ML models require testing to evaluate their
1-20
Соседние файлы в папке Библиотека им академика М.И. Перельмана
