Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5525_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Editor biographies
- •Yi Wang
- •X. Sharon Qi
- •List of contributors
- •1.2.3 Feature engineering and representation
- •1.2.4 Linear separability
- •1.2.5 Classical models
- •1.1 A brief introduction to AI
- •1.2 Machine learning basics
- •1.2.1 Learning paradigms
- •1.3 Artificial neural networks
- •1.3.1 Feed-forward neural networks
- •1.3.2 Recurrent neural networks
- •1.3.3 Convolutional neural networks
- •1.3.4 Attention
- •1.3.5 Training neural networks
- •1.3.6 Applications and use cases of deep learning
- •1.4 Model training and evaluation
- •1.4.1 Hyperparameters
- •1.4.2 Data split
- •1.4.3 Evaluation metrics
- •1.5 Generative models
- •1.5.1 Generative adversarial networks
- •1.5.2 Diffusion models
- •1.5.3 Applications and use cases
- •1.6 Ethical consideration and bias
- •1.6.1 Transparency and explainability
- •1.6.2 Bias and fairness
- •1.6.3 Data privacy violation
- •1.6.4 Risk and misuse
- •1.7 Summary
- •References
- •2.1 Introduction
- •2.1.2 Staff roles in radiation therapy
- •2.2 Overview of AI in radiation therapy
- •2.2.1 Patient evaluation and dose prescription
- •2.2.2 Treatment simulation
- •2.2.3 Contouring
- •2.2.4 Treatment planning
- •2.2.5 Quality assurance
- •2.2.6 Treatment delivery
- •2.2.7 Response assessment and toxicity management
- •2.3 Summary
- •3.1 Introduction
- •3.1.1 Introduction of clinical decision making and AI
- •3.1.2 The role of AI in clinical decision making
- •3.2 AI algorithms for clinical decision making
- •3.2.2 Radiomics
- •3.2.3 Data integration by AI
- •3.2.4 Interpretability of AI models
- •3.3 Application of AI in clinical decision making
- •3.3.1 Diagnosis and disease phenotyping
- •3.3.2 Personalized treatment
- •3.3.3 Treatment outcome and prognosis prediction
- •3.4 Challenges and future directions of AI in clinical decision making
- •3.4.1 Challenges and concerns
- •3.4.2 Future directions
- •3.5 Summary
- •References
- •4.1 Introduction
- •4.2 Imaging for treatment planning
- •4.2.1 CT simulation
- •4.2.2 4D-CT
- •4.2.3 PET/CT
- •4.2.4 MRI
- •4.3 Imaging for treatment guidance
- •4.3.1 Portal imaging
- •4.3.2 CBCT
- •4.3.3 CT-on-rail and CT-linac
- •4.3.4 MR-linac
- •4.3.5 PET-linac
- •4.4 Imaging for motion management
- •4.4.1 ExacTrac
- •4.4.2 Varian triggered imaging
- •4.4.3 4D-CBCT
- •4.4.4 Cine MRI
- •4.4.5 4D-MRI
- •4.4.6 Surface imaging
- •4.5 Imaging for treatment assessment
- •4.5.1 Contrasted CT
- •4.5.2 PET/CT
- •4.5.3 Functional MRI
- •4.6 Summary
- •5.1 Introduction to big data in radiation oncology
- •5.1.1 Overview of big data
- •5.1.2 Sources of big data in radiation oncology
- •5.1.3 Big data and AI in radiation oncology
- •5.2 Big data lifecycle in radiation oncology
- •5.2.1 Data aggregation and storage
- •5.2.2 Data sharing and security
- •5.4 The application of big data in radiation oncology
- •5.4.1 Medical image segmentation
- •5.4.2 Automatic treatment planning
- •5.4.3 Treatment response prediction
- •5.4.4 Quality assurance and patient safety
- •5.4.5 Clinical decision support
- •5.2.3 Data visualization
- •5.2.4 Knowledge creation and implementation
- •5.2.5 Data archiving and deletion
- •5.3 Big data analytics with AI
- •5.3.1 Data processing and integration
- •5.3.2 AI modeling
- •5.5 Challenges and future perspectives
- •5.6 Summary
- •Reference
- •6.1 The road to ART
- •6.1.1 3D conformal radiotherapy (3DCRT)
- •6.1.2 Intensity modulated radiotherapy (IMRT)
- •6.1.3 Image-guided radiotherapy (IGRT)
- •6.1.4 Adaptive radiotherapy (ART)
- •6.2 ART workflow and implementation
- •6.2.2 Current practice
- •6.2.3 Clinical impact
- •6.3 Considerations for implementing online ART
- •6.3.1 Time as a limiting factor
- •6.3.2 Implications for fast and reliable re-planning
- •6.3.4 Clinical considerations
- •6.4 Summary
- •7.1 Components of ART workflow
- •7.1.1 Simulation
- •7.1.2 Pre-planning
- •7.1.3 Online imaging and daily re-planning
- •7.1.4 Quality assurance
- •7.2 AI-driven ART
- •7.2.1 Simulation
- •7.2.2 Pre-planning
- •7.2.3 AI for delivery
- •7.3 Outlook and future directions
- •7.3.1 Real-time ART with AI
- •7.3.2 Dose escalation and functional adaption with AI
- •7.4 Summary
- •References
- •8.1 Introduction
- •8.2 Synthetic CT: deep learning methods
- •8.2.1 Conventional methods
- •8.2.2 U-Net
- •8.2.3 Generative adversarial networks
- •8.2.4 Denoising diffusion probabilistic model
- •8.3 Synthetic CT from CBCT
- •8.3.1 Noise and artifact reduction
- •8.3.2 Online dose calculation
- •8.3.3 Online image segmentation
- •8.4 Synthetic CT from MRI
- •8.4.1 Synthetic image accuracy
- •8.4.2 Dose calculation in MR-only radiation therapy
- •8.4.3 PET attenuation correction
- •8.4.4 Image registration
- •8.5 Discussion and outlook
- •8.6 Summary
- •References
- •9.1 AI-based image registration and segmentation for ART
- •9.1.1 Adaptive radiation therapy
- •9.2 Artificial intelligence
- •9.2.1 What is machine learning?
- •9.2.2 What is deep learning?
- •9.3 Deep learning: the basic components
- •9.3.1 Convolutional neural networks: looking at the picture
- •9.3.2 Pooling layers: keeping what matters most
- •9.3.3 Fully connected (dense) layers: bringing it all together
- •9.3.4 Activations
- •9.3.5 Loss: driving the model
- •9.3.6 Auto-encoders: remove the noise
- •9.3.7 Supervised versus unsupervised learning
- •9.3.8 Pre-trained convolutional neural networks
- •9.4 Image registration: bringing two images together
- •9.4.1 Registration similarity metrics
- •9.4.2 Types of registrations
- •9.5 AI-based image registration
- •9.5.1 Supervised learning
- •9.5.2 Unsupervised learning
- •9.5.3 Registration in ART
- •9.5.4 Commonalities in architectures
- •9.6 Image segmentation
- •9.6.1 Introduction: coloring by the numbers
- •9.6.2 Segmentation networks
- •9.6.3 Best practices
- •9.7 Summary
- •10.1 Introduction
- •10.1.1 Overview of chapter content
- •10.2 The landscape of AI-assisted dose prediction
- •10.2.1 Traditional machine learning for dose prediction
- •10.2.2 Deep learning-based dose prediction
- •10.2.3 Challenges in AI-assisted dose prediction
- •10.3 Re-planning workflows powered by AI
- •10.3.1 Deep learning for re-planning pipelines
- •10.4 Future directions of AI-assisted dose prediction and re-planning
- •10.5 Summary
- •11.1 Introduction
- •11.2.1 Imaging-based motion monitoring
- •11.2.2 Delivery system actions
- •11.2.3 Challenges for real-time ART implementation
- •11.3 AI in real-time ART workflows
- •11.3.1 Improving intrafraction motion monitoring through AI
- •11.3.2 Mitigating system latency through AI
- •11.4 AI for ART delivery: future directions
- •11.4.1 Management of non-respiratory motion
- •11.4.2 Training AI models with small or unpaired datasets
- •11.4.4 Biology-guided ART delivery
- •11.5 Summary
- •References
- •12.1 Introduction
- •12.2 Patient QA
- •12.2.1 Pre-planning QA
- •12.2.2 Pre-treatment plan QA
- •12.2.3 On-treatment QA
- •12.3 Treatment delivery systems and instruments
- •12.3.1 Machine commissioning
- •12.3.2 Machine QA
- •12.3.3 Dosimetry tool QA
- •12.4 Summary
- •References
- •13.1 Data resources for response modeling in radiotherapy
- •13.1.1 Clinical data
- •13.1.2 Imaging (radiomics)
- •13.1.3 Treatment planning (dosiomics)
- •13.1.4 Multiomics
- •13.2 Radiotherapy treatment outcome modeling
- •13.2.1 TCP/NTCP in radiotherapy
- •13.2.2 Clinical outcomes versus PROs
- •13.2.3 Machine learning response prediction
- •13.2.4 Explainability of ML response models
- •13.2.5 Sample use cases
- •13.3 AI response-based adaptive radiotherapy
- •13.3.1 Requirements and challenges
- •13.3.2 Prediction versus treatment optimization
- •13.3.3 Sample use cases
- •13.4 Challenges and recommendations
- •13.5 Summary
- •Acknowledgments
- •References
- •14.1 Overview of challenges in AI-driven ART
- •14.2 Data challenges
- •14.2.1 Data availability
- •14.2.2 Data quality
- •14.2.3 Data privacy
- •14.3 Technical challenges
- •14.3.2 Model robustness and generalizability
- •14.3.3 Model explainability and interpretability
- •14.4 Challenges associated with online and real-time workflows
- •14.4.1 Image quality
- •14.4.2 Dose calculation
- •14.4.3 Real-time ART
- •14.5 Operational challenges
- •14.5.1 Clinical validation
- •14.5.3 Staff training
- •14.5.4 User experiences
- •14.5.5 Quality management program
- •14.5.6 Financial challenges
- •14.6 Ethical, regulatory, and legal challenges
- •14.6.1 Ethical issues
- •14.6.2 Regulatory and legal issues
- •14.7 Summary
- •References
- •15.1 Clinical considerations for CT-based offline ART
- •15.1.1 Patient and site selection
- •15.1.2 Re-simulation
- •15.1.3 Re-planning
- •15.1.4 Plan summation and evaluation
- •15.1.6 Limitations and future directions
- •15.2 Clinical considerations for CBCT/CT-based online ART
- •15.2.2 Patient and site selection
- •15.2.3 Simulation
- •15.2.4 Pre-planning review
- •15.2.5 Reference planning
- •15.2.9 Limitations and future directions
- •15.3 Summary
- •References
- •16.1 Introduction
- •16.2 Overview of MRI-guided ART systems
- •16.3 MRI-guided ART workflow
- •16.4 AI applications for MRI-guided ART
- •16.4.1 Synthetic CT generation
- •References
- •16.4.2 Auto-segmentation
- •16.4.3 Image registration
- •16.4.4 Others
- •16.4.5 Future AI development and implementation
- •16.5 Summary
- •17.1 Functional PET-guided ART
- •17.1.1 PET-based functional imaging overview
- •17.1.2 From anatomy to function: the power of PET in radiation therapy
- •17.1.5 Conclusions and future prospects
- •17.2 Functional MRI-guided ART
- •17.2.1 From anatomy to function: the power of functional MRI in radiation therapy
- •17.2.4 Conclusion and future prospects
- •17.3 Summary
- •References
- •18.1 Proton ART
- •18.1.1 Clinical context and necessity
- •18.1.2 Patient populations
- •18.1.4 Rationale for AI in proton ART
- •18.2 AI in proton ART
- •18.2.1 Imaging
- •18.2.2 Deformable and rigid registration
- •18.2.3 Contour propagation
- •18.2.4 Dose calculations
- •18.2.5 Plan optimization
- •18.2.6 Other developments
- •18.3 Implementation of adaptive proton therapy
- •18.4 Summary
- •References
- •19.1 Designing clinical trials with AI
- •19.1.1 The essential role of clinical trials
- •19.1.2 Trial protocols and methodologies
- •19.1.3 AI-driven clinical trial design and execution
- •19.1.4 Incorporation of digital twins (DTs) in clinical trials
- •19.2 Implementation of AI in ongoing clinical trials
- •19.2.1 Integration with existing clinical trial frameworks
- •19.2.2 Quality assurance, compliance, and standardization
- •19.3 Case studies of AI in adaptive radiotherapy trials
- •19.3.1 Overview of guidance for advanced radiotherapy in clinical trials
- •19.3.2 AI in the radiotherapy clinical trial quality assurance processes
- •19.4 Ethical and regulatory considerations
- •19.4.1 Patient consent and data privacy
- •19.4.2 Bias, fairness, and transparency
- •19.4.3 Regulatory guidelines and compliance
- •19.5 Future directions and challenges
- •19.5.1 Emerging technologies and techniques
- •19.5.2 Alternative strategies
- •19.6 Conclusion
- •19.7 Summary
- •References
- •20.1 Risk management
- •20.1.1 Prospective risk assessments
- •20.1.2 Root cause analysis

Artificial Intelligence in Adaptive Radiation Therapy
9.2.2 What is deep learning?
Deep learning can be defined as a subset within machine learning. Whereas machine
learning combines several features directly into an output, ‘deep’ learning is often
associated as having several intermediary or ‘hidden’ layers. A single hidden layer,
as shown in figure 9.1, would most appropriately be referred to as a ‘shallow’
network, with ‘deep’ requiring two or more hidden layers. The benefit of multiple
hidden layers is the ability to represent predictions beyond simple linear regression
when non-linear activation functions are used (activation functions are spoken of
more in the following sections). The downside to the increase in the prediction
capabilities of the model is the decrease in interpretability of the model predictions.
Deep learning is often cited as being a ‘black box’, although this does not mean it is
impossible to crack open the box and take a peek inside now and again.
9.3 Deep learning: the basic components
While deep learning has evolved considerably over the past decades, most models
can be broken down into a combination of several key components. Having a solid
understanding of each part will help one understand better how more complex
models are generated in the future.
9.3.1 Convolutional neural networks: looking at the picture
Convolutional layers form the bedrock of deep learning image classification,
segmentation, and registration models. The basis of these three tasks all relies on
the same thing: the computer needs to understand what is occurring within an image.
Humans, as well as most animals, have the benefit of eyes which enable us to capture
a large amount of information within our field of view. For a computer this can be a
relatively daunting task.
The first thing which needs to be defined is the computer’s receptive field of view.
A convolution is very simply the creation of a new matrix by multiplication and
resultant addition of elements on an image. A visual representation of this concept is
shown in figure 9.2. Here, the input image is convolved with a simple three-by-three
kernel to create a resultant feature image.
If the kernel were a vertical line surrounded by zeros, the output feature map
would show where vertical lines are present in the image.
One of the downsides of convolutions is that they are locally dependent. This means
that since a kernel size is fixed, images of varying zoom or scale might not properly
align with a previously defined kernel. A natural question which arises from this, is
why not push towards large kernels which are able to cover large portions of the input
image? There are two main concerns with this approach: (i) the larger the kernel, the
more computational resources are required to perform the convolution and addition
and (ii) larger kernels have more variables present, where an n by m matrix will have n
× m + 1 variables. A three-by-three kernel has nine variables (plus a bias to make ten
total), while a nine-by-nine kernel has 81 variables (plus a bias to make 82).
9-4

Artificial Intelligence in Adaptive Radiation Therapy
Figure 9.2. General representation of a convolution. The input image (left) convolved with the 3 × 3 kernel
(middle) creates a feature map (right).
Figure 9.3. Example of a two-by-two max-pooling layer.
This modest increase in kernel size creates significantly more variables to train. A
remedy to this local dependency issue comes in the form of pooling layers.
9.3.2 Pooling layers: keeping what matters most
Pooling layers are a way of taking only the most important aspect of a generated
feature map. While they come in a variety of forms, a very commonly used pooling
layer is max-pooling. Pooling layers are like convolutional layers in that they are a
matrix, typically of size two-by-two for two-dimensional images, but rather than
multiplying and adding, a max-pooling layer will simply take the maximum value
within the defined matrix size, figure 9.3.
One of the most immediate implications of this is that the feature map itself has
now decreased in size by a factor of 2 in each direction. This can either be seen as
beneficial or detrimental, depending on the desired output. From a positive
perspective, max-pooling layers have multiple benefits. The first of these benefits
is that the most intense, or important, feature is maintained, while presumably
less important features are removed. Second, subsequent convolutions on the
9-5

Artificial Intelligence in Adaptive Radiation Therapy
post-max-pooled image require significantly less memory—as stated before, the
feature map has been reduced in size. A 2 × 2 max-pooled image will be 1/4 the
original image size, meaning less memory is required for the convolutional
operation. Third, subsequent convolutions now cover a relative area that is larger
than the original image size. A three-by-three kernel on an image with resolution of
1 mm × 1 mm per voxel would have a receptive field size of 3 mm by 3 mm. After
max-pooling, the resolution of the image is 2 mm × 2 mm per voxel, so a three-bythree kernel has a relative receptive field of view of 6 mm by 6 mm. This is
particularly useful in images where the scale/zoom of the image can vary. If the
model is designed to identify that a dog is present within an image, it should not
matter how large or small the dog presented is.
Despite these benefits for image classification, if the user wishes to identify
individual voxels relating to a class (segmentation), the pooling layers can lead to a
significant decrease in resolution. A representation of max-pooling followed by
bilinear up sampling to the original image size is shown in figure 9.4. Note how the
fine resolution text is quickly lost by the fourth pooling, while general features such
sa color can be maintained for multiple layers. Maintaining the higher-resolution
information for later decision making is a main factor in the popularity of the U-Net
style architecture discussed later in this chapter.
One thing we have yet to cover is how these features can lead to a prediction. A
model needs a method of combining these features into a final prediction; this
combination of features is defined as a fully connected layer.
Figure 9.4. Illustration of multiple max-pooling layers and bilinear resampling to illustrate how fine resolution
information (text) is quickly lost, while large information (general colors) remains.
9-6

)
Artificial Intelligence in Adaptive Radiation Therapy
9.3.3 Fully connected (dense) layers: bringing it all together
Fully connected, or dense, layers can be imaged as the opposites of convolutional
layers. While convolutional layers are locally dependent, the fully connected layer
will combine every aspect of the input into an output, figure 9.5. The number of
inputs and outputs in a dense layer can vary and are user defined. The mathematical
value of each output is
out
i
out
=×+
Fh W B
(
i
−
i
1
ii
, where F is an activation function,
W is the weighting matrix, and B is the bias.
Within convolutional neural networks there can be any number of convolutions
and pooling layers. These maps can be represented as several n by m matrices, where
n and m are likely smaller than the original image dimensions. In order to feed these
features into a dense layer they must first be ‘flattened’. This can be imaged as taking
the feature image and laying it into a single vector: converting the n-by-m matrix
into an n × m vector. These features can then be used in the final classification. For
the example in figure 9.6, we might say that if ears, whiskers, and paws are present,
then a dog is present in the image.
The low-level and abstract features of lines, curves, or fuzziness are most likely
to appear early in the architecture. Higher level features, such as ears and whiskers,
can be imagined arising deeper, as a combination of overlapping low-level
features. Finally, all of these features can be combined together after flattening
to identify if a dog is present somewhere in the image. Note that our output does
nothing to identify where the dog is present. Furthermore, if the model is truly
identifying whether ears, whiskers, and paws are present, it very likely confuses a
cat for a dog.
Figure 9.5. A fully connected, or ‘dense’ layer. Output features are a function of weights from each input
feature.
9-7

Artificial Intelligence in Adaptive Radiation Therapy
Figure 9.6. Generalized example of a convolutional neural network for the identification of a dog. The first
convolutional layers are very low-level; lines, edges, fuzziness. Further down, these features combine to
identify ears, paws, or whiskers, before finally feeding into the decision-making process of a fully connected
layer.
9.3.4 Activations
Throughout the previous sections we have discussed features, convolutions, and
dense layers, but have purposefully been neglecting a very important step in the
process. Part of the true power in deep learning comes not only from these critical
parts, but also from the final step after each: the activation function.
9.3.4.1 Linear activations
The activation function is a deceivingly simple thing. A simple statement can
summarize its purpose: using a defined function, map an input value to an output
value. One of the most basic forms of activation is a linear activation: f(x) = mx + b,
which can have a scaling (m) and bias (b) impact on the outputs. While linear
activations can be a useful way of scaling and shifting values, stacking multiple
linearly activated layers does not add any new information to the system. To
demonstrate this, imagine two linear feed-forward layers, where x is the original
values:
=+ymxb
11
1
=+ymyb.
2
2
2
1
The entirety of this network y2can be defined and simplified as
=++→++→+ymmxbbmmxmbbMxB,
()
21 1 2 21 21 2
2
where m2× b1+ b2can be combined as a single bias, B, and m2× m1can be a single
scalar, M. This means that regardless of how many linearly activated layers are
placed together, they are only capable of expressing linear relationships. For
modeling something with non-linear such as gravity, they would inevitably fall
short.
9-8

Artificial Intelligence in Adaptive Radiation Therapy
9.3.4.2 Non-linear activations
The power of multiple layers becomes more apparent as we transition towards nonlinear activation functions. Stacking multiple non-linearly activating layers can add
increasing complexity to the model’s predictive abilities. For example, if the
activation function were f(x) = mx
representational ability of several polynomials: x
2
+ gx + c, two such layers would have the
4
, x3, x2, x, and bias.
A set of four commonly seen non-linear activation functions is shown in
figure 9.7.
Figure 9.7. List of common activation functions, sigmoid, rectified linear unit (ReLU), leaky ReLU, and
exponential linear unit (ELU).
9.3.4.3 Sigmoid activation
The sigmoid activation function, while not the most commonly seen in the
intermediate layers of deep learning architectures, is a common final activation
function for a binary prediction. This function is commonly expressed as being
synonymous with neurons firing within the brain. Given a number of inputs, if the
value exceeds some critical threshold the neuron will fire. A sigmoid activation is an
appropriate activation for a binary classification system such as our dog model
shown in figure 9.6, where ‘0’ would indicate no dog, and ‘1’ would indicate the
presence of a dog.
In the case of multiple classification options, the sigmoid activation is often
exchanged for the soft-max activation. This provides a probability of a class scaled
to the sum of all class’s probability:
z
c
==
ycx
∣)
(
C
∑
=
j
e
,
z
j
e
1
where c is a particular class and C is the total number of classes.
9.3.4.4 Non-linear activations: rectified linear
The rectified linear unit activation (ReLU) was one of the first and most popular
activation functions used in modern deep learning architectures. The function is
relatively simple: if x is less than 0, the returned value is 0, otherwise the returned
value is x. The ‘leaky’ ReLU and exponential linear unit (ELU) are both variations
9-9

Artificial Intelligence in Adaptive Radiation Therapy
on the ReLU, where the differences focus on what occurs when values less than 0 are
presented to the function. One of the major issues with ReLU was ‘dying’ nodes,
meaning if a kernel initialized at a value less than 0, all gradient through the node
‘died’. The leaky and ELU activation allow at least some gradient.
An important question to raise at this point is ‘Why use these non-linear
activations?’ The answer to this lies in the most interesting part of deep learning:
loss and back-propagation.
9.3.5 Loss: driving the model
When making a prediction, we need some method of identifying the ‘correctness’ of
the model prediction. This metric is referred to as the loss. There are several ways of
expressing a model loss, although, for mathematical reasons, it is desired that this
loss is something that the model is attempting to minimize. An in-depth discussion of
back-propagation and gradient descent is beyond the scope of this chapter, please
refer to other texts discussed at the beginning of this chapter [25, 26] for a detailed
explanation.
9.3.6 Auto-encoders: remove the noise
If one were to be told to draw a bicycle, remove while playing a game of
Telestrations, they would likely draw out two simple circles as wheels, a frame
connecting them, and a set of handlebars. Perhaps peddles and spokes on the wheels
would be included for the more artistically inclined, and yet, this drawing of a
bicycle would be relatively far removed from the appearance of an actual bicycle.
While there are hundreds, if not thousands, of different types of bicycles, most
people would be able to review this crude drawing and conclude that the image is
supposed to be that of a bicycle. The human brain is amazing in this regard: a
bicycle, in all its many shapes and forms, can be condensed down to just a few simple
aspects. This ‘condensing of important features or removing of unnecessary parts
(noise) is the basis of auto-encoders [27].
The goal of an auto-encoder is to simplify, or compress, a set of inputs into its
most important aspects, figure 9.8. By reducing the number of features and
accurately reproducing the input, the model is able to identify what are the most
important aspects of the input. This can similarly be used as a form of principal
component analysis (PCA) [28] and is a common building block of both supervised
and unsupervised learning models. The so-called ‘stacked’ auto-encoder is simply a
combination of auto-encoders with significantly noted improvement in model
predictive capability [29]. The training of these models is often pieced together,
with each layer trained separately before being combined and fine-tuned in the final
model.
9.3.7 Supervised versus unsupervised learning
Distinctions between supervised and unsupervised learning can be simplified as a
statement about data labeling. If the user understands what the desired output
should be, e.g. ‘is a stop sign present in this image?’ and uses those labels to train a
9-10

Artificial Intelligence in Adaptive Radiation Therapy
Figure 9.8. Basic example of an auto-encoder.
model, we are operating under ‘supervised’ learning. However, if we simply had
many pictures of stop signs and other signs with no ‘ground truth’ labeling, we
would require an unsupervised learning technique. This is not to say that unsuper-
vised learning does not require an equal amount of work for training as a supervised
learning model.
9.3.8 Pre-trained convolutional neural networks
Much of the previous discussion revolves around feature extraction (convolutions),
refinement (pooling), and combination of those features (dense layers), all with the
understanding that relevant features are being identified by the model. However, this
is not a guarantee. When first creating a convolutional network, such as one for
image classification by the visual geometry group 16 (VGG-16) [30], figure 9.9, every
convolutional kernel is going to be a randomly generated distribution of numbers.
Deep learning researchers all suffer from this issue. Updating these random
kernels based on correct predictions is a labor intensive and sensitive process that
occupies most, if not all, of a researcher’s time and effort when beginning a new
classification process. The VGG-16 model was trained on the ImageNet challenge,
which contains more than 14 million images. This is far above what most medical
researchers have available for training their own models. However, this can still be
used to our advantage with the concept of using this pre-trained model.
9-11

Artificial Intelligence in Adaptive Radiation Therapy
Figure 9.9. Representation of the visual geometry group (VGG-16) classification architecture.
The fundamental basis of using a pre-trained model is that certain low-level
features are ubiquitous. As humans, we are able to identify many different objects all
of which are built on the same fundamental pieces: lines, curves, edges, etc. We are
then able to take this pre-trained model that has been trained on thousands of
different images and utilize everything before the fully connected layers as a feature
extraction network. This saves potentially countless hours of training; by freezing
the convolution layers, we can simply train the end classification. Note: Even after a
new classification model is trained, it is common to un-freeze the earlier layers in a
process of fine tuning.
9.4 Image registration: bringing two images together
Throughout the course of a patient’s care there is a potential for multiple forms of
medical imaging to be acquired for various purposes, e.g. defining tumor boundaries, identifying surrounding critical normal tissue, and characterizing tumor and
normal tissue function. For example, a magnetic resonance image (MRI) can exhibit
exquisite contrast of soft tissues in ways that computed tomography (CT) scans
cannot, and an FDG positron emission tomography (PET) scan is able to identify
regions of metabolic hyperactivity. Currently, standard practice within radiation
oncology includes a CT scan enabling the visualization of a desired treatment site
and enabling radiation dose calculation, although there is significant research into
MRI based generation of electron density.
An ability to refer to multiple imaging modalities to evaluate tumor extent or
normal tissues is vital to the decision-making and treatment planning process.
Unfortunately, these PET, MRI, and CT scans are rarely acquired at the same time
or with exactly consistent positioning of the patient. These changes in time,
positioning, and/or anatomy have led to the necessity of tools to align different
modalities/images about a region of interest in the anatomy. This aligning of
different modalities or images is referred to as image registration.
This is not to say that image registration is only important during the planning
process. Accurate alignment of the planning images with treatment imaging (conebeam CT, fan-beam CT, MRI) is equally vital to a safe and successful treatment.
9-12

Artificial Intelligence in Adaptive Radiation Therapy
With ART, it is often vital to understand the distribution of previously delivered
radiation for guidance in future planning. This understanding relies on accurate
representation and registration of the new imaging with previous plans. It is equally
important to understand how a registration is evaluated and functions, the
‘registration metric’, per AAPM Task Group 132 [31]. After all, when registering
two images, how does one know when the registration is ‘good’?
There are multiple ways in which a registration can be evaluated. Often, the
registration is focused on a particular region of clinical relevance. Identifying the
region of relevance is the first and most important aspect of image registration.
9.4.1 Registration similarity metrics
There are broadly two ways of evaluating/driving a registration: intensity-based and
feature-based.
9.4.1.1 Intensity based registration
Intensity based registration can be defined as a comparison of intensities between the
two images, either as a whole, or about a defined sub-region of interest in the image
[32–34]. When operating between two images of the same modality, simple
comparisons of the intensity values between the two can offer a reasonable
evaluation of the registration, possible as the sum of squared differences (SSD)
[35, 36]:
N
1
=−
()
IISSD
XY
∑
N
=
i
1
2
ii
,
where I is the intensity and subscript X and Y refer to two images across the total
number of voxels being evaluated. While this metric can be very useful, it suffers
when the voxel intensities throughout the two images vary (a multi-phase contrast
enhanced CT). To this end, the normalized cross-correlation (NCC) [36–38]isa
wonderful metric, accounting for the differences in intensity between the two images,
although assuming that relatively high and low intensities correlate with each other:
n
−−
xxyy
()()
i
∑
=
i
=
CC .
1
n
−−
xx yy
()()
i
∑∑
==
i
1
i
n
2
i
1
2
i
An important note from both methods is that they require the intensities between the
images to be correlated: a bright spot in one image corresponds to a bright spot in
the other, etc. For multi-modality images (CT to MR), this is not a guarantee.
A commonly used metric for the registration of images of different modalities is
mutual information [39, 40]. This metric is based on the mutual probabilities
between the two images and has no reliance on the absolute intensity of the images:
9-13
Соседние файлы в папке Библиотека им академика М.И. Перельмана
