Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5525_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Editor biographies
- •Yi Wang
- •X. Sharon Qi
- •List of contributors
- •1.2.3 Feature engineering and representation
- •1.2.4 Linear separability
- •1.2.5 Classical models
- •1.1 A brief introduction to AI
- •1.2 Machine learning basics
- •1.2.1 Learning paradigms
- •1.3 Artificial neural networks
- •1.3.1 Feed-forward neural networks
- •1.3.2 Recurrent neural networks
- •1.3.3 Convolutional neural networks
- •1.3.4 Attention
- •1.3.5 Training neural networks
- •1.3.6 Applications and use cases of deep learning
- •1.4 Model training and evaluation
- •1.4.1 Hyperparameters
- •1.4.2 Data split
- •1.4.3 Evaluation metrics
- •1.5 Generative models
- •1.5.1 Generative adversarial networks
- •1.5.2 Diffusion models
- •1.5.3 Applications and use cases
- •1.6 Ethical consideration and bias
- •1.6.1 Transparency and explainability
- •1.6.2 Bias and fairness
- •1.6.3 Data privacy violation
- •1.6.4 Risk and misuse
- •1.7 Summary
- •References
- •2.1 Introduction
- •2.1.2 Staff roles in radiation therapy
- •2.2 Overview of AI in radiation therapy
- •2.2.1 Patient evaluation and dose prescription
- •2.2.2 Treatment simulation
- •2.2.3 Contouring
- •2.2.4 Treatment planning
- •2.2.5 Quality assurance
- •2.2.6 Treatment delivery
- •2.2.7 Response assessment and toxicity management
- •2.3 Summary
- •3.1 Introduction
- •3.1.1 Introduction of clinical decision making and AI
- •3.1.2 The role of AI in clinical decision making
- •3.2 AI algorithms for clinical decision making
- •3.2.2 Radiomics
- •3.2.3 Data integration by AI
- •3.2.4 Interpretability of AI models
- •3.3 Application of AI in clinical decision making
- •3.3.1 Diagnosis and disease phenotyping
- •3.3.2 Personalized treatment
- •3.3.3 Treatment outcome and prognosis prediction
- •3.4 Challenges and future directions of AI in clinical decision making
- •3.4.1 Challenges and concerns
- •3.4.2 Future directions
- •3.5 Summary
- •References
- •4.1 Introduction
- •4.2 Imaging for treatment planning
- •4.2.1 CT simulation
- •4.2.2 4D-CT
- •4.2.3 PET/CT
- •4.2.4 MRI
- •4.3 Imaging for treatment guidance
- •4.3.1 Portal imaging
- •4.3.2 CBCT
- •4.3.3 CT-on-rail and CT-linac
- •4.3.4 MR-linac
- •4.3.5 PET-linac
- •4.4 Imaging for motion management
- •4.4.1 ExacTrac
- •4.4.2 Varian triggered imaging
- •4.4.3 4D-CBCT
- •4.4.4 Cine MRI
- •4.4.5 4D-MRI
- •4.4.6 Surface imaging
- •4.5 Imaging for treatment assessment
- •4.5.1 Contrasted CT
- •4.5.2 PET/CT
- •4.5.3 Functional MRI
- •4.6 Summary
- •5.1 Introduction to big data in radiation oncology
- •5.1.1 Overview of big data
- •5.1.2 Sources of big data in radiation oncology
- •5.1.3 Big data and AI in radiation oncology
- •5.2 Big data lifecycle in radiation oncology
- •5.2.1 Data aggregation and storage
- •5.2.2 Data sharing and security
- •5.4 The application of big data in radiation oncology
- •5.4.1 Medical image segmentation
- •5.4.2 Automatic treatment planning
- •5.4.3 Treatment response prediction
- •5.4.4 Quality assurance and patient safety
- •5.4.5 Clinical decision support
- •5.2.3 Data visualization
- •5.2.4 Knowledge creation and implementation
- •5.2.5 Data archiving and deletion
- •5.3 Big data analytics with AI
- •5.3.1 Data processing and integration
- •5.3.2 AI modeling
- •5.5 Challenges and future perspectives
- •5.6 Summary
- •Reference
- •6.1 The road to ART
- •6.1.1 3D conformal radiotherapy (3DCRT)
- •6.1.2 Intensity modulated radiotherapy (IMRT)
- •6.1.3 Image-guided radiotherapy (IGRT)
- •6.1.4 Adaptive radiotherapy (ART)
- •6.2 ART workflow and implementation
- •6.2.2 Current practice
- •6.2.3 Clinical impact
- •6.3 Considerations for implementing online ART
- •6.3.1 Time as a limiting factor
- •6.3.2 Implications for fast and reliable re-planning
- •6.3.4 Clinical considerations
- •6.4 Summary
- •7.1 Components of ART workflow
- •7.1.1 Simulation
- •7.1.2 Pre-planning
- •7.1.3 Online imaging and daily re-planning
- •7.1.4 Quality assurance
- •7.2 AI-driven ART
- •7.2.1 Simulation
- •7.2.2 Pre-planning
- •7.2.3 AI for delivery
- •7.3 Outlook and future directions
- •7.3.1 Real-time ART with AI
- •7.3.2 Dose escalation and functional adaption with AI
- •7.4 Summary
- •References
- •8.1 Introduction
- •8.2 Synthetic CT: deep learning methods
- •8.2.1 Conventional methods
- •8.2.2 U-Net
- •8.2.3 Generative adversarial networks
- •8.2.4 Denoising diffusion probabilistic model
- •8.3 Synthetic CT from CBCT
- •8.3.1 Noise and artifact reduction
- •8.3.2 Online dose calculation
- •8.3.3 Online image segmentation
- •8.4 Synthetic CT from MRI
- •8.4.1 Synthetic image accuracy
- •8.4.2 Dose calculation in MR-only radiation therapy
- •8.4.3 PET attenuation correction
- •8.4.4 Image registration
- •8.5 Discussion and outlook
- •8.6 Summary
- •References
- •9.1 AI-based image registration and segmentation for ART
- •9.1.1 Adaptive radiation therapy
- •9.2 Artificial intelligence
- •9.2.1 What is machine learning?
- •9.2.2 What is deep learning?
- •9.3 Deep learning: the basic components
- •9.3.1 Convolutional neural networks: looking at the picture
- •9.3.2 Pooling layers: keeping what matters most
- •9.3.3 Fully connected (dense) layers: bringing it all together
- •9.3.4 Activations
- •9.3.5 Loss: driving the model
- •9.3.6 Auto-encoders: remove the noise
- •9.3.7 Supervised versus unsupervised learning
- •9.3.8 Pre-trained convolutional neural networks
- •9.4 Image registration: bringing two images together
- •9.4.1 Registration similarity metrics
- •9.4.2 Types of registrations
- •9.5 AI-based image registration
- •9.5.1 Supervised learning
- •9.5.2 Unsupervised learning
- •9.5.3 Registration in ART
- •9.5.4 Commonalities in architectures
- •9.6 Image segmentation
- •9.6.1 Introduction: coloring by the numbers
- •9.6.2 Segmentation networks
- •9.6.3 Best practices
- •9.7 Summary
- •10.1 Introduction
- •10.1.1 Overview of chapter content
- •10.2 The landscape of AI-assisted dose prediction
- •10.2.1 Traditional machine learning for dose prediction
- •10.2.2 Deep learning-based dose prediction
- •10.2.3 Challenges in AI-assisted dose prediction
- •10.3 Re-planning workflows powered by AI
- •10.3.1 Deep learning for re-planning pipelines
- •10.4 Future directions of AI-assisted dose prediction and re-planning
- •10.5 Summary
- •11.1 Introduction
- •11.2.1 Imaging-based motion monitoring
- •11.2.2 Delivery system actions
- •11.2.3 Challenges for real-time ART implementation
- •11.3 AI in real-time ART workflows
- •11.3.1 Improving intrafraction motion monitoring through AI
- •11.3.2 Mitigating system latency through AI
- •11.4 AI for ART delivery: future directions
- •11.4.1 Management of non-respiratory motion
- •11.4.2 Training AI models with small or unpaired datasets
- •11.4.4 Biology-guided ART delivery
- •11.5 Summary
- •References
- •12.1 Introduction
- •12.2 Patient QA
- •12.2.1 Pre-planning QA
- •12.2.2 Pre-treatment plan QA
- •12.2.3 On-treatment QA
- •12.3 Treatment delivery systems and instruments
- •12.3.1 Machine commissioning
- •12.3.2 Machine QA
- •12.3.3 Dosimetry tool QA
- •12.4 Summary
- •References
- •13.1 Data resources for response modeling in radiotherapy
- •13.1.1 Clinical data
- •13.1.2 Imaging (radiomics)
- •13.1.3 Treatment planning (dosiomics)
- •13.1.4 Multiomics
- •13.2 Radiotherapy treatment outcome modeling
- •13.2.1 TCP/NTCP in radiotherapy
- •13.2.2 Clinical outcomes versus PROs
- •13.2.3 Machine learning response prediction
- •13.2.4 Explainability of ML response models
- •13.2.5 Sample use cases
- •13.3 AI response-based adaptive radiotherapy
- •13.3.1 Requirements and challenges
- •13.3.2 Prediction versus treatment optimization
- •13.3.3 Sample use cases
- •13.4 Challenges and recommendations
- •13.5 Summary
- •Acknowledgments
- •References
- •14.1 Overview of challenges in AI-driven ART
- •14.2 Data challenges
- •14.2.1 Data availability
- •14.2.2 Data quality
- •14.2.3 Data privacy
- •14.3 Technical challenges
- •14.3.2 Model robustness and generalizability
- •14.3.3 Model explainability and interpretability
- •14.4 Challenges associated with online and real-time workflows
- •14.4.1 Image quality
- •14.4.2 Dose calculation
- •14.4.3 Real-time ART
- •14.5 Operational challenges
- •14.5.1 Clinical validation
- •14.5.3 Staff training
- •14.5.4 User experiences
- •14.5.5 Quality management program
- •14.5.6 Financial challenges
- •14.6 Ethical, regulatory, and legal challenges
- •14.6.1 Ethical issues
- •14.6.2 Regulatory and legal issues
- •14.7 Summary
- •References
- •15.1 Clinical considerations for CT-based offline ART
- •15.1.1 Patient and site selection
- •15.1.2 Re-simulation
- •15.1.3 Re-planning
- •15.1.4 Plan summation and evaluation
- •15.1.6 Limitations and future directions
- •15.2 Clinical considerations for CBCT/CT-based online ART
- •15.2.2 Patient and site selection
- •15.2.3 Simulation
- •15.2.4 Pre-planning review
- •15.2.5 Reference planning
- •15.2.9 Limitations and future directions
- •15.3 Summary
- •References
- •16.1 Introduction
- •16.2 Overview of MRI-guided ART systems
- •16.3 MRI-guided ART workflow
- •16.4 AI applications for MRI-guided ART
- •16.4.1 Synthetic CT generation
- •References
- •16.4.2 Auto-segmentation
- •16.4.3 Image registration
- •16.4.4 Others
- •16.4.5 Future AI development and implementation
- •16.5 Summary
- •17.1 Functional PET-guided ART
- •17.1.1 PET-based functional imaging overview
- •17.1.2 From anatomy to function: the power of PET in radiation therapy
- •17.1.5 Conclusions and future prospects
- •17.2 Functional MRI-guided ART
- •17.2.1 From anatomy to function: the power of functional MRI in radiation therapy
- •17.2.4 Conclusion and future prospects
- •17.3 Summary
- •References
- •18.1 Proton ART
- •18.1.1 Clinical context and necessity
- •18.1.2 Patient populations
- •18.1.4 Rationale for AI in proton ART
- •18.2 AI in proton ART
- •18.2.1 Imaging
- •18.2.2 Deformable and rigid registration
- •18.2.3 Contour propagation
- •18.2.4 Dose calculations
- •18.2.5 Plan optimization
- •18.2.6 Other developments
- •18.3 Implementation of adaptive proton therapy
- •18.4 Summary
- •References
- •19.1 Designing clinical trials with AI
- •19.1.1 The essential role of clinical trials
- •19.1.2 Trial protocols and methodologies
- •19.1.3 AI-driven clinical trial design and execution
- •19.1.4 Incorporation of digital twins (DTs) in clinical trials
- •19.2 Implementation of AI in ongoing clinical trials
- •19.2.1 Integration with existing clinical trial frameworks
- •19.2.2 Quality assurance, compliance, and standardization
- •19.3 Case studies of AI in adaptive radiotherapy trials
- •19.3.1 Overview of guidance for advanced radiotherapy in clinical trials
- •19.3.2 AI in the radiotherapy clinical trial quality assurance processes
- •19.4 Ethical and regulatory considerations
- •19.4.1 Patient consent and data privacy
- •19.4.2 Bias, fairness, and transparency
- •19.4.3 Regulatory guidelines and compliance
- •19.5 Future directions and challenges
- •19.5.1 Emerging technologies and techniques
- •19.5.2 Alternative strategies
- •19.6 Conclusion
- •19.7 Summary
- •References
- •20.1 Risk management
- •20.1.1 Prospective risk assessments
- •20.1.2 Root cause analysis

Artificial Intelligence in Adaptive Radiation Therapy
Figure 11.4. Results of pancreatic target localization on x-ray images acquired at various projection angles
using a deep learning approach. The deep learning model accurately predicts target location with and without
fiducial markers (FMs). (Reproduced with permission from [
61]. Copyright 2019 Elsevier.)
process variable-sized input images. Utilizing data-driven features learned by a
neural network, deep learning models achieve superior performance in lung tumor
motion monitoring compared to template matching-based methods. Deep learningbased methods also show promise in prostate motion monitoring, where the low soft
tissue contrast has challenged direct x-ray-based motion monitoring. To improve
network localization accuracy, Zhao et al [61] proposed concatenating a region
proposal network with a target detection network, where the latter network takes the
output of the region proposal network as input. As shown in figure 11.4, the network
successfully predicted the pancreas target position without fiducial implants, despite
the poor soft tissue contrast of the x-ray projection images.
Direct 2D-to-3D registration through AI has also been investigated, where AI
models output registration results between the 2D images acquired during treatment
delivery and a reference 3D image. The 3D target motion can be inferred from the
registration results. Hou et al proposed a convolutional neural network to learn a
regression function that maps 2D images to their correct position and orientation in
3D space [62]. The network was trained and evaluated for direct registration of
digitally reconstructed radiographs (DRR)-to-CT images for thoracic motion
monitoring. Foote et al combined a patient-specific motion subspace with a
convolutional neural network to predict 3D motion from a single x-ray projection
[63]. The low-rank motion subspace was constructed through principal component
analysis (PCA) of deformation fields between 4DCT scans of different breathing
phases and a reference CT image. The network was trained to take DRR projections
calculated from deformed reference CT images as inputs, and outputs subspace
coordinates that define the deformations associated with the DRR projections. The
model is therefore not limited to predicting rigid translation and rotation of the
target; instead, voxel-wise deformation over the entire image volume can be
obtained.
11-9

Artificial Intelligence in Adaptive Radiation Therapy
11.3.1.3 Real-time volumetric imaging through AI
Volumetric imaging provides 3D visualization of patient anatomy and is routinely
used for treatment planning and pre-treatment patient set-up. However, the
acquisition of volumetric imaging is time-consuming. To reconstruct high-quality
CBCT images without artifacts, hundreds of kV projections typically need to be
acquired. For MRI imaging, sampling of the Fourier space needs sufficient density
to meet the resolution limits set by the Nyquist sampling theorem, which can take up
to several minutes depending on scanning volume, contrast and spatial resolution.
As a result, despite the advantage of fully characterizing both rigid and non-rigid
motions over the entire patient anatomy, volumetric imaging using conventional
sampling and reconstruction techniques is not suitable to guide motion monitoring
during ART delivery, where changes in patient position and/or configuration need to
be determined in real time.
Sparse imaging has been investigated to accelerate volumetric image acquisition,
where only a limited set of data samples is acquired for image reconstruction. Direct
image reconstruction from the sparse samples will result in subsampling artifacts.
Model-based image reconstruction utilizes prior knowledge of the imaging subject to
regularize the reconstruction process and remove image artifacts. Manually crafted
prior knowledge models, including total variation, low rank, and sparsity in the
transform domain, have been investigated. However, the acceleration factors
achieved using these methods are still insufficient to support sub-second imaging
for motion monitoring. Furthermore, to solve the model-based image reconstruction
problem, iterative optimization is needed, which also adds latency to an imagingguided ART delivery system.
AI models learn prior knowledge in a data-driven manner and have demonstrated
advantages over manually crafted models. Zhang et al proposed a neural network
that combines a DenseNet with a deconvolution-based network [64]. The network
takes CT images generated using a sparse set of projections and reconstructed
through conventional filtered back projection as inputs and outputs the corresponding fully sampled CT images. Compared to total variation model-based image
reconstruction, the network-reconstructed images showed reduced streaking artifacts and increased structural similarity with fully sampled images. The work by
Shen et al investigated CT imaging with a single x-ray projection, making it possible
for real-time CT-based motion monitoring during ART delivery [65]. The network
comprises a 2D representation network that learns feature representations from
x-ray projections and a 3D generation network that produces volumetric CT images
from the learned features. A transformation module was introduced between the 2D
and 3D networks to bridge the 2D and 3D feature spaces.
AI models have also been investigated for sparse MRI reconstruction. Zhu et al
showed that a deep neural network can learn a direct mapping between sensor and
image domain MRI data [66]. As shown in figure 11.5, their model (AUTOMAP) is
flexible in learning reconstruction transforms for various sparse MRI sampling
strategies and demonstrated superior image quality when reconstructing from sparse
MRI samples as compared to conventional models. Aggarwal et al combined the
known signal model of MRI formation with a deep neural network for MRI
11-10

Artificial Intelligence in Adaptive Radiation Therapy
Figure 11.5. MRI acquired with various sparse sampling strategies. AUTOMAP-based reconstruction
achieves superior image quality compared to conventional reconstruction. (Reproduced with permission
from [
66]. Copyright 2018 Springer Nature.)
reconstruction [67]. The deep neural network learns useful features in the image
domain and serves as a regularization prior for the model-based image reconstruction process. By incorporating the forward signal model, the authors showed that a
smaller network can be trained with fewer training data samples without compromising image reconstruction quality. To achieve real-time volumetric MRI from a
single radial projection, Feng et al proposed a signature matching-based method
[68]. During an initial offline learning stage, a 4D MRI library consisting of different
breathing motion states and the corresponding breathing motion signatures was
constructed. During the subsequent online matching stage, new motion signatures
were acquired in real time and matched to the pre-learnt motion signatures
associated with known motion states. The method showed promise in real-time
MRI-guided 3D motion monitoring for ART delivery, although handling outlier
motion states that were not learned during the offline stage remains a challenging
topic. By incorporating prior knowledge of MRI acquisition strategy and the sensorto-image domain transform into a deep neural network, Liu et al demonstrated
model robustness to longitudinal patient anatomy changes and eliminated the need
for 4D MRI collection before each MRI-guided ART delivery session [69]. The
model was trained and tested on MRI datasets acquired months apart and
11-11

Artificial Intelligence in Adaptive Radiation Therapy
reconstructed high-quality volumetric MRI from sparse samples with sub-second
acquisition times.
11.3.2 Mitigating system latency through AI
AI has also begun to show promise to reduce system latency for motion monitoring
during ART delivery. AI-based sparsely sampled imaging has reduced not only
image acquisition time but also reconstruction time. Compared to conventional
methods that take several minutes to iteratively solve the sparse image reconstruction problem [70, 71], a trained AI model can produce images within seconds.
In addition to reducing the imaging time, AI models have also been investigated
to support system decision-making during real-time ART. By directly estimating
motion information from the acquired image samples, AI models can remove the
need of image registration to accelerate system decision-making based on the motion
information, thus reducing system latency. The motion monitoring error caused by
system latency can also be mitigated through AI-based motion prediction ahead of
the acquisition time.
11.3.2.1 Direct motion estimation through AI
Conventionally, motion information is estimated through registration of images
acquired during ART delivery with a reference image acquired before treatment.
The motion information ranges from low-dimensional vectors describing rigid
motion (shift and rotation) up to high-dimensional deformation vector fields
(DVFs) with granularity possibly as high as a different transform for each voxel
location. Decisions to adapt the treatment are then made based on the motion
information. The image registration step adds additional time between motion
occurrence and system action to account for the motion. As this time increases, for
example for situations where complex registration, such as deformable registration,
is needed, the true configuration of the patient has increasing potential to deviate
from that measured, leading to reduced effectiveness of motion monitoring which
would need to be offset by attempting to add accurate motion predictions to
measurements.
To address the temporal efficiency of motion prediction, efforts have been made
to directly estimate deformation vector fields (DVFs) from sparse image samples,
with the goal of reducing both image acquisition and motion estimation time.
Terpstra et al proposed a deep learning model to directly estimate 3D DVFs from
under-sampled MRI [72]. They collected a population dataset of 4D MRI from 27
patients and calculated ground truth DVFs through deformable registration
between fully sampled MRIs at different breathing motion phases. A neural network
was trained to take retrospectively under-sampled images as inputs and output the
corresponding DVFs (figure 11.6). The authors reported a total time of 200 ms for
sparse image acquisition and DVF estimation, with an estimation error of less than
2 mm. Stemkens et al introduced a patient-specific subspace model for 3D DVF
estimation from 2D MRI [ 73]. Deformable registration was performed across a
patient-specific 4D MRI dataset. The subspace model was constructed by
11-12

Artificial Intelligence in Adaptive Radiation Therapy
Figure 11.6. (a)–(c) DVFs estimated by a neural network from under-sampled MRI, (d)–(f) compared to
DVFs estimated through deformable registration of high-quality images. (Reproduced from [
72]. CC BY 4.0.)
performing principal component analysis on the obtained DVFs. The problem of
estimating 3D DVF was then reduced to the estimation of principal component
coefficients that best match the sparse MRI and the deformed reference MRI at
sampled locations. They reported a temporal resolution of less than 500 ms with an
average error of 1.45 mm for DVF estimation. Motion estimation from MRI sensor
domain samples without image reconstruction has also been explored. Huttinga et al
incorporated a forward signal model that relates MR sensor domain data to the
underlying image subject for motion estimation [74]. A low-rank motion representation was constructed from a pre-treatment 4D dataset to reduce the problem of
estimating high-dimensional DVFs to estimating eigenvector coefficients. The
method was validated on an MR-linac system, with a reported total latency of
170 ms. Shao et al investigated direct motion estimation from sparse sensor domain
samples without ground truth DVFs [75]. Instead, the network was trained to output
DVFs that optimally wrap a fully sampled prior image to the target sparse samples.
The data fidelity loss was calculated in the sensor domain by undersampling the
deformed prior image using the same sparse sampling pattern as the target. A
regularization loss was also introduced to enforce the smoothness of the DVFs.
Real-time motion estimation from sparse x-ray samples has also been investigated. Li et al explored a Bayesian approach to estimate 3D target location from
sparse x-ray projections [76]. Probability density functions of tumor locations were
estimated from 2D projections acquired with different x-ray imaging directions
during patient set-up. During treatment, 2D tumor location information on the
imager was used to update the likelihood function and the third dimension of tumor
location along the direction of imaging x-ray was estimated by maximizing the
posterior probability distribution. Before treatment, the 2D location information
11-13

Artificial Intelligence in Adaptive Radiation Therapy
was calculated directly from the acquired projections, while the third dimension of
motion information was estimated from a probabilistic function that utilizes
previously acquired x-ray projections as the prior. Shao et al proposed a neural
network that estimates deformation vector fields (DVFs) from x-ray projections
acquired at arbitrary angles [77]. Similar to motion estimation from sparse MRI
samples, a 4DCT dataset was used to train the network with sparse x-ray projections
as inputs and ground truth DVFs as outputs. To enforce angle agnosticism, a
geometry-informed x-ray feature pooling layer was developed to allow the network
to extract angle-dependent image features for motion estimation. The authors
reported a liver target localization error of less than 2 mm with the proposed
method.
For ultrasound-guided ART delivery, Liu et al investigated landmark position
tracking using a cascaded Siamese network [78]. Through the cascaded design, the
model can increasingly focus on regions with the highest probability of containing
the target. The network was trained in an unsupervised manner, where a pseudo
tracking target was determined through corner point detection. They reported a
tracking error of less than 2.5 mm. Mezheritsky et al explored 3D respiratory motion
modeling from 2D ultrasound images using convolutional autoencoders [79]. The
model comprises a rigid alignment module and a deformable motion module to
estimate the 3D DVFs associated with the 2D images. The network was trained
using a population database to encode the motion priors into a low-dimensional
latent space and decode the corresponding DVFs from the latent codes. They
reported a mean tracking error of 3.5 mm.
11.3.2.2 Motion prediction through AI
Utilizing temporal priors of motion patterns, AI models have also been developed to
predict motion ahead of time to address issues of latency. For the prediction of lowdimensional motion signals, Teo et al implemented a perceptron neural network to
predict 1D breathing motion trajectories [80]. The network was trained and tested
using breathing motion signals collected from CyberKnife lung patients. It took
motion samples acquired over the previous 4 s to predict the future breathing motion
signal 650 ms ahead, with submillimeter accuracy. Ruan et al proposed a kernel
density estimation-based approach to predict 1D breathing motion signals acquired
from surface imaging [81]. The joint probability distributions of observed and future
motion signals were estimated through Gaussian kernel regression. The authors
reported a normalized root mean square error of less than 1 mm under various data
sampling strategies and prediction horizons. Ren et al developed an autoregressive
moving-average model to predict breathing motion in both the superior–inferior and
anterior–posterior directions, where motion observations were obtained from
marker-based imaging [82]. They also reported that the prediction accuracy depends
on the lookahead time and imaging rates, with submillimeter accuracy achieved for
a prediction horizon of 200 ms and an imaging rate of 10 Hz.
The prediction of high-dimensional motion information has also been investigated. Ginn et al explored an image regression model to predict motion states
based on 2D cine MRI acquired on an MR-linac system [83]. Future MR images
11-14

Artificial Intelligence in Adaptive Radiation Therapy
Figure 11.7. Example predicted motion traces and target contours via image regression. (Reproduced from
[
83] with permission from John Wiley & Sons. Copyright 2019 American Association of Physicists in
Medicine.)
with a prediction horizon of 250–330 ms were generated through Gaussian kernel
regression over historical image samples acquired at a 4 Hz frame rate. As shown in
figure 11.7, future motion information can be calculated from the predicted images
to support ART delivery, and the authors reported an averaged gating decision
accuracy of 95.8%. Romaguera et al predicted in-plane deformation fields via
discriminative spatial transformers [84]. The model aimed to learn deformations
between consecutive images and extrapolate over time, utilizing image features
learned over multiple scales. Submillimeter target position prediction accuracy was
achieved when evaluated on MRI, CT, and ultrasound datasets.
While predicting high-dimensional motion information has the advantage of
modeling complex motion beyond rigid shifts, it also poses challenges for
prediction accuracy. Lombardo et al compared three models with similar network
structures that predict 2D tumor centroid positions, 2D tumor contours, and
DVFs associated with the tumor, respectively, using cine MRI acquired on an
MR-linac [85]. Their results sugge sted that shif ti ng the tumor contour based on the
simple 2D centroid position achieved the highest accuracy in predicting tumor
boundary, evaluated using Hausdorff distances between the ground truth and
predicted tumor contours.
11-15

Artificial Intelligence in Adaptive Radiation Therapy
11.4 AI for ART delivery: future directions
11.4.1 Management of non-respiratory motion
While most AI models have focused on managing respiration-induced anatomy
changes, recent research has demonstrated that non-respiratory motions, such as
peristalsis and slow baseline anatomy drifts also significantly alter patient anatomy
and contribute to treatment delivery uncertainty [6–8]. Modeling non-breathing
motion to guide ART delivery is challenged by two factors. First, such motion is
usually non-cyclic, meaning prior knowledge learnt from historic observations may
have limited utility in modeling future motion states. Second, such motion may
occur simultaneously with respiration, for example, in the abdomen. The significant
magnitudes of respiratory motion can confound the analysis and modeling for nonrespiratory motion.
Liu et al have investigated simultaneous modeling of breathing motion and slow
baseline anatomy drifts for abdomen patients [86]. A multi-temporal resolution
image time series was reconstructed from radial MRI samples, where fast breathing
motion was extracted from high-temporal-resolution samples and slow drifting
motion was extracted from a breathing motion-corrected low-temporal-resolution
image time series. A motion prediction scheme was further investigated based on
low-rank motion models constructed from principal component analysis of DVFs
between different motion states. Model coefficients for future motion states were
predicted through Gaussian kernel regression and linear regression, for breathing
and slow drifting motions, respectively. Different prediction horizons were also
investigated, where breathing motion was predicted 340 ms ahead of time and slow
drifting motion was predicted 8.5 s ahead of time (figure 11.8). This longer
prediction horizon for slow drifting motion is permitted by the slower motion
velocity and is chosen to mitigate the latency resulted from slow drifting motion
modeling without compromising the prediction accuracy (submillimeter in centroid
position estimation).
Prediction of stomach contractile motion has also been investigated, following a
similar approach that separates stomach motion from breathing motion through
multi-temporal resolution radial MRI reconstruction [87]. A stomach motion model
consisting of ten motion phases was built from a breathing-motion-corrected time
Figure 11.8. Example slow drifting motion prediction results. A DVF was predicted to deform a reference
image to reflect motion states 8.5 s ahead of the image acquisition time. (Reproduced with permission from
[
86]. Copyright 2021 Institute of Physics and Engineering in Medicine.)
11-16

Artificial Intelligence in Adaptive Radiation Therapy
series of 2–5 min. Prediction of future stomach motion was achieved through linear
extrapolation of motion phases determined using sparse MRI samples. The authors
reported submillimeter prediction accuracy using sparse samples of 10 radial MRI
spokes with a prediction horizon of 5 s.
While preliminary investigations have demonstrated the feasibility of monitoring
non-breathing motion during treatment delivery, how to account for such motion
during treatment remains an open question. Zhang et al have developed a dose
accumulation tool to assess the accumulated dose in gastrointestinal organs [88].
Variations in motion states during treatment delivery were reconstructed retrospectively through multi-temporal resolution reconstruction of radial MRI samples.
Delivered dose was calculated for each motion state and accumulated to estimate the
total delivered dose. One solution to accounting partially for dose variations due to
such motions involves adapting the treatment plan for subsequent treatments based
on the accumulated dose from previous fractions. Other motion management
schemes, such as MLC tracking or treatment gating, may also help mitigate
treatment uncertainty introduced by non-breathing motion. However, accounting
for multiple motions simultaneously will complicate the decision-making process for
treatment adaptation, add system latency, and elongate treatment time.
11.4.2 Training AI models with small or unpaired datasets
The development of AI models typically requires a training dataset to optimize
model parameters. For AI models developed for ART delivery, most training
datasets are population-based and consist of motion information (volumetric images
or low-dimensional motion signals) collected from a patient population. For the
modeling of cyclic motion, such as respiration, patient-specific training datasets have
also been used. These patient-specific datasets consist of samples acquired from the
same patients but at different motion states. The quality of the training dataset, in
terms of both image quality and quantity, directly impacts the model quality.
However, due to the variance in data collection protocols and the protected health
information associated with medical data, collecting high-quality training data can
be time-consuming and may present a bottleneck for clinical implementation.
Shen et al investigated a novel machine learning technique, implicit neural
representation (INR) learning to develop a patient-specific AI model for sparse
imaging [89]. The model required only a single prior image for training and
demonstrated the potential of high-quality image reconstruction from sparse
samples of 20 CT projections or 40 radial MRI spokes. Liu et al leveraged INR
to further accelerate volumetric imaging for real-time motion monitoring [90]. The
method utilized two volumetric MRIs acquired at different breath-hold states for
model training and reconstructed volumetric MRI from sparse samples of two
orthogonal cine slices. The method was evaluated for free-breathing abdominal
patients and reconstructed high-quality MRI with breathing motion states in
between or outside the training images (figure 11.9). As illustrated in figure 11.10,
during INR-based sparse image reconstruction, a prior model was trained by
optimizing a perceptron network to learn a mapping from prior image coordinates
11-17

Artificial Intelligence in Adaptive Radiation Therapy
Figure 11.9. INR reconstructs volumetric MRI with motion states in between (patient 1) and outside (patient
2) the 2 prior images. (Reproduced from [
American Association of Physicists in Medicine.)
90] with permission from John Wiley & Sons. Copyright 2024
Figure 11.10. Framework of INR learning for real-time volumetric MRI. (Reproduced from [90] with
permission from John Wiley & Sons . Copyright 2024 American Association of Physicists in Medicine.)
to prior image voxel values. The optimized network parameterization served as a
regularization for the subsequent sparse image reconstruction task, where the
network weights were further fine-tuned to fit the sparse samples of the testing
images. Volumetric reconstruction was obtained by querying the fine-tuned network
at desired 3D image coordinates. By regarding each pair of image coordinate and
image voxel value as a training sample, INR requires significantly less training data
than conventional machine learning methods.
There is growing interest in population-based learning with unpaired training
data samples. Generative models, such as generative adversarial networks and
diffusion models, aim to learn a mapping between two probabilistic distributions
instead of a mapping between data points. Therefore, the model inputs and ground
truth outputs do not have to be perfectly aligned. Generative models have found
11-18
Соседние файлы в папке Библиотека им академика М.И. Перельмана
