Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5525_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
29 Мб
Скачать
Artificial Intelligence in Adaptive Radiation Therapy
Figure 11.4. Results of pancreatic target localization on x-ray images acquired at various projection angles using a deep learning approach. The deep learning model accurately predicts target location with and without ducial markers (FMs). (Reproduced with permission from [
61]. Copyright 2019 Elsevier.)
process variable-sized input images. Utilizing data-driven features learned by a neural network, deep learning models achieve superior performance in lung tumor motion monitoring compared to template matching-based methods. Deep learning­based methods also show promise in prostate motion monitoring, where the low soft tissue contrast has challenged direct x-ray-based motion monitoring. To improve network localization accuracy, Zhao et al [61] proposed concatenating a region proposal network with a target detection network, where the latter network takes the output of the region proposal network as input. As shown in gure 11.4, the network successfully predicted the pancreas target position without ducial implants, despite the poor soft tissue contrast of the x-ray projection images.
Direct 2D-to-3D registration through AI has also been investigated, where AI models output registration results between the 2D images acquired during treatment delivery and a reference 3D image. The 3D target motion can be inferred from the registration results. Hou et al proposed a convolutional neural network to learn a regression function that maps 2D images to their correct position and orientation in 3D space [62]. The network was trained and evaluated for direct registration of digitally reconstructed radiographs (DRR)-to-CT images for thoracic motion monitoring. Foote et al combined a patient-specic motion subspace with a convolutional neural network to predict 3D motion from a single x-ray projection [63]. The low-rank motion subspace was constructed through principal component analysis (PCA) of deformation elds between 4DCT scans of different breathing phases and a reference CT image. The network was trained to take DRR projections calculated from deformed reference CT images as inputs, and outputs subspace coordinates that dene the deformations associated with the DRR projections. The model is therefore not limited to predicting rigid translation and rotation of the target; instead, voxel-wise deformation over the entire image volume can be obtained.
11-9
Artificial Intelligence in Adaptive Radiation Therapy
11.3.1.3 Real-time volumetric imaging through AI
Volumetric imaging provides 3D visualization of patient anatomy and is routinely used for treatment planning and pre-treatment patient set-up. However, the acquisition of volumetric imaging is time-consuming. To reconstruct high-quality CBCT images without artifacts, hundreds of kV projections typically need to be acquired. For MRI imaging, sampling of the Fourier space needs sufcient density to meet the resolution limits set by the Nyquist sampling theorem, which can take up to several minutes depending on scanning volume, contrast and spatial resolution. As a result, despite the advantage of fully characterizing both rigid and non-rigid motions over the entire patient anatomy, volumetric imaging using conventional sampling and reconstruction techniques is not suitable to guide motion monitoring during ART delivery, where changes in patient position and/or conguration need to be determined in real time.
Sparse imaging has been investigated to accelerate volumetric image acquisition, where only a limited set of data samples is acquired for image reconstruction. Direct image reconstruction from the sparse samples will result in subsampling artifacts. Model-based image reconstruction utilizes prior knowledge of the imaging subject to regularize the reconstruction process and remove image artifacts. Manually crafted prior knowledge models, including total variation, low rank, and sparsity in the transform domain, have been investigated. However, the acceleration factors achieved using these methods are still insufcient to support sub-second imaging for motion monitoring. Furthermore, to solve the model-based image reconstruction problem, iterative optimization is needed, which also adds latency to an imaging­guided ART delivery system.
AI models learn prior knowledge in a data-driven manner and have demonstrated advantages over manually crafted models. Zhang et al proposed a neural network that combines a DenseNet with a deconvolution-based network [64]. The network takes CT images generated using a sparse set of projections and reconstructed through conventional ltered back projection as inputs and outputs the correspond­ing fully sampled CT images. Compared to total variation model-based image reconstruction, the network-reconstructed images showed reduced streaking arti­facts and increased structural similarity with fully sampled images. The work by Shen et al investigated CT imaging with a single x-ray projection, making it possible for real-time CT-based motion monitoring during ART delivery [65]. The network comprises a 2D representation network that learns feature representations from x-ray projections and a 3D generation network that produces volumetric CT images from the learned features. A transformation module was introduced between the 2D and 3D networks to bridge the 2D and 3D feature spaces.
AI models have also been investigated for sparse MRI reconstruction. Zhu et al showed that a deep neural network can learn a direct mapping between sensor and image domain MRI data [66]. As shown in gure 11.5, their model (AUTOMAP) is exible in learning reconstruction transforms for various sparse MRI sampling strategies and demonstrated superior image quality when reconstructing from sparse MRI samples as compared to conventional models. Aggarwal et al combined the known signal model of MRI formation with a deep neural network for MRI
11-10
Artificial Intelligence in Adaptive Radiation Therapy
Figure 11.5. MRI acquired with various sparse sampling strategies. AUTOMAP-based reconstruction achieves superior image quality compared to conventional reconstruction. (Reproduced with permission from [
66]. Copyright 2018 Springer Nature.)
reconstruction [67]. The deep neural network learns useful features in the image domain and serves as a regularization prior for the model-based image reconstruc­tion process. By incorporating the forward signal model, the authors showed that a smaller network can be trained with fewer training data samples without compro­mising image reconstruction quality. To achieve real-time volumetric MRI from a single radial projection, Feng et al proposed a signature matching-based method [68]. During an initial ofine learning stage, a 4D MRI library consisting of different breathing motion states and the corresponding breathing motion signatures was constructed. During the subsequent online matching stage, new motion signatures were acquired in real time and matched to the pre-learnt motion signatures associated with known motion states. The method showed promise in real-time MRI-guided 3D motion monitoring for ART delivery, although handling outlier motion states that were not learned during the ofine stage remains a challenging topic. By incorporating prior knowledge of MRI acquisition strategy and the sensor­to-image domain transform into a deep neural network, Liu et al demonstrated model robustness to longitudinal patient anatomy changes and eliminated the need for 4D MRI collection before each MRI-guided ART delivery session [69]. The model was trained and tested on MRI datasets acquired months apart and
11-11
Artificial Intelligence in Adaptive Radiation Therapy
reconstructed high-quality volumetric MRI from sparse samples with sub-second acquisition times.
11.3.2 Mitigating system latency through AI
AI has also begun to show promise to reduce system latency for motion monitoring during ART delivery. AI-based sparsely sampled imaging has reduced not only image acquisition time but also reconstruction time. Compared to conventional methods that take several minutes to iteratively solve the sparse image reconstruc­tion problem [70, 71], a trained AI model can produce images within seconds.
In addition to reducing the imaging time, AI models have also been investigated to support system decision-making during real-time ART. By directly estimating motion information from the acquired image samples, AI models can remove the need of image registration to accelerate system decision-making based on the motion information, thus reducing system latency. The motion monitoring error caused by system latency can also be mitigated through AI-based motion prediction ahead of the acquisition time.
11.3.2.1 Direct motion estimation through AI
Conventionally, motion information is estimated through registration of images acquired during ART delivery with a reference image acquired before treatment. The motion information ranges from low-dimensional vectors describing rigid motion (shift and rotation) up to high-dimensional deformation vector fields (DVFs) with granularity possibly as high as a different transform for each voxel location. Decisions to adapt the treatment are then made based on the motion information. The image registration step adds additional time between motion occurrence and system action to account for the motion. As this time increases, for example for situations where complex registration, such as deformable registration, is needed, the true conguration of the patient has increasing potential to deviate from that measured, leading to reduced effectiveness of motion monitoring which would need to be offset by attempting to add accurate motion predictions to measurements.
To address the temporal efciency of motion prediction, efforts have been made to directly estimate deformation vector elds (DVFs) from sparse image samples, with the goal of reducing both image acquisition and motion estimation time. Terpstra et al proposed a deep learning model to directly estimate 3D DVFs from under-sampled MRI [72]. They collected a population dataset of 4D MRI from 27 patients and calculated ground truth DVFs through deformable registration between fully sampled MRIs at different breathing motion phases. A neural network was trained to take retrospectively under-sampled images as inputs and output the corresponding DVFs (gure 11.6). The authors reported a total time of 200 ms for sparse image acquisition and DVF estimation, with an estimation error of less than 2 mm. Stemkens et al introduced a patient-specic subspace model for 3D DVF estimation from 2D MRI [ 73]. Deformable registration was performed across a patient-specic 4D MRI dataset. The subspace model was constructed by
11-12
Artificial Intelligence in Adaptive Radiation Therapy
Figure 11.6. (a)–(c) DVFs estimated by a neural network from under-sampled MRI, (d)(f) compared to DVFs estimated through deformable registration of high-quality images. (Reproduced from [
72]. CC BY 4.0.)
performing principal component analysis on the obtained DVFs. The problem of estimating 3D DVF was then reduced to the estimation of principal component coefcients that best match the sparse MRI and the deformed reference MRI at sampled locations. They reported a temporal resolution of less than 500 ms with an average error of 1.45 mm for DVF estimation. Motion estimation from MRI sensor domain samples without image reconstruction has also been explored. Huttinga et al incorporated a forward signal model that relates MR sensor domain data to the underlying image subject for motion estimation [74]. A low-rank motion represen­tation was constructed from a pre-treatment 4D dataset to reduce the problem of estimating high-dimensional DVFs to estimating eigenvector coefcients. The method was validated on an MR-linac system, with a reported total latency of 170 ms. Shao et al investigated direct motion estimation from sparse sensor domain samples without ground truth DVFs [75]. Instead, the network was trained to output DVFs that optimally wrap a fully sampled prior image to the target sparse samples. The data delity loss was calculated in the sensor domain by undersampling the deformed prior image using the same sparse sampling pattern as the target. A regularization loss was also introduced to enforce the smoothness of the DVFs.
Real-time motion estimation from sparse x-ray samples has also been inves­tigated. Li et al explored a Bayesian approach to estimate 3D target location from sparse x-ray projections [76]. Probability density functions of tumor locations were estimated from 2D projections acquired with different x-ray imaging directions during patient set-up. During treatment, 2D tumor location information on the imager was used to update the likelihood function and the third dimension of tumor location along the direction of imaging x-ray was estimated by maximizing the posterior probability distribution. Before treatment, the 2D location information
11-13
Artificial Intelligence in Adaptive Radiation Therapy
was calculated directly from the acquired projections, while the third dimension of motion information was estimated from a probabilistic function that utilizes previously acquired x-ray projections as the prior. Shao et al proposed a neural network that estimates deformation vector elds (DVFs) from x-ray projections acquired at arbitrary angles [77]. Similar to motion estimation from sparse MRI samples, a 4DCT dataset was used to train the network with sparse x-ray projections as inputs and ground truth DVFs as outputs. To enforce angle agnosticism, a geometry-informed x-ray feature pooling layer was developed to allow the network to extract angle-dependent image features for motion estimation. The authors reported a liver target localization error of less than 2 mm with the proposed method.
For ultrasound-guided ART delivery, Liu et al investigated landmark position tracking using a cascaded Siamese network [78]. Through the cascaded design, the model can increasingly focus on regions with the highest probability of containing the target. The network was trained in an unsupervised manner, where a pseudo tracking target was determined through corner point detection. They reported a tracking error of less than 2.5 mm. Mezheritsky et al explored 3D respiratory motion modeling from 2D ultrasound images using convolutional autoencoders [79]. The model comprises a rigid alignment module and a deformable motion module to estimate the 3D DVFs associated with the 2D images. The network was trained using a population database to encode the motion priors into a low-dimensional latent space and decode the corresponding DVFs from the latent codes. They reported a mean tracking error of 3.5 mm.
11.3.2.2 Motion prediction through AI
Utilizing temporal priors of motion patterns, AI models have also been developed to predict motion ahead of time to address issues of latency. For the prediction of low­dimensional motion signals, Teo et al implemented a perceptron neural network to predict 1D breathing motion trajectories [80]. The network was trained and tested using breathing motion signals collected from CyberKnife lung patients. It took motion samples acquired over the previous 4 s to predict the future breathing motion signal 650 ms ahead, with submillimeter accuracy. Ruan et al proposed a kernel density estimation-based approach to predict 1D breathing motion signals acquired from surface imaging [81]. The joint probability distributions of observed and future motion signals were estimated through Gaussian kernel regression. The authors reported a normalized root mean square error of less than 1 mm under various data sampling strategies and prediction horizons. Ren et al developed an autoregressive moving-average model to predict breathing motion in both the superior–inferior and anterior–posterior directions, where motion observations were obtained from marker-based imaging [82]. They also reported that the prediction accuracy depends on the lookahead time and imaging rates, with submillimeter accuracy achieved for a prediction horizon of 200 ms and an imaging rate of 10 Hz.
The prediction of high-dimensional motion information has also been inves­tigated. Ginn et al explored an image regression model to predict motion states based on 2D cine MRI acquired on an MR-linac system [83]. Future MR images
11-14
Artificial Intelligence in Adaptive Radiation Therapy
Figure 11.7. Example predicted motion traces and target contours via image regression. (Reproduced from [
83] with permission from John Wiley & Sons. Copyright 2019 American Association of Physicists in
Medicine.)
with a prediction horizon of 250–330 ms were generated through Gaussian kernel regression over historical image samples acquired at a 4 Hz frame rate. As shown in gure 11.7, future motion information can be calculated from the predicted images to support ART delivery, and the authors reported an averaged gating decision accuracy of 95.8%. Romaguera et al predicted in-plane deformation elds via discriminative spatial transformers [84]. The model aimed to learn deformations between consecutive images and extrapolate over time, utilizing image features learned over multiple scales. Submillimeter target position prediction accuracy was achieved when evaluated on MRI, CT, and ultrasound datasets.
While predicting high-dimensional motion information has the advantage of modeling complex motion beyond rigid shifts, it also poses challenges for prediction accuracy. Lombardo et al compared three models with similar network structures that predict 2D tumor centroid positions, 2D tumor contours, and DVFs associated with the tumor, respectively, using cine MRI acquired on an MR-linac [85]. Their results sugge sted that shif ti ng the tumor contour based on the simple 2D centroid position achieved the highest accuracy in predicting tumor boundary, evaluated using Hausdorff distances between the ground truth and predicted tumor contours.
11-15
Artificial Intelligence in Adaptive Radiation Therapy

11.4 AI for ART delivery: future directions

11.4.1 Management of non-respiratory motion
While most AI models have focused on managing respiration-induced anatomy changes, recent research has demonstrated that non-respiratory motions, such as peristalsis and slow baseline anatomy drifts also signicantly alter patient anatomy and contribute to treatment delivery uncertainty [68]. Modeling non-breathing motion to guide ART delivery is challenged by two factors. First, such motion is usually non-cyclic, meaning prior knowledge learnt from historic observations may have limited utility in modeling future motion states. Second, such motion may occur simultaneously with respiration, for example, in the abdomen. The signicant magnitudes of respiratory motion can confound the analysis and modeling for non­respiratory motion.
Liu et al have investigated simultaneous modeling of breathing motion and slow baseline anatomy drifts for abdomen patients [86]. A multi-temporal resolution image time series was reconstructed from radial MRI samples, where fast breathing motion was extracted from high-temporal-resolution samples and slow drifting motion was extracted from a breathing motion-corrected low-temporal-resolution image time series. A motion prediction scheme was further investigated based on low-rank motion models constructed from principal component analysis of DVFs between different motion states. Model coefcients for future motion states were predicted through Gaussian kernel regression and linear regression, for breathing and slow drifting motions, respectively. Different prediction horizons were also investigated, where breathing motion was predicted 340 ms ahead of time and slow drifting motion was predicted 8.5 s ahead of time (gure 11.8). This longer prediction horizon for slow drifting motion is permitted by the slower motion velocity and is chosen to mitigate the latency resulted from slow drifting motion modeling without compromising the prediction accuracy (submillimeter in centroid position estimation).
Prediction of stomach contractile motion has also been investigated, following a similar approach that separates stomach motion from breathing motion through multi-temporal resolution radial MRI reconstruction [87]. A stomach motion model consisting of ten motion phases was built from a breathing-motion-corrected time
Figure 11.8. Example slow drifting motion prediction results. A DVF was predicted to deform a reference image to reect motion states 8.5 s ahead of the image acquisition time. (Reproduced with permission from [
86]. Copyright 2021 Institute of Physics and Engineering in Medicine.)
11-16
Artificial Intelligence in Adaptive Radiation Therapy
series of 2–5 min. Prediction of future stomach motion was achieved through linear extrapolation of motion phases determined using sparse MRI samples. The authors reported submillimeter prediction accuracy using sparse samples of 10 radial MRI spokes with a prediction horizon of 5 s.
While preliminary investigations have demonstrated the feasibility of monitoring non-breathing motion during treatment delivery, how to account for such motion during treatment remains an open question. Zhang et al have developed a dose accumulation tool to assess the accumulated dose in gastrointestinal organs [88]. Variations in motion states during treatment delivery were reconstructed retro­spectively through multi-temporal resolution reconstruction of radial MRI samples. Delivered dose was calculated for each motion state and accumulated to estimate the total delivered dose. One solution to accounting partially for dose variations due to such motions involves adapting the treatment plan for subsequent treatments based on the accumulated dose from previous fractions. Other motion management schemes, such as MLC tracking or treatment gating, may also help mitigate treatment uncertainty introduced by non-breathing motion. However, accounting for multiple motions simultaneously will complicate the decision-making process for treatment adaptation, add system latency, and elongate treatment time.
11.4.2 Training AI models with small or unpaired datasets
The development of AI models typically requires a training dataset to optimize model parameters. For AI models developed for ART delivery, most training datasets are population-based and consist of motion information (volumetric images or low-dimensional motion signals) collected from a patient population. For the modeling of cyclic motion, such as respiration, patient-specic training datasets have also been used. These patient-specic datasets consist of samples acquired from the same patients but at different motion states. The quality of the training dataset, in terms of both image quality and quantity, directly impacts the model quality. However, due to the variance in data collection protocols and the protected health information associated with medical data, collecting high-quality training data can be time-consuming and may present a bottleneck for clinical implementation.
Shen et al investigated a novel machine learning technique, implicit neural representation (INR) learning to develop a patient-specic AI model for sparse imaging [89]. The model required only a single prior image for training and demonstrated the potential of high-quality image reconstruction from sparse samples of 20 CT projections or 40 radial MRI spokes. Liu et al leveraged INR to further accelerate volumetric imaging for real-time motion monitoring [90]. The method utilized two volumetric MRIs acquired at different breath-hold states for model training and reconstructed volumetric MRI from sparse samples of two orthogonal cine slices. The method was evaluated for free-breathing abdominal patients and reconstructed high-quality MRI with breathing motion states in between or outside the training images (gure 11.9). As illustrated in gure 11.10, during INR-based sparse image reconstruction, a prior model was trained by optimizing a perceptron network to learn a mapping from prior image coordinates
11-17
Artificial Intelligence in Adaptive Radiation Therapy
Figure 11.9. INR reconstructs volumetric MRI with motion states in between (patient 1) and outside (patient
2) the 2 prior images. (Reproduced from [ American Association of Physicists in Medicine.)
90] with permission from John Wiley & Sons. Copyright 2024
Figure 11.10. Framework of INR learning for real-time volumetric MRI. (Reproduced from [90] with permission from John Wiley & Sons . Copyright 2024 American Association of Physicists in Medicine.)
to prior image voxel values. The optimized network parameterization served as a regularization for the subsequent sparse image reconstruction task, where the network weights were further ne-tuned to t the sparse samples of the testing images. Volumetric reconstruction was obtained by querying the ne-tuned network at desired 3D image coordinates. By regarding each pair of image coordinate and image voxel value as a training sample, INR requires signicantly less training data than conventional machine learning methods.
There is growing interest in population-based learning with unpaired training data samples. Generative models, such as generative adversarial networks and diffusion models, aim to learn a mapping between two probabilistic distributions instead of a mapping between data points. Therefore, the model inputs and ground truth outputs do not have to be perfectly aligned. Generative models have found
11-18