Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5525_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
29 Мб
Скачать
Artificial Intelligence in Adaptive Radiation Therapy
[106] Han B et al 2023 Characterization of biology-guided radiotherapy accuracy as a function of
PET tracer uptake Int. J. Radiat. Oncol. Biol. Phys.
[107] Ouyang J, Chen K T, Gong E, Pauly J and Zaharchuk G 2019 Ultra-low-dose PET
reconstruction using generative adversarial network with feature matching and task-specic perceptual loss Med. Phys.
[108] Lim H, Chun I Y, Dewaraja Y K and Fessler J A 2019 Improved low-count quantitative
PET reconstruction with an iterative neural network Med. Phys.
[109] Fu J et al 2022 Patient-specic mean teacher UNet for enhancing PET image and low-dose
PET reconstruction on ReeXion X1 biology-guided radiotherapy system arXiv:
46 3555–64
117 e668–e9
46 3512–22
2209.05665
11-29
IOP Publishing
Artificial Intelligence in Adaptive Radiation Therapy
Yi Wang and X. Sharon Qi
Chapter 12
Artificial intelligence for quality assurance
in adaptive radiation therapy
Sang Kyu Lee and Maria Chan
Quality assurance (QA) within the realm of radiation oncology physics plays a crucial role in ensuring that all patient care processes adhere to predened quality standards [1]. This comprehensive approach encompasses various policies and procedures designed to establish these standards and outline the methods for monitoring. While articial intelligence (AI) has the capability to enhance this process by analysing complex data and interpreting system outputs in a comprehen­sible manner, it is important to recognize that AI serves as a support tool rather than a replacement for QA. The role of AI is to assist in the interpretation and processing of intricate information, thereby facilitating the QA process but not supplanting the foundational principles and practices of QA.

12.1 Introduction

In this chapter, we categorize quality assurance (QA) within the eld of radiation therapy into two main divisionsradiotherapy plan QA and machine/instrumentation QA. We delve deeper into radiotherapy plan QA by mapping it across the care chain of a patient, identifying several key processes where medical physicists contribute significantly (figure 12.1). Following this, we explore various scholarly works on the application of artificial intelligence (AI) that can support and enhance QA practices across each sub-category. This review work aims to highlight the potential of AI as a tool to assist in the precision and efciency of QA processes in radiation therapy.

12.2 Patient QA

12.2.1 Pre-planning QA
12.2.1.1 Contouring QA
Several early investigations used the distribution of the morphological [24]or texture [5] features of the contours to detect a contouring error by nding signicant
doi:10.1088/978-0-7503-6119-4ch12 12-1 ª IOP Publishing Ltd 2025. All rights,
including for text and data mining (TDM), artificial intelligence (AI) training, and similar technologies, are reserved.
Artificial Intelligence in Adaptive Radiation Therapy
Figure 12.1. Simplied QA workow in radiation oncology physics. Targets of the QA and QA elements are represented in parallelograms and round-edged rectangles, respectively.
deviation from the ground truth distribution derived from veried contours. Predened thresholding on the deviation from the mean was applied by Chen et al [3] to detect anomalies. Alternatively, a supervised machine-learning approach, such as conditional random forest [2] can be used. A deep-learning-based approach, which had been introduced as an auto-contouring tool [6], can also be used for contouring QA: for example, a convolutional neural network (CNN) was used for verifying atlas-based auto-contours [7].
Care must be taken when adopting the AI contouring QA tools. Many published algorithms are trained to maximize quantitative accuracy metrics such as the Dice similarity coefcient (DSC) or Hausdorff distance (HD). However, these metrics do not always agree with each other [8] or may not be sensitive enough to detect localized errors that could be clinically relevant [7]. Although a correlation between clinical acceptability and surface DSC has been demonstrated [9], other comple­mentary evaluation methods, such as Likert scales [10] or the Turing test [11] could be considered. Moreover, it is important to consider potential mistakes and inter­operator variability in manual contours that served as training datathe use of a high-quality contour dataset, such as the manually curated multi-institutional dataset [12] should be considered. Alternatively, tools are available to generate consensus contours from multiple sets of human-generated contours [13]. As such, evaluation of AI contours should be conducted in multiple domains, rather than relying on a single metric [14](figure 12.2).
12.2.1.2 Image fusion (registration)
Importance of image fusion (registration) QA is growing, particularly in adaptive radiotherapy where registration error between simulation and daily CT propagates
12-2
Artificial Intelligence in Adaptive Radiation Therapy
Figure 12.2. Different methods of evaluation of AI-generated contours. (Reproduced with permission from [
14]. Copyright 2021 Elsevier.)
to contouring and dose accumulation accuracy and thus negatively impacts plan quality. Image registration QA is conducted mainly at two levels: (i) at the commissioning stage for the registration system and (ii) at a patient-specic level as part of radiotherapy workow, when image fusion is required for target delineation or plan adaptation. Patient-specic evaluation of image registration is challenging due to a lack of ground truth [17], sources of uncertainties such as anatomical changes or image quality [18], and a lack of consensus on evaluation methods [19]. Consequently, QA practices vary between institutions [19, 20], many of which use qualitative visual inspection in a clinical setting [20]. Although visual inspection is an important element of the patient-specic registration QA, as recommended by TG-132 [15], it is not standardized and relies on the expertise level of a user [16]. Quantitative assessment can be done by identifying and calculating the displacement between homologous anatomical landmarks or con­tours, or examination of a deformation vector eld (DVF) [16]. AI can help in the extraction of homologous structures, which would have been a labor-intensive process if done manually. For this purpose, in addition to the existing image transformation methods [21, 22], the use of machine learning such as the decision tree [23] or deep learning [24] has been investigated. AI can also play an important role in QA of deformable registration and can augment the DVF-based metrics that were already proposed as quantitative QA [2527]. For example, Neylon et al [28] trained a supervised neural network model to calculate the registration error in the physical distance from the registered image and a biomechanical model. Similarly, CNN-based supervised models were used to predict registration accuracy from patches of images [29, 30]. Smolders et al [31] trained a deep-learning-based DVF uncertainty prediction model, which can be used for comparing different registration algorithms or estimating uncertainty in accumulated dose.
12-3
Artificial Intelligence in Adaptive Radiation Therapy
12.2.2 Pre-treatment plan QA
12.2.2.1 Plan quality QA
One of the most important QAs of plan quality is the planned dose distribution meeting clinical goals in terms of target coverage and organs-at-risk (OARs) sparing. Reviewers often use a table of scoresheets that compares the calculated dose–volume histogram (DVH) parameters for targets and OARs versus the preset criteria that were established based on clinical experience or published guidelines. However, it is not obvious from the scoresheets that the best possible trade-off between target coverage and OAR sparing is achieved, which is determined by complex interactions between patient anatomy, treatment modalities, and other patient-specic considerations. AI can unravel these complex relationships to inform the reviewers whether the planned dose distribution is optimal given these con­straints. Knowledge-based planning (KBP) learns from high-quality plans using various AI approaches to estimate a range of achievable dose. A thorough review of KBP is outside the scope of this chapter and can be found in Ge et al [32]. The concept of KBP can be used not only for assisting IMRT optimization, but also for automated QA of the completed plans. Tol et al [33], and subsequently Cao et al [34], used a commercial KBP tool RapidPlan (Varian Medical Systems, Palo Alto, CA) to predict a range of DVH parameters for OARs and showed that the plans for which OAR doses exceeded the predicted range can be improved by further optimization. Stanhope et al [35] used their DVH prediction model to identify the clinically unacceptable patient-specic QA (PSQA) results that translate into the DVH falling outside the predicted DVH range. A deep-learning-based approach [3638] aims at directly predicting 3D dose distribution from simulation CT (gure 12.3); DVH parameters can then be derived from the predicted 3D distribution and used for agging suboptimal plans [38]
Use of dose prediction models for automatic plan quality QA comes with caveats: in order to achieve precise and accurate prediction and identify more suboptimal plans, it is important to train a model from the carefully selected high-quality plans with consistent OARs sparing [34]. Also, as pointed out in [34, 38], automatic QA does not always align with physicians judgment, due to the subjective nature of plan quality review where physicians have diverse preferences over target coverage and OAR sparing.
12.2.2.2 Pre-treatment chart review
Pre-treatment chart review is a comprehensive review of a radiotherapy plan by a qualied medical physicist before the treatment begins. Physics chart review comprises various aspects of a plan, including (i) data transfer integrity, (ii) dose calculation accuracy, (iii) plan quality, (iv) consistency of prescription including image guidance requests, and (v) other special considerations [39]. It was shown to be the most effective safety barrier to prevent radiotherapy incidents [40]. However, studies show that a sizable portion of errors pass through the physics chart review [4042]. Automation and standardization are listed as two major paths to enhance the effectiveness of physics chart review [ 39, 41]. To this end, software solutions have
12-4
Artificial Intelligence in Adaptive Radiation Therapy
Figure 12.3. Examples of head-and-neck plans agged by a deep-learning-based 3D dose prediction model by Gronberg et al for suboptimal dose to (A) a spinal cord and (B) oral cavity. For (C), the model was not able to predict suboptimal dose to the esophagus. (Reproduced with permission from [
38]. Copyright 2023 Elsevier.)
been developed to assist the manual plan checking process [4346]. These software tools typically apply a set of predened rules for a given treatment type. As previously mentioned, however, there are many circumstances that these predened rules cannot apply, and it becomes intractable to set separate rules for every permutation of different scenarios. AI is benecial in that it can evaluate the plan in the context that it learns from data.
Error detection in chart review can be seen as anomaly detectionidentication of the samples that deviate from the rest of the data. A qualitative approach applies data exploration techniques, such as clustering and principal component analysis (PCA), to visually identify outliers. An example is the study by Azmandian et al [47] who used the data exploration techniques to detect a gross error in the beam parameters for prostate four-eld box treatments (gure 12.4).
One possible quantitative approach is to learn a joint probability distribution of the parameters characterizing a radiotherapy plan (e.g. anatomical sites, fractiona­tion, treatment modality). If the probability model returns a small probability value for a test plan having a certain parameter set, the plan can be agged for possible presence of errors. A Bayesian network (BN) can calculate joint probability using: (i) a directed acyclic graph (DAG) representing dependent or independent relationships between variables and (ii) conditional probability values that can be learned from training data. BNs have been used for error detection in radiotherapy plans [4851]. Obtaining network topology for a BN model such as gure 12.5 remains a challenge, often requiring a group of experts to manually dene the dependency relationships.
12-5
Artificial Intelligence in Adaptive Radiation Therapy
Figure 12.4. Four-eld box prostate plans grouped into clusters based on their energy and monitor units. The plans are projected to a two-dimensional plane dened by the rst two principal components and clustered using k-means clustering. (Reproduced with permission from [
47]. Copyright 2007 IOP Publishing Ltd.)
Kalet et al [49] proposed the use of heuristics derived from radiation oncology ontology data to reduce the manual work in learning the network. A combination of data-driven and heuristic approaches can also be used for learning BN DAG [52]. In addition to BN, other machine-learning methods have been applied to detecting outlier radiotherapy plans. The forest-based method, used by Liu et al [ 53], makes it attractive for plan QA in that it can handle categorical variables and is robust to high-dimensional data [54].
When implementing these error detection models in clinic, a data drift, a change in data distribution over time, has to be considered. Data drift is prevalent in the medical eld [55], including radiation oncology, where treatment practices, such as fractionation schemes, constantly change over time [56]. Thus, these decision
12-6
Artificial Intelligence in Adaptive Radiation Therapy
Figure 12.5. A Bayesian network topology by representing the dependency relationships (represented by arrows) between the variables that characterize a treatment plan (represented by nodes). (Reproduced with permission from [
48]. Copyright 2015 Institute of Physics and Engineering in Medicine.)
support tools need be regularly monitored and re-calibrated to update the criteria discriminating erroneous plans, in a similar fashion to a study by Ruan et al [57]
12.2.2.3 Patient-specific QA
Pre-treatment PSQA was introduced to ensure that IMRT plans involving complex movements of multi-leaf collimators (MLC) are properly calculated, transferred to a machine, and accurately delivered [58]. PSQA is carried out by delivering a plan to a detector and evaluating the agreement between calculation and measurement. The agreement is typically computed using two-dimensional gamma analysis [59]to accommodate two-dimensional detectors. A plan is regarded satisfactory if the gamma passing rate (GPR) is above the xed threshold: TG-218 [60] recommends a passing rate of 95% under 3%/2 mm criteria. The power of AI can be harnessed to predict GPR from IMRT QA even before delivering the QA plan, based on plan, machine, and measurement device characteristics. Such a prediction system has three potential clinical benets, as follows. (i) Plan-specic issues, such as the over­optimization use of TPS beyond its capability, can be identied and mitigated ahead of time, thereby minimizing interruption in patient care. (ii) AI can be used to establish a GPR threshold specic to a certain type of plan, enhancing the sensitivity and specicity of the PSQA. (iii) The prediction model can illuminate systematic factors in TPS or delivery systems that can lead to dose discrepancy between calculation and measurement.
12-7
Artificial Intelligence in Adaptive Radiation Therapy
Figure 12.6. Left: Clustering of the examined plans using PCA to demonstrate the plans with low passing rate (circled) can be isolated by a combination of plan characteristics. Right: Agreement between the actual gamma passing rate and prediction by the virtual IMRT QA model. (Reproduced from [ Wiley & Sons. Copyright 2016 American Association of Physicists in Medicine.)
70] with permission from John
Several plan characteristics have been proposed as predictors of PSQA. First of all, IMRT or VMAT plan complexity metrics that quantify the irregularities in eld apertures, MLC motion [6163], or a uence map [64], were shown to be correlated with robust deliverability [65] and PSQA GPR [66], and thus have been used for many IMRT QA prediction models. In addition, dosiomics analysis—texture analysis of dose distribution [67]—was proposed as another indicator of plan complexity that can contribute to dose delivery inaccuracies [68], and was shown to have benets in predicting IMRT QA results [69]. Many machine-learning methods were used for putting together these predictive features into a PSQA prediction model. A seminal work by Valdes et al [70], coined as ‘virtual IMRT QA’ built a Poisson regression-based prediction model that predicts PSQA GPR from 78 plan complexity metrics (gure 12.6). This model was validated in multiple external datasets [71]. Since their work, different machine-learning methods were investigated to enhance GPR prediction, such as CNN [7274], support vector machine [75], boosted trees [76], and random forest [69, 77].
Several studies highlight the limitation of PSQA, especially in its reliance on GPR, due to lack of sensitivity [78, 79], correlation with dose to targets and OARs [80, 81], or independent audits [82]. One of the criticisms is the use of a xed GPR threshold for pass/fail, based on empirical evidence that the GPR from PSQA does not follow the Gaussian distribution [70, 83] as assumed for the TG-119 recom­mendation [58]. AI can be used to overcome the traditional one-size-ts-all approach to better explainthe 2D gamma distribution or GPR, by unraveling complex relationships between PSQA results and various sources of uncertainties, such as calculation engines, delivery systems, and detectors. For example, CNN has been used to predict MLC errors from features derived from EPID images [84]ora3D detector array [
85]. Another limitation of GPR is the loss of spatial information.
This neural network-based method was extended to detect other types of errors, such as phantom set-up errors [86], monitoring unit scaling errors [87], or MLC modeling in TPS [88].
12-8
Artificial Intelligence in Adaptive Radiation Therapy
Some investigators attempted to directly predict two-dimensional gamma dis­tributions by framing them as a deep-learning-based image synthesis problem. For example, Matsuura et al [89] used a generative adversarial network (GAN) to predict gamma distribution for EPID-based PSQA using the calculated uence map as input (gure 12.7). Similarly, Mahdavi et al [90] created an articial neural network (ANN)-based model to convert measured EPID uences to the equivalent 2D gamma agreement between TPS and 2D array measurement.
There are considerations when training or evaluating the machine-learning-based PSQA models. First, the data (whether GPR or voxels with gamma > 1) the models are trained on are highly skewed towards passing. To address the class imbalance, a large sample size is needed for model training, or remedies such as oversampling (SMOTE) have to be taken. Second, the model should be interpretable so that root causes for poor dose agreement are fed back to the planning or machine QA process as actionable information.
Figure 12.7. Examples of two-dimensional gamma distributions from IMRT QA synthesized by a GAN from an input uence map. (Reproduced from [ American Association of Physicists in Medicine.)
89] with permission from John Wiley & Sons. Copyright 2023
12-9