Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5545_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
15 Мб
Скачать
☆
14.6 Common Models ofAI
Fig. 14.3 Illustration of hierarchy of ANN
AI
ML
DL
RNN
257
14.6.1 Training, Validation andTesting
Most AI models, if not all, undergo three common steps for a successful application: training, validation, and testing, which are discussed below.
Training is an essential component of many AI platforms and involves feeding a large and diverse dataset on a specic topic of interest to the algorithm that learns the pattern of the data and makes an accurate prediction in response to a query. Dataset of interest is retrieved from the database of the specic task from commer­cial rms or created independently and fed into the AI algorithm. The model gets the training by scrounging through the dataset and recognizing the pattern in it. The training phase is like a teacher guiding a student on a problem-solving project. When a query is posed to the algorithm, it looks for the pattern it has learned during training and applies it to predict the answer to the query. As mentioned, the dataset is used in training, but for the sake of convenience and relatively high importance of the training step, dataset is normally disproportionately split like 80:20 or 90:10in favor of training against validation, whereas testing is often carried out using a sepa­rate dataset, except on occasion, a portion of the same dataset is used.
Following training, the model is subjected to the validation step, where the out­put of the training step is veried as to its validity. The data used for validation is a smaller portion of the dataset mentioned above. Validation of the AI model is a key step that ensures the AI model is reliable, safe, accurate, and secure. Datasets used include high-quality, relevant data in the validation step.
The testing phase veries if the output from the model is appropriate and accept­able by comparing it with a previously unseen dataset. Testing helps recognize the
258
14 Basics ofArticial Intelligence
data biases, imbalances, and errors prompting for corrections. The testing step can nd out how the model is working.

14.7 Machine Learning

Machine learning (ML) is an algorithmic subset of ANN in which the software uses known datasets to recognize their pattern and make responses to the questions asked. It distinguishes itself from ANN in that while ANN commonly solves com­plex problems, namely, summarizing documents, medical image analysis, or recog­nizing faces, with greater accuracy, ML is a broader category of algorithms that can learn from training data provided by humans without any explicit programming and make accurate decisions.
As mentioned, ML is one of many AI algorithms that require a massive dataset for proper training and validation. ML is a continuous upgrading process, as new information is added to the database, and it improves the answer to a question asked. It is like learning bicycle riding where after falling a few times starts understanding how to balance and pedal simultaneously, and even learns how to ride without hold­ing the bike handles. As more data are added, ML continues to improve signicantly in making predictions, and the ML model is evaluated to assess its ability to more accurately make predictions or decisions.
ML falls into two categories: supervised and unsupervised. In supervised ML systems, initially an extraction technique is employed in the ML algorithm to iden­tify the most useful features of the input data (an image or a video), which are then fed into the ML program to obtain the best model parameters (Apostolopoulos etal.
2023). In the supervised model, databases are labeled, meaning the data are prop-
erly formatted in tables and columns such that the pattern of the data is easily rec­ognizable by the machine to predict an accurate answer. Note that for an algorithm to learn the pattern of the dataset means what the information is about—is it a pic­ture of a tiger, is it about a text document, is it about a song, or is it about a social event? As mentioned above, it involves feeding an appropriate dataset to the com­puter for training purposes. In these circumstances, graphical processing units (GPUs) are critical components of computers to provide the enormous computing power to perform rapidly zillions of calculations. In the nal testing phase, the out­put predicted by the model is assessed for its appropriateness by comparing the output with a previously unknown dataset. Supervised models are easy to apply and ideal for spam detection, weather forecasting, pricing predictions, etc.
Unsupervised ML models, in contrast, deal with unlabeled data and require com­paratively much more data than supervised models to learn the pattern of the data. They uncover the hidden pattern on their own from the unlabeled data without human guidance, though at times they still require some human intervention to vali­date output variables. In unsupervised ML, however, the machine performs AI applications without any training or human guidance. An example of unsupervised ML is medical imaging, where data on scan images are unlabeled, and the machine
14.7 Machine Learning
259
learns on its own from different possibilities within the image data to make an inter­pretation without further direct supervision.
Common examples of machine learning are: image recognition such as facial recognition, autonomous self-driving cars, medical imaging to diagnose and plan for treatment, speech recognition to convert speech to text using natural language processing (NLP) (see below), scam detection, and use in chatbots (see later).

14.7.1 Decision Tree

A decision tree is a supervised machine learning algorithm used primarily for clas­sication and regression tasks. It displays a tree-like owchart made of nodes rep­resenting features, branches showing likely outcomes, and leaf nodes indicating the nal output. The algorithm looks for the best features in the dataset and uses them to split the tree based on the results. The process continues until a stopping condi­tion is met, such as reaching maximum depth or no further spitting, and all data are classied. Decision trees are easy to interpret and visualize, but overtting and small data changes may be causes for concern.

14.7.2 Random Forest

Random forest is a popular supervised machine learning algorithm commonly used for classication and regression tasks. Breiman (2001) is credited for its introduc­tion that creates many decision trees using different randomly chosen parts of the dataset. During the build-up of a tree, the algorithm randomly picks up a few fea­tures for use in splitting the dataset. Each decision tree is different from each other, thus making separate predictions based on what it learned from its part of the data. For classications, the decision trees “vote” to nalize the prediction by majority, whereas in regression, the average of all tree outputs is used.
The advantages of random forest are the high accuracy of the predictions because the nal output is better than the individual decision tree output. It works perfectly even if some data are missing or outliers. It performs well with big and complex datasets, since it needs a large amount of data for prediction. Since decision trees are independent of each other making their own predictions, the combined output becomes more accurate. Combining multiple decision trees helps reduce the risk of overtting due to averaging.
There are some disadvantages of the random forest paradigm. Because many decision trees are employed in the random forest model, it needs a large memory, making it slower in operation and computationally expensive. The output of the random forest model is more complex to interpret than that of individual decision trees due to the smaller size of the latter. Also, it is not ideal for real-time applica­tions where speed is critical.
Examples of random forest applications are medical diagnosis, predicting dis­eases, fraud detection, image, and speech classication.
260
14 Basics ofArticial Intelligence
Note The above Decision Tree and Random Forest sections are written based on the Outputs generated by OpenAI’s ChatGPT and Microsoft’s Copilot, an AI com­panion created by Microsoft.

14.7.3 Support Vector Machine

Support vector machine (SVM), introduced by Vapnik and Chervonenkis (1995), is a supervised machine learning algorithm that does an excellent task of classica­tion. The primary function of SVM is to nd an optimal hyperplane in an N-dimensional space to separate the data points into different classes by maximiz­ing the margin between them. It is accomplished by maximizing the distance between the hyperplane and the closest data points from each class (support vec- tors). SVM is good for distinguishing between spam and non- spam, image classi­cation and handwriting recognition. In nuclear medicine, SVM can assist in differentiating between infarct and ischemia on SPECT and PET images, based on its capability to separate myocardial perfusion image features like pixel intensity, texture, color, etc.
Essential features of SVM are support vectors, which are important for dening the hyperplane. Maximizing the margin between the classes improves the operation of SVM.The model works well with high-dimensional spaces that tend to prevent overtting. It has the disadvantage of a high cost of computing, and also is not ideal for large datasets impeding the speed of computation.

14.7.4 Computer Vision

Computer vision is a branch of articial intelligence that uses machine learning and neural networks to extract information from images, videos, and other inputs. CNN, transfer learning, and deep learning frameworks are commonly utilized for this pur­pose. Among the wide range of applications of computer vision are the real-world tasks, facial recognition, analysis of X-rays, MRIs in medical imaging, object detec­tion, robotic automation, and quality control in manufacturing, inventory manage­ment in retail business, and recognition of trafc signs and lights.
Since computer vision relies on real-world events, robust hardware is an essen­tial requirement. Low-quality components in the computer cause artifacts, and high­denition live video streaming quality and higher frame rates are essential in real-time processing in the computer vision algorithm. As mentioned before, GPUs are the heart and soul of AI, so are for computer vision. Nvidia GeForce GTX and AMD Radeon HD offer much faster and more accurate real-time processing by computer vision. Yet the hardware developers should strive for more powerful units, addressing the issues of low power consumption and high performance.
Since computer vision is data-hungry, requiring massive data, it is quite likely that mislabeled data and missing labels may creep into the dataset, causing dif­culty for the model to predict accurately. Noisy data with errors and overtting in

14.8 Deep Learning

261
the model cause confusion in the model to perform poorly. All these factors point to the fact that the quality of datasets should be a prime consideration in com­puter vision.
14.8 Deep Learning
Like ML, deep learning (DL) belongs to the broad category of ANN and is an upgraded variation of machine learning, which can perform more complex tasks using large volumes of data (Wang etal. 2020). It is used in speech recognition, chess playing, and patient care requiring deep thinking. It differs from supervised ML in that the algorithm uses raw (unstructured or unlabeled) data that are, unlike supervised ML, not organized into a database. The DL algorithm shes through the massive, unorganized data, learns the pattern of data, and makes accurate predic­tions without human help. While ML training takes a short time (a few seconds to hours), DL training requires more than a week due to the large number of parame­ters in the DL algorithm (Sarker 2021). However, DL testing time is much shorter than the ML testing time. In the testing stage, previously unseen data, like pictures or images, are fed into the trained model for classication, which is a relatively short process.
DL is useful in medical imaging, particularly in the acquisition of images and identication of abnormalities within them. Zhang and Qie (2023) have published an excellent review on the useful application of DL in medical imaging and its future potential. DL is a wide-ranging group of AI algorithms that can be described in three subcategories, namely convolutional neural networks (CNNs), recurrent neural networks (RNNs), and generative neural networks (GANs). The architecture and functional aspects of these models are briey described below.

14.8.1 Convolutional Network

CNN is a type of neural network introduced by Lecun et al. (1989) that is commonly used for image recognition and computer vision tasks. An excellent review by Yamashita etal. (2018) provides a detailed description of CNN on its basic concept and various applications. Structurally, a CNN consists of multiple layers of data that are broadly categorized into three groups, namely convolutional layers, pooling lay­ers, and fully connected layers (Fig.14.4).
Convolution mathematically means a combination of two functions to create a new function and executes a specialized linear operation for feature extraction from an image using a lter or kernel (a small matrix of weights). The input is assumed to be an array of arbitrary numbers, called a tensor. The kernel slides over the entire input tensor vertically and horizontally. An example of creating an output feature map is illustrated in Fig.14.5. The gure shows a kernel of size 2×2 (light green color) and an input tensor of size 4×4. As the kernel slides over the tensor image, it covers a 2×2 section of the tensor indicated by the light blue color. The product
262
Fig. 14.4 An example of a convolutional neural network (CNN) including multiple convolutions and pooling layers. (Used with permission of Springer Nature from Sarker (2021); permission conveyed through Copyright Clearance Center, Inc.)
14 Basics ofArticial Intelligence
of each element of the kernel and each corresponding element of the input tensor is calculated at every location of the input tensor and summed to obtain an output value at the corresponding location of the output tensor, termed a feature map (light orange color). The kernel can move one pixel, two pixels, or more at a time, which are termed stride 1, stride 2, etc. Stride 1 offers better resolution and is commonly used. A parameter “padding” is used where rows and columns of zeros are added on each side of the input tensor so as to t the center of the kernel on the outermost element. This ltering of the image with the kernel detects the specic features of the image, namely edges, corners, and textures, resulting in a feature map. Next, this feature map is fed as input into the next layer, and the process continues until a pat­tern is learned and the convolution stops.
Following convolution operation, the feature map is then fed into the pooling layer of the CNN structure to reduce the dimension of the input image for the ef­ciency of operation. Max pooling is the most common form of pooling operation. The process is called the downsampling of data points on the image, which means reducing the pixel points in the image. While the pooling process of data has the drawback of losing some information from the image due to loss of pixel points, it offers an efcient technique for object detection and image classication.
Following the pooling operation, the feature map is attened, i.e., transformed into a one-dimensional array of numbers and fed into the fully connected layer, which nalizes it to an output. This facilitates the scope of the fully connected layer to classify images based on the features extracted in the previous layers. It integrates the various features from previous convolutional and pooling layers and formats them in a map for classication and detection (Wang etal. 2020). The application of CNN in a clinical case is presented in Fig.14.6.
Many CNN architectures such as LeNet (Lecun etal. 1998), AlexNet (Krizhevsky etal. 2017), DenseNet (Huang etal. 2017), etc. have been devised and successfully applied to various AI tasks such as image segmentation, classication, detection, and registration (Greenspan etal. 2016). CNN is useful in face recognition, medical image analysis (Tajbakhsh et al. 2016), road and trafc sign recognition, (automa­tion of vehicles), text analysis, and creative artwork. All these CNN models are pretrained and available for application in a variety of essential tasks mentioned
14.8 Deep Learning
263
Fig. 14.5 Feature map graphically illustrates the primary calculations executed at each step. In this gure, the light green color represents the 2×2 kernel, while the light blue color represents the similar-sized area of the input image. Both are multiplied; the end result after summing up the resulting product values (marked in a light orange color) represents an entry value to the output feature map. (Reproduced from Alzubaidi etal. (2021). This is an open access article)
264
Fig. 14.6 Basic structure of CNN, in which network extracts radiomic features, produces a con­volution function, pools data through a kernel, and attens the pooled feature map for input into fully connected hidden layers of the neural network. (ReLU=rectied linear unit). This research was originally published in JNMT. (Reproduced with permission from Currie (2019b). © SNMMI)
14 Basics ofArticial Intelligence
above. When prompted for background information on these pre-trained CNN mod­els, ChatGPT generated the Table14.2 given below. Note that U-Net is not pre­trained, although its encoder and decoder are pretrained (Ronneberger et al. 2015).

14.8.2 Recurrent Neural Network

RNNs are a group of DL algorithm models, which are commonly used to perform tasks that require processing and converting sequential time-dependent input data to a sequential data output. The architecture of RNN is similar to that of ANN, consist­ing of input nodes, hidden layers, and an output. RNN undergoes training following the same principle using weights, forward propagation, backpropagation, and itera­tion. However, it has an additional feature of self-looping or a recurrent workow. Examples are natural language processing (speech), translating a text from one lan­guage to another, and video analysis (Padoy 2012). The key feature of RNN is its ability to remember the information from the previous steps in the sequence and use it to make future predictions.
Endoscopic videos are successfully analyzed by using an RNN.In this instance, RNN tracks and analyzes temporal changes in the videos, allowing the surgeon to identify abnormal tissues and ultimately to remove them by surgical incision. Although RNN is successfully applied in specic tasks such as image captioning and video analysis, its deployment has been limited. Currently, RNNs are being replaced by more efcient programs like transformer-based AI (ChatGPT) and large language model (LLM). Because of this, a detailed discussion of RNN is omitted.
14.8 Deep Learning
265
Table 14.2
CNN Year Key feature LeNet 1998 One of the earliest CNNs, designed for digit
AlexNet 2012 Deeper and more powerful than LeNet ResNet 2015 Introduced residual connections to allow deeper
DenseNet 2016 Introduced dense connectivity between layers
GoogleNet 2014 Introduced the Inception structure focusing on
MobileNet 2017 Designed for mobile and an embedded vision
U-Net 2015 Mainly used for image segmentation, Not pretrained,
a
Provided by ChatGPT of Open AI in response to my query; “Brief history of pre-trained CNN
models?”
Some useful pre-trained CNN models
recognition (MNIST)
networks
computational efciency
application
but its encoder and decoder are pretrained
a
Pre-trained models available?
✓ (on MNIST, usually)
✓ (on ImageNet) ✓ (e.g., ResNet-18,
ResNet-50 on ImageNet)
✓ (on ImageNet and more)
✓ (on ImageNet) and depth

14.8.3 Generative Adversarial Network

A generative adversarial network (GAN) is a DL AI model based on neural net­works rst introduced by Goodfellow etal. in 2014. GAN has two parts: (1) the generator (generative) that produces new data following modication of input data, and (2) the discriminator that veries if the output is correct against the actual input. This process of generating new data from input data and verication by the dis­criminator of the veracity of the new data continues until the generated data is indis­tinguishable from the real data. For example, imagine a counterfeiter trying to produce fake currency and an FBI agent trying to detect it. As both try to improve their strategy, fake money and real money become indistinguishable.
The training is adversarial, since the generator tries to fool the discriminator, whereas the discriminator tries to detect fakes. The two networks are trained simul­taneously using a minimax optimization framework (not discussed further here due to complexity and referred to other references) and forward-propagation and back­propagation as in an ANN.The goal of training is to have (1) the generator minimize the probability of fake diagnosis by the discriminator and (2) the discriminator maximize the probability of making a correct distinction between real and fake samples. Because the two networks constantly compete with each other, training the model often becomes difcult.
GAN is applied to generate realistic artwork, create human-like voices, and help improve medical imaging and fashion design. It is also used to create synthetic data for ML training and synthetic medical images for training and research (Frid-Adar etal. 2018). They can also create deepfakes of seemingly realistic images and vid­eos that are used in movies and videos.
266
14 Basics ofArticial Intelligence

14.8.4 Transfer Learning

A variation in DL called transfer learning is employed where a pre-trained model like ImageNet for images, or GPT for language is reused for a new task using the knowledge learned from the initial training (Pan and Yang 2010). The pre-trained model has already learned to recognize edges, shapes, textures, etc. of an input image and the last few layers need to be trained as in ANN.This technique is useful in cases where data are limited making training faster. Transfer learning has been successfully applied to medical imaging in the diagnosis and classication of diseases.

14.9 Radiomics

Radiomics is a technique to retrieve quantitative features from medical images such as CT, ultrasound, MRI, or PET images, which describe their characteristics, namely, shape, volume, density, and texture of normal or abnormal tissues (Koçak etal. 2019). In common practice, these features are assessed visually by the practi­tioners, which may lead to errors due to intra- and inter-observer variations along with missing some hidden features in the images. Also, radiomics discovers patterns in the image not visible to the human eye, which are processed by an applicable AI algorithm to provide a clinical outcome.
Applying AI in radiomics plays a crucial role in unfolding these features, thus facilitating image interpretation and disease diagnosis more accurately. Several AI models can be applied in radiomics, namely ML, DL, Random forests, Support vec­tor machines (SVM), which are discussed earlier. Currently, there are several well­designed AI programs to extract features for use in radiomics, but they are not as popular as expected. Also, one can use coding techniques to extract essential fea­tures of images from two popular platforms, MATLAB and Python, which have vast libraries of these features.
Many quantitative features are extracted from the 2D or 3D SPECT or PET images, which can then be analyzed by AI algorithms to nd correlations with cer­tain diseases (Song etal. 2020). To extract the characteristic features, images are often segmented for the convenience of easy analysis. Target tissues are segmented, whereby regions of interest (ROI) are created, and features are identied by apply­ing radiomics. Radiomics is commonly used in cancer diagnosis, but it has been applied to other diseases, too. Despite its useful contribution to image analysis, it faces several challenges. Different institutions use different scanners and different scanning protocols, which leads to variations in imaging results, thus compromising the accuracy of the extracted features and affecting the strength of the trained mod­els across different centers. Noise, artifacts, and inter-observer variations in manual segmentation add another challenge to the operation of radiomics.