Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5545_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Preface
- •Contents
- •1: Structure of Matter
- •2: Radioactive Decay
- •2.1 Spontaneous Fission
- •1.1.1 Radiation
- •1.2 The Atom
- •1.2.3 Nuclear Binding Energy
- •1.3 Nuclear Nomenclature
- •1.5 Questions
- •Suggested Readings
- •2.2 Isomeric Transition
- •2.2.1 Gamma (γ)-Ray Emission
- •2.2.2 Internal Conversion
- •2.2.2.1 Problem 2.1
- •2.2.2.2 Answer
- •2.3 Alpha (α)-Decay
- •2.4 Beta (β−)-Decay
- •2.5 Positron (β+)-Decay
- •2.6 Electron Capture
- •2.7 Questions
- •Suggested Readings
- •3.1 Radioactive Decay Equation
- •3.1.1 General Equation
- •3.1.2 Half-Life
- •3.1.3 Mean Life
- •3.1.4 Effective Half-Life
- •3.2 Units of Radioactivity
- •3.3 Specific Activity
- •3.4 Calculation
- •3.5 Successive Decay Equations
- •3.5.1 General Equation
- •3.5.2 Transient Equilibrium
- •3.5.3 Secular Equilibrium
- •3.6 Questions
- •Suggested Readings
- •4.5 Poisson Distribution
- •4.6 Gaussian Distribution
- •4.7 Chi-Square Test
- •4.8 Minimum Detectable Activity
- •4.10 Questions
- •Suggested Readings
- •5.1 Cyclotron-Produced Radionuclides
- •5.2 Reactor-Produced Radionuclides
- •5.2.1 Fission or (n, f) Reaction
- •5.2.2 Neutron Capture or (n, γ) Reaction
- •5.6 Radionuclide Generators
- •5.8 Questions
- •Suggested Readings
- •6.1.1 Specific Ionization
- •6.1.2 Linear Energy Transfer
- •6.1.3 Range
- •6.1.4 Bremsstrahlung
- •6.1.5 Positron Annihilation
- •6.2.1.1 Photoelectric Effect
- •6.2.1.2 Compton Scattering
- •6.2.1.3 Pair Production
- •6.2.1.4 Raleigh Scattering
- •6.2.1.5 Photodisintegration
- •6.3.2 Half-Value Layer
- •6.5 Questions
- •Suggested Readings
- •7: Gas-Filled Detector
- •7.1 Principles of Gas-Filled Detector
- •7.2 Ionization Chamber
- •7.2.1 Ion Chamber Survey Meter
- •7.2.2 Dose Calibrator
- •7.2.2.1 Constancy
- •7.2.2.2 Accuracy
- •7.2.2.3 Linearity
- •7.2.2.4 Geometry
- •7.2.3 Pocket Dosimeter
- •7.3 Proportional Counter
- •7.4 Geiger–Müller Counter
- •7.5 Questions
- •Suggested Readings
- •8.1 Scintillation Counter
- •8.4.3 Characteristic X-Ray Peak
- •8.4.4 Backscatter Peak
- •8.4.5 Iodine Escape Peak
- •8.2 Solid Scintillation Detector
- •8.2.1 NaI (Tl) Detector
- •8.2.2 Bismuth Germanate Detector
- •8.2.3 Barium Fluoride Detector
- •8.2.4 Lutetium Oxyorthosilicate Detector
- •8.2.5 Gadolinium Oxyorthosilicate Detector
- •8.2.6 Yttrium Oxyorthosilicate Detector
- •8.2.7 Yttrium Aluminum Perovskite Detector
- •8.2.8 Lutetium Yttrium Oxyorthosilicate Detector
- •8.2.9 Lanthanum Bromide Detector
- •8.3 Solid-State Detector
- •8.3.2 Cadmium–Zinc–Tellurium Detector
- •8.3.3 Cesium Iodide (CsI(Tl)) Detector
- •8.3.4 Solid Scintillation Counter
- •8.3.4.1 NaI(Tl) Detector
- •8.3.4.2 Photomultiplier Tube
- •8.3.4.3 Preamplifier
- •8.3.4.4 Linear Amplifier
- •8.3.4.5 Pulse-Height Analyzer
- •8.3.4.6 Display or Storage
- •8.4 Gamma-Ray Spectrometry
- •8.4.1 Photopeak
- •8.4.6 Positron Annihilation Peak
- •8.4.7 Coincidence Peak
- •8.5 Liquid Scintillation Counter
- •8.5.1 Quenching
- •8.6.1 Energy Resolution
- •8.6.2 Detection Efficiency
- •8.6.2.1 Intrinsic Efficiency
- •8.6.2.2 Photopeak Efficiency or Photofraction
- •8.6.2.3 Geometric Efficiency
- •8.6.3 Dead Time
- •8.7 Gamma Well Counter
- •8.8 Thyroid Probe
- •8.8.1 Thyroid Uptake Measurement
- •8.9 Questions
- •Suggested Readings
- •9: Gamma Camera
- •9.1 Gamma Camera
- •9.1.2 Detector
- •9.1.3 Collimator
- •9.1.4 Photomultiplier Tube
- •9.1.5 X-, Y-Positioning Circuit
- •9.1.6 Pulse-Height Analyzer
- •9.2 Digital Camera
- •9.2.1 Solid State Digital Camera
- •9.3 Questions
- •Suggested Readings
- •10.1.1 Spatial Resolution
- •10.1.1.1 Intrinsic Resolution
- •10.1.1.2 Collimator Resolution
- •10.1.1.3 Scatter Resolution
- •10.1.2.1 Bar Phantom
- •10.1.2.2 Line-Spread Function
- •10.1.2.3 Modulation Transfer Function
- •10.1.3 Sensitivity
- •10.1.3.1 Collimator Efficiency
- •10.1.4 Uniformity
- •10.1.5 Pulse-Height Variation
- •10.1.6 Nonlinearity
- •10.1.7 Edge Packing
- •10.2 Gamma Camera Tuning
- •10.4 Contrast
- •10.4.1 Count Density
- •10.4.2 Image Noise
- •10.4.4 High Count Rate
- •10.4.6 Patient Motion
- •10.5.1 Daily Checks
- •10.5.1.2 Uniformity
- •10.5.2 Weekly Checks
- •10.5.3 Monthly Checks
- •10.5.3.1 High-Count Uniformity Calibration
- •10.5.3.2 Collimator Integrity
- •10.5.4 Annual, Semiannual, or As-Needed Checks
- •10.6 Questions
- •References and Suggested Readings
- •11.1.1 Central Processing Unit
- •11.1.2 Computer Memory
- •11.1.3 External Storage Device
- •11.1.4 Input/Output Device
- •11.1.7 Digital-to-Analog Conversion
- •11.1.8 Digital Image
- •11.2.1 Digital Data Acquisition
- •11.2.2 Static Study
- •11.2.3 Dynamic Study
- •11.2.4 Gated Study
- •11.2.7 Display
- •11.3.1 PACS
- •11.4 Questions
- •Suggested Readings
- •12: Single Photon Emission Computed Tomography
- •12.1 Tomographic Imaging
- •12.2 Single Photon Emission Computed Tomography
- •12.2.1 Data Acquisition
- •12.2.2 Image Reconstruction
- •12.2.2.1 Simple Backprojection
- •12.2.2.2 Filtered Backprojection
- •12.2.2.3 The Convolution Method
- •12.2.2.4 The Fourier Method
- •12.2.2.6 Iterative Reconstruction
- •12.3 SPECT/CT Scanner
- •12.4 Factors Affecting SPECT
- •12.4.1 Photon Attenuation
- •12.4.2 Attenuation Correction Methods
- •12.5 Partial-Volume Effect
- •12.5.2 Sampling
- •12.5.3 Scattering
- •12.6.1 Spatial Resolution
- •12.6.2 Sensitivity
- •12.6.3 Other Parameters
- •12.7.1 Daily Tests
- •12.7.2 Weekly Tests
- •12.7.2.1 Spatial Resolution
- •12.9 Questions
- •References and Suggested Readings
- •13: Positron Emission Tomography
- •13.1 Introduction
- •13.2 PET Radiopharmaceuticals
- •13.3.2 Block Detector
- •13.5 Coincidence Timing Window
- •13.6 PET/CT Scanner
- •13.7 PET/MR Scanner
- •13.7.2 MR Scanner
- •13.7.3 Commercial PET/MR Scanner
- •13.8 Mobile PET or PET/CT Scanner
- •13.9 Micro-PET Scanner
- •13.11 Data Acquisition
- •13.12 Image Reconstruction
- •13.13 Factors Affecting PET
- •13.13.1 Normalization
- •13.13.2 Photon Attenuation Correction
- •13.13.4 Random Coincidences
- •13.13.5 Scatter Coincidences
- •13.13.6 Dead Time
- •13.13.7 Radial Elongation
- •13.14.1 Spatial Resolution
- •13.14.2 Sensitivity
- •13.14.2.1 Noise Equivalent Count Rate
- •13.15.1 Daily Tests
- •13.15.1.1 Sinogram Check
- •13.15.2 Weekly Tests
- •13.15.2.1 Normalization
- •13.18 Questions
- •References and Suggested Reading
- •14.1 Background
- •14.5 Artificial Neural Network
- •14.7 Machine Learning
- •14.7.1 Decision Tree
- •14.7.2 Random Forest
- •14.7.3 Support Vector Machine
- •14.7.4 Computer Vision
- •14.8 Deep Learning
- •14.8.1 Convolutional Network
- •14.8.2 Recurrent Neural Network
- •14.8.3 Generative Adversarial Network
- •14.8.4 Transfer Learning
- •14.9 Radiomics
- •14.10 Natural Language Processing
- •14.11 Large Language Model
- •14.12 Generative Artificial Intelligence
- •14.13.1 Prompt
- •14.13.2 Token
- •14.13.3 Hallucination
- •14.13.4 Deepfake
- •14.13.5 Overfitting
- •14.15 Chatbot
- •14.18 Legal Implication
- •14.20 Questions
- •References
- •15.1 Introduction
- •15.2.1 Scheduling
- •15.2.2 Image Acquisition
- •15.2.3 Image Processing
- •15.2.4 Interpretation
- •15.2.5 Reporting
- •15.3.1 Oncology
- •15.3.2 Cardiovascular Disease
- •15.3.3 Bone Scintigraphy
- •15.3.4 Thyroid Imaging
- •15.5 Drug Development
- •15.6 Questions
- •References and Suggested Reading
- •16: Internal Radiation Dosimetry
- •16.1 Radiation Unit
- •16.1.1 Roentgen
- •16.1.2 Rad
- •16.1.3 Gray
- •16.1.4 Rem
- •16.1.5 Radiation Weighting Factor
- •16.1.6 Quality Factor
- •16.1.7 Sievert
- •16.2 Dose Calculation
- •16.2.1 Radiation Dose Rate
- •16.2.2 Cumulative Radiation Dose
- •16.2.3 Factors Affecting Ã
- •16.2.4 The S Values
- •16.4 Pediatric Dosage
- •16.5 Questions
- •References and Suggested Readings
- •17: Radiation Biology
- •17.1 The Cell
- •17.2.1 DNA Molecule
- •17.2.2 Chromosome
- •17.5 Cell Survival Curves
- •17.6 Factors Affecting Radiosensitivity
- •17.6.1 Dose Rate
- •17.6.2 Linear Energy Transfer
- •17.6.4 Chemicals
- •17.7 Radiosensitizer
- •17.7.1 Oxygen
- •17.7.2 Pyrimidine
- •17.7.3 Others
- •17.8 Radioprotector
- •17.9 Apoptosis
- •17.13.1 Hematopoietic Syndrome
- •17.13.2 Gastrointestinal Syndrome
- •17.13.3 Cerebrovascular Syndrome
- •17.14.1 Somatic Effects
- •17.14.1.1 Carcinogenesis
- •17.14.1.3 Dose–Response Relationship
- •17.14.1.5 Leukemia
- •17.14.1.6 Breast Cancer
- •17.14.1.7 Other Cancers
- •17.14.1.10 Nonspecific Life-Shortening
- •17.14.1.11 Cataractogenesis
- •17.14.2 Genetic Effects
- •17.14.2.1 Spontaneous Mutation
- •17.14.2.2 Doubling Dose
- •17.14.2.3 Genetically Significant Dose
- •17.17 Questions
- •References and Suggested Readings
- •18.1 Introduction
- •18.2 Radiation Protection
- •18.2.3 Occupational Dose Limits
- •18.2.4 ALARA Program
- •18.2.5.1 Time
- •18.2.5.2 Distance
- •18.2.5.3 Shielding
- •18.2.5.4 Activity
- •18.2.6 Personnel Monitoring
- •18.2.6.1 Film Badge
- •18.2.6.2 Thermoluminescent Dosimeter
- •18.2.6.3 Optically Stimulated Luminescence Dosimeter
- •18.3 Radiation Regulations
- •18.3.1 License
- •18.3.1.1 General License
- •18.3.1.2 Specific License of Limited Scope
- •18.3.1.3 Specific Licenses of Broad Scope
- •18.3.2 Radiation Safety Committee
- •18.3.3 Radiation Safety Officer
- •18.3.4.3 Supervision
- •18.3.4.4 Mobile Nuclear Medicine Service
- •18.3.4.5 Written Directives
- •18.4 Bioassay
- •18.6 Radioactive Waste Disposal
- •18.6.2 Release into Sewerage Systems
- •18.6.4 Other Disposal Methods
- •18.7 Radioactive Spill
- •18.8 Recordkeeping
- •18.10 Dirty Bombs
- •18.11 Types of Accidental Radiation Exposure
- •18.12 Protective Measures in Case of Explosion of a Dirty Bomb
- •18.13 Verification Card for Radioactive Patients
- •18.14 Radiation Phobia
- •18.15 European Regulations Governing Radiation
- •18.16 Questions
- •References and Suggested Readings
- •Index

14.6 Common Models ofAI
Fig. 14.3 Illustration of
hierarchy of ANN
AI
ML
DL
RNN
257
14.6.1 Training, Validation andTesting
Most AI models, if not all, undergo three common steps for a successful application:
training, validation, and testing, which are discussed below.
Training is an essential component of many AI platforms and involves feeding a
large and diverse dataset on a specic topic of interest to the algorithm that learns
the pattern of the data and makes an accurate prediction in response to a query.
Dataset of interest is retrieved from the database of the specic task from commercial rms or created independently and fed into the AI algorithm. The model gets
the training by scrounging through the dataset and recognizing the pattern in it. The
training phase is like a teacher guiding a student on a problem-solving project.
When a query is posed to the algorithm, it looks for the pattern it has learned during
training and applies it to predict the answer to the query. As mentioned, the dataset
is used in training, but for the sake of convenience and relatively high importance of
the training step, dataset is normally disproportionately split like 80:20 or 90:10in
favor of training against validation, whereas testing is often carried out using a separate dataset, except on occasion, a portion of the same dataset is used.
Following training, the model is subjected to the validation step, where the output of the training step is veried as to its validity. The data used for validation is a
smaller portion of the dataset mentioned above. Validation of the AI model is a key
step that ensures the AI model is reliable, safe, accurate, and secure. Datasets used
include high-quality, relevant data in the validation step.
The testing phase veries if the output from the model is appropriate and acceptable by comparing it with a previously unseen dataset. Testing helps recognize the

258
14 Basics ofArticial Intelligence
data biases, imbalances, and errors prompting for corrections. The testing step can
nd out how the model is working.
14.7 Machine Learning
Machine learning (ML) is an algorithmic subset of ANN in which the software uses
known datasets to recognize their pattern and make responses to the questions
asked. It distinguishes itself from ANN in that while ANN commonly solves complex problems, namely, summarizing documents, medical image analysis, or recognizing faces, with greater accuracy, ML is a broader category of algorithms that can
learn from training data provided by humans without any explicit programming and
make accurate decisions.
As mentioned, ML is one of many AI algorithms that require a massive dataset
for proper training and validation. ML is a continuous upgrading process, as new
information is added to the database, and it improves the answer to a question asked.
It is like learning bicycle riding where after falling a few times starts understanding
how to balance and pedal simultaneously, and even learns how to ride without holding the bike handles. As more data are added, ML continues to improve signicantly
in making predictions, and the ML model is evaluated to assess its ability to more
accurately make predictions or decisions.
ML falls into two categories: supervised and unsupervised. In supervised ML
systems, initially an extraction technique is employed in the ML algorithm to identify the most useful features of the input data (an image or a video), which are then
fed into the ML program to obtain the best model parameters (Apostolopoulos etal.
2023). In the supervised model, databases are labeled, meaning the data are prop-
erly formatted in tables and columns such that the pattern of the data is easily recognizable by the machine to predict an accurate answer. Note that for an algorithm
to learn the pattern of the dataset means what the information is about—is it a picture of a tiger, is it about a text document, is it about a song, or is it about a social
event? As mentioned above, it involves feeding an appropriate dataset to the computer for training purposes. In these circumstances, graphical processing units
(GPUs) are critical components of computers to provide the enormous computing
power to perform rapidly zillions of calculations. In the nal testing phase, the output predicted by the model is assessed for its appropriateness by comparing the
output with a previously unknown dataset. Supervised models are easy to apply and
ideal for spam detection, weather forecasting, pricing predictions, etc.
Unsupervised ML models, in contrast, deal with unlabeled data and require comparatively much more data than supervised models to learn the pattern of the data.
They uncover the hidden pattern on their own from the unlabeled data without
human guidance, though at times they still require some human intervention to validate output variables. In unsupervised ML, however, the machine performs AI
applications without any training or human guidance. An example of unsupervised
ML is medical imaging, where data on scan images are unlabeled, and the machine

14.7 Machine Learning
259
learns on its own from different possibilities within the image data to make an interpretation without further direct supervision.
Common examples of machine learning are: image recognition such as facial
recognition, autonomous self-driving cars, medical imaging to diagnose and plan
for treatment, speech recognition to convert speech to text using natural language
processing (NLP) (see below), scam detection, and use in chatbots (see later).
14.7.1 Decision Tree
A decision tree is a supervised machine learning algorithm used primarily for classication and regression tasks. It displays a tree-like owchart made of nodes representing features, branches showing likely outcomes, and leaf nodes indicating the
nal output. The algorithm looks for the best features in the dataset and uses them
to split the tree based on the results. The process continues until a stopping condition is met, such as reaching maximum depth or no further spitting, and all data are
classied. Decision trees are easy to interpret and visualize, but overtting and
small data changes may be causes for concern.
14.7.2 Random Forest
Random forest is a popular supervised machine learning algorithm commonly used
for classication and regression tasks. Breiman (2001) is credited for its introduction that creates many decision trees using different randomly chosen parts of the
dataset. During the build-up of a tree, the algorithm randomly picks up a few features for use in splitting the dataset. Each decision tree is different from each other,
thus making separate predictions based on what it learned from its part of the data.
For classications, the decision trees “vote” to nalize the prediction by majority,
whereas in regression, the average of all tree outputs is used.
The advantages of random forest are the high accuracy of the predictions because
the nal output is better than the individual decision tree output. It works perfectly
even if some data are missing or outliers. It performs well with big and complex
datasets, since it needs a large amount of data for prediction. Since decision trees are
independent of each other making their own predictions, the combined output
becomes more accurate. Combining multiple decision trees helps reduce the risk of
overtting due to averaging.
There are some disadvantages of the random forest paradigm. Because many
decision trees are employed in the random forest model, it needs a large memory,
making it slower in operation and computationally expensive. The output of the
random forest model is more complex to interpret than that of individual decision
trees due to the smaller size of the latter. Also, it is not ideal for real-time applications where speed is critical.
Examples of random forest applications are medical diagnosis, predicting diseases, fraud detection, image, and speech classication.

260
14 Basics ofArticial Intelligence
Note The above Decision Tree and Random Forest sections are written based on
the Outputs generated by OpenAI’s ChatGPT and Microsoft’s Copilot, an AI companion created by Microsoft.
14.7.3 Support Vector Machine
Support vector machine (SVM), introduced by Vapnik and Chervonenkis (1995), is
a supervised machine learning algorithm that does an excellent task of classication. The primary function of SVM is to nd an optimal hyperplane in an
N-dimensional space to separate the data points into different classes by maximizing the margin between them. It is accomplished by maximizing the distance
between the hyperplane and the closest data points from each class (support vec-
tors). SVM is good for distinguishing between spam and non- spam, image classication and handwriting recognition. In nuclear medicine, SVM can assist in
differentiating between infarct and ischemia on SPECT and PET images, based on
its capability to separate myocardial perfusion image features like pixel intensity,
texture, color, etc.
Essential features of SVM are support vectors, which are important for dening
the hyperplane. Maximizing the margin between the classes improves the operation
of SVM.The model works well with high-dimensional spaces that tend to prevent
overtting. It has the disadvantage of a high cost of computing, and also is not ideal
for large datasets impeding the speed of computation.
14.7.4 Computer Vision
Computer vision is a branch of articial intelligence that uses machine learning and
neural networks to extract information from images, videos, and other inputs. CNN,
transfer learning, and deep learning frameworks are commonly utilized for this purpose. Among the wide range of applications of computer vision are the real-world
tasks, facial recognition, analysis of X-rays, MRIs in medical imaging, object detection, robotic automation, and quality control in manufacturing, inventory management in retail business, and recognition of trafc signs and lights.
Since computer vision relies on real-world events, robust hardware is an essential requirement. Low-quality components in the computer cause artifacts, and highdenition live video streaming quality and higher frame rates are essential in
real-time processing in the computer vision algorithm. As mentioned before, GPUs
are the heart and soul of AI, so are for computer vision. Nvidia GeForce GTX and
AMD Radeon HD offer much faster and more accurate real-time processing by
computer vision. Yet the hardware developers should strive for more powerful units,
addressing the issues of low power consumption and high performance.
Since computer vision is data-hungry, requiring massive data, it is quite likely
that mislabeled data and missing labels may creep into the dataset, causing difculty for the model to predict accurately. Noisy data with errors and overtting in

14.8 Deep Learning
261
the model cause confusion in the model to perform poorly. All these factors point to
the fact that the quality of datasets should be a prime consideration in computer vision.
14.8 Deep Learning
Like ML, deep learning (DL) belongs to the broad category of ANN and is an
upgraded variation of machine learning, which can perform more complex tasks
using large volumes of data (Wang etal. 2020). It is used in speech recognition,
chess playing, and patient care requiring deep thinking. It differs from supervised
ML in that the algorithm uses raw (unstructured or unlabeled) data that are, unlike
supervised ML, not organized into a database. The DL algorithm shes through the
massive, unorganized data, learns the pattern of data, and makes accurate predictions without human help. While ML training takes a short time (a few seconds to
hours), DL training requires more than a week due to the large number of parameters in the DL algorithm (Sarker 2021). However, DL testing time is much shorter
than the ML testing time. In the testing stage, previously unseen data, like pictures
or images, are fed into the trained model for classication, which is a relatively
short process.
DL is useful in medical imaging, particularly in the acquisition of images and
identication of abnormalities within them. Zhang and Qie (2023) have published
an excellent review on the useful application of DL in medical imaging and its
future potential. DL is a wide-ranging group of AI algorithms that can be described
in three subcategories, namely convolutional neural networks (CNNs), recurrent
neural networks (RNNs), and generative neural networks (GANs). The architecture
and functional aspects of these models are briey described below.
14.8.1 Convolutional Network
CNN is a type of neural network introduced by Lecun et al. (1989) that is commonly
used for image recognition and computer vision tasks. An excellent review by
Yamashita etal. (2018) provides a detailed description of CNN on its basic concept
and various applications. Structurally, a CNN consists of multiple layers of data that
are broadly categorized into three groups, namely convolutional layers, pooling layers, and fully connected layers (Fig.14.4).
Convolution mathematically means a combination of two functions to create a
new function and executes a specialized linear operation for feature extraction from
an image using a lter or kernel (a small matrix of weights). The input is assumed
to be an array of arbitrary numbers, called a tensor. The kernel slides over the entire
input tensor vertically and horizontally. An example of creating an output feature
map is illustrated in Fig.14.5. The gure shows a kernel of size 2×2 (light green
color) and an input tensor of size 4×4. As the kernel slides over the tensor image,
it covers a 2×2 section of the tensor indicated by the light blue color. The product

262
Fig. 14.4 An example of a convolutional neural network (CNN) including multiple convolutions
and pooling layers. (Used with permission of Springer Nature from Sarker (2021); permission
conveyed through Copyright Clearance Center, Inc.)
14 Basics ofArticial Intelligence
of each element of the kernel and each corresponding element of the input tensor is
calculated at every location of the input tensor and summed to obtain an output
value at the corresponding location of the output tensor, termed a feature map (light
orange color). The kernel can move one pixel, two pixels, or more at a time, which
are termed stride 1, stride 2, etc. Stride 1 offers better resolution and is commonly
used. A parameter “padding” is used where rows and columns of zeros are added on
each side of the input tensor so as to t the center of the kernel on the outermost
element. This ltering of the image with the kernel detects the specic features of
the image, namely edges, corners, and textures, resulting in a feature map. Next, this
feature map is fed as input into the next layer, and the process continues until a pattern is learned and the convolution stops.
Following convolution operation, the feature map is then fed into the pooling
layer of the CNN structure to reduce the dimension of the input image for the efciency of operation. Max pooling is the most common form of pooling operation.
The process is called the downsampling of data points on the image, which means
reducing the pixel points in the image. While the pooling process of data has the
drawback of losing some information from the image due to loss of pixel points, it
offers an efcient technique for object detection and image classication.
Following the pooling operation, the feature map is attened, i.e., transformed
into a one-dimensional array of numbers and fed into the fully connected layer,
which nalizes it to an output. This facilitates the scope of the fully connected layer
to classify images based on the features extracted in the previous layers. It integrates
the various features from previous convolutional and pooling layers and formats
them in a map for classication and detection (Wang etal. 2020). The application of
CNN in a clinical case is presented in Fig.14.6.
Many CNN architectures such as LeNet (Lecun etal. 1998), AlexNet (Krizhevsky
etal. 2017), DenseNet (Huang etal. 2017), etc. have been devised and successfully
applied to various AI tasks such as image segmentation, classication, detection,
and registration (Greenspan etal. 2016). CNN is useful in face recognition, medical
image analysis (Tajbakhsh et al. 2016), road and trafc sign recognition, (automation of vehicles), text analysis, and creative artwork. All these CNN models are
pretrained and available for application in a variety of essential tasks mentioned

14.8 Deep Learning
263
Fig. 14.5 Feature map graphically illustrates the primary calculations executed at each step. In
this gure, the light green color represents the 2×2 kernel, while the light blue color represents
the similar-sized area of the input image. Both are multiplied; the end result after summing up the
resulting product values (marked in a light orange color) represents an entry value to the output
feature map. (Reproduced from Alzubaidi etal. (2021). This is an open access article)

264
Fig. 14.6 Basic structure of CNN, in which network extracts radiomic features, produces a convolution function, pools data through a kernel, and attens the pooled feature map for input into
fully connected hidden layers of the neural network. (ReLU=rectied linear unit). This research
was originally published in JNMT. (Reproduced with permission from Currie (2019b). © SNMMI)
14 Basics ofArticial Intelligence
above. When prompted for background information on these pre-trained CNN models, ChatGPT generated the Table14.2 given below. Note that U-Net is not pretrained, although its encoder and decoder are pretrained (Ronneberger et al. 2015).
14.8.2 Recurrent Neural Network
RNNs are a group of DL algorithm models, which are commonly used to perform
tasks that require processing and converting sequential time-dependent input data to
a sequential data output. The architecture of RNN is similar to that of ANN, consisting of input nodes, hidden layers, and an output. RNN undergoes training following
the same principle using weights, forward propagation, backpropagation, and iteration. However, it has an additional feature of self-looping or a recurrent workow.
Examples are natural language processing (speech), translating a text from one language to another, and video analysis (Padoy 2012). The key feature of RNN is its
ability to remember the information from the previous steps in the sequence and use
it to make future predictions.
Endoscopic videos are successfully analyzed by using an RNN.In this instance,
RNN tracks and analyzes temporal changes in the videos, allowing the surgeon to
identify abnormal tissues and ultimately to remove them by surgical incision.
Although RNN is successfully applied in specic tasks such as image captioning
and video analysis, its deployment has been limited. Currently, RNNs are being
replaced by more efcient programs like transformer-based AI (ChatGPT) and large
language model (LLM). Because of this, a detailed discussion of RNN is omitted.

14.8 Deep Learning
265
Table 14.2
CNN Year Key feature
LeNet 1998 One of the earliest CNNs, designed for digit
AlexNet 2012 Deeper and more powerful than LeNet
ResNet 2015 Introduced residual connections to allow deeper
DenseNet 2016 Introduced dense connectivity between layers
GoogleNet 2014 Introduced the Inception structure focusing on
MobileNet 2017 Designed for mobile and an embedded vision
U-Net 2015 Mainly used for image segmentation, Not pretrained,
a
Provided by ChatGPT of Open AI in response to my query; “Brief history of pre-trained CNN
models?”
Some useful pre-trained CNN models
recognition (MNIST)
networks
computational efciency
application
but its encoder and decoder are pretrained
a
Pre-trained models
available?
✓ (on MNIST,
usually)
✓ (on ImageNet)
✓ (e.g., ResNet-18,
ResNet-50 on
ImageNet)
✓ (on ImageNet
and more)
✓ (on ImageNet)
and depth
14.8.3 Generative Adversarial Network
A generative adversarial network (GAN) is a DL AI model based on neural networks rst introduced by Goodfellow etal. in 2014. GAN has two parts: (1) the
generator (generative) that produces new data following modication of input data,
and (2) the discriminator that veries if the output is correct against the actual input.
This process of generating new data from input data and verication by the discriminator of the veracity of the new data continues until the generated data is indistinguishable from the real data. For example, imagine a counterfeiter trying to
produce fake currency and an FBI agent trying to detect it. As both try to improve
their strategy, fake money and real money become indistinguishable.
The training is adversarial, since the generator tries to fool the discriminator,
whereas the discriminator tries to detect fakes. The two networks are trained simultaneously using a minimax optimization framework (not discussed further here due
to complexity and referred to other references) and forward-propagation and backpropagation as in an ANN.The goal of training is to have (1) the generator minimize
the probability of fake diagnosis by the discriminator and (2) the discriminator
maximize the probability of making a correct distinction between real and fake
samples. Because the two networks constantly compete with each other, training the
model often becomes difcult.
GAN is applied to generate realistic artwork, create human-like voices, and help
improve medical imaging and fashion design. It is also used to create synthetic data
for ML training and synthetic medical images for training and research (Frid-Adar
etal. 2018). They can also create deepfakes of seemingly realistic images and videos that are used in movies and videos.

266
14 Basics ofArticial Intelligence
14.8.4 Transfer Learning
A variation in DL called transfer learning is employed where a pre-trained model
like ImageNet for images, or GPT for language is reused for a new task using the
knowledge learned from the initial training (Pan and Yang 2010). The pre-trained
model has already learned to recognize edges, shapes, textures, etc. of an input
image and the last few layers need to be trained as in ANN.This technique is useful
in cases where data are limited making training faster. Transfer learning has been
successfully applied to medical imaging in the diagnosis and classication of
diseases.
14.9 Radiomics
Radiomics is a technique to retrieve quantitative features from medical images such
as CT, ultrasound, MRI, or PET images, which describe their characteristics,
namely, shape, volume, density, and texture of normal or abnormal tissues (Koçak
etal. 2019). In common practice, these features are assessed visually by the practitioners, which may lead to errors due to intra- and inter-observer variations along
with missing some hidden features in the images. Also, radiomics discovers patterns
in the image not visible to the human eye, which are processed by an applicable AI
algorithm to provide a clinical outcome.
Applying AI in radiomics plays a crucial role in unfolding these features, thus
facilitating image interpretation and disease diagnosis more accurately. Several AI
models can be applied in radiomics, namely ML, DL, Random forests, Support vector machines (SVM), which are discussed earlier. Currently, there are several welldesigned AI programs to extract features for use in radiomics, but they are not as
popular as expected. Also, one can use coding techniques to extract essential features of images from two popular platforms, MATLAB and Python, which have vast
libraries of these features.
Many quantitative features are extracted from the 2D or 3D SPECT or PET
images, which can then be analyzed by AI algorithms to nd correlations with certain diseases (Song etal. 2020). To extract the characteristic features, images are
often segmented for the convenience of easy analysis. Target tissues are segmented,
whereby regions of interest (ROI) are created, and features are identied by applying radiomics. Radiomics is commonly used in cancer diagnosis, but it has been
applied to other diseases, too. Despite its useful contribution to image analysis, it
faces several challenges. Different institutions use different scanners and different
scanning protocols, which leads to variations in imaging results, thus compromising
the accuracy of the extracted features and affecting the strength of the trained models across different centers. Noise, artifacts, and inter-observer variations in manual
segmentation add another challenge to the operation of radiomics.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
