Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5545_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
15 Мб
Скачать
☆
14.12 Generative Articial Intelligence
267

14.10 Natural Language Processing

Natural language processing (NLP) is a branch of computer science that can under­stand and communicate with human language. It uses machine learning for this purpose to recognize texts and voices, although deep learning may be used for com­plex problems. NLP is trained using a large amount of text or voice data in the machine learning model to learn the patterns and associative relationships of the data. Once trained, the model can generate texts on new data. NLP is trained on a continuous basis as new data are accumulated.
NLP works in several ways depending on the context it is used. For text process­ing, initially, the data are organized in a format for the model to understand. NLP, however, cannot process the data in the input text format, and so it is broken into small units called “tokens” such as words, sentences, phrases, etc., a process called “tokenization.” The model then processes and interprets the data to its liking to cre­ate a meaningful text. It is the powerful force behind entities like Google Assistant, Microsoft’s AI assistant, Amazon’s Alexa, etc. Speech recognition is a task of con­verting speech to text, but a successful conversion is achieved if the speech is clean and not altered by obscure dialect, mispronunciation of words, incorrect grammar.

14.11 Large Language Model

A large language model (LLM) is a deep-learning algorithm that is pre-trained on a massive data set to generate, translate, and process texts using natural language processing (NLP). They are trained using unsupervised learning on a vast amount of data and texts collected from media and literature, and the trained model recognizes the previously unknown pattern from the unlabeled data. LLM training is quite laborious and time-consuming due to the requirement of a massive amount of data, which includes zillions of parameters that attributes to the term Large Language Model. For example, LLM-GPT 4 has billions of parameters. LLM can write, code, read, and draw, providing support for many industries in their operation. The scope of LLM is so vast that many enterprises and institutions have adopted it to improve their strategy. However, the requirement of massive data and concomitant large capacity computers, along with experts in high-level computing OpenAI’s GPT-4 and Google’s Gemini are good examples of LLM that understand queries and gener­ate text inputs based on them.

14.12 Generative Artificial Intelligence

A popular form of AI, Generative AI (GenAI), uses a machine learning AI model that is trained to analyze the patterns in a large volume of available data and repli­cate those patterns to create new data, content, or information like text, images, audios, and videos. It employs a neural network for training purposes, and following training, it creates a new item similar to the item it was trained on. For example, if
268
14 Basics ofArticial Intelligence
a poem is entered as input to the GenAI, the latter is trained on the data of the poem and creates a new poem. Similarly, new texts, images, videos, or audios are created in response to the user’s queries, resembling real-world samples. To generate high­quality outputs, GenAI requires the training data to be comprehensive and diverse, a robust model architecture, and an appropriate training process along with evalua­tion strategies. In this model, data storage and training of GenAI are carried out a priori to identify and establish the patterns within the dataset.
Once the model is trained, prompts (see below) are then fed into the GenAI’s algorithm to create new data. Prompts address the issues of reasons for the model’s usage and the expected output from such use. For example, if the desired output is a video of an event, the prompt may include any distinctive features of people pres­ent, decorative patterns, specic features of the event, etc. for appropriate produc­tion of a new video.
Evaluation of the GenAI output is typically carried out by using a different vali­dation dataset, which is not used for training purposes. The purpose of using such an unseen dataset is to determine how well the model performs with new, previously unseen data. If the output is not up to expectation, additional data may be required for retraining, or the model’s architecture may be ne-tuned.
GenAI is widely used in many applications, such as creating texts and images, and translating text from one language to another. Because of its versatility, its use in a variety of disciplines has increased dramatically over the years. OpenAI ini­tially introduced the GenAI models like ChatGPT and DALL-E, which are chatbots discussed below. Nowadays, big tech companies like Google, Microsoft, Amazon, and Meta have launched their own GenAI tools to capitalize on the technology’s rapid growth. Google’s Gemini, Microsoft’s Copilot, and OpenAI’s ChatGPT-4 are examples of very useful GenAI models for nding a quick answer to a specic task.
GenAI has become the model of choice for content generation, creation of new text, artwork, videos, audios, even personalized content, and interestingly, many video games. It has facilitated the eld of drug discovery by predicting the drug­target interaction, leading to the development of new drugs in a relatively short time. Businesses, nancial institutions, industries, educational institutions, and social media are adopting this model for their successful operation.
However, GenAI has its pitfalls too, like bad actors creating and spreading mis­information, deepfakes, hallucination (see below), mistrust, and infringement of copyright and intellectual property. Moreover, if AI systems are not adequately secured, they could become a target for cyberattacks, creating additional data secu­rity concerns.
14.13 Additional Terms ofInterest inAI

14.13.1 Prompt

An AI prompt is a query to infuse as input to LLM through GenAI to nd its answer. A prompt can be a question, command, statement, code, etc. The query users enter
14.13 Additional Terms ofInterest inAI
to ask what they want from GenAI programs like ChatGPT (see below) or its image­making equivalents, like OpenAI’s Dall-E.GenAI scrounges through the data it is already trained on and nds the appropriate response to the query. For an accurate response from AI, prompts must be precise, effective, and understandable to the AI model. An intelligent prompt can generate a story of interest, can create a blog post on a specic topic of interest, and create advertising materials as prompted by busi­ness entrepreneurs, etc., and many other tasks. It can even generate codes for a specic task.
269

14.13.2 Token

In articial intelligence (AI), a token is the smallest unit of text or data that an AI model requires to process the stored data to provide an answer to a question. Tokens can be words, characters, subwords, or punctuation marks, and are the centerpiece items in NLP.For the convenience of processing, text data is broken into smaller units, considered tokens, to be processed by AI models. They can be a letter or a whole phrase, which is fed into LLM for training to learn the pattern of data to nd an answer to a query. According to OpenAI, a token contains roughly four charac­ters of text. Tokens help enhance search algorithms, improve text classication, and promote sentiment analysis.

14.13.3 Hallucination

In AI, hallucination occurs when a response to a query generated by an AI model, especially an LLM, appears to be plausible but, in fact, incorrect, fake, or com­pletely irrelevant to the input query. It can cause detrimental harm if one seeks reli­able information, particularly in medical diagnosis. Hallucination is caused by insufcient or poor quality data and limitation in the scope of the model.

14.13.4 Deepfake

Deepfakes are Images, photos, or videos created by AI algorithms designed to fool people into thinking they are real. They use ML algorithms to analyze large amounts of photos or recordings of a person. The algorithms learn to produce output that resembles the examples they were fed and trained on.
Deepfakes can replace faces, change facial expressions, and create synthetic faces. They can portray deceased actors in movies and spread lies against adversar­ies. False information and fake pornography can be generated by deepfakes. Because it is difcult to decipher truth from falsehood, they can be used to the advantage of perpetrators to sway public opinion and inuence an election. Deepfakes can cause
270
havoc in the nancial world with irreparable damage. However, it can hold some promise for counterterrorism.
14 Basics ofArticial Intelligence

14.13.5 Overfitting

When an AI model ts too close to the training dataset, the model cannot make accurate predictions with any other dataset except the training data itself, the over­tting occurs. In overtting, the model performs well on the training set but poorly on the test/validation set. When training continues for a long time, it tends to learn irrelevant information within the dataset that results in overtting. Small or noisy dataset also causes this problem. The model cannot perform well in the classica­tion and prediction of an intended job. Several steps are often taken to prevent over­tting. Using more and cleaner data is an option to prevent overtting. Next is to pause training early to avoid noises in the model. Proper feature selection in build­ing a model may also be helpful in minimizing overtting.
14.13.6 Encoder andDecoder
Encoders and decoders are components of neural network architectures used to trans­form data from one format to another by compressing to a lower-dimensional entity. Normally, an encoder transforms original input data (e.g., image, text, audio etc.) to an output encoded representation (usually a vector). For example, an autoencoder compresses an image into a latent code. Decoder, on the other hand, reconstructs an output (a reconstructed image, translated sentence) from the encoded entity. A com­bined encoder-decoder structure is commonly used for various AI applications.
14.14 Computer andSoftware
To handle the vast quantity of data efciently and accurately, very fast computers (supercomputers) with high-speed processing units and enormous memory capacity are required. In personal and business computers, data processing is carried out by a unit called the central processing unit (CPU). The CPUs have been described in detail in Chap. 11. However, in AI application, more efcient units with vast storage and high memory capacity (RAM) are required. Graphics processing units (GPUs) are commonly used for the purpose, which are built by combining many (thou­sands) CPUs in parallel conguration. High quality GPUs are expensive, but ef­cient in computing. Besides commonly known computer’s internal storage and ash drives, high-tech corporations offer long-term permanent cloud storages, which are extremely useful in deep learning (DL) algorithms. Network-attached storages and magnetic tapes are alternative choices for long-term permanent storage of AI data. IBM’s Watson is a typical example of a superfast high capacity computer, which was used in the gameshow, Jeopardy, against competitive players. Hewlett-Packard,

14.15 Chatbot

271
Microsoft, and Super Micro Computer Inc. are just a few of many manufacturers who are competing with one another to stay ahead in the game in building superfast computers.
Computers have dual purposes—one to store a massive amount of collected data and the other to carry out the software (algorithm) programs to achieve the intended answer. The more information on a given topic is stored in the computer storage, the more accurate the answer by AI to a question on that topic. High-capacity RAMs are crucial for speedy access to and real-time processing of the data in AI applications. The software needs to be robust, speedy, trustworthy, and cyberattack-proof. Also, the software, and hence the computer, needs to be extremely fast to accomplish a variety of tasks by different AI models.
14.15 Chatbot
Many of us are familiar with the term chatbot, which consists of a set of instructions that simulate conversation with humans through texts or voice interactions. Besides offering text or voice responses, they can build websites and codes, generate images, and analyze documents. Chatbots are language models and use natural language processing (NLP) as the core technology in guiding different chatbots (Bansal and Khan 2018). Since the initial introduction of chatbot on November 30, 2022, by OpenAI, several high-tech companies have introduced chatbots such as OpenAI’s ChatGPT-3.5, ChatGPT-4, and ChatGPT 5, Google’s Gemini (formerly Bard), Microsoft’s Copilot, Amazon’s Alexa, and Salesforce’s Einstein. ChatGPT-3.5 is the initial version of the chatbot from OpenAI, where GPT stands for generative pretrained transformer (Bhayana et al. 2023). This means the model works like LLM using a neural network, is trained, and generates texts by predicting the next most likely word or phrase.
ChatGPT-4o is an upgraded version of ChatGPT-3.5 along with a small default version GPT-4o mini. Recently, OpenAI has introduced an upgraded version of GPT 4.1 along with GPT 4.1 mini. While GPT 4.1 mini remains the default model for all users, including those with a free account, ChatGPT-4.1 is available to only paid subscribers. Initially, chatbots used to make erroneous responses, but with time, AI technology has improved signicantly, resulting in more accurate chatbot responses. All chatbots do not give the same answer to a query, because of nuances in the use of datasets. Currie (2023) made a comparative study of relative perfor­mance by ChatGPT-3.5 and ChatGPT-4in examination sets and written assign­ments of nuclear medicine subjects, found that GPT-4 outperforms ChatGPT-3.5. While ChatGPT-3.5 carries a minimal risk to academic integrity, ChatGPT-4 remains vulnerable to hallucination and fabrication. The latest version ChatGPT-5 has been released in 2025 by Open AI, offering better reasoning, accuracy and memory. It handles texts, codes, images, audios, videos more efciently than GPT-4.
The operational principle of a chatbot is to ask one of these chatbots in a com­puter a question of specic interest, it will search the relevant dataset for the most appropriate answer, which is then displayed on the computer screen. Those who are
272
14 Basics ofArticial Intelligence
novices in computer operation but have some basic knowledge can enjoy the fun and beauty of AI technology in a couple of simple steps. First, download and install a basic chatbot like OpenAI’s ChatGPT, Microsoft’s Copilot, or Google’s Gemini, on the desktop. However, it requires registration and login to access the program. Some are free, while others charge a fee. There are upgraded versions available for these chatbots for an additional charge. When the icon is clicked, a bar opens up and the chatbot asks how it can help you. For simpler questions, login is not required to access the program, but for a complex question the chatbox prompts you to login. With login or no login, one types in a question of interest in the bar, and the com­puter displays the answer in seconds. If the response is not up to expectations, the chatbot offers the option of different answers.
14.15.1 Example ofChatGPT
For the sake of interest, I am quoting below a simple example of the response (ital­ics) given by ChatGPT when I asked for tips on a trip to the Taj Mahal in India.
A trip to the Taj Mahal is an extraordinary experience, lled with history, beauty, and cultural richness. Located in Agra, India, this iconic monument is one of the Seven Wonders of the World, known for its stunning white marble architecture and intricate detailing.
Here’s what you can expect and plan for a visit:
Best Time to Visit
October to March: Ideal weather with cooler temperatures, making it comfort­able to explore.
Early Morning or Late Afternoon: The Taj Mahal looks particularly magical during sunrise or sunset when the marble reects beautiful colors.
Getting There
By Air: Agra has a small airport, but most international tourists y into Delhi and then drive or take a train to Agra (about 3–4hours away).
By Train: The Gatimaan Express or Shatabdi Express are good train options from Delhi.
By Car: A road trip from Delhi is common, via the Yamuna Expressway.
Tickets and Timings
Opening Hours: Sunrise to sunset, closed on Fridays.
Tickets: Available online or at the gate. There are different rates for international tourists, Indian citizens, and children.
Taj Mahal Night Viewing: Available on full moon nights and the two days before and after, offering a mesmerizing view of the monument under moonlight.
What to See
Main Mausoleum: The tomb of Mumtaz and Shah Jahan, with exquisite marble work and precious stone inlay.
Gardens: Mughal-style gardens with symmetrical layouts, fountains, and lush greenery.
14.17 Disadvantages ofAI
The Mosque and Guest House: Flanking the main structure, add to the architec­tural beauty.
Yamuna River View: Behind the Taj Mahal, the river provides a serene backdrop.
Nearby Attractions
Agra Fort: Another UNESCO World Heritage site, located a short distance from the Taj Mahal.
Fatehpur Sikri: A historic Mughal city, about an hour from Agra.
Mehtab Bagh: For a fantastic view of the Taj Mahal from across the river.
Tips
Shoes: You will need to remove your shoes or wear shoe covers inside the mausoleum.
Dress Modestly: Since the Taj Mahal is a tomb, respectful attire is recommended.
Crowds: It can get quite crowded, so arriving early in the day helps avoid the rush.
Would you like to know more details about the Taj Mahal or how to plan a perfect trip there?
273
14.16 Benefits ofAI
Automation of routine and repetitive tasks, solution of complex problems reducing human errors are the obvious benets, which minimize human engagement in the task and let humans focus on other priorities. Examples of AI benets are data entry, customer service in the business, automation of driving (self-driving), diagnosis and treatment of diseases, robot application in various institutions, automated self­service, etc. While the ML model in AI can solve basic problems, the DL model is very efcient in solving more complex tasks. Robotic automation can provide accu­rate results and reduce the time to complete a task. AI can improve customer rela­tionships by providing prompt responses to customer queries. AI in manufacturing can reduce assembly errors and production times while increasing production and worker safety. Weather forecast (accurate prediction of hurricanes, tornadoes, ood­ing, snowfall, etc.) benets tremendously from AI applications. Other entities that benet from the use of AI include nancial institutions, marketing, educational institutions, gaming, and even military.
14.17 Disadvantages ofAI
Despite the tremendous benet of AI to human society in different walks of life, it comes with risks and potential dangers. One of the concerns alludes to the displace­ment or even elimination of jobs for humans, the impact of which experts cannot predict yet. Another concern reects biased human decisions that may be discrimi­natory against certain demographics. Also, AI can be used to generate fake news, spreading disinformation, compromising social trust, and creating chaos.
274
Furthermore, AI-generated material has the potential to infringe upon people’s copyright and intellectual property rights.
14 Basics ofArticial Intelligence

14.18 Legal Implication

Because of the prolic growth of AI and its concomitant effects on human lives, both benecially and adversely, lawmakers around the world are seriously engaged in regulating its application and development. European Union passed a sweeping Articial Intelligence Act to ensure safety, transparency, traceability, and non­discriminatory aspects of AI.China and Brazil have followed suit. In the USA, the Biden administration introduced an AI Bill of Rights followed by an Executive Order on Safe, Secure, and Trustworthy AI in 2023, which was repealed by President Trump in January 2025. Despite many attempts, Congress has failed to come up with robust legislation.
14.19 Future ofAI
At present, despite immense development, the AI machine cannot yet talk, think, or function like a human being. With the enormous ingenuity of human beings, it is believed that AI will undergo continuous upgrades over the next several decades making day-to-day life much easier and more comfortable, and thus improving the quality of human life. A day will come when an AI machine will think, talk, and respond like a normal human being.

14.20 Questions

1. Describe the principles of articial intelligence.
2. Describe the hierarchical relationship among articial intelligence, machine
learning, deep learning and convolutional neural network.
3. Database is a core requirement in articial intelligence. How is it generated?
4. What is the difference between supervised and unsupervised machine learning?
5. Describe the neural network and its operation in articial intelligence.
6. Indicate which AI model is appropriate to apply for the following tasks: A. Chess game B. Face of a man C. Driverless autodriving D. Translating a text
7. Does the generative AI use neural network? What is its distinct
characteristics?
8. Bias is used to adjust the function of a node by moving it up or down. True ____
or False_________

References

275
9. Testing of an AI model is commonly performed by using (a) Same dataset as the one used in training (b) A new previously unseen dataset (c) No dataset
10. Describe how backpropagation is carried out in AI application.
11. What is the loss function used in backpropagation, and how is it used?
12. What the difference between CPU and GPU?
13. What are the main attributes of radiomics in AI?
14. Explain the following terms: prompt, token, hallucination, deepfake.
15. What are the ethical and legal challenges in AI application?
16. How does generative adversarial network work?
17. Describe the function of a chatbot? Name some of the chatbots introduced by
several tech companies.
18. Describe how a feature map in CNN is generated?
19. What is the function of computer vision?
20. Explain overtting in AI and how it can be rectied.
21. Elucidate the benets and disadvantages of AI.
22. What are pretrained AI models and give some examples.
23. Explain the function of radiomics.
References
Alzubaidi, L., Zhang, J., Humaidi, A.J. etal. Review of deep learning: concepts, CNN archi-
tectures, challenges, applications, future directions. J Big Data. 8, 53 (2021). https://doi.
org/10.1186/s40537- 021- 00444- 8 (This is an open access article).
Apostolopoulos ID, Papandrianos NI, Feleki A. etal. Deep learning-enhanced nuclear medicine
SPECT imaging applied to cardiac studies. EJNMMI Physics. 2023; 10:6 Bansal H, and Khan R.A review paper on human computer interaction. Int. J.Adv. Res. Comput.
Sci. Softw. Eng. 2018; 8:53. Bhatnagar S, Alexandrova A, Avin S, etal. Mapping intelligence: Requirements and possibilities.
In Müller VC (Ed.), Philosophy and theory of articial intelligence 2017, Springer, Berlin
2018, pp.117–135. Bhayana R, Bleakene RR, Krishna S.GPT-4in radiology: improvements in advanced reasoning.
Radiology. 2023; 307:e230987. Breiman L. Random Forest. Machine Learning. 2001; 45: 5–32. https://doi.org/10.102
3/A:101093340432.
Coronel C and Morris S. Database Systems: Design, Implementation, & Management. 13th
Edition, Cengage Learning; 2018. Currie G.Intelligent imaging: Anatomy of Machine Learning and Deep Learning, J Nucl Med
Technol. 2019a; 47:273–281. Currie GM. Intelligent Imaging: Articial Intelligence Augmented Nuclear Medicine. J Nucl.
Med. Technol, 2019b; 47(3):217–222. Currie GM.GPT-4in Nuclear Medicine Education: Does It Outperform GPT-3.5? J Nucl Med
Technol. 2023; 51(4): 314–317. Frid-Adar M, Diamant I, Klang E etal. GAN-based synthetic medical image augmentation for
increased CNN performance in liver lesion classication. Neurocomputing 2018, 321, 321–331. Goodfellow I, Pouget-Abadie J, Mirza M, etal., Generative adversarial nets. In: Advances in neu-
ral information processing systems. 2014; p.2672–680.
276
Greenspan H, van Ginneken B, Summers RM.Guest Editorial Deep Learning in Medical Imaging:
Overview and Future Promise of an Exciting New Technique. IEEE Trans. Med. Imaging 2016,
35, 1153–1159. Huang G, Liu Z., Van Der Maaten, L. et al. Densely connected convolutional networks. In
Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu,
HI, USA, 21–26 July 2017; pp.4700–4708. Koçak K, Durmaz SD, Ateş E and Kılıçkesmez Ö. Radiomics with articial intelligence: a practi-
cal guide for beginners. Diagn Interv Radiol. 2019; 5(6):485–495. Krizhevsky A, Sutskever I, Hinton GE Imagenet classication with deep convolutional neural net-
works. Commun. ACM 2017, 60, 84–90. Kussul E, Baidyk T. “Improved method of handwritten digit recognition tested on MNIST
database”. Image and Vision Computing. 2004; 22 (12): 971–981. https://doi.org/10.1016/j.
imavis.2004.03.008.
LeCun Y, et al. Backpropagation Applied to Handwritten Zip Code Recognition. Neural
Computation 1989; 1: 541. https://doi.org/10.1162/neco.1989.1.4.541. Lecun Y, Bottou L, Bengio Y, Haffner P.Gradient-based learning applied to document recognition.
Proc. IEEE 1998, 86, 2278–2324. McCarthy J, Minsky ML, Rochester N, and Shannon CE.A proposal for the Dartmouth Summer
Research Project on Articial Intelligence, August 31, 1955. AI Mag. 2006; 27:12–14. Padoy, N.Towards automatic recognition of surgical activities. In Proceedings of the International
Conference on Medical Image Computing and Computer-Assisted Intervention, Nice, France,
1–5 October 2012; pp.267–274. Pan SJ and Yang Q.A survey on transfer learning. IEEE Trans. Knowl. Data Eng. 2010, 22,
1345–1359. Patni JC and Pinjarkar L.Kickstart Database Management System Fundamentals: Key Concepts,
Principles, and Advanced Techniques for Modern Database Design, Management, and
Optimization, 2024; Orange Education Pvt Ltd., India. Ronneberger O, Fischer P, Brox T (2015). U-Net: Convolutional Networks for Biomedical Image
Segmentation. arXiv (2015):1505.04597 (cs.CV) Sarker IH.Deep Learning: A Comprehensive Overview on Techniques, Taxonomy, Applications
and Research Directions. SN Computer Science. 2021; 2:420 Song J, Yin Y, Wang H., etal. A review of original articles published in the emerging eld of
radiomics. Eur J Radiol, 2020; 127: 108991 Tajbakhsh N, Shin JY, Gurudu SR, etal. Convolutional neural networks for medical image analy-
sis: Full training or ne-tuning? IEEE Trans. Med. Imaging 2016; 35: 1299–1312. Vapnik V.N., and Chervonenkis, A. Support vector method for function approximation, regres-
sion estimation, and signal processing. Advances in Neural Information Processing Systems
(NIPS). 1995; 8: 281–287. Wang W, Liang D, Chen Q, etal. Medical image classication using deep learning. In: Chen, YW,
Jain, L. (eds) Deep Learning in Healthcare. Intelligent Systems Reference Library, vol 171.
Springer, Cham
Appl. 2020, 33–51. Yamashita R, Nishio M, Do RKG, and Togashi K.Convolutional neural networks: an overview and
application in radiology. Insight Imaaging.2018; 9: 611–629. Zhang H and Qie Y.Applying Deep Learning to Medical Imaging: A Review. Appl Sci. 2023,
13(18): 10521.
https://doi.org/10.1007/978- 3- 030- 32606- 7_3 Deep. Learn. Healthc. Paradig.
14 Basics ofArticial Intelligence