Добавил:
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Exam Time for IT. Ч.2. Практикум

.pdf
Скачиваний:
0
Добавлен:
12.08.2026
Размер:
1 Мб
Скачать

Topic 5. Data Processing

Data processing is any computer process that converts data into information or knowledge. The processing is usually assumed to be automated and running on a computer. Because data are most useful when well-presented and actually informative, data-processing systems are often referred to as information systems to emphasize their practicality. Nevertheless, both terms are roughly synonymous, performing similar conversions; data-processing systems typically manipulate raw data into information, and likewise information systems typically take raw data as input to produce information as output.

Data

Process

Information

Data are defined as numbers or characters that represent measurements from observable phenomena. A single datum is a. single measurement from observable phenomena. Measured information is then algorithmically derived and/or logically deduced and/or statistically calculated from multiple data. Information is defined as either a meaningful answer to a query or a meaningful stimulus that can cascade into further queries. For example gathering seismic data leads to alteration of seismic data to suppress noise, enhance signal and migrate seismic events to the appropriate location in space. Processing steps typically include analysis of velocities and frequencies, static corrections, deconvolution, normal moveout, dip moveout, stacking, and migration, which can be performed before or after stacking. Seismic processing facilitates better interpretation because subsurface structures and reflection geometries are more apparent.

More generally, the term «data processing» can apply to any process that converts data from one format to another, although data conversion would be the more logical and correct term. From this perspective, data processing becomes the process of converting information into data and also the converting of data back into information. The distinction is that conversion doesn’t require a question (query) to be answered. For example, information in the form of a string of characters forming a sentence in English is converted or encoded meaningless hardware-oriented data to evermore-meaningful information as the processing proceeds toward the human being.

Embedded System.

Conversely, that simple example for pedagogical purposes here is usually described as an embedded system (for the software resident in the keyboard itself) or as (operating-)systems programming, because the information is derived from a hardware interface and may involve overt control of the hardware through that interface by an operating system. Typically control of hardware by a device driver manipulating ASIC or FPGA registers is not viewed as part of data processing proper or information systems proper, but rather as the domain of embedded systems or (operating-)systems programming. Instead, perhaps a more conventional example of the established practice of using the term data processing is that a business has collected numerous data concerning an aspect of its operations and that this multitude of data must be presented in meaningful, easy-to-access presentations for the managers who must then use that information to increase revenue or to decrease cost. That conversion and presentation of data as information is typically performed by a data-processing application.

Data Analysis.

When the domain from which the data are harvested is a science or an engineering, data processing and information systems are considered too broad of terms and the more specialized

21

term data analysis is typically used, focusing on the highly-specialized and highly-accurate algorithmic derivations and statistical calculations that are less often observed in the typical general business environment. In these contexts data analysis packages like DAP, gretl or PSPP are often used. This divergence of culture is exhibited in the typical numerical representations used in data processing versus numerical; data processing’s measurements are typically represented by integers or by fixed-point or binary-coded decimal representations of numbers whereas the majority of data analysis’s measurements are often represented by floating-point representation of rational numbers.

Processing.

Practically all naturally occurring processes can be viewed as examples of data processing systems where «observable» information in the form of pressure, light, etc. are converted by human observers into electrical signals in the nervous system as the senses we recognize as touch, sound, and vision. Even the interaction of nonliving systems may be viewed in this way as rudimentary information processing systems. Conventional usage of the terms data processing and information systems restricts their use to refer to the algorithmic derivations, logical deductions, and statistical calculations that recur perennially in general business environments, rather than in the more expansive sense of all conversions of real-world measurements into realworld information in, say, an organic biological system or even a scientific or engineering system.

In data processing or information processing, a data processing system or data processing unit or data processor is a system which processes data which has been captured and encoded in a format recognizable by the data processing system or has been created and stored by another unit of an information processing system.

A data entry is a specialized component or form of an information processing (sub)system. Its chief difference is that it tends to perform a dedicated function (i.e., its program is not readily changeable). Its dedicated function is normally to perform some (intermediate) step of converting input (‘raw’ or unprocessed) data, or semi-processed information, in one form into a further or final form of information through a process called decoding/encoding or formatting or re-formatting or translation or data conversion before the information can be output from the data processor to a further step in the information processing system.

For the hardware data processing system, this information may be used to change the sequential states of a (hardware) machine called a computer. In all essential aspects, the hardware data processing unit is indistinguishable from a computer’s central processing unit

(CPU), i.e. the hardware data processing unit is just a dedicated computer. However, the hardware data processing unit is normally dedicated to the specific computer application of format translation.

A software code compiler (e.g., for Fortran or Algol) is an example of a software data processing system. The software data processing system makes use of a (general purpose)

22

computer in order to complete its functions. A software data processing system is normally a standalone unit of software, in that its output can be directed to any number of other (not necessarily as yet identified) information processing (sub)systems.

Elements of Data Processing.

In order to be processed by a computer, the data needs first to be converted into a machine readable format. Once data is in digital format, various procedures can be applied on the data to get useful information. Data processing includes all the processes from data entry up to data mining: data entry, data cleaning, data coding, data translation, data summarization, data aggregation, data validation, data tabulation, statistical analysis, computer graphics, data warehousing, data mining.

Questions for self-check:

What are data defined as?

What distinction is there between data processing and data conversion?

What is the established practice of using the term data processing? Why?

What does the term of data analysis focuse on?

How are data processing measurements typically represented?

What processes can be viewed as examples of data processing systems?

What is the usage of the terms data processing and information systems restricted to refer to?

What is a data processing system?

What functions does data entry perform?

How can a software data processing system be exemplified?

What elements does data processing include?

Exercises:

1. Give English-Russian equivalents of the following words and expressions:

 

нечто целое; десятичный;

 

apparent; divergence;

очевидный, явный, открытый;

нахождение

предполагать;

assume; integer;

оригинала

скорость, быстрота;

roughly; conversely;

функции;

объединение, соединение;

rudimentary;

долговременное

доход, выручка;

validation; overt;

хранение;

грубо, приблизительно,

deconvolution;

добыча;

примерно; исходные данные;

decimal; raw data;

сведение в

получать, извлекать; наоборот;

query; capture;

таблицы;

выводить, прослеживать;

tabulation; derive;

несоответствие,

видимый, несомненный,

deduce; perennially;

расхождение;

очевидный;

mining; warehousing;

элементарный;

проверка данных;

aggregation; revenue;

всегда, постоянно;

запрос (критерий поиска

velocity.

собирать (данные).

объектов в базе данных), вопрос

 

 

 

 

 

23

2.Find the word belonging to the given synonymic group among the words and word combinations from the previous exercise:

a)basic, elementary, simple, undeveloped;

b)suppose, believe, presume, take for granted, imagine, think, guess;

c)justification, confirmation, examination, control, verification, testing;

d)approximately, about, more or less, generally, almost, something like, just about, in the region of;

e)collect, secure, attain, gain, acquire, obtain (data);

f)income, profits, returns, proceeds, takings;

g)obvious, clear, evident, noticeable, perceptible, visible;

h)get, receive, draw from, take, gain;

i)inquiry, question;

j)unconcealed, explicit, open, plain, obvious, clear;

k)trace, monitor, observe, figure out, work out;

l)the whole, total, unit;

m)speed, rate, rapidity, swiftness, pace, haste, quickness;

n)on the other hand, on the contrary, in opposition;

o)output, extraction, production, getting

24

Topic 6. Data transmission

Data transmission is essentially the same thing as digital communications, and implies physical transmission of a message as a digital bit stream, represented as an electro-magnetic signal, over a physical point-to-point or point-to-multipoint communication channel. Examples of such channels are copper wires, optical fibers, wireless communication channels, and storage media.

Data transmission is a subset of the field of data communications, which also includes computer networking or computer communication applications and networking protocols, for example, routing and switching.

Applications and History.

The first data transmission applications in modern time were telegraphy (1809) and teletypewriters (1906). The fundamental theoretical work in data transmission and information theory by Harry Nyquist, Ralph Hartley, Claude Shannon and others during the early 20th century, was done with these applications in mind.

Data transmission is utilized in computers in computer buses and for communication with peripheral equipment via parallel ports and serial ports such us RS-232 (1969), Firewire (1995) and USB (1996). The principles of data transmission are also utilized in storage media for error detection and correction since 1951. Data transmission is utilized in computer networking equipment such as modems (1940), local area networks (LAN) adapters (1964), repeaters, hubs, microwave links, wireless network access points (1997), etc.

In telephone networks, digital communication is utilized for transferring many phone calls over the same copper cable or fiber cable by means of Pulse Code Modulation (PCM), i.e. sampling and digitalization, in combination with Time Division Multiplexing (TDM) (1962). Telephone exchanges have become digital and software controlled, facilitating many value added services. For example the first AXE telephone exchange was presented in 1976. Since late 1980th, digital communication to the end user has been possible using Integrated Services Digital Network (ISDN) services. Since the end of 1990th, broadband access techniques such as ADSL, Cable modems, fiberto-the-building (FTTB) and fiber-to-the-home (FTTH) have become wide spread to small offices and homes. The current tendency is to replace traditional telecommunication services by packet mode communication such as IP telephony and IPTV.

Protocols and Handshaking.

A protocol is an agreed-upon format for transmitting data between two devices, e.g.: computer and printer. All communications between devices require that the devices agree on the format of the data. The set of rules defining a format is called a protocol.

The protocol determines the following:

the type of error checking to be used if any, e.g.: check digit (and what type/ what formula to be used);

data compression method, if any, e.g.: zipped files if the file is large, like transfer across the Internet, LANs and WANs;

how the sending device will indicate that it has finished sending a message, e.g.: in a Communications port a spare wire would be used, for serial (USB) transfer start and stop digits maybe used;

how the receiving device will indicate that it has received a message;

rate of transmission (in baud or bit rate);

whether transmission is to be synchronous or asynchronous.

In addition, protocols can include sophisticated techniques for detecting and recovering from transmission errors and for encoding and decoding data.

Handshaking is the process by which two devices initiate communications, e.g.: a certain ASCII character or an interrupt signal/ request bus signal to the processor along the Control Bus. Handshaking begins when one device sends a message to another device indicating that it wants

25

to establish a communications channel. The two devices then send several messages back and forth that enable them to agree on a communications protocol. Handshaking must occur before data transmission as it allows the protocol to be agreed.

Asynchronous and synchronous data transmission.

Asynchronous and synchronous communication refers to methods by which signals are transferred in computing technology. These signals allow computers to transfer data between components within the computer or between the computer and an external network. Most actions and operations that take place in computers are carefully controlled} and occur at specific times and intervals. Actions that are measured against a time reference, or a clock signal, are referred to as synchronous actions. Actions that are prompted as a response to another signal, typically not governed by a clock signal, are referred to as asynchronous signals.

Typical examples of synchronous signals include the transfer and retrieval of address information within a computer via the use of an address bus. For example, when a processor places an address on the address bus, it will hold it there for a specific period of time. Within this interval, a particular device inside the computer will identify itself as the one being addressed and acknowledge the commencement of an operation related to that address.

In such an instance, all devices involved in ensuing bus cycles must obey the time constraints applied to their actions — this is known as a synchronous operation. In contrast, asynchronous signals refer to the operations that are prompted by an exchange of signals with one another, and are not measured against a reference time base. Devices that cooperate asynchronously usually include modems and many network technologies, both of which use a collection of control signals to notify intent in an information exchange. Asynchronous signals, or extra control signals, are sometimes referred to as handshaking signals because of the way they mimic two people approaching one another and shaking hands before conversing or negotiating.

Within a computer, both asynchronous and synchronous protocols are used. Synchronous protocols usually offer the ability to transfer information faster per unit time than asynchronous protocols. This happens because synchronous signals do not require any extra negotiation as a prerequisite to data exchange. Instead, data or information is moved from one place to another at instants in time that are measured against the clock signal being used. This signal is usually comprised of one or more high frequency rectangular shaped waveforms, generated by special purpose clock circuitry. These pulsed waveforms are connected to all the devices that operate synchronously, allowing them to start and stop operations with respect to the clock waveform.

In contrast, asynchronous protocols are generally more flexible, since all the devices that need to exchange information can do so at their own natural rate — be these fast or slow. A clock signal is no longer necessary; instead the devices that behave asynchronously wait for the handshaking signals to change state, indicating that some transaction is about to commence. The handshaking signals are generated by the devices themselves and can occur as needed, and do not require an outside supervisory controller such as a clock circuit that dictates the occurrence of data transfer.

Asynchronous and synchronous transmission of information occurs both externally and internally in computers. One of the most popular protocols for communication between computers and peripheral devices, such as modems and printers, is the asynchronous RS-232 protocol. Designated as the RS-232C by the Electronic Industries Association (ElA), this protocol has been so successful at adapting to the needs of managing communication between computers and supporting devices, that it has been pushed into service in ways that were not intended as part of its original design. The RS-232C protocol uses an asynchronous scheme that permits flexible communication between computers and devices using byte-sized data blocks each framed with start, stop, and optional parity bits on the data line. Other conductors carry the handshaking signals and possess names that indicate their purpose — these include data terminal ready, request to send, clear to send, data set ready, etc.

26

Another advantage of asynchronous schemes is that they do not demand complexity in the receiver hardware. As each byte of data has its own start and stop bits, a small amount of drift or imprecision at the receiving end does not necessarily spell disaster since the device only has to keep pace with the data stream for a modest number of bits. So, if an interruption occurs, the receiving device can re-establish its operation with the beginning of the arrival of the next byte. This ability allows for the use of inexpensive hardware devices.

Although asynchronous data transfer schemes like RS-232 work well when relatively small amounts of data need to be transferred on an intermittent basis, they tend to be sub-optimal during large information transfers. This is so because the extra bits that frame incoming data tend to account for a significant part of the overall inter-machine traffic, hence consuming a portion of the communication bandwidth.

An alternative is to dispense with the extra handshaking signals and overhead, instead synchronizing the transmitter and receiver with a clock signal or synchronization information contained within the transmitted code before transmitting large amounts of information. This arrangement allows for collection and dispatch of large batches of bytes of data, with a few bytes at the front-end that can be used for the synchronization and control. These leading bytes are variously called synchronization bytes, flags, and preambles. If the actual communication channel is not a great distance, the clocking signal can also be sent as a separate stream of pulses. This ensures that the transmitter and receiver are both operating on the same time base, and the receiver can be prepared for data collection prior to the arrival of the data.

An example of a synchronous transmission scheme is known as the High-level Data Link Control, or HDLC. This protocol arose from an initial design proposed by the IBM Corporation. HDLC has been used at the data link level in public networks and has been adapted and modified in several different ways since.

A more advanced communication protocol is the Asynchronous Transfer Mode (ATM), which is an open, international standard for the transmission of voice, video, and data signals. Some advantages of ATM include a format that consists of short, fixed cells (53 bytes) which reduce overhead in maintenance of variable-sized data traffic. The versatility of this mode also allows it to simulate and integrate well with legacy technologies, as well as offering the ability to guarantee certain service levels, generally referred to as quality of service (QoS) parameters.

Questions for self-check:

What are the first examples of data transmission application?

Where is data transmission utilized?

What is the protocol?

When does handshaking occur?

What signals are referred to as asynchronous? Give some

examples.

Which devices cooperate asynchronously?

What do synchronous protocols usually offer? Why?

Why are asynchronous protocols generally more flexible?

What are the advantages of asynchronous schemes?

Why do asynchronous data transfer schemes work well when relatively small amounts of data need to be transferred on an intermittent basis? What is an alternative?

27

Exercises:

1. Give Russian equivalents of the following words and expressions:

optional, drift. Permit. point-to-point, dispatch, routing, commence, clock signal, intermittent, value, added, dispense with, check, digit, bandwidth, handshaking, optical fiber, parity, clock circuitry, prerequisite

2.Replace the underlined words and expressions with synonyms from the previous exercise:

a)Data transmission implies physical transmission of a message as a digital bit stream, represented as an electro-magnetic signal, over a physical double-point or multipoint communication channel.

b)Examples of such channels are copper wires, light pipes, wireless communication channels, and storage media.

c)Telephone exchanges have become digital and software controlled, facilitating many cost attached services.

d)Connection acknowledgement must occur before data transmission as it allows the protocol to be agreed.

e)As each byte of data has its own start and stop bits, a small amount of divergence or imprecision at the receiving end does not necessarily spell disaster since the device only has to keep pace with the data stream for a modest number of bits.

f)Data transmission is a subset of the field of data communications, which also includes computer networking or computer communication applications and networking protocols, for example tracing and switching.

g)Synchronous signals do not require any extra negotiation as a precondition to data exchange.

h)The extra bits that frame incoming data tend to account for a significant part of the overall inter-machine traffic, hence consuming a portion of the communication throughput.

i)Actions that are measured against a time reference, or a beat wave, are referred to as synchronous actions.

j)The devices that behave asynchronously wait for the handshaking signals to change state, indicating that some transaction is about to start.

k)The handshaking signals are generated by the devices themselves and can occur as needed, and do not require an outside supervisory controller such as a synchronization diagram that dictates the occurrence of data transfer.

l)The RS-232C protocol uses an asynchronous scheme that allows flexible communication between computers and devices using byte-sized data blocks each framed with start, stop, and additional even bits on the data line.

m)Although asynchronous data transfer schemes like RS-232 work well when relatively small amounts of data need to be transferred on a time-dependent basis, they tend to be suboptimal during large information transfers.

n)To do without the extra handshaking signals and overhead, instead synchronizing the transmitter and receiver with a clock signal or synchronization information contained within the transmitted code before transmitting large amounts of information allows for collection and large batches of bytes of data sending, with a few bytes at the frontend that can be used for the synchronization and control.

28

3. Translate the words/expressions into English:

отправка, отправление; двухточечный, двухпунктовый; оптоволокно; начинать(ся); дополнительный, необязательный; допускать, позволять; маршрутизация (в сети); схема синхронизации; обходиться без чего-л.; пропускная способность; предпосылка, предварительное условие; нестационарный (о сигнале); синхросигнал, тактовый сигнал; обмен с квитированием; подтверждение установления/квитирование связи; четность; добавленная стоимость; отклонение, смещение; контрольный разряд.

29

Topic 7. Information Retrieval

Information retrieval is a wide, often loosely-defined term but in these pages we shall be concerned only with automatic information retrieval systems: automatic as opposed to manual and information as opposed to data or fact. Unfortunately, the word «information» can be very misleading. In the context of information retrieval (IR), information, in the technical meaning given in Shannon’s theory of communication, is not readily measured (Shannon and Weaver). In fact, in many cases one can adequately describe the kind of retrieval by simply substituting

‘document’ for ‘information’. Nevertheless, ‘information retrieval’ has become accepted as the science of searching for documents, for information within documents and for metadata about documents, as well as that of searching relational databases and the World Wide Web. There is an overlap in the usage of the terms data retrieval, document retrieval, information retrieval, and text retrieval, but each also has its own body of literature, theory, praxis and technologies. IR is interdisciplinary, based on computer science, mathematics, library science, information science, information architecture, cognitive psychology, linguistics, statistics and physics.

To make clear the difference between data retrieval (DR) and information retrieval (IR), some of the distinguishing properties of data and information retrieval are listed in the table:

 

Data Retrieval (DR)

Information Retrieval (IR)

Matching

Exact match

Partial match, best match

Inference

Deduction

Induction

Model

Deterministic

Probabilistic

Classification

Monothetic

Polythetic

Query language

Artificial

Natural

Query specification

Complete

Incomplete

Items wanted

Matching

Relevant

Error response

Sensitive

Insensitive

Let us now take each item in the table in turn and look at it more closely. In data retrieval we are normally looking for an exact match, that is, we are checking to see whether an item is or is not present in the file. In information retrieval this may sometimes be of interest but more generally we want to find those items which partially match the request and then select from those a few of the best matching ones.

The inference used in data retrieval is of the simple deductive kind, that is, aRb and bRc then aRc. In information retrieval it is far more common to use inductive inference; relations are only specified with a degree of certainty or uncertainty and hence our confidence in the inference is variable. This distinction leads one to describe data retrieval as deterministic but information retrieval as probabilistic. Frequently Bayes’ Theorem is invoked to carry out inferences in IR, but in DR probabilities do not enter into the processing.

Another distinction can be made in terms of classifications that are likely to be useful. In DR we are most likely to be interested in a monothetic classification, that is, one with classes defined by objects possessing attributes both necessary and sufficient to belong to a class. In IR such a classification is one the whole not very useful, in fact more often a polythetic classification is what is wanted. In such a classification each individual in a class will possess only a proportion of all the attributes possessed by all the members of that class. Hence no attribute is necessary or sufficient for membership to a class.

The query language for DR will generally be of the artificial kind, one with restricted syntax and vocabulary, in IR we prefer to use natural language although there are some notable exceptions. In DR the query is generally a complete specification of what is wanted, in IR it is invariably incomplete. This last difference arises partly from the fact that in IR we are searching for relevant documents as opposed to exactly matching items. The extent of the match in IR is

30