Добавил:
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Learn statistics in English. Учебно-практическое пособие

.pdf
Скачиваний:
0
Добавлен:
12.08.2026
Размер:
969 Кб
Скачать

Практикум

hydraulic heads known only to be above the land surface (artesian wells on old maps).

6.Seasonal patterns. Values tend to be higher or lower in certain seasons of the year.

7.Autocorrelation. Consecutive observations tend to be strongly correlated with each other. For the most common kind of autocorrelation in water resources (positive autocorrelation), high values tend to follow high values and low values tend to follow low values.

8.Dependence on other uncontrolled variables. Values strongly covary with water discharge, hydraulic conductivity, sediment grain size, or some other variable.

Methods for analysis of water resources data, whether the simple summarization methods such as those in this text, or the more complex procedures, should recognize these common characteristics.

141

Learn statistics in English

Read the text. Illustrate how the existence of differences relates to the absence of complete determinism. Distinguish the physical sources of randomness in real world.

Roughly примерно

To split into расщеплять на, делить

To equalize сравнять, уравнивать

Requisite нужный, требуемый

Color-blind дальтоник

To furnish предоставлять

Quantum-mechanical квантово-механический

To ascertain констатировать, доказывать

Solid твердый, прочный

Lattice решётка, плетение

Particles частицы

Suspended свободно плавающий

Discreteness дискретность

Permeation проникание

Impurities примеси

Unpredictability and Randomness

A large number of phenomena exist that are neither completely determinate nor completely chaotic. To describe them, one may use a system of nonidentical but mutually comparable objects and then classify them into several groups. Of interest to us might be to what group a given object belongs. We shall illustrate how the existence of differences relates to the absence of complete determinism. Suppose that we are interested in the sex of newborn children. It is known that roughly half of births are boys and half are girls. In other words, the «things» being considered split into two groups. If a strictly valid law existed for the birth of a boy or girl, then it would still be impossible to produce the mechanism which would continually equalize the sexes of babies being born in the requisite proportion (without assuming the effect of the results of prior births on succeeding births, such a premise is meaningless). One may give numerous examples of valid statements like

142

Практикум

«such a thing happens in such and such fraction of the cases», for instance, «1% of males are color-blind». As in the case of the sex of babies, the phenomenon cannot be explained on the basis of determinate laws. It is advantageous to view a set-up of things as a sequence of events proceeding in time.

The absence of determinism means that future events are unpredictable. Since events can be classified in some sort of way, one may ask to what class will a future event belong? But once again (determinism not being present), one cannot furnish an answer in advance. The question is still posed in the given situation. The examples cited suggest a proper way to state the question: how often will a phenomenon of a given class occur in the sequence? We shall speak about chance in precisely such situations and it will be natural to raise such questions and to find answers for them.

Sources of Randomness

We shall now point out a few of the most important existing physical sources of randomness in the real world.

(1)Quantum-mechanical laws. The laws of quantum mechanics are statements about the wave functions of micro-objects. According to these laws, we can specify, for instance, just the wave function of an electron in a field of force. Based on the wave function, only the probability of detecting the electron in some particular region of space may be found – to predict its position is impossible. In exactly the same way, one cannot ascertain the energy of an electron and it is only possible to determine a discrete number of possible energy levels and the probability that the energy of the electron has a specified value. We perceive that the fundamental laws of the microworld make use of the language of probability and thus phenomena in the microworld are random. An important example of a random phenomenon in the microworld is the emission of a quantum of light by an excited atom. Another important example is nuclear reactions.

(2)Thermal motion of molecules. The molecules of any substance are in constant thermal motion. If the substance is a solid, then the molecules range close to positions of equilibrium in a

143

Learn statistics in English

crystal lattice. But in fluids and gases, the molecules perform rather complex movements changing their directions of motion frequently as they interact with one another. The presence of such a motion may be ascertained by watching the movement of microscopic particles suspended in a fluid or gas (this is so-called Brownian motion). This motion is of a random nature and the energies of the individual molecules are also random, that is, the energies of the molecules can assume different values and so one talks about the fraction of molecules having an energy within narrow specified bounds. This is the familiar Maxwell distribution in physics. A simple experiment will convince one that the energies of the molecules are different. Take the phenomenon of boiling water: if all of the molecules had the same energy, then the water would become steam all at once, that is, with an explosion, and this does not happen.

(3) Discreteness of matter. The discreteness of matter leads to the occurrence of randomness in another way. Items (1) and (2) also considered material particles. The following fact should now be noted: the laws of classical physics have been formulated for macrobodies just as if matter filled up space continuously. The discreteness of matter leads to the occurrence of deviations of the actual values of physical quantities from those predicted by the laws. These deviations or «fluctuations» are of a random nature and they affect the course of a process substantially. Thus, the discreteness of the carriers of electricity in metallic conductors – the electrons – is the source of fluctuation currents which are the reason for internal noise in radios. The discreteness of matter results in the mutual permeation of substances. Furthermore the absence of pure substances, that is, the existence of impurities, also results in random deviations from the calculated flow of phenomena.

(d) Cosmic radiation. Experimentation shows that it is irregular (aperiodic and unpredictable) but it conforms to laws that can be studied by probability theory.

144

Практикум

Read the text. Present the basic idea of icon plots. What phases does the analysis of icon plots consist of? Define two categories of icon plots.

Icon рисунок, символ, пиктограмма

Plot график

Multidimensional многомерный

To capitalize on использовать

Consistent совместимый, подходящий

Sequence последовательность

Non-salient не существенный

To assign сопоставить, принять

To drop исключить

Taxonomy классификация, таксономия

Spoked wheel колесо со спицами

Hub центр, ядро

Consecutive ступенчатый, идущий подряд

Icon Plots

Icon Graphs represent cases or units of observation as multidimensional symbols and they offer a powerful although not easy to use exploratory technique. The general idea behind this method capitalizes on the human ability to "automatically" spot complex (sometimes interactive) relations between multiple variables if those relations are consistent across a set of instances (in this case "icons"). Sometimes the observation (or a "feeling") that certain instances are "somehow similar" to each other comes before the observer (in this case an analyst) can articulate which specific variables are responsible for the observed consistency (Lewicki, Hill, & Czyzewska, 1992). However, further analysis that focuses on such intuitively spotted consistencies can reveal the specific nature of the relevant relations between variables.

The basic idea of icon plots is to represent individual units of observation as particular graphical objects where values of variables are assigned to specific features or dimensions of the objects (usually one case = one object). The assignment is such that

145

Learn statistics in English

the overall appearance of the object changes as a function of the configuration of values.

Thus, the objects are given visual "identities" that are unique for configurations of values and that can be identified by the observer. Examining such icons may help to discover specific clusters of both simple relations and interactions between variables.

Analyzing Icon Plots

The "ideal" design of the analysis of icon plots consists of five phases:

1.Select the order of variables to be analyzed. In many cases a random starting sequence is the best solution. You may also try to enter variables based on the order in a multiple regression equation, factor loadings on an interpretable factor, or a similar multivariate technique. That method may simplify and "homogenize" the general appearance of the icons which may facilitate the identification of nonsalient patterns. It may also, however, make some interactive patterns more difficult to find. No universal recommendations can be given at this point, other than to try the quicker (random order) method before getting involved in the more time-consuming method.

2.Look for any potential regularities, such as similarities between groups of icons, outliers, or specific relations between aspects of icons (e.g., "if the first two rays of the star icon are long, then one or two rays on the other side of the icon are usually short"). The Circular type of icon plots is recommended for this phase.

3.If any regularities are found, try to identify them in terms of the specific variables involved.

4.Reassign variables to features of icons (or switch to one of the sequential icon plots) to verify the identified structure of relations (e.g., try to move the related aspects of the icon closer together to facilitate further comparisons). In some cases, at the end of this phase it is recommended to drop the variables that appear not to contribute to the identified pattern.

5.Finally, use a quantitative method (such as a regression method, nonlinear estimation, discriminate function analysis, or cluster analysis) to test and quantify the identified pattern or at least some aspects of the pattern.

146

Практикум

Taxonomy of Icon Plots

Most icon plots can be assigned to one of two categories: circular and sequential.

Circular icons. Circular icon plots (star plots, sun ray plots, polygon icons) follow a "spoked wheel" format where values of variables are represented by distances between the center ("hub") of the icon and its edges.

Those icons may help to identify interactive relations between variables because the overall shape of the icon may assume distinctive and identifiable overall patterns depending on multivariate configurations of values of input variables.

In order to translate such "overall patterns" into specific models (in terms of relations between variables) or verify specific observations about the pattern, it is helpful to switch to one of the sequential icon plots which may prove more efficient when one already knows what to look for.

Sequential icons. Sequential icon plots (column icons, profile icons, line icons) follow a simpler format where individual symbols are represented by small sequence plots (of different types).

The values of consecutive variables are represented in those plots by distances between the base of the icon and the consecutive break points of the sequence (e.g., the height of the columns). Those plots may be less efficient as a tool for the initial exploratory phase of icon analysis because the icons may look alike. However, as mentioned before, they may be helpful in the phase when some hypothetical pattern has already been revealed and one needs to verify it or articulate it in terms of relations between individual variables.

147

Learn statistics in English

Final Test

Part I

Match the English terms on the left with the Russian ones on the right.

1. random sample

1.

оценка параметра

2.

correlation

2.

долгосрочный

3.

coefficient of determination

3.

обрабатывать

4.

rank

4.

корреляция

5.

handle

5.

случайная выборка

6.

parameter estimation

6.

ранжировать

7.

bell-curve

7.

вероятностное

 

long-term

 

моделирование

8.

8.

квадрат смешанной

 

 

 

корреляции

9.

interval scale

9.

многомерное измерение

10.

probabilistic modeling

10.

колокообразная кривая

11.

joining (tree clustering)

11.

интервальная шкала

12.

residual value

12.

независимая переменная

13.

least squares

13.

метод наименьших

 

predictor variable

 

квадратов

14.

14.

остаточное значение

15.

multiple dimension

15. объединение (древовидная

 

 

 

кластеризация)

Match the Russian terms on the left with the English ones on the right.

16.

квадраты отклонений

1.

residual value

17.

модель скользящей средней

2.

value on dimensions

18.

оценка по методу наимень-

3.

squared deviations

 

ших квадратов

 

 

19.

остаточное значение

4.

time series

20.

квадрат когерентности

5.

smoothing

21.

остаточная дисперсия

6.

dependent variable

22.

значения параметров

7.

slope

23.

временной ряд

8.

residual variance

148

 

 

 

Final Test

24.

зависимая переменная

9.

square root

25.

сглаживание

10.

squared coherency

26.

дисперсионный анализ

11.

negative exponentially

 

 

 

weighted smoothing

27.

квадратный корень

12.

least squares estimation

28.

отрицательно взвешенное

13.

variance analysis

 

экспоненциальное

 

 

 

сглаживание

 

two-way joining

29.

наклон

14.

30.

двувходовое объединение

15.

moving average model

Part II

Fill the gaps with the words or word combinations from the given list on the right.

31.

The critical features of … are ensuring anonymi-

1. probability

 

ty, the presence of an objective third party, and

 

 

constructive feedback to the organization.

 

32.

… are data collected at the same or approximate-

2. Pearson cor-

 

ly the same point in time.

relation

33.

The exact shape of the … is defined by a func-

3. the ARIMA

 

tion which has only two parameters: mean and

methodology

 

standard deviation.

 

34.

… is derived from the verb to probe meaning to

4. organization

 

find out, what is not too easily accessible or un-

records

 

derstandable.

 

35.

… assumes that the two variables are measured

5. measure-

 

on at least interval scales, and it determines the

ment scale

 

extent to which values of two variables are

 

 

«proportional» to each other.

 

36.

Factor that determines the amount of informa-

6. cross-

 

tion that can be provided by a variable is its type

sectional data

 

of … .

 

37. The most common techniques is … which replaces

7. survey ques-

 

each element of the series by either the simple or

tionnaires

 

weighted average of n surrounding elements,

 

 

where n is the width of the smoothing «window».

 

 

 

149

Learn statistics in English

 

38.

… include strategic plans, absentee lists, griev-

8. outliers

 

ances field, units of performance per person, and

 

 

costs of production.

 

39.

… have a profound influence on the slope of the

9. normal dis-

 

regression line and consequently on the value of

tribution

 

the correlation coefficient.

 

40.

We may use the neighbors across clusters that

10. joining (tree

 

are furthest away from each other; this method

clustering)

 

is called … .

 

41.

The deviation of a particular point from the re-

11. multiple

 

gression line (its predicted value) is called … .

regression

42.

… allows us to uncover the hidden patterns in

12. the residual

 

the data and to generate forecasts.

value

43.

The purpose of … is to join together objects into

13. hierarchical

 

successively larger clusters, using some measure

tree

 

of similarity or distance.

 

44.

… allows the researcher to ask (and hopefully

14. complete

 

answer) the general question «what is the best

linkage

 

predictor of».

 

45.

A typical result of joining (tree clustering) is the

15. moving av-

 

… .

erage

 

 

smoothing

Part III

Choose the definitions to the terms on the left.

46.

quantitative data

1.

it is a very common general type of pattern in

 

 

 

time series data, where the amplitude of the

 

 

 

seasonal changes increases with the overall

 

 

 

trend (i.e. the variance is correlated with the

 

 

 

mean over the segments of the series).

47.

trend

2.

expresses the degree to which two or more

 

 

 

predictors are related to the dependent va-

 

 

 

riable.

48.

residual value

3.

people assembled in a series of groups pos-

 

 

 

sess certain characteristics and provide data of

 

 

 

qualitative nature in a focused discussion.

150