Learn statistics in English. Учебно-практическое пособие
.pdf
Практикум
hydraulic heads known only to be above the land surface (artesian wells on old maps).
6.Seasonal patterns. Values tend to be higher or lower in certain seasons of the year.
7.Autocorrelation. Consecutive observations tend to be strongly correlated with each other. For the most common kind of autocorrelation in water resources (positive autocorrelation), high values tend to follow high values and low values tend to follow low values.
8.Dependence on other uncontrolled variables. Values strongly covary with water discharge, hydraulic conductivity, sediment grain size, or some other variable.
Methods for analysis of water resources data, whether the simple summarization methods such as those in this text, or the more complex procedures, should recognize these common characteristics.
141
Learn statistics in English
Read the text. Illustrate how the existence of differences relates to the absence of complete determinism. Distinguish the physical sources of randomness in real world.
Roughly – примерно
To split into – расщеплять на, делить
To equalize – сравнять, уравнивать
Requisite – нужный, требуемый
Color-blind – дальтоник
To furnish – предоставлять
Quantum-mechanical – квантово-механический
To ascertain – констатировать, доказывать
Solid – твердый, прочный
Lattice – решётка, плетение
Particles – частицы
Suspended – свободно плавающий
Discreteness – дискретность
Permeation – проникание
Impurities – примеси
Unpredictability and Randomness
A large number of phenomena exist that are neither completely determinate nor completely chaotic. To describe them, one may use a system of nonidentical but mutually comparable objects and then classify them into several groups. Of interest to us might be to what group a given object belongs. We shall illustrate how the existence of differences relates to the absence of complete determinism. Suppose that we are interested in the sex of newborn children. It is known that roughly half of births are boys and half are girls. In other words, the «things» being considered split into two groups. If a strictly valid law existed for the birth of a boy or girl, then it would still be impossible to produce the mechanism which would continually equalize the sexes of babies being born in the requisite proportion (without assuming the effect of the results of prior births on succeeding births, such a premise is meaningless). One may give numerous examples of valid statements like
142
Практикум
«such a thing happens in such and such fraction of the cases», for instance, «1% of males are color-blind». As in the case of the sex of babies, the phenomenon cannot be explained on the basis of determinate laws. It is advantageous to view a set-up of things as a sequence of events proceeding in time.
The absence of determinism means that future events are unpredictable. Since events can be classified in some sort of way, one may ask to what class will a future event belong? But once again (determinism not being present), one cannot furnish an answer in advance. The question is still posed in the given situation. The examples cited suggest a proper way to state the question: how often will a phenomenon of a given class occur in the sequence? We shall speak about chance in precisely such situations and it will be natural to raise such questions and to find answers for them.
Sources of Randomness
We shall now point out a few of the most important existing physical sources of randomness in the real world.
(1)Quantum-mechanical laws. The laws of quantum mechanics are statements about the wave functions of micro-objects. According to these laws, we can specify, for instance, just the wave function of an electron in a field of force. Based on the wave function, only the probability of detecting the electron in some particular region of space may be found – to predict its position is impossible. In exactly the same way, one cannot ascertain the energy of an electron and it is only possible to determine a discrete number of possible energy levels and the probability that the energy of the electron has a specified value. We perceive that the fundamental laws of the microworld make use of the language of probability and thus phenomena in the microworld are random. An important example of a random phenomenon in the microworld is the emission of a quantum of light by an excited atom. Another important example is nuclear reactions.
(2)Thermal motion of molecules. The molecules of any substance are in constant thermal motion. If the substance is a solid, then the molecules range close to positions of equilibrium in a
143
Learn statistics in English
crystal lattice. But in fluids and gases, the molecules perform rather complex movements changing their directions of motion frequently as they interact with one another. The presence of such a motion may be ascertained by watching the movement of microscopic particles suspended in a fluid or gas (this is so-called Brownian motion). This motion is of a random nature and the energies of the individual molecules are also random, that is, the energies of the molecules can assume different values and so one talks about the fraction of molecules having an energy within narrow specified bounds. This is the familiar Maxwell distribution in physics. A simple experiment will convince one that the energies of the molecules are different. Take the phenomenon of boiling water: if all of the molecules had the same energy, then the water would become steam all at once, that is, with an explosion, and this does not happen.
(3) Discreteness of matter. The discreteness of matter leads to the occurrence of randomness in another way. Items (1) and (2) also considered material particles. The following fact should now be noted: the laws of classical physics have been formulated for macrobodies just as if matter filled up space continuously. The discreteness of matter leads to the occurrence of deviations of the actual values of physical quantities from those predicted by the laws. These deviations or «fluctuations» are of a random nature and they affect the course of a process substantially. Thus, the discreteness of the carriers of electricity in metallic conductors – the electrons – is the source of fluctuation currents which are the reason for internal noise in radios. The discreteness of matter results in the mutual permeation of substances. Furthermore the absence of pure substances, that is, the existence of impurities, also results in random deviations from the calculated flow of phenomena.
(d) Cosmic radiation. Experimentation shows that it is irregular (aperiodic and unpredictable) but it conforms to laws that can be studied by probability theory.
144
Практикум
Read the text. Present the basic idea of icon plots. What phases does the analysis of icon plots consist of? Define two categories of icon plots.
Icon – рисунок, символ, пиктограмма
Plot – график
Multidimensional – многомерный
To capitalize on – использовать
Consistent – совместимый, подходящий
Sequence – последовательность
Non-salient – не существенный
To assign – сопоставить, принять
To drop – исключить
Taxonomy – классификация, таксономия
Spoked wheel – колесо со спицами
Hub – центр, ядро
Consecutive – ступенчатый, идущий подряд
Icon Plots
Icon Graphs represent cases or units of observation as multidimensional symbols and they offer a powerful although not easy to use exploratory technique. The general idea behind this method capitalizes on the human ability to "automatically" spot complex (sometimes interactive) relations between multiple variables if those relations are consistent across a set of instances (in this case "icons"). Sometimes the observation (or a "feeling") that certain instances are "somehow similar" to each other comes before the observer (in this case an analyst) can articulate which specific variables are responsible for the observed consistency (Lewicki, Hill, & Czyzewska, 1992). However, further analysis that focuses on such intuitively spotted consistencies can reveal the specific nature of the relevant relations between variables.
The basic idea of icon plots is to represent individual units of observation as particular graphical objects where values of variables are assigned to specific features or dimensions of the objects (usually one case = one object). The assignment is such that
145
Learn statistics in English
the overall appearance of the object changes as a function of the configuration of values.
Thus, the objects are given visual "identities" that are unique for configurations of values and that can be identified by the observer. Examining such icons may help to discover specific clusters of both simple relations and interactions between variables.
Analyzing Icon Plots
The "ideal" design of the analysis of icon plots consists of five phases:
1.Select the order of variables to be analyzed. In many cases a random starting sequence is the best solution. You may also try to enter variables based on the order in a multiple regression equation, factor loadings on an interpretable factor, or a similar multivariate technique. That method may simplify and "homogenize" the general appearance of the icons which may facilitate the identification of nonsalient patterns. It may also, however, make some interactive patterns more difficult to find. No universal recommendations can be given at this point, other than to try the quicker (random order) method before getting involved in the more time-consuming method.
2.Look for any potential regularities, such as similarities between groups of icons, outliers, or specific relations between aspects of icons (e.g., "if the first two rays of the star icon are long, then one or two rays on the other side of the icon are usually short"). The Circular type of icon plots is recommended for this phase.
3.If any regularities are found, try to identify them in terms of the specific variables involved.
4.Reassign variables to features of icons (or switch to one of the sequential icon plots) to verify the identified structure of relations (e.g., try to move the related aspects of the icon closer together to facilitate further comparisons). In some cases, at the end of this phase it is recommended to drop the variables that appear not to contribute to the identified pattern.
5.Finally, use a quantitative method (such as a regression method, nonlinear estimation, discriminate function analysis, or cluster analysis) to test and quantify the identified pattern or at least some aspects of the pattern.
146
Практикум
Taxonomy of Icon Plots
Most icon plots can be assigned to one of two categories: circular and sequential.
Circular icons. Circular icon plots (star plots, sun ray plots, polygon icons) follow a "spoked wheel" format where values of variables are represented by distances between the center ("hub") of the icon and its edges.
Those icons may help to identify interactive relations between variables because the overall shape of the icon may assume distinctive and identifiable overall patterns depending on multivariate configurations of values of input variables.
In order to translate such "overall patterns" into specific models (in terms of relations between variables) or verify specific observations about the pattern, it is helpful to switch to one of the sequential icon plots which may prove more efficient when one already knows what to look for.
Sequential icons. Sequential icon plots (column icons, profile icons, line icons) follow a simpler format where individual symbols are represented by small sequence plots (of different types).
The values of consecutive variables are represented in those plots by distances between the base of the icon and the consecutive break points of the sequence (e.g., the height of the columns). Those plots may be less efficient as a tool for the initial exploratory phase of icon analysis because the icons may look alike. However, as mentioned before, they may be helpful in the phase when some hypothetical pattern has already been revealed and one needs to verify it or articulate it in terms of relations between individual variables.
147
Learn statistics in English
Final Test
Part I
Match the English terms on the left with the Russian ones on the right.
1. random sample |
1. |
оценка параметра |
|
2. |
correlation |
2. |
долгосрочный |
3. |
coefficient of determination |
3. |
обрабатывать |
4. |
rank |
4. |
корреляция |
5. |
handle |
5. |
случайная выборка |
6. |
parameter estimation |
6. |
ранжировать |
7. |
bell-curve |
7. |
вероятностное |
|
long-term |
|
моделирование |
8. |
8. |
квадрат смешанной |
|
|
|
|
корреляции |
9. |
interval scale |
9. |
многомерное измерение |
10. |
probabilistic modeling |
10. |
колокообразная кривая |
11. |
joining (tree clustering) |
11. |
интервальная шкала |
12. |
residual value |
12. |
независимая переменная |
13. |
least squares |
13. |
метод наименьших |
|
predictor variable |
|
квадратов |
14. |
14. |
остаточное значение |
|
15. |
multiple dimension |
15. объединение (древовидная |
|
|
|
|
кластеризация) |
Match the Russian terms on the left with the English ones on the right.
16. |
квадраты отклонений |
1. |
residual value |
17. |
модель скользящей средней |
2. |
value on dimensions |
18. |
оценка по методу наимень- |
3. |
squared deviations |
|
ших квадратов |
|
|
19. |
остаточное значение |
4. |
time series |
20. |
квадрат когерентности |
5. |
smoothing |
21. |
остаточная дисперсия |
6. |
dependent variable |
22. |
значения параметров |
7. |
slope |
23. |
временной ряд |
8. |
residual variance |
148
|
|
|
Final Test |
24. |
зависимая переменная |
9. |
square root |
25. |
сглаживание |
10. |
squared coherency |
26. |
дисперсионный анализ |
11. |
negative exponentially |
|
|
|
weighted smoothing |
27. |
квадратный корень |
12. |
least squares estimation |
28. |
отрицательно взвешенное |
13. |
variance analysis |
|
экспоненциальное |
|
|
|
сглаживание |
|
two-way joining |
29. |
наклон |
14. |
|
30. |
двувходовое объединение |
15. |
moving average model |
Part II
Fill the gaps with the words or word combinations from the given list on the right.
31. |
The critical features of … are ensuring anonymi- |
1. probability |
|
ty, the presence of an objective third party, and |
|
|
constructive feedback to the organization. |
|
32. |
… are data collected at the same or approximate- |
2. Pearson cor- |
|
ly the same point in time. |
relation |
33. |
The exact shape of the … is defined by a func- |
3. the ARIMA |
|
tion which has only two parameters: mean and |
methodology |
|
standard deviation. |
|
34. |
… is derived from the verb to probe meaning to |
4. organization |
|
find out, what is not too easily accessible or un- |
records |
|
derstandable. |
|
35. |
… assumes that the two variables are measured |
5. measure- |
|
on at least interval scales, and it determines the |
ment scale |
|
extent to which values of two variables are |
|
|
«proportional» to each other. |
|
36. |
Factor that determines the amount of informa- |
6. cross- |
|
tion that can be provided by a variable is its type |
sectional data |
|
of … . |
|
37. The most common techniques is … which replaces |
7. survey ques- |
|
|
each element of the series by either the simple or |
tionnaires |
|
weighted average of n surrounding elements, |
|
|
where n is the width of the smoothing «window». |
|
|
|
149 |
Learn statistics in English |
|
|
38. |
… include strategic plans, absentee lists, griev- |
8. outliers |
|
ances field, units of performance per person, and |
|
|
costs of production. |
|
39. |
… have a profound influence on the slope of the |
9. normal dis- |
|
regression line and consequently on the value of |
tribution |
|
the correlation coefficient. |
|
40. |
We may use the neighbors across clusters that |
10. joining (tree |
|
are furthest away from each other; this method |
clustering) |
|
is called … . |
|
41. |
The deviation of a particular point from the re- |
11. multiple |
|
gression line (its predicted value) is called … . |
regression |
42. |
… allows us to uncover the hidden patterns in |
12. the residual |
|
the data and to generate forecasts. |
value |
43. |
The purpose of … is to join together objects into |
13. hierarchical |
|
successively larger clusters, using some measure |
tree |
|
of similarity or distance. |
|
44. |
… allows the researcher to ask (and hopefully |
14. complete |
|
answer) the general question «what is the best |
linkage |
|
predictor of». |
|
45. |
A typical result of joining (tree clustering) is the |
15. moving av- |
|
… . |
erage |
|
|
smoothing |
Part III
Choose the definitions to the terms on the left.
46. |
quantitative data |
1. |
it is a very common general type of pattern in |
|
|
|
time series data, where the amplitude of the |
|
|
|
seasonal changes increases with the overall |
|
|
|
trend (i.e. the variance is correlated with the |
|
|
|
mean over the segments of the series). |
47. |
trend |
2. |
expresses the degree to which two or more |
|
|
|
predictors are related to the dependent va- |
|
|
|
riable. |
48. |
residual value |
3. |
people assembled in a series of groups pos- |
|
|
|
sess certain characteristics and provide data of |
|
|
|
qualitative nature in a focused discussion. |
150
