Learn statistics in English. Учебно-практическое пособие
.pdf
Theme VI. Multiple regression
Extreme, a экстремальный. Square root – квадратный корень.
Assume, v предполагать, допускать. Assumption, n предположение. Affect, v воздействовать, влиять.
Minor, a малый.
Prudent, a благоразумный.
Bivariate, a двумерный. Curvature, n кривизна.
Robust, a устойчивый.
Violation, n нарушение, отклонение.
Limitation, n ограничение. Ascertain, v установить.
Causal, a причинный. Seductive, a соблазнительный. Plug in – включать.
Estimate, n оценка. Replicate, v повторять.
Multicollinearity, n мультиколлинеарность.
Matrix, n матрица.
Redundant, a излишний, избыточный. Apparent, a видимый, истинный.
Tolerance, n допустимое отклонение от чего-либо, толерант- ность.
Semi-partial – получастный.
Remedy, n средство.
Centered polynomial model – центрированная полиномиальная модель.
Polynomial, n многочлен, полином.
Subtract, v вычитать.
Final assignments to the text
Choose 5 terms from the text, write them down, translate and remember.
91
Learn statistics in English
Choose the definitions to the terms on the left, translate them and learn.
multiple regression |
is an indicator of how well the model |
|
fits the data. |
least squares estimation |
is referred to as the slope. |
intercept |
the signs of the regression we look at to |
|
interpret the direction of the relation- |
|
ship between variables. |
regression coefficient |
general purpose of this analysis is to |
|
learn more about the relationship be- |
|
tween several independent or predictor |
|
variables and a dependent or criterion |
|
variables. |
partial correlation |
a procedure when the line is computed |
|
so that the squared deviations of the |
|
observed points from that line are mi- |
|
nimized. |
residual value |
is also referred to as the constant. |
R-square value |
is a type of correlation when variable x1 |
|
is correlated with the y variable, after |
|
controlling for all other independent |
|
variables. |
correlation coefficient R the deviation of a particular point from
|
the regression line (its predicted value). |
B-coefficients |
expresses the degree to which two or |
|
more predictors are related to the de- |
|
pendent variable. |
92 |
|
Theme VI. Multiple regression
Give the definitions to the following terms:
R-square value; partial correlation; multiple regression; residual value; B-coefficient; negative correlation; least squares estimation.
Translate into English:
1.Общее назначение множественной регрессии состоит в анализе связи между несколькими независимыми пере- менными и зависимой переменной.
2.Процедура множественной регрессии может определить какие позиции лежат ниже линии регрессии, какие ле- жат выше линии регрессии, а какие адекватны.
3.Множественная регрессия позволяет исследователю задать вопросотом, что является лучшимпредиктором для.
4.Термин «множественная» указывает на наличие не- скольких предикторов или регрессоров, которые исполь- зуются в модели.
5.Общая вычислительная задача, которую требуется решать при анализе методом множественной регрессии, состоит в подгонкепрямой линии к некоторому наборуточек.
6.Программа строит линию регрессии так, чтобы миними- зировать квадраты отклонений этой линии от наблюдае- мых точек и на эту общую процедуру иногда ссылаются как на оцениваниепометоду наименьших квадратов.
7.Прямая линия на плоскости задается уравнением: пере- менная Y может быть выражена через константу (a) и уг- ловой коэффициент (b), умноженный на переменную X.
8.Константу называют свободным членом, а угловой ко- эффициент – регрессионным или В-коэффициентом.
9.Регрессионные коэффициенты представляют независи- мые вклады каждой независимой переменной в предска- зание зависимой переменной. Другими словами пере-
менная X1 коррелирует с переменной Y после учета влияния всех других независимых переменных. Этот тип корреляции также называют частной корреляцией.
10.Если одна величина коррелирована с другой, то это может быть отражением того факта, что они обе коррелированы с третьейвеличиной или совокупностью величин.
93
Learn statistics in English
11.Линия регрессии выражает наилучшее предсказание зави- симой переменной (Y) по независимым переменным (X).
12.Отклонение отдельной точки от линии регрессии (от предсказанного значения) называется остатком.
13.Если связь между переменными X и Y отсутствует, то от- ношение остаточной изменчивости переменной Y к ис- ходной дисперсии равно 1.0.
14.Если X и Y жестко связаны, то остаточная изменчивость отсутствует, и отношение дисперсий будет равно 0.0.
15.В большинстве случаев отношение будет лежать где-то между значениями 0.0 и 1.0. 1.0 минус это отношение называется R-квадратом или коэффициентом детерми- нации.
16.Значение R-квадрата является индикатором степени подгонки модели к данным.
17.Степень зависимости двух или более предикторов (неза- висимых переменных или переменных X) с зависимой переменной (Y) выражается с помощью коэффициента детерминации.
18.Для интерпретации направления связи между перемен- ными смотрят на знаки (+ или -) регрессионных коэф- фициентов или В-коэффициентов.
19.Если В-коэффициент положителен, то связь этой пере- менной с зависимой переменной положительна; если В- коэффициент отрицателен, то и связь носит отрица- тельный характер; если В-коэффициент равен 0, связь между переменными отсутствует.
20.Если нелинейность связи очевидна, то можно рассмот- реть или преобразования переменных или явно допус- тить включение нелинейных членов.
21.Основное концептуальное ограничение всех методов регрессионного анализа состоит в том, что они позволя- ют обнаружить только числовые зависимости, а не ле- жащие в их основе причинные связи.
22.Необходимо использовать, по крайней мере, от 10 до 20 наблюдений на одну переменную, в противном случае оценки регрессионной линии будут, вероятно, очень
94
Theme VI. Multiple regression
ненадежными и, скорее всего, невоспроизводимыми для желающих повторить исследование.
23.Подгонка полиномов высших порядков от независимых переменных с ненулевым средним может создать боль- шие трудности с мультиколлинеарностью.
24.Решением в данном случае является процедура центри- рования независимой переменной, т.е. вначале вычесть из переменной среднее, а затем вычислять многочлены.
25.Выбросы (т.е. экстремальные наблюдения) могут вызвать серьезное смещение оценок, «сдвигая» линию регрессии в определенном направлении и тем самым, вызывая смещение регрессионных коэффициентов.
Test
1. Match the English terms on the left with the Russian ones on the right.
1. subtract |
1. |
коэффициент детерминации |
|
2. multiple regression |
2. |
двумерный |
|
3. multicollinearity |
3. |
вычитать |
|
4. regression equation |
4. |
множественная регрессия |
|
5. least squares |
5. |
в среднем |
|
6. bivariate |
6. |
частная корреляция |
|
7. residual value |
7. |
метод наименьших квадратов |
|
8. predictor variable |
8. |
независимая переменная |
|
9. polynomial |
9. |
остаточное значение |
|
10. R-square |
10. |
многочлен |
|
11. slope |
11. |
уравнение регрессии |
|
12. on the average |
12. |
мультиколлинеарность |
|
13. dependent variable |
13. |
угловой коэффициент |
|
14. partial correlation |
14. |
зависимая переменная |
|
2. Match the Russian terms on the left with the English ones on the right.
1. |
значения параметров |
1. independent variable |
2. |
центрированная |
2. outlier |
|
полиномиальная модель |
|
95
Learn statistics in English |
|
||
3. |
квадраты отклонений |
3. residual variance |
|
4. |
квадратный корень |
4. negative correlation |
|
5. |
выброс |
5. squared deviations |
|
6. |
получастный |
6. centered polynomial model |
|
7. |
отклонение |
7. value on dimensions |
|
8. |
оценка по методу |
8. semi-partial |
|
|
наименьших квадратов |
|
|
9. |
независимая переменная |
9. square root |
|
10. |
остаточная дисперсия |
10. constant |
|
11. константа |
11. deviation |
||
12. |
отрицательная корреляция |
12. intercept |
|
13. |
свободный член |
13. confidence interval |
|
14. |
доверительный интервал |
14. least squares estimation |
|
3.Fill the gaps with the words or word combinations from the given list.
1. the residual value; 2. the polynomials; 3. multiple regression; 4. the correlation coefficient R; 5. multicollinearity; 6. linear regression; 7. the square root; 8. prediction; 9. to subtract; 10. the values of dimensions;
11.R-square value; 12. regression equation; 13. B coefficients; 14. a constant (a); 15. a slope (b).
1.To interpret the direction of the relationship between variables, we look at the signs of regression or … .
2.… will be highly correlated due to the mean of the primary independent variable.
3.The deviation of a particular point from the regression line (its predicted value) is called … .
4.The solution of multicollinearity is to center the independent variable i.e. … the mean and then to compute the polynomials.
5.… allows the researcher to ask (and hopefully answer) the general question «what is the best predictor of».
6.The degree to which two or more predictors are related to the dependent variable is expressed in … , which is … of R- square.
7.… is common problem in many correlation analyses.
96
Theme VI. Multiple regression
8.The smaller the variability of the residual values around the regression line relative to the overall variability, the better is our … .
9.… can be used in a multiple regression analysis to build … .
10.A line in a two dimensional or two variable space is defined by the equation: Y = a + b * X. The Y can be expressed in terms of … and … times the X variables.
11.… is an indicator of how well the model fits the data.
12.The goal of … procedures is to fit a line through the points.
97
Learn statistics in English
Theme VII.
Time Series Analysis
Read the text, translate it with the help of the vocabulary and be ready to speak about the main idea of each part.
We will first review techniques used to identify patterns in time series data (such as smoothing and curve fitting techniques and autocorrelations), then we will introduce a general class of models that can be used to represent time series data and generate predictions (autoregressive and moving average models). Finally, we will review some simple but commonly used modeling and forecasting techniques based on linear regression.
General Introduction
We will review techniques that are useful for analyzing time series data, that is, sequences of measurements that follow nonrandom orders. Unlike the analyses of random samples of observations that are discussed in the context of most other statistics, the analysis of time series is based on the assumption that successive values in the data file represent consecutive measurements taken at equally spaced time intervals.
Two Main Goals
There are two main goals of time series analysis: (a) identifying the nature of the phenomenon represented by the sequence of observations, and (b) forecasting (predicting future values of the time series variable). Both of these goals require that the pattern of observed time series data is identified and more or less formally described. Once the pattern is established, we can interpret and integrate it with other data (i.e., use it in our theory of the investigated phenomenon, e.g., sesonal commodity prices). Regardless of the depth of our understanding and the validity of our interpretation (theory) of the phenomenon, we can extrapolate the identified pattern to predict future events.
98
Theme VII. Time Series Analysis
Identifying Patterns in Time Series Data
Systematic Pattern and Random Noise
As in most other analyses, in time series analysis it is assumed that the data consist of a systematic pattern (usually a set of identifiable components) and random noise (error) which usually makes the pattern difficult to identify. Most time series analysis techniques involve some form of filtering out noise in order to make the pattern more salient.
Two General Aspects of Time Series Patterns
Most time series patterns can be described in terms of two basic classes of components: trend and seasonality. The former represents a general systematic linear or (most often) nonlinear component that changes over time and does not repeat or at least does not repeat within the time range captured by our data (e.g., a plateau followed by a period of exponential growth). The latter may have a formally similar nature (e.g., a plateau followed by a period of exponential growth), however, it repeats itself in systematic intervals over time. Those two general classes of time series components may coexist in real-life data. For example, sales of a company can rapidly grow over years but they still follow consistent seasonal patterns (e.g., as much as 25% of yearly sales each year are made in December, whereas only 4% in August).
This general pattern is well illustrated in a "classic" Series G data set representing monthly international airline passenger totals (measured in thousands) in twelve consecutive years from 1949 to 1960. If you plot the successive observations (months) of airline passenger totals, a clear, almost linear trend emerges, indicating that the airline industry enjoyed a steady growth over the years (approximately 4 times more passengers traveled in 1960 than in 1949). At the same time, the monthly figures will follow an almost identical pattern each year (e.g., more people travel during holidays then during any other time of the year). This example data file also illustrates a very common general type of pattern in time series data, where the amplitude of the seasonal changes increases with the
99
Learn statistics in English
overall trend (i.e., the variance is correlated with the mean over the segments of the series). This pattern which is called multiplicative seasonality indicates that the relative amplitude of seasonal changes is constant over time, thus it is related to the trend.
Trend Analysis
There are no proven "automatic" techniques to identify trend components in the time series data; however, as long as the trend is monotonous (consistently increasing or decreasing) that part of data analysis is typically not very difficult. If the time series data contain considerable error, then the first step in the process of trend identification is smoothing.
Smoothing. Smoothing always involves some form of local averaging of data such that the nonsystematic components of individual observations cancel each other out. The most common technique is moving average smoothing which replaces each element of the series by either the simple or weighted average of n surrounding elements, where n is the width of the smoothing "window. Medians can be used instead of means. The main advantage of median as compared to moving average smoothing is that its results are less biased by outliers (within the smoothing window). Thus, if there are outliers in the data (e.g., due to measurement errors), median smoothing typically produces smoother or at least more "reliable" curves than moving average based on the same window width. The main disadvantage of median smoothing is that in the absence of clear outliers it may produce more "jagged" curves than moving average and it does not allow for weighting.
In the relatively less common cases (in time series data), when the measurement error is very large, the distance weighted least squares smoothing or negative exponentially weighted smoothing techniques can be used. All those methods will filter out the noise and convert the data into a smooth curve that is relatively unbiased by outliers. Series with relatively few and systematically distributed points can be smoothed with bicubic splines.
Fitting a function. Many monotonous time series data can be adequately approximated by a linear function; if there is a clear
100
