Добавил:
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Learn statistics in English. Учебно-практическое пособие

.pdf
Скачиваний:
0
Добавлен:
12.08.2026
Размер:
969 Кб
Скачать

Theme VII. Time Series Analysis

near trend slope; 23. mean squared error; 24. the cyclical component; 25. a function minimization algorithm; 26. additive; 27. spectrum analysis;

28.the mean error value; 29. multiplicative; 30. exponential smoothing.

1.is based on the assumption that successive values in the data file represent consecutive measurements taken at equally spaced time intervals.

2.There are two main goals of time series analysis : represented by sequence of observations, and .

3.In time series analysis it is assumed that the data consist of and which usually makes the pattern difficult to identify.

4.Most time series patterns can be described in terms of two basic classes of components: and seasonality.

5.indicates that the relative amplitude of seasonal changes is constant over time, thus it is related to the trend.

6.The most common techniques is which replaces each element of the series by either the simple or weighted average of n surrounding elements, where n is the width of the smoothing «window».

7.can be used instead of means.

8.When the measurement error is very large, or negative exponentially weighted smoothing techniques can be used.

9.Series with relatively few and systematically distributed points can be smoothed with .

10.If there is a clear monotonous nonlinear component, the data first need to be transformed to remove the nonlinearity and usually can be used.

11.is defined as correlational dependency of order k between each i’th element of the series and the i-k’th element and measured by autocorrelation; k is usually called .

12.Seasonal patterns of the time series can be examined via .

13.The correlogram displays graphically and numerically the autocorrelation function, that is, for consecutive lags in a specified range of lags.

14.is similar to autocorrelation, except that when calculating it, the autocorrelations with all elements within the lag are partialled out.

131

Learn statistics in English

15.allows us to uncover the hidden patterns in the data and to generate forecasts.

16.The input series for ARIMA needs to be , that is, it should have a constant mean, variance, and autocorrelation through time.

17.If the series is differenced once, and there are no autoregressive parameters in the model, then the constant represents the mean of the differenced series, and therefore of the un-differenced series.

18.During the parameter estimation phase is used to maximize the likelihood of the observed series, given the parameter values.

19.has become very popular as a forecasting method for a wide variety of time series data.

20.is simply computed as the average error value.

21.is computed as the sum of the squared error values.

22.Seasonal components can be in nature or .

23.is different from the seasonal component, it is longer than one season, and different cycles can be of different lengths.

24.The purpose of is to identify the seasonal fluctuations of different lengths.

25.The square root of the sum of the squared cross-density and quad-density values is called .

26.is computed by the dividing the cross-amplitude value by the spectrum density estimates for one of the two series in the analysis.

132

Практикум

Практикум

Reading Comprehension Practice

Read the text and give the short review of Quetelet’s ideas. Why he demonstrated the importance of statistics to social science?

Forerunner предшественник

Contemporaries современники Prior до

Evidence доказательство Accidental стихийный, случайный Random беспорядочный

Betterment улучшение, совершенствование Mentor наставник

Celestial небесный

Treatise трактат, рассуждение

Cross-tabulation составление многокоординатной таблицы Mean среднее значение

A priori заранее, предварительно

To ascertain констатировать, определять, устанавливать Species категория, вид

Savant учёный

Hereditarianism точка зрения основывающаяся на призна- нии ведущей роли наследственности в формировании пове- дения

To access достигать, иметь доступ Inductive побуждающий, вводный

Quetelet’s view on Statistics in the Social Sciences

Adolphe Quetelet (1796-1874) was a Belgian social statistician and a forerunner in demonstrating the importance of statistics to social science. Quetelet was convinced that knowledge of causes influenced the course of human affairs more profoundly than his

133

Learn statistics in English

contemporaries appreciated. Prior to the 1820s this type of knowledge was generally regarded in European intellectual circles as evidence of God's hand in ordering the universe. Quetelet argued that the perfection of science could be judged by the ease in which it could be approached by calculation. Statistics emphasize the regularity of social processes, eliminating accidental and random elements in social events to enable a discovery of underlying laws governing phenomena. In this way Quetelet was a pioneer in developing a whole new methodology to be used in the social sciences. He felt that using statistics to gather social knowledge was the solution for the betterment of society.

In 1820, Quetelet was elected to the Royal Academy of Sciences in Brussels. During this time he was engaged in a variety of projects where he applied algebra and geometry to demographic tables. In one instance, he utilized Belgian birth and mortality tables as the basis for the construction of insurance rates. He learned of the potential for such application of calculation to social matter from Malthus's Essay on Population, Fourier's statistical research on Paris and its environs in the early 1820s, and, most of all, from the work of his friend and mentor, Laplace, on celestial mechanics (mechanique celeste), on the principles of probabilistic theory, and on the method of least squares. Originating from the gaming tables of the seventeenth century, the subject of probability had moved rapidly ahead by the early nineteenth century and the stage was nicely set for Quetelet to begin to apply statistical methods on a wide scale.

In his publication «A Treatise On Man», Quetelet calculated the average weight and height of subjects and cross-tabulated these with sex, age, occupation, and geographical region. In combination these average values produced a statistically arrived at fabrication which Quetelet termed `the average man.' `Everything occurs then as though there existed a type of man, from which all other men differed more or less . . . Each people presents its mean, and the different variations from this mean in numbers that may be calculated ‘a priori'. The average man is a concept that is critical to understanding Quetelet's writings because it remains central to all of his statistical

134

Практикум

studies of society. He believed that `If the average man were ascertained for one nation, he could represent the type of that nation. If he could be ascertained according to the mass of men, he would represent the type of human species altogether'.

Quetelet details an important and extensive role for his average man. He felt that many scientists would find this property useful, `The artist, the man of literature, and the savant, will afterwards choose from among these materials best suited to the subject of their studies'. He goes on to detail benefits to the physician, naturalist, and politician. In these fields it is impossible to discuss or make judgments upon individuals without using comparisons of a perceived `normal' condition, which of course is reflected in the average man. The statistics Quetelet gathered have great historical significance. Doctors, criminologists, and anthropologists could analyze records of the physical properties of man. Biological variations suggested various means of identification. Quetelet provided the foundations for a deterministic criminology that was subsequently adopted by Lombroso who emphasized biologism and mental hereditarianism. Registers of mental traits were used in the first studies of experimental psychology and later also in the science of education. All moral statisticians could access this type of data to aid the inductive study of social life.

Read the text. Distinguish the main characteristics of Structural Equation Modeling. Comment upon them. What approaches does SEM use?

Structural equation modeling моделирование структурного уравнения

Nonlinearity нелинейность

Path analysis пат-анализ, анализ троп

Analysis of covariance ковариационный анализ

Mediating variables промежуточные переменные

Susceptible допускающий

Misspecification неправильная спецификация

135

Learn statistics in English

Robust устойчивый

Goodness-of-fit критерий адекватности Variance дисперсия

Covariance ковариация

Consistent совместимый

Causal причинно обусловленный Deficient несовершенный, дефективный

Post-hoc апостериорный (по полученным результатам) Cross-validation перекрестная проверка

Ambiguity неточность, двойственность Utmost наиболее отдаленный

Structural Equation Modeling

Structural equation modeling (SEM) grows out of and serves purposes similar to multiple regression, but in a more powerful way which takes into account the modeling of interactions, nonlinearities, correlated independents, measurement error, correlated error terms, multiple latent independents each measured by multiple indicators, and one or more latent dependents also each with multiple indicators. SEM may be used as a more powerful alternative to multiple regression, path analysis, factor analysis, time series analysis, and analysis of covariance. That is, these procedures may be seen as special cases of SEM, or, to put it another way, SEM is an extension of the general linear model (GLM) of which multiple regression is a part.

Advantages of SEM compared to multiple regression include more flexible assumptions (particularly allowing interpretation even in the face of multicollinearity), use of confirmatory factor analysis to reduce measurement error by having multiple indicators per latent variable, the attraction of SEM's graphical modeling interface, the desirability of testing models overall rather than coefficients individually, the ability to test models with multiple dependents, the ability to model mediating variables rather than be restricted to an additive model (in OLS regression the dependent is a function of the Var1 effect plus the Var2 effect plus the

136

Практикум

Var3 effect, etc.), the ability to model error terms, the ability to test coefficients across multiple between-subjects groups, and ability to handle difficult data (time series with autocorrelated error, nonnormal data, incomplete data). Moreover, where regression is highly susceptible to error of interpretation by misspecification, the SEM strategy of comparing alternative models to assess relative model fit makes it more robust.

SEM is usually viewed as a confirmatory rather than exploratory procedure, using one of three approaches:

1.Strictly confirmatory approach: A model is tested using SEM goodness-of-fit tests to determine if the pattern of variances and covariances in the data is consistent with a structural (path) model specified by the researcher. However as other unexamined models may fit the data as well or better, an accepted model is only a notdisconfirmed model.

2.Alternative models approach: One may test two or more causal models to determine which has the best fit. There are many good- ness-of-fit measures, reflecting different considerations, and usually three or four are reported by the researcher. Although desirable in principle, this AM approach runs into the real-world problem that in most specific research topic areas, the researcher does not find in the literature two well-developed alternative models to test.

3.Model development approach: In practice, much SEM research combines confirmatory and exploratory purposes: a model is tested using SEM procedures, found to be deficient, and an alternative model is then tested based on changes suggested by SEM modification indexes. This is the most common approach found in the literature. The problem with the model development approach is that models confirmed in this manner are post-hoc ones which may not be stable (may not fit new data, having been created based on the uniqueness of an initial dataset). Researchers may attempt to overcome this problem by using a cross-validation strategy under which the model is developed using a calibration data sample and then confirmed using an independent validation sample.

Regardless of approach, SEM cannot itself draw causal arrows in models or resolve causal ambiguities. Theoretical insight and judgment by the researcher is still of utmost importance.

137

Learn statistics in English

SEM is a family of statistical techniques which incorporates and integrates path analysis and factor analysis. In fact, use of SEM software for a model in which each variable has only one indicator is a type of path analysis. Use of SEM software for a model in which each variable has multiple indicators but there are no direct effects (arrows) connecting the variables is a type of factor analysis. Usually, however, SEM refers to a hybrid model with both multiple indicators for each variable (called latent variables or factors), and paths specified connecting the latent variables. Synonyms for SEM are covariance structure analysis, covariance structure modeling, and analysis of covariance structures. Although these synonyms rightly indicate that analysis of covariance is the focus of SEM, be aware that SEM can also analyze the mean structure of a model.

Read the text. Comment on characteristics of water resources data.

To gain извлекать Inconclusive неубедительный To convey передавать

Sulfate сульфат Percentile процентиль

Robust робастный Median медиана Aquifer водоносный Skewness асимметрия

Interquartile интерквартильный, вероятный Bound норма, граничная оценка Lognormal логарифмически нормальный

Density function плотность распределения, функция плот- ности

Threshold предельная величина Censored data усеченные данные

Water discharge водосброс

138

Практикум

Statistical Methods in Water Resources

When determining how to appropriately analyze any collection of data, the first consideration must be the characteristics of the data themselves. Little is gained by employing analysis procedures which assume that the data possess characteristics which in fact they do not. The result of such false assumptions may be that the interpretations provided by the analysis are incorrect, or unnecessarily inconclusive. Therefore we begin with a discussion of the common characteristics of water resources data. These characteristics will determine the selection of appropriate data analysis procedures.

One of the most frequent tasks when analyzing data is to describe and summarize those data in forms which convey their important characteristics. "What is the sulfate concentration one might expect in rainfall at this location"? "How variable is hydraulic conductivity"? "What is the 100 year flood" (the 99th percentile of annual flood maxima)? Estimation of these and similar summary statistics are basic to understanding data. Characteristics often described include: a measure of the center of the data, a measure of spread or variability, a measure of the symmetry of the data distribution, and perhaps estimates of extremes such as some large or small percentile. This text discusses methods for summarizing or describing data.

This text also quickly demonstrates the use of robust and resistant techniques. The reasons why one might prefer to use a resistant measure, such as the median, over a more classical measure such as the mean, are explained.

The data about which a statement or summary is to be made are called the population, or sometimes the target population. These might be concentrations in all waters of an aquifer or stream reach, or all stream flows over some time at a particular site. Rarely are all such data available to the scientist. It may be physically impossible to collect all data of interest (all the water in a stream over the study period), or it may just be financially impossible to collect them.

Instead, a subset of the data called the sample is selected and measured in such a way that conclusions about the sample may be extended to the entire population. Statistics computed from the sample are only inferences or estimates about characteristics of the

139

Learn statistics in English

population, such as location, spread, and skewness. Measures of location are usually the sample mean and sample median. Measures of spread include the sample standard deviation and sample interquartile range. Use of the term "sample" before each statistic explicitly demonstrates that these only estimate the population value, the population mean or median, etc. As sample estimates are far more common than measures based on the entire population, the term "mean" should be interpreted as the "sample mean", and similarly for other statistics used in this text. When population values are discussed they will be explicitly stated as such.

Characteristics of Water Resources Data

Data analyzed by the water resources scientist often have the following characteristics:

1.A lower bound of zero. No negative values are possible.

2.Presence of 'outliers', observations considerably higher or lower than most of the data, which infrequently but regularly occur. Outliers on the high side are more common in water resources.

3.Positive skewness, due to items 1 and 2. An example of a skewed distribution, the lognormal distribution, is presented in figure 1.1. Values of an observation on the horizontal axis are plotted against the frequency with which that value occurs. These density functions are like histograms of large data sets whose bars become infinitely narrow. Skewness can be expected when outlying values occur in only one direction.

4.Non-normal distribution of data, due to items 1–3 above. Figure 1.2 shows an important symmetric distribution, the normal. While many statistical tests assume data follow a normal distribution as in figure 1.2, water resources data often look more like figure 1.1. In addition, symmetry does not guarantee normality. Symmetric data with more observations at both extremes (heavy tails) than occurs for a normal distribution are also non-normal.

5.Data reported only as below or above some threshold (censored data). Examples include concentrations below one or more detection limits, annual flood stages known only to be lower than a level which would have caused a public record of the flood, and

140