AÖF Soru Bankası

Econometrıcs II (ENG)Ünite 2 Soru-Cevap

Econometrıcs II (ENG) (IKT326U) soru-cevapları.

What are endogenous regressors, why do they pose a problem in econometric modeling, and how can these problems be effectively addressed? 

 

The term “endogenous regressors” implies that certain variables within the model are determined within the system itself. Although our model may consist of only one equation, some variables can be influenced and shaped by hidden or ignored equations operating within the system. This interdependence can lead to the endogeneity problem, impacting the reliability of our results. Our exploration then proceeds to identify situations where endogeneity problems naturally arise. Understanding these scenarios is crucial as it lays the foundation for addressing such issues effectively. To tackle endogeneity, we introduce a two-step methodology utilizing instrumental variables – exogenous variables that are unaffected by the endogenous regressors.

What causes the endogeneity problem in regression models, how does it relate to the violation of the zero conditional expectation assumption, and what are the consequences for interpreting regression coefficients?

 

Endogeneity problem in regression models occurs when the Gauss-Markov assumption on zero conditional expectation is violated. Consider a simple linear regression model Y=β0+β1X+u. The zero conditional expectation assumption can be written as E[u | X ] = cov(u, X ) = 0. In this case, we can say that the error term is independent of the regressor; in other words, the error term does not carry any common information with the regressor. This assumption is critical because it is required to ensure the ceteris paribus assumption when we interpret the coefficient β1. If this assumption fails, we cannot control other variables since the error term and regressor move together. Moreover, the presence of such a situation engenders bias, sometimes called endogeneity bias. Consequently, the endogeneity problem (bias) results from dependence between the model’s error term and regressor(s). Notice that we can generalize this problem to the multiple linear regression model with the same principles.

Why is the zero conditional expectation assumption often violated in econometric models, and what are some common sources and examples of endogeneity problems in economics and finance?

 

The crucial question is: why and how is this assumption violated in econometric models? First, endogeneity is not a rare problem in economics, finance, and other social sciences. On the contrary, it is one of the primary matters that a practitioner should focus on when proposing her/his regression model. For instance, the endogeneity problem is a natural consequence of phenomena such as errors in variables and simultaneous equations in the cross-sectional data. As indicated in Gujarati (2022), in many examples of economic or financial regression one way or unidirectional cause-and-effect relationship is not meaningful. The authors claim that in such models there exist more than one equation that determines the whole system (Gujarati, 2022). In these equations, the dependent variables may also appear as regressors. Among many examples, “Demand and Supply Model”, Keynesian Consumption (or saving) Function”, “Wage and Price Equations”, and “IS-LM model” are popular examples (Gujarati, 2022).

How do the Supply and Demand model, the Investment-Savings identity, and the IS-LM model illustrate the presence of simultaneous equations in economics, and why does this structure naturally lead to endogeneity problems in parameter estimation?

Some examples of simultaneous equations in economics can be listed as follows

1) Supply and Demand Equations: In a simple market, you may have two simultaneous equations representing the supply and demand relationship for a certain good or service where the Demand Equation is given by QD=a-bP and the Supply Equation is given by QS=c-dP. We define QD and QS as the quantity demanded and the quantity, respectively; P is the price of the good/service a, b, c, and d are constants that represent the demand and supply parameters. The solution to this system of equations will give you the equilibrium price (P) and quantity (Q) at which the quantity demanded equals the quantity supplied. One may also want to estimate these parameters, however by its nature, this estimation is prone to the endogeneity problem as we explain later in this subsection.

2) Investment and Savings Equations: In a closed economy without government involvement, investment (I ) and savings (S) are usually assumed to be equal. Therefore, the following simultaneous equation holds:

I = S, where I is the total investment in the economy and S is the total savings in the economy. This equation represents the savings-investment equilibrium, where the total savings in the economy are being invested. Similarly, if the variables are represented with two separate equations, one can estimate the unknown system parameters. However, the endogeneity problem is again inevitable in this case.

3) IS-LM Model: The IS-LM model is a macroeconomic model that analyzes the relationship between interest rates (i) and output or income (Y). The model consists of two simultaneous equations namely, IS Curve: Y = C + I (Y, i) and LM Curve: M / P = L(i, Y), where Y is the national income or output; C is consumption expenditure; I is investment expenditure, which is a function of income (Y) and interest rates (i); G is government expenditure; M/P represents the real money supply (M) divided by the price level (P);

L is the demand for real money balances, which is a function of interest rates (i) and income (Y). The solution to the IS-LM model provides the equilibrium income/output (Y) and the interest rate (i) in the economy. Similarly, if one wants to estimate the system parameters via OLS, the endogeneity problem potentially appears.

What is meant by “errors in variables” in economics and finance, how do measurement errors in dependent and explanatory variables differ, and why do they lead to endogeneity problems in regression analysis?

 

In economics and finance, “errors in variables,” also known as measurement error, refers to the situation where one or more variables used in a statistical analysis are measured imprecisely or with error. This measurement error can occur in both the dependent variable and the explanatory variables in a regression analysis. It can lead to biased and inefficient parameter estimates, potentially impacting the validity of the statistical analysis and the conclusions drawn from it. The bias of the parameter estimates in this case is due to naturally emerging endogeneity problems. We discuss this issue after introducing two main types of errors in variables.

First, we consider Measurement Error in Dependent Variables. This situation appears when the dependent variable is measured with error, it means that the observed values of Y differ from the true values. This could be due to data collection problems, survey errors, or other inaccuracies in measuring the variable of interest. Measurement error in the dependent variable can lead to biased and inconsistent coefficient estimates, potentially affecting the results of regression analysis.

Second, we consider Measurement Error in Explanatory Variable which appears when the explanatory variable is measured with error, the observed values of X differ from their true values. This type of error is also known as “exogenous measurement error” or “errors-in-variables” (EIV). Measurement error in explanatory variables can lead to biased coefficient estimates and can also lead to attenuation bias, where the true relationship between X and Y is underestimated or obscured. This situation is a typical example of an endogeneity problem.

How does measurement error in income lead to an endogeneity problem in the Keynesian consumption function, and how is this reflected in the observed regression model? 

To explain the connection between errors in variables and Endogeneity problem, we consider a simple economic model, in which some variables may be mismeasured or obtained from a noisy source. Now the Keynesian consumption function:

C* i = β0 + β1 Y * i + u* i ,

where, C* i is the actual personal consumption, Y * i is the actual personal income of individual i; β0 is the autonomous consumption, and β1 is the marginal propensity to consume; u* i is the unobserved error term. Now, suppose that individuals misreport their income and consumptions. That is, the observed consumption is given as Ci = C* i + eci, and the observed income is given as Yi = Y * i + eyi where eci and eyi are measurement errors for consumption and income, respectively. For brevity, we set eci =0 and eyi ~ N (0, σ2y) with σ2y>0. In this case, only income is mismeasured. The new regression model is as following:

Ci = β0 + β1 Yi + ui .

With a simple algebra, we can show that ui = u* i - β1 eyi . Notice that even if u* i and eyi are independent, Yi and ui are dependent on each other. This situation is a clear violation of the zero conditional mean assumption; hence the OLS estimators are biased and inconsistent.

How does the violation of the zero conditional mean assumption in a multiple linear regression model lead to endogeneity bias, and how are endogenous and exogenous regressors defined in this context?

 

We mathematically formulate the endogeneity bias in a multiple linear regression framework. Consider the multiple linear regression model with k +1 ≥ 1 coefficients:

Yi = β0 + β jX ji + ui j=1k Σ

where, we have n observation from the sample {(Yi , X1i ,..., Xki )}n i=1. For further simplification, we define k vector Xi = [ X1i ,..., Xki ], and X = [ X1 ,..., Xn ]' is the matrix containing all observations of the regressors. We can state the zero conditional mean assumption as E[εi | X ] = 0 . Notice that this condition implies that the error term is jointly independent of all regressors. Accordingly, if at least one regressor is not independent of the error term, we say that the zero conditional mean assumption is violated. As a result of this violation, the endogeneity problem arises. Now, we discuss how the violation of the zero conditional mean assumption engenders bias in the OLS estimation. Moreover, the condition mentioned above is called the “strong form exogeneity” of the regressors. Additionally, we can express the “weak form exogeneity” as cov(Xji , ui) = 0 ∀j ∈{1,..., k} . If there exists a regressor, say cov(Xji , ui) ≠ 0, we call the variable Xj “endogenous” regressor, otherwise, it is called “exogenous regressor”. In a regression model, we may have endogenous and exogenous regressors at the same time.

What is the Keynesian consumption function, and why is it important for understanding aggregate demand, economic stability, and fiscal policy in macroeconomic analysis?

 

This subsection depicts the endogeneity bias in a simple simultaneous equation model with two endogenous variables. We consider the simple Keynesian consumption function, which is an important concept in economics, particularly in the field of macroeconomics. It was proposed by the renowned economist John Maynard Keynes as part of his general theory of income, employment, and output. The consumption function describes the relationship between disposable income and consumer spending in an economy.

Keynesian consumption function is particularly important since it helps us to understand aggregate demand. In macroeconomics, aggregate demand represents the total demand for goods and services in an economy. Consumer spending is a significant component of aggregate demand. The Keynesian consumption function helps economists understand how changes in disposable income influence consumer spending and, in turn, impact overall aggregate demand.

Moreover, Keynesian economics emphasizes the importance of maintaining stable aggregate demand to achieve full employment and economic stability. The consumption function plays a key role in this regard, as it highlights the relationship between income and consumption. By understanding how changes in consumption respond to changes in income, policymakers can devise measures to stabilize the economy during periods of recession or inflation.

To sum up, the Keynesian consumption function is a fundamental tool for understanding the relationship between income and consumer spending, guiding fiscal policy decisions, and providing insights into the dynamics of aggregate demand and economic stability. It continues to be a valuable concept for macroeconomic analysis and policy formulation.

Why are reduced form equations problematic for OLS estimation in the presence of endogenous variables, and what are the implications for identifying structural parameters like α1 and α2?

 

Equations in are known as Reduced Form Equations, parameters δ1 and δ2 are called Reduced Form Parameters, and error terms vp and vq are Reduced Form Error Terms.

Notice that both Equations in Price are endogenous, since it is correlated with the error terms and appears as a regressor. On the one hand, the OLS estimators are biased in this case whatever equation (demand or supply) is used for the estimation. On the other hand, one cannot identify α1 and α2 at the same time even though β1 is identifiable. As a result, the OLS does not produce reliable results under the original supply and demand equations, and also under the reduced form equations.

How does the Two-Stage Least Squares (2SLS) method use instrumental variables to address endogeneity and provide consistent estimates in econometric models?

 

The tool that we utilize is called the “Two-Stage Least Squares” method (2SLS). As its name suggests, this method requires two sequential estimations of linear regression models. The first stage aims to clean out the problematic parts of the endogenous regressors by using variables called “instruments” (or instrumental variables). The selection of the instruments relies on the condition that instruments depend on the endogenous regressors but do not depend on the error term of the original regression. After cleaning the endogenous parts of the regressors, we apply OLS again to obtain final coefficient estimates.

The 2SLS method provides a way to obtain consistent and unbiased estimates when dealing with endogeneity problems. First, the 2SLS method addresses the endogeneity problem by using instrumental variables (IVs) to proxy for the endogenous variables and identify the causal effect more accurately. By finding valid instrumental variables, the 2SLS method allows researchers to isolate the exogenous variation in the endogenous variables and obtain consistent estimates of the causal relationship. Moreover, the 2SLS method is particularly useful when dealing with simultaneous equations models where there is endogeneity and simultaneous causation between variables. It allows researchers to break the simultaneity problem and obtain consistent estimates of the underlying relationships. Consequently, the two-stage least squares method is an essential tool in econometrics for addressing endogeneity and obtaining reliable estimates of causal relationships. By using instrumental variables and a two-stage procedure, the 2SLS method provides a robust and effective solution to the challenges posed by endogeneity in regression analysis, making it a valuable tool for empirical research in economics and related fields.

How does the Two-Stage Least Squares (2SLS) method separate endogenous regressors into exogenous and endogenous parts, and how are fitted values used in the second stage to eliminate endogeneity?

In Equation , we can say that e2i and e3i are unobserved error terms that creates the endogeneity problem, respectively for Y2i and Y3i. Notice that we can separate each endogenous regressor into an exogenous and endogenous part. The part that is explained by exogenous regressor X1 and the instruments Zj for j∈{1,..., p} constitute the exogenous part naturally. Moreover, the remaining parts e2i and e3i, then, can be considered as the components that create endogeneity. Accordingly, if we remove the impact of these error terms, we can obtain a variable that has a similar role in the original regression and is also free from endogeneity. In this regard, we obtain the fitted values from the first stage regressions as:

Ŷ2i = ˆ α20 + ˆ α21 X1i + ˆ α22 Z1i + ... + ˆ α2(1+ p) Zpi ,

Ŷ3i = ˆ α30 + ˆ α31 X1i + ˆ α32 Z1i + ... + ˆ α3(1+ p) Zpi .

These fitted values are free from endogeneity problems by our construction, under the assumptions made for the instruments, and the properties of the OLS estimation.

In the second stage, we regress the dependent variable on the exogenous variables and the fitted values of the endogenous regressors from the first stage. For our example, we have the following regression equation:

Y1i = β0 + β1Ŷ2i + β2Ŷ3i + β3X1i + ui ,

where we simply change the endogenous regressors with the fitted values obtained in the first stage regression.

What are the main limitations of the Two-Stage Least Squares (2SLS) method, and how can the weak instrument problem be detected and addressed?

 

Although the 2SLS (or instrumental variable [IV]) method is intuitively appealing, it has some crucial constraints. In this subsection, we list some of these issues. In the first part of the previous section, we discuss the order condition as a necessary requirement for consistency of the 2SLS estimators. However, we need a sufficient condition to complete the theoretical aspects of the 2SLS method. This sufficient condition is known as rank condition, and it requires matrix algebra to check. Since matrix algebra for the OLS is beyond the scope of this course, we skip the details of this condition.

Another crucial issue is the weak instrument problem that emerges when the instruments weakly depend on the endogenous regressors. To see this, consider the first stage regression. If the instruments are too weak to explain the endogenous regressor, then the fitted values obtained in this stage are very close to the observed value of the endogenous regressor. In this case, the instruments are unable to clear out the endogenous components, which causes estimation bias. To avoid this case, Stock and Watson (2003) suggest conducting an F test in the first stage regressions. If the instruments are jointly significant, then we can safely use these instruments in the 2SLS procedure. Otherwise, Baltagi (2007) suggests using alternative instruments or another estimation method such as Limited Information Maximum Likelihood estimation, which is beyond the scope of the current course.

Why is testing for endogeneity important in econometric analysis, and how does it affect causal inference, policy decisions, and model accuracy? 

Testing for endogeneity is essential in econometric and statistical analysis for several reasons. In many economic and social science studies, the main goal is to establish a causal relationship between the explanatory and dependent variables. However, endogeneity can create spurious connections, making it difficult to determine the true direction of causality. By testing for endogeneity, researchers can identify potential problems with causality and employ appropriate methods to establish causal relationships more accurately. Moreover, Endogeneity can lead to biased and inconsistent parameter estimates in regression analysis as we discussed earlier. If endogeneity is not addressed, the results obtained from the analysis may not reflect the true relationships between variables. By testing for endogeneity and using appropriate techniques to address it, researchers can obtain more reliable and valid inference from their models.

Furthermore, in applied economics and policy analysis, the results of econometric models often guide policy recommendations. If endogeneity is not properly addressed, policymakers may implement policies based on flawed or misleading results leading to suboptimal outcomes. By testing for endogeneity, policymakers can make more informed decisions and develop effective policies. In addition, testing for endogeneity can enhance model accuracy. Econometric models are used to explain and predict real-world phenomena. When endogeneity is present, the model’s accuracy and predictive power may be compromised. By testing for endogeneity and employing appropriate techniques to account for it, researchers can improve the model’s accuracy and predictive performance.

Why are the second-stage OLS estimators in 2SLS free from endogeneity bias, and how should residuals be computed to correctly estimate the error variance for standard errors and t-statistics?

 

The OLS estimators from the second stage regression are called 2SLS and they are free from endogeneity bias problem, since the fitted values of the endogenous regressors are not endogenous anymore. Additionally, the standard errors and t-statistics for the regression coefficients can be obtained from the second stage regression after fixing the variance of the error terms. Instead of directly using the residuals from the second stage regression to estimate error variance, one should replace the second stage coefficient estimates in the original regression and compute the residuals. In our example, we compute the residuals as ˆ ui = Y1i – ˆβ 0 ˆβ 1Y2i – ˆβ 2Y3i – ˆβ 3X1i ,instead of ˆ ui = Y1i – ˆβ 0 ˆβ 1Ŷ2i – ˆβ 2Ŷ3i – ˆβ 3X1i .

Why is testing for endogeneity essential in econometric analysis, and how does it contribute to the validity and reliability of research conclusions and policy recommendations?

 

Overall, testing for endogeneity is a critical step in econometric analysis to ensure the validity, reliability, and accuracy of the results. By identifying and addressing endogeneity, researchers can draw more accurate conclusions, establish causal relationships, and make more informed policy recommendations.

How is the presence of endogeneity tested using the F test in regression models, and what does the null hypothesis H₀: δ₁ = δ₂ = 0 imply about the variables Y₂ and Y₃?

 

The endogeneity testing procedure relies on the hypothesis H0 : δ1 = δ2= 0 . Under this null hypothesis, the error term u does not depend on the suspected endogenous components of Y2 and Y3. As a result, if we reject this null hypothesis, we detect the presence of an endogeneity problem. However, if we do not reject the null hypothesis, the original regression model is free from endogeneity, and Y2 and Y3 are exogenous variables. We can conduct this test with a standard F test or LM test framework. For the F test, we need to run restricted and unrestricted regressions derived from Equation . The unrestricted model is given in Equation . We can show the restricted model as Y1i = β0 + β1Y2i + β2Y3i + β3X1i vi . Let SSRu and SSRr be the sum of squared residuals obtained from unrestricted and restricted, respectively. The test statistic is given as

Fend = SSRr SSRu ( )/ m SSRu / n k −1() ,

where m is the number of endogenous variables, k+1 is the number of coefficients in Equation . This test statistic is distributed as F(m, n–k–1) under the null hypothesis, and we reject the null if the test statistic is larger than the upper quantile of the distribution (for given significance level).

What are the two key conditions for instrument validity in 2SLS estimation, and how is the overidentification restriction tested to ensure instruments are exogenous?

One particularly important question about the endogeneity testing 2SLS estimation is about the validity of the instruments, which has two dimensions coming from the assumptions on instruments. The first dimension called “relevance” requires that the instruments are related to the endogenous variables. We have already discussed this issue under the “weak instruments problem”. The second issue called “exogeneity” of the instruments dictates that the instruments should not be related to the error term of the original regression. This issue is investigated under the name “overidentification restriction”. The main idea in testing overidentification restriction is to check whether any instrument is endogenous.

To construct such a test, we utilize 2SLS estimation. As in Wooldridge (2005), in the first step, we obtain the residuals û2SLSi from 2SLS estimation of the original model. Next, we regress û2SLSi on a constant term and all instruments. For instance, in our previous example, we need to estimate the model:

û2SLSi = α0 + α1 Z1i + ... + αpZpi +ri ,

Where ri is the error term, αj for j ∈ {0, ... , p} are unknown coefficients. After estimating this auxiliary regression, we can compute an F statistic to test the null hypothesis H0 : α1 = ... = αp = 0 similar as in the previous subsection. If we reject the null hypothesis, there is at least one instrument that fails the endogeneity conditions. As a result, overidentification restriction fails. However, this test does not tell us which instrument(s) is(are) problematic. One can apply this test on any subset of the instruments to check for the problematic instrument. After the detection of such an instrument, one can omit it from the instrument list.

What does the change in the intercept term and similarity in MPC results between OLS and 2SLS suggest, and how can an endogeneity test help interpret this outcome?

 

Although the results for the mpc is numerically similar to the OLS results, the intercept term switches sign, but it is still insignificant. The reason for this finding may be a weaker link between the Income and the error term. We can check this condition with an endogeneity test described in the previous sections. To perform this test, we require the residuals from the first stage estimation. Whereas we have already run this regression labeled as “OLS.1st” in the R code.

Why does endogeneity bias OLS estimators, and how does the Two-Stage Least Squares (2SLS) method address this issue effectively?

 

The presence of endogeneity poses a significant challenge as it introduces bias in the Ordinary Least Squares (OLS) coefficient estimators. To mitigate this problem, we introduce a two-stage procedure known as “two-stage Least squares” (2SLS). This approach involves modifying the classical OLS model to effectively address endogeneity concerns. The elegance of the 2SLS method lies in its simplicity and the ability to circumvent endogeneity-related biases.

How does the use of the 2SLS method and hypothesis testing help address endogeneity bias and strengthen the credibility and relevance of econometric research?

 

By adopting the 2SLS method and conducting the hypothesis tests, researchers can effectively address endogeneity bias and enhance the credibility of their regression models. The practical application of these methodologies to real-world and simulated data enriches the understanding of the concepts presented, further solidifying their relevance in econometric research.

Bu ünitenin sorularını uygulamada çözŞıklar, doğru cevaplar ve süreli sınav modu AÖF Soru Bankası uygulamasında