Repository navigation
Conversation
| ) | ||
| s = score_cov.weight_matrix(x, x_loo, eps) | ||
| self._cov_config = score_cov.config | ||
| bread = inv(x_loo.T @ x / nobs) |
| self._cov_config = score_cov.config | ||
| bread = inv(x_loo.T @ x / nobs) | ||
| c = bread @ s @ bread.T / nobs | ||
| return (c + c.T) / 2 |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #750 +/- ##
==========================================
+ Coverage 99.57% 99.58% +0.01%
==========================================
Files 108 110 +2
Lines 19634 20327 +693
Branches 1565 1600 +35
==========================================
+ Hits 19550 20243 +693
Misses 31 31
Partials 53 53
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Implements the JIVE estimator of Angrist, Imbens & Krueger (1999),
which eliminates the many-instruments bias of 2SLS by using
leave-one-out first-stage predictions.
Leave-one-out first stage:
X̃_i = [(P_Z X)_i - h_i X_i] / (1 - h_i), h_i = z_i'(Z'Z)^{-1}z_i
Estimator:
β̂_JIVE = (X̃'X)^{-1} X̃'y
Leverage scores computed in O(n·k_instr) without forming the n×n hat
matrix. Covariance is a heteroskedasticity-robust sandwich estimator
using X̃ as the score instrument.
* linearmodels/iv/model.py: IVJIVE class (inherits _IVModelBase) and
_JIVECovariance helper; from_formula support.
* linearmodels/iv/__init__.py: export IVJIVE.
* linearmodels/iv/tests/test_jive.py: 25 tests — output shapes, PSD
covariances, leverage score properties, consistency vs KF, many-IV
bias reduction, from_formula, weighted estimation, multiple endog.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The value of kappa in LIML only depends on the data, and the Anderson-Rubin and Basmann F tests of IVResults use it for any estimator, not just the k-class estimators. Move _estimate_kappa from _IVLSModelBase to _IVModelBase so that the other estimators can provide it. There is no change in behavior.
Review of #702. The estimates match a literal leave-one-out implementation and the covariance is a correct sandwich, but several things did not work. cov_type and cov_config were accepted by fit and ignored. For one model, cov_type="unadjusted", "kernel" and "clustered" all returned the robust standard error 0.096795 and reported cov_type "robust", so fit(cov_type="clustered", clusters=...) silently gave heteroskedasticity robust results. JIVE is an IV estimator with the instruments X~, so its covariance is (X~'X)^{-1} S (X'X~)^{-1}, and S is now computed by the estimators of the covariance of the moment conditions that the GMM models use. unadjusted, robust, kernel (Bartlett, Parzen and quadratic spectral) and clustered are all supported, along with debiased, with the same options as the other models. The homoskedastic estimator uses the uncentered residual variance, as 2SLS and Stata do, and not the centered variance that the weight matrix classes use, which differs when the residuals of a weighted model do not have mean zero. This convention is checked against Stata for weighted 2SLS. Two computations were not accurate when the instruments are correlated or the regressors have different scales. The leverages and the first-stage predictions were computed from inv(Z'Z), which squares the condition number of Z. For instruments with condition numbers of 2e5, 2e6 and 2e8 the parameters had relative errors of about 1e-7, 2e-6 and 4.5e-3 compared with explicit leave-one-out refits. They are now computed from a QR decomposition of Z, with errors of about 1e-10 or less in the same cases. The covariance first reused the GMM covariance with an identity weight matrix, which forms the inverse of (X'X~)(X~'X) and squares the condition number of X'X~. The standard errors of the housing example (Stata's hsng2) were wrong by a factor of about 4000, and those of the Card example were NaN, although the parameters were right. The bread (X~'X)^{-1} is now formed directly. anderson_rubin and basmann_f raised an AssertionError because the results did not have the LIML kappa, which only depends on the data and is now provided. If it cannot be computed because of a LinAlgError, which cannot happen for valid data, the two tests are NaN and the model is still estimated. kappa was reported as 1, as for 2SLS, and is now None since JIVE is not a k-class estimator. IVResults was annotated to only accept the k-class models and now accepts any IV model, which fixes the type error in the constructor call. The error for a leverage of 1, which happens if a dummy instrument is 1 for a single observation, now says what is wrong and how many observations are affected. The docstring said that JIVE eliminates the bias of 2SLS, and now describes the estimator accurately: it reduces the many instrument bias but can have a larger variance, and the usual standard errors can understate the uncertainty when there are many instruments. The example in from_formula failed because brthord has missing values. The tests were in linearmodels/iv/tests, which does not exist on main and is not collected by the Linux CI jobs, which is why Codecov reported 30% coverage of the patch. They are in linearmodels/tests/iv and have been rewritten. Most of them only checked shapes and signs, and one could not fail (the JIVE error was less than the 2SLS error plus 0.3). The new tests compare with - values from R in results/jive-reference.R, which computes JIVE with an explicit leave-one-out first stage for 42 combinations of data (simulated data, Stata's hsng2 housing data and Card's 3010 observations), model (one to three endogenous variables, no constant, several exogenous variables, 24 instruments with a high leverage, an exactly identified model), covariance estimator (unadjusted, robust, clustered and kernel with the Bartlett, Parzen and quadratic spectral kernels), weights and debiased, for the parameters, standard errors, R-squared and model F statistic. Before it writes the values it checks that its covariance code reproduces 22 Stata 2SLS results in the repository (the four covariance estimators, with and without weights and small sample adjustments, for the simulated and the housing data), and that the first stage equals the leverage shortcut. There are no results from Stata or R for JIVE itself in the repository, so these values do not come from an independent implementation of JIVE, - an explicit implementation of the definition in numpy for 30 random models, which vary the number of regressors, endogenous variables and instruments, the constant, the weights and the clusters, for every covariance estimator, - the fact that JIVE without endogenous variables is OLS, so every covariance estimator has to equal IV2SLS, with and without weights, - the properties of the estimator: it does not change with an invertible transformation of the instruments, the order of the observations or the regressors, the scale of the variables or the weights, and it shifts the constant when the dependent variable is shifted, - error handling, missing values, pandas input, the formula interface, the results methods, compare and the first stage. Of the 145 tests, 87 fail on the original implementation. The tests give 100% line coverage of the JIVE code, and a test fails for 19 of 20 single edits of it that were tried. The other edit only changes round-off.
The other tests show that the estimates and the covariance estimators are what they are defined to be. These show that they behave as they should statistically, using a fixed seed. With 3 strong instruments and 300 observations, the coverage of nominal 95% confidence intervals is 0.935 (unadjusted) and 0.945 (robust) with homoskedastic errors. With errors whose variance depends on the instruments the robust intervals cover 0.953 of the time and the unadjusted 0.865. With 30 clusters and errors and instruments that have a cluster component, the clustered intervals cover 0.918 and the robust 0.700. A covariance estimator is accurate where it is appropriate and is not where it is not, so the option has to be used. Also check the many instrument bias with 30 weak instruments: the median error of 2SLS is 0.165 and that of JIVE is -0.013, and JIVE has a larger spread. With 2 or 10 strong instruments JIVE is close to 2SLS. On the original implementation, which ignored the covariance estimator, the tests with heteroskedastic and cluster correlated errors fail.
Add IVJIVE to the module reference, the list of IV estimators in the introduction and the home page, and the type aliases used when building the documentation.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements the JIVE estimator of Angrist, Imbens & Krueger (1999),
which eliminates the many-instruments bias of 2SLS by using
leave-one-out first-stage predictions.
Leave-one-out first stage:
X̃_i = [(P_Z X)_i - h_i X_i] / (1 - h_i), h_i = z_i'(Z'Z)^{-1}z_i
Estimator:
β̂_JIVE = (X̃'X)^{-1} X̃'y
Leverage scores computed in O(n·k_instr) without forming the n×n hat
matrix. Covariance is a heteroskedasticity-robust sandwich estimator
using X̃ as the score instrument.
_JIVECovariance helper; from_formula support.
covariances, leverage score properties, consistency vs KF, many-IV
bias reduction, from_formula, weighted estimation, multiple endog.
closes #702