When a Regression Looks Excellent — and Still May Be Wrong
What money, nominal GDP, gold, and the U.S. economy from 1990 to 2011 can teach us about the difficult art of specifying an econometric model.
A regression can look spectacular and still be asking the wrong question. It can explain almost all of the observed variation in a variable, produce highly significant coefficients, draw a fitted line that hugs the data remarkably closely—and nevertheless contain a structural flaw. That tension is the real subject of Gómez Julián’s study of hypothesis testing and the optimal specification of a regression model. The empirical setting is monetary: the paper examines the U.S. money supply between 1990 and 2011 in relation to nominal GDP and the international price of gold. But the deeper lesson concerns econometric reasoning itself. How do we move from a model that merely fits to one whose specification we have serious statistical reasons to trust?
That distinction matters far beyond monetary economics. Political scientists estimate the effects of institutions on growth, epidemiologists study relationships between exposures and outcomes, sociologists model inequality, and economists routinely estimate relations among prices, production, investment, employment, money, and financial variables. In every case, a regression equation compresses a complex process into a manageable mathematical form. The danger is obvious: a convenient equation can conceal dynamics that the equation itself does not represent.
The paper treats hypothesis tests not as decorative statistics added after a regression has already been estimated, but as instruments for interrogating the model itself. A high coefficient of determination is not the end of the investigation. It is the beginning of a sequence of questions: Are the residuals independent? Is their variance stable? Are the parameters stable? Is the assumed functional form adequate? Are the time series stationary? Could an apparently excellent fit be spurious?
- The economic hypothesis behind the regression
- Why the United States, and why 1990–2011?
- The deceptively impressive first OLS result
- Why residual diagnostics change the story
- Ramsey RESET and the problem of functional form
- ARCH, volatility, and dependence hidden in squared errors
- Unit roots, spurious regression, and cointegration
- What the investigation ultimately establishes—and what remains open
1. The economic model under examination
The empirical exercise does not begin from a purely data-driven search for correlations. It begins from an economic proposition inherited from an earlier 2016 study. In the formulation used here, the quantity of money in circulation is treated as a function of the prices of commodities in the economy and of the international price of gold, under simplifying conditions concerning the velocity of circulation and the relative stability of the variables during the period under examination.
For the U.S. application, commodity prices are represented empirically by nominal GDP—called Current GDP in the dataset—while the second explanatory variable is the international gold price. Money in circulation, denoted by M, becomes the dependent variable.
Current GDP → nominal output, used in the paper as an empirical expression of commodity prices
Gold → international price of gold
This direction of explanation is important. The paper is not merely asking whether money, output, and gold happen to move together. Its theoretical framework assigns a specific economic meaning to their relationship. The statistical model is therefore supposed to represent a prior conception of how the monetary process works.
That distinction between statistical relationship and economic meaning later becomes one of the paper’s recurring themes. A regression can identify mathematical dependence among measured variables, but the interpretation of that dependence does not emerge automatically from the regression coefficients. It is supplied by the theoretical framework within which the statistical model has been constructed.
This is especially important for readers outside economics. Imagine that two political indicators move almost perfectly together over twenty years. The regression does not know whether one causes the other, whether both respond to a third process, whether their common trend creates a misleading correlation, or whether the relation changes after an institutional break. Those are questions about the structure of the phenomenon—and therefore about model specification.
2. Why study the United States from 1990 to 2011?
The paper deliberately uses the same broad historical interval as the earlier study while moving the empirical setting from El Salvador to the United States. The choice is not presented as a neutral convenience. Gómez Julián argues that the social and economic structure of the United States makes the period particularly useful because it contains episodes that could expose weaknesses in a simple regression specification.
The interval includes the long expansion associated with the post-Cold-War era, the rise and collapse of the dot-com boom, the financial crisis that erupted in 2008, and the first years of its aftermath. In particular, the observations for 2009, 2010, and 2011 are treated as analytically valuable because they incorporate the macroeconomic disturbance following the financial crisis.
A model may behave quite well during relatively tranquil periods and reveal its weaknesses only when the system is disturbed. Including 2009–2011 therefore acts, in effect, as a small stress test: does the relationship estimated from the preceding years survive when the economy moves through an unusually turbulent episode?
The paper repeatedly insists that economic time series are records of historical processes rather than interchangeable observations floating in abstract statistical space. The year 2009 is not simply “observation number 20.” It follows 2008. The history that produced it matters. This idea becomes crucial once autocorrelation, structural stability, unit roots, and cointegration enter the discussion.
This historical perspective also explains why the paper resists treating every statistical problem as something that can simply be removed through data cleaning. If an economic series contains persistence, structural change, or stochastic trends, those properties may reflect features of the economic process itself. A model should therefore try to understand them before mechanically eliminating them.
3. The first regression looks almost too good
The starting empirical model is a conventional multiple linear regression estimated by ordinary least squares. In simplified notation,
For a reader encountering econometrics for the first time, the idea is straightforward. The regression chooses the values of the coefficients β that make the predicted money supply as close as possible—in the least-squares sense—to the observed money supply. The residual ε for each year is the part that the equation fails to explain.
And at first sight, the results are extremely impressive.
| Term / statistic | Estimate | Reported p-value |
|---|---|---|
| Intercept | −2.130 × 10¹² | 7.67 × 10⁻⁶ |
| Current GDP | 0.8392 | 6.29 × 10⁻¹⁴ |
| Gold | 2.044 × 10⁹ | 8.10 × 10⁻⁵ |
| R² | 0.9834 | — |
| Adjusted R² | 0.9817 | — |
| Overall F statistic | 563.3 | < 2.2 × 10⁻¹⁶ |
An R² of roughly 0.983 means that, within this fitted linear specification, approximately 98.3 percent of the observed variation in the dependent variable is accounted for by the regressors. Both slope coefficients also appear extremely significant under the conventional t tests. If one judged the regression only by these headline numbers, the temptation would be to declare victory.
Why? Because statistical significance is conditional. A t statistic has the interpretation we normally give it only if the probabilistic structure assumed by the model is sufficiently appropriate. An R² can be enormous because two variables genuinely share an economically meaningful relation—but it can also become enormous when several non-stationary time series drift through time in broadly similar directions.
Likewise, a visually excellent fit does not prove that the functional form is correct. A straight line may approximate a curved relationship surprisingly well over a short sample. Nor does it establish that the unexplained component behaves as required by the classical regression framework.
The residuals therefore become witnesses. If they contain systematic temporal structure, changing variance, nonlinear information, or persistent patterns, then the model has left something important behind.
4. What does “well specified” actually mean?
This is the conceptual bridge between the economic question and the statistical machinery that occupies the rest of the study. A regression model is not adequately specified merely because its fitted values resemble the observations. Its mathematical form must also be compatible with important features of the data-generating process.
Is the assumed linear relationship sufficient, or do squared, cubic, logarithmic, or otherwise nonlinear terms contain systematic information that the original model misses?
Are regression errors approximately independent and homoscedastic, or does the past help predict their magnitude or variance?
Are the variables stationary, or do persistent stochastic trends make conventional regression inference potentially misleading?
The paper therefore deploys a broad diagnostic battery: the Ramsey RESET specification test, ARCH tests, White’s heteroskedasticity test, variance inflation factors for collinearity, residual autocorrelation tests, residual normality tests, CUSUM and CUSUM-of-squares procedures for parameter stability, and—later in the investigation—unit-root and cointegration techniques.
The philosophical point behind that battery is worth emphasizing. No single diagnostic test is treated as an oracle. Different tests examine different potential failures, and they can even produce apparently conflicting indications. Model evaluation therefore becomes cumulative: one result raises a suspicion, another narrows the possibilities, a modification removes one defect but perhaps exposes another, and the investigation gradually moves toward a more adequate representation.
Instead of asking only, “Is this coefficient statistically significant?”, the paper encourages a broader sequence of questions: What assumption is being tested? What would rejection imply about the model? If a defect is detected, what new specification should be considered? And after modifying the model, what new diagnostics become necessary?
This transforms hypothesis testing from a ritual of stars and p-values into a method of model criticism. In that sense, the tests are not merely judging individual coefficients. They are helping the researcher decide whether the mathematical representation itself should survive in its present form.
5. The first crack: the errors remember the past
The paper’s subsequent diagnostics show why the impressive R² cannot be taken at face value. In the empirical discussion, a significant problem of residual autocorrelation emerges. In plain language, the unexplained error associated with one year is not behaving as though it were disconnected from errors in surrounding years.
That matters because ordinary least squares is at its cleanest when the error process does not carry precisely this kind of systematic temporal dependence. If residuals are autocorrelated, the equation is effectively saying: “I have explained the systematic part of the process,” while the residuals answer: “No—you have left some systematic time structure inside me.”
In a genuinely historical system this is hardly an exotic possibility. Production in one year depends partly on previous production; investment decisions inherit previous expectations and balance sheets; monetary magnitudes accumulate institutional and financial history. Time links observations together.
The crucial econometric question is therefore not whether such dependence can exist—it obviously can—but whether the chosen model has represented it adequately. If it has not, temporal dependence migrates into the residuals.
The study responds by introducing an additional term constructed from changes in nominal GDP. This modification improves the residual autocorrelation diagnostics substantially. Yet the author explicitly acknowledges an important limitation: the choice is empirically motivated rather than derived from an independent theoretical justification, and the paper does not claim to have searched exhaustively over the optimal number of lags, compared every plausible dynamic specification, or estimated the full family of VAR-type alternatives.
The modified specification should therefore be understood as a provisional diagnostic solution, not as the final uniquely correct model. Its purpose is methodological: to show how one statistical test can reveal a problem, motivate a modification, and force the researcher to test the new model again.
And that is exactly what happens. Fixing one problem does not end the investigation. Once the residual autocorrelation appears to improve, other diagnostics—especially the behavior of squared errors and the Ramsey specification tests—raise a more difficult possibility:
That is where the paper’s two central diagnostic tools enter the story: the ARCH test, which asks whether the variance of the errors carries temporal structure, and the Ramsey RESET test, which asks whether nonlinear combinations contain explanatory information that the original regression has failed to capture.
Those tests are the subject of the next stage.
6. A model can pass one test and fail another
Once the residuals become the object of investigation, an important principle emerges: there is no single thing called “the regression assumptions.” Different diagnostic tests ask different questions about different aspects of the model. Passing one of them therefore does not imply passing the others.
This distinction becomes unusually clear in the paper’s U.S. regression. The residuals look reasonably well behaved in several respects. The normality test does not provide sufficient evidence to reject Gaussian residuals. The variance inflation factors are modest, so severe multicollinearity does not appear to be driving the estimates. White’s test likewise provides no strong evidence of ordinary heteroskedasticity.
And yet a residual-autocorrelation test detects temporal dependence.
| Diagnostic | Null hypothesis | Reported result | Reading at 5% |
|---|---|---|---|
| White | No heteroskedasticity | p ≈ 0.232 | Do not reject |
| Breusch–Godfrey LM, order 1 | No residual autocorrelation | p ≈ 0.023 | Reject |
| ARCH, order 1 | No ARCH effect | p ≈ 0.991 | Do not reject |
| Residual normality | Residuals are normally distributed | p ≈ 0.105 | Do not reject |
There is nothing logically contradictory about that combination. The errors may have an approximately constant variance and an approximately normal marginal distribution while still being serially dependent. In other words, the size and shape of the error distribution can look acceptable even though errors occurring close together in time are related.
White asks whether the variance changes systematically. Breusch–Godfrey asks whether errors are serially correlated. ARCH asks whether the conditional variance itself displays temporal dependence. A normality test asks about the shape of the residual distribution. These are related questions, but they are not interchangeable.
This is one reason diagnostic work cannot be reduced to a checklist in which every test is interpreted as another version of “good model” versus “bad model.” Each test illuminates a different part of the statistical structure.
7. Ramsey RESET: asking whether the equation has the wrong shape
Residual autocorrelation raises the possibility that something systematic has escaped the regression. But it does not by itself tell us exactly what. One possibility is temporal dynamics. Another is an omitted variable. A third is that the assumed mathematical form of the relationship is itself inadequate.
The Ramsey Regression Specification Error Test—RESET—is designed as a broad diagnostic for precisely this sort of misspecification.
The intuition is easier than the name suggests. Begin with a fitted regression:
From that equation we obtain fitted values, M̂. RESET then asks whether nonlinear transformations—such as M̂² or M̂³, or related nonlinear transformations of the regressors—contain explanatory information that the original equation failed to capture.
REJECTION → the original specification has left systematic nonlinear information unexplained.
This does not mean that RESET discovers the uniquely correct alternative equation. A rejection does not tell us, by itself, “use a quadratic model” or “add precisely this missing variable.” RESET is better understood as a smoke detector: it can tell us that something about the specification deserves attention without necessarily identifying the source of the smoke.
The results are revealing—but not uniform
The paper runs several versions of RESET in R. This matters because the conclusion depends somewhat on how the additional nonlinear information is constructed.
| RESET variant in R | Power | p-value | At α = 0.05 |
|---|---|---|---|
| Fitted values | 2 | 0.06463 | Do not reject |
| Regressors | 2 | 0.1216 | Do not reject |
| Principal-component form | 2 | 0.04188 | Reject |
| Fitted values | 3 | 0.01402 | Reject |
| Regressors | 3 | 0.1177 | Do not reject |
| Principal-component form | 3 | 0.05739 | Do not reject |
That pattern deserves a careful interpretation. There is evidence of specification trouble, but not a unanimous verdict across every implementation. Two versions reject adequate specification at the conventional 5 percent threshold, while several others do not. Some of the non-rejections are nevertheless close to conventional thresholds.
The Gretl analysis tells a similar story. A specification test using squares and cubes jointly reports a p-value of roughly 0.037, providing evidence against the null of adequate specification. A logarithmic nonlinearity test also rejects linearity at approximately the 5 percent level, while other nonlinear formulations do not.
This distinction is fundamental. If several plausible transformations reveal predictive structure that the original regression leaves unexplained, the investigator no longer has good reason to regard the initial linear form as obviously sufficient.
And notice how far we have moved from the seductive R² of 0.983. The regression still fits the observed money series extremely closely. Nothing has happened to erase that fit. What has changed is our interpretation of what that fit proves.
8. ARCH: when the errors are calm today because they were calm yesterday
The second test emphasized by the paper is the ARCH test associated with Robert Engle’s work on autoregressive conditional heteroskedasticity.
Its central insight is that ordinary autocorrelation and volatility dependence are different phenomena. Suppose the signs of successive errors are difficult to predict: a positive error may be followed by a negative one, then another positive one. Standard autocorrelation could therefore be weak.
But suppose large errors tend to arrive near other large errors, while small errors tend to cluster around small errors. The signs may appear irregular while the magnitudes clearly remember the past.
If ε² at time t is systematically related to earlier ε² values, the variance contains temporal structure.
This is the core of the ARCH idea. The paper introduces it partly through the language of heavy-tailed distributions, but the technically decisive point is more specific: an ARCH test is a test for dynamics in conditional variance, not simply a generic test for heavy tails. Heavy tails and volatility clustering can be related empirically, but they are not the same property.
In the initial 1990–2011 regression, the reported first-order ARCH p-value is approximately 0.991. At the conventional 5 percent level, there is therefore no evidence for rejecting the null hypothesis of no first-order ARCH effect.
This produces an instructive combination:
The ordinary residuals show evidence of first-order autocorrelation.
White’s test does not provide evidence of general heteroskedasticity.
The squared residuals do not show evidence of first-order ARCH dependence.
So the temporal problem detected at this stage concerns the residuals themselves rather than a clear clustering of their squared magnitudes. That diagnostic distinction becomes important when the model is modified.
9. CUSUM versus CUSUM of squares: two views of stability
Another question is whether the regression coefficients remain stable throughout the sample. If the relationship among money, nominal GDP, and gold changes substantially over time, a single set of coefficients estimated for 1990–2011 may be concealing several different regimes.
The paper approaches this through CUSUM and CUSUM of squares. Both procedures examine cumulative residual behavior, but they do not have identical statistical properties or identical sensitivity to different forms of instability.
And here the results appear to disagree.
The ordinary CUSUM analysis indicates parameter instability beginning around the transition from 2006 to 2007. The CUSUM-of-squares graph, by contrast, remains within its confidence boundaries and therefore points toward parameter stability over the sample.
The paper does not simply average these conclusions. It turns to methodological literature comparing the tests through Monte Carlo simulations and gives greater analytical weight to CUSUM of squares, which the cited research describes as more reliable under several forms of misspecification and more resistant in settings involving non-predetermined regressors, cointegration, or weak serial correlation.
This is another useful lesson in statistical practice. Two tests can disagree without either calculation being mechanically “wrong.” Tests have different power against different alternatives. They react differently to the location, form, and magnitude of structural change. Their assumptions also matter.
Still, the ordinary CUSUM result remains economically suggestive. The timing is difficult to ignore: instability appears near the years immediately preceding the 2008 crisis. The paper treats this not as proof that the crisis has been statistically detected in advance, but as an observation deserving further investigation.
That caution is important. A diagnostic graph can reveal an unusual pattern. Explaining the historical cause of that pattern requires more than the graph itself.
10. Repairing the model creates a new model—and therefore new problems
Having identified residual autocorrelation, the paper experiments with an additional GDP-related term. This is where the logic of hypothesis testing becomes especially visible: diagnosis leads to respecification, and respecification immediately requires another round of diagnosis.
The paper describes the added variable in places as involving the second differences of nominal GDP, while also identifying it as Current GDP (t−2) and naming the Gretl variable CurrentGDPt2. These are not mathematically identical operations: a two-period lag, GDPt−2, is different from a second difference, Δ²GDPt. For that reason, the safest reading is to preserve the paper’s reported variable and results while recognizing that its verbal description and notation are not fully aligned.
The author is explicit about another limitation: there is no independent theoretical argument offered for choosing this particular modification. It is adopted because the empirical results improve relative to an alternative tried in the study. Nor does the investigation claim to have performed an exhaustive search over optimal lag structures, Monte Carlo lag selection, VAR models, or every competing dynamic specification.
That restraint matters. The added term is a provisional econometric repair, not a declaration that the final structural model has been discovered.
The modified regression nevertheless produces another strikingly high fit:
| Statistic | Modified model | Interpretive point |
|---|---|---|
| R² | ≈ 0.9904 | Fit improves still further |
| Adjusted R² | ≈ 0.9888 | Remains extremely high |
| Durbin–Watson | ≈ 0.879 | Still far below the benchmark value near 2 |
| White test | p ≈ 0.054 | Borderline, but not rejected at 5% |
The Breusch–Godfrey sequence becomes especially revealing. For lag orders two through six, the reported p-values rise above 0.05. Yet the first-order test remains significant, with a p-value of roughly 0.020.
A careful reading therefore requires a qualification. The paper subsequently describes the modification as solving the autocorrelation problem across the examined lag structure, but the displayed first-order LM result still rejects the null of no autocorrelation at 5 percent. The safest conclusion from the numerical output is consequently that the modification substantially changes and weakens the detected autocorrelation pattern, but does not unambiguously eliminate first-order residual autocorrelation.
A researcher should privilege the reported statistical output over the desire for a tidy narrative. If one diagnostic remains problematic, the model remains diagnostically open—even when most of the other numbers improve.
Something else happens as well. In the modified model, RESET becomes much more forceful. Gretl reports p-values of approximately 0.0036 for a squares-only specification test and 0.0014 for tests involving cubes and the joint squares-and-cubes formulation.
In other words, adding temporal structure may alleviate one problem while making another easier to see. The improved fit does not rescue the linear specification. On the contrary, the nonlinear diagnostics now give stronger evidence that the equation is still leaving systematic structure unexplained.
The ARCH diagnostics also become more interesting. The first-order ARCH test in the modified specification reports a p-value of approximately 0.069. That is not significant at 5 percent, but it is below 10 percent. The higher-order tests reported for orders two and three are much less suggestive of an ARCH effect.
This is why the paper regards the altered ARCH behavior, together with the persistently low Durbin–Watson statistic, as a warning rather than as a clean resolution of the model’s problems.
11. Hypothesis testing as an iterative method of criticism
At this point the investigation has reached a result more important than any individual p-value.
The original model looked exceptionally strong by conventional headline measures. Diagnostics then revealed residual dependence. A modified specification changed that dependence but did not make every warning disappear. RESET continued—and in some versions became even more emphatic—in suggesting that the linear form was incomplete.
The workflow therefore looks less like a one-shot estimation problem and more like a sequence:
This is the methodological heart of the paper. Hypothesis tests are valuable not because a p-value mechanically declares what reality is, but because carefully chosen tests can reveal where a particular mathematical representation ceases to behave as it should.
And the next problem is more fundamental still.
Everything discussed so far presumes that conventional regression inference is meaningful for these time series. But GDP, the money supply, and gold prices evolve historically and often trend strongly through time. If those series are non-stationary, then an extraordinarily high R² and spectacular t statistics may arise even when a conventional regression is statistically deceptive.
That takes the investigation from regression diagnostics into one of the deepest problems in applied time-series econometrics: unit roots, spurious regression, and cointegration.
And there the apparently simple monetary regression becomes a much more interesting object.
12. The deeper problem: what if the series themselves do not stand still?
The specification diagnostics have already complicated the apparently excellent OLS regression. But the paper now introduces a problem that reaches deeper than residual autocorrelation or an omitted nonlinear term: non-stationarity.
To understand why this matters, consider what a conventional regression quietly hopes for. If we observe a variable over time, we would like its probabilistic behavior to be sufficiently stable that observations from different years remain meaningfully comparable. In the weak form of stationarity emphasized in the paper, the mean and variance remain constant through time, while the covariance between two observations depends on the distance separating them rather than on the calendar date at which they occur.
Economic series frequently make this assumption difficult. Nominal GDP grows over decades. Monetary aggregates expand. Asset prices can wander for long periods. A shock today can alter not merely the next observation but the level around which future observations evolve.
This is the setting in which the concept of a unit root becomes important.
Take the simplest autoregressive process:
If |ρ| is comfortably below one, shocks tend to fade. But when ρ = 1, the equation becomes
In the lag-operator notation discussed in the paper, this is
The practical consequence is crucial. A unit-root process carries a stochastic trend. Its path depends cumulatively on past shocks. A disturbance is not merely temporary noise around an unchanged trajectory; it becomes part of the history from which the next observation begins.
This is one reason the paper devotes considerable attention to the historical character of economic time series. Its broader argument is that persistence should not automatically be treated as dirty data to be removed before “the real” analysis begins. Persistent movement may itself contain information about the underlying economic process.
13. A macroeconomic dispute becomes a lesson in time-series econometrics
At this point the paper makes a long historical detour through the debate surrounding the 2008–2009 recession, especially the disagreement involving Gregory Mankiw, Paul Krugman, Brad DeLong, and others over whether output losses associated with recessions should be expected to disappear through unusually rapid subsequent growth.
The technical issue beneath the polemics is highly relevant to the paper. Imagine that output collapses during a recession. There are two very different statistical stories one might tell.
The recession pushes output temporarily below an underlying path. Once the shock disappears, unusually strong growth can return output toward that old trajectory.
The recession changes the level from which the economy subsequently evolves. Growth may resume, but part of the lost output is not automatically recovered.
Deciding between these stories is not semantic. It changes both forecasting and the interpretation of macroeconomic shocks.
The paper sides strongly with the view that economists should not presume automatic return to an old trend. It invokes both the unit-root literature and later evidence that, on average across countries, output following crises may remain permanently below its previous trend path.
For the monetary regression, however, the importance of this debate is methodological rather than merely historical. If macroeconomic series can inherit shocks permanently, then fitting one trending level series against another requires considerable caution.
And that leads to a famous statistical trap.
14. The beautiful regression that may mean nothing
Suppose two completely unrelated variables both rise persistently with time. A regression of one on the other can produce a high R², apparently significant t statistics, and an elegant fitted line—even though there is no meaningful relationship connecting them.
This is the classic problem of spurious regression.
When unrelated non-stationary variables share persistent trends, ordinary regression machinery may mistake common trending behavior for substantive explanatory power. The usual t and F inference can then become unreliable.
This possibility suddenly makes the original result—R² ≈ 0.983—look rather different. Earlier it appeared to be the model’s greatest achievement. Now its sheer magnitude, combined with a relatively low Durbin–Watson statistic, becomes part of the reason for investigating whether the fit might be deceptive.
The paper invokes the traditional rule of thumb associated with this problem: when an estimated time-series regression displays a very high R² together with a much lower Durbin–Watson statistic, suspicion of spurious regression is warranted. The rule is diagnostic rather than a proof, so the analysis cannot stop there.
This is perhaps the most important reversal in the investigation. The same statistic can look reassuring at the beginning of an analysis and suspicious later, once we know more about the temporal properties of the data.
15. Dickey–Fuller: testing whether shocks die or accumulate
To investigate the problem, the paper turns to the Dickey–Fuller family of unit-root tests, especially the Augmented Dickey–Fuller test (ADF) and the ADF–GLS / DF–GLS variant.
A convenient way to see the logic is to rewrite the autoregressive process in changes:
If γ is consistent with zero, the previous level does not generate the type of mean reversion required to reject the unit-root hypothesis.
The lagged differences are included to absorb additional serial dependence.
The test can be estimated under different deterministic specifications: without a constant or deterministic trend, with a constant, or with both constant and trend. That choice is not cosmetic. A series that looks non-stationary under one specification may behave differently once a deterministic trend is explicitly modeled.
The paper also emphasizes lag selection and employs information criteria—including variants of AIC and BIC—because too few lags can leave residual autocorrelation behind, while too many can consume substantial degrees of freedom in an already tiny sample.
In standard hypothesis-testing language, a p-value is not the probability that the null hypothesis is true, nor the probability that rejecting it would constitute a Type I error. It is the probability, assuming the null hypothesis and the test model are true, of obtaining a test statistic at least as extreme as the one observed. Throughout this guided reading, p-values are interpreted in that conventional frequentist sense.
What happens to the variables?
The resulting picture is not clean enough to permit a one-line declaration that every series behaves identically.
For the money supply, the paper reports unit-root evidence in levels under the specifications it examines. The behavior of nominal GDP and the GDP-related term differs across the tests and specifications, while the gold-price series also changes classification depending on whether a constant and deterministic trend are included.
The important point for the argument is therefore not that “all four variables possess exactly the same order of integration.” The important point is that non-stationarity is sufficiently pervasive to make the validity of a levels regression a serious econometric question.
The paper then examines first differences.
If a series is I(1), differencing it once should ordinarily produce an I(0), or stationary, process. But the paper’s results are again mixed. Some first-difference specifications become stationary, while others continue to exhibit unit-root behavior depending on the deterministic terms included.
| Series discussed | First-difference result reported by the paper |
|---|---|
| Nominal GDP | Results depend on the specification; the paper finds stationary behavior in relevant variants but not uniformly in every formulation. |
| Money supply | Unit-root behavior without deterministic terms; stationarity appears when constant and/or trend specifications are considered. |
| Gold price | Unit root in the specifications without trend; stationarity is obtained when drift and deterministic trend are incorporated. |
This prevents differencing from serving as a simple universal cure in the paper’s exercise. The investigation consequently returns to the levels relation and asks a more interesting question.
16. Two non-stationary series can still belong together
Non-stationarity does not automatically imply that a regression is spurious.
Suppose two variables wander over time, each carrying a stochastic trend. If they wander independently, regressing one on the other may indeed create a meaningless statistical relationship. But suppose their long-run movements are linked so that a particular linear combination of them is stationary.
Then the two series are said to be cointegrated.
A useful image is two hikers walking along an irregular mountain path while connected by a rope. Neither remains in a fixed location; both may wander far from where they began. But they cannot drift arbitrarily far apart. Their individual coordinates are non-stationary, while the distance imposed by the rope remains constrained.
This distinction rescues an important class of long-run economic relationships from the spurious-regression problem. A levels regression among non-stationary variables can be meaningful if those variables share a genuine cointegrating relation.
The paper therefore moves from asking about the variables individually to asking about their joint stochastic structure.
17. Engle–Granger: look at what remains after the long-run relation
The first route explored is the Engle–Granger logic. Estimate a candidate long-run relation and obtain its residuals:
The intuition is powerful. If money, GDP, and gold merely share unrelated trends, their regression error can itself wander indefinitely. If instead their combination represents a stable long-run relationship, deviations from that relationship should not drift without bound.
After fitting the proposed long-run relationship, do the residuals behave like another non-stationary process, or do they return toward a statistically stable range?
The paper reports an awkward result. Some of the individual series do not satisfy the unit-root conditions in the same way, and the residual-based evidence does not yield the completely clean configuration the author was seeking. A reported p-value of approximately 0.099 becomes part of this inconclusive picture at the 5 percent significance level.
Rather than treating that ambiguity as the end of the exercise, the paper moves to a multivariate method: the Johansen approach.
The paper describes one “cointegrating regression” as a regression that explicitly includes time and calls time the variable that cointegrates the series. That is not the standard definition of Engle–Granger cointegration. In conventional econometrics, cointegration concerns whether a linear combination of integrated variables—typically examined through the residuals of their levels relationship—is stationary. A deterministic time trend may be included in a specification, but time itself is not thereby “the cointegrating variable.” The empirical sequence reported in the paper is retained here, while this distinction is made explicit to avoid conflating the two concepts.
18. Johansen: stop examining one equation at a time
Engle–Granger is naturally built around a single candidate cointegrating relation. Johansen’s method approaches the problem as a system. Instead of choosing one variable as the sole dependent variable and relegating the others to the right-hand side, the variables are embedded in a vector autoregressive structure.
The paper applies Johansen’s procedure to a four-variable system containing the money supply, current GDP, the additional GDP-related term, and gold.
The central object can be represented through a vector error-correction form:
The rank of the matrix Π contains the essential information.
There is no stationary long-run linear combination detected among the integrated variables.
One or more independent long-run relations constrain the joint movement of the variables.
In the standard framework, the system does not have the structure of a collection of I(1) variables requiring cointegration analysis in the usual sense.
Johansen supplies two familiar testing strategies: the trace statistic and the maximum-eigenvalue statistic. They approach the rank question slightly differently, but both ask how many independent long-run relations are supported by the system.
For the purpose of this paper’s argument, the most important null hypothesis is
The paper tests more than one deterministic specification
Johansen results are reported under alternative assumptions about deterministic components: a system treated essentially as a pure random walk, another incorporating drift, and another allowing drift together with a deterministic trend.
Across these cases, the paper emphasizes the same finding: the null hypothesis of zero cointegrating relations is rejected.
| Johansen setting discussed | Key hypothesis | Paper’s interpretation |
|---|---|---|
| No constant / pure random-walk setting | r = 0 | Rejected |
| Specification with drift | r = 0 | Rejected |
| Drift + deterministic trend | r = 0 | Rejected |
The displayed output for the first case is particularly emphatic: the trace statistic for r = 0 is reported with a p-value effectively equal to zero. The other deterministic specifications likewise reject the zero-rank hypothesis.
The paper therefore takes the Johansen results as evidence that the strong comovement observed in the regression is not merely an accidental correlation among unrelated trending series.
19. So was the original regression spurious?
The paper’s final answer to this particular question is no.
That answer is not obtained from the original R², from the significance of the GDP and gold coefficients, or from the visually close correspondence between predicted and observed money. All of those were available near the beginning of the investigation.
The conclusion arrives only after a much longer chain:
On the paper’s interpretation of the Johansen evidence, rejecting r = 0 means that at least one long-run relation exists within the system. This removes the specific fear that the original high fit is nothing more than the classic nonsense regression created by unrelated stochastic trends.
But—and this is essential—“not spurious” does not mean “correctly specified.”
The two questions are logically distinct.
Johansen evidence is used by the paper to answer: apparently not; a cointegrating relation is present.
No. RESET and the residual diagnostics have already provided reasons to doubt the simple linear formulation.
No. The paper explicitly leaves VAR, SVAR, VEC and related dynamic approaches as avenues for subsequent work.
This is the point at which the investigation becomes much more interesting than a standard “significant coefficient” exercise. The monetary variables may possess a genuine long-run statistical relationship while the particular linear equation initially chosen to represent that relationship remains incomplete.
Cointegration can defend a long-run relationship against the accusation of spurious regression without defending the exact regression specification used to represent that relationship.
A theoretically meaningful relationship can be real and yet nonlinear. It can be real and yet dynamic. It can be real and yet require several equations rather than one. It can survive the test for spuriousness while still failing a specification test.
That distinction prepares the final question of the paper—and of this guided reading:
That is where the final stage begins.
20. After all those tests, what has actually survived?
We can now return to the deceptively simple regression with which the investigation began. Money in circulation was modeled as a function of nominal GDP and the international price of gold. Ordinary least squares produced an extraordinarily high R², highly significant coefficients, and fitted values that visually followed the observed monetary series very closely.
Had the investigation stopped there, the conclusion would have been easy: the model fits remarkably well.
But it did not stop there.
Residual diagnostics raised questions about serial dependence. RESET tests showed that nonlinear information could matter. CUSUM and CUSUM-of-squares did not tell exactly the same story about parameter stability. A modification involving the GDP-related term changed some diagnostic results while generating new questions. Unit-root tests then exposed a deeper danger: the variables themselves contained enough non-stationary behavior to make a conventional levels regression potentially misleading.
Finally, cointegration analysis provided the paper with a way out of the spurious-regression problem. The Johansen tests rejected the hypothesis of zero cointegrating relations under the deterministic specifications examined. The author therefore concludes that the regression is not merely the artificial product of unrelated trending time series.
The paper finds evidence that the variables share a genuine long-run statistical relationship, while simultaneously finding that the particular linear regression used to represent that relationship is not yet an optimally specified model.
That combination is not a contradiction. It is arguably the most important substantive lesson of the entire exercise.
21. A relationship can survive even when a model does not
It is useful to separate three propositions that are easily collapsed into one.
The estimated regressions display very strong statistical comovement among money, nominal GDP, and gold.
The Johansen results are interpreted as evidence of cointegration and therefore against the hypothesis that the levels relationship is simply a nonsense regression among unrelated trends.
This proposition does not survive. Specification diagnostics leave substantial reason to search for a better mathematical representation.
This distinction is fundamental in applied econometrics. Empirical criticism of a particular equation does not necessarily refute the underlying relationship that motivated it. Conversely, evidence that a relationship exists does not confer immunity on the equation chosen to describe it.
That is why the paper’s methodological emphasis on hypothesis testing is so important. The tests progressively separate questions that a single regression output tends to blur together.
22. What the study does not establish
The safest way to understand the contribution is also to specify its limits.
First, the paper does not discover a final, uniquely correct dynamic equation for the U.S. monetary process. The additional GDP-related variable is explicitly presented as an empirically motivated modification rather than as the consequence of a completed theoretical derivation. The author also acknowledges that the analysis does not exhaustively determine the optimal lag structure through simulation or systematically compare all possible dynamic alternatives.
Second, the existence of cointegration does not by itself settle every substantive economic question. It indicates that the variables possess a long-run statistical structure that constrains their joint evolution. The economic meaning and direction assigned to that structure come from the theoretical framework in which the regression was constructed.
The paper itself makes precisely this broader methodological distinction earlier in the analysis: regression models represent relationships among variables, while the meaning and direction of those relationships must be interpreted through the relevant theoretical framework.
Rejecting the hypothesis of zero cointegrating relations helps answer the specific statistical question of whether the apparent long-run comovement can be dismissed as a spurious regression. It does not automatically identify the best functional form, the complete dynamic structure, or every causal mechanism operating among the variables.
Third, the empirical sample is small. The principal period contains only twenty-two annual observations, from 1990 through 2011. That matters particularly once additional regressors, lag structures, nonlinear terms, and multivariate time-series procedures enter the analysis. Every extra parameter consumes scarce information.
The study should therefore be read as an econometric investigation of specification and diagnosis—not as the final possible model of the monetary structure of the United States.
23. Why 2009–2011 turned out to matter
One of the paper’s more interesting empirical exercises is the comparison between specifications containing the first years after the 2008 crisis and versions ending in 2008.
A naïve expectation might be that adding three extraordinarily turbulent years would simply make the regression worse. But the diagnostics do not yield such a simple result.
For example, the paper reports a Durbin–Watson statistic of approximately 0.91 for 1990–2011 compared with roughly 0.51 for 1990–2008. Neither value is particularly reassuring in isolation, but the longer sample moves closer to the conventional benchmark near two.
The CUSUM comparison is also notable. The paper identifies five observations associated with instability in the 1990–2011 version, versus ten in the shorter 1990–2008 exercise, while CUSUM of squares indicates stability in both cases.
And after the provisional GDP-related modification, the longer period again performs better in some autocorrelation diagnostics: the paper reports disappearance of higher-order autocorrelation earlier in the 1990–2011 sample than in the shortened sample.
Adding crisis observations does not mechanically “contaminate” an econometric model. Sometimes precisely those observations reveal the structure of the process more clearly—or alter the diagnostics in ways that force the investigator to rethink what the apparently tranquil years were telling them.
This reinforces the paper’s historical orientation. A crisis is not merely an inconvenient outlier standing outside the system. For an economic analysis, crisis observations can be part of the phenomenon one is trying to understand.
24. Why the next natural step is a system rather than another isolated equation
By the end of the investigation, several clues point in the same direction.
The variables possess substantial temporal persistence. Residual autocorrelation appears in important specifications. Nonlinearity remains a concern. The possibility of endogenous interaction is raised. Cointegration indicates that the variables may be tied together over the long run. And Johansen’s framework itself naturally treats them as members of a multivariate dynamic system.
For that reason, the paper points toward the VAR family and related approaches—VAR, structural VAR, near-VAR, and vector error-correction models—as natural candidates for further investigation.
This is particularly relevant because the original theoretical framework concerns a social process whose variables need not be independent in any simple contemporaneous sense. Money, production, prices, financial conditions, and monetary institutions evolve together through time.
If the variables are cointegrated, a vector error-correction model becomes especially natural because it can combine short-run changes with the long-run relation identified by cointegration.
These alternatives are proposed as future directions rather than estimated as the definitive replacement model in the present study. The paper explicitly leaves VAR, SVAR, VEC and related specifications for subsequent investigation.
This matters because intellectual restraint is part of good model building. Discovering that the current model is incomplete does not imply that one has already discovered its successor.
25. The real protagonist of the paper is not money—it is hypothesis testing
At first glance, a reader could reasonably classify the study as a paper about money, nominal GDP, gold, and the United States between 1990 and 2011. Those variables provide the empirical material, and the economic framework determines what the regression is intended to mean.
But structurally, the study is about something broader: how an econometric model is criticized.
The sequence matters.
↓
Initial regression specification
↓
OLS estimation
↓
Residual and stability diagnostics
↓
Detection of autocorrelation and specification problems
↓
Provisional respecification
↓
New diagnostics
↓
Unit-root analysis
↓
Suspicion of spurious regression
↓
Cointegration analysis
↓
Rejection of zero cointegration
↓
A relationship that survives—but a model still awaiting improvement
Seen this way, hypothesis testing is not a final bureaucratic stage appended to an otherwise completed model. It is one of the mechanisms by which the model evolves.
A null hypothesis establishes a provisional benchmark. Data are used to confront that benchmark. Rejection—or failure to reject—changes what questions are sensible to ask next. Another specification is then examined. New assumptions become testable. New weaknesses appear.
The procedure is therefore iterative rather than ceremonial.
This is also why apparently contradictory diagnostics need not be embarrassing. White can fail to detect heteroskedasticity while Breusch–Godfrey detects autocorrelation. CUSUM can flag instability while CUSUM of squares remains inside its boundaries. A modification can improve one diagnostic and worsen another. Unit-root evidence can initially threaten a levels regression while cointegration later provides evidence that the long-run relation is not spurious.
That is not statistical chaos. It is what happens when different tests are aimed at different properties of a complicated object.
26. What non-specialists should take away
For a reader without formal econometric training, the deepest lesson can be stated with almost no mathematics.
Suppose someone presents a regression claiming that variable X explains variable Y. The graph looks excellent. The R² is 0.98. The coefficients have several stars beside them.
None of those facts is meaningless. But none of them completes the argument.
You would still want to know whether the errors contain patterns, whether the relationship changes through time, whether the assumed straight line is hiding curvature, whether the variables are simply trending together, and whether the apparent relationship survives methods designed specifically for non-stationary time series.
This paper is, in effect, an extended demonstration of that principle.
R² and fitted values help answer this question.
Residual, heteroskedasticity, autocorrelation, stability, and specification tests address this layer.
Unit-root and cointegration analysis address the deeper time-series problem.
Only after all three levels have been considered does the meaning of an impressive regression become substantially clearer.
27. What economists and econometricians should take away
For a technical reader, the paper is most interesting when understood as a worked example of specification uncertainty.
The progression from a static OLS equation toward questions of lag structure, serial correlation, conditional variance, functional form, structural stability, integration, and cointegration illustrates an important fact about empirical work: model choice is rarely separable from model criticism.
The paper does not deliver the last word on that process. Indeed, its own endpoint makes clear that further specification work is required.
That incompleteness is not incidental to the argument. It is part of the argument.
An “optimal” specification cannot be identified simply by maximizing fit. The researcher must progressively examine whether the statistical assumptions, temporal structure, functional form, and multivariate relationships implied by the model are compatible with the observed data.
This also explains why no single information criterion, goodness-of-fit statistic, residual test, or cointegration statistic can substitute for the entire modeling process. Each operates within a defined question. Model adequacy emerges from the architecture of those questions taken together.
28. The larger epistemological lesson
There is a broader idea running beneath the econometrics.
A statistical model is an abstraction. Its usefulness depends on what it preserves from the process it represents and what it leaves outside. The residual is therefore not merely mathematical garbage produced after estimation. It is the quantitative trace of everything the fitted equation did not capture.
Sometimes that remainder behaves approximately like noise. Sometimes it remembers yesterday. Sometimes its variance clusters. Sometimes it reveals a nonlinear structure. Sometimes the entire apparent relationship is generated by common stochastic trends. And sometimes variables that wander individually nevertheless reveal a stable long-run connection when studied jointly.
Hypothesis tests give the investigator systematic ways to distinguish among those possibilities.
That change of question alters the role of statistics. Instead of being used only to confirm an equation, statistical inference becomes a method for attempting to falsify inadequate representations and forcing the researcher toward better ones.
The resulting model may therefore become more complicated as the investigation progresses. But that complexity is not necessarily a defect. If the process itself is dynamic, historically dependent, nonlinear, and jointly determined, a simpler equation can achieve elegance only by omitting part of what needs to be explained.
29. Final assessment: an excellent fit, a real relation, an unfinished model
We can now describe the paper’s journey in three phrases.
The initial OLS regression explains an exceptionally large proportion of the observed variation in the money series.
Johansen cointegration results lead the paper to reject the interpretation that the relationship is merely spurious.
Nonlinearity and dynamic structure remain unresolved, so the linear regression cannot be treated as the final model.
The last point is crucial because the temptation in empirical work is often to search for a specification that finally produces reassuring statistics. This investigation points in another direction. Every apparently satisfactory result can generate another legitimate question.
A high R² leads to residual analysis. Residual dependence leads to dynamic modification. Modification leads to another specification test. Trending variables lead to unit-root analysis. Unit roots lead to the possibility of spurious regression. That possibility leads to cointegration. Cointegration, finally, saves the existence of a long-run relation without saving the original equation from further criticism.
The paper therefore ends somewhere more intellectually interesting than “the model works” or “the model fails.”
And that is precisely why hypothesis testing matters.
Its deepest role is not to decorate a regression table with asterisks. It is to organize doubt: to convert suspicions about a model into questions that can be confronted with data, to reveal which assumptions remain defensible, and to indicate where the next round of modeling must begin.
In the study examined here, that process starts with three variables—money, nominal GDP, and gold—and ends with a much broader lesson. Good econometrics is not the art of finding an equation that looks convincing. It is the discipline of discovering how far an equation can be trusted before the evidence forces us to build a better one.
Gómez Julián’s investigation begins with an OLS model of the U.S. money supply as a function of nominal GDP and the international price of gold for 1990–2011. Despite an exceptionally high R² and highly significant coefficients, diagnostic testing reveals residual and specification problems that motivate further investigation. Ramsey RESET raises doubts about linearity; autocorrelation and stability tests reveal temporal complications; unit-root analysis raises the danger of spurious regression; and first-difference results do not provide a completely simple resolution. The analysis therefore moves to cointegration. After an inconclusive intermediate stage, Johansen tests reject zero cointegration under the specifications examined, leading the paper to conclude that the regression is not spurious. Yet the same investigation also concludes that the model remains statistically improvable, particularly in its specification, and points toward richer dynamic approaches such as VAR, SVAR and VEC models. The methodological lesson is consequently larger than the monetary application: hypothesis testing is most useful when treated as an iterative instrument for criticizing, revising and progressively improving a statistical model rather than as a final ritual for attaching significance stars to coefficients.


Leave a Comment/Deja un Comentario