---
title: 'Signal v. Noise: Examining the Copper-to-Gold Ratio as a Predictor of Treasury Yield Curve Shifts'
source: 'https://www.youtube.com/watch?v=sI0UMDO2nJE'
video_id: 'sI0UMDO2nJE'
date: 2026-09-17
duration_sec: 1476
channel: 'Brandon Kaplan'
---

# Signal v. Noise: Examining the Copper-to-Gold Ratio as a Predictor of Treasury Yield Curve Shifts

> Source: [Signal v. Noise: Examining the Copper-to-Gold Ratio as a Predictor of Treasury Yield Curve Shifts](https://www.youtube.com/watch?v=sI0UMDO2nJE)

## Summary

This video presents a final project for FinTech 534, investigating the predictive power of the copper-to-gold ratio on the 10-year minus 2-year treasury yield spread. The analysis uses a Vector Autoregression (VAR) model to test the relationship, finding that while the model is robust, it fails to accurately forecast the spread due to the nonlinear and volatile nature of the underlying data. The presenter concludes that a more complex model, such as a nonlinear autoregressive model or neural network, may be necessary to capture the true dynamics.

### Key Points

- **Project Introduction** [00:00] — Brandon Kaplan presents his final project for FinTech 534, titled 'Simulverse Noise', which examines the copper-to-gold ratio as a predictor of treasury yield growth shifts.
- **Scope and Purpose** [00:15] — The goal is to empirically test the claim that the copper-to-gold ratio is a leading indicator for the 10-year minus 2-year treasury yield spread, investigating the veracity and conditions of this economic theory.
- **Copper as an Economic Indicator** [00:49] — Copper is an industrial necessity used in construction, electronics, and green energy. Its demand and price are correlated with economic activity, making it a potential leading indicator for economic growth.
- **Gold as a Safe Haven** [01:35] — Gold is a financial plunder, not an industrial commodity. It is used to store wealth and acts as a hedge against inflation and market volatility, as shown by KP et al. and Nguyen et al.
- **Model Error Analysis** [19:00] — The RMSE is 0.145, but the MAPE is 291.7%, indicating that the model's predictions vary widely as a percentage of actual values. The small RMSE is misleading because positive and negative errors cancel out.
- **Model Limitations** [20:55] — The VAR model fails to capture the nonlinear dynamics and volatility of the copper-to-gold ratio and yield spread. The relationship appears to be nonlinear, and the model is not equipped to handle this.
- **Future Improvements** [22:36] — Suggestions for improvement include using more granular data, incorporating macroeconomic variables, and employing nonlinear models like GARCH or neural networks (e.g., LSTM) to better handle volatility and nonlinearity.

### Conclusion

The copper-to-gold ratio does not show a significant linear relationship with the treasury yield spread in this analysis. The model's poor forecasting ability suggests that the relationship is nonlinear and may require more sophisticated modeling techniques to uncover.

## Transcript

Hi, my name is Brandon Kaplan and I am presenting my final project for FinTech 534 titled Simulverse Noise, examining the copper to gold ratio as a predictor of treasury yield growth shifts.
So let's begin with a brief overview of the scope of this project. The goal here was to take an empirical look into the copper to gold ratio's predictive power on the 10 minus the 2-year treasury yield spread.
The purpose of this analysis is ultimately to investigate the veracity and the conditions of claims in economic theory that the copper to gold price ratio is in fact leading indicator of treasury yields. So I'll start with background and motivation, get into the methodology of my analysis, then the empirical investigation itself with results and discussion, and then finally a conclusion.
So starting with some background. We'll talk about copper. Copper is an industrial necessity, an indispensable commodity in any modern economy, according to Parnes.
Now, copper is definitely used in a variety of economic activities, from construction to electronics and machinery, and more recently, green energy. Its demand, as well as its price, because price rises with demand, is heavily correlated to economic activity.
When things are being built, copper is usually going to be involved. Jahnke showed that in some rich countries, demand can actually be used to forecast economic growth. On the other hand, gold is considered a financial plunder and not an industrial commodity.
Its scarcity and its softness result in its limited industrial applications. In fact, throughout history, gold has been a resource in which to store wealth. Its value also tends to appreciate in terms of economic instability.
It's been considered a safe haven and a hedge against inflation as shown by KP et al. In addition to being a hedge against inflation, it's also just a hedge against unfavorable volatility in the market as shown by Nguyen et al.
Where if you have money in stocks and the market seems very volatile and you don't want to be exposed to it, take the money out of the stocks, put it in the hard asset like gold and it's safe. Thus, broadly speaking, in the copper to gold ratio, the copper price is going to be indicative
of periods of stability or expansion and less volatility when increasing. And gold price is going to be indicative of turbulence, downturn, and uncertainty when increasing. Now, treasury bonds, these are debt securities issued by the U.S. government.
The yield that we're going to be talking about is just the interest that the U.S. government pays on an annual basis to the bondholder. Now the T10Y2Y is an abbreviation that the Fed uses for the 10 minus the 2-year treasury
yield spread and the difference between the yield on the 10-year and the 2-year treasury bonds. Normally, during normal economic circumstances, the yield curve is upward sloping and this is indicating that the yield on the 10-year is higher than the yield on the 2-year, so
it's positive. Now, this makes sense because a long-term bond like a 10-year should have a higher yield than the 2-year bond because if you're an investor and you purchase a bond, then you're
going to expect greater interest payments for purchasing the bond that matures after a longer period of time. With more time, there's going to come more risk and you should be compensated for that
accordingly. In anticipation of times of crisis, turbulence, deflation, investors are usually going to be anticipating that the Fed slashes interest rates.
Since long-term bonds are typically having higher yields than short-term, locking in a higher rate for a longer-term bond before interest rates are cut is a smart decision. Otherwise, if you don't do that, and say you put your money in a short-term bond, when
it matures you might have to reinvest that principal at a lower rate because the rates are slushed. In this case the fixed income from the long-term bond yield is a hedge against inflation. We can sort of understand why the curve
inverts that is it goes below this here black line at zero in these times as follows. So if interest rates are expected to fall or are falling then the demand
for existing bonds with high yields like the 10-year is going to increase so that you can lock in the best rate possible. If the demand for these bonds increases, this implies like, wealth demand price increases, and
thus this price increases, yield is inversely correlated, inversely related to price, then the yield drops when the price goes up. If the yield on the 10-year is dropping, regardless of what the 2-year does, it stays the same
and that they both go down at a proportional rate, which it won't, then you're going to have a negative difference when the 10-year is less than the 2-year.
And in these cases, you have a negative yield, which is when the blue line goes below the black line. Now you'll notice that from the stats from the Federal Reserve Economic Database, each of these gray chunks is an economic recession.
First is the dot-com, this is the Great Recession, and then this is the COVID-19 pandemic. Before each of them, there is a yield curve inversion. Getting into my motivation behind citing this is that I'm interested in the relationship
between the copper to gold ratio and the 10-year line of the tiered yield spread on account of the wide range of implications of the latter. From mortgage rates, bank loans, other activities, and fixed income, just to the general market
as a whole the treasury yield is indicative of investor sentiment across a wide range of activity I also interested in looking at a forward indicator not as opposed to a lagging or backwards or a coincidence or real indicator
which get a lot more attention in the literature. There's definitely a broad range of appeal here beyond financial experts. I, for one, am just generally interested in fixed income securities, and I think people that are interested in the macroeconomy as a whole would take
interest in this. Given the post-COVID turbulence this is also somewhat current and relevant. Now getting into my methodology, I'm going to look here first at the data courses that I'm processing.
The copper data, which is for Grade A copper, this is its spot price on the London Metal Exchange in dollars per metric ton. I got this from the National Institute of Statistics and Economic Studies or INSEE and this is
downloaded on an average per month basis. Now the gold data which was the gold multi-contributor price per troy ounce in US dollars. Troy ounce is about 10%
heavier than the standard ounce. This is a priori from the Refinitum on a per day basis. The yield data I got that from the Federal Reserve Economic Database on an average percent per month basis. So what I had to do with the data to clean it up
was that the daily gold spots had to be converted to an average per month price from the daily price the copper spots had to be converted to price per pound from price per metric ton and then the copper to gold ratio itself was just computed by taking
the quotient of the copper spot per pound over the gold spot for triumphs the treasury yield spread i didn't compute myself the federal reserve did but they just set the difference between the 10-year yield and the two-year yield and average dip per month.
In terms of the analysis, the data frame spans from January 1995 to the end of October 2023 at a monthly frequency. To train and then test my model, I split the data up into training test sets
accordingly. The time series data was non-stationary meaning that like its mean and variance we're not classic over time. To handle this for the model that I test to use, I have to
difference the data, which means to subtract the current observation from the previous, and then you help remove the trends and stabilize the needs. The stationarity I tested before and after differencing using an augmented Dickey-Fuller test. Afterwards, when I knew that the data was
stationary, I set a vector autoregression or VAR model to the difference data. So a VAR model is a kind of stochastic, randomly determined process model. It can handle multivariate
time series. So the model itself is inherently endogenous and this is going to mean that there isn't a clear independent or dependent variable as you might think in a traditional regression where you increase A and B response. So instead a VAR model is going to treat every
variable in the system as a function of the path values of all the other variables in the system. This means that each variable has its own equation within the model with path values, or what we'll refer to as lags, of itself and all other variables as predictors. The term
autoregressive just indicates that the model is based on the premise that path values have a regressive effect on current values. This is going to be important for time series data where we're looking at historical data and running to influence future. The number of lags or path values was
was determined using the AIC, which is the information criterion. And here are the number of lags that we choose and want to optimize it, because it's important in just balancing complexity with overfitting.
We don't want to get the best staff if we're fitting too many basically. For instance, a VAR model with one lag just implies that each variable is predicted from the previous value of itself
and the previous value of each other variable in the model. Lastly, I did an impulse response analysis, forecast error variance, and Granger cloud reality tests.
For the forecasting, I just did an out-of-sample forecast conducted on the test set using the model I fitted. I plotted and found some metrics, mean absolute error, root mean square error, mean absolute
percentage error. We'll get into those a little bit later. In terms of the limitations and the assumptions that this model makes, so the data only encompasses 23 years for training but is also on a monthly frequency so the frequency was not very granular.
Likewise the test data overlapped with the brief but intense COVID recession and general turbulence So I was using data from a time of turbulence to test something that was fitted over a wider
range of time. The outliers in the time series data can't be easily removed without messing up the entire data set. Moreover, in terms of the model, it's a theoretical, so there's not really a theoretical structure,
so it might not capture the underlying economic theory or causal relationships at hand. have the endogeneity as I mentioned previously this is intentional and part of the model's design basically you're trying to capture the interdependencies
among the variables and not so much say just one predicts the other in terms of the impulse response analysis the shocks to the model are same to your orthogonal meaning just each shock or disturbance to one variable is
considered to be statistically independent from shocks to the others in In real life, that might not be the case. You have some sort of economic emergency that affects both the treasury yield and the copper to gold ratio.
Here we're going into sort of the empirical investigation with some of my results. This is a graph I created of the copper to gold ratio over time You can see here that in the dark gray we have the same three recessions and in each dark gray point there is a dip in the copper to gold ratio meaning that either copper price decreased or gold price increased during this time or both
And so overall, the ratio went down. Note, though, there are decreases in the copper to gold ratio and many other points that are at recessions. Just at each recession, there is a dip. So after I trained the AR model I got a summary and the summary shows that the AIT found that
the optimal number of lags is going to be 8. The model itself was statistically robust and had a p-value much smaller than the 0.05 necessary for its physical significance. In terms of the different path values or lags of the ratio variable, 3 and 7 were demonstrating
very marginal significance in predictive power but better than anything else. 0.058 and 0.71 both pretty close but not less than 0.05. Overall we can't say which includes its specific significance but there was predictive power in the ratio
on the treasury yield spread based on this model. The R-squared value of 0.23 implies the model explains 23% of the variance in spread time series data. This is generally considered a pretty weak correlation and R squared value close to 1 is
considered strong. However, the correlation residuals of 0.038, those are small and this suggested the model did account for most of the inter-variable dynamics between the ratio and the spread. For the Granger Causality Test, the F-statistic p-value was 0.4455, much greater
than 0.05. It's not significant and the ratio did not greater cause the spread to move. In this graph we're looking at the impulse response analysis. How a sudden and unexpected
change in the ratio here diff ratio or difference ratio would affect our spread over the next 20 years, 20 periods rather. The dashed line around the impulse response line, the impulse response
line or impulse response quotient is the black line, the dashed line here in red, those represent the bounce of the 95% confidence interval generated through bootstrapping. The interval is just providing a range of values in which we can be 95% confident that the true impulse
response lies according to the data that we sent, not like the actual real life reality, just the true according to the model. So we're pretty sure that the black line lies between these red lines.
Now the significance of the response or whether it's significant statistically is going to be assessed by checking whether the confidence intervals include zero which is this red line
that is not dashed. So if the confidence interval does not include zero implies that you can be fairly confident that the observed response or impact of the shock is different from zero and therefore statistically significant. To determine if the change in the spread
falls in the same direction as the shock and ratio, we're looking at the function, the black line. Our hypothesis is that, the null hypothesis, right, is that if ratio experiences a positive shock, spread will also increase. However, we can see the
lines motion rates above and below zero all within the confidence interval over time. This is going to say we can't confidently say that there's a relationship between ratio and spread. So this one our current model of
shock and ratio doesn't lead to a predictable or significant change in spread that one can act reliably upon. Here we're looking at a chart that shows a forecast error variance decouple which helps us understand how much of the
future uncertainty in our variable of interest, diff spread, which is just the difference treasury spread yield spread is due to its own shocks which I explained in the previous slide versus shocks to diff ratio so that here we have
percentage which is just percentage representing the proportion of the forecast error variance horizontal axis is more confusing the horizon which is
just the number of time steps ahead for which the variance is being decomposed. Now each bar here represents a point in the future, so one point on the horizon, two,
etc. where a forecasting for the spread. The light portion here, according to this legend right, is showing how much the amount of forecast error variance in the spread is explained by its own shots.
the darker portion, which is so small you can't even see many of them except for as it goes farther out on the horizon, is explained by shocks to diff ratio. So you can observe here that the light portion is obviously dominant in all tickets. It
basically means that the majority of the forecast error variance in spread is due to its own shocks. The darker segments are very small and that indicates that shocks to diff ratio explain only a minor part of the forecast error variance in diff
spread. This analysis basically says that while the ratio has some influence on the forecast or a variance of spread, the majority of the variability in our forecast for spread is actually due to the variables of our own shocks. And this is important in understanding the
dynamics of the spread versus the ratio in terms of leveraging it.
now here's the results from the actual forecast in terms of the plot and the metrics so visually it's pretty clear here that the model forecast on that sample test data from Jan 2019 to October 2023 which is the dashed red line
is pretty awful considering how much the forecast spread differs from the actual spread here the actual spread in blue goes up and and down So the metrics that we looked at were mean absolute error which is MAE
This is the average magnitude of the errors between predicted and actual values. This doesn't consider the direction, positive or negative, of the errors, and they're all treated the same. So a lower end MAE is going to indicate that the model's predictions are generally closer
to the actual observed value. Pretty straightforward measure of accuracy. Now the RFME or root mean squared error is similar but gets extra weight to larger errors. So if you're here,
this error is weighted larger than if you're here. So this is significant because useful when large errors are particularly undesirable, which in this case that might be the case.
lower RSME is better and it suggests that the model frequently is not going to make large errors when it's predicting. Lastly we have the MAPE and this is measuring the MAE as a percentage of the actual values which provides a more relative
measure of error. So if you want to express the MAE proportional, proportionate to the size relative to the actual measurements, this is useful. So we have here like 0.0112 MAE. Now that itself seems small, but if we compare it to the fact that the actual spread only deviates from at most 0.3-ish to at most 0.-3-ish, 0.1 is very significant.
And you can see that in MAP, which is 291.7. That's huge. This basically means that the model predictions vary widely as a percentage of the actual values.
And that's especially true because the different spread is very close to zero. We also have the RSME of 0.145. And this on the surface suggests that the predictions are quite reliable,
but it's also not taking into account that about half of the predictions are positive errors and half of them or more of us are negative errors and so the positives and the negatives are mostly canceling out and we're getting a
root-straining error that's pretty small so on the surface the M&E and the RSME are a little deceiving but in context does back up the fact that the forecast is just poor overall you can see it's only right at these points where the blue
line and the dashed line intersect and it's close some other times but more or less it doesn't capture this sort of random up and down movement. This is related to my
conclusions. Basically, while we've seen that the model itself is robust, it definitely lacks the ability to forecast accurately. So it seems to me that the model does not
capture enough about the dynamics of the copper to gold ratio and the 10 year minus of two-year times your yield spread, and this indicates that the relationship might be better assessed by a nonlinear model. The ratio and the yield themselves are satiric processes that are both volatile and nonlinear,
and the DAR model, while it can handle nonlinear data, can't really represent a nonlinear relationship, and it also doesn't handle volatility all too well. So, in addition, we have this power virus that I mentioned in the assumptions.
They are common to both ratio and needle spread and because they're difficult to remove, we would need a way to treat them more robustly because they might just be created externally from the bulb.
Since there are larger macroeconomic determinants beyond ratio and spread themselves that are quite here over long periods of time, the large structure itself, where we're looking backwards might have trouble mitigating this. Introducing more time series variables that would
you know account for these macroeconomic factors this would not be advantageous because I believe that this would just introduce more side co-linearity and I would just avoid that.
If you did that then you might have independent variables that are possibly correlated and you get something that is empirically not valuable. Nonetheless there are going to be other factors that influence the ratio on the spread that we need to
account for and to do that you might consider in the future a different type of model like a augmented VAR model with a generalized autoregressive conditional heteroscedar 50 model or GARCH where we can model the conditional
variance and better handle the volatility of the data. And while these findings show no significant relationship between the ratio and the spread themselves, this doesn't preclude that from the
indication reality. I also looked at a very specific frequency of data and window which is long term with less granular data. Very likely the case that the actual
indication of the ratio and the spread is on a much shorter term scale and so I would need to look at much more granular times through data to get a sense of that. I would definitely like to look at this further where I did more data transformation to more robustly handle any of
the outliers and at the end of the day incorporate a non-linear model would be my ideal situation such as a nonlinear autoregressive model or a neural network model like long
short-term memory networks. Overall this is important in the sense that the VAR model is not equipped to handle this kind of time series data even though it is
pretty good at handling other time series data and that suggests that the relationship itself between the ratio and the yield is nonlinear. Here are the in terms of the research that I provided.
And thank you very much for listening to my presentation.
