Elm Wealth Research

Posts about:

Uncategorized

Merton Share Derivations: What’s in your denominator?

May 2, 2024

Uncategorized

Merton Share Derivations: What’s in your denominator?

By Jeffrey M. Rosenbluth and James White 1

1.  Introduction

The Review of Economics and Statistics published a pair of companion papers in 1969. “Lifetime Portfolio Selection by Dynamic Stochastic Programming”, by Paul Samuelson and “Lifetime Portfolio Selection under Uncertainty: The Continuous-Time Case”, by Robert Merton. Both deal with the question of how to allocate one’s portfolio between a risk-less and risky asset in a multi-period setting. The Samuelson paper considers the discrete time case and Merton’s the continuous time one. Merton solves this problem and provides a closed form solution of the highly stylized case where the risky asset rate of return follows a Brownian Motion (so that its price follows a Geometric Brownian Motion), the risk-less rate is constant, the utility function is CRRA (Constant Relative Risk Aversion) and the investor re-balances the portfolio continuously. Under these assumptions, he also shows that portfolio selection is myopic, that is, independent of the investment horizon and in fact the investor keeps a constant fraction of her wealth in the risky asset. We call this fraction the Merton Share. It is important to note that the Merton Share formula would be different under alternative sets of assumptions and that some of Merton’s assumptions are unrealistic: in particular, continuous re-balancing and Geometric Brownian Motion for asset prices. Nevertheless, we believe that using the Merton Share as a rule of thumb makes good sense and will often be close to the correct solution.

Merton used the theory of optimal control, and the Bellman principle of optimality in particular, to derive a partial differential equation (the Hamilton-Jacobi-Bellman equation) to solve the problem. In general, finding a closed-form solution to the HJB equation is rare and solutions are typically found numerically. In this note, we motivate and derive the Merton Share several different ways that are hopefully easier mathematically and provide more intuition as to why the formula makes sense. In doing so, we often deal with a single-period model and sometimes need to use approximations to derive the formula. The framework we will be using throughout to derive this formula is Expected Utility maximization; we will also often be assuming CRRA utility.

U ( W ) = { W 1 − γ 1 – γ if  γ ≠ 1 log ⁡ ( W ) if  γ = 1

The defining feature of CRRA utility functions and hence its name, Constant Relative Risk Aversion, is that relative risk aversion is constant:

R ( W ) = − W U ” ( W ) U ′ ( W ) = γ

We denote by k̂ the optimal fraction of wealth to invest in the risky asset. The Merton Share formula is:

k ^ = μ γ σ 2

where μ is the expected excess return,2 that is the return on the risky asset minus the risk-less rate and σ is its standard deviation. This formula certainly passes the smell test, the Merton Share is higher when excess return is higher, and lower when standard deviation and risk aversion are higher. The variance term in the denominator σ2 as opposed to perhaps σ may seem less intuitive though.

2.  Motivation

Let’s start to reveal why the denominator in the Merton Share is variance as opposed to standard deviation. We actually have another relation for k that relates it to the standard deviation of the portfolio. If we invest a fraction k of wealth in the risky asset with standard deviation of return σ, then the standard deviation σp of the portfolio is kσ.

σ p = k σ

Rearranging, we have:

k = σ p σ

If we choose k = k̂ (the Merton Share), and let σ̂p denote the standard deviation of the portfolio at k̂, the we obtain:

σ ^ p σ = μ γ σ 2

so that:

σ ^ p = μ γ σ

This, hopefully, provides some intuition for why variance in the denominator of our Merton Share formula makes sense. It says that the risk (standard deviation) of the optimal portfolio is the ratio of excess return to standard deviation of the risky asset divided by the coefficient of relative risk aversion γ. The ratio of excess return to standard deviation is called the Sharpe Ratio, and is a commonly-used metric of the quality of a risky asset or trade.

Our intuition didn’t lead us far astray. It’s the risk of the optimal portfolio – not the optimal fraction – that is proportional to the Sharpe Ratio.

2.1  Myopic Portfolio Choice

Let’s approach the question of “Why variance in the denominator?” from another angle. First, in addition to CRRA with relative risk aversion γ, we also make the more restrictive assumption that asset returns are independent over time.3 Consider an investor with a two-period horizon. At the end of the first period the investor is faced with a single period optimization problem and since her risk aversion does not depend on wealth and returns are independent, the solution to this problem does not depend on how much was invested in the risky asset in Period 1. Now, the time 0 portfolio choice problem does not depend on time 1 wealth. Hence, the investor makes a single period portfolio choice at time 0 as well. When an investor makes the same portfolio decisions regardless of horizon, we say portfolio choice is myopic. By backward induction the above argument can be applied to any number of periods. This shows that with CRRA utility and time independent returns that portfolio choice is myopic. In our case we have in fact an even stronger result, constant portfolio choice over time, in which the investor holds the same fraction of wealth in the risky asset in each period. Let’s state this a a theorem and prove it more formally.

Theorum 1. If returns follow a stochastic process with independent increments, then for investors with CRRA utility of wealth, portfolio choice is constant over time.4

Proof. By the scale invariance property of CRRA utility, we know that k̂ does not depend on Wt that is wealth at any time t. From the independent increments assumption, we know that future risky asset prices do not depend on past wealth or past choices of k̂. Therefore, portfolio choice is myopic and k̂ is constant.

How does this help us to motivate the use variance in the denominator of the Merton Share? If portfolio choice is myopic, that means we would invest the same fraction of wealth in the risky asset for any horizon t. Suppose we have the function:

k ( X t ) = μ t ρ ( X t )

and we are choosing between standard deviation and variance for the operator ρ. We know that if our choice is myopic, then k(Xt) will not depend on t. For this to be true, its denominator must be a factor of t so that the t‘s will cancel. If Xt has independent increments, as do the majority of the stochastic processes employed to model excess returns, then StDev(Xt) = σ √t, so ρ can’t be standard deviation. On the other hand, variance Var(Xt) = σ2 t works just fine.

3.  Derivations

We provide core derivations (and two more in appendix) that are designed to motivate different aspects of the portfolio choice problem as it relates to the Merton Share.

3.1  Static Approximation

In this section, we assume the risky asset excess return is identically distributed over periods of the same length and that they are uncorrelated. In this case, both mean and variance are proportional to the horizon. The utility function U(W) is required to be twice differentiable and concave. We approximate this utility function with a Taylor series, resulting in a formula that is only valid for short horizons 5. We then specialize this result to the CRRA utility case.

Let W be the value of the initial portfolio. For a portfolio return Y, let U(W(1 + Y)) be the utility after one period with horizon t. Since U is twice differentiable we can approximate it with a second order Taylor series about Y = 0:

U ( W ( 1 + Y ) ) ≈ U ( W ) + U ′ ( W ) Y W + 1 2 U ” ( W ) ( Y W ) 2 E [ U ( W ( 1 + Y ) ] ≈ U ( W ) + U ′ ( W ) E [ Y ] W + 1 2 U ” ( W ) E [ Y 2 ] W 2 = U ( W ) + U ′ ( W ) E [ Y ] W + 1 2 U ” ( W ) ( Var [ Y ] + E [ Y ] 2 ) W 2

Notice what is happening here, the combination of approximating utility by a Taylor series and taking its expected value introduces the moments of the probability distribution into the equation! If we take more terms of the Taylor series for a better approximation, then we need more moments. This should gives us additional comfort in choosing variance, not standard deviations, in the Merton Share formula.

In our case, the portfolio with a fraction k of wealth invested in the risky asset and the remainder in the risk free asset Y = (r + kX)t, where t is the horizon of the investment. The excess return X has mean μ t and variance σ2 t as per our assumption, and r is the risk free rate of return. So E[Y] = (r + kμ)t and Var[Y] = k2σ2 t. Since E[Y]2 = (r + kμ)2 t2, it can be ignored for small t.

We want to maximize:

U ( W ) + ( r + k μ ) t U ′ ( W ) W + 1 2 k 2 σ 2 t U ” ( W ) W 2

We differentiate with respect to k to obtain the first order condition:

0 = μ t U ′ ( W ) W + k σ 2 t U ” ( W ) W 2 = μ U ′ ( W ) + k σ 2 U ” ( W ) W

Hence:

k ^ = − μ U ′ ( W ) σ 2 W U ” ( W )

Recall from Section 1 the coefficient of relative risk aversion:

R ( W ) = − W U ” ( W ) U ′ ( W ) 1 R ( W ) = − U ′ ( W ) W U ” ( W )

Substituting this in to the above formula for k, we arrive at:

k ^ = μ R ( W ) σ 2

This is a fairly general result, we have made very few assumptions about the utility function and asset return distribution.

Specializing to the CRRA utility case R(W) = γ so that:

k ^ = μ γ σ 2

the Merton Share.

3.2  Asset Prices follow a Geometric Brownian Motion

In this section, we derive the Merton Share using assumptions similar to the ones Merton himself used. We assume CRRA utility, and have a risky asset St that follows a Geometric Brownian Motion (GBM) and a risk-less asset Bt with continuously compounded return r. That is:

d S t S t = ( r + μ ) d t + σ d Z t d B t B t = r d t

where Zt is a Standard Brownian Motion (i.e μ = 0, σ = 1).

This setup is very common in finance. It is also very different from the derivation above, in that we are now have a dynamic optimization problem. Hence, k̂ is now a stochastic process that depends on the price path of the asset and time t, – call it k̂(St, t). Solving for k̂(St, t) is a problem in Stochastic Control6 which is beyond the scope of this note and requires quite a bit more mathematical machinery 7. But by employing Theorem 1, we know k̂ is constant and hence we can side step the stochastic control problem. Note that we still require the portfolio to be re-balanced to contain a fraction of wealth k̂ in the risky asset at every moment in time.

Given the above differential equations we can write down the stochastic differential equation (SDE) for wealth. We can think of this as saying that instantaneous returns on the wealth portfolio are k times the instantaneous return on the risky asset plus 1 – k times the return on the riskless asset.

d W t W t = ( 1 – k ) d B t B t + k d S t S t = ( 1 – k ) r d t + k ( r + μ ) d t + k σ d Z t = ( r + k μ ) d t + k σ d Z t

Just as the risky asset is following Geometric Brownian Motion, we can see that the portfolio also is following GBM, i.e. the portfolio is also expressed as an SDE for GBM. The difference now is that the drift is r + kμ and the diffusion term is kσ, hence:

W t = W exp ⁡ ( ( r + k μ – 1 2 k 2 σ 2 ) t + k σ Z t )

Without loss of generality, we can let W = 1, letting Rt = (r + kμ -½ k2 σ2)t + k σ Zt. We have:

E [ W t 1 – γ 1 – γ ] = E [ exp ⁡ ( ( 1 – γ ) R t ) 1 – γ ] = exp ⁡ ( ( r + k μ – 1 2 k 2 σ 2 ) t + 1 2 ( 1 – γ ) k 2 σ 2 t / 2 )

where we have used the fact that the mean of a log-normal random variable with drift m and diffusion term s is:

exp ⁡ ( m + 1 2 s 2 )

For γ > 1, maximizing this expression is the same as minimizing:

( r + k μ – 1 2 k 2 σ 2 ) + 1 2 ( 1 – γ ) k 2 σ 2

The first order condition is:

μ – k σ 2 + ( 1 – γ ) k σ 2 = μ − γ k σ 2 = 0

Solving for k gives:

k ^ = μ γ σ 2

Appendix

Normal Returns and Constant Absolute Risk Aversion (CARA) Utility

CARA utility and normally-distributed returns provide the only case where the Merton Share is an exact formula in the single-period world. Normal returns are undesirable since they allow negative asset prices and can’t be used for both sub-period and total period returns. The CARA (exponential) utility function exhibits constant absolute risk aversion A, which is also unrealistic. Despite these shortcomings, this case provides an instructive example. The CARA (exponential) utility function is:

U ( W ) = − exp ⁡ ( − A W ) A

To maximize expected utility, we can minimize the negative of U. Letting W be the starting wealth, we have:

min k E [ exp ⁡ ( − A ( 1 + r + k X ) W ) ] = E [ exp ⁡ ( − A ( 1 + r ) W ) exp ⁡ ( − k A X W ) ] = exp ⁡ ( − A ( 1 + r ) W ) E [ exp ⁡ ( − k A X W ) ]

The expectation of the log-normal random variable:

E [ exp ⁡ ( − k A X W ) ] = exp ⁡ ( − k A μ W + k 2 A 2 σ 2 W 2 2 )

So our minimization problem becomes:

min k exp ⁡ ( − A ( 1 + r ) W ) exp ⁡ ( − k A μ W + k 2 A 2 σ 2 W 2 2 )

which is the same as:

max k A μ W – k 2 A 2 σ 2 W 2 2

The first order condition is:

0 = A μ W – k A 2 σ 2 W 2 = μ – k A σ 2 W

Solving for k gives:

k ^ = μ A W σ 2 = μ R ( W ) σ 2

This is effectively the Merton Share formula, and is the same result we obtained in Section 3.1.

Let’s explore this result a bit further. Since we assume that a utility function U is strictly concave, we know from Jensen’s inequality that:

E [ U ( W 1 ) ] < U ( E [ W 1 ] )

We can think of this as an equality:

E [ U ( W 1 ) ] = c U ( E [ W 1 ] )

for some c > 1. For most combinations of utility function and wealth distribution, we do not know what c is explicitly, but for the combination of exponential utility and normal returns we do. It’s exp(k2 A2 σ2 W2 /2). This shows that our maximization problem is a trade-off between mean μ and variance σ2.

Quadratic Utility

The quadratic utility function

U ( W ) = − 1 2 ( a – W ) 2

is not very realistic in that it has increasing absolute risk aversion and a “satisfaction” point beyond which more wealth lowers utility.

Its Arrow-Pratt Measure of Absolute Risk Aversion is:

A ( W ) = 1 a – W

It’s often used to demonstrate a utility function whose portfolio selection fraction depends only on mean and variance regardless of the distribution of returns.

As usual, we start with the expected utility maximization problem:

max k E [ − 1 2 ( a – ( 1 + r + k X ) W ) 2 ]

Differentiating with respect to k and setting to 0:

0 = E [ W X ( a − ( 1 + r – k X ) W ) ] = a μ W − ( 1 + r ) μ W 2 – k ( σ 2 + μ 2 ) W 2 = μ ( a − ( 1 + r ) W ) – k ( σ 2 + μ 2 ) W k ( σ 2 + μ 2 ) W = μ ( a − ( 1 + r ) W ) k = μ ( a − ( 1 + r ) W ) ( σ 2 + μ 2 ) W = μ σ 2 + μ 2 ( 1 – r W A ( W ) W A ( W ) )

Static Approximation Revisited

When we derived the Merton Share back in Section 3.1, we made the assumptions that excess returns are identically distributed over periods of the same length and that they are uncorrelated. We needed to do this to ensure that both portfolio return and variance scale with horizon t. This is what allowed us to approximate the solution for small t. It turns out we can drop this restriction if instead we assume that the mean excess return is small. We can always write the excess return X as the sum of its expected return and a random variable with zero mean and the same standard deviation as X, say Z:

X = μ + Z

This lets us take k̂ to be a function of μ, k̂(μ) then we can use a first order Taylor expansion about 0 to estimate it.

k ^ ( μ ) ≈ k ^ ( 0 ) + μ k ^ ′ ( 0 )

And since the optimal investment in a risky asset with zero return is 0.

k ^ ( μ ) ≈ μ k ^ ′ ( 0 )

Let W1 = (1 + r)W and w̃ = W1 + k̂(μ)(μ + Z)W. At the optimum, k̂, the first order condition must be 0.

E [ ( μ + Z ) W U ′ ( W ( 1 + r + k ^ ( μ ) ( μ + Z ) ) ) ] = E [ ( μ + Z ) W U ′ ( W ~ ) ] = 0

We use this to calculate k̂'(0) by implicit differentiation. Differentiating the first order condition with respect to μ, then setting μ = 0:

0 = E [ ( μ + Z ) W ( k ^ ( μ ) W + k ^ ′ ( μ ) ( μ + Z ) W ) U ” ( W ~ ) + W U ′ ( W ~ ) ] = E [ Z 2 W 2 k ^ ′ ( 0 ) U ” ( W 1 ) + W U ′ ( W 1 ) ] = E [ Z 2 W k ^ ′ ( 0 ) U ” ( W 1 ) + U ′ ( W 1 ) ] = σ 2 W k ^ ′ ( 0 ) U ” ( W 1 ) + U ′ ( W 1 ) k ^ ′ ( 0 ) = − U ′ ( W 1 ) σ 2 W U ” ( W 1 ) μ k ^ ′ ( 0 ) = μ R ( W ) σ 2

which is the same result we found in Section 3.1. In this case, we see that small means a first order Taylor expansion of k̂ is sufficient, i.e. μ is close to 0.


Further Reading and References

  • Paul Samuelson. (1969). “Lifetime Portfolio Selection by Dynamic Stochastic Programming”, The Review of Economics and Statistics, 51 (3).
  • Robert Merton. (1969). “Lifetime Portfolio Selection under Uncertainty: The Continuous-Time Case”, The Review of Economics and Statistics, 51 (3).
  • Jonathan Ingersoll. (1987). Theory of Financial Decision Making, Rowman & Littlefield.
  • Tomas Bjork. (1998) Arbitrage Theory in Continuous Time, Oxford University Press.

  1. This not is not an offer or solicitation to invest. Past returns are not indicative of future performance.
  2. Most authors use μ to denote the risky asset expected return and μ – r do denote the expected excess return, we find it less cumbersome to use μ for the excess risky asset return.
  3. This is true if the risky asset price follows a Geometric Brownian Motion.
  4. It is interesting to note that if γ = 1 (i.e. log utility), then the independence assumption can be dropped. This follows from the fact that the log of a product is the sum of the logs and the linearity of Expectation.
  5. See Theory of Financial Decision Making Part I, Chapter 8 for a more in depth treatment
  6. There is another approach called the Martingale Method which is also beyond the scope of this note; see Arbitrage Theory in Continuous Time.
  7. See Arbitrage Theory in Continuous Time Part IV for an exposition.
Read More

The Chemistry of 10% More

November 9, 2017

Uncategorized

The Chemistry of 10% More

By Victor Haghani 1

A few nights ago at the dinner table, I was ritually lamenting my unsuccessful efforts to lose a few pounds. My son Mark, now a high school senior, posed a good question – the kind I’ve learned to think about very carefully. If I lost 5 pounds of fat, where would it go? After letting me flail around a bit, he gave me the answer (hint: neither heat release nor bathroom visits are correct answers.2). Turns out we lose most of the weight by exhaling CO2. Simple. So, if you want to lose weight, don’t hold your breath!

This reminded me of a similarly “simple” question my friend Larry Hilibrand asked me about 10 years ago that also involved a “conservation of mass” principle. In a world with two assets – cash and equities – how can investors in aggregate increase their allocation to equities, given that for every buyer there has to be a seller (assuming companies aren’t issuing or buying back shares), and so cash can’t come into or leave the market?

I recall some floundering back then too, before the obvious came into view: the price of equities must go up. I thought I was off the hook, but there was one more question: Just how much do equities need to go up if investors in aggregate have 50% in equities and decide they want to increase that to 60%? I encourage you to give it a quick guess before doing the math.

There are a few ways to figure it out, one being to consider that an investor holding $50 in equities and $50 in cash will need the original holding of cash to become 40% of the new total portfolio value after equities go up in price. So, her new total portfolio value would need to be $50 / 0.4 = $125 . The only way for that to happen is for equities to rise in value to be worth $75. That’s a 50% gain, as you can see in the chart below. I don’t know about you, but my guess was a lot lower than 50%.

And, notice from the chart below that a 50% starting allocation is where this impact is the smallest (see the footnote below if you want a formula).3

Of course, the same math holds if investors all want to reduce their allocation to equities. From the same starting allocation of 50%, equities would need to drop by 33.3% for everyone to wind up with a 40/60 equity/cash allocation.4

In practice, the impact would be dampened by companies issuing or buying back shares, although it’s worth noting that IPOs in the US have averaged only 0.25% of total US market capitalization annually over the past ten years.5

However, in the short term, return-chasing investors, who estimate future expected returns based on historical performance, may have the opposite impact and actually put more fuel on the fire.6

This paradigm is a gross over-simplification of reality, but I’ve found it useful, and unexpected.


  1. Victor is the Founder and CIO of Elm Partners. Past returns are not indicative of future performance. This not is not an offer or solicitation to invest.
    A big thank you to Larry Hilibrand, who provided the idea for this note, as well as valuable comments in drafting it, and to James White for helping me make it shorter.
  2. Heat release would imply a bodily nuclear reaction, and the stuff that leaves us was largely never part of us to lose. And here’s a YouTube video titled Mathematics of Weight Loss if you want a more complete answer. Also, thank you to Lasse Pedersen who pointed me to the even more beautiful chemistry and physics of the question, “Where does a tree get its mass from?” In this clip from BBC series Fun to Imagine (1983), the late Richard Feynman, with his infectious exuberance, explains
    “People look at a tree and think it comes out of the ground, that plants grow out of the ground…if you ask, where does the substance [of the tree] come from? You find out…trees come out of the air!”
  3. The equity price change needed to change aggregate asset allocation by ∆% from a starting allocation of E0 is: ((1 – E0)/(1 – (E0 + ∆)) – (1 – E0)) / E0 – 1 ≈ for small ∆ , and min occurs at E0 = 50% .
  4. The answer is symmetric in log-space, as ln(1.5) = -ln(0.67) , suggesting a 33.3% drop in is the same amount as a 50% gain in log space.
  5. “Value of initial public offerings (IPOs) in the United States from 2000 to 2021”.
  6. At Elm, we strongly believe in estimating equity returns looking forward, rather than backwards, as discussed in this video, The Most Important Number You Won’t Find in the WSJ. Even investors who use a long historical window for their estimate will find that a doubling of the market will increase the historical 20-year return by over 3.5% per annum.
Read More

What Didn’t Happen on Black Monday 1987

October 17, 2017

Uncategorized

What Didn’t Happen on Black Monday 1987

By Victor Haghani 1

“…I don’t know.”

That was the only answer I had to each of my dad’s questions. It was very dark as we wandered the streets near my Upper East Side apartment in the early hours of October 20th, 1987, discussing the events of the prior day – Black Monday.

I was 25 years old and had been on John Meriwether’s Arb Desk for just over a year, following two years working in Salomon’s Bond Portfolio Analysis Group. My memories from the day the US stock market dropped 22% didn’t seem very noteworthy when my friend Rich Dewey asked to interview me for an article, Black Monday Revisited, he was writing for Bloomberg to mark its 30th anniversary. After all, compared to the other people Rich was interviewing—Paul Tudor Jones, Howard Marks, Stanley Druckenmiller, Ed Thorp and my former Salomon colleagues Eric Rosenfeld and Michael Lewis – what could I add?

I’ve been thinking that my recollection of some things that didn’t happen might be interesting. First of all, Salomon’s Arb Group didn’t do a single trade on October 19th, or pretty much for that whole week. It wasn’t that we didn’t see great opportunities, but rather that it was clear to everyone that this was a crisis and a time to preserve capital. Salomon was first and foremost a financial intermediary. Our capital was limited, and it would be sorely needed for providing liquidity to clients and projecting financial strength to all our counterparties. Salomon had about $3.5 billion of capital supporting a balance sheet of $100 billion.

We borrowed money to finance our long positions in bonds primarily through the repo market, but we also had to borrow securities to support our short positions. When we shorted a bond, such as the 9.25% of 2/15/2016 that we were running a big position in at the time, we needed to borrow that specific bond from someone who held it in order to make good on the sale. We could borrow money from anyone– a client, a bank or as a last resort even the Fed. But the only party who could lend you a security was someone who owned it free and clear, and hadn’t lent it already. And the lending market was mostly overnight, so we had to roll every day.2

Our biggest fear was that clients who were lending us the securities we were short might ask for them back, forcing us to cover our shorts and unwind our trades. It’s almost axiomatic that when you’re forced to unwind a trade, you lose money on it. With these concerns in mind, John holed us up in a room off the desk, and we put our efforts into methodically triaging our portfolio. We kept the trades we’d be most able to hold to convergence, and cut the ones that were least defensible. We had been having a solidly profitable year up until October, but gave back most of our gains in the few days around Black Monday. The firm’s decision to let us keep our best positions was rewarded with a very profitable 1988.

The defensive orientation on our desk was echoed across the whole firm, and it was probably the same at Goldman, Bankers Trust and all the other trading houses on the Street. As a highly levered financial firm, dependent on short-term funding of our balance sheet, our crisis mentality was about surviving the storm, not trying to profit from it. While the most popular 1987 crash stories celebrate the trading acumen of the likes of Tudor Jones, Druckenmiller or Taleb, the less publicized story of how so much capital was constrained or frozen—an essential ingredient in all market panics– might be the more important takeaway.

Another thing that I don’t remember happening was an economic depression following the stock market crash. Well, I guess that’s because it didn’t! It was supposed to though, just as the Great Depression followed the Black Tuesday of October 1929. In fact, in December 1987, 33 prominent economists (5 with Nobel prizes) issued a statement predicting that “the next few years could be the most troubled since the 1930s.”3 The fact that we moved forward with barely a blip to the real economy makes us feel that October’s stock market crash was bogus, a market move that had nothing to do with fundamentals. But it didn’t have to turn out that way. It’s important to remember what didn’t happen: an alternative future in which the Fed didn’t act as it did (would Volcker have reacted as Greenspan did?), the stock markets fell even further, financial firms started failing, and we got a long and deep recession.4

Could it happen again? Of course it could. And it has. Extreme market moves of the magnitude of Black Monday’s 20-times normal daily move have occurred periodically since then, just not in the US equity market or on the one-day time scale. For example, two years ago, the Swiss Franc put in a 40 times daily upward move against the Euro when the Swiss National Bank suddenly abandoned the policy of capping its value.5 If we look at more arcane, but still important, markets, we find further examples, such as changes in swap spreads or long-dated equity volatility in October 1998 (what is it with October anyway?), or diversified equity momentum trading strategies that lost close to 90% in 2009, or the melt-down of equity quant strategies the week of August 6th, 2007.

Plus ça change. Humans, with all our behavioral foibles, are still important players in the markets. While the presumed cause of Black Monday, Portfolio Insurance,6 is now defunct, it has been superseded by vast amounts of capital dedicated to algorithmic trading or trend-following, both strategies expressly designed to make money, not to stabilize markets. Risk management systems based on VaR (recent volatility of positions) or a tight stop-loss discipline are inherently destabilizing too. And then we have the Volcker Rule and other post-financial crisis regulatory changes, which have the unintended consequence of dramatically reducing the ability and incentive of the banks to provide liquidity in normal market conditions, let alone in a crisis. What do we have against all this? More circuit-breakers and a tradition of Central Bank intervention to stabilize markets in every crisis since Black Monday.7 Hopefully, they will continue to do so, but we should be prepared for when they don’t.

You may wonder, how did this prepare me for another tumultuous October, eleven years later, when I was a partner at LTCM? Stay tuned, its 20th anniversary is just twelve months away.


  1. Victor is the founder and CIO of Elm Partners. Past returns are not indicative of future performance. This not is not an offer or solicitation to invest.
  2. We always appreciated the important role played by the repo desk, but during the crisis these guys, who sat at the edge of the trading floor and who at other firms on the Street were treated as back-office operational staff, earned hero status and a mainstream position on the trading floor thereafter.
  3. “Group of 7, Meet the Group of 33”
  4. One more thing that didn’t happen in those tumultuous days was our normal afternoon Liar’s Poker session, but it did come back before long.
  5. January 15, 2015. Within hours it had recovered to being down ‘just’ 20%. Other examples: over a longer horizon, we have the equity market sell-off from late 2007 to early 2009, which was a 60% drop (88% in lognormal terms), a very unlikely move given the typical annual variability of equity markets. In 2013 the US bond market experienced its “Taper Tantrum” when 10-year US Treasury rates almost doubled, to 3%, over a 6-month period. There was also the ‘flash rally’ of October 15, 2014, when the 10-year Treasury rate declined by 0.37% in about 30 minutes, before quickly going most of the way back to where it had been.
  6. For my younger readers, Portfolio Insurance. The Brady Commission Report concluded that Portfolio Insurance was the proximate cause of the stock market decline. It was estimated (here) that $60-90 BB of assets were following a portfolio insurance strategy, and they needed to sell $10-15 BB of equities on Black Monday. To put this in perspective, the $ value of US equity trading volume is 15x higher and the market capitalization of US publicly listed companies is 10x larger today than it was 30 years ago (World Bank).
  7. Except for the Swiss, who didn’t get the memo.
Read More

Some Clarity on Risk Parity

October 16, 2017

Uncategorized

Some Clarity on Risk Parity

By Victor Haghani and James White 1
This post was first published on Bloomberg Prophets.

Google “risk parity” and you’ll see a grab bag of conflicting results: articles and posts trying to explain what it means, why it reduces stock-market volatility, why it increases stock-market volatility, why it’s less risky than a traditional portfolio, or why it’s more risky, among other things.2 We’ll try to cut through this confusion to show that risk parity and traditional portfolios are closely related in philosophy.

Risk parity is all about how an investor allocates risk, not capital, typically with the use of leverage and with the idea that an equal risk allocation to various asset classes increases the benefits of diversification. Risk parity and traditional portfolios are usually presented as being philosophically miles apart, and hard to compare or analyze side by side except by looking at the historical record, such as in the chart below. The problem, though, is that financial market history is limited in that it reflects only a very specific set of historical conditions and, as we know, past performance isn’t indicative of future results.

What’s an investor to think? Both types of portfolios come out of the same theory of portfolio construction, but with different sets of basic starting assumptions. The theoretical toolkit we’re talking about here is the “Optimal Expected Utility” framework applied to financial markets by Paul Samuelson and his protégé Robert Merton starting in the 1960s.3 Their work helped them both garner a Nobel Prize, and produced a set of practical tools for determining how much of one’s wealth should be allocated to different investments with the understanding that the future is uncertain. The tools are primarily based on an investment’s expected excess return, volatility, and the investor’s personal level of risk-aversion.4

The basic idea is that an investor’s utility doesn’t keep going up as investment size – and thus risk – increases, but instead there’s an optimal investment size that maximizes expected utility given one’s personal level of risk-aversion. This simple idea leads to some powerful results. With the five assumptions below, the utility toolkit tells us that the portfolio that maximizes expected utility is the one with the highest Sharpe ratio – a common measure of risk-adjusted returns5 – levered or de-levered to an optimal level of risk.

  • Assumptions for Optimality of Basic Risk Parity Portfolio:
    • All assets follow a random walk and are continually tradeable
    • All assets have equal pair-wise correlation with each other
    • All assets have equal Sharpe ratio over a long horizon6
    • Unlimited leverage is available at the risk-free rate
    • No fees, transactions costs, or other drags on return

To build that portfolio with uncorrelated, equal-Sharpe ratio assets, we’d hold an amount of each asset that is inversely proportional to its volatility, resulting in each asset contributing an equal amount of risk to the portfolio. This is why it’s called “risk parity” investing.7 Using some stylized risk/return assumptions, the table below shows how this works with two risky assets for an investor with a “typical” amount of risk aversion.8 The first three rows show arbitrary allocations, and the final three rows show utility-optimal allocations corresponding to the given portfolio assumptions.

Portfolio Stocks Bonds Expected Excess Return Risk Sharpe
Ratio
Risk-Adjusted Return
Stocks 100% 0% 4.0% 16.0% 0.25 0.2%
Bonds 0% 100% 1.0% 4.0% 0.25 0.8%
Traditional 60% 40% 2.8% 9.7% 0.29 1.4%
Risk Parity
Unlevered
20% 80% 1.6% 4.5% 0.35 1.3%
Risk Parity
Levered
50% 200% 4.1% 11.6% 0.35 2.1%
Risk Parity + 0.6% Extra Borrow Cost 45% 85% 2.5% 8.0% 0.31 1.5%

The Merton toolkit suggests our investor would optimally want to own, via leverage, $250 of the equal-risk portfolio for every $100 of savings, resulting in the “Risk Parity Levered” portfolio in the table. The performance of this portfolio is quite a bit better on a risk-adjusted basis, 0.7 percent a year to be exact, than the traditional 60/40 stock/bond portfolio.9

We made some strong assumptions to get this result. Let’s see what happens when we loosen just one of them and assume that leverage isn’t available at the risk-free rate, but at a rate 0.6 percent higher? That cuts the risk-adjusted return of the optimal portfolio to 1.5 percent per year. This portfolio is very close, both in risk-adjusted return and Sharpe ratio, to the traditional 60/40 portfolio, making the traditional portfolio functionally optimal given the assumptions. We get the same result if we assume the investor doesn’t want to use leverage, regardless of the rate. Changing this assumption isn’t some abstract technicality: Leverage in real markets is not freely available at all times or at consistent rates, and there are many reasons an investor may choose to eschew leverage.10

The portfolios we see here represent two ends of a spectrum – but the range is surprisingly narrow. In our admittedly stylized two-asset example, only 0.7 percent per year of risk-adjusted return separates the fully-levered risk parity portfolio from the unlevered traditional one, which gives a sense for the level of fees, trading costs and extra borrowing expense a risk parity strategy could plausibly support.11 If the five assumptions above seem reasonable, risk parity portfolios may make sense for you, but if not, a more traditional portfolio may be a better fit, and is just as consistent with good finance theory.12


  1. Victor is the founder and CEO of Elm Partners, a HNW Robo-investment manager. James works with Elm Partners in addition to pursuing his own research and investment interests.
  2. There are a number of overviews of risk parity online. Here’s one: “Understanding Risk Parity”.
  3. 3 One of the most seminal papers is Robert C. Merton’s “Lifetime Portfolio Selection under Uncertainty: the Continuous-Time Case.” The Review of Economics and Statistics (51), 1969.
  4. A core result is that, for one risky asset following a random walk, the optimal investment size is where (μ – r)/(n * σ2) is the asset’s excess return, σ its volatility, and n the investor’s coefficient of risk-aversion. We often refer to this result as the “Merton Rule.”
  5. Sharpe Ratio is the ratio of an asset’s excess return to volatility: SR = (μ – r)/ σ)
  6. This is consistent with both the historical record and what many equilibrium models would predict.
  7. Assuming risk and volatility are interchangeable for the purposes of this discussion. For uncorrelated assets with different Sharpe ratios, the solution is to scale each proportional to Sharpe ratio and inversely proportional to risk.
  8. We’ll define typical here as that degree of risk aversion that would maximize expected utility by investing 100 percent of savings in a stock/bond portfolio with a 60/40 mix. With the numbers in our illustration, this implies a coefficient of risk aversion in the Merton model of 3. For readers familiar with the Kelly Criterion, this means our investor is 3 times as risk averse as a Kelly bettor. Also, typical risk parity implementations include four or more assets, including commodities and credit.
  9. Risk-adjusted return = Expected Return – ½ * σ2 * n , where n is the coefficient of risk aversion.
  10. A few reasons that come to mind: 1) real markets may not follow pure random walks but can also gap, 2) leverage may not be easily adjusted once set, 3) terms other than rate may not be attractive, or 4) whenever the investor hears the word “leverage” he or she suffers painful flashbacks.
  11. And they are separated by only 0.06 in terms of Sharpe ratio, although some back-tests suggest a difference of close to 0.2. As discussed in “What’s Past is Not Prologue,” we cannot rely on historical data on its own to support or reject the existence of this amount of difference in Sharpe ratio.
  12. The authors would like to thank Larry Hilibrand and Vlad Ragulin for their help.
Read More

When (if Ever) Has it Paid to Wait for a Stock Market Correction?

August 7, 2017

Uncategorized

When (if Ever) Has it Paid to Wait for a Stock Market Correction?

By Victor Haghani and James White 1

A few weeks ago, one of our investors told us she had some cash that she wanted to add to her account with us, but that she felt the market was at such a high level that she should wait for a market correction before putting it to work. She was confident that a market drop would happen before long, and she saw little downside in waiting for it. We’ve been hearing this general sentiment from many of our investors, and potential investors, and thought it timely to share some thoughts and analysis on this topic.

While we share our investor’s concern about the current valuation of the US market, we told her that in fact the downside to waiting is more than meets the eye. While the probability of a correction at some point is indeed quite high, the correction may happen from a much higher market level, or may take so long to happen that she stops waiting and gets “pushed in” at a higher level. In fact, under the assumption that the market follows a random walk, as long as you believe the expected return of the market is higher than what you’d earn on your cash, the expected opportunity cost of waiting for a correction exceeds the expected benefit of investing after the correction.

In practice, few investors believe markets efficiently follow a random walk, even though it’s a pillar of finance theory. Investor beliefs do tend to be strongly influenced by history though, so let’s look at Yale Professor Robert Shiller’s US stock market data going back to the late 1800s.2 Before describing what we found in this data, we should lay down a few caveats: 1) historical market behavior is not necessarily indicative of the future, 2) even 115 years of market data isn’t enough to draw statistically significant conclusions on many interesting questions, especially those involving relatively rare events, and 3) the US stock market is not the whole stock market, and the rest of the global market, which at Elm forms a significant fraction of client portfolios, looks a lot less frothy than the US market.

With these caveats in mind, the first question we ask the data is: during times when the market has been “expensive” what has been the average cost or benefit of waiting for a correction of 10% from the starting price level, rather than investing right away? We define “expensive” as times when the stock market had a CAPE (Cyclically-adjusted Price-Earnings Ratio) that was more than one standard deviation above its historical average level. While the CAPE of the US Stock Market is currently hovering around two standard deviations above average, as per our second caveat above, there aren’t enough equivalent periods in the historical record to make a statistically significant data analysis conditioned on the current state of the market. We focus on a comparison over a 3-year horizon, which we choose because the investor is unlikely to wait indefinitely for the hoped-for correction. The most salient findings from this analysis are:

  • From a given “expensive” starting point, there was a 56% probability that the market had a 10% correction within 3 years, waiting for which would result in about a 10% return benefit vs. having invested right away.3
  • In the 44% of cases where the correction doesn’t happen, there’s an average opportunity cost of about 30% – much higher than the average benefit.
  • Putting these together, the mean expected cost of waiting for a correction was about 8% versus investing right away.

We note that while the probability of a correction inside the horizon is somewhat likely, it’s far from certain – and that when it doesn’t happen, the expected opportunity cost is much higher than the expected benefit. We suspect that the perception that waiting for a correction is a good strategy arises primarily for three reasons. First, while a correction occurring is indeed more likely than not, investors may confuse the chance of a correction from peak-to-trough with the lower chance of a correction from a fixed price level. For example, the historical probability of a 10% correction happening any time during a 3-year window is 88%, significantly higher than the 56% occurrence of that correction from the market level at the start of the period. Second, the cost of waiting and not achieving the correction is a “hidden” opportunity cost, and we humans have a well-documented bias to underweight opportunity costs relative to realized costs. Finally, investors may believe they can wait indefinitely for the correction to happen, but in practice few investors have that sort of staying power. In fact, as we’ll see below, the longer you’re able to (stubbornly) wait for the correction, the greater the average opportunity cost you would have suffered.

We’ve made some very specific assumptions here: that the investor is waiting for a 10% correction, has a 3-year horizon, etc. So, we repeated this historical analysis with correction ranges from 1% to 10%, horizons of 1-year and 5-years, and with different criteria for what makes the market look “expensive.” 4 Each point in the chart below represents a combination of these four assumptions. You can see that across all scenarios there has been a material cost for waiting. The longer the horizon that you’d have been willing to wait for the correction to occur (the red points represent a 5-year horizon), the higher the average cost.

Now shifting focus from the historical record to looking forward, it’s true that the lower one’s expectation of the stock market return, the lower the expected cost of waiting for a correction. If you believe the stock market has a negative expected return to a particular horizon, then waiting for a correction to invest makes sense. However, at least as far as the historical record for the US stock market goes, higher market valuations are consistent with lower prospective long-term returns, but not negative expected returns.

We’re not suggesting that looking at the historical record closes the book on the question of whether or not to wait for a correction. As you know, we firmly believe in the axiom that past returns are not indicative of future returns – and there are many other personal and circumstantial factors an investor should consider before deciding whether entering the market now is the right thing to do. However, we have the impression that some investors, like our client who got us thinking about this question in the first place, are waiting for a correction specifically because they think that history shows that waiting pays. We hope it’s been useful to show that history, for what it’s worth, doesn’t support that view.

For further related reading, you may enjoy our recent note: What’s the Best Way to Get Invested in the Market?


  1. Victor is the Founder and CIO of Elm Partners, and James is Elm’s CEO. This not is not an offer or solicitation to invest, nor should this be construed in any way as tax advice. Past returns are not indicative of future performance.
    Thanks to Larry Hilibrand, Vladimir Ragulin, Chi-fu Huang, Andy Morton and John Glazer for their input.
  2. Professor Shiller’s S&P data is presented as monthly averages, which we felt might be more relevant in exploring the topic of this note.
  3. The cost or benefit is calculated as the return difference over the full 3 years between being fully invested for the entire period versus waiting for the 10% correction and re-entering the market when and if it occurs. Carry from dividends and cash yield are both included in the differential return calculation.
  4. In addition to our CAPE-based definition of “expensive” we also looked at waiting for a correction from times when the market was at an all-time high at the start of the period.
Read More

What does family planning have to do with investing?

November 22, 2016

Uncategorized

What does family planning have to do with investing?

We recently launched a new sub-section of our website, Elm Labs, as a place to experiment with interactive research and educational tools.

We launched with a simple (yet tricky) family planning puzzle – and so far, we’ve had answers from over 1000 people. Assuming that there was no self-selection bias in terms of who answered it, it’s a fair estimation that the vast majority of them were finance professionals in the 45 to 65 year old age group, with about half having graduate degrees, and about two dozen being current or former professors in finance.

Of these, 44% got the right answer, and of those who didn’t, 54% thought that the imbalance in the population would favor girls over boys (which is what I would have guessed too, as that seemed to be the objective of the family planning strategy to begin with).

I received about two dozen emails with comments (actually, mostly were objections). The most common remark was that the (unrealistic) assumption that a couple could have as many children as needed in order to have a girl was what made the expected number of girls and boys equal. Not true: if we assume, for example, that a family stops trying after 5 boys in a row, that doesn’t change the answer that the expected number of girls and boys is still equal. I pasted a table below which shows that, just in case you’re wondering.

One friend (a former stats professor) pointed me to an area of probability theory that probes questions involving stopping rules much more deeply (see Doob’s optional sampling theorem for a taste). As is often the case, simple problems can reveal layers and layers of more profound questions when un-ravelled by inquisitive minds.

Stay tuned for more on the question of stopping rules. In particular, we’ve been working on some research that explores the question of whether a stop-loss approach to investing (more fully: cut your losses early and let your profits run) in and of itself has been a profitable investing and trading approach, and if so, why.

Girls Boys Prob E(Girls) E(Boys)
1 0 0.5 0.5 0
1 1 0.25 0.25 0.25
1 2 0.125 0.125 0.25
1 3 0.0625 0.0625 0.1875
1 4 0.03125 0.03125 0.125
0 5 0.03125 0 0.15625
Sum 1 0.96875 0.96875
Read More