Elm Wealth Research

Posts by:

James White

Financial Times full page article about Elm (and Victor)

May 8, 2017

In the News

Financial Times full page article about Elm (and Victor)

“Mr. Haghani has been on an intellectual journey that, in many ways, mirrors the evolution and central debates of the modern investment industry…His mission is supported by research he and his team at Elm publish to illustrate the common ways in which investors damage their own interests by being too active.”

Read the full article below, or on FT.com.

Read More

What’s Past is NOT Prologue

April 11, 2017

Featured Insights

What’s Past is NOT Prologue

By James White, Jeff Rosenbluth and Victor Haghani 1

Thank you to the 702 people who read and interacted with our note exploring when past returns are indicative of future returns. We received so much thought-provoking feedback that we decided a follow-up note was in order.

We started by asking readers to guess how many flips they’d want to see in order to be able to discern, with 95% confidence, a fair coin from a coin biased 60% to land on heads.2

Below is a histogram of the guesses.

The median guess was 40 flips. While lower than the full-credit answer of 143, it does show that our readers appreciate it takes a really long time to identify an investment with this kind of risk/reward simply by track record. In Appendix I, we include the calculation used to arrive at 143.3. Our readers are a pretty mathematical bunch, and we’re sure that if they took their time to calculate an answer, rather than giving a quick guess as we requested, most would have come close to the correct answer. But the point of the exercise was to illustrate how when we are thinking fast, we tend to put too much weight on small samples: a full 30% of respondents, the single largest bucket, thought 10 flips or less was sufficient. This built-in tendency to overweight small samples can easily lead us to ignore the dictum that “past performance is not indicative of future results.”

Our readers generally agreed that a 60/40 coin would represent an attractive investment opportunity. At a rate of one flip per year, this would equate to an investment with a Sharpe Ratio of 0.2, roughly comparable to most broad public markets.4 In fact, a few suggested that 95% was too high a confidence test given the attractiveness of the opportunity. They remarked that they’d be happy to invest half their money on each coin, and not worry about figuring out which was which. This comment highlights the fact that in the two-coin problem we presented, you’ve got a 50% chance of picking the right coin without learning anything through flipping them. However, if we make the problem more realistic by asking how many flips are needed to discern a 60/40 biased coin from 3 fair coins, the number of flips required jumps to 220 (see Appendix I). So even if we had a lower confidence test but with more potential coins, as would be likely in the real world, it still takes a long time to figure out which is the good coin.

Perhaps the most thought-provoking suggestion arising from our original note came from our friend Andy Morton. He proposed a more realistic setup of the problem wherein you can invest in 100 active fund managers, but only 15% are expected to generate a post-fee return of 1% a year in excess of their benchmark, while the other 85% are expected to lose 1% a year vs. their benchmark after expenses. We’ll assume that each has an annual risk vs. the benchmark of 10%, which makes the outperformance of the good funds vs. the bad funds similar in Sharpe Ratio to a 60/40 coin vs. a fair coin. If anything, this may be optimistic given S&P Dow Jones reports that only 1 in 10 active managers outperformed their benchmarks over relatively short horizons, while with our assumptions, just under 50% of funds would outperform their benchmark each year.5

Let’s explore the ramifications of “chasing” returns with this setup. You start off by putting 1% of your portfolio into each fund, because you don’t yet have data to tell them apart. Your starting expected return is -0.7% p.a. versus the benchmark (85% * -1% + 15% * 1% ). Each year you move more of your portfolio to the funds that have been doing well, using a Bayesian update of the probability of each fund being one of the good ones, and that helps your expected return improve from the starting point (see Appendix II for more detail). Alas, even by the end of 5 years – a reasonable “lookback” window for real-world fund evaluation – the expected return versus the benchmark on your portfolio will still be -0.66%. Extending out to 10 years doesn’t help much either – you only improve to -0.6% expected return.6

It’s generally believed that Warren Buffett-like investors are very rare – much rarer than 1 in 1,000 – but regardless let’s say all 15 of our “good” fund managers are like Buffet and produce an excess Sharpe Ratio of 0.45 (roughly Berkshire’s Sharpe Ratio of excess returns versus the S&P 500 since 1980). In this case, after 5 years we’d still only be breaking even across the whole portfolio, and would only have about 50% of our capital allocated to the 15 Warren Buffetts! We can clearly see that in the real world, where Buffetts are much rarer, 5 to 10 years of track record just doesn’t tell us that much when dealing with a portfolio composed mostly of mere mortals.

As we discussed in our note from six months ago, “What I learned from my daughter (about investing)”, our attempt to build a portfolio of rare investment gems – a clutch of young Buffetts – also has to contend with the drag of the false positive. Even if we can identify a good from a not-so-good investment with 90% accuracy, if the good ones represent just 5% of the possible investments, the best we can expect is a portfolio comprised of 32% good ones and 68% not-so-good ones. The cognitive bias that makes it hard to see this result is called base-rate bias.

In the real world, any kind of return-chasing strategy will also face additional headwinds, amongst them the fact that when an investment strategy was truly extraordinary in the past, its very discovery will diminish future performance (or even make future returns negative), as more capital is drawn towards it. This may be the the prime reason why investor returns (i.e. dollar-weighted returns) tend to be significantly below fund returns (i.e. time-weighted returns).7 This doesn’t happen with coin flips.

Another comment worth sharing, from the marketing capo of a well-established hedge fund group, was, “I don’t know anyone who would look at a fund investment without looking at the past track record.” We think the reason for this, and perhaps the central problem with investing in active fund managers, is that there is virtually no other way to develop a forward-looking expectation other than to look at the past track record. Unfortunately, as we’ve illustrated with our two-coin example, there is very limited statistical power in the typical historical return dataset which is often limited to 3-5 years due to issues around manager turnover.

Perhaps when a fund picker says, “We only will look at funds that have at least a 3-year track record,” he is thinking that if he observes the funds on a weekly basis, that gives him a lot more data points to draw upon to reach a good level of confidence. Unfortunately, it doesn’t help because when we observe the managers on a more frequent basis, it’s requiring us to discern a coin that is closer to 50/50, which takes exactly the extra flips you get by sampling more frequently. Bob Merton pointed this out in a 1980 paper titled, “On Estimating the Expected Return on the Market,” “…nothing is gained in terms of accuracy of the expected return estimate by choosing finer observation intervals for the returns.” 8

One more remark we think worth sharing was, “An investment in a fund is always about the manager.” That at least recognizes the limited value of past returns, especially in isolation and over typical, relatively short horizons. The problem is it’s very hard to separate our qualitative evaluation of a manager’s character from our knowledge of their track record – most managers with an apparently sterling, but often short-term, track record also present as highly confident and competent people, in a way which might not be the case with a different frame.

Once you accept the very low discriminating power of past returns in most practical investing contexts, what are you to do? Of course, you’ll want to look at all the information at your disposal regarding any investment, and track record should be part of the due diligence of any investment. If it’s particularly long and distinguished, it might have some impact on your forecast, when used together with other factors. It just shouldn’t be the primary or isolated driver of your return forecast. And if you’re ever unsure whether you’re being unduly influenced by a good track record over a normal (relatively short) time horizon, ask yourself whether you’d still make the investment if the return had been half as bad as it was good. If you answer yes, you’re probably giving about the right weight to track record in your decision process.

All this discussion on the low value of historical returns in most investment contexts may leave you feeling either depressed or liberated, depending on your perspective and occupation. If you’re an investor, it’s really not so bad: just stick to investments where you don’t need to rely on the track record to make your decision. That leaves almost all of the direct, cost-efficient, investible universe open to you– bonds, equities, real estate and any strategy where you can produce a reasonable forward-looking return estimate without relying on past returns.


Appendix I: Flipping a biased coin and one or more fair coins

We would like to calculate how many flips of 2 coins, one biased and one fair we need to be 95% confident that the coin with more heads is the biased one. Above, we assumed that the biased coin has probability of heads of 0.6, here we will be a bit more general and represent this probability by p . Since each coin has a binomial distributions and the coins are assumed independent, the joint probability mass function of the two coins after n flips is the product of two binomial probability mass functions. Denote by Q(n,k,j) the probability of k heads for the biased coin and j heads for the fair coin after n flips, thus

  Q(k,j,n) = (nk)  pk  (1 – p)n – k  (nj)12j12n – j

This allows us to calculate the probability that the coin with more heads is biased as

  1 2n Σ j < k k ≤ n ( n k ) ( n j )  pk  (1 – p)n – k

Choosing n to be 143 and p = 0.6 , we obtain 95.01%.

If we have more than one fair coin, the analysis is very similar. We calculate the joint probability mass function for the biased coin and the fair coins simply by multiplying together the individual probability mass functions. Then we sum over all of the cases where the biased coin has more heads than the maximum heads of any fair coin. All though this is a closed form solution, the summation has way too many terms to handle by hand. Therefore, we enlist the aid of a computer to calculate it.


Appendix II: Bayesian Capital Allocation for 100 Funds

In our 100 fund example from above, we used a “return-chasing” capital-allocation rule based on Bayesian learning, and here we explain the logic behind that in more detail.

Bayes’ Theorem (sometimes also called Bayes’ Rule) provides a powerful tool for updating our beliefs about the world as new information is received. In the 100 funds example, for each fund we would have a level of belief about whether that fund is in the Good group or the Bad group. Before we have any data, our belief that a given fund is good would be at 15% for each fund. As we have incremental data about fund performance, we can use Bayes Theorem to update our level of belief for each fund, in a way that’s rigorous and consistent with the information-value contained by the new data. Specifically, we update our level of belief by the relative likelihood of seeing the new data if our hypothesis is true versus if it’s false. To give a simple example: if we get new data in, and that particular data is no more likely when our hypothesis is true versus when it’s false, then our level of belief stays the same. On the other hand, if we get in new data and that data is highly likely if our hypothesis is true, but highly unlikely if it’s false, then our level of belief will go up substantially.

Bayes’ Theorem: For a hypothesis H with prior probability P(H) and new data D with P(D) ≠ 0 then,

  P(H|D) = P(D|H)P(H) P(D)

Using the mechanics of Bayes’ Theorem and assuming that capital is allocated proportionally to our level of belief that a fund is in the Good group, we can derive a closed-form solution for the mathematical expectation of capital that would be allocated to the Good funds at each point in time as more fund data becomes known. We can also see in closed form the mathematical expectation of the updated probability that a Good fund is good and a Bad Fund is bad, at each point in time:

As we can see from the above chart, even with this fairly sophisticated update strategy, relatively short periods of time don’t meaningfully impact our probability estimate that a Good fund is Good, or that a Bad fund is Bad.


  1. James is Elm Partner’s CEO, and Victor is the Founder and CIO. Past returns are not indicative of future performance. This not is not an offer or solicitation to invest.
  2. Example inspired by: “Good and bad properties of the Kelly criterion,” MacLean, Thorp, Ziemba, 2010, page 8.
  3. Another version of the problem, posed and solved by Costas Kaplanis, lets the observer stop the flipping whenever desired. In this version of the problem, the solution involves using Bayes’ Theorem to update the probabilities regarding the identities of each coin. The result is that you usually need significantly less than 143 flips to be 95% confident you’ve identified the correct coin, but there is still a significant probability that you will need well in excess of 143 to get to that level of confidence. Stay tuned for more from Costas on this topic.
  4. We did have some respondents who suggested that they only invest in opportunities with a Sharpe Ratio of 1 or higher. In our experience, investments that appear to have such high Sharpe Ratios are closed to outsiders, implicitly bear significant, hidden negative tail risk, are ephemeral or, infrequently, are fraudulent. For a more academic treatment, see Harvey, Liu, Zhu, “…and the Cross-Section of Expected Returns,” (2015), on SSRN.
  5. See SPIVA Persistence Scorecard here for S&P Dow Jones report.
  6. We can make the example more realistic by introducing some mean-reversion into the picture in order to capture the effect that when an investment strategy truly was extraordinary in the past, its very discovery will tend to diminish future performance (or even make future returns negative), as more capital is drawn towards it.
     

    This doesn’t happen with coin flips. The result is that a moderate amount of this effect means that even if the good managers are generating 4% more return than the not-so-good managers, even after 50 years the expected return of your portfolio still won’t be positive.

     

    To give this result some context, this amount of excess return is similar in risk-adjusted terms to Warren Buffett’s past 40 year track record relative to the S&P 500. The amount of mean-reversion we introduced is such that each fund is expected over the coming year to make back or give up 20% of the return it lost or made relative to the benchmark and to its underlying expected return over the past five years.

     

    So, if a fund outperformed by 10% over the past five years, it would be expected to do 2% worse than its normal expected return over the next year. Of course, a strategy set up to explicitly take advantage of the mean reversion would do better, but that is the topic for another note.

  7. See our Elm paper on return-chasing here. Also, see Morningstar’s “Mind the Gap” notes and Dalbar’s annual investor behaviour reports.
  8. Merton, Robert, C., “On Estimating the Expected Return on the Market: An Exploratory Investigation,” Journal of Financial Economics 8 (1980) 323-361. See Appendix 1, pages 355-357. To see why you can’t get around this inconvenient truth, note that in continuous time, the number of years of observation, using the annual Sharpe Ratio (SR ) as an input is 2 * (1.645 / SR)2 , where 1.645 is the 95% cumulative probability level in a Normal distribution. If we sample more frequently, f times per year, the required number of periods goes to 2 * (1.645 / (SR√f))2 = 2f * (1.645 / SR)2 .
     

    So, we’ll need to observe f times as many periods as when we look annually, which is exactly how many more periods we get to observe by breaking the year into f intervals. Note this exact result depends on a strong assumption about the distribution returns are being drawn from.

Read More

What’s the Best Way to Get Invested in the Market?

March 8, 2017

Investing 101

What’s the Best Way to Get Invested in the Market?

By Victor Haghani and James White 1

A few friends recently asked me how to go about putting more of their savings into the stock market. Is it better to jump in all at once, to average-in over time, or to wait for a market correction? They’ve been held back by a litany of worries—the durability of this eight-year bull market, the distortionary effects of Quantitative Easing, technological disruption, populist political movements (Trump, Brexit, Le Pen), the potential dissolution of the Euro, pension deficits and demographic time-bombs, to name just a few.

I wondered if some of Elm’s recent research might offer fresh perspectives, beyond the useful axiom that “time in the market is more valuable than timing the market.”

Step 1: Determine your desired long-term allocation

One of my father’s favorite jokes was about the village simpleton, Mullah Nasruddin, who was out one sunny morning searching for something in front of his house. A passer-by asked him what he was doing. “Looking for my watch,” he said. After some time looking around, the perplexed onlooker asked him if he was sure that he lost it there, to which Mullah Nasruddin replied, “No, I lost it in my shed, but it’s too dark in there to find it.”

In that spirit, let’s think about stock market returns where we have the clearest view – in the long term. With a long enough horizon, equity returns primarily boil down to earnings and dividends, which are easier to predict than sentiment-driven changes in valuation. The cyclically-adjusted earnings yield of the global stock market today is about 5%, and the dividend yield is about 2.5%, which suggests a long-term expected compound real return of about 4 – 5%.2 This is consistent with a survey we recently conducted of 120 financially sophisticated friends-of-Elm who on average expected a compound return of 4.2% above inflation for US equities. Given the substantially cheaper valuation of non-US equities, I suspect the survey average would have been 5% if we’d posed the question for global equities.

Once you’ve decided on your forecast, the next step is to think about how much you’d like to have allocated to equities given that forecast. To do this, you’ll need to consider the risk of equities and your personal degree of risk aversion.3 We’ve recently circulated three short notes on this topic,4 but it’s also just fine to think about the problem more holistically and intuitively, and arrive at an optimal allocation that feels right without getting too technical. By way of illustration, I expect global equities to deliver a long-term compound expected return of about 5% above inflation, resulting in my personal optimal allocation to equities of about 70%.5

It’s important to realize this is not a binary decision; your optimal allocation will be higher or lower depending on your forecast of the return of equities, and only if your expected return is zero (or negative) would you want to have no equities at all.

Step 2: How much to worry about the short term?

Unfortunately, there’s a lot of evidence warning us that it’s very difficult to forecast short-term equity returns.6 It’s true that when equities are cheap, measured by any of the common valuation multiples, their returns over subsequent periods are higher than when they are expensive by those same measures. Although we’re already taking account of this in our long-term forecasts, some investors worry that in the short-term the performance is more exaggerated from the effect of valuation reverting to “fair” value.

While we don’t have enough historical data to draw precise conclusions, what we do have suggests that the commonly used valuation multiples don’t give us much extra predictable valuation change in the short term. For example, if we look back over the past 130 years of US equity market experience (using Yale Professor Robert Shiller’s dataset), when PEs were in the highest decile, subsequent one-year real returns were lower than average, but still had a mean positive annual return of 3.4%.7 This is not to say that you can’t have an expectation that equities are going to go down 10% next year, in which case you really shouldn’t own any equities at all. It’s just that if you do happen to think equities are going down next year, based primarily on equities being over-valued, it’s useful to know that history is not on your side.8

If your short-term forecast is different to your long-term one, it is the short-term one that should primarily drive your decision.9 For example, if I thought next year’s return of equities was 3% instead of 5%, then my desired allocation would be 50% rather than 70%.10

The Benefits (and Cost) of Wading in Gently

Once you determine the allocation that’s right for you, how should you go about getting there? Most people who ask this question probably already know that cold logic dictates that we should immediately make the adjustment and move our allocation to the optimal point. An incremental adjustment over time, i.e. “averaging-in,” is not financially optimal. However, averaging-in does provide a psychic benefit for those of us (most of us) that are predisposed to seeing the glass as half full when it comes to our past decisions. If the markets go up over the period that we’re averaging-in, then we can be glad we got some of our savings invested at the beginning, and if the markets go down, then we can be comforted that we didn’t invest the full amount right at the start.

For example, if an investor put money to work in the stock market over the course of a year in four equal allocations, there would be an 80% chance that at least one of those purchases would be in profit when looking back from the end of the year.11 Viewed from a slightly different angle, there’s only a roughly 1 in 3 chance that the lowest value of the stock market occurred right at the start of the averaging-in period.

When does ½ = ¾ ?

Building up your allocation to the desired level over time does leave some expected return on the table. But, if you subscribe to the expected utility framework that we’ve recently been writing about, you’ll find that the expected cost is quite low. This is because as we increase our allocation to equities, our expected gain goes up in proportion to our allocation, but the cost of risk grows faster than in proportion to our allocation. In the case of the most popular family of utility functions, it goes up quadratically, that is, in proportion to the allocation squared.12 In other words, we should require four times the compensation to bear 2 times the risk. This relationship between return and risk results in Expected Utility as a function of our allocation taking the shape of a parabola. As you can see in the chart, the curve flattens out as we approach the optimal point, which means that going the last part of the way to the optimal point doesn’t get us a lot of extra expected utility.

Thus, moving ½ way to our optimal allocation gets us ¾ of the value of going the whole way, and going ⅔ of the way to optimal gets us about 90% of the value of going the full distance (see chart again).

Recall also that at the optimal point, the risk-adjusted return of our investment is equal to half the gross expected return. Combining these two effects means that averaging-in to your optimal allocation costs relatively little in terms of risk-adjusted return. For example, the expected cost of averaging-in over a year from a 0% to a 50% allocation to equities has an expected risk-adjusted cost of only 0.40%. I suspect this is a cost many investors would find worth bearing to get the comfort of wading into the market gradually. Of course, the cost to you of adopting such an approach will depend on your desired allocation and the averaging-in program you choose.

Warning: not all averaging-in plans are created equal

The averaging-in we’ve been discussing is a disciplined program of investing a fraction of your savings into the market regularly, over a pre-defined period of time.13 Some investors are attracted to the idea of contingent averaging-in, where they plan to increase their allocation to equities only if they fall below some target price level. The problem here is that the investor is taking the risk of the market going up and never (or at least for generations) coming back down again to the target level, and thereby incurring a very significant opportunity cost. To illustrate, let’s consider an investor who adopts a plan of setting a target 10% below today’s market to take his allocation to equities from 0 to 50%. Assuming an expected compound return of 5% above inflation, and evaluating this to a 10-year horizon, there’s a roughly 1/3 chance that he never gets a chance to invest, and the result is he is incurring an expected, risk-adjusted opportunity loss of about 12% of his savings. This is the expected outcome. For an idea of a worst-case outcome, we need only think about an investor who in early 2009 set a target 10% below the level of the markets, and is still waiting for Godot while investors who averaged-in to the equity markets nearly tripled their investment.

Conclusion

Once you’ve decided on your target allocation to equities, perhaps with greater weight on the long-term return driven by earnings than on predicted short-term expected changes in sentiment, the next step is to decide how to get to that target. Choosing a plan is ultimately a matter of personal preference, as any approach other than moving straight to your optimal allocation is financially sub-optimal. However, some plans are less sub-optimal than others. For example, averaging-in over one year has quite a low expected cost. More efficient still would be to immediately move about 50% of the way to optimal (thereby getting 75% of the benefit) and then moving the rest of the way over the course of a year. At the other extreme, following a contingent plan that only goes into action when the price of equities drops by some threshold amount is a risky approach with a high expected cost. In the end, any plan is a good one if it overcomes the inertia and anxiety that hold us back from our chosen investment destination. And, once we’re in for a while, another powerful human trait, forgetfulness, will make us wonder why we spent as much time as we did contemplating the plunge.


  1. This not is not an offer or solicitation to invest, nor should this be construed in any way as tax advice. Past returns are not indicative of future performance.
    Thank you to Jeff Rosenbluth, Vlad Ragulin, Andy Morton, James W. White, Amir Mossanen, Arjun Krishnamachar, Josh Haghani and my colleagues at Elm Partners for their useful comments on this note.
  2. Many observers feel the US dividend yield is an under-estimate of the true dividend yield, in that US companies make heavy use of stock buybacks in lieu of traditional dividend payments. For more discussion of the expected return of equities, see our video: The most important number you won’t find in the Wall Street Journal.
  3. You also need to consider other investment opportunities that are different than equities. For this note, we will assume that the only two assets are equities and some risk-less asset (e.g. US T-bills).
  4. Here, here and here.
  5. Using our modified Merton rule, my 70% optimal allocation is consistent with an expected compound real return of 5%, risk of 18% and a coefficient of risk aversion of 3 (i.e. 1/3 the risk tolerance of a Kelly bettor).
  6. See, for example, Morningstar and Dalbar’s research showing that time-weighted returns have exceeded dollar-weighted returns of US mutual funds by several percent per year over long horizons.
  7. The dispersion around that mean estimate was 18%. There were 160 monthly data points in the richest decile of monthly CAPE. We used overlapping data which further reduces the precision of the forecast, as does our belief that these draws do not come from a stationary distribution.
  8. While it is a small dataset, it is interesting to note that in 111 out of the 160 months in this richest decile of CAPE, momentum as measured with reference to the one-year moving average was positive, and in those cases, the average return was 8.1% pa. In the 49 cases of negative momentum, the next year’s return was -7.3%, suggesting that momentum is more powerful a predictor than valuation in the short-term.
  9. Under assumptions of stocks following a random walk and an investor with constant relative risk aversion, the asset allocation decision is completely myopic (determined by the short-term expectation). In cases where equities are mean-reverting, or display momentum, the results are different. See work by Merton, Campbell or Kritzman for further discussion.
  10. Using the modified Merton rule again. See footnote 4, and substitute 3% for 5%.
  11. Calculated via simulation, using the risk and return assumptions previously stated. The four allocations would take place at the beginning of the year, and then every 3 months thereafter.
  12. The family known as constant relative risk aversion functions, or CRRA for short.
  13. There are several ways to average-in: depending on whether the focus is dollars, fraction of savings or number of shares. In this note, we use fraction of savings as the variable. The other approaches to averaging-in do not materially change the conclusions we reach.
Read More

What Our Market Return Forecasts Really Mean: Equity Convexity and Investment Sizing

February 14, 2017

Risk and Return

What Our Market Return Forecasts Really Mean: Equity Convexity and Investment Sizing

By Victor Haghani and James White 1

“The key issue in investments is estimating expected return.”
  – Fischer Black

Introduction

You’re probably familiar, at least in passing, with the “convexity” of long-term bonds – i.e. that yields dropping 1% produce a bigger price move than yields rising 1%. A significant amount of brainpower has gone into understanding all the ramifications of this convexity in the fixed income markets, and the various issues and opportunities that arise are now very well understood. Equities, on the other hand, aren’t typically regarded as convex instruments.2 We tend to think of equities directly in terms of their price, rather than their “yield” as we do with bonds, and often think primarily about their annual return, which moves lockstep with price. But, as we discuss below, equities do have important convexity properties, and they tie into two themes we think deserve more attention: how investors think about long-term returns, and how to properly size portfolios and investments. Our story of how equity convexity, return forecasts, and investment sizing all tie together starts in the late 1960s with a remarkable result from Robert C. Merton.

The Merton Rule

A typical first step in building an investment portfolio is to forecast long-term returns, and identify the basket of investments with the most attractive return relative to their risk.3 With this accomplished, we still need to decide what percentage of our wealth to invest in that basket. In 1969, as part of his PhD dissertation under the guidance of Paul Samuelson, a 25-year old Robert C. Merton gave us an elegant and powerful rule for making that decision. Subject to a few important assumptions, the rule is simple and exact, and its simplicity and intuitive appeal make it a valuable rule of thumb. It brings together the three main variables that we would expect to be critical to the sizing decision: the basket’s expected risk and return, and the investor’s personal degree of risk aversion. Remarkably given its usefulness and beauty, this rule does not have a widely-agreed moniker; we hope we won’t cause any offense if we call it the “Merton Rule:” 4

Optimal Wealth Fraction to Invest = μ – r nσ2 = Sharpe Ratio nσ

where µ − r is the basket’s expected real return,5 σ its annual volatility, and n is the investor’s degree of risk-aversion.6 To illustrate, if your estimated expected real return of your basket is 4% and its risk (standard deviation) is 16%, and your level of risk aversion is n = 3 , then the formula says the optimal fraction of your wealth to invest in the basket is 52%:   4% (3)(16%)2

While you may not have heard of the Merton Rule, you might well have come across the Kelly Criterion,7 which can be thought of as a special case of the Merton Rule where n = 1 and the asset can only take two discreet future values, like a coin flip.8 The Merton Rule and Kelly Criterion are closely related, but they were developed independently – the Kelly Criterion in the context of gambling, and the Merton Rule in the context of investment-portfolio decision making under uncertainty.

Now we’d like to point out an important aspect of the Merton Rule which is central to the purpose of this note. Owing to its derivation, the input real return µ − r must be the forecast arithmetic mean of future returns. This is an important detail that can have a sizable impact on the Merton Rule’s result, so it’s worth quickly exploring the definition of the Arithmetic Mean return, and its relation to the other widely used return metric, the Geometric Mean return.

Arithmetic vs. Geometric Mean Return and Convexity Return

The Arithmetic Mean of a sequence of rates is the simple average of those rates. In contrast, the Geometric Mean return is the single rate which, when compounded, produces the same outcome as earning that sequence of rates period by period.

The Arithmetic Mean return (AM) is always greater than or equal to the Geometric Mean return (GM), and there is a simple formula that’s a good approximation for their difference:

AM – GM ~ 1 2 σ2

where σ is the standard deviation of the sequence of rates in question. We call this difference between the arithmetic and geometric mean returns the “Convexity Return,” as the difference fundamentally arises from the non-linear (convex) relationship between investment value and compound rate of return,9 or more prosaically, that equities can go up a lot but cannot go down by more than 100%.

The chart above illustrates this non-linear relationship. We show an approximation of the Expected Value of the investment by averaging the better-than-expected and worse-than-expected outcomes of 1% and 7% compound returns. The Convexity Return of about 1.3% in this case is the difference between the Geometric Mean (GM) return of 4% and the Arithmetic Mean (AM) return of about 5.3% that corresponds to the Expected Value (EV) of the investment.

As you can see, the Convexity Return can be significant. The 1.3% Convexity Return we get for equities assuming 16% annual volatility is meaningful. Although the issue of arithmetic or geometric mean return may seem like a technical detail, it represents an amount of return we wouldn’t ignore or “sweep under the rug” in any other context.

To get accurate results from the Merton Rule, it’s important we’re very clear about exactly what type of mean excess return we’re forecasting, and what that forecast really means.10 When the Merton Rule was formulated in the late 60s, we suspect most academics and practitioners looked primarily to historical returns to feed their future long-term return forecasts – and in this context assuming the forecast is an arithmetic mean makes perfect sense, as taking a simple average is a natural thing to do with a series of historical return data. Now however, people use a variety of methods to generate forecasts of future returns. This got us wondering whether most people today, when they’re forecasting expected long-term returns, think of that forecast (explicitly or implicitly) as a forecast of the arithmetic or geometric mean. So, we did a survey.11

The Survey

Seeking the proverbial “wisdom of the crowd,” at the start of 2017 we asked 118 experienced finance professionals (average age about 55), who were frequent readers of Elm’s blog posts, three questions about their views of the long-term return distribution of the US stock market via an online survey. The questions, and answers from our 118 respondents, are:

  1. What is your estimate of the investment return you would earn, expressed as an annual return above inflation, on a broad US equity market index fund starting today, and holding for 30 years, reinvesting all dividends, and ignoring taxes? 4.2% as in the histogram.
  2. You chose x% in Q1. At x% for 30 years, $1mm would grow into $1mm ∗ (1 + x%)30 in inflation-adjusted dollars. Do you agree that there’s a roughly 50% chance that your investment turns out better than this? 83% agreed.
  3. Still within the context of an investment in the broad US equity market: Which do you agree with more?
    1. Over a 30-year horizon for your equity investment, realizing an outcome double your estimated investment value or half your estimated investment value are about equally likely. 77% agreed.
    2. Over a 30 year horizon for your equity investment, realizing an outcome double your estimated investment value or losing all your money are about equally likely. 23% agreed.

Interpretation of Survey Results

What we learned from this survey was far more interesting than that the average real return estimate of our respondents was 4.2%. The chart shown to the right of a normal distribution represents more-or-less how the typical respondent would see future stock market rate-of-return outcomes.12 We have centered the distribution at the rough average of our respondents estimates, 4% real return, to reflect the view of the 83% of our respondents who described their estimate as lying at the 50% point (median) in the distribution of outcomes.13 The chart also is drawn to reflect the belief of 77% of our respondents that there was about an equal chance of the investment doubling or halving versus the central outcome.14

That 83% of respondents saw their estimate as the median return is highly significant, because for a log-normally distributed asset the median return is equal to the geometric mean return. The arithmetic mean return is much higher than the median in this case; for an asset with 16% volatility, there is only a 33% chance of exceeding the arithmetic mean.

We acknowledge that, by necessity, the survey was both brief and somewhat indirect. But we conclude the results support a hypothesis that most investors are forecasting a geometric mean return, not an arithmetic mean.

Adapting the Merton Rule, with Big Impact on Optimal Allocation

But now we have a problem – the Merton Rule wants an arithmetic mean, but most forecasters are estimating a geometric mean. What to do?

Simple: to use the Merton Rule in a way consistent with how most survey respondents are estimating returns, we just need to add the Convexity Return to respondents’ return estimate. When we do this, we arrive at a “Modified” Merton Rule:

Modified Merton Rule = (μ + ½σ2) – r nσ2 = Sharpe Ratio nσ + 1 2n

The modification is quite intuitive.15 We simply add the Convexity Return (½σ2 ) to the return estimate in the numerator. So for our typical respondent, he should set µ − r equal to his 4% excess return estimate plus 1.28% for the Convexity Return, instead of just putting in 4% as might seem natural in using the out-of-the-box Merton rule. To illustrate, using the same numbers as we did in Section 2, the modified Merton rule suggests an optimal fraction of wealth to invest of 69%, a significant increase over the 52% we get with the wrong input.

When we simplify, we see that the extra allocation above the Merton rule is solely a function of the degree of risk aversion, n, of the investor; more volatility generates more Convexity Return which is exactly the amount of return we require for the extra risk that generates it: very neat and tidy!

We can see it makes a big difference. For the typical respondent to our survey (assuming risk-aversion index n = 3 ), using the Modified Merton Rule would represent a roughly 33% increase in their optimal allocation to equities. For the respondents who were least optimistic about future equity returns,16 taking account of the Convexity Return would indicate a roughly 66% increase in optimal allocation.

Historical Context

So why does the mainstream academic literature make the assumption that investors already include the Convexity Return in their expected return estimates? We think there are two main factors at work here.

First, in his seminal papers on portfolio choice, Robert C. Merton makes the implicit assumption that the mean return “input” into the model dynamics is an arithmetic mean of annual returns. Later writers tend to follow the pioneers in a field, and through that tendency this assumption became standard in the literature. There’s nothing wrong with this assumption per se, but as we’ve discovered from our survey results, most investors today implicitly estimate a geometric mean (without Convexity Return) when they think about future market returns – so this estimate needs to be adjusted when using classic tools such as the Merton Rule, which is what we have explicitly done in our suggested modified rule-of-thumb. We suspect that at the time Merton was writing, the most common technique for estimating future returns was looking at historical returns, in which case the estimate representing an arithmetic mean is perfectly natural. Today however, many people prefer using forward-looking return estimates,17 which may more naturally be thought of as forecasts of the Geometric Mean, or Compound, Return.

Conclusion

We suspect that some readers may see Convexity Return, and the attendant Modified Merton Rule, as financial alchemy. It is not. The impact on experienced returns is of similar magnitude, and should attract a similar degree of investor attention, as the fees charged for and the Alpha promised by active investment management. While you can’t feed your family with expected returns – let alone Convexity Returns – you can’t make sound decisions under uncertainty without accurately taking the full measure of the distribution of possible outcomes.

Other readers may view the propositions of this note as mainly semantic, saying “as long as I think of the arithmetic mean return when making investment decisions, I don’t need an updated scaling heuristic.” True, but they should also recognize that their way of looking at the future is uncommon relative to the respondents to our survey.

We realize it is challenging enough for an investor to settle on a central, base-case estimate for the long-term real return and risk of equities together with an estimate of his individual degree of risk aversion. But if you are like the vast majority of the 118 people who took our survey, we suggest it could be worth the extra effort to explicitly factor the often-overlooked Convexity benefit of owning equities into your investment decisions.


Appendix I: Modeling Implications of the Survey Results

In the Portfolio Choice literature pioneered by Robert C. Merton, and most of the related literature which followed, the standard process for a risky asset is written:

dSt St = μ dt + σ dZt

where Zt is a standard Weiner Process. We call this the “Native Geometric” choice of process. Using Itô’s Lemma, this is equivalent to:

dlnSt = ( μ – 1 2 σ2) dt + σ dZt

with solution:

St = S0e(μ – ½ σ2 )t + σ Zt

Now take

g(xt) = ln(( xt x0 )1/t)

as the function mapping price to geometric mean (aka compound) rate of return (adjusted to continuous-compounding), and:

a(xt) = ln( 1 t t Σ i = 1 ( x_i xi – 1 ))

as the function mapping price to the arithmetic mean rate of return. We show below a few features of the Native Geometric process:

  • 𝔼 [ST] = S0eμ T
  • g(𝔼 [ST]) = μ
  • 𝔼 [a(ST)] = μ
  • 𝔼 [g(ST)] = μ – ½ σ2
  • CDFg(ST)(μ – ½ σ2) = 50%18

Despite being the classic choice of process, this doesn’t agree very well with how our survey respondents see the world. Most respondents said they see their return estimate μ̂  as being the return which there is a 50% chance of exceeding, i.e. CDF(μ̂ ) = 50%. But we see above that, for the Native Geometric process, the 50% point not only doesn’t equal μ, but more importantly depends on the process variance σ2. It’s true that for a given μ̂  and σ we can pick a μ s.t. CDF(μ̂ ) = 50%, but μ is then implicitly also a function of σ2 and effectively we have a new SDE. This strongly suggests that the Native Geometric choice is not the best model for the process described by survey respondents. Ideally, we’d like the observed return estimate to map directly onto a model parameter without additional calibration or adjustment.

As a model candidate potentially more consistent with respondents’ views, consider the process:

dlnSt = μ dt + σ dZt

with solution:

St = S0e μ t + σ Zt

We call this the “Native Exponential” process choice. The main features of this process are:

  • 𝔼 [ST] = S0e(μ + ½ σ2 ) T
  • g(𝔼 [ST]) = μ + ½ σ2
  • g(𝔼 [a(ST)]) = μ + ½ σ2

  • 𝔼 [g(ST)] = μ
  • CDFg(ST)(μ) = 50%

This seems to line up much more naturally with the survey results – we can simply set μ = μ̂ . But using a Native Exponential process has important implications for the optimal scaling of risky assets in a portfolio. As we’ll see in Appendix II, the result is materially different from the classic Merton scaling rule.


Appendix II: “Modified” Merton Rule Derivation

Start with a portfolio consisting of two assets – cash earning a riskless rate r, and a risky asset S with SDE:19

dlnSt = μ dt + σ dZt

The investor has CRRA utility:

u(x) = { x1 – n 1 – n ln(x) n ≠ 1 n = 1

and wishes to find the fraction of wealth κ to invest20 in the risky asset which optimizes utility to horizon T. As a consequence of continuously holding the fraction of wealth κ, we have:

dPt = θt dSt + (1 – κ) r Pt dt

θt = κ Pt St

which gives us the SDE for the portfolio value P parameterized by κ:

dP P = (r + κ (μ – r) + ½κσ2)dt + κ σ dZ

which by Itô’s Lemma is equivalent to:

dlnP = (r + κ (μ – r) + ½κσ2 – ½κ2σ2)dt + κ σ dZ

with solution:

Pt = P0 e(r + κ(μ – r) + ½κσ2 – ½κ2σ2)T + κ σ Zt

where Zt ~ N(0,√T). Taking the n ≠ 1 case, we wish to maximize expected utility:21

𝔼 [u(Pt)] = P01 – n 1 – n eRT

where:

R = (1 – n)(r + κ(μ – r) + ½κσ2 – ½κ2σ2) + ½(1 – n)2 κ2 σ2

and:

∂ R ∂ κ = (1 – n)(μ – r) + ½(1 – n)σ2 – (1 – n)σ2κ + (1 – n)2σ2κ

and we do this with the standard method of taking the partial wrt κ and setting it equal to zero:

∂ 𝔼 [u(Pt)] ∂ κ = P01 – n 1 – n eRT( ∂ R ∂κ )T = 0

(μ – r) + ½σ2 – σ2κ + (1 – n)σ2κ = 0

κ = μ + ½σ2 – r nσ2 = μ – r nσ2 + 1 2n

which is what our intuition expected, i.e. we just replaced μ with μ + ½σ2. Applying the same technique to the n = 1 case shows the same rule holds true for that case also.


Appendix III: Merton Rule Derivation

Start with a portfolio consisting of two assets – cash earning a riskless rate r, and a risky asset S with SDE:

dS S = μ dt + σ dZt

The investor has CRRA utility:

u(x) = { x1 – n 1 – n ln(x) n ≠ 1 n = 1

and wishes to find the fraction of wealth κ to invest22 in the risky asset which optimizes utility to horizon T. As a consequence of continuously holding the fraction of wealth κ, we have:

dPt = θt dSt + (1 – κ) r Pt dt

θt = κ Pt St

Which gives us the SDE for the portfolio value P parameterized by κ:

dP P = (r + κ (μ – r))dt + κ σ dZ

which by Itô’s Lemma is equivalent to:

dlnP = (r + κ (μ – r) – ½κ2σ2)dt + κ σ dZ

with solution:

Pt = P0 e(r + κ(μ – r) – ½κ2σ2)T + κ σ Zt

where Zt ~ N(0,√T). Taking the n ≠ 1 case, we wish to maximize expected utility:23

𝔼 [u(Pt)] = P01 – n 1 – n eRT

where:

R = (1 – n)(r + κ(μ – r) – ½κ2σ2) + ½(1 – n)2 κ2 σ2

and

∂ R ∂ κ = (1 – n)(μ – r) – (1 – n)σ2κ + (1 – n)2σ2κ

and we do this with the standard method of taking the partial wrt κ and setting it equal to zero:

∂ 𝔼 [u(Pt)] ∂ κ = P01 – n 1 – n eRT( ∂ R ∂κ )T = 0

(μ – r) – σ2κ + (1 – n)σ2κ = 0

κ = μ – r nσ2

which is the classic Merton Rule. Applying the same technique to the n=1 case shows the same rule holds true for that case also.


Further Reading and References:

  • MacLean, Thorp, Ziemba. “Good and Bad Properties of the Kelly Criterion,” Berkeley, January 2010.
  • Robert C. Merton. “Lifetime Portfolio Selection under Uncertainty: the Continuous-Time Case,” The Review of Economics and Statistics (51), 1969.
  • Robert C. Merton. “Optimum Consumption and Portfolio Rules in a Continuous-time Model,” Journal of Economic Theory (31), 1971.
  • Fischer Black, Myron Scholes. “The Pricing of Options and Corporate Liabilities,” Journal of Political Economy (81), 1973.
  • Fischer Black. “Estimating Expected Returns,” FAJ, 1993.
  • Mark Kritzman. Puzzles of Finance Chapter 3: Why the Expected Return Is Not To Be Expected. Wiley, 2000

  1. This not is not an offer or solicitation to invest, nor should this be construed in any way as tax advice. Past returns are not indicative of future performance.
  2. e.g. Modern Portfolio Theory suggests this is the Market Portfolio, but in general it could be whatever portfolio the investor decides is most attractive.
  3. In this note, we use equities as our main exemplar, but the effects we discuss pertain as well to other asset-classes.
  4. See Appendix III for more information on the specific assumptions and a derivation of the result.
  5. Or, more generally, its excess return measured relative to the investors risk-free benchmark or numeraire.
  6. For a formal definition of n , the coefficient of risk-aversion, see equation 2 in Appendix II. The authors suggest n = 2 to n = 4 is a reasonable range for most investors.
  7. Especially if you’ve been reading some of the recent notes from the authors, such as “Optimal Trade Sizing in a Game with Favourable Odds: The Stock Market,” Haghani and Morton, Dec 2017, SSRN.
  8. For those more familiar with fractional Kelly betting than the Merton Rule, we can think of n, the coefficient of investor risk aversion, as implying an optimal position-size of NO relative to the full Kelly (i.e. log-utility) investor.
  9. To illustrate with a two period example, say an investment returns 25% for a year and then loses 15% in the second year. The arithmetic mean return is 5% (the average of 25% and -15%), while the geometric mean return is 3.08% = (1.25 ∗ 0.85)1/2 − 1 . The difference of almost 2% in this case is equal to the Convexity Return.
  10. Please forgive the pun.
  11. We suspect most people don’t explicitly think “this is a forecast of the arithmetic/geometric mean”, but we can infer which mean they’re implicitly forecasting through asking questions about the expected properties of their forecast.
  12. Assuming 16% annual standard deviation in price (or about 3% standard deviation in the compound return to the 30 year horizon.), and that the stock market follows a random walk. The results of this analysis are not materially impacted by some long-term mean reversion or short-term momentum in equity prices.
  13. This is the median of the distribution, and for the normal distribution we have used, it is also the mean of the compound return distribution. See Appendix I for more discussion. Of the 17% who didn’t see it as the 50% point, we learned from a sample follow-up that about 15% of those thought the median was higher than their estimate.
  14. i.e. that the distribution of prices is better described as log-normal than normal, which implies that the distribution of compound rates is normal.
  15. See Appendix II for the formal derivation.
  16. the lowest decile of estimates.
  17. In the case of equities typically based on dividends, P/E ratios, etc.
  18. CDF is the Cumulative Density Function.
  19. We motivate this choice of SDE in Appendix I. An investor whose desired “input” return is intrinsically a continuously-compounded mean (in which case the arithmetic and geometric means are equal) would also find this the natural SDE.
  20. and continuously rebalance
  21. Which fortunately we can write analytically for this particular process.
  22. and continuously rebalance
  23. Which fortunately we can write analytically for this particular process.
Read More

Tax-Efficient Investing for US Citizens Long-Term Resident in the UK

February 3, 2017

Tax Matters

Tax-Efficient Investing for US Citizens Long-Term Resident in the UK

By Victor Haghani 1

US citizens who have been living in the UK for a long time, like me, are facing significant changes in how we will be taxed in the UK. Unfortunately, we got caught in the cross-fire when the UK set its sights on about 100,000 long-term UK residents who don’t pay tax on investment income to the UK, under the UK’s generous non-domicile rules, or to their country of domicile either, as almost no country in the world taxes its citizens when they live abroad. No country, that is, but one: the US. These new UK rules mean that US citizens who are long-term resident in the UK face the risk of being punitively taxed, or even double-taxed, by the UK and the US.

I’ve attended half a dozen meetings with accountants, lawyers and financial advisors who have done their best to explain the new rules to me. Even more complex than the rules are some of the remedies being proposed.

At Elm Partners, we think we’ve got a simple and tax efficient investment solution for US-UK investors. We’ve taken our Separately Managed Account offering, which is already cost efficient and tax efficient for most US investors, and made some modifications that should also make it tax efficient from a UK perspective for US-UK investors like me.


But, before telling you more:

I’ve got to remind you that this is an area of tax law that is complex, and each of us has a different tax situation. Even within a family, different pools of savings are taxed differently.

You should always seek professional advice. I’m not qualified to give, and this note is not, tax advice. The information in this note is purely illustrative and for purposes of encouraging discourse and further research.


What’s changing?

From April 2017, individuals who have been resident in the UK for 15 of the last 20 years will be deemed to be UK-domiciled for income tax and capital gains tax purposes.

Many US-UK investors currently elect the remittance basis of taxation, which means that for an annual remittance fee, they are only taxed on income and gains that they bring into the UK. The remittance basis of taxation allows US-UK investors to focus almost exclusively on the US tax efficiency of their investments, and to give limited attention to UK taxation. As the IRS taxes offshore investments made by US citizen punitively (e.g. PFIC taxation), US-UK investors tend to invest in US-domiciled products such as LLCs, LPs, or US-listed equities, mutual funds and ETFs.

Unfortunately, from April 2017, US-UK investors who remain in the UK will be stripped of the safe harbour of the remittance basis, and will have to think about whether their investments are tax efficient from both a US and UK perspective.

What’s the Problem?

Most (but not all) US-domiciled funds are taxed unfavourably in the UK. In a typical UK-domiciled fund, income and capital gains are treated separately. Income is taxed as it arises at the rate of 45%2 and Capital Gains is taxed at realisation at a lower rate of 20%.

Just as the US is tough on US investors investing in offshore vehicles, the UK is too. And just as the US tends to treat the UK as if it were a tax haven, the UK treats US-domiciled funds as “Offshore Funds” too.3 This means that a UK investor in a US domiciled fund will have all returns, whether Income or Capital Gains, treated as Income and taxed as they arise at the higher rate of 45%, which isn’t good as it’s higher than the rate at which they’d be taxed in the US.4

This isn’t so bad if the US-UK investor intends to keep his investment in the US fund until he leaves the UK, as that investment will not be subject to tax at all in the UK if there are no distributions and the investment is not sold.5 However, if he wants to liquidate that investment or if it’s a private equity fund that makes distributions, then the capital gains are taxed at the higher rate of 45%, which will be well above the long-term capital gains rate in the US and hence leave our investor with a tax inefficient outcome.

The problem is more complex, and can potentially be worse, if the investment is in a US LLC, which is a typical structure for many hedge fund and private equity investments. In the US, an LLC is treated as a partnership and investors are taxed on an arising basis (via a k-1) whether or not there have been distributions. By contrast, the UK will likely tax this income as it is received at 45% as a dividend. The UK will not allow an offset for US taxes paid in the past on that income, and it may be difficult to get an offset in the US on the tax paid, due to timing and character differences.6

Benefits of US funds under the HMRC’s Offshore Fund Reporting Status Regime:

To mitigate the onerous treatment of Offshore Funds, the HMRC allows funds to apply for “Reporting Status”. An offshore fund which benefits from reporting status is taxed similarly to a UK-domiciled fund. Under the regime, a Reporting Fund reports its income (whether distributed or not), which is then taxed on an arising basis at 45%, and the remaining Capital Gains are taxed on realisation at 20%. A full list of Offshore Funds which benefit from the HMRC’s Reporting Status can be found here.
At Elm, we’ve been searching through this list (about 50,000 line items) and have identified a sufficient range of low-cost US-listed funds and ETFs which have UK Reporting Status to be able to build a balanced and diversified portfolio that should be tax efficient from both a US and UK perspective. Through our existing relationship with our brokerage in the US, we can offer this solution to US-UK investors through our Separately Managed Accounts, which we can help open with minimal hassle.


Summary: Thoughts on the efficiency of investment options from a US-UK perspective

Individual equities and bonds:
This is probably the most tax efficient solution7 for US-UK investors, but it is difficult and expensive to build and maintain a portfolio of the hundreds (or thousands) of individual equities and bonds you would need to be well-diversified, unless you’ve got a really big portfolio to invest.

US listed Index funds and ETFs with UK Reporting Status:
Tax efficient from a US and UK perspective, and a cost efficient manner of creating and managing a globally diversified balanced portfolio. Elm Partners offers Separately Managed Accounts for US-UK investors constructed of these US listed ETFs with UK Reporting Status.8 See this short slide deck for details on all the changes we made to our basic SMA program to be a better fit for US-UK investors. Vanguard is currently the only blue-chip provider whose ETFs benefit from Reporting Status, but we would expect more providers to follow suit. We are in discussions with a number of index fund and ETF sponsors encouraging this to happen.

US domiciled funds without UK reporting status:
For US-UK investors who are confident they will not sell their holdings until they leave the UK, these investments may be tax efficient from a US-UK perspective, as they should not generate taxable gains in the UK if they are not sold.9 Income, however, would be taxable in the UK, but as long as interest rates and dividend yields remain low, the overall tax efficiency should remain high. Elm Partners has a non-distributing Delaware fund which fits this category. However, if the investor has a change of plans, or heart, and liquidates the investment at a gain while still subject to UK tax, this may turn out to be tax inefficient.

Hedge funds and Private Equity:
These investments tend to be relatively tax inefficient for US investors from a solely US perspective (see our note on hedge fund vs long only equity taxation here). The tax inefficiency of these investments may be compounded from a combined US-UK perspective, depending on the exact structure of the investment and the circumstances of the investor.10 These investments also tend to be cost inefficient and low on diversification and liquidity.


  1. Victor is the Founder and CIO of Elm Partners. Past returns are not indicative of future performance. This not is not an offer or solicitation to invest.
  2. I’m using the highest marginal rate as I’m assuming the typical US-UK investor is HNW.
  3. An investment in a limited partnership might be treated differently under UK tax rules, with all income being deemed to pass to the investor as it arises and maintaining the character of the income. Limited Partnerships tend to be less commonly used as pooled investment vehicles in the US.
  4. For the purpose of this note, I’m going to ignore the effect of withholding tax on dividends from equities, and use of those withholding taxes in US and UK tax returns.
  5. For investors who intend to leave the UK before realising their investment, there is a bright side to being invested in a non-distributing LLC, as the HMRC should treat the LLC as a company and will only tax income as it is distributed. However, you need to find an LLC investment that doesn’t make distributions. You won’t be surprised that at Elm Partners, we can offer that option to US-UK investors too.
  6. Even this is somewhat complicated by the recent “Anson vs HMRC” case, in which an investor in a US LLC, Mr Anson, successfully argued that the LLC should be treated as transparent by HMRC. HMRC has stated that it intends to continue to treat LLCs as companies, except for in special circumstances matching the Anson case. See here for more details. The implications of this case are not fully known, so it is important to speak to your tax advisors if you are in this situation.
  7. It is possible that owning individual equities may be the best way to be able to reclaim dividend withholding taxes on a UK tax return. We have not found a definitive answer on this question.
  8. Mutual funds are another good option but may be difficult to buy, as some brokerages such as the one we use, will only allow mutual fund purchases if the investor has a US address.
  9. Funds in the form of LPs may not be treated this way by HMRC, but rather as a pass-through.
  10. As noted above, whether these investments take the form of an LLC or LP can have a significant impact. We believe there is nothing to stop an LLC applying for UK reporting status, but it may actually prove to be more detrimental to investors, as the payment of income tax would be brought forward to an arising basis. Worse still, the payment of tax in this way may not be able to benefit from double tax relief due to the mis-alignment of tax treatment between the US and the UK.
Read More

A Sharper Lens for Sizing Up Nickels and Steamrollers

January 24, 2017

Investing 101

A Sharper Lens for Sizing Up Nickels and Steamrollers

By Victor Haghani and James White 1

In a world of low rates and high stock prices, it’s natural that many investors are looking for ways to earn a good return with limited exposure to equities. However, many candidate strategies have return distributions which are significantly different from the Normal and Log-normal distributions that serve as reasonable approximations for the return profile of most typical portfolio asset classes. They require an upgraded set of tools to analyze and incorporate into our portfolio, maintaining a good balance of risk and reward.

As an example, one strategy that we’ve been hearing about a lot recently is buying short-term, high-yielding bonds, particularly financial issues.2 On the face of it, this strategy is appealing: these bonds have 4 – 5% yields, seem unlikely to default, and don’t come with a lot of daily market risk. For investors who feel, as many do, that equities are expensive and will generate relatively meagre returns, these bonds appear to provide a lower-risk alternative without sacrificing yield. However, these bonds are quite different animals from the diversified debt/equity portfolios most investors normally think about, and provide an excellent practical illustration of the shortcomings of the most commonly employed investment analysis heuristics.

The first problem is our well-documented tendency, when thinking fast and intuitively, of putting too much weight on highly likely “headline” outcomes and ignoring ones that are very unlikely but extreme.3 We are apt to see the promised return of a 5% bond as the expected return, effectively setting to zero the probability of loss from default. Of course, we know there’s no such thing as a free lunch; 5% bonds can’t be risk-free in a world of 0-1% interest rates. Assuming a 3% probability of losing 65% in a default or restructuring brings the resultant expected return on these bonds to 3%, below the headline yield of 5% but still a healthy 60% of it.4

Using the heuristic of expected return, or even the ratio of expected return to risk (aka Sharpe ratio), can lead to a significant mis-evaluation when applied to a case like this. Neither metric gives us an adequate way to weigh the small risk of a large loss, nor do either tell us how much of these bonds we should optimally hold. What’s needed is a more fundamental and versatile tool, a sharper lens for sizing up the proverbial nickels and steamrollers. Expected Utility fits the bill.

It’s a concept that’s been around for a long time, but surprisingly is hardly used by present-day analysts and investors. Tellingly, professional gamblers rely on it heavily (well, those who don’t tend to change profession quickly). It’s the general framework behind the “Kelly Criterion,” a betting and investing guideline used to good effect by Ed Thorp, a pioneer in applying probability theory to gambling and investing, and others such as Warren Buffet and James Simons.5

The basic idea is that instead of thinking of our wealth in terms of dollars, we’re going to think in terms of how much our wealth is really worth to us, which we call its utility. For most of us, increasing the dollar value of our wealth delivers less and less utility. A consequence of this decreasing marginal utility of wealth is that we are risk averse: a loss hurts more than the same gain, and we need to be paid to take symmetric risks. Each person’s risk aversion is a matter of individual preference. Starting to think in terms of utility instead of dollar value can be challenging, but we have no choice: to make better decisions we have to get our arms around it.6

For the rest of this note, we’re going to assume a level of risk aversion that we think may be typical for high net worth investors. It’s a degree of risk aversion that would leave an investor indifferent between accepting or declining a payment of 4% of wealth in order to take a 50-50 risk of making or losing 20% of wealth.7

So, coming back to these short-term bonds – what’s the expected utility of owning them? The answer depends on what fraction of our wealth we invest in them. Let’s look at this chart of expected value and expected utility as a function of what fraction of our wealth is invested: 8

As you can see, if we put 45% of our wealth into these bonds, we’ve increased our expected utility by the same amount as if we had received a risk-free payment equal to 0.5% of our wealth. A greater or smaller allocation to these bonds is of less utility to us, and beyond a 75% allocation, we’d actually be better off doing nothing at all.

As is clear from the chart, expected utility not only tells us the “price of risk” – the difference between the Expected Return and Expected Utility lines in the chart – but it also allows us to determine how much of the asset is optimal to hold. Even better, we can compare across assets with very different distributions, knowing we’re pricing risk consistently. The heuristics of Expected Return or Sharpe ratio are silent on these critical questions and do not allow for appropriately comparing assets with dramatically different return distributions.9 By linking investment inputs and characteristics to optimal size, Expected Utility also gives us an intuitive framework for translating the uncertainty in our assumptions into investment-sizing decisions.

Our analysis so far has been highly stylized, but using this utility framework is a good starting point from which to build in more complex, real-world factors. For example, we could extend the analysis to look at a portfolio of these bonds or the full menu of alternative investments, rather than just a single bond. For taxable investors, we’d want to take account of interest being taxed at high ordinary rates, while the periodic capital losses cannot be offset or carried back. A robust analysis might also bring in other factors such as the timing of potential losses relative to when those losses hurt the most.

To a greater or lesser degree, many popular investment strategies, such as selling puts on the stock market, holding concentrated individual stock portfolios or investing in leveraged hedge fund strategies which don’t or can’t employ tight stop-loss limits, are also good candidates for the application of the Expected Utility framework. Our message here is not that utility is the only thing you need – but rather that for enterprising investors looking at potential investments with highly skewed and/or non-linear return distributions, the standard toolbox of Expected Return and Sharpe Ratio is inadequate. Although Expected Utility analysis is rarely seen in the mainstream, we hope we have illustrated that it is both practical and useful.


Appendix: Variation on a Theme

An extension to an interesting and related investment challenge is how to scale a Hedge-fund (HF) investment. In the case of a HF that has an expected excess return of 5% and follows a Brownian random-walk with Sharpe Ratio of 1, then for an investor twice as risk averse as Kelly, expected utility is maximized by investing 1000% of wealth in the HF.10 This doesn’t pass a basic sanity test, let alone one of prudence! So what’s wrong? Is the utility-corrected lens we’re looking through flawed somehow, or is it that we’re looking at an inaccurate representation of the HF return distribution? We think the lens is fine and that an investment that is fully described by a Brownian random-walk with Sharpe Ratio of 1, with no risk of blow-up, is a truly amazing investment. But, as at least one of your authors can attest, HFs do have a real risk of blow-up with low recovery,11 even if the probability is relatively remote.

What does our utility framework suggest once we incorporate this tail risk into our analysis? Let’s now assume that most of the time our HF follows the 1-Sharpe random walk as above, but also has a 0.1% / year12 chance of blowing up with zero recovery. We can see this is just a slightly more complex variation of the short-term bonds example above. As shown in the chart below this blow-up risk, though very small, has a dramatic impact on the optimal scaling recommendation.

Including the small chance of a blow-up yields approximately the same scaling recommendation as for a 0.3-Sharpe Brownian random-walk with a 5% expected return, which is similar to the stock-market. In contrast, including the 0.1%/year chance of blow-up only moves the Sharpe Ratio from 1 to 0.83, a much smaller difference than suggested by the utility framework. Here again, we see that for highly skewed or non-linear return distributions, the toolbox of Expected Return and Sharpe ratio is inadequate, and an expected utility analysis is both useful and intuitive.


  1. This not is not an offer or solicitation to invest, nor should this be construed in any way as tax advice. Past returns are not indicative of future performance.
  2. These bonds are usually subordinated or issued at the holding company level, or may even be preferred stock.
  3. Daniel Kahneman, Thinking Fast and Slow (2010).
  4. For reference, 3% is the fair-market default rate assuming a 4% equity risk premium and a credit β of 50%.
  5. See Ed Thorpe’s Beat the Dealer (1966).
  6. See one of your authors’ notes for more information on calibrating personal utility, available at SSRN: “Practical Utility, Risk Aversion, and Investment Sizing”.
  7. There are many models for utility, but here we use the most common model   U(x) = x1 – n – 11 – n and assume a typical value of n = 2 for our model investor. In general this leads to betting 50% as much as the “classic” Kelly criterion suggests. This form of utility function is wealth-scale independent.
  8. Assuming a 5% 1-year bond with default probability 3%, recovery 35%, and a 1% risk-free rate subtracted out so we can just look at excess return.
  9. Assets with dramatically different distributions can have very different utility and intuitively very different risk profiles, but identical Sharpe ratios.
  10. Assuming the only investment options are the HF and a riskless bond, and the investment is continuously rebalanced to 1000% of current wealth – generally not a realistic option for HF investments. There’s a classic result that the utility-optimal scaling under these conditions for an asset following a Geometric Brownian random walk is:   κ = μ – rnσ2 = 1000% in this case.
  11. A result of the leverage typically employed.
  12. Chosen not for perfect accuracy, but to demonstrate the impact of even quite a small blow-up likelihood, in this case 1 HF out of a pool of 1,000 each year.
Read More

Vic’s TEDx Talk: Quitting is For Winners

January 6, 2017

In the News

Vic’s TEDx Talk: Quitting is For Winners

My academic and early on-the-job training at Salomon Brothers ingrained in me a basic principle of value investing: when an investment becomes cheaper, try to add to your position. The last thing you should want to do is sell out. Yet, this is precisely what you’ll do if you’re a follower of the “Cut your losses early and let your profits run” creed of investing. And you’d be in very good company, as celebrated speculators like George Soros, Paul Tudor Jones, and Jesse Livermore1 attribute much of their success to abiding by this tenet.

In this short TEDx talk, I explore why the use of a pre-defined stop-loss can be valuable in important life decisions. In terms of investing, I explain why the stop-loss approach generates profits in addition to reducing losses (hint: think momentum).

Hope you find these ideas useful as you navigate the year to come, and beyond.2


  1. Reminiscences of a Stock Operator (1923).
  2. At Elm Partners, while we don’t use stop-losses per se, our rules-based long-only investment approach does incorporate exposure to momentum in asset prices combined with an attention to valuation. This was influenced in part by the long-term success of traders who followed the “cut your losses early and let your profits run” principle. This TEDx talk is meant solely for informational purposes and does not constitute an offer or solicitation to invest, or investment advice.
Read More

Podcast: Ignore Investing’s Mathematical Underpinnings at Your Peril

December 15, 2016

In the News

Podcast: Ignore Investing’s Mathematical Underpinnings at Your Peril

I recently had the pleasure of featuring on Bloomberg’s Odd Lots podcast, with Joe Weisenthal and Tracy Alloway. They asked me some great questions about mathematics in investing.

You can listen to the full 25 minutes here, or alternatively read a transcript of our discussion below.


Tracy Alloway: Let’s bring in Victor Haghani. Like I said, he was at LTCM, but he is now the CEO of Elm Partners, which is basically a portfolio of low cost index and exchange traded funds. Victor, thanks so much for joining us.

Victor Haghani: Thank you very much for having me.

TA: Victor, we actually brought you on after reading a paper that you did basically about what coin tossing and the probabilities involved in coin tossing can teach us about investing. Can you tell us about that paper?

VH: It came out of an experiment that I did with a colleague of mine who I’d worked with at Elm Partners, Rich Dewey. We had heard about some research that had been done involving coin flipping and how people managed situations where they were given a favourable odds investment opportunity. Sometimes you can’t quite remember where the ideas come from, but we decided to do this experiment where we would give subjects real money and allow them to flip a coin that was biased to be 60% likely to come up heads, 40% tails, and we told them that to begin with, and we gave them half an hour to flip, to bet as much of their starting $25 as they wanted, and at the end, however much money they had left in their bank, we would pay them up to a maximum amount of $250.

What we found was that our participants, who were pretty quantitatively trained young men and women, didn’t do very well and they didn’t get some of the basic concepts of decision making under uncertainty or they didn’t quite get the independent nature of the flips and the fact that it just made sense to keep betting heads, to bet some modest constant proportion of how much they had in their bank at any point in time on heads. It was really interesting to think about how people were having trouble with that and to give us some ideas for trying to help with education, as well, on that topic.

Joe Weisenthal: Explain really quickly the exact mechanics. They had $25, and they were supposed to bet what? Explain to us what the nature of the bet is. Then what did the lesson show about mistakes that people might or might not make when they invest?

VH: Sure. The exact mechanics of it were that we told the people to come for a lecture, and then we asked them to get out their laptops and to play this game. We gave them $25. That turned up on their screen in their bank accounts or their bankroll. Then they could bet up to the $25, or however much they had in their bank, on the flip of a coin and they could do it repeatedly. Some people flipped the coin 300 times in the 30 minutes that they had. If they won the flip, then their bankroll would go up and vice versa, and however much they were left with at the end, we told them, and we did pay them out as a check or cash, which was especially for a bunch of college students, which were the majority of our subjects, was very welcome.

We told them that there was a maximum payout to begin with, but it was only when they got to a point where they could reach the $250, for example if they had $225 in their bank accounts and they were betting $30 on heads, we would say, “By the way, the most we’ll pay you is 250, so you might want to reduce your bet from $30 to $25 because there’s no point in winning $255. We won’t pay you that.”

The most surprising thing in a way was the fact that people would relatively frequently bet on tails. Even though we told them it was 60% likely to be heads, even though in general heads was coming up more frequently for most people after they had flipped it a number of times, and particularly after a string of heads like four heads in a row, they were then more likely to bet on tails. Not everybody.

JW: That seems like a deep failure of numeracy to ever bet on tails, even to think that the past streak of flips has any bearing on the next flip.

VH: It is, but it’s like this deep-seated need that we have to see a story in random things. Given that half of the subjects at some point bet on tails, and 30% of them bet on tails a fair amount of the time, there’s something deep-seated in there. I had my mom do the experiment, and we talked about it afterwards, and she said to me, “I know that I should have never bet on tails, but I just couldn’t resist.” She knew it. She knew it didn’t make any sense, but she just couldn’t resist. It was interesting.

We did another experiment following up on the famous interview question about family planning, that if everybody in a society wants to have a girl, and so each family keeps having children until they have a girl, does that change the expected number of boys and girls? Most people feel that it does even though when thought of as a coin flip, you can see that they’re independent and there’s really nothing much you can do to change the expected number of boys being equal to the expected number of girls to any finite horizon.

TA: The point of those types of experiments is essentially that the optimal investment strategy is dictated by maths, and yet, we choose to ignore it for whatever reason because we instinctively don’t understand probabilities or there’s some emotional thing going on.

VH: I think people understand it. Our subjects were really quantitatively trained. They understood all of this. Some of them were even mathematicians at one of the universities where we did it, and some of the subjects were also investment professionals that had a lot of math and econ and finance training, so they understand it, but I think there’s something deep-seated that comes up and steers us off the path. It’s quite a lot of statistics training is probably what’s needed to get people to be disciplined. To be disciplined is not a lot of fun. Think about you’re sitting there flipping a coin for a half an hour and you’re just trying to bet 15% of your bankroll on it and keep betting on heads.

JW: It reminds me of reading about professional poker players who know that they can make a steady profit playing limit poker, which is a very mathematical, very little bluffing version of the game of poker, but they’re just bored out of their minds when they play it. No limit is more fun. It’s more exciting. It’s a little less mathematical and more based on emotion. They are more inclined to lose. These sure things are not very enjoyable practises.

VH: Think about index investing. The most boring thing you could do is take all of your savings and to put it into two index funds. Very few people really do that, and very few people do that and stick to it. People will do it, and then they will come back and feel that they need to change it because there was an election or there was a change in interest rates or something. It’s fighting that urge to leave it alone. Fighting the urge to be active is difficult in a lot of different contexts.

TA: We’re all suckers for a sense of control.

JW: Let’s talk about a different mathematical concept that’s incredibly important to investing, and that is compounding. I forget who said it. Maybe it was Einstein. Someone famous said something about compounding being one of the most powerful forces on Earth.

VH: I think that they say Einstein may have said something like that, as strange as it may be.

JW: I don’t know why he would have been talking about it, but I think he did say something about it for whatever reason. What is it people don’t understand? Why is compounding such an important concept to understand? What do people get wrong about this?

VH: I think that in these investing things or math things, in general, one of the things that really gets us is non-linearities, things that are not proportional. Compounding is one of those things, that the growth of your money doesn’t go up in a straight line. It goes up in this exponential line. It starts off growing slowly, and then as it gets bigger, it’s growing faster in terms of the amount of money by which it’s growing. The rate of growth, let’s say, stays the same. When you start to look at relatively long periods of time, which are the kinds of periods of time that are relevant to us in terms of building savings for retirement or personal security, the effects become large.

Those long-term horizons are important, and compounding of small effects really magnify out there. The one that we hear a lot about is the effect of fees, that if you’re compounding at a 5% return, because you’re paying 2% fees, compared to if you’re compounding at a 7% return with low fees, that what you wind up with at the end is not proportional to 7 over 5, but rather that 7% winds up giving you a lot more at the end because it’s 1.07 being raised to a power divided by 1.05 being raised to a power. Everything gets magnified by compounding.

Another thing that’s similar to fees is taxes. If we can invest in a way where we don’t pay tax until the end of our investment horizon, we wind up with a lot more money than if we’re paying the same rate of tax on our growth every year as we go along. An example of that would be, let’s say that you have an investment that has an 8% rate of return, and let’s say tax rates are 50%, just to make the math simple. After 30 years, if you’re paying tax every year, then your 8% return is only a 4% after tax return. If you have $100,000 and you’re investing it, then after tax, that $100,000 has grown to $324,000 after 30 years at this 4% rate of growth, half of the 8%. But if instead you’re deferring your tax to the end, then you’re growing at 8% because you’re not paying any tax on it, but at the end, you have to pay 50% tax on all your gain. When you do that, you wind up with close to double the money after 30 years. You wind up with $550,000-

JW: Wow.

VH: And almost a 6% instead of a 4% rate of return. That stuff really kicks in over these long horizons and is important. Small differences wind up being big differences because of compounding.

TA: What’s your favourite financial formula for investing if you had to choose one?

VH: I don’t know. I guess one of the simplest ones, one that’s been on my mind lately …… I think if I had more time to think of it, I’d find a better one, but it’s been on my mind a bit, is what’s known as Sharpe’s Equality from a paper that William Sharpe, the Nobel Prize winner, wrote in the early 1990s. I think the paper was called The Arithmetic of Active Investing. In that, he just made the very simple statement that the return on the average actively managed dollar has to equal the return of the market minus fees on the active stuff, and that comes about because the market return must equal a weighted average of the returns of the passive and active segments of the market. If the total market return is the same as the index return of the passive part, it’s like saying 2 equals 1 plus 1, then 2 minus 1 equals 1.

It’s a very simple…It’s like in physics, the idea of the conservation of energy-

JW: What are the practical ramifications of that from an investor standpoint? It sounds like an identity essentially. How does that manifest itself practically in terms of making investing decisions?

VH: It just helps us a lot in terms of thinking about what we’re doing when we choose active strategies, for an active strategy to be working for us, that we have to believe that there’s some other active strategy that’s losing money, and we have to be able to identify why and who that’s likely to be, that if we think that we’re going to make money, we really have to be sure of who we’re making the money from.

TA: It’s a zero sum game essentially.

VH: Yes, within that space. I think that it’s a good first approximation. It’s a valid identity. There’s some caveats and so on that people would bring into it, but I like that it’s simple. It reminds us of Bill Sharpe who is a really cool guy. I think it’s a really useful one to remember.

TA: I promised a potential LTCM question. I guess one of the other things we’ve observed in markets recently is the rise of smart Beta, but also risk parity strategies, and some people have likened risk parity to the old Black-Scholes portfolio insurance of the 1980s, and some people have connected LTCM’s collapse with Black-Scholes, so I guess I’m just curious how you feel about risk parity and how you feel about the down-sides of mathematics in finance.

VH: For me, the really short answer is leverage…My LTCM experiences just made me not want to use leverage explicitly in any sort of investment strategy for myself or anybody that I would be trying to help. Leverage has its place in our financial system. It has its place perhaps within the investment community, but personally, that was the primary cause of the problems at LTCM, so for me anyway, I know the arguments for risk parity may well be that the aversion to leverage by people like me is what makes using a moderate amount of leverage a good idea. That’s what some people who are proponents of risk parity would argue, that it’s an inefficiency that a bunch of people like me now are averse to using leverage, but I’m averse to using it.

I’m not a fan of risk parity because I just don’t feel that I need to use leverage to get better quality returns. I think that the returns afforded by the marketplace without using leverage and the risks attached thereto, are all sufficient for me, and then I can go to sleep and not worry about having to reduce exposures because my leverage is causing me to do that.

TA: What about financial formulas, in general, and maths in investing? What are the down-sides?

VH: Models used in investing are very useful. One of my colleagues once said to just think about the yield to maturity of a bond. Think about that as a model. At some point in time, yield to maturity wasn’t really used. People used to talk about the price of a bond. They used to talk about the current yield, the coupon divided by the price, and then people started yield to maturity. Yield is just a much more useful thing to use in thinking about comparing different bonds with each other; implied volatility is a more useful way of comparing stock options to each other.

There’s nothing magical. It doesn’t tell you what to do, but it’s just a more useful lens. These models are a useful way of decomposing things into more intuitive quantities that we can use in our decision making. I think that math in finance is useful, for sure. There’s no doubt about that, but when we start to try to optimise things too much using math, trying to become too optimal and following narrow mathematical rigour too far is extremely dangerous. You come up with a whole portfolio of different investments and you look at an optimization of that and it tells you to do things that common-sense would tell you probably don’t make sense to do.

Taken to an extreme, I think that mathematical models can lead us to dangerous places sometimes. That’s a great question. I wish I had more time to think about it and give you a better answer to it.

JW: That’s a great answer, and Victor Haghani of Elm Funds, really appreciate you coming on. Fascinating conversation. Looking forward to reading and learning more about some of these concepts, and I think our listeners will have learned a lot from this.

TA: Thank you.

VH: Thank you very much. It was a pleasure.

Read More

How much of a good thing is best for you?

December 9, 2016

How Elm Works

How much of a good thing is best for you?

By Victor Haghani and Andrew Morton 1

A thought experiment:

You can invest your wealth in only two assets: a risk-free one and a market portfolio of all public equities. Your investment choices, however, are limited to: A) put 100% in the risk-free asset, or B) 10% in the risk-free asset and 90% in equities. You cannot mix A) and B) – you must choose one or the other. What is the lowest return (above the risk-free rate) you would need to expect from equities for you to choose B? 2

Do you have your number in mind? Great. Now imagine you’re still in this two asset world and you wake up one day and find that the expected return of equities is, in fact, exactly equal to your answer above. But now you’re no longer limited to just options A and B; you’re completely free to invest however much you like in equities, from 0% to 100% or more. How would you invest now?

We’ve asked about a dozen friends this question, all financial professionals. If you’re like them, you’ve likely read the question twice, trying to understand exactly what we’re getting at. You may feel like you just answered that question, and isn’t 90% the answer?

Well, no, we don’t think it is. Please read on as we try to explain why we think 45% is about the right answer, and how we can use the perspective of this problem to answer some other interesting questions.3

You can get an immediate intuition for the problem and its solution by replacing our first question with: how much ketchup on your fries would be so much that you’d be indifferent between no ketchup and that much ketchup? 4 And then replace our follow-up question with: how much ketchup is your optimal amount? Our first question was calibrating your risk aversion via indifference points, and the follow-up question involved choosing an optimal point anywhere in between. The halfway point is a pretty good estimate, and exact in some investing models.

Putting a price on risk:

The properties of risk aversion are central to this problem, so let’s analyze the simplest type of risk, a 50/50 coin flip. A side payment is needed to make a typical, risk-averse person indifferent to risking a fraction f of his wealth on such a flip. How should this required side payment vary with f ? There is a very good reason to suggest that it should be proportional to f2 , meaning we should demand four times the side payment for twice the risk. To see this, compare a single flip risking √2% of your wealth with two consecutive flips each risking 1%. In both cases the mean and variance of total wealth change are the same (0 and 2).5 It seems reasonable that the total side payment should be the same too, which is only the case if the side payment is proportional to f2 .

Thus the problem of side payments for coin flips boils down to choosing the specific multiple of f2 as your indifference point. Although it’s a matter of individual preference, we suggest a reasonable range for this multiple is 1 to 2.6 Let’s use 1, meaning you are indifferent to an offer of 1% compensation to take on a ±10% of wealth flip, or 4% to take on a ±20% flip and so on. You would accept coin flip offers with higher side payments than this, and decline those with lower.

To connect the coin example to the stock market, let’s simplistically model investing 90% of savings in the stock market for one year as risking 18% of wealth on the flip of a coin, meaning over a year we’ll either lose 18% or gain 18%.7 The rule suggests that we’d require a side payment of 3.2% (i.e. 0.182) of our wealth, or 3.6% on the 90% we have invested in equities, to be indifferent.8 You can see that in answering the question about what expected return you would need to be indifferent between putting 90% of your savings into equities or 0% in equities, you were calibrating your function of risk aversion.

Let’s now turn to the question of how much you should invest given the freedom to choose any trade size you like. Consider what happens as you increase your investment from 0. Your expected gain goes up proportionally, but your risk – and hence, the required reward for bearing it – goes up with the square of the investment. Adding these two effects gives the diagram below, illustrating why the optimal point is halfway between the indifference points of 0% and 90%.9 So, someone who is indifferent to investing 90% in equities if they had a 3.6% expected return would optimally invest 45% of his or her wealth in equities at that expected return. In general, optimal size is half the indifference point size.

What kind of questions can we answer with this simple framework?

  • What expected return on equities would you need for your optimal allocation to be 75%?
    Recall that the required minimum return you said would make you indifferent to being 90% invested in equities, R90% , is also the expected return for which 45% is your optimal allocation. Optimal allocation varies in proportion to expected return, and so for it to be 75%, the expected return needs to be 75%/45% higher than your R90% . For our prototypical investor who has R90% of 3.6%, the expected return of equities at which a 75% allocation would be optimal is 6%.10
  • What is the effect of over-investing?
    As can be seen from the diagram, if you invest double your optimal allocation, you’ve thrown away all the benefit of the investment, and things get worse at an increasing rate from there. This means we should err on the side of taking less risk in the face of uncertainty about the probabilities of future returns.
  • At the other extreme, what about small investments? For example: a $100 coin flip for an investor whose net worth is $100,000?
    The f2 rule of thumb suggests being indifferent at a side payment of about (0.0012) x $100,000 = 10 cents! That seems crazy. But if that investor has 45% in the stock market and the rest in the riskless asset, over about one hour his or her wealth will fluctuate by about $100 (with stock volatility of 20%). Using the same 3.6% excess return of equities as earlier, the expected gain in that hour is about 19 cents. So if optimally invested wealth is producing 19 cents for $100 of risk, he or she is roughly indifferent to a side payment of half that. The framework therefore suggests (if needed!) accepting smaller rewards for small risks than might at first seem intuitive.
  • Should a passive investor in the stock market have a static allocation to equities?
    Investment advisers typically subject their clients to a risk evaluation to determine an appropriate allocation to equities, often taking the form of the kind of indifference questions we posed. If you believe that the expected return and/or risk of the equity market change over time, then your optimal allocation to equities should also change.11
  • Does horizon matter?
    To the extent investments are like coin flips, following a random walk, with the risk we care about (variance) and the compensation we’re being offered to accept that risk both growing proportionately with time, then, whether we choose a month, a year or a decade as the horizon for our investment won’t affect our choice of the optimal amount we should invest in it. Whether we’re flipping a biased coin once or 300 times, the amount we put at risk as a fraction of wealth should be the same.
  • How much is it worth to be able to invest in equities?
    We don’t need to believe that we can beat the market for us to put a positive value on the opportunity to invest in equities. Recall that we’re choosing an allocation to equities that maximizes the surplus of what we’re expecting to be paid over what we need to get paid to accept that risk. That surplus is equal to half of the risk premium multiplied by our optimal allocation. So, for example, if we see equities priced to deliver a 5% expected return, and our optimal allocation is 60%, then our surplus is 1⁄2 * 60% * 5% = 1.5% pa . This tidy sum should make us feel pretty good, even grateful, about the existence of the equity market and our ability to freely choose how much of it we’d like. We’re being invited to play a game with favorable odds, like betting on the flip of a coin that has a 60% heads bias, or playing in a poker game where the other players have more money than skill.12
  • What is “Kelly” betting, and how does it relate? The Kelly criterion, named after the scientist at Bell Labs credited with formulating it in 1956, tells us how to bet to maximize the expected growth rate of our wealth. Implicit in the Kelly criterion is a level of risk aversion that is half of what we’ve been using in these examples. So a Kelly investor in the stock market would need only 1.8% excess return to be indifferent to holding 90% of wealth in the market versus nothing – equivalently, he or she would only need to expect 3.6% to optimally be 90% invested in equities.13 This is a much more aggressive posture than most investors we’ve met seem comfortable with.

Conclusion:

We hope this discussion has given you some simple but versatile tools for thinking about a broad range of investment-related questions. Seeing the risk-taking decision more like tuning the dial on a radio, rather than flipping an on-off switch, enables us to get the most out of any game where the odds are in our favor, including investing in the stock market. The trick is to tune your portfolio to the point where the cost of risk, which is increasing quadratically, is just about to grow faster than the expected return you’re being paid to take that risk, which is increasing linearly. We hope you’ve found some “utility” in this brief summary, and that we haven’t taken too many liberties in attempting to distill what are some of the most valuable insights in finance from the last 400 years, from Daniel Bernoulli to Bob Merton, and many brilliant minds in between.


Further Reading and References:

  • Arrow, Kenneth J. “Alternative Approaches to the Theory of Choice in Risk-Taking Situations,” Econometrica, Oct 1951.
  • Bernoulli, Daniel. “Exposition of a New Theory on the Measurement of Risk,” (1738). Translated in Econometrica, (1954).
  • Haghani, Victor and Richard Dewey. “Rational Decision-Making under Uncertainty: Observed Betting Patterns on a Biased Coin,” SSRN, 2016.
  • Merton, Robert C. “Continuous-Time Finance.” 1990.
  • Norstad, John. “An Introduction to Utility Theory,” Norstad.org, 1999.

You can also see our pieces here, here and here on estimating the long-term expected return of the stock market and its risk.


  1. This not is not an offer or solicitation to invest, nor should this be construed in any way as tax advice. Past returns are not indicative of future performance.
  2. We will always use “the expected return of equities” to mean the expected return above the risk-free rate.
  3. If you answered 45%, you can pat yourself on the back and get back to teaching your finance students.
  4. Sorry, but we’re assuming you’re North American and do like ketchup on your fries.
  5. To be precise, if the second flip is for stakes of 1% of the new wealth, the variance of wealth change is 2.0001%.
  6. A multiple of 1 corresponds to that of an investor with power utility function with relative risk aversion parameter n = 2 . The power utility function is given by: u(w) = (w1 – n – 1)/(1 – n) for n ≠ 1 , u(w) = ln(w) for n = 1 . With n = 2 , we won’t accept a fair coin flip where we lose half of our wealth, no matter how big the upside if we win.
  7. Investing in the stock market isn’t the same as flipping a coin with known probability of outcomes. There is uncertainty in addition to risk concerning the distribution of outcomes. The setup of the problem in this note leaves out many real world practicalities, such as the value of one’s human capital, the correlation of future consumption with the market, and the existence of other assets, to name just a few. We also have modeled the stock market as normally, rather than log-normally, distributed to the one year horizon.
  8. By the same logic, the risk of the stock market is 18% / 0.9 = 20% .
  9. The net value for any allocation, k , to equities is: kα f2 – k(α f)2 , where α is the proportion committed to the risky investment, and k is the coefficient of risk aversion. The objective is maximized with respect to α when α = 1⁄2 .
  10. Let’s use the notation (P,f) for earning a side payment of fraction P of wealth in exchange for risking fraction f on a fair coin flip. What is the optimal fraction of wealth, k , of this risk we should choose, assuming we are indifferent to flips of (f2,f) ? To be indifferent, we need to choose k such that (kf2 = kP , or k = P / f2 . The optimal choice is half of that, so  k = 12 P / f2 You can see that the optimal allocation, k , moves in proportion to expected return, P , and in inverse proportion to the square of risk, f . So, if stock market risk dropped by 10% (say from 20% to 18% in our example), then your optimal allocation to equities would go up by  1(1 – 10%)2~24%
  11. Except if the changes in the return and risk of the market are such as to keep the ratio of return to variance constant.
  12. We could also view this 1.5% a year as the risk-free payment we’d need to receive to forego being able to invest in the equity market, and only be able to invest in the risk-free asset. We’ve asked a version of this question to about 30 of our friends. We’re still conducting this survey, so stay tuned for the results in a future note.
  13. As above, assuming 20% annual stock market volatility. The Kelly criterion in continuous time gives an optimal fraction to invest,  f = Rσ2 where σ is the annual risk of the stock market.
Read More

Do Index Buyers Make Over-Valued Stocks More Over-Valued?

December 2, 2016

Investing 101

Do Index Buyers Make Over-Valued Stocks More Over-Valued?

For more than 60 years, the capitalization-weighted market portfolio has been a cornerstone of the modern theory of investing. As it inexorably takes its place as a cornerstone of investment practice, it has been subject to a litany of increasingly strident criticisms, such as the recent Sanford C. Bernstein paper, “The Silent Road to Serfdom: Why Passive Investing is Worse Than Marxism.” If the purpose of such notes is to get press attention, they are successful, but if the purpose is to help our understanding of how markets function, they generally fall short.1

In this note, I want to dispel one frequently voiced myth about indexing, namely that when investors put their money into broad, market-cap weighted index funds or ETFs, it has the unintended consequence of increasing the aggregate mis-valuation in the market, by making over-valued equities more over-valued and under-valued ones more under-valued.2 For example, Timothy O’Neill, the global co-head of Goldman Sachs’ investment-management division, called indexing “a bubble machine” which “guarantees that the most valuable company stays the most valuable, and gets more valuable and keeps going up.”3 It’s a good story, but on closer inspection, it’s not true.

To see why, think of the equity market as being owned by two kinds of investors: a) market cap indexers, who own every equity in the market in proportion to its market capitalization, and, b) active managers, each of whom owns a portfolio of equities that matches his or her assessment of which are the best ones to own or avoid. We can see that while each active manager will own a portfolio that diverges from market cap weights, the holdings of the active managers in aggregate must by definition be in proportion to market cap weights.

As William Sharpe explained in his seminal 1991 note, “The Arithmetic of Active Management”:4 by definition, since the portfolio representing the total market and the portfolio of indexers are both capitalization weighted, it follows that active managers have to be capitalization weighted in aggregate too.5

Let’s dig deeper into these flows that worry the indexing critics. There are two ways that an investor can buy an index fund: a) with fresh money, for example from a pay-check, or b) from the sale of an actively managed holding of equities. In the case of a fresh money purchase, if the seller of the index fund is an indexer too, then the whole index as a package changes hands, and that shouldn’t have any impact on relative prices of individual equities. Furthermore, since it’s spread out over the whole market, an index trade should also have less price impact than an investor buying a concentrated subset of the equity market. If, however, the sale comes from a cross-section of active managers, then the new indexer needs to buy five times as much of a stock that has five times the market cap compared to another.

There are two reasons why this should not increase aggregate market mis-valuation: 1) to the extent that this buying has price impact, wouldn’t our best guess be that a stock that is five times as big as another can absorb five times the buying with the same percentage price impact? And, if you don’t agree with that, then 2) assuming that many stocks are mis-valued, why should we expect that big companies are more over-valued in percentage terms than small companies? Isn’t it more plausible that some large companies are over-valued and some are under-valued, and likewise for smaller companies? While it is true that an over-valued company has a market value that is larger than its fair value, for any given equity we don’t know a priori whether it is over-valued or under-valued, a subtle but critical distinction.

This reasoning was explained in greater depth by Harvard Professor André Perold in a 2007 note defending market cap weighted indexing: “Holding a stock in proportion to its capitalization weight does not change the likelihood that the stock is overvalued or undervalued. The notion that capitalization weighting imposes an intrinsic drag on performance is, accordingly, false.” 6

Let’s turn to the second case, that of an investor who decides to sell his actively managed holdings and move into an index fund. We should only be concerned here if we believe the investor somehow manages to sell the active managers who hold relatively under-valued equities. Even though all active managers think they own under-valued equities, we’ve already shown this cannot be the case in aggregate, because active managers as a group can’t do anything other than own the market index.

The intransigent critic of indexing might counter that that the investor who moves into an index fund will tend to sell his active holdings that have performed the worst recently, which would have the effect of pushing those stocks down even further. This argument requires us to believe that the active manager who is doing the worst is also somehow most likely to be the best at identifying under-valued equities. If we believe that investors in actively managed portfolios tend to chase returns (and we do7), then the effect of the return chaser moving into an index fund results in less of the pressure that the critic of indexing is concerned about. This is because the return chaser is doing half the chasing by not buying the actively managed portfolio that has done well recently. Better still, by giving up return chasing by becoming an index investor, he will be less likely to cause mis-valuation in the future.

In this note, we have used a simplified representation of the marketplace to explain why the argument that investor flows into broad, market-cap weighted indexes make misvalued equities more mis-valued is not correct. Of course, the real world is not so simple; investors frequently use index funds and ETFs to get exposure to narrowly defined indexes, such as utilities or REITs, or as part of active asset allocation approach (see our recent note on Active Index Investing). These uses of index products are worthy of attention (and are the main focus of our business at Elm Partners), but they don’t turn the fallacy into a truth in the case of broad, market cap weighted index funds, which has been the focus of this note, and represent the vast majority of index fund and ETF assets.


  1. See this WSJ article (Nov 24th, 2016) by Burton Malkiel (Princeton professor and author of “A Random Walk Down Wall Street”) for a critique of the Sanford C. Bernstein note.
  2. In the case of flows into narrowly defined “index” baskets, such as utilities or REITs, as we pointed out in our note “What’s up with REITs?” in July, this is better thought of as active management rather than broad market cap indexing and as with all actively managed investing, it can indeed impact relative valuations.
  3. James Ledbetter, “Is Passive Investment Actively Hurting the Economy?” The New Yorker, March 9, 2016.
  4. William Sharpe, “The Arithmetic of Active Management,” Financial Analysts’ Journal, 1991: “Each passive manager will obtain precisely the market return, before costs. From this, it follows (as the night from the day) that the return on the average actively managed dollar must equal the market return. Why? Because the market return must equal a weighted average of the returns on the passive and active segments of the market. If the first two returns are the same, the third must be also.”
  5. Lasse Pedersen, in “Sharpening the Arithmetic of Active Management”, SSRN.com, (2016), argues that Sharpe’s equality does not hold in general. In the case of a well-constructed, well-managed index, the effects that Pedersen lists are small enough to ignore for the purpose to which we are applying Sharpe’s equality.
  6. André Perold, “Fundamentally Flawed Indexing,” Financial Analysts Journal, volume 63, number 6, November/December 2007. Professor Perold was responding to the theory proposed by Robert Arnott (and others, including Jeremy Siegel) that,
    “No longer must investors suffer a performance drag by settling for an index that inherently overweights every overvalued company and underweights every undervalued one. With due respect to the pioneers in finance theory and the cap-weighted indexers, there is a better way.”

    Page 41: Arnott, Robert. 2006. “An Overwrought Orthodoxy.” Institutional Investor (18 December):36–41.

  7. For a more detailed discussion of return chasing, see our research paper on SSRN.com and this blog post: Return Chasing Can Be Hazardous to Your Wealth.
Read More