Does announcing trades still hurt index investors?
MSCI usually announces the stocks that will enter and leave its minimum-volatility indexes nine trading days before the changes take effect, and implements the changes at the closing prices of the day before. Huij and Kyosev (2016), two Robeco researchers, found that from 2010 to 2015 the stocks entering these indexes outperformed the other index stocks between the announcement and that close, and the stocks leaving underperformed them. Robeco calls this price move the cost of transparency, because the changes are public and other investors can trade on them before they are implemented.
I adapt their approach to the US index, MSCI USA Minimum Volatility, from 2012 to 2025, and take the stocks entering and leaving from the holdings of the iShares ETF that tracks the index. In 2012 to 2015, the part of their period that the ETF’s holdings cover, the stocks entering returned 0.57 percentage points more than the average index stock between the announcement and the rebalance, and the stocks leaving 1.19 points less. Over 2012 to 2025, the differences are −0.01 points for the stocks entering and +0.11 points for the stocks leaving, both close to zero.
The test
- Index changes. MSCI rebalances the index in May and November, and from August 2025 every quarter. At each rebalance, I compare the holdings of USMV, the iShares ETF that tracks the index, two trading days before and two trading days after the rebalance day. An addition, a stock entering the index, is held after and not before, and a deletion is held before and not after. November 2011 and May 2017 are left out, because iShares has no holdings files for the days needed, which leaves 28 rebalances from May 2012 to November 2025.
- Exclusions. As in Huij and Kyosev, I leave out the stocks that the parent index, MSCI USA, added or deleted at the same rebalance, because funds tracking MSCI USA trade them too, and the stocks delisted within three months, mostly takeovers. EUSA, the iShares ETF that tracks MSCI USA with equal weights, gives the parent index’s stocks. The stock data covers US-incorporated companies only, so companies incorporated elsewhere are left out. This leaves 369 additions and 326 deletions.
- Days. The rebalance day (RD) is the last day of the old index: MSCI implements the changes at its closing prices, and they take effect on the next trading day. MSCI’s methodology states that the changes are announced, in general, nine trading days before they take effect, so the announcement day (AD) is eight trading days before RD, as in Huij and Kyosev. I have not verified the announcement date of each rebalance.
- Return. A stock’s daily total return minus the average total return of the ETF’s holdings with returns in the stock data, added up from the close of AD to the close of RD. Unlike Huij and Kyosev, I leave out AD’s own return and give each rebalance the same weight.
- Volume. The value traded in a stock relative to its normal level, adjusted for trading in the market as a whole.
- Samples. 2012 to 2015, with eight rebalances, and 2012 to 2025, with all 28.
Starting point
The test uses daily total returns and value traded from LSEG Workspace, for US-incorporated stocks, and the ETF holdings that iShares publishes for past dates. Because the LSEG data is licensed, the raw data stays off this page, and I show the code without running it here, with the output of my run. Each calculation’s code is in an expandable block beside its explanation, and it runs after the code in the appendix, which reads the stock data, downloads the holdings and finds the index changes.
Step 1. The abnormal return
For each stock that enters or leaves the index, I compare its daily return with the average return of the index’s stocks on the same day:
\[AR_{i,t} = R_{i,t} - \bar{R}_t,\]
where \(R_{i,t}\) is the total return of stock \(i\) on day \(t\), the price change plus dividends, and \(\bar{R}_t\) is the average total return that day of the ETF’s holdings with returns in the stock data: the holdings before the rebalance up to RD, and the holdings after it from the next day on. Days are counted from AD, so AD is day 0 and RD is day 8.
# Days counted from AD (day 0), where RD is day 8: the days a stock needs a return on, the days
# shown and the normal days. np.arange(a, b) gives the whole numbers from a up to b, without b
CHECKED_DAYS = np.arange(-10, 25)
DAYS = np.arange(-5, 14)
NORMAL_DAYS = np.arange(-50, -10)
# The daily rows of the stocks needed: USMV's stocks, and EUSA's stocks before each rebalance
needed = holdings.loc[(holdings["fund"] == "USMV") | (holdings["side"] == "before"), "ric"]
panel = stocks.loc[stocks["Instrument"].isin(needed),
["Instrument", "Date", "Daily Total Return", "Daily Value Traded"]]
panel = panel.astype({"Instrument": str, "Date": str}).rename(
columns={"Instrument": "ric", "Daily Value Traded": "value_traded"})
panel["return"] = panel["Daily Total Return"] / 100 # from % to a decimal
panel["day"] = pd.to_datetime(panel["Date"]).map(day_number)
panel = panel.dropna(subset=["day"]).astype({"day": int}) # trading days only
# how="cross" pairs every rebalance with every day of the list
all_days = rebalances[["rd", "ad_number"]].merge(
pd.DataFrame({"day_from_ad": np.arange(-50, 25)}), how="cross")
all_days["day"] = all_days["ad_number"] + all_days["day_from_ad"]
# The index's average return: the old stocks up to RD (day 8), the new ones from day 9 on
index_days = all_days.loc[all_days["day_from_ad"].isin(CHECKED_DAYS)].copy()
index_days["side"] = np.where(index_days["day_from_ad"] <= 8, "before", "after")
index_days = index_days.merge(usmv[["rd", "side", "ric"]], on=["rd", "side"])
index_days = index_days.merge(panel[["ric", "day", "return"]], on=["ric", "day"])
index_return = index_days.groupby(["rd", "day"])["return"].mean().rename("index_return")
# Each index change's days; how="left" keeps a day even when the stock has no row that day
event_days = events.merge(all_days, on="rd")
event_days = event_days.merge(panel[["ric", "day", "return", "value_traded"]], on=["ric", "day"],
how="left")
event_days = event_days.merge(index_return.reset_index(), on=["rd", "day"], how="left")
event_days["abnormal_return"] = event_days["return"] - event_days["index_return"]Step 2. The abnormal volume
The volume shows how much more than usual these stocks trade around the rebalance. The abnormal volume of stock \(i\) on day \(t\) is
\[AV_{i,t} = \frac{V_{i,t} / \bar{V}_i}{V_{m,t} / \bar{V}_m},\]
where \(V_{i,t}\) is the stock’s value traded in dollars and \(\bar{V}_i\) its average over the 40 trading days from AD−50 to AD−11, before the announcement. \(V_{m,t}\) and \(\bar{V}_m\) are the same for the market, the parent index’s stocks together. Dividing by the market’s ratio adjusts for days on which the whole market trades more than usual, so 1 is a normal day. As in Huij and Kyosev, a stock needs a return on every trading day from AD−10 to AD+24 and a value traded on at least 10 of its 40 normal days.
# The market's value traded: EUSA's stocks before the rebalance, added up
market_days = all_days.merge(eusa.loc[eusa["side"] == "before", ["rd", "ric"]], on="rd")
market_days = market_days.merge(panel[["ric", "day", "value_traded"]], on=["ric", "day"])
market_traded = market_days.groupby(["rd", "day"])["value_traded"].sum().rename("market_traded")
event_days = event_days.merge(market_traded.reset_index(), on=["rd", "day"], how="left")
# Normal volume: the stock's and the market's average value traded on the normal days
normal_days = event_days.loc[event_days["day_from_ad"].isin(NORMAL_DAYS)]
normal = normal_days.groupby(["rd", "ric"]).agg(stock_normal=("value_traded", "mean"),
volume_days=("value_traded", "count"),
market_normal=("market_traded", "mean"))
checked = event_days.loc[event_days["day_from_ad"].isin(CHECKED_DAYS)]
checked = checked.merge(normal.reset_index(), on=["rd", "ric"])
checked["abnormal_volume"] = ((checked["value_traded"] / checked["stock_normal"])
/ (checked["market_traded"] / checked["market_normal"]))
# Keep the index changes with a return on every checked day and enough normal days, then crop
# to the days shown
checks = checked.groupby(["rd", "ric"]).agg(return_days=("abnormal_return", "count"),
volume_days=("volume_days", "first")).reset_index()
has_all_days = (checks["return_days"] == len(CHECKED_DAYS)) & (checks["volume_days"] >= 10)
complete = checks.loc[has_all_days, ["rd", "ric"]]
shown = checked.merge(complete, on=["rd", "ric"])
shown = shown.loc[shown["day_from_ad"].isin(DAYS)]
print(f"{len(complete)} of {len(checks)} index changes have complete data")695 of 695 index changes have complete data
All 695 index changes have complete data, so this step leaves none out.
Step 3. The two samples
Then I average the abnormal returns of each day, first over the stocks of each rebalance and then over the rebalances:
\[\overline{AR}_t = \frac{1}{K}\sum_{k=1}^{K}\frac{1}{N_k}\sum_{i \in k} AR_{i,t},\]
where \(K\) is the number of rebalances in the sample and \(N_k\) the number of additions, or of deletions, at rebalance \(k\). Averaging within each rebalance first gives every rebalance the same weight, because the stocks that change on the same day share the same market news. Each line in the video adds up these averages and is zero at AD:
\[CAR(t) = \sum_{s=-5}^{t} \overline{AR}_s - \sum_{s=-5}^{0} \overline{AR}_s,\]
where the first sum runs from AD−5, the first day shown, to day \(t\), and subtracting the second sum puts the line at zero at AD. So \(CAR(8)\) is the return relative to the average index stock from the close of AD to the close of RD. I start at the close of AD and leave out AD’s own return, because the time of day at which MSCI publishes is not known, and AD’s return can come before the announcement. The volume is averaged the same way, with additions and deletions together, and shown in % above its normal level.
# The two samples: the rebalances up to EARLY_SAMPLE_END, and all of them
early = shown.loc[shown["rd"].dt.year <= EARLY_SAMPLE_END].assign(sample="2012–2015")
samples = pd.concat([early, shown.assign(sample="2012–2025")])
# Each rebalance counts once: the average of its stocks first, then the average over rebalances
per_rebalance = samples.groupby(["sample", "change", "rd", "day_from_ad"])[
"abnormal_return"].mean()
average_return = per_rebalance.groupby(["sample", "change", "day_from_ad"]).mean()
# One column per sample and direction; cumsum adds up the days, and subtracting the value at AD
# makes each line zero at the close of AD
lines = average_return.unstack(["sample", "change"]).cumsum() * 100
lines = lines - lines.loc[0]
volume_per_rebalance = samples.groupby(["sample", "rd", "day_from_ad"])["abnormal_volume"].mean()
volume = volume_per_rebalance.groupby(["sample", "day_from_ad"]).mean().unstack("sample")
volume = (volume - 1) * 100 # in % above the normal level
counts = samples.drop_duplicates(["sample", "rd", "ric"]).groupby(["sample", "change"]).size()
print("Index changes used:")
print(counts.unstack().to_string())
print()
print("Return relative to the index's stocks from the close of AD to the close of RD (%):")
print(lines.loc[8].unstack().round(2).to_string())
print()
print("Value traded on RD, above its normal level (%):")
print(volume.loc[8].round(0).to_string())Index changes used:
change Addition Deletion
sample
2012–2015 93 62
2012–2025 369 326
Return relative to the index's stocks from the close of AD to the close of RD (%):
change Addition Deletion
sample
2012–2015 0.57 -1.19
2012–2025 -0.01 0.11
Value traded on RD, above its normal level (%):
sample
2012–2015 73.0
2012–2025 88.0
Over the eight rebalances of 2012 to 2015, the stocks entering returned 0.57 percentage points more than the average index stock from the close of AD to the close of RD, and the stocks leaving 1.19 points less, in the direction Huij and Kyosev describe. Over all 28 rebalances, the differences are −0.01 points for the stocks entering and +0.11 points for the stocks leaving. The value traded on RD was 73% above its normal level in 2012 to 2015 and 88% above it over 2012 to 2025. So, over the longer sample, the stocks entering and leaving returned about as much as the other index stocks before the changes were implemented, and trading in them on RD was still far above normal.
The video
The video shows both samples. The block below draws the 2012 to 2015 lines first, pauses, and then moves the lines and the volume bars to the 2012 to 2025 estimates.
TEAL, ORANGE, BAR_GREY = "#17868A", "#D2822B", "#B9AFA3"
BG, INK, GRID, SHADE = "#FCEFE3", "#1f1f1f", "#EADCCC", "#F4E2CF"
WIDTH, HEIGHT, DPI, FPS = 1440, 960, 120, 60 # pixels, dots per inch, frames per second
DAY_LABELS = ["−5", "−4", "−3", "−2", "−1", "AD", "+1", "+2", "+3", "+4", "+5", "+6", "+7",
"RD", "+1", "+2", "+3", "+4", "+5"]
early_lines, full_lines = lines["2012–2015"], lines["2012–2025"]
# The axis limits: the lowest and highest values rounded out to a half, plus half a percent
y_low = np.floor(lines.min().min() * 2) / 2 - 0.5
y_high = np.ceil(lines.max().max() * 2) / 2 + 0.5
volume_low = min(-20, np.floor(volume.min().min() / 10) * 10)
volume_high = max(30, np.ceil(volume.max().max() / 20) * 20 + 10)
MIN_GAP = (y_high - y_low) * 0.065 # the smallest gap between two labels
# The figure and its two charts; add_axes takes [left, bottom, width, height] as shares of the
# figure, and sharex gives both charts the same days
fig = plt.figure(figsize=(WIDTH / DPI, HEIGHT / DPI), dpi=DPI)
fig.patch.set_facecolor(BG)
top = fig.add_axes([0.10, 0.39, 0.735, 0.49])
bottom = fig.add_axes([0.10, 0.16, 0.735, 0.18], sharex=top)
for axis in [top, bottom]:
axis.set_facecolor(BG)
axis.axvspan(0, 8, color=SHADE, zorder=0) # the days from AD to RD
axis.axhline(0, color="#C9BCAB", lw=1, zorder=1)
axis.axvline(0, color="#8C8173", lw=1.3, ls=(0, (3, 3)), zorder=2) # AD, dashed
axis.axvline(8, color="#5B534A", lw=1.4, zorder=2) # RD, solid
axis.grid(axis="y", color=GRID, lw=1)
axis.set_axisbelow(True)
axis.spines[["top", "right"]].set_visible(False)
axis.spines[["left", "bottom"]].set_color("#D5C6B4")
axis.tick_params(labelsize=12.5, colors="#6C645B", length=0)
top.set_xlim(-5.5, 13.5)
top.set_ylim(y_low, y_high)
bottom.set_ylim(volume_low, volume_high)
top.tick_params(axis="x", labelbottom=False)
bottom.set_xticks(DAYS, DAY_LABELS)
bottom.tick_params(axis="x", pad=10, labelsize=13)
for label in bottom.get_xticklabels():
if label.get_text() in ["AD", "RD"]:
label.set(fontsize=15, fontweight="bold", color=INK)
top.set_ylabel("Return relative to index stocks (%)", fontsize=14, color="#4a4a4a", labelpad=9)
bottom.set_ylabel("Volume above\nnormal (%)", fontsize=12.5, color="#4a4a4a", labelpad=9)
sample_label = fig.text(0.94, 0.955, "", ha="right", va="center", fontsize=27, color=INK)
# One line, one name, one dot and one value for each direction. plot returns a list of lines,
# and [0] takes the one line. The value's cream outline keeps it readable over the lines
curves, names, dots, values = {}, {}, {}, {}
for change, color, name in [("Addition", TEAL, "Stocks entering"),
("Deletion", ORANGE, "Stocks leaving")]:
curves[change] = top.plot([], [], color=color, lw=3.4, solid_capstyle="round", zorder=4)[0]
names[change] = top.text(13.8, 0, name, color=color, fontsize=14, va="center")
dots[change] = top.plot([], [], "o", color=color, ms=9, zorder=5)[0]
values[change] = top.text(7.55, 0, "", color=color, fontsize=16, fontweight="bold",
ha="right", va="center", zorder=6,
path_effects=[patheffects.withStroke(linewidth=4, foreground=BG)])
bars = bottom.bar(DAYS, np.zeros(len(DAYS)), color=BAR_GREY, width=0.72, zorder=3)
def spread(heights):
"""Two label heights, moved MIN_GAP apart when they are closer, in the same order."""
if abs(heights["Addition"] - heights["Deletion"]) >= MIN_GAP:
return heights
middle = heights.mean()
if heights["Addition"] >= heights["Deletion"]:
return pd.Series({"Addition": middle + MIN_GAP / 2, "Deletion": middle - MIN_GAP / 2})
return pd.Series({"Addition": middle - MIN_GAP / 2, "Deletion": middle + MIN_GAP / 2})
# The plan of the video, one row per drawing: how much of each line is drawn (0 to 1), how far
# the lines have moved from the early to the full sample (0 to 1), whether the values at RD
# show, and for how many frames the drawing stays. s * s * (3 - 2 * s) starts and ends slowly
s = np.arange(1, 101) / 100
drawing = s * s * (3 - 2 * s)
s = np.arange(1, 181) / 180
moving = s * s * (3 - 2 * s)
plan = pd.concat([
pd.DataFrame({"drawn_in": [0.0], "drawn_out": 0.0, "moved": 0.0, "show_values": False,
"frames": 30}),
pd.DataFrame({"drawn_in": drawing, "drawn_out": 0.0, "moved": 0.0, "show_values": False,
"frames": 1}),
pd.DataFrame({"drawn_in": 1.0, "drawn_out": drawing, "moved": 0.0, "show_values": False,
"frames": 1}),
pd.DataFrame({"drawn_in": [1.0], "drawn_out": 1.0, "moved": 0.0, "show_values": True,
"frames": 90}),
pd.DataFrame({"drawn_in": 1.0, "drawn_out": 1.0, "moved": moving, "show_values": False,
"frames": 1}),
pd.DataFrame({"drawn_in": [1.0], "drawn_out": 1.0, "moved": 1.0, "show_values": True,
"frames": 300}),
], ignore_index=True)
video_path = FOLDER / "nr75_usmv_video.mp4"
# avc1 is H.264, the format web browsers play. On Windows, OpenCV can first print that it cannot
# load OpenH264; it then writes H.264 with Windows' own encoder
video = cv2.VideoWriter(str(video_path), cv2.VideoWriter_fourcc("a", "v", "c", "1"), FPS,
(WIDTH, HEIGHT))
assert video.isOpened(), "OpenCV could not open the video writer"
for step in plan.itertuples():
# The lines and the bars between the two samples: moved 0 is 2012-2015, moved 1 2012-2025
shown_lines = (1 - step.moved) * early_lines + step.moved * full_lines
shown_volume = (1 - step.moved) * volume["2012–2015"] + step.moved * volume["2012–2025"]
sample_label.set_text("2012–2015" if step.moved < 0.5 else "2012–2025")
end_heights = spread(shown_lines.loc[13])
rd_heights = spread(shown_lines.loc[8])
for change, drawn in [("Addition", step.drawn_in), ("Deletion", step.drawn_out)]:
# The line ends between two days while it is drawn, so that it grows smoothly;
# np.interp gives the line's height at that point
x_end = -5 + drawn * 18 # day -5 when drawn is 0, day 13 when 1
x = np.append(DAYS[DAYS < x_end], x_end)
curves[change].set_data(x, np.interp(x, DAYS, shown_lines[change]))
names[change].set_y(end_heights[change])
names[change].set_visible(drawn == 1)
# {:+.2f} writes two decimals with the sign; replace puts in a typographic minus
dots[change].set_data([8], [shown_lines.loc[8, change]])
values[change].set_text(f"{shown_lines.loc[8, change]:+.2f}%".replace("-", "−"))
values[change].set_y(rd_heights[change])
dots[change].set_visible(step.show_values)
values[change].set_visible(step.show_values)
for bar, height in zip(bars, shown_volume):
bar.set_height(height)
# Draw the figure, turn its pixels from matplotlib's RGBA into OpenCV's BGR order, and
# write the same picture for as many frames as the plan says
fig.canvas.draw()
picture = cv2.cvtColor(np.asarray(fig.canvas.buffer_rgba()), cv2.COLOR_RGBA2BGR)
for repeat in range(step.frames):
video.write(picture)
video.release()
plt.close(fig)
print(f"Saved {video_path.name}: {plan['frames'].sum() / FPS:.1f} seconds")Saved nr75_usmv_video.mp4: 13.3 seconds
The teal line is the stocks entering and the orange line the stocks leaving, each relative to the average index stock and zero at AD. The shaded days run from AD to RD, and the dots give each line’s value at RD while the video pauses. The grey bars are the value traded above its normal level, for both groups together. The years in the corner name the sample. So, the teal and orange lines that separate before RD in 2012 to 2015 are both close to zero at RD over 2012 to 2025.
Limitations
The announcement day is assumed. I take AD from MSCI’s general rule and have not checked the announcement date of each rebalance. If MSCI announced a day earlier, part of the price move happens before AD, outside the measure. If it announced a day later, the measure starts before the news.
Eight rebalances in 2012 to 2015. The early averages rest on eight rebalance days, and the stocks that change on one day share that day’s market news. So the 2012 to 2015 numbers are imprecise, and the comparison with 2012 to 2025 rests on few early observations.
A price move is not a trading cost. The measure is how the stocks moved relative to the average index stock, which is not what the funds paid. Part of the move on RD itself can come from index funds’ own buying and selling at the closing prices.
Conclusion
For this US index, the stocks entering outperformed and the stocks leaving underperformed before the rebalance in 2012 to 2015, as Huij and Kyosev describe. Across 2012 to 2025, both average differences are close to zero. So, in the longer sample, prices did not move against index investors on average between the announcement and the rebalance.
Appendix: building the index changes
The code below runs first. It reads the stock data, downloads the holdings and finds the index changes that the steps above use.
Settings
The location of the stock file and the rebalance days, which are the implementation days that MSCI’s review press releases state.
from pathlib import Path
import re
import time
import cv2 # OpenCV: writes the frames into an MP4 video
import matplotlib.pyplot as plt
from matplotlib import patheffects
import numpy as np
import pandas as pd
import requests # downloads the iShares files
FOLDER = Path.cwd() # files are read and written beside this notebook
US_STOCKS = Path(r"C:\Users\alexa\Dropbox\Claude_Code_Test\Betting against beta\US_stocks.csv")
# The implementation days (RD), from MSCI's review press releases at
# app2.msci.com/eqb/pressreleases/archive/. USMV started in October 2011
REBALANCE_DAYS = ["2011-11-30", "2012-05-31", "2012-11-30", "2013-05-31", "2013-11-26",
"2014-05-30", "2014-11-25", "2015-05-29", "2015-11-30", "2016-05-31",
"2016-11-30", "2017-05-31", "2017-11-30", "2018-05-31", "2018-11-30",
"2019-05-28", "2019-11-26", "2020-05-29", "2020-11-30", "2021-05-27",
"2021-11-30", "2022-05-31", "2022-11-30", "2023-05-31", "2023-11-30",
"2024-05-31", "2024-11-25", "2025-05-30", "2025-08-26", "2025-11-24"]
EARLY_SAMPLE_END = 2015 # the early sample ends in 2015, close to Robeco's sampleStock data
US_stocks.csv has one row per stock and trading day, with the stock’s LSEG code (Instrument). One call reads the six columns the test needs, which takes about a minute and 3 GB of memory. A trading day is a date with rows for at least 500 stocks, and the trading days are numbered so that RD minus 8 is AD.
# Text columns are read as "category", which stores each different text once, so the 27 million
# rows take about 0.7 GB
stocks = pd.read_csv(US_STOCKS,
usecols=["Instrument", "Company Name", "Date", "Ticker",
"Daily Total Return", "Daily Value Traded"],
dtype={"Instrument": "category", "Company Name": "category",
"Date": "category", "Ticker": "category"})
# The trading days in date order, numbered 0, 1, 2, ..., so that eight trading days after
# day n is day n + 8
rows_per_date = stocks["Date"].value_counts()
busy_dates = rows_per_date.index[rows_per_date >= 500]
trading_days = pd.to_datetime(busy_dates.astype(str)).sort_values()
day_number = pd.Series(range(len(trading_days)), index=trading_days)
rebalances = pd.DataFrame({"rd": pd.to_datetime(REBALANCE_DAYS)})
rebalances["rd_number"] = rebalances["rd"].map(day_number)
rebalances["ad_number"] = rebalances["rd_number"] - 8 # AD is eight trading days before RD
# {:,} writes commas between thousands, and {:%Y-%m-%d} writes a date as year-month-day
print(f"{len(stocks):,} rows; {len(trading_days):,} trading days from "
f"{trading_days[0]:%Y-%m-%d} to {trading_days[-1]:%Y-%m-%d}")27,422,289 rows; 6,574 trading days from 2000-01-03 to 2026-02-27
Holdings
I use the files two trading days before and after RD, because iShares has dated its files in two ways: until 2014 the file of RD shows the old holdings, and later files show the new ones. When iShares has no usable file for a day, I try the next day further out, up to four trading days. A file is usable when its first line names the fund, its second line gives the date asked for and its table lists stocks. The code saves every file that iShares sends for a fund, with or without holdings, so a later run downloads nothing. A rebalance without all four files, USMV and EUSA before and after, is left out.
FUNDS = {"USMV": 239695, "EUSA": 239693} # iShares' number for each fund
FUND_NAMES = {"USMV": "iShares MSCI USA Min Vol Factor ETF", # the first line of each file
"EUSA": "iShares MSCI USA Equal Weighted ETF"}
URL = ("https://www.ishares.com/varnish-api/blk-one01-product-data/product-data/api/v1/"
"get-fund-document?appType=PRODUCT_PAGE&appSubType=ISHARES&targetSite=us-ishares"
"&locale=en_US&portfolioId={portfolio}&userType=individual&component=holdings"
"&asOfDate={date:%Y%m%d}")
(FOLDER / "holdings").mkdir(exist_ok=True) # one saved file per fund and date
def holdings_on(fund, date):
"""The stocks a fund held on a date, from the saved file or else downloaded and saved.
Returns None when the file is not usable: it names another fund or date, or lists no stocks.
"""
path = FOLDER / "holdings" / f"{fund}_{date:%Y%m%d}.csv"
if not path.exists():
response = requests.get(URL.format(portfolio=FUNDS[fund], date=date),
headers={"User-Agent": "Mozilla/5.0 (academic research)"},
timeout=60)
response.raise_for_status() # stop if iShares answers with an error
content = response.content
time.sleep(1) # at most one request a second
# Save every file that names the fund, with or without holdings for the date, so that a
# later run downloads nothing. utf-8-sig drops the mark some files start with
if content.decode("utf-8-sig").startswith(FUND_NAMES[fund]):
path.write_bytes(content)
else:
return None
# The first line names the fund, and the second gives the date of the holdings, or "-" when
# iShares has none for that date
lines = path.read_text(encoding="utf-8-sig").splitlines()
if (len(lines) < 10 or lines[0] != FUND_NAMES[fund]
or lines[1] != f'Fund Holdings as of,"{date:%b %d, %Y}"'):
return None
# The table starts on the tenth line, and thousands="," reads 1,494.50 as 1494.5
table = pd.read_csv(path, skiprows=9, thousands=",")
stocks_held = table.loc[table["Asset Class"] == "Equity", ["Ticker", "Name", "Weight (%)"]]
if stocks_held.empty:
return None
return stocks_held
# The four files of a rebalance, each on the nearest trading day with a usable file: two trading
# days from RD, else three, else four
FILES = [("USMV", "before", [-2, -3, -4]), ("USMV", "after", [2, 3, 4]),
("EUSA", "before", [-2, -3, -4]), ("EUSA", "after", [2, 3, 4])]
holding_tables = []
for rebalance in rebalances.itertuples():
for fund, side, offsets in FILES:
for offset in offsets:
date = trading_days[rebalance.rd_number + offset]
table = holdings_on(fund, date)
if table is not None:
holding_tables.append(table.assign(rd=rebalance.rd, fund=fund, side=side,
date=date))
break # stop at the first usable file
holdings = pd.concat(holding_tables, ignore_index=True)
# Keep the rebalances with all four files: USMV and EUSA, before and after
files_found = holdings.drop_duplicates(["rd", "fund", "side"]).groupby("rd").size()
has_all_files = rebalances["rd"].isin(files_found.index[files_found == 4])
print("Rebalances left out, files missing:",
rebalances.loc[~has_all_files, "rd"].dt.strftime("%Y-%m-%d").tolist())
rebalances = rebalances.loc[has_all_files]
# copy() makes a separate table, so the columns added below change only this one
holdings = holdings.loc[holdings["rd"].isin(rebalances["rd"])].copy()Rebalances left out, files missing: ['2011-11-30', '2017-05-31']
LSEG codes
The iShares files give each stock’s ticker and name, but no ISIN, and they show today’s ticker even in old files. So I give each holding the LSEG code that trades on the file’s date, in three steps:
- Same ticker. LSEG’s ticker equals the holding’s ticker, counting letters and digits only. A code that LSEG marks as delisted, with ^, must also share a word of the company name. Otherwise ACE Ltd, which took the ticker CB when it bought Chubb in 2016, would get the old Chubb Corp’s code in the files before 2016.
- Same name. A holding without a code from step 1 gets the code with the same company name, ignoring punctuation and common words such as INC, CORP and CLASS.
- One to one. A holding with two codes, or a code found for two holdings on the same date, gets no code.
# Each LSEG code's company name, which the stock file gives on the code's first row only
named_rows = stocks.dropna(subset=["Company Name"]).drop_duplicates("Instrument")
lseg_name = named_rows.astype(str).set_index("Instrument")["Company Name"]
# The codes trading on the days of the holdings files
on_file_days = stocks["Date"].isin(holdings["date"].dt.strftime("%Y-%m-%d"))
trading = stocks.loc[on_file_days, ["Instrument", "Date", "Ticker"]].astype(str)
trading = trading.rename(columns={"Instrument": "ric"})
trading["date"] = pd.to_datetime(trading["Date"])
COMMON_WORDS = {"INC", "CORP", "CORPORATION", "CO", "COMPANY", "COMPANIES", "LTD", "PLC", "LLC",
"LP", "THE", "CLASS", "CL", "A", "B", "C", "NV", "SA", "AG", "REIT", "HOLDINGS",
"HOLDING", "HLDGS", "GROUP", "INTL", "INTERNATIONAL"}
def name_words(name):
"""A company name in capitals without punctuation and common words: 'Apple, Inc.' -> 'APPLE'."""
words = re.sub(r"[^A-Z0-9 ]", " ", str(name).upper()).split()
return " ".join(word for word in words if word not in COMMON_WORDS)
# Tickers in capitals, with letters and digits only
holdings["key"] = holdings["Ticker"].str.upper().str.replace(r"[^A-Z0-9]", "", regex=True)
trading["key"] = trading["Ticker"].str.upper().str.replace(r"[^A-Z0-9]", "", regex=True)
holdings["holding_words"] = holdings["Name"].map(name_words)
# A code without a company name gets a missing value (NaN)
trading["lseg_words"] = trading["ric"].map(lseg_name.map(name_words))
# Step 1: the same ticker on the same date. A delisted code must also share a word of the
# name; set(...) & set(...) keeps the words that two names have in common
by_ticker = holdings.merge(trading[["date", "key", "ric", "lseg_words"]], on=["date", "key"])
shares_a_word = []
for holding_words, lseg_words in zip(by_ticker["holding_words"], by_ticker["lseg_words"]):
shares_a_word.append(len(set(holding_words.split()) & set(str(lseg_words).split())) > 0)
is_delisted = by_ticker["ric"].str.contains("^", regex=False)
has_lseg_name = by_ticker["lseg_words"].notna()
by_ticker = by_ticker.loc[~is_delisted | ~has_lseg_name | np.array(shares_a_word)]
# Step 2: the same name on the same date, for the holdings without a code from step 1.
# indicator=True adds the column _merge, which says whether a row was found in both tables
found = holdings.merge(by_ticker[["date", "Ticker", "Name"]].drop_duplicates(),
on=["date", "Ticker", "Name"], how="left", indicator=True)
without_code = found.loc[found["_merge"] == "left_only"].drop(columns="_merge")
by_name = without_code.merge(trading[["date", "ric", "lseg_words"]].dropna(),
left_on=["date", "holding_words"], right_on=["date", "lseg_words"])
# Step 3: one code per holding and one holding per code on each date. duplicated(keep=False)
# marks every row whose key appears more than once
codes = pd.concat([by_ticker, by_name])[["date", "Ticker", "Name", "ric"]].drop_duplicates()
one_code = ~codes.duplicated(["date", "Ticker", "Name"], keep=False)
one_holding = ~codes.duplicated(["date", "ric"], keep=False)
codes = codes.loc[one_code & one_holding]
holdings = holdings.merge(codes, on=["date", "Ticker", "Name"], how="left")
# {:.1%} writes a share as a percentage with one decimal
print(f"{holdings['ric'].notna().mean():.1%} of the holdings have an LSEG code")87.1% of the holdings have an LSEG code
A holding without a code has no returns in the stock file, so it cannot enter the test.
Additions and deletions
A stock counts as held on both dates when the other file has its code or its ticker and name, which keeps a stock whose code was found on one date only from counting as a change. The same comparison of EUSA’s files gives the parent index’s changes. A stock is delisted within three months when LSEG marks its code as delisted and its last row in the stock file is at most three months after the effective day, the trading day after RD.
def holdings_not_in(fund_holdings, side, other_side):
"""The holdings of one side that the other side does not hold, by code or by ticker and name.
how="left" keeps every holding of this side, and indicator adds a column that says whether
the other side has the same code or the same ticker and name ("both") or not ("left_only").
"""
these = fund_holdings.loc[fund_holdings["side"] == side]
others = fund_holdings.loc[fund_holdings["side"] == other_side]
these = these.merge(others[["rd", "ric"]].dropna().drop_duplicates(), on=["rd", "ric"],
how="left", indicator="same_code")
these = these.merge(others[["rd", "Ticker", "Name"]].drop_duplicates(),
on=["rd", "Ticker", "Name"], how="left", indicator="same_name")
return these.loc[(these["same_code"] == "left_only") & (these["same_name"] == "left_only")]
# Additions: held after RD and not before. Deletions: held before RD and not after
usmv = holdings.loc[holdings["fund"] == "USMV"]
usmv_changes = pd.concat([holdings_not_in(usmv, "after", "before").assign(change="Addition"),
holdings_not_in(usmv, "before", "after").assign(change="Deletion")])
# The parent index's changes at the same rebalance
eusa = holdings.loc[holdings["fund"] == "EUSA"]
eusa_changes = pd.concat([holdings_not_in(eusa, "after", "before"),
holdings_not_in(eusa, "before", "after")])
# Leave out the USMV changes that the parent index also made: the same code, or else the same
# ticker and name, among EUSA's changes
events = usmv_changes[["rd", "ric", "Ticker", "Name", "change"]].merge(
eusa_changes[["rd", "ric"]].dropna().drop_duplicates(), on=["rd", "ric"], how="left",
indicator="parent_code")
events = events.merge(eusa_changes[["rd", "Ticker", "Name"]].drop_duplicates(),
on=["rd", "Ticker", "Name"], how="left", indicator="parent_name")
events = events.loc[(events["parent_code"] == "left_only")
& (events["parent_name"] == "left_only")]
# A stock without an LSEG code has no returns in the stock file
events = events.dropna(subset=["ric"])[["rd", "ric", "change"]]
# Leave out the stocks delisted within three months: LSEG marks a delisted code with ^, and the
# month of the stock's last row is at most three months after the month of the effective day,
# the trading day after RD. Dates written as year-month-day sort in date order as text
event_rows = stocks.loc[stocks["Instrument"].isin(events["ric"]), ["Instrument", "Date"]]
last_day = event_rows.astype(str).groupby("Instrument")["Date"].max()
events = events.merge(rebalances[["rd", "rd_number"]], on="rd")
events["effective_month"] = trading_days[events["rd_number"] + 1].to_period("M")
events["last_month"] = pd.to_datetime(events["ric"].map(last_day)).dt.to_period("M")
delisted_soon = (events["ric"].str.contains("^", regex=False)
& (events["last_month"] <= events["effective_month"] + 3))
events = events.loc[~delisted_soon, ["rd", "ric", "change"]]
print(f"{len(events)} index changes: "
f"{(events['change'] == 'Addition').sum()} additions and "
f"{(events['change'] == 'Deletion').sum()} deletions")695 index changes: 369 additions and 326 deletions
Read next
- Why the 60/40 portfolio is not dead. Another result that depends on the sample period: the stock-bond frontier across five-year windows.
- How to do factor investing How a factor portfolio is selected, weighted and rebalanced, with its trading costs.
- Trading costs more than halved Swedish momentum’s ending value What trading costs do to a strategy that rebalances often.
Disclaimer: an empirical study for discussion purposes, not investment advice. The design is descriptive and not causal. Past performance does not guarantee future returns.
Sources: Joop Huij and Georgi Kyosev, Price Response to Factor Index Additions and Deletions, 2016. Robeco, Guide to factor investing in equity markets, 2018. Rebalance days from MSCI’s index review press releases, and the announcement rule from MSCI’s methodology for its minimum-volatility indexes. Holdings of the iShares MSCI USA Min Vol Factor ETF (USMV) and the iShares MSCI USA Equal Weighted ETF (EUSA) from iShares. Daily total returns and value traded from LSEG Workspace (licensed); I do not reproduce the raw data here.