Institutional crowding, measured from 13 years of SEC Form 13F filings — and whether it predicts how securities behave under stress.
When many managers hold the same position at high portfolio weight, a forced seller imposes losses on everyone else holding it, and their losses force more selling. That exposure is invisible in any single filing. It is visible across all of them.
Full panel, every quarter, no event selection. Fama-MacBeth over 82,191 security-quarters across 52 quarters, 2,896 tickers:
| Dependent variable | Ranking | t | Reading |
|---|---|---|---|
| Next-quarter return | within-sector | +2.04 | crowded names outperform in calm |
| Worst-10-day return | within-sector | −5.93 | the same names underperform in stress |
Negative in 84% of 51 quarters for the stress measure. Crowding pays in calm and costs in stress — a conditional risk premium, and the actual fire-sale signature.
Corroborated out-of-sample on three unrelated shocks, with positioning publicly filed before each window opened: COVID 2020 (t = −2.35), the 2022 rate shock (t = −3.79), and the 2024 yen-carry unwind (t = −4.20).
The measure has one central defect: a sovereign wealth fund holding 75% of a company looks identical to forty hedge funds in the same trade. Both score high; only one sells into a drawdown.
No formula change fixes this, because the answer is not in the number — it is in who the holders are. So an agent investigates: it pulls the holder list, classifies the holders that matter, and re-derives the score without committed capital.
| Security | Headline | Adjusted | Reading |
|---|---|---|---|
| GlobalFoundries | 97% | 21% | Mubadala is 75.3% of the institutional base |
| Morgan Stanley | 95% | 24% | overstated |
| Symbotic | 97% | 97% | genuine crowding |
| Figma | 98% | 99% | genuine crowding |
The loop branches on evidence — how many holders get classified depends on how concentrated ownership turns out to be, and whether it simulates a liquidation depends on what the classifier returned. The policy layer is pluggable: a transparent rule-based policy generated every published trace with no API key; an LLM policy implements the same two methods.
13F's VALUE field is reported in dollars by some filers and thousands by others,
with no flag. Microsoft at 31 Dec 2019, true close $157.70:
| Filer | VALUE / shares | Implied price | Reporting in |
|---|---|---|---|
| UBS Asset Management | 2,308,218,628 / 14,636,770 | $157.70 | dollars |
| Vanguard Group | 95,895,803 / 608,090,061 | $0.16 | thousands |
Vanguard holds the largest position in the security and appears 1000× too small. Any
model on raw VALUE is wrong, and biased — the understaters are the largest index
managers. The fix recovers scale per row from implied price and validates against known
closes. The correction rate is regime-dependent (84% of rows in 2019Q4, 7% by 2024Q2,
following the SEC's 2023 change), so a hardcoded multiplier fails in one direction or
the other.
python -m venv .venv && .venv/bin/pip install -r requirements.txt
.venv/bin/python crowdrisk/ingest/fetch_all.py # ~2.9 GB, 53 quarters
.venv/bin/python crowdrisk/ingest/build_panel.py # ~8 min
.venv/bin/python crowdrisk/ingest/holdings_store.py
.venv/bin/python crowdrisk/model/score.py
.venv/bin/python crowdrisk/agent/investigator.py GFSSEC requires a User-Agent header identifying you on every request — replace the
placeholder contact in parse.py and figi.py before running at volume.
DESIGN.md covers the full architecture: every data trap and its fix, the measure and why it is rank-transformed within sector, all validation, the agent's tools and loop, and the hypotheses that failed — including a liquidity term that entered with the wrong sign and was dropped rather than kept for the story.
Everything is SEC public domain, free and redistributable. Total cost: $0.
A research and engineering exercise. Not investment advice.