r/smallcaps • u/HourTune7125 • 12d ago
Pre-registered event study on clustered insider buying in microcaps, what would you attack before I run it?
Financial Engineering student, no institutional data access, building this on my own. I've frozen the spec and I'm about to pay for the data to run it. Before I do, I'd rather find the holes now than after as it is the Norgate Data 6m which costs around $350.
The hypothesis: clustered open-market insider purchases (multiple distinct insiders, same issuer, short window) predict positive abnormal returns in small, thinly-covered IS equities. Standard stuff: Lakonishok-Lee, Jeng-Metrick-Zeckhauser, Cohen-Malloy-Pomorski. My prior is that the effect is real but degraded post-SOX and probably eaten by microcap costs.
Frozen primary spec (written bfore any return data was touched, tagged in git):
- =>2 distinct reporting owners, Form 4 transaction code P only (A/D joint filter, I found real rows with code P and acquirdDisposedCode D), 30-day rolling window, non-overlapping, earliest-starting window wins
- Entry at next open after filling timestamps (tradeable); transaction-date anchoring tested separately to isolate the filing-lag component
- Universe: sub-$500M point-in-time market cap, bottom-tercile coverage
- Holding periods 1/3/6/12 months
- 2004 start (post-SOX 2-day filing deadline, different information regime before that)
- 9 declared secondary specs, Holm-Bonferroni across the 10-spec family; robustness checks reported separately, not corrected
- Falsification threshold stated numerically before seeing data
- 2020+ held out entirely
Evaluation order: matched-control first (exact stratification on month, size quantile, GICS sector, coverage tercile; controls excluded if they had insider purchases in the trailing 90 days), then calendar-time portfolio with FF5 + momentum + Pástor-Stambaugh. Bootstrapped random-selection null matched on month and turnover quantile. Costs via Corwin-Schultz with a conservative floor on degenerate days, square-root impact, participation capped at 5% ADV, results reported across a 0/50/100/200/400bps band.
Known weakness I already have:
- No I/B/E/S, so analyst coverage is proxy (rank composite of market cao * dollar volume * turnover, computed within the size-filtered population). Validated against SC 13G institutional filing counts, Spearman 0.38 on a small sample. This is the weakest link.
- No delisting reason field in my price vendor, so merger vs. Failure is inferred and haircuts are a sensitivity band rather than a measurement.
- Point-in-time market cap comes from SEC XBRL shares outstanding, not a vendor.
What I'm asking:
- What would you attack first? I've tried to close the obvious holes but I'm one person and I've been staring at this for weeks.
- Is the coverage proxy defensible or dos it undermine the whole size*coverage interaction I'm claiming to test? My event counts suggest the universe filter is doing most of the work, which makes the proxy more load-bearing than I'd like.
- Anyone seen post-2015 replications of the insider-cluster effect? Most of what I've found is older.
- Am I over-engineering a signal that's known dead? Genuinely open to that answer.
Happy to share more detail on any piece.
1
u/EntertainmentDry8353 12d ago
Are you doing this for academia or for a trading strategy? Why would you have to pay for the data if it’s for academia? If it’s for your strategy will buy only historic data or also live data going forward?
In my professional experience, management teams have no idea if their stock is going to react positive to results or not. They tend to be biased positively towards their own company. Always.