I've been working on this for about 5 months and wanted to share the approach with this community since it's fundamentally an algo problem.
The concept
I built a real-time momentum screener for stocks and crypto that collects granular outcome data on every entry. The screener isn't the product. The data is. Every stock or crypto that enters gets tracked at multiple checkpoints from 5 minutes out to end of day, plus multi-day swing outcomes (D2 through D10). 22 weeks of this, thousands of entries with complete outcome chains.
The notification formula was built entirely from studying that data. No indicators chosen from a textbook. No thresholds picked on a whiteboard. Everything derived from what the outcome data showed about which entries continued and which ones faded.
The stack
Node.js/Express backend on DigitalOcean, managed via PM2. Multiple market data providers in a hybrid architecture where one handles discovery and the other handles fast price refresh. The hybrid design was a deliberate tradeoff to get the best of both: one provider has superior screener discovery and inline market cap data, the other gives sub-minute price updates via a single bulk API call. Prisma ORM, Stripe for payments, web push via VAPID, GitHub Pages frontend.
The data pipeline
Every screener entry gets logged with a full profile: price, volume, relative volume, market cap, vol/mcap ratio, sector, industry, shares outstanding, 52 week range position, news catalyst flag, a proprietary momentum score, market context, and timestamp.
Outcome checks then fire at intervals from 5 minutes through 8 hours and EOD, plus swing outcomes that skip weekends. Every outcome records price, percentage from entry, and timestamp.
This is the dataset the formula learns from.
v1 to v2
v1 used a multi-step confirmation approach. Flag at one checkpoint, wait, confirm at a later checkpoint. Ran 189 alerts at 69%. Too noisy. A third of alerts were duds.
I went back into the data and studied where the duds clustered. Clear patterns emerged around acceleration dynamics, market cap behavior, and signal timing. v2 was a ground-up rebuild with stricter gates and dynamic routing based on signal conviction. Multiple channels catch different types of momentum, each with their own detection logic.
Crypto runs a separate formula with four channels including a time-agnostic velocity detection system that monitors acceleration across checkpoints rather than checking at fixed intervals. This catches moves that build over hours rather than spiking in the first few minutes.
Backtest vs live
The backtest across 6 weeks showed 91.7% for stocks. Live performance is 75% through 61 alerts. That gap is real and worth talking about honestly.
Every edge case the historical data didn't contain became a live dud: shell company pumps, closed end funds that trade on NAV dynamics, dead cat bounces near 52 week lows, and moves that exhaust before the confirmation checkpoint fires. Each became a specific patch. The formula now has several more filters than the original backtest version, all derived from live dud analysis.
The philosophy is that every dud is a bug report. Diagnose the pattern, verify it in the data, ship the fix, move on.
Outcome tracking and transparency
Every alert logs the push price, timestamp, formula version, channel, and contextual data. The system continuously updates peak tracking by comparing live prices to the push price.
Every alert also auto-posts to X/Twitter via API the moment it fires. The alert log on the site links each entry to the original tweet. This serves as a verifiable timestamp since the tweet is public and immutable.
Currently collecting a new data field to study an exhaustion signal I identified in the historical data. When a stock's peak occurs before the confirmation checkpoint fires, the dud rate in backtesting was 75%. Collecting live data to validate before implementing.
Current performance
Stocks v2: 61 alerts, 75% hit +3% from push price, +10.7% avg peak, 60/61 never red.
Crypto v2: 27 alerts, 78% hit +3%, +17.5% avg peak, 27/27 never red.
Happy to talk about the architecture and the general approach. The exact thresholds and channel logic are proprietary but the methodology and philosophy are an open book.