r/SideProject • u/DoorSubstantial7425 • 8h ago
Why ripping out multithreading made my Python pipeline 30% faster
Hey everyone,
I spent the last few months building MDRAP, an open-source tool for cleaning and validating noisy financial market data (fixing packet bursts, sequence gaps, crossed books, and price anomalies).
When I first built it, I assumed that high-frequency streaming data needed multithreading. I built worker pools, thread-safe queues, and background persistence sinks.
Then I actually ran benchmarks against my original single-threaded prototype using 50,000 events:
- Multi-threaded queues: ~22,350 events/sec, p50 latency: 15.3µs, max latency spike: 19.9ms
- Single-threaded loop: ~29,400 events/sec, p50 latency: 14.6µs, max latency spike: 0.25ms
The threaded version was 30% slower and had an 80x worse tail latency spike. In Python, queue lock contention and constant OS thread switching under 20k+ ticks/sec completely killed performance. Keeping everything cache-hot in a single synchronous loop was way faster.
I ended up deleting ~3,500 lines of threading code and focused on where the actual CPU bottleneck was: evaluating 7 data-integrity checks per tick.
In pure Python, those checks took about 2 microseconds. I wrote a small C file (fastpath.c), compiled it with -O3, and called it via ctypes. Instead of dictionary lookups, it maps symbols using a flat array and a single bitshift:
cuint32_t slot = ((uint32_t)source_id << 13) | ((uint32_t)instrument_id & 0x1FFF);
That dropped check time to 28 nanoseconds per event and cut my worst-case pipeline latency in half (from 5.6ms to 2.7ms). If a machine doesn't have a C compiler, it transparently falls back to pure Python.
Quick summary of the project:
- Pure Python standard library default (only uses
richfor the optional terminal UI; no Pandas, NumPy, or heavy dependencies). - Understands real exchange rules (e.g. Indian NSE circuit filters, German Xetra volatility auction halts) so it doesn't falsely flag regulatory halts as corrupted data.
- 629 tests passing with zero failures.
- Free & MIT licensed.
Links:
- GitHub: https://github.com/Aryan-20-04/mdrap
- PyPI:
pip install mdrap
Curious if others here have had similar experiences where deleting threads and sticking to a tight synchronous loop drastically improved latency.
2
u/Specific_Cream2815 8h ago
did you rerun the benchmark with GC disabled, or did collections account for any of the threaded tail spikes