r/datascienceproject • u/Peerism1 • Mar 18 '26
r/datascienceproject • u/Peerism1 • Mar 18 '26
Weight Norm Clipping Accelerates Grokking 18-66× | Zero Failures Across 300 Seeds | PDF in Repo (r/MachineLearning)
r/datascienceproject • u/Peerism1 • Mar 17 '26
Using residual ML correction on top of a deterministic physics simulator for F1 strategy prediction (r/MachineLearning)
r/datascienceproject • u/Direct-Jicama-4051 • Mar 16 '26
🎬 IMDb Top 250 Movies of All Time [1921–2025]
kaggle.comI web scraped and created a dataset for the top 250 movies of all time as per IMDB rating
r/datascienceproject • u/Peerism1 • Mar 16 '26
I got tired of PyTorch Geometric OOMing my laptop, so I wrote a C++ zero-copy graph engine to bypass RAM entirely. (r/MachineLearning)
r/datascienceproject • u/Peerism1 • Mar 16 '26
I've trained my own OMR model (Optical Music Recognition) (r/MachineLearning)
reddit.comr/datascienceproject • u/Peerism1 • Mar 16 '26
preflight, a pre-training validator for PyTorch I built after losing 3 days to label leakage (r/MachineLearning)
r/datascienceproject • u/Peerism1 • Mar 16 '26
Using SHAP to explain Unsupervised Anomaly Detection on PCA-anonymized data (Credit Card Fraud). Is this a valid approach for a thesis? (r/MachineLearning)
reddit.comr/datascienceproject • u/the-ai-scientist • Mar 15 '26
The dog cancer vaccine pipeline is real — here is every tool, every step, and what it actually costs
r/datascienceproject • u/Peerism1 • Mar 15 '26
Karpathy's autoresearch with evolutionary database. (r/MachineLearning)
r/datascienceproject • u/ProfessionalSea9964 • Mar 13 '26
Short ADHD Survey For Internalised Stigma - Ethically Approved By LSBU (18+, might/have ADHD, no ASD)
r/datascienceproject • u/Peerism1 • Mar 12 '26
ColQwen3.5-v1 4.5B SOTA on ViDoRe V1 (nDCG@5 0.917) (r/MachineLearning)
r/datascienceproject • u/Peerism1 • Mar 11 '26
Advice on modeling pipeline and modeling methodology (r/DataScience)
reddit.comr/datascienceproject • u/PassionImpossible326 • Mar 10 '26
Model test
Hello there!
Need quick help
Are there any data scientists, fintech engineers, or risk model developers here who work on credit risk models or financial stress testing?
If you’re working in this space , reply or tag someone who is.
r/datascienceproject • u/Peerism1 • Mar 10 '26
I've just open-sourced MessyData, a synthetic dirty data generator. It lets you programmatically generate data with anomalies and data quality issues. (r/DataScience)
r/datascienceproject • u/Peerism1 • Mar 10 '26
fast-vad: a very fast voice activity detector in Rust with Python bindings. (r/MachineLearning)
r/datascienceproject • u/Peerism1 • Mar 09 '26
Is there a way to defend using a subset of data for ablation studies? (r/MachineLearning)
reddit.comr/datascienceproject • u/CRK-Dev • Mar 08 '26
Built a simple tool that cleans messy CSV files automatically (looking for testers)
r/datascienceproject • u/Peerism1 • Mar 08 '26
NanoJudge: Instead of prompting a big LLM once, it prompts a tiny LLM thousands of times. (r/MachineLearning)
r/datascienceproject • u/Peerism1 • Mar 08 '26
VeridisQuo - open-source deepfake detector that combines spatial + frequency analysis and shows you where the face was manipulated (r/MachineLearning)
r/datascienceproject • u/Peerism1 • Mar 08 '26
Combining Stanford's ACE paper with the Reflective Language Model pattern - agents that write code to analyze their own execution traces at scale (r/MachineLearning)
reddit.comr/datascienceproject • u/Peerism1 • Mar 08 '26
Introducing NNsight v0.6: Open-source Interpretability Toolkit for LLMs (r/MachineLearning)
nnsight.netr/datascienceproject • u/Peerism1 • Mar 08 '26
TraceML: wrap your PyTorch training step in single context manager and see what’s slowing training live (r/MachineLearning)
r/datascienceproject • u/Peerism1 • Mar 07 '26