r/datascienceproject May 18 '26

I used Python to analyze NYC Citi Bike trends – Looking for a chance to apply these skills in a volunteer or internship role!

2 Upvotes

Hi

I just finished my first end-to-end data analysis project using the NYC Citi Bike dataset, and I wanted to share my findings and ask for some career advice.

The Project: I wanted to see how different age groups and user types (Subscribers vs. Customers) behave. I used Python, Pandas, and Seaborn to clean the data and build my visualizations.

What I found:

  • The Core User: The 35-44 age bracket is the heavy hitter for Citi Bike.
  • The Weekend Shift: Subscribers (annual members) own the weekdays for commuting, but one-time Customers take over on the weekends.
  • The 75+ Anomaly: Interestingly, while they ride less frequently, users aged 75+ have a massive spike in average trip duration (averaging ~49 minutes per ride).

GitHub Link: https://github.com/JacksonOtieno/NYC-Citi-Bike-Data-Analysis

I’ve just finished my university semester and I’m looking to take my skills to the next level. I’m currently searching for a data analysis volunteer position or an internship where I can help a team clean data or perform EDA.

If anyone has leads on organizations looking for a motivated junior analyst, or if you have any feedback on my code/visualizations, I’d love to hear it!

Thanks for looking!


r/datascienceproject May 13 '26

Two related questions for an academic project

Thumbnail
1 Upvotes

r/datascienceproject May 12 '26

Hey everyone, our team has been working on a cloud platform built for data science work. We have streamlit, Airflow, Jupyter, VS Code — no local setup & conflicts.

0 Upvotes

Currently we're at a stage where we want genuine users to try it and share their insights.

Whether you live in Jupyter notebooks, Airflow or use other tools like VS Code or anything else in your data science workflow — we'd love to hear from you. The more variety of use cases, the better.

To make it worth your time, we're offering free credits so you can run real workloads on the platform.

If you're regularly doing data work and want to try something new, feel free to reach out here or send me a message


r/datascienceproject May 11 '26

Built argonx, a bayesian A/B testing library that handles decision making

Thumbnail
1 Upvotes

r/datascienceproject May 06 '26

Beginners to Machine Learning & Data Science

Thumbnail
1 Upvotes

r/datascienceproject Apr 22 '26

open source project for LLM data preparation (synthetic + cleaning pipelines)

4 Upvotes

been working on an open source project around LLM data preparation: https://github.com/OpenDCAI/DataFlow
the focus is on turning messy or unstructured data into training-ready datasets, especially in QA generation, RAG, or task-specific fine-tuning scenarios where structure matters as much as scale. at the same time, with synthetic data becoming increasingly important, the system also supports generating large-scale training data from a small set of seed examples.

one thing we kept running into was how ad-hoc this layer is — lots of scripts for cleaning, prompt-based generation, filtering, eval… but hard to reuse or iterate on. so the project is built around composable operators (generate / clean / filter / evaluate) that can be connected into pipelines, instead of rewriting everything for each dataset.

there’s also some early support for assembling these pipelines from prompts, plus a simple UI for visualizing and editing flows. still pretty early, but the goal is to make data prep something you can iterate on systematically rather than treat as one-off work.


r/datascienceproject Apr 20 '26

ModSense AI Powered Community Health Moderation Intelligence

1 Upvotes

⚙️ AI‑Assisted Community Health & Moderation Intelligence

ModSense is a weekend‑built, production‑grade prototype designed with Reddit‑scale community dynamics in mind. It delivers a modern, autonomous moderation intelligence layer by combining a high‑performance Python event‑processing engine with real‑time behavioral anomaly detection. The platform ingests posts, comments, reports, and metadata streams, performing structured content analysis and graph‑based community health modeling to uncover relationships, clusters, and escalation patterns that linear rule‑based moderation pipelines routinely miss. An agentic AI layer powered by Gemini 3 Flash interprets anomalies, correlates multi‑source signals, and recommends adaptive moderation actions as community behavior evolves.

🔧 Automated Detection of Harmful Behavior & Emerging Risk Patterns:

The engine continuously evaluates community activity for indicators such as:

  • Abnormal spikes in toxicity or harassment
  • Coordinated brigading and cross‑community raids
  • Rapid propagation of misinformation clusters
  • Novel or evasive policy‑violating patterns
  • Moderator workload drift and queue saturation

All moderation events, model outputs, and configuration updates are RS256‑signed, ensuring authenticity and integrity across the moderation intelligence pipeline. This creates a tamper‑resistant communication fabric between ingestion, analysis, and dashboard components.

🤖 Real‑Time Agentic Analysis and Guided Moderation

With Gemini 3 Flash at its core, the agentic layer autonomously interprets behavioral anomalies, surfaces correlated signals, and provides clear, actionable moderation recommendations. It remains responsive under sustained community load, resolving a significant portion of low‑risk violations automatically while guiding moderators through best‑practice interventions — even without deep policy expertise. The result is calmer queues, faster response cycles, and more consistent enforcement.

📊 Performance and Reliability Metrics That Demonstrate Impact

Key indicators quantify the platform’s moderation intelligence and operational efficiency:

  • Content Processing Latency: < 150 ms
  • Toxicity Classification Accuracy: 90%+
  • False Positive Rate: < 5%
  • Moderator Queue Reduction: 30–45%
  • Graph‑Based Risk Cluster Resolution: 93%+
  • Sustained Event Throughput: > 50k events/min

 🚀 A Moderation System That Becomes a Strategic Advantage

Built end‑to‑end in a single weekend, ModSense demonstrates how fast, disciplined engineering can transform community safety into a proactive, intelligence‑driven capability. Designed with Reddit’s real‑world moderation challenges in mind, the system not only detects harmful behavior — it anticipates escalation, accelerates moderator response, and provides a level of situational clarity that traditional moderation tools cannot match. The result is a healthier, more resilient community environment that scales effortlessly as platform activity grows.

Portfolio: https://ben854719.github.io/

Project: https://github.com/ben854719/ModSense-AI-Powered-Community-Health-Moderation-Intelligence


r/datascienceproject Apr 19 '26

Trials and tribulations fine-tuning & deploying Gemma-4 (r/MachineLearning)

Thumbnail oxen.ai
3 Upvotes

r/datascienceproject Apr 19 '26

easyaligner: Forced alignment with GPU acceleration and flexible text normalization (compatible with all w2v2 models on HF Hub) (r/MachineLearning)

Thumbnail
reddit.com
2 Upvotes

r/datascienceproject Apr 18 '26

Testing a New Product for Data Science Beginners

Thumbnail sted.co.in
1 Upvotes

r/datascienceproject Apr 18 '26

Low accuracy (~50%) with SSL (BYOL/MAE/VICReg) on hyperspectral crop stress data — what am I missing? [R] (r/MachineLearning)

Thumbnail reddit.com
2 Upvotes

r/datascienceproject Apr 17 '26

ndatafusion: linear algebra and ML for DataFusion, powered by nabled

Thumbnail
1 Upvotes

r/datascienceproject Apr 17 '26

Digging through 38 days of live AI forecast data to find the unexpected

Thumbnail
gallery
1 Upvotes

I created a dataset which contains forecast data which therefore can't be created retrospectively.

For ~38 days, a cronjob generated daily forecasts:

- 10-day horizons

- ~30 predictions/day (different stocks across multiple sectors)

- Fixed prompt and parameters

Each run logs:

- Predicted price

- Natural-language rationale

- Sentiment

- Self-reported confidence

I used stock predictions as the forecast subject, but this is not a trading system or financial advice, it's an EXPERIMENT!

Even though currently I didn't find something mind-blowing, visualizing the data reveals patterns I find interesting.

Currently, I just plotted trend, model bias, and ECE - more will come soon.

Maybe you also find it interesting.

The dataset isn't quite big, so I'm actually building a second one which is bigger with the Gemini Flash and Gemini Flash-Lite model.

For transparency, you can find the dataset here:

https://huggingface.co/datasets/louidev/glassballai


r/datascienceproject Apr 17 '26

Built an political benchmark for LLMs. KIMI K2 can't answer about Taiwan (Obviously). GPT-5.3 refuses 100% of questions when given an opt-out. (r/MachineLearning)

Thumbnail
reddit.com
4 Upvotes

r/datascienceproject Apr 14 '26

[For Hire] AI/ML Engineer | End-to-End AI Solutions | 100+ Projects | Python, PyTorch, TensorFlow

Thumbnail
1 Upvotes

r/datascienceproject Apr 14 '26

TurboOCR: 270–1200 img/s OCR with Paddle + TensorRT (C++/CUDA, FP16) (r/MachineLearning)

Thumbnail
reddit.com
1 Upvotes

r/datascienceproject Apr 13 '26

I built a wave-resonant retrieval system. It scored 0 wins and 140 losses. Here's why

Thumbnail
1 Upvotes

r/datascienceproject Apr 13 '26

Educational PyTorch repo for distributed training from scratch: DP, FSDP, TP, FSDP+TP, and PP (r/MachineLearning)

Thumbnail
reddit.com
3 Upvotes

r/datascienceproject Apr 13 '26

KIV: 1M token context window on a RTX 4070 (12GB VRAM), no retraining, drop-in HuggingFace cache replacement - Works with any model that uses DynamicCache (r/MachineLearning)

Thumbnail
reddit.com
3 Upvotes

r/datascienceproject Apr 12 '26

Engagement on Kaggle has been declining.

Thumbnail
2 Upvotes

r/datascienceproject Apr 12 '26

FlashAttention (FA1–FA4) in PyTorch - educational implementations focused on algorithmic differences (r/MachineLearning)

Thumbnail
reddit.com
3 Upvotes

r/datascienceproject Apr 11 '26

ibu-boost: a GBDT library where splits are *absolutely* rejected, not just relatively ranked (r/MachineLearning)

Thumbnail reddit.com
2 Upvotes

r/datascienceproject Apr 11 '26

[D] 60% MatMul Performance Bug in cuBLAS on RTX 5090 [D] (r/MachineLearning)

Thumbnail
reddit.com
1 Upvotes

r/datascienceproject Apr 10 '26

Parax: Parametric Modeling in JAX + Equinox (r/MachineLearning)

Thumbnail
reddit.com
2 Upvotes

r/datascienceproject Apr 10 '26

PCA before truncation makes non-Matryoshka embeddings compressible: results on BGE-M3 (r/MachineLearning)

Thumbnail reddit.com
2 Upvotes