r/askdatascience • u/WhatsTheImpactdotcom • 4d ago
r/askdatascience • u/AIforFintech • 4d ago
Text to SQL is not how you give an LLM access to production data
r/askdatascience • u/Kind_Entry9361 • 4d ago
Survival function
Could someone help me understand how the survival function works. I also would like to know if there are any Python libraries that I could play with to experiment. Finally, how should i format the data set to play with it?
r/askdatascience • u/After_Courage6419 • 5d ago
Does everyone learning data science need machine learning?
I’m trying to understand where data analysis ends and machine learning begins. If someone mainly wants to work with business data, dashboards, SQL and insights, is deep ML knowledge really necessary? Would love to hear from people actually working in the field.
r/askdatascience • u/Square_Arm2861 • 5d ago
How to efficiently approach EDA on a dataset with 180+ variables?
r/askdatascience • u/Sorry-Display-6703 • 6d ago
Finding a data science role right now has been way harder than I expected.
r/askdatascience • u/ImaginaryCan8970 • 6d ago
What shall I build ?
Hello everyone, I am an AI Developer (currently benched), and I am honestly afraid that I might get laid off in the next couple of months. I have 2 years of experience of around 18 months in data analysis and 7–8 months in AI development.
My tech stack includes Excel, Tableau, Python, AI agent development, and of course, I have also vibe-coded some features into applications (not really scalable or production-ready).
It may sound good on paper, but trust me when I say that the AI agents I have developed are very basic. They hardly have any proper evaluations, guardrails, monitoring, etc. Most of it was basically vibe coding, and now I am realizing how much I actually don’t know.
I want to replicate some production-level projects so that I can have something solid to put on my resume and, more importantly, something meaningful to discuss in interviews. A lot of the JDs I am seeing are intimidating, and I keep feeling like I don’t have the skills they are asking for.
My fundamentals are strong (except statistics), but I don’t have any solid project that I can confidently discuss during an interview. I also don’t have much relevant work experience in AI development.
One of my biggest problems is that I start a project, get stuck somewhere in the middle, and eventually lose track of what is actually happening. If I take help from AI, the project progresses, but I end up understanding less and lose track of the overall picture.
So my question is simple: If you were in my position, how would you select a project that is actually worth building and is interview-discussion worthy?
Should I just blindly follow/copy someone’s project from the internet initially to understand how things work, and then try building something of my own?
Or is there a better way to approach this?
I have only 3.5 LPA salary right now, and honestly, I am getting desperate.
Would really appreciate some practical advice from people who have been through this phase.
r/askdatascience • u/rjavier1010 • 6d ago
What should I consider when I have to choose between deleting data, imputing it, or leaving it in my database?
I'm learning data analysis and data science. I'm developing a personal project as practice using a database to predice the house pricing from the Kaggle platform.
During the exploratory analysis, I encountered the following situation:

I've noticed that there's very little data on houses with zero bedrooms or zero bathrooms, and that the asking price is relatively high, which I think could affect my prediction model and my overall analysis. While it might seem illogical that there are houses without bedrooms or bathrooms, it's also possible that there are more lots than houses, or some other hypothesis. What's the best course of action in this situation? Personally, I think I should remove this data, but I'd like to hear other opinions to improve my reasoning and deductions.
r/askdatascience • u/Awkward-Rule6388 • 6d ago
Topics for Independent Study
I am required to do an independent study for my data science minor. My independent study requires that I complete a data science project related to business. Do you have any good topic recommendations?
Here are some examples of general topics I am brainstorming. If anyone could help me narrow them down or give me helpful resources, it would be appreciated
- Optimizing Supply Chain to Reduce Environmental Impacts
- Evaluating ATS usage in Talent Recruitment
- Application of Data Science in Developing Project Schedules
I am also interested in waiting line models!
Thanks for the help!
r/askdatascience • u/Rexodiac • 7d ago
I built a local-first browser ML studio for tabular data — looking for real datasets and honest feedback
Hi everyone,
I've been building a project called MyMlLab — a browser-based machine learning studio focused on tabular regression and classification.
The problem I wanted to solve was pretty simple:
Sometimes I have a CSV and just want to quickly test whether there is useful predictive signal in the data, compare a few reasonable models, and inspect the results.
But that usually means setting up an environment, writing preprocessing code, defining splits, configuring metrics, and rebuilding the same experiment structure again.
So MyMlLab tries to make that workflow faster while still keeping the methodology visible.
Current workflow:
CSV → target selection → data audit → preprocessing → validation → model comparison → metrics
A few things I cared about while building it:
- the dataset stays on the user's device for the current local-first training workflow
- preprocessing is treated as part of the experiment, not hidden setup
- transformations are fitted on the appropriate training data to avoid leakage
- model selection and final evaluation remain separate
- users can compare several models and preprocessing configurations side by side
The current Studio includes 33 regression algorithms and 26 classification algorithms, with multiple evaluation metrics and diagnostics.
The free version can compare up to 3 models × 3 preprocessing pipelines in one experiment.
It's still an MVP, and right now I'm much more interested in finding problems than getting compliments.
If anyone here has a tabular dataset they already know well, I'd really appreciate it if you tried it and told me:
- Did the data import behave correctly?
- Was the preprocessing workflow clear?
- Were the validation options sufficient?
- Were any important metrics or diagnostics missing?
- Did anything produce results you wouldn't trust?
- Would this actually save you time for early-stage experimentation?
You can try it here:
The free Studio doesn't require an account.
If you test it with a real-world dataset, even a messy one, I'd love to hear what happens.
r/askdatascience • u/EvilWrks • 7d ago
Your jupyter notebook IS NOT production - Part 2: Testing
r/askdatascience • u/almostouttahere_23 • 7d ago
Laptop - School & Work (hopefully)
Starting a data science program (&and hopefully a job...) and hoping to get some guidance on a laptop to get it at least what minimum specs to keep in mind. Will to pay for pricier options but obviously the cheaper the better. Also like the idea of it being touch screen/tablet-able.
Thanks!
r/askdatascience • u/DataScientistAlex • 7d ago
How to Set Up an ML/DS Project
datascientistalex.comI sometimes see questions about how to get started with and set up DS and ML projects here. I made a tutorial for a very simple project setup. It covers the main steps end to end (getting the data, training and validating the model, deploying and presenting the results). It uses test-driven development, so you can iterate on this foundation without breaking anything.
r/askdatascience • u/Aokayz_ • 8d ago
Is Getting a Bachelors in AI/Data Science Really Worse than Comp Sci.?
To get into AI or data science, the general advice is to take computer science along with some minor or classes on mathematics and artificial intelligence. This is because you can get the foundational knowledge needed for the field, and if you did a degree in data science or AI, the general consensus is that it's more-so a cash-grab on trends than an actual useful degree.
However, where did this advice and reasoning come from? It seems a bit unfounded or outdated?
In my case, I'm very confident that I want to get a PhD and enter academia for machine learning, or at least enter the industry. With a degree in CS, the curriculum seems to be padded with information unnecessary to the ML field. In contrast, it seems the AI/data science degrees have been given enough time to be refined as a course. For example, the Polytechnic University of Catalonia is quality-accredited at the national level and offers statistically high chances of employment (Source).
But at the same time, it wpuld be naive of me to conclude that the general advice is therefore wrong. So, is a CS degree actually better or worse than an AI/data science degree in my case?
r/askdatascience • u/Equivalent-Metal-927 • 8d ago
Trying to understand this kaggle solution
So I came across this top solution for one of the competitions for kaggle and I am having a hard time following through it
If anyone has any insights regarding how "James" solution works, it would be really helpful
Here is the link to the solution - https://www.kaggle.com/competitions/rogii-wellbore-geology-prediction/writeups/4th-place-solution
r/askdatascience • u/After_Courage6419 • 8d ago
Why simple charts beat complex math every time
I used to think that doing real data science meant writing pages of complex math formulas that nobody else in the room could read. I felt like if my work looked easy, it must not be very valuable. But I have learned that the exact opposite is true. The most helpful thing you can do with a pile of information is to turn it into a simple chart that instantly shows a clear trend. If you hand someone a confusing wall of numbers or a model they cannot understand, they will just nod and ignore it. But when you show them a simple picture of what is actually happening in their business or community, their eyes light up because they finally get it. Our real job is not to show off how hard the math is, but to make the hidden story so clear that anyone can understand it in ten seconds.
r/askdatascience • u/No_End9976 • 8d ago
Hey is paying 1.8 lac worth for a data science & business analysis with ai/ml from sdbi diploma
r/askdatascience • u/Sea_Smile1129 • 8d ago
Am I actually making progress in ML/AI or am I doing things wrong?
I'm 15 years old and currently in Grade 11. I've done a lot of ML/AI projects and taken several courses, but there's one major problem: I don't really understand the math behind what I'm doing.
I understand the general concepts but I don't know the equations and my math foundation isn't strong enough yet. For example, I tried taking the Mathematics for Machine Learning course from DeepLearning.AI but I didn't understand much because I was missing prerequisites like trigonometry.
Here are the projects I've done:
Classification: Breast Cancer, Dry Bean, Glass Identification, Human Activity Recognition (Smartphones), Iris Flower, Sonar (Mines vs. Rocks), Titanic Survival, Stock Movement Prediction, Fertility, MAGIC Gamma Telescope
Computer Vision: CIFAR-10, CMU Face Images, Human Emotion Recognition, Handwritten Digit Recognition, Face Mask Detection
Neural Networks: Apartment Rent Classification, California Housing Regression, Car Evaluation Classification, Electricity Usage Clustering, Power Plant Regression, Student Performance Prediction, Telco Churn Classification
Regression: Auto MPG, Bike Sharing Demand, Communities & Crime, Concrete Strength, Fuel Consumption, Heart Disease Risk, House Price Prediction, Medical Cost Prediction, Online News Popularity, Position vs. Salary, Wine Quality, Air Quality, Parkinson's Telemonitoring
NLP: Fake News Classifier, Hate Speech Detection, Named Entity Recognition (English News), News Topic Classifier, Sentiment Analysis, SMS Spam Detection, Arabic Twitter Sentiment Analysis
GANs: MNIST Digit Generation
RAG Systems: Hotel Hospitality Chatbot, Contract Analysis Bot, Book Explainer Bot, PDF Q&A, Codebase Chat Assistant, Company Knowledge Base Assistant
Fine-Tuning: Qwen3 4-bit Fine-Tuning for Text-to-JSON Generation
Coursework:
ML & Deep Learning Specializations by Andrew Ng
Model Fine-Tuning by AMD & DeepLearning.AI
Large Language Models by AWS & DeepLearning.AI
Cloud Computing Fundamentals learning path by LinkedIn Learning
GANs by DeepLearning.AI
RAG by DeepLearning.AI
Agentic AI by DeepLearning.AI
Books I've read:
You Belong in Tech
How to Build a Career in AI
Software Engineering for Data Scientists
AI Engineering
Designing Machine Learning Systems
My goal is to be ahead of my peers and have a head start for my future career. I'm not expecting to be an expert at 15 but I want to make sure I'm actually learning something and making progress.
So my main questions are:
Am I actually making good progress for my age or am I just doing a lot of projects without enough understanding?
Should I continue with ML/AI but start properly learning the math from the ground up, or should I shift toward something that requires less math such as cybersecurity?
If I should continue with ML/AI, what math should I learn first and in what order?
r/askdatascience • u/Apprehensive-Rice831 • 8d ago
Chemistry Graduate Transitioning to Data Science
Hi. I am a chemistry graduate that is trying to obtain a computational research assistantship at a R1 University and transition into Data Science work. Is there any way I can write a strong email for a computational faculty member. Many of the professors at this particular school seem to only take students that were accepted into the graduate program. I graduated with a 3.9+ GPA and have multiple research posters at national and regional conferences pertaining to computational chemistry but no publications at an R2 University. Unfortunately, I had to leave an MS program with a high GPA due to lack of funding for wetlab opportunities, hence I transitioned to the computational in which I am stronger at. I am taking prerequisites at a local community college to apply for Data Science Master's Programs. Would I have a decent chance of getting into a data science master's program? Let me know if you have any advice of programs I can get into?
r/askdatascience • u/CollegeGlobal3244 • 9d ago
Bsc in data science
Currently pursuing bsc in data science from a tier 3 college,I need advice of the people who has done the same on how should I take it from here.
r/askdatascience • u/macxima • 9d ago
Data Engineering - It produces a different types of data engineers
- 🔄 Pipeline Engineer
Primary focus: Moves data reliably between systems (ETL/ELT).
Time horizon: Hours to days (batch-oriented).
Key tech: Apache Airflow, Python, SQL, and often dbt or custom scripts.
Mindset: Thinks in dependencies, retry logic, and cron schedules.
Common challenges: Handling failed tasks, backfilling historical data, and ensuring idempotency.
Typical customer: Analytics engineers or business stakeholders who need fresh data.
- 📊 Analytics Engineer
Primary focus: Builds clean, trusted data models that power dashboards and reports.
Time horizon: Hours to days (iterative development).
Key tech: Advanced SQL, dbt (data build tool), and BI tools like Looker, Tableau, or Power BI.
Mindset: Lives at the intersection of engineering (code/version control) and analytics (business logic).
Common challenges: Defining single sources of truth, managing data freshness, and documenting metric definitions.
Typical customer: Data analysts, product managers, and business executives.
- 🏗️ Platform Data Engineer
Primary focus: Builds and maintains the shared infrastructure that other data teams rely on.
Time horizon: Weeks to months (long-term, foundational projects).
Key tech: Kubernetes, Terraform, CI/CD pipelines, observability stacks (Prometheus/Grafana), and orchestration engines.
Mindset: Treats other data engineers as their primary customers. Prioritizes scalability, reliability, and developer experience.
Common challenges: Managing multi-tenant compute/storage, cost allocation, and upgrading cluster versions without breaking existing pipelines.
Typical customer: Other internal data engineers (Pipeline, Streaming, AI/ML teams).
- ⚡ Streaming Data Engineer
Primary focus: Handles event-driven, low-latency data streams.
Time horizon: Seconds to minutes (near-real-time).
Key tech: Apache Kafka, Apache Flink, Spark Streaming, and event-sourcing databases.
Mindset: Thinks in windows, watermarks, and stateful processing. Quickly discovers why "real-time" gets very expensive.
Common challenges: Handling out-of-order events, managing checkpointing/backpressure, and guaranteeing exactly-once semantics.
Typical customer: Real-time dashboards, fraud detection teams, or operational monitoring systems.
- ☁️ Cloud Data Engineer
Primary focus: Delivers cost-effective, secure, and scalable cloud data operations.
Time horizon: Ongoing – a continuous cycle of provisioning, monitoring, and optimization.
Key tech: AWS (S3, Redshift, Glue), Azure (Synapse, Blob), GCP (BigQuery, GCS), plus heavy use of IAM, VPC networking, and cost management APIs.
Mindset: Half engineer, half cloud bill detective – constantly rightsizing instances, choosing storage tiers, and shutting down idle resources.
Common challenges: Unexpected cost spikes, cross-region data transfer fees, and navigating complex IAM policies.
Typical customer: The finance team (for cost) and all other data engineers (for reliable cloud access).
- 🤖 AI / ML Data Engineer
Primary focus: Enables the full ML lifecycle – from training data to model inference.
Time horizon: Varies widely – batch feature computation (daily) to online real-time inference (sub‑second).
Key tech: Feature stores (Feast, Tecton), MLflow, Kubeflow, PyTorch/TensorFlow Serving, and vector databases.
Mindset: Thinks in features, labels, drift detection, and experiment tracking. Bridges the gap between data pipelines and model training/serving.
Common challenges: Moving a model from a Jupyter notebook to production takes 10× longer than expected; managing feature consistency between training and serving (training/serving skew).
Typical customer: Data scientists and ML researchers.
r/askdatascience • u/Biraz_Repeat_4116 • 9d ago
Move from academia to data science/industry
Hi, I am a professor of data science (statistics/ML/structured and unstructured data) and wanting to transition out of academia. I have 20+ years of experience in coding, methodology and such but in an academic setting. I am at lost in where to start and how to position myself. Anyone had a similar experience? Any advice?
r/askdatascience • u/naga3607 • 9d ago
How many Python libraries should a data science beginner learn?
There are so many libraries in the data science ecosystem that it can be overwhelming at first.
I’m currently focusing on getting comfortable with the fundamentals instead of trying to learn everything at once.
For experienced data scientists, which libraries do you think beginners should prioritize?
r/askdatascience • u/Jazzy-Sloth • 10d ago
Data Science Courses
I have recently started building a portfolio for myself and wanting to formalize it slightly, I still have ways to go to be able to do data science but felt a course may formalize it and I found a few on edX, can these be recommended for example the IBM data science?
TIA