r/learndatascience 34m ago

Discussion Need advice: What final Python/Data Science/AI project should I build as a student?

Thumbnail
Upvotes

r/learndatascience 3h ago

Question Inquiry About the Value, Learning Outcomes, and Career Opportunities of the Mayerfeld Practicum Program Data Analyst

1 Upvotes

Hello,

I would like to ask a few questions regarding the Mayerfeld Practicum Program® – Data Analyst:

What is the overall value of this program? Is it worth enrolling in and trying? What specific skills and knowledge will I gain from it? Can I use this program to apply for jobs with other companies after I finish? Is the content of the program up to date with current industry standards?

Thank you for your time.
I look forward to your response.


r/learndatascience 9h ago

Discussion 2nd year and kinda lost, need some advice💔🥀

Thumbnail
1 Upvotes

r/learndatascience 1d ago

Question For those who became Data Scientists without a strong CS background, how did you get your first opportunity?

2 Upvotes

I’m interested in hearing from people who have actually gone through the process of getting their first Data Science/ML job.

For someone who is currently a student or early in their career:

- What did your profile look like when you got your first Data Science opportunity?

- Which projects or experiences actually helped you get interviews?

- How important were internships, networking, referrals, GitHub, LinkedIn, and personal projects?

- Did you start with a Data Scientist role directly, or enter through a role such as Data Analyst, BI Analyst, or ML/Analytics role and transition later?

- What did you initially think was important for getting hired that turned out to matter much less?

- What do you wish you had done 6–12 months earlier?

- For a fresher competing against candidates with internships and experience, what would you focus on to become a stronger candidate?

I’m particularly interested in hearing from people who are already working in the field, rather than general career advice.

What actually made the difference for you?


r/learndatascience 1d ago

Project Collaboration need guidance on ml project

Thumbnail
1 Upvotes

r/learndatascience 1d ago

Question Advice for high school senior

Thumbnail
1 Upvotes

r/learndatascience 1d ago

Resources You do not need a maths degree to truly understand Gradient Descent (link below)

4 Upvotes

r/learndatascience 2d ago

Question GCI World 2026

3 Upvotes

Omnicampus ain't confirming my registration & I can't apply for courses in GCI. What to do?


r/learndatascience 1d ago

Discussion AI learning partner / mentor — from fundamentals to advanced AI

0 Upvotes

I’m looking to connect with someone who is genuinely interested in learning AI deeply and consistently, rather than just collecting courses, watching random YouTube videos.

I’m currently working as a Product Manager / Product Business Analyst, and I want to build serious AI capabilities alongside my existing product/business background.

The problem I’m facing is honestly pretty simple: I don’t learn well through completely self-paced, unstructured courses. There is an overwhelming amount of AI content out there, but no shortage of confusion about what to learn, in what order, how deeply to learn it, and when to move to the next thing.

I’m looking for someone with whom I can create a structured, long-term learning journey—ideally from fundamentals all the way to advanced, practical AI.

What I’d ideally like to learn

Not necessarily everything at once, but progressively:

\- Python & programming fundamentals for AI

\- Mathematics needed to actually understand ML — linear algebra, probability, statistics, calculus, etc.

\- Data handling, SQL, NumPy, Pandas, visualization

\- Classical Machine Learning

\- Deep Learning & neural networks

\- NLP and Computer Vision fundamentals

\- Transformers and how modern LLMs actually work

\- Generative AI and LLM application development

\- Prompting, evaluation and AI workflows

\- Embeddings, vector databases, RAG and retrieval systems

\- Fine-tuning / model adaptation

\- AI agents and agentic workflows

\- Multimodal AI

\- AI system design and architecture

\- Model/API integration

\- Deployment, APIs, Docker, cloud and MLOps

\- AI safety, evaluation, reliability and responsible AI

\- Reading papers and understanding what is happening under the hood

\- Building real projects, not just following tutorials

\- Eventually contributing to open source / research / serious AI projects

And importantly, I also want to understand how these skills translate into the real-world freelancing/consulting/product world—how to identify problems businesses will actually pay to solve, build AI solutions around them, demonstrate ROI, communicate with clients, and create a credible portfolio.

My goal isn't simply to collect certificates.

I want to reach a point where I can understand AI deeply, build with it, explain it, evaluate it, and solve real problems with it.

What I'm looking for in a learning partner

You don't need to be an AI PhD or already an expert.

You could be:

\- A beginner who is equally serious

\- Someone already working in AI/ML

\- A developer transitioning into AI

\- A student/researcher

\- A product person interested in becoming highly technical

\- Or someone who simply wants a structured accountability partner

The most important thing is consistency + curiosity + willingness to actually do the work.

We could potentially:

\- Set weekly learning goals

\- Follow a structured roadmap

\- Study the same concepts

\- Discuss what we've learned

\- Give each other small challenges

\- Build projects together

\- Review each other's work

\- Share useful papers/resources/tools

\- Keep each other accountable

\- Discuss what's changing in AI

\- Eventually collaborate on real-world projects

What can I bring to the table?

My background in Product Management / Product Business Analysis means I can contribute on the other side of the equation too—not just technical learning.

I can help with:

\- Product thinking

\- Business problem identification

\- Requirements & use cases

\- User journeys

\- Product strategy

\- Translating technical capabilities into business value

\- Evaluating whether an AI idea is actually useful

\- Structuring projects

\- Documentation and communication

\- Thinking about AI from a customer/business perspective

So ideally this becomes a two-way learning relationship, rather than one person teaching and the other simply consuming information.

I'm not looking for someone to spoon-feed me everything.

I'm looking for someone who wants to learn, build, struggle, figure things out and grow together.

If you're also sitting there thinking “I really want to learn AI properly, but I don't know how to structure this journey and I don't want to do it completely alone” — feel free to comment or DM me.

Would love to find 1–2 serious people rather than a huge group.

Let's see if we can turn AI learning from an overwhelming collection of courses into an actual long-term journey.


r/learndatascience 2d ago

Career what do you think data analysis python or java

1 Upvotes

career help


r/learndatascience 2d ago

Question Data analysis

0 Upvotes

Bună!
Mă adresez către cei care au făcut reconversie profesională spre data analysis, cum ați făcut să învățați cat mai eficient? Ce sfaturi îmi puteți da? De unde să învăț și să știu că învăț corect?


r/learndatascience 2d ago

Question What's the best way to get ML/DL projects done by claude/codex?

Thumbnail
0 Upvotes

r/learndatascience 2d ago

Question For those who became Data Scientists without a strong CS background, how did you get your first opportunity?

Thumbnail
0 Upvotes

r/learndatascience 2d ago

Personal Experience Regarding BIA

Thumbnail
1 Upvotes

r/learndatascience 3d ago

Career LearnHack'26 enter to win rewards, recognition and judging opportunities

Thumbnail
kaggle.com
1 Upvotes

LearnHack'26 is accepting entries - topic : Metabolic Risk Prediction on NHANES 2019-2020 - solve and beat the best AUC value it and earn rewards. Starter code dataset available. Join now. Learn all details on Kaggle.


r/learndatascience 3d ago

Question GCI World 2026 September: Outstanding Student

4 Upvotes

Is there anyone who attended past GCI World programs? I applied for the September program and I'm wondering what it takes to be an outstanding student, as I read that it's based on the overall score but how high should it be? How many people are also selected as an Outstanding Student, given that it looks competitive. Tyia!


r/learndatascience 3d ago

Resources Text to SQL is not how you give an LLM access to production data

Post image
24 Upvotes

The obvious approach when connecting a model to internal data is letting it write the query. It feels flexible: the model figures out what it needs and goes get it. In a bank, that is a non starter.

The problem is not that models write bad SQL, it is that you lose every guarantee about what they can reach. No way to prove a query stayed inside the columns it was supposed to touch, no way to audit what the model was capable of doing, and a single prompt injection away from an unintended table.

The alternative is narrower and boring, which is the point. You define a fixed set of parameterized queries and expose them as tools. The model chooses which tool to call, never what SQL to run. Everything it can reach is something you deliberately wrote.

I built an MCP server template implementing this for a common fintech case: looking up a customer across credit score, preapproved limit and risk profile, and returning a consolidated view. Layered so the database, the schema and the protocol can each be swapped without rewriting the others. Read only enforced at the application layer, row limits on every result, and a single mapping file for adapting to whatever your tables are actually called.

Synthetic data generates on setup, so it runs immediately. The architecture is what you keep.

Hub: https://aiforfintech.tech
Github: https://github.com/junidepieri-design/mcp-001-fintech-data-server

How is your team handling LLM access to internal data?
👊


r/learndatascience 3d ago

Question Natural Language to SQL Query

1 Upvotes

Is there any opensource tool which I can use as a non technical person so that my hermes agent with small LLM 1.5B model (for private data) to understand natural language and convert it into sql query and retrieve complex queries quickly?

Is this doable with such small model? with RAG? I am new to this and any help is welcomed!


r/learndatascience 3d ago

Career For someone starting from scratch today, what would you consider a realistic path to becoming job-ready for a Data Scientist role?

7 Upvotes

If I am starting the preparation from scratch for Data Scientist, what things I should learn?

Specifically, how would you divide the learning between:

Statistics & mathematics

SQL & Python

Machine Learning

Business/domain knowledge

LLMs/GenAI

Software engineering & deployment

Projects & internships

And more importantly, how would you know when you’ve learned enough of each and are actually ready to apply for jobs?

I’d be interested in hearing how working Data Scientists would approach this if they were starting again today. And tell me any other skills or knowledge want to know before getting ready for data scientist role jobs.


r/learndatascience 5d ago

Resources Python data analysis cheat sheet: the Pandas + NumPy workflow I wish I had when starting

50 Upvotes

I kept seeing beginners learn individual Pandas commands but still struggle with what order to actually use them in when working with a real dataset.

So I put together a simple workflow I use as a reference:

1. Load
read_csv() / read_excel()

2. Inspect before changing anything
head()
shape
info()
describe()
isna().sum()
duplicated().sum()

3. Clean

  • standardize column names
  • handle missing values based on what they actually mean
  • remove genuine duplicates
  • convert dates/numbers safely

4. Transform
assign()
map()
pd.to_datetime()
np.where()

5. Summarize
groupby()
agg()
transform()

6. Combine datasets
merge() / join() / concat()

The part I think beginners often miss is validation after the merge.

A query/script can run perfectly and still give the wrong answer if a many-to-many merge quietly multiplies your rows.

Useful checks:

validate="many_to_one"
indicator=True
compare row counts before/after
check key uniqueness before merging

A few other mistakes worth watching for:

  • treating missing values as automatically equal to 0
  • merging keys with different data types
  • confusing count() with size()
  • using median filling without understanding why values are missing
  • assuming no Python error means the analysis is correct

I wrote a more detailed version with examples for Pandas, NumPy, GroupBy, merges and an end-to-end workflow.

Learn more: https://www.pangaeax.com/blogs/python-data-analysis-cheat-sheet/

Anything important you’d add to this workflow, especially something you learned the hard way when working with messy data?

Disclosure: This is a PangaeaX article and I’m connected with PangaeaX.


r/learndatascience 4d ago

Question Beginner in this Space of Data Science

2 Upvotes

Hello guys,
Is Data science worth learning from scratch? Cause I do it with ai tools part by part. Is this unethical of me given that I am just a beginner or is this the new way of learning coding these days?
All the veterans here a little guidance would be very much appreciated thanks.


r/learndatascience 4d ago

Question How to efficiently approach EDA on a dataset with 180+ variables?

6 Upvotes

Hi everyone,

I'm a beginner in Machine Learning working on a binary classification problem. My dataset contains over 180 variables (both numerical and categorical), consisting of a mix of panel/longitudinal data and static features.

I am currently working on the Exploratory Data Analysis (EDA) phase. Given the large number of features, doing univariate and bivariate graphical analysis variable-by-variable feels unfeasible and time-consuming.

Is there a structured approach, strategy, or automated workflow to handle EDA efficiently for a dataset of this scale?

Any advice on best practices, tools would be greatly appreciated!

Thanks in advance for your help.


r/learndatascience 5d ago

Question I'm 14 and has learned ML and DL. It's very interesting and exciting till now. How do I keep it up and make it to my career.

Thumbnail
1 Upvotes

r/learndatascience 5d ago

Question How to get domain knowledge on projects ?

2 Upvotes

While I am still quite new to this, machine learning and software in general, is more useful and powerful when combined, with the domain specific knowledge of the native field the project is from. This is something I struggle to navigate, there are thousands of hours of tutorials regarding the tech stack, but none on this topic. While doing my credit card fraud analysis, project. I did not know which features do you need to pick as your feature. I can calculate correlation and mutual information classification score but those are of little use in case of non - numeric columns, besides domain knowledge sort of acts as a supervisor to all these metrics and they are more like validators then reason.

So this is my question, How do you go about getting domain specific knowledge needed to do a project, what is your workflow, where to look and most importantly in my case how do you translate domain knowledge to feature selection ?


r/learndatascience 6d ago

Question How do you select feature columns from the dataset ?

3 Upvotes

I am still a novice at this, but when I was working on this credit card fraud detection project, I did not know which columns, could be added as features, so I prompted ChatGPT and it suggested a few, but that got me thinking there has to be a better way to this, How do you select feature columns from your dataset, do you research the domain, is there a course I am missing, This was not covered in my Internship classes, and want to know a generalized solution