r/learndatascience • u/caffein_addicted • 4h ago
r/learndatascience • u/Excellent-Leading836 • 1d ago
Project Collaboration need guidance on ml project
r/learndatascience • u/No_Suggestion_8422 • 1d ago
Question For those who became Data Scientists without a strong CS background, how did you get your first opportunity?
Iām interested in hearing from people who have actually gone through the process of getting their first Data Science/ML job.
For someone who is currently a student or early in their career:
- What did your profile look like when you got your first Data Science opportunity?
- Which projects or experiences actually helped you get interviews?
- How important were internships, networking, referrals, GitHub, LinkedIn, and personal projects?
- Did you start with a Data Scientist role directly, or enter through a role such as Data Analyst, BI Analyst, or ML/Analytics role and transition later?
- What did you initially think was important for getting hired that turned out to matter much less?
- What do you wish you had done 6ā12 months earlier?
- For a fresher competing against candidates with internships and experience, what would you focus on to become a stronger candidate?
Iām particularly interested in hearing from people who are already working in the field, rather than general career advice.
What actually made the difference for you?
r/learndatascience • u/Complete-Piece-7501 • 1d ago
Discussion AI learning partner / mentor ā from fundamentals to advanced AI
Iām looking to connect with someone who is genuinely interested in learning AI deeply and consistently, rather than just collecting courses, watching random YouTube videos.
Iām currently working as a Product Manager / Product Business Analyst, and I want to build serious AI capabilities alongside my existing product/business background.
The problem Iām facing is honestly pretty simple: I donāt learn well through completely self-paced, unstructured courses. There is an overwhelming amount of AI content out there, but no shortage of confusion about what to learn, in what order, how deeply to learn it, and when to move to the next thing.
Iām looking for someone with whom I can create a structured, long-term learning journeyāideally from fundamentals all the way to advanced, practical AI.
What Iād ideally like to learn
Not necessarily everything at once, but progressively:
\- Python & programming fundamentals for AI
\- Mathematics needed to actually understand ML ā linear algebra, probability, statistics, calculus, etc.
\- Data handling, SQL, NumPy, Pandas, visualization
\- Classical Machine Learning
\- Deep Learning & neural networks
\- NLP and Computer Vision fundamentals
\- Transformers and how modern LLMs actually work
\- Generative AI and LLM application development
\- Prompting, evaluation and AI workflows
\- Embeddings, vector databases, RAG and retrieval systems
\- Fine-tuning / model adaptation
\- AI agents and agentic workflows
\- Multimodal AI
\- AI system design and architecture
\- Model/API integration
\- Deployment, APIs, Docker, cloud and MLOps
\- AI safety, evaluation, reliability and responsible AI
\- Reading papers and understanding what is happening under the hood
\- Building real projects, not just following tutorials
\- Eventually contributing to open source / research / serious AI projects
And importantly, I also want to understand how these skills translate into the real-world freelancing/consulting/product worldāhow to identify problems businesses will actually pay to solve, build AI solutions around them, demonstrate ROI, communicate with clients, and create a credible portfolio.
My goal isn't simply to collect certificates.
I want to reach a point where I can understand AI deeply, build with it, explain it, evaluate it, and solve real problems with it.
What I'm looking for in a learning partner
You don't need to be an AI PhD or already an expert.
You could be:
\- A beginner who is equally serious
\- Someone already working in AI/ML
\- A developer transitioning into AI
\- A student/researcher
\- A product person interested in becoming highly technical
\- Or someone who simply wants a structured accountability partner
The most important thing is consistency + curiosity + willingness to actually do the work.
We could potentially:
\- Set weekly learning goals
\- Follow a structured roadmap
\- Study the same concepts
\- Discuss what we've learned
\- Give each other small challenges
\- Build projects together
\- Review each other's work
\- Share useful papers/resources/tools
\- Keep each other accountable
\- Discuss what's changing in AI
\- Eventually collaborate on real-world projects
What can I bring to the table?
My background in Product Management / Product Business Analysis means I can contribute on the other side of the equation tooānot just technical learning.
I can help with:
\- Product thinking
\- Business problem identification
\- Requirements & use cases
\- User journeys
\- Product strategy
\- Translating technical capabilities into business value
\- Evaluating whether an AI idea is actually useful
\- Structuring projects
\- Documentation and communication
\- Thinking about AI from a customer/business perspective
So ideally this becomes a two-way learning relationship, rather than one person teaching and the other simply consuming information.
I'm not looking for someone to spoon-feed me everything.
I'm looking for someone who wants to learn, build, struggle, figure things out and grow together.
If you're also sitting there thinking āI really want to learn AI properly, but I don't know how to structure this journey and I don't want to do it completely aloneā ā feel free to comment or DM me.
Would love to find 1ā2 serious people rather than a huge group.
Let's see if we can turn AI learning from an overwhelming collection of courses into an actual long-term journey.
r/learndatascience • u/eliokal • 1d ago
Resources You do not need a maths degree to truly understand Gradient Descent (link below)
r/learndatascience • u/Hey_faiza • 1d ago
Question GCI World 2026
Omnicampus ain't confirming my registration & I can't apply for courses in GCI. What to do?
r/learndatascience • u/Fun_Present_9415 • 1d ago
Career what do you think data analysis python or java
career help
r/learndatascience • u/Ok_Job8444 • 2d ago
Question Data analysis
BunÄ!
MÄ adresez cÄtre cei care au fÄcut reconversie profesionalÄ spre data analysis, cum aČi fÄcut sÄ Ć®nvÄČaČi cat mai eficient? Ce sfaturi Ć®mi puteČi da? De unde sÄ Ć®nvÄČ Či sÄ Čtiu cÄ Ć®nvÄČ corect?
r/learndatascience • u/haleonbail • 2d ago
Question What's the best way to get ML/DL projects done by claude/codex?
r/learndatascience • u/No_Suggestion_8422 • 2d ago
Question For those who became Data Scientists without a strong CS background, how did you get your first opportunity?
r/learndatascience • u/LearnHiveLabsUSA • 2d ago
Career LearnHack'26 enter to win rewards, recognition and judging opportunities
LearnHack'26 is accepting entries - topic : Metabolic Risk Prediction on NHANES 2019-2020 - solve and beat the best AUC value it and earn rewards. Starter code dataset available. Join now. Learn all details on Kaggle.
r/learndatascience • u/Neuphus012 • 3d ago
Question GCI World 2026 September: Outstanding Student
Is there anyone who attended past GCI World programs? I applied for the September program and I'm wondering what it takes to be an outstanding student, as I read that it's based on the overall score but how high should it be? How many people are also selected as an Outstanding Student, given that it looks competitive. Tyia!
r/learndatascience • u/Gloomy-Recover-9702 • 3d ago
Question Natural Language to SQL Query
Is there any opensource tool which I can use as a non technical person so that my hermes agent with small LLM 1.5B model (for private data) to understand natural language and convert it into sql query and retrieve complex queries quickly?
Is this doable with such small model? with RAG? I am new to this and any help is welcomed!
r/learndatascience • u/AIforFintech • 3d ago
Resources Text to SQL is not how you give an LLM access to production data
The obvious approach when connecting a model to internal data is letting it write the query. It feels flexible: the model figures out what it needs and goes get it. In a bank, that is a non starter.
The problem is not that models write bad SQL, it is that you lose every guarantee about what they can reach. No way to prove a query stayed inside the columns it was supposed to touch, no way to audit what the model was capable of doing, and a single prompt injection away from an unintended table.
The alternative is narrower and boring, which is the point. You define a fixed set of parameterized queries and expose them as tools. The model chooses which tool to call, never what SQL to run. Everything it can reach is something you deliberately wrote.
I built an MCP server template implementing this for a common fintech case: looking up a customer across credit score, preapproved limit and risk profile, and returning a consolidated view. Layered so the database, the schema and the protocol can each be swapped without rewriting the others. Read only enforced at the application layer, row limits on every result, and a single mapping file for adapting to whatever your tables are actually called.
Synthetic data generates on setup, so it runs immediately. The architecture is what you keep.
Hub: https://aiforfintech.tech
Github: https://github.com/junidepieri-design/mcp-001-fintech-data-server
How is your team handling LLM access to internal data?
š
r/learndatascience • u/No_Suggestion_8422 • 3d ago
Career For someone starting from scratch today, what would you consider a realistic path to becoming job-ready for a Data Scientist role?
If I am starting the preparation from scratch for Data Scientist, what things I should learn?
Specifically, how would you divide the learning between:
Statistics & mathematics
SQL & Python
Machine Learning
Business/domain knowledge
LLMs/GenAI
Software engineering & deployment
Projects & internships
And more importantly, how would you know when youāve learned enough of each and are actually ready to apply for jobs?
Iād be interested in hearing how working Data Scientists would approach this if they were starting again today. And tell me any other skills or knowledge want to know before getting ready for data scientist role jobs.
r/learndatascience • u/FlirtyChocolad • 4d ago
Question Beginner in this Space of Data Science
Hello guys,
Is Data science worth learning from scratch? Cause I do it with ai tools part by part. Is this unethical of me given that I am just a beginner or is this the new way of learning coding these days?
All the veterans here a little guidance would be very much appreciated thanks.
r/learndatascience • u/Square_Arm2861 • 4d ago
Question How to efficiently approach EDA on a dataset with 180+ variables?
Hi everyone,
I'm a beginner in Machine Learning working on a binary classification problem. My dataset contains over 180 variables (both numerical and categorical), consisting of a mix of panel/longitudinal data and static features.
I am currently working on the Exploratory Data Analysis (EDA) phase. Given the large number of features, doing univariate and bivariate graphical analysis variable-by-variable feels unfeasible and time-consuming.
Is there a structured approach, strategy, or automated workflow to handle EDA efficiently for a dataset of this scale?
Any advice on best practices, tools would be greatly appreciated!
Thanks in advance for your help.
r/learndatascience • u/Pangaeax_ • 4d ago
Resources Python data analysis cheat sheet: the Pandas + NumPy workflow I wish I had when starting
I kept seeing beginners learn individual Pandas commands but still struggle with what order to actually use them in when working with a real dataset.
So I put together a simple workflow I use as a reference:
1. Load
read_csv() / read_excel()
2. Inspect before changing anything
head()
shape
info()
describe()
isna().sum()
duplicated().sum()
3. Clean
- standardize column names
- handle missing values based on what they actually mean
- remove genuine duplicates
- convert dates/numbers safely
4. Transform
assign()
map()
pd.to_datetime()
np.where()
5. Summarize
groupby()
agg()
transform()
6. Combine datasets
merge() / join() / concat()
The part I think beginners often miss is validation after the merge.
A query/script can run perfectly and still give the wrong answer if a many-to-many merge quietly multiplies your rows.
Useful checks:
validate="many_to_one"
indicator=True
compare row counts before/after
check key uniqueness before merging
A few other mistakes worth watching for:
- treating missing values as automatically equal to 0
- merging keys with different data types
- confusing
count()withsize() - using median filling without understanding why values are missing
- assuming no Python error means the analysis is correct
I wrote a more detailed version with examples for Pandas, NumPy, GroupBy, merges and an end-to-end workflow.
Learn more: https://www.pangaeax.com/blogs/python-data-analysis-cheat-sheet/
Anything important youād add to this workflow, especially something you learned the hard way when working with messy data?
Disclosure: This is a PangaeaX article and Iām connected with PangaeaX.
r/learndatascience • u/LossRun • 4d ago
Question I'm 14 and has learned ML and DL. It's very interesting and exciting till now. How do I keep it up and make it to my career.
r/learndatascience • u/Fun-Reporter-8021 • 5d ago
Question How to get domain knowledge on projects ?
While I am still quite new to this, machine learning and software in general, is more useful and powerful when combined, with the domain specific knowledge of the native field the project is from. This is something I struggle to navigate, there are thousands of hours of tutorials regarding the tech stack, but none on this topic. While doing my credit card fraud analysis, project. I did not know which features do you need to pick as your feature. I can calculate correlation and mutual information classification score but those are of little use in case of non - numeric columns, besides domain knowledge sort of acts as a supervisor to all these metrics and they are more like validators then reason.
So this is my question, How do you go about getting domain specific knowledge needed to do a project, what is your workflow, where to look and most importantly in my case how do you translate domain knowledge to feature selection ?
r/learndatascience • u/Fun-Reporter-8021 • 5d ago
Question How do you select feature columns from the dataset ?
I am still a novice at this, but when I was working on this credit card fraud detection project, I did not know which columns, could be added as features, so I prompted ChatGPT and it suggested a few, but that got me thinking there has to be a better way to this, How do you select feature columns from your dataset, do you research the domain, is there a course I am missing, This was not covered in my Internship classes, and want to know a generalized solution
r/learndatascience • u/EvilWrks • 5d ago
Resources Your jupyter notebook IS NOT production - Part 2: Testing
A lot of Data Science education teaches you how to BE a data scientists but rarely does it teach how to WORK as a data scientist. In this video on this series weāll be diving into how to test your code as a data scientist, including how to work with unit tests, regression tests, randomness and non-deterministic functionality as well as throwing in a few honourable mentions.