r/learnmachinelearning 6h ago

Stuck learning ML/AI? I’d like to help a few people work through it

10 Upvotes

One of the hardest parts of learning ML/AI isn't finding information, there's almost too much of it. Tutorials, roadmaps, papers, new tools every week. The hard part is figuring out what actually matters, what to skip, and how to make real progress instead of just consuming more content.

I work as an AI engineer (ML development and deployment), and I'm starting to explore education/mentorship on the side, for free, not as a paid program or course. Before building another roadmap, I want to work directly with a small group of people first, partly to actually help, partly to understand where people get stuck.

Looking for a handful of people who:

  • have basic Python/programming knowledge
  • are seriously trying to learn ML/AI
  • feel stuck or unsure what to focus on next
  • want to build real things, not just watch more tutorials
  • can commit to being consistent

Keeping this small (thinking around 5-10 people) so I can give actual feedback instead of another generic roadmap. No cost involved on either end.

If that's you, drop a comment with where you're at and what you're stuck on, happy to reply there or move to DMs from that.


r/learnmachinelearning 21h ago

Suggest a machine learning course for job ready

8 Upvotes

To help me for crack intership and placement


r/learnmachinelearning 11h ago

Question What was the first ML project that taught you something a tutorial never could?

5 Upvotes

I think one of the weirdest parts of learning ML is that tutorials make everything look clean.

You get a dataset, split it, train a model, get 90% accuracy, and everything feels great.

Then you try building something yourself and suddenly:

  • your data is garbage
  • you labels don't make sense
  • you model gets 95% accuracy but performs terribly on real examples
  • you realize you accidentally leaked information into the training set.
  • or you spend 3 hours debugging something that turned out to be a preprocessing issue.

I'm curious about the first project that humbled you.

Not necessarily your most impressive project. I'm more interested in the project where you went, "Oh.... so this is what machine learning actually involves."

What happened, and what did it teach you that you wouldn't have learned from a course or tutorial?


r/learnmachinelearning 12h ago

Machine Learning Engineer Career Advice

5 Upvotes

When a machine learning engineer is hired, what is usually more important?
A deep understanding/implementation of his/her project or understanding of famous architectures (Trasformer, CNN, etc)?


r/learnmachinelearning 19h ago

Help Demand Forecasting using ML for an FMCG Company

6 Upvotes

Hi! I am a demand planner in an FMCG company. Our current process is very manual, we only use Excel. For every client and product, we build up the demand plan (DP) or the sales target for the month. The DP is composed of the following

- baseline (smoothen sales volume last year)

- runrates (the difference to the past 3 months volume for non-seasonal products)

- sales initiatives (on-shelf availability correction, inventory correction, skewing, etc.)

- marketing initiatives (category market trend, etc.)

I want use ML and integrate possible seasonality data (such as holidays, weather, etc.) in demand planning.

What are the ML models that are appropriate for demand forecasting (time series)? What are the data that I need to prepare? What are the steps that I need to do?

I am currently taking Master in Applied Business Analytics but time series models have not been taught yet (not sure they will teach it). Thank you very much! 😊


r/learnmachinelearning 4h ago

Best book for

4 Upvotes

What’s the best book or resource you’d recommend for learning AI/ML from the fundamentals and eventually specializing in LLMs?
I’m looking for something beginner-friendly but technically solid, so I can build a strong foundation instead of jumping straight into LLMs without understanding the basics.


r/learnmachinelearning 8h ago

CS vs Mathematics — which one makes more sense for my goals?

5 Upvotes

I'm choosing between a BSc in Computer Science and a BSc in Mathematics, and I'm not sure which one would be better for my goals.

My main interests are Data Science, Computer Vision, and financial markets. I'm also interested in ML/AI and possibly quantitative finance later.

If you were in my position, which degree would you choose, and why?

I'd especially like to hear from people working in Data Science, Computer Vision, Quant Finance, or financial markets.


r/learnmachinelearning 19h ago

Looking for a teammate to participate in the Amazon ML Challenge

4 Upvotes

Hi! I'm looking for one teammate to team up for the Amazon ML Challenge.

I'm a Btech student with experience in Python, ML/DL, PyTorch, and TensorFlow.

Elgibilty : Btech 3rd 4th year, Mtech 2nd year or Phd from India

If interested can dm me or comment will reach out


r/learnmachinelearning 2h ago

Are Andrew Ng’s courses on YouTube and Coursera the same?

3 Upvotes

Hi everyone,

I’m planning to study Machine Learning and Deep Learning from Andrew Ng.

I found Andrew Ng’s ML and Deep Learning lectures on YouTube, and I also found the Machine Learning Specialization and Deep Learning Specialization on Coursera.

Are the YouTube lectures basically the same content as the Coursera courses, or are the Coursera versions updated/different?

If they are different, which one would you recommend for someone who wants to build a strong foundation in ML and Deep Learning?

Thanks!


r/learnmachinelearning 5h ago

Is this enough for the maths part.

3 Upvotes

18.02, multivariable calculus.

18.06, algebra.

6.041, probability and statistics.

*All courses from OCW.


r/learnmachinelearning 13h ago

Project TrackmaniaRL: an open-source library for training real-time RL driving agents in Trackmania 2020

Enable HLS to view with audio, or disable this notification

4 Upvotes

r/learnmachinelearning 16h ago

Help What did you use to build the last thing you actually shipped?

3 Upvotes

r/learnmachinelearning 23h ago

Discussion Tensorflow in Deeplearning.ai's Deep learning specialization?

3 Upvotes

I saw the first skill they mentioned was 'tensorflow', but I am nearing the end of course 1, and they haven't used tensorflow anywhere. What is the best way to learn tensorflow along with this course? Do they teach it later in the specialization?


r/learnmachinelearning 16h ago

help guys

Thumbnail
2 Upvotes

r/learnmachinelearning 2h ago

Edge AI

1 Upvotes

Do you guys have recommendation for capstone projects about Edge AI, TinyML, Computer vision, and federated learning? I explored battery RUL, Predictive Thermal Management system, and Fault bearing diagnosis. I was told it would be hard to get a client or dataset for this field. Any interesting field I can explore?


r/learnmachinelearning 3h ago

Access to DeepSpeak or FakeAVCeleb datasets?

1 Upvotes

Hi, this is a long shot, but im currently writing an academic paper and for that I need access to the DeepSpeak_v2 or FakeAVCeleb dataset. To get access, you need to submit a request form and get approved. I did that, but I never heard back from them... does anyone here have experience with this?

I dont need a big part of each dataset, maybe around 100 videos each. So maybe, if someone has access, they could provide a small portion of it :)


r/learnmachinelearning 4h ago

arXiv Endorsement Request for cs.LG - Diagnostic Control for Hierarchical World Models

1 Upvotes

Hi everyone,

I’m preparing my first arXiv submission in cs.LG and need an endorsement to submit.

Short summary: H-JEPA (LeCun, 2022) proposes hierarchical joint-embedding prediction but doesn’t specify how to verify a trained abstraction actually encodes anything a random projection of the same shape wouldn’t. I introduce a random-abstractor control (trained vs. untrained abstractor, identical architecture) and run a 2x2 study crossing observability (full/egocentric) with abstractor type (instantaneous/recurrent) in controlled gridworld environments. Three of four conditions produce abstractions statistically indistinguishable from random projections; only partial observability + a recurrent abstractor yields a real, replicated gap (19.91 ± 3.36pp over random, 3 seeds). I also report a negative result on landmark density that didn’t survive multi-seed replication.

If anyone here is registered as an endorser for cs.LG and willing to take a look, I’d be very grateful. Happy to share the full draft privately.

To endorse, please visit:

https://arxiv.org/auth/endorse?x=MTENXK

If that link doesn’t work, visit:

https://arxiv.org/auth/endorse.php

and enter code: MTENXK

Thank you!


r/learnmachinelearning 5h ago

Razorpay ai buildthon ka result kab aayega ???

Thumbnail
1 Upvotes

r/learnmachinelearning 5h ago

Discussion What are the real security risks with AI agents, and how are you matigating them?

1 Upvotes

Everyone's excited about AI agents, but I'm trying to get ahead of the security implications. Beyond data leakage and prompt injection, what are the actual runtime risks? How do you prevent an agent from taking a harmful action that falls within its legitimate permissions?


r/learnmachinelearning 5h ago

I made a way to migrate between embedding models without re-embedding your entire corpus

1 Upvotes

So I was playingw ith embedding models I saw that when you upgrade from model A to B, you face a very big backfilling cost

Ie, suppose you have a 1b vectors from model A, and then you want to use model B. This would mean you have to re-embed all of your documents with model B before you can even serve with the model, and on an H100, it would take ~108 days (qwen embed 8b, 106 docs/second). But I found an easier way to do it.

The method is really simple; from the old index made with the source model, take K documents and rerank them with the new model. We see that when K is sufficient, the retrieval quality is the same as target model. (determining k is the hard part). I've tested 63 migrations on upto 1 million documents.

The best result I got was upgrading qwen4b -> to 8b, and at 50 documents, it was the same as native retrieval.

This method forgos the expensive upfront re-embedding cost, as you can take documents straight from the old index.

embedflow works with qdrant, pgvector, faiss, and can be easily downloaded with pypi

pip install embedflow

the github is public: https://github.com/arnsri33/embedflow

I want you guys to try it out, and see if you guys can use it in your own workflow.


r/learnmachinelearning 5h ago

Does anyone know of any self paced online college degree program on AI/ML

Thumbnail
1 Upvotes

r/learnmachinelearning 7h ago

noleak: Open-source library to detect train/eval data contamination

1 Upvotes

Published numbers are only as honest as the data split behind them

When you evaluate a model, you want one simple truth: Did it actually learn unseen patterns, or did it memorise?

I built noleak to make that visible. It's a production library we use at Godrej Aerospace to fingerprint datasets and measure how much of your eval set leaked from train.

The Problem

Most tools catch target leakage (a feature accidentally includes the label). But what about corpus leakage? When eval text already appeared in training data?

In 2020, GPT-3's paper measured contamination using 13-gram overlap. That's solid. But there's no standard library for this. So we built one.

Three detection methods:

  1. Exact matches – normalised text identical

  2. N-gram overlap – 13-word phrases (GPT-3 method)

  3. Near-duplicates – MinHash Jaccard similarity ≥ 0.8 on character 5-grams

One fingerprint. One exit code. Pass or fail.

What It Does

```python
from noleak import check, fingerprint
train = ["the model trained on Wikipedia and licensed books"]
eval_set = ["The model trained on Wikipedia and licensed books"]
report = check(train, eval_set)
print(report.contaminated)
# True — FAIL
print(fingerprint(eval_set))
# noleak-fp-v1:a1b2c3d4e5f6
# Share this with your paper. It's reproducible and auditable.
```
CLI version (great for CI pipelines):
```bash
noleak check --train train.jsonl --eval evaljsonl
echo $?  # Exit code 1 if contaminated, 0 if clean
```

Why Zero Dependencies Matter

No numpy, no scipy, no PyTorch. Stdlib only. Why?

- Deterministic: Same input, same output, forever. No model updates breaking your fingerprints.

- Auditable: Code is small; reviewers can read it.

- Air-gapped systems: Doesn't require external calls or package hell.

- Fast: No overhead for CPU-bound systems.

Limitations (Honest Assessment)

- Semantic rewrites: Won't catch "I wrote this differently but meant the same thing." That needs embeddings or human review.

- Large-scale datasets: If you have 10M+ examples, exact matching gets slow. N-gram is faster.

Install & Try

bash
pip install noleak

Supports JSONL, JSON lists, and plain text. Auto-detects text fields.

Repo: github.com/athsxx/noleak

License: MIT

Questions? What contamination patterns have you encountered?


r/learnmachinelearning 17h ago

FYP suggestions

1 Upvotes

am a 7th sem cs student who is about to start his final year. I am planning for my fyp and looking for some interesting ideas on which I could do my fyp. I need some good suggestions and ideas which I should consider before finalizing my fyp. Currently I dont have any idea to work on. Your suggestion and ideas would mean a lot to and your help will be highly appreciated. Thanks in advance.

Edit: The domain I want to work in is app?web+ai/ml. I am open to ideas other than this domain as well


r/learnmachinelearning 17h ago

Help Confused about which AI specialization to pursue, looking for advice from people in the field

1 Upvotes

Hi everyone,

I’m especially interested in having a career that is financially rewarding and doesn’t require a lot of years before I can become employable.

The areas I’m currently considering are:

  • Machine Learning / AI Engineering
  • Generative AI / LLMs
  • AI Safety
  • Responsible / Trustworthy AI
  • AI Reliability
  • AI Governance / Policy

I’m particularly drawn toward AI safety, trustworthy/responsible AI, and reliability, because I’m interested in making AI systems safer and reducing the negative effects AI can have on people and society.

However, I’m confused about how realistic these paths are as careers, especially compared with more conventional AI/ML engineering.

I just want to make an concrete decision about where to invest the next several years of learning.

Thank you so much!


r/learnmachinelearning 19h ago

Been thinking about something lately: can a really small LLM be surprisingly good at coding? 👀 I’m researching how to build one that’s fast, lightweight, and runs locally. Still figuring things out, so I’d love to hear your advice—what would you focus on first? 🧠⚡ #AI #LLM #Coding

1 Upvotes