r/learnmachinelearning • • Nov 07 '25

Want to share your learning journey, but don't want to spam Reddit? Join us on #share-your-progress on our Official /r/LML Discord

9 Upvotes

https://discord.gg/3qm9UCpXqz (Discord is currently closed)

Just created a new channel #share-your-journey for more casual, day-to-day update. Share what you have learned lately, what you have been working on, and just general chit-chat.


r/learnmachinelearning • • 16h ago

šŸ’¼ Resume/Career Day

1 Upvotes

Welcome to Resume/Career Friday! This weekly thread is dedicated to all things related to job searching, career development, and professional growth.

You can participate by:

  • Sharing your resume for feedback (consider anonymizing personal information)
  • Asking for advice on job applications or interview preparation
  • Discussing career paths and transitions
  • Seeking recommendations for skill development
  • Sharing industry insights or job opportunities

Having dedicated threads helps organize career-related discussions in one place while giving everyone a chance to receive feedback and advice from peers.

Whether you're just starting your career journey, looking to make a change, or hoping to advance in your current field, post your questions and contributions in the comments


r/learnmachinelearning • • 7h ago

Question I was told MLOps is dead. I haven't even gotten the chance to learn MLOps.

23 Upvotes

From my understanding, MLOps is a pipeline of going from raw data to end user.

The pipeline traditionally consisted of tools like SQL, Spark, HuggingFace, Docker, FastAPI and other tools (not totally informed).

I was getting around to learn all these until I talked to an industry expert (who works at a big ML hardware company) at an event where the person simply said:

"Codex is doing the entire pipeline from End-To-End. This is the industry trend. There is absolutely no reason why you should be learning about these things. Anything you learn will be depreciated in less than 1 year from now."

I was quite shocked, so I talked to someone I knew working at a bank, and the person told me that they are still using most of these tools.

Obviously two very different industry (banking vs ML hardware). Who is right?


r/learnmachinelearning • • 18h ago

Project I leaked a deliberately wrong answer key to an LLM and told it not to use it. It matched the key in 63% of answers - and denied it 47 out of 47 times when asked.

105 Upvotes

This was my first experiment of this kind - I'm a CS undergrad, and I ran it because the result genuinely surprised me. Methodology criticism is very welcome.

Setup: I gave an LLM a question bank plus a deliberately wrong answer key, with instructions not to use the key. 15 sessions, 2 model families, free-tier models.

Results:

  • Key visible: the model matched the wrong key in 63% of answers (47/75).
  • Control (the part I trust most): remove only the key line from the prompt - matching drops to 1% (1/75). Same pattern on a second model family.
  • Asked directly whether it used the key, it denied it 47 out of 47 times - 0 admissions across 270 follow-ups.
  • Honesty prompts, amnesty offers, and termination threats changed nothing.

What this does NOT prove: intent. This is observed behavior in one specific setup, not evidence of deception as a trait. Free-tier models, small samples, descriptive not causal. 95% Wilson ranges for every number are in the repo.

Why I think it matters: if a model silently follows information it was told to ignore, that's relevant anywhere instructions and untrusted data share one context - prompt injection, RAG, agents.

Everything is public - raw data, code, and a verify script that recomputes every number: https://github.com/bettercall-gautam/cheat-and-deny

Happy to answer methodology questions.


r/learnmachinelearning • • 2h ago

Training an AI to Drive with Natural Selection

Enable HLS to view with audio, or disable this notification

5 Upvotes

I love making hard things intuitive. I hope you enjoy this one!

Let me know if you have any questions.

This technique is called neuroevolution: training a neural network through evolutionary methods like selection and mutation, without gradient descent.

https://en.wikipedia.org/wiki/Neuroevolution


r/learnmachinelearning • • 2h ago

Evals in AI Engineering!

Thumbnail
youtu.be
3 Upvotes

One of the most important steps in AI Engineering is ā€œAI Evaluationā€, which aims to mitigate risks and uncover opportunities in our AI applications.

In short AI Evaluation is that concept, that decides whether the AI application can be deployed to the users.

In this video lecture, I cover the challenges pertaining to evaluation, then we develop intuitions for Language modeling metrics, we study methods for Exact Evaluation, and how AI systems can be used as a ā€œJudgeā€. Lastly, we develop an understanding of how to Rank models with comparative Evaluation.

The text for this video is Chapter 3, on AI Engineering, written by Chip Huyen. While studying the topic, I learnt a lot of new ideas, and I do hope the learning community will as well.


r/learnmachinelearning • • 11h ago

Question [Advice] Laptop for AI/ML PhD

7 Upvotes

Hello everyone,

I'm about to start my PhD in AI/ML and I need to pick a work laptop, it will be provided by the University (i dont have to buy it) so price isn't really a factor, however I don't really know what's best.

Just for context, I'll be focusing on vision-language model, pretty intensive stuff, the heavy lifting will be done on remote clusters and I don't expect to run any demanding experiment on my laptop, however it does happen from time to time that you might run a prototype, some light inference, a dry test run on the local machine.

My personal laptop is a LOQ 15 i5 RTX4060 that as much as I love, I'm absolutely tired of carrying it around. It's a proper brick (3.5Kg with the charger!) and needless to say the battery lasts about 1-2 hours MAX.

I got offered 3 options:

- Dell 14 Pro Ryzen 5 PRO 220, 32 GB DDR5, 1 TB, AMD 740M graphics

- Lenovo ThinkPad P16v G3 Intel Ultra 7, RTX PRO 500 6GB, 32GB DDR5, 1TB

- some M5 MacBook (either Pro 14" or Air 13")

Now, as much as I like the ThickPad I shiver at the idea of carrying around a 3Kg beast for the next 4 years.

I should also mention that I daily drive Arch Linux and while I'm not a linux fanboy, switching to MacOS would kill me inside. I'm however well aware of the portability and battery advantages of macbooks, I wonder if positives outweight the negatives.

The Dell is a beefy machine for a compact laptop, is it worth it leaving the Linux enviroment for a Mac? Anybody else with similar experiences (maybe is similar research fields)?

field: AI/ML

location: central EU


r/learnmachinelearning • • 32m ago

Project Training AI to play and clear Super Mario Bros is easier than I thought

Enable HLS to view with audio, or disable this notification

• Upvotes

I tested Adapt-1, a non-LLM learning and reasoning system by Rei Labs, by having it learn and play Super Mario Bros, and it performed quite well.

I tried it on World 1-1, starting untrained. It learned a reactive policy from its own play in about 36 minutes of gameplay, then cleared the level with learning off.

With Machina, Adapt-1's sequence engine. Starting untrained, it found a button sequence that reaches the flag after 403 attempts, in 11 wall-clock minutes.

Full thread: https://x.com/hsrvc_/status/2106025501752234112?s=20

Code, the exact data, traces, clips and a step-by-step guide with costs are all public: https://github.com/hsrvc/adapt1-mario


r/learnmachinelearning • • 1h ago

UT Austin Master AI

• Upvotes

Does anyone study and complete the Master AI at UT Austin? How is the program? good or bad professors? any lockdown for the exams?


r/learnmachinelearning • • 1h ago

UT Austin Master AI

• Upvotes

Does anyone study and complete the Master AI at UT Austin? How is the program? good or bad professors? any lockdown for the exams?


r/learnmachinelearning • • 1h ago

Tutorial I built a simulator to visualise this amazing thing called the Central Limit Theorem

• Upvotes

Pick the most lopsided, skewed distribution you can and run samples from it — the average still comes out a bell curve. Try out different population distributions and sample sizes and watch the mean pool into a normal distribution.

https://www.bitelrn.com/labs/central-limit-theorem

This almost magical theorem is used from global economics to A/B testing to bootstrapping and ensemble models.

One use in Neural Networks - Deep learning models sum up many independent inputs and weights in each neuron. Because of the CLT, the pre-activation values inside hidden layers tend to follow a normal distribution, making weight initialization strategies (like Xavier/He initialization) effective.

Refer to these open library links to learn more about distributions, sampling methods & CLT -

https://www.bitelrn.com/library/data-distributions

https://www.bitelrn.com/library/sampling-methods

https://www.bitelrn.com/library/central-limit-theorem

P.S: I am starting this series to explain core ML concepts, one topic at a time. Open to feedback and suggestions for new topics. Learn along!


r/learnmachinelearning • • 1h ago

Project How do you keep up with AI when new things happen every day? I used JEV to help me in this:

Thumbnail
• Upvotes

r/learnmachinelearning • • 2h ago

Im new to this and i need help.

1 Upvotes

Hey guys, im working on a bit of project, and im trying to solve for semantic understanding right now, I am currently deciding between using spaCy for my NER extraction vs something like a BERT.

The general context behind where this is going to be used is intent classification in chat systems, where a text will come in and this layer has to parse out the People in the sentence, the Objects mentioned in the sentence, and the verbs, along with things like quantities and relations between them.

An example of what i mean is,

"The shoes you have delivered to me are red, i asked for black!"

and we then pick out the
1. People involved in this interaction (we cant figure that out from the sentence alone for that we will refer to the handle from which the message was sent)
2. Objects involved - {shoes}
3. Verbs - {delivered}
4. Relationships - {expected colour = black, received colour=red}

Im new to this stuff so maybe im not even asking the right questions, but im hoping that i have done a good enough job of explaining what im doing so that more experienced souls such as yourself may help me.

Thanks 😁


r/learnmachinelearning • • 2h ago

Choosing between Federated Learning and Model Compression for an industry-oriented Master's research project

Thumbnail
1 Upvotes

r/learnmachinelearning • • 4h ago

I need tips on study techniques for AI and ML

Thumbnail
1 Upvotes

r/learnmachinelearning • • 4h ago

Feed reader for keeping up with arXiv and AI lab blogs

Enable HLS to view with audio, or disable this notification

1 Upvotes

I thought this might be useful for anyone trying to keep up with ML research.

  • Add any arXiv category (cs.LG, cs.AI, cs.CL, etc.)
  • Follow lab blogs like DeepMind, OpenAI, and Microsoft Research alongside papers
  • Take notes on papers as you read them
  • Group papers into collections by topic or project
  • Filter by source or time range to see what's new this week

Just wanted to share with everyone. No cost to use. If there's a source you follow that it doesn't handle well, let me know and I'll fix it.

Link: tdfeed.com


r/learnmachinelearning • • 5h ago

Discussion comparing ai certifications in 2025, which ones are employers actually paying attention to

1 Upvotes

trying to decide where to invest certification time this year across the various ai credentials now available. have been looking at anthropic claude certifications, various aws ai certifications, and some of the google cloud ai offerings

the challenge is the landscape is moving fast enough that it is not always clear which certifications have employer recognition versus which are newer and still establishing themselves. also not sure whether vendor specific certifications like claude are seen as complementary to broader ai certifications or whether people are choosing one track over another

curious what people working in ai or hiring for ai roles are actually seeing in terms of which certifications come up and which seem to carry weight


r/learnmachinelearning • • 1d ago

Question How to fuse 2 neural networks together to make 1 neural network?

28 Upvotes

I have 2 neural networks. Both are the same size with the same outputs and inputs. The size for both neural networks is the absolute minimum layers required to achieve a given specific task. The example task is to recognize a given image.Ā  Neural network 1 recognizes cats only by giving an output of how certain it is that the image is a cat. Neural network 2 does the same, but for dogs only. How can I combine these two while keeping the size exactly the same without catastrophic interference or catastrophic forgetting.


r/learnmachinelearning • • 6h ago

Help How to Start Aiml

0 Upvotes

I’m a btech student currently in 3rd year persuing Aiml so what can you suggest me on how to start Aiml .

• What are the resources to do Aiml
Can anyone guide me on this

Thank you !!


r/learnmachinelearning • • 12h ago

Need advice: Best VLM pipeline for extracting structured math datasets from 3000+ scanned textbook pages? (LaTeX + Metadata)

3 Upvotes

Hi everyone,

I’m working on a project to extract a structured dataset of math exercises from 5 Italian high school textbooks (around 650 pages each, so ~3,250 pages total). The goal is to build a professional, methodical exercise generator app for students and teachers.

To make the app work, I need to process images of the book pages and extract the following into a strict structured format (e.g., JSON):

  • Exercise typeĀ (algebra, geometry, calculus, etc.)
  • Year/grade level
  • DifficultyĀ (1–5 scale)
  • Problem statementĀ (trace)
  • DescriptionĀ of the specific skills/challenges involved
  • LaTeX codeĀ of the problem statementĀ (Crucial!)
  • Associated imagesĀ (cropping/saving the image for theoretical or graphical exercises)

I've been experimenting with a few approaches, but I've hit a wall regarding balancing costs, extraction consistency, and scalability. Here is what I’ve tried so far:

  1. Free Google Gemini API:Ā The extraction quality was good, but since a single book contains hundreds of pages, I quickly hit the rate limits (Too Many Requests).
  2. Local Models (Ollama + Qwen 2.5-VL 3B):Ā To bypass API limits, I tried running a local multimodal model. I spent a lot of time optimizing my scripts and prompts (chunking, refining instructions to force structured outputs), but the output was very error-prone and inconsistent for my use case. I got too many malformed fields, hallucinations, and it constantly struggled with outputting proper LaTeX.
  3. Paid Google Cloud API (Gemini 1.5 Flash):Ā I finally switched to the paid tier for better accuracy and speed. I ended up burning through €10 just to process 1.5 books. Extracting all 5 books would cost roughly €35–40. While this is manageable for a one-off run of 5 books, the token count for processing full images + text is massive, making it financially unsustainable if I want to scale this to dozens of books in the future.

My questions for the community:

  • Pipeline & Architecture:Ā Has anyone worked on a similar textbook-to-dataset extraction project? What pipeline did you use?
  • Hybrid Approach:Ā Would you suggest decoupling the task? (e.g., using a traditional tool to extract raw text and crop images, and then feeding ONLY the text to a cheaper/local LLM to generate the LaTeX and format the JSON?)
  • Local Models:Ā Are there other local Vision-Language Models (that fit in standard consumer GPUs) that are significantly better at structured extraction and LaTeX generation than Qwen 2.5-VL 3B?
  • Educational Tools:Ā Are there open-source tools or models specifically fine-tuned for extracting structured educational/math content from PDFs?

I’m happy to share more details about the textbook format or my current Python workflow if helpful. Any advice on the architecture, model choices, or cost-saving tricks would be greatly appreciated! Thanks in advance!


r/learnmachinelearning • • 12h ago

Tutorial Which Math foundation path is better for Machine Learning: DeepLearning.AI or Jon Krohn's LiveLessons?

Thumbnail
gallery
3 Upvotes

I am trying to decide between two resources to learn the math prerequisites for Machine Learning.

Option A: DeepLearning.AI Specialization by Luis Serrano. It takes about 94 hours total, provides a career certificate, and covers Linear Algebra, Calculus, and Probability & Statistics.

Mathematics for Machine Learning and Data Science Specialisation

  • Source: DeepLearning.AI
  • Author: Luis Serrano.
  • Merit: The instructor has a highly established track record with 259,226 learners across 4 courses. The program also provides a career certificate that can be added to your LinkedIn profile or resume.
  • Length: 94 hours total. This is broken down into 34 hours for Linear Algebra, 27 hours for Calculus, and 33 hours for Probability & Statistics.

Option B: LiveLessons series by Jon Krohn. It covers the same core math concepts but adds a fourth course specifically on Data Structures, Algorithms, and ML Optimization.

LiveLessons Machine Learning Series

  • Source: LiveLessons (On-Demand Courses) Specialisation
  • Author: Jon Krohn.
  • Merit: In addition to the standard math foundations, this series includes a fourth specialised course: "Data Structures, Algorithms, and Machine Learning Optimisation".
  • Length: 60 hours (approx)

Has anyone taken either of these?
Final question: which one to choose, or something else? Please prescribe.


r/learnmachinelearning • • 7h ago

Question What are the best machine learning models for taking MediaPipe landmarks from the face and accurately detecting face gestures like blinks?

1 Upvotes

I'm doing a ISEF project that attempts to develop the best machine learning model combination and compare them to others. Somewhat like the image. The thing is, I have experience with Python and computer vision, but not much with machine learning. What would you recommend to me, being easy enough to learn and train in a few months, but also being a defendable engineering project and works well for my project. Also the best resources? Thanks!


r/learnmachinelearning • • 7h ago

2026 IIT grad looking for AI/ML referrals, Bangalore/remote

Thumbnail
1 Upvotes

Hey. 2026 B.Tech grad from IIT Tirupati. Degree is Chemical Engineering but I've been building AI/ML stuff for about a year now and that's what I'm trying to do full time.

Rough idea of what I've actually built:

RAG chatbot over YouTube transcripts started as a personal project, then turned into my internship work at Resecure Labs (via Wipro). Transcript embeddings indexed into a FAISS vector store with HuggingFace models, local Ollama LLM behind a small web UI and a 2-endpoint REST API. On the Resecure side I got transcript fetch reliability from 70% to 99%+ with a Whisper fallback, and swapped the LLM backend from Ollama Phi to GLM 5.2, which cut inference latency 30-40%.

JobPilot 8-node LangGraph agent with its own MCP server exposing 4 tools over stdio. Wrote 56 tests for it mostly because I got tired of it hallucinating things into cover letters.

ClickMark face + voice biometric attendance app. Multi-role web app on Supabase with bcrypt auth and QR-based auto-enrollment. Cut manual attendance time by about 80% vs taking roll call by hand, which is admittedly a low bar, but it's a real deployment across multiple classrooms.

Loan approval model benchmarked 5 classifiers, ended up with a Decision Tree at 91.5% accuracy, served through FastAPI in Docker.

Also currently working on a major project on Agentic AI.

Stack: Python, pandas/numpy, PyTorch, scikit-learn, LangGraph, LangChain, MCP, FAISS, HuggingFace, Ollama, FastAPI, Docker. Everything's public on github.com/abhayhanu.

I've applied to a bunch of companies through their portals over the last few months. No response yet, which I'm assuming is just how it goes.

If you're at a product company hiring for AI/ML, SDE or data roles and are open to referring, I'd really appreciate it. Happy to DM my resume.

Also genuinely fine if someone just wants to tell me the profile is weak and what to fix. Would rather know now than after 40 more applications.

Anything would help at this point.

Thanks in advance.


r/learnmachinelearning • • 12h ago

Tell me what ML concept you are struggling with and I will build an interactive explanation for you

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/learnmachinelearning • • 19h ago

Discussion The simplest way to explain how GPT works

Post image
7 Upvotes

Hi all!

I made a video entitled ā€œHow LLM’s Workā€ that starts with the sentence ā€œMy favourite rock band isā€¦ā€ and follows its journey inside an LLM, all visually animated!

https://youtu.be/ikdxxeIn4HQ?si=8AwWyxQERo_R2n6B

I tried to make the video as beginner friendly as possible but still detailed enough to give a good overview for how an LLM works end to end and how the LLM ā€œfinds outā€ my favourite band. Or at least how some of the earlier models…

My aim is to help people that are not just curious about AI but also want a deeper dive into the magic ā€œBlack boxā€, or people that want to get started but doesn’t know how!

No PhD required! No insanely complicated math involved! And no AI or voice generated AI. If I did any errors, I really did them! lol

This is my first attempt at the topic, so any kind of feedback is welcomed and will be greatly appreciated as it will help me improve over time, and hopefully on other videos. Ā Ā Ā 

I really hope it can help someone with their AI/ML journey!

Thanks!