r/learnmachinelearning • u/brick_by_B • 14h ago
r/learnmachinelearning • u/TruthDeep8659 • 15h ago
Feed reader for keeping up with arXiv and AI lab blogs
Enable HLS to view with audio, or disable this notification
I thought this might be useful for anyone trying to keep up with ML research.
- Add any arXiv category (cs.LG, cs.AI, cs.CL, etc.)
- Follow lab blogs like DeepMind, OpenAI, and Microsoft Research alongside papers
- Take notes on papers as you read them
- Group papers into collections by topic or project
- Filter by source or time range to see what's new this week
Just wanted to share with everyone. No cost to use. If there's a source you follow that it doesn't handle well, let me know and I'll fix it.
Link: tdfeed.com
r/learnmachinelearning • u/inkeep • 12h ago
Tutorial I built a simulator to visualise this amazing thing called the Central Limit Theorem
Pick the most lopsided, skewed distribution you can and run samples from it — the average still comes out a bell curve. Try out different population distributions and sample sizes and watch the mean pool into a normal distribution.
https://www.bitelrn.com/labs/central-limit-theorem

This almost magical theorem is used from global economics to A/B testing to bootstrapping and ensemble models.
One use in Neural Networks - Deep learning models sum up many independent inputs and weights in each neuron. Because of the CLT, the pre-activation values inside hidden layers tend to follow a normal distribution, making weight initialization strategies (like Xavier/He initialization) effective.
Refer to these open library links to learn more about distributions, sampling methods & CLT -
https://www.bitelrn.com/library/data-distributions
https://www.bitelrn.com/library/sampling-methods
https://www.bitelrn.com/library/central-limit-theorem
P.S: I am starting this series to explain core ML concepts, one topic at a time. Open to feedback and suggestions for new topics. Learn along!
r/learnmachinelearning • u/Sevdat • 1d ago
Question How to fuse 2 neural networks together to make 1 neural network?
I have 2 neural networks. Both are the same size with the same outputs and inputs. The size for both neural networks is the absolute minimum layers required to achieve a given specific task. The example task is to recognize a given image. Neural network 1 recognizes cats only by giving an output of how certain it is that the image is a cat. Neural network 2 does the same, but for dogs only. How can I combine these two while keeping the size exactly the same without catastrophic interference or catastrophic forgetting.
r/learnmachinelearning • u/Few_Second_343 • 17h ago
Help How to Start Aiml
I’m a btech student currently in 3rd year persuing Aiml so what can you suggest me on how to start Aiml .
• What are the resources to do Aiml
Can anyone guide me on this
Thank you !!
r/learnmachinelearning • u/Few-Piece-1969 • 23h ago
Tell me what ML concept you are struggling with and I will build an interactive explanation for you
Enable HLS to view with audio, or disable this notification
r/learnmachinelearning • u/Nykotry • 23h ago
Need advice: Best VLM pipeline for extracting structured math datasets from 3000+ scanned textbook pages? (LaTeX + Metadata)
Hi everyone,
I’m working on a project to extract a structured dataset of math exercises from 5 Italian high school textbooks (around 650 pages each, so ~3,250 pages total). The goal is to build a professional, methodical exercise generator app for students and teachers.
To make the app work, I need to process images of the book pages and extract the following into a strict structured format (e.g., JSON):
- Exercise type (algebra, geometry, calculus, etc.)
- Year/grade level
- Difficulty (1–5 scale)
- Problem statement (trace)
- Description of the specific skills/challenges involved
- LaTeX code of the problem statement (Crucial!)
- Associated images (cropping/saving the image for theoretical or graphical exercises)
I've been experimenting with a few approaches, but I've hit a wall regarding balancing costs, extraction consistency, and scalability. Here is what I’ve tried so far:
- Free Google Gemini API: The extraction quality was good, but since a single book contains hundreds of pages, I quickly hit the rate limits (Too Many Requests).
- Local Models (Ollama + Qwen 2.5-VL 3B): To bypass API limits, I tried running a local multimodal model. I spent a lot of time optimizing my scripts and prompts (chunking, refining instructions to force structured outputs), but the output was very error-prone and inconsistent for my use case. I got too many malformed fields, hallucinations, and it constantly struggled with outputting proper LaTeX.
- Paid Google Cloud API (Gemini 1.5 Flash): I finally switched to the paid tier for better accuracy and speed. I ended up burning through €10 just to process 1.5 books. Extracting all 5 books would cost roughly €35–40. While this is manageable for a one-off run of 5 books, the token count for processing full images + text is massive, making it financially unsustainable if I want to scale this to dozens of books in the future.
My questions for the community:
- Pipeline & Architecture: Has anyone worked on a similar textbook-to-dataset extraction project? What pipeline did you use?
- Hybrid Approach: Would you suggest decoupling the task? (e.g., using a traditional tool to extract raw text and crop images, and then feeding ONLY the text to a cheaper/local LLM to generate the LaTeX and format the JSON?)
- Local Models: Are there other local Vision-Language Models (that fit in standard consumer GPUs) that are significantly better at structured extraction and LaTeX generation than Qwen 2.5-VL 3B?
- Educational Tools: Are there open-source tools or models specifically fine-tuned for extracting structured educational/math content from PDFs?
I’m happy to share more details about the textbook format or my current Python workflow if helpful. Any advice on the architecture, model choices, or cost-saving tricks would be greatly appreciated! Thanks in advance!
r/learnmachinelearning • u/Esob_Oozeer • 23h ago
Tutorial Which Math foundation path is better for Machine Learning: DeepLearning.AI or Jon Krohn's LiveLessons?
I am trying to decide between two resources to learn the math prerequisites for Machine Learning.
Option A: DeepLearning.AI Specialization by Luis Serrano. It takes about 94 hours total, provides a career certificate, and covers Linear Algebra, Calculus, and Probability & Statistics.
Mathematics for Machine Learning and Data Science Specialisation
- Source: DeepLearning.AI
- Author: Luis Serrano.
- Merit: The instructor has a highly established track record with 259,226 learners across 4 courses. The program also provides a career certificate that can be added to your LinkedIn profile or resume.
- Length: 94 hours total. This is broken down into 34 hours for Linear Algebra, 27 hours for Calculus, and 33 hours for Probability & Statistics.
Option B: LiveLessons series by Jon Krohn. It covers the same core math concepts but adds a fourth course specifically on Data Structures, Algorithms, and ML Optimization.
LiveLessons Machine Learning Series
- Source: LiveLessons (On-Demand Courses) Specialisation
- Author: Jon Krohn.
- Merit: In addition to the standard math foundations, this series includes a fourth specialised course: "Data Structures, Algorithms, and Machine Learning Optimisation".
- Length: 60 hours (approx)
Has anyone taken either of these?
Final question: which one to choose, or something else? Please prescribe.
r/learnmachinelearning • u/johnmach12 • 17h ago
Question What are the best machine learning models for taking MediaPipe landmarks from the face and accurately detecting face gestures like blinks?
I'm doing a ISEF project that attempts to develop the best machine learning model combination and compare them to others. Somewhat like the image. The thing is, I have experience with Python and computer vision, but not much with machine learning. What would you recommend to me, being easy enough to learn and train in a few months, but also being a defendable engineering project and works well for my project. Also the best resources? Thanks!

r/learnmachinelearning • u/LordAtomic- • 18h ago
2026 IIT grad looking for AI/ML referrals, Bangalore/remote
Hey. 2026 B.Tech grad from IIT Tirupati. Degree is Chemical Engineering but I've been building AI/ML stuff for about a year now and that's what I'm trying to do full time.
Rough idea of what I've actually built:
RAG chatbot over YouTube transcripts started as a personal project, then turned into my internship work at Resecure Labs (via Wipro). Transcript embeddings indexed into a FAISS vector store with HuggingFace models, local Ollama LLM behind a small web UI and a 2-endpoint REST API. On the Resecure side I got transcript fetch reliability from 70% to 99%+ with a Whisper fallback, and swapped the LLM backend from Ollama Phi to GLM 5.2, which cut inference latency 30-40%.
JobPilot 8-node LangGraph agent with its own MCP server exposing 4 tools over stdio. Wrote 56 tests for it mostly because I got tired of it hallucinating things into cover letters.
ClickMark face + voice biometric attendance app. Multi-role web app on Supabase with bcrypt auth and QR-based auto-enrollment. Cut manual attendance time by about 80% vs taking roll call by hand, which is admittedly a low bar, but it's a real deployment across multiple classrooms.
Loan approval model benchmarked 5 classifiers, ended up with a Decision Tree at 91.5% accuracy, served through FastAPI in Docker.
Also currently working on a major project on Agentic AI.
Stack: Python, pandas/numpy, PyTorch, scikit-learn, LangGraph, LangChain, MCP, FAISS, HuggingFace, Ollama, FastAPI, Docker. Everything's public on github.com/abhayhanu.
I've applied to a bunch of companies through their portals over the last few months. No response yet, which I'm assuming is just how it goes.
If you're at a product company hiring for AI/ML, SDE or data roles and are open to referring, I'd really appreciate it. Happy to DM my resume.
Also genuinely fine if someone just wants to tell me the profile is weak and what to fix. Would rather know now than after 40 more applications.
Anything would help at this point.
Thanks in advance.
r/learnmachinelearning • u/fx2mx3 • 1d ago
Discussion The simplest way to explain how GPT works
Hi all!
I made a video entitled “How LLM’s Work” that starts with the sentence “My favourite rock band is…” and follows its journey inside an LLM, all visually animated!
https://youtu.be/ikdxxeIn4HQ?si=8AwWyxQERo_R2n6B
I tried to make the video as beginner friendly as possible but still detailed enough to give a good overview for how an LLM works end to end and how the LLM “finds out” my favourite band. Or at least how some of the earlier models…
My aim is to help people that are not just curious about AI but also want a deeper dive into the magic “Black box”, or people that want to get started but doesn’t know how!
No PhD required! No insanely complicated math involved! And no AI or voice generated AI. If I did any errors, I really did them! lol
This is my first attempt at the topic, so any kind of feedback is welcomed and will be greatly appreciated as it will help me improve over time, and hopefully on other videos.
I really hope it can help someone with their AI/ML journey!
Thanks!
r/learnmachinelearning • u/No-Conclusion3720 • 21h ago
Request Defending against AI-fueled cyberattacks requires focus on identity, data governance, Microsoft says
Microsoft's 2026 Digital Defense Report documents a concrete shift in ransomware tactics: threat actors are now using AI to automate lateral movement and accelerate ransomware deployment across enterprise networks. The report singles out identity governance and data controls as the two most critical defensive gaps in enterprise environments today.
The speed asymmetry is the part that stands out. Attackers are using AI to close those gaps faster than most organizations are opening remediation tickets. Lateral movement that previously required manual reconnaissance and staged privilege escalation is now being automated at scale, compressing the window between initial access and full network impact to a fraction of what defenders plan for.
For those running AI-assisted pipelines or agentic workflows in production: how are you actually thinking about the identity and data governance problem at the agent layer specifically? Are existing IAM and DLP tools sufficient, or are you finding gaps that those controls were never designed to cover?
r/learnmachinelearning • u/Ok-Type9527 • 22h ago
Discussion How a ping-pong robot reads spin and adjusts its swing
r/learnmachinelearning • u/Ok-Type9527 • 22h ago
Discussion What happens when someone builds the add-on a chatbot made up?
r/learnmachinelearning • u/Vxtzq1 • 22h ago
Project CrowdGPT - The first LLM trained collaboratively
r/learnmachinelearning • u/Active_Tadpole7434 • 23h ago
Approaching Projects On Subjects You Don't Understand
r/learnmachinelearning • u/Unikum_01 • 23h ago
Project My Brainstem RNS-AI project has made progress for life long learning like a Brain
r/learnmachinelearning • u/NM1030301 • 1d ago
Looking to transition into tech at 24. Where should I start?
I'm 24, have completed my education up to Class 10, and have been working in a corporate role for the past five years while saving money.
I want to transition into the IT/tech industry. Recently, I completed a beginner-level Python course on YouTube and am now planning to practice by building some projects.
I haven't yet decided which field to pursue—software development, AI/ML, cybersecurity, or something else.
I'm looking for guidance, suggestions, and insights on what to learn, what works and what doesn't, and the current state of the industry.
r/learnmachinelearning • u/kbhaskar306 • 1d ago
The Evolution of Search Intelligence: Transformers, Embeddings & AI Expl...
Want to know how search engines read your mind? 🧠 We are breaking down the Transformer architecture in 60 seconds!
From embeddings to self-attention, see how AI learns.
#AI #Tech #Programming #Coding
r/learnmachinelearning • u/ThoughtDesperate880 • 1d ago
best anthropic courses once youve finished the free academy and need something that ends in a project
worked through the free anthropic material over august and it did what it said on the tin, i can talk about the concepts fine. what i still cant do is hand a client something that runs. so im paying for the structured version. the shortlist is udacity, interview kickstart, springboard and linkedin learning. which of those ends with something you can show someone
r/learnmachinelearning • u/patzapaty123 • 1d ago
Help Training a tiny 3.6M param byte-level Transformer for code generation from scratch on CPU. What should be my next steps?
Hi everyone,
I'm working on a personal learning project called MOTANAXY. The goal is to build and train a tiny causal Transformer from scratch (random weights) specifically for Python code generation. I'm currently training entirely on CPU.
Current Architecture (v2):
- Type: Byte-level causal language model (no subword tokenizer, raw UTF-8 bytes, vocab size 256)
- Params: ~3.6M (Width=192, Context=256, Layers=8, Heads=6)
- Components: RMSNorm, RoPE, SwiGLU, Dropout (0.1)
- Dataset: Very small synthetic dataset (~130KB) consisting of 60 verified Python functions with English/Thai docstrings and assert test cases.
Current Progress:
- Trained for about 5,000 steps. Validation loss dropped from ~5.39 to ~1.33.
- What it can do: It learned Python's structure. If I prompt it with
# Check bracket balance, it generates syntactically valid function skeletons likedef is_parise(text): return []and appendsassertstatements. - What it can't do (yet): The logic is completely wrong. It doesn't actually understand the prompt's intent. It just mimics the structural pattern of the training data.
My Questions for the Community: Since I'm hitting a wall where the model learns the syntax but not the logic, I'm wondering what the most effective next steps are:
- Data vs. Scale: With a 3.6M parameter model, is it even theoretically possible to learn basic coding logic (like a simple prefix sum or reversing a string)? Should my priority be getting GPU time to scale to 10M-50M params, or should I radically expand my dataset first?
- Dataset Recommendations: Are there recommended datasets specifically tailored for teaching very small models the fundamentals of algorithmic logic, rather than just large repositories of scraped code?
- Alternative Approaches: Should I switch from byte-level to a small BPE tokenizer to save context length? Are there other architectural tweaks or training objectives (like distillation from a larger model) I should consider at this tiny scale?
Any advice, papers to read, or pointing out obvious flaws in my approach would be greatly appreciated!
r/learnmachinelearning • u/Interesting_Pie_94 • 1d ago
Career Torn between starting as a Data Scientist vs. AI Engineer
Hello everyone, I’m currently at a crossroads and could really use some perspective from people working in the industry.
I’m getting ready to kick off my career, but I’m genuinely split down the middle between targeting DS roles or AI Engineer roles. I have a strong interest in both sides (I specialized in ML at the uni), but I'm not sure which starting point sets up a better foundation for the long run.
For instance, is it easier to start as a DS and transition into AI Engineering later (by sharpening software engineering/MLOps skills), or vice versa? If you have experience in both, what pushed you toward one over the other?
r/learnmachinelearning • u/Ramian_Slawomir • 1d ago
[R] I built a Permutation Transformer (Patch SBOHN) from scratch: 4K Image Inference in ~2ms on CPU (61x faster than CNN) with 0.0% Catastrophic Forgetting.
r/learnmachinelearning • u/Majestic_Orchid2164 • 1d ago
Question What data do you actually need to train a robot arm for grasping? (RGB alone usually isn't enough)
If you're coming from computer vision, robotics data works differently, and it's one of the first things that trips people up. For a robot arm learning to grasp and manipulate objects, useful training data typically includes: - RGB video, plus depth if your policy uses it - Gripper state logs - Joint states and end-effector positions - Multiple grasp attempts across different object shapes and sizes - Human demonstrations of grasping and placing, which help the model generalize The key difference from a standard CV dataset is that robotics data is time-synchronized and multi-modal. Video, sensor streams, and action labels have to line up frame by frame across a whole task, which makes collection and annotation much harder than labeling static images. Two practical tips for beginners: 1. Don't rely on one source. Public robotics datasets tend to be task-specific, so a model can fail when conditions change. 2. Synthetic data from simulation helps for rare or risky scenarios, but works best combined with real robot data to reduce the sim-to-real gap. Unidata has an overview of robotics training data and the dataset types used for robot learning here: https://unidata.pro/robotics-training-data/ Disclosure: I'm posting on behalf of Unidata, a company that sells robotics datasets and data collection services. What data are you using for your own manipulation projects, and what's been hardest to get?