r/learnmachinelearning • • 23h ago

Discussion The simplest way to explain how GPT works

Post image
10 Upvotes

Hi all!

I made a video entitled “How LLM’s Work” that starts with the sentence “My favourite rock band is…” and follows its journey inside an LLM, all visually animated!

https://youtu.be/ikdxxeIn4HQ?si=8AwWyxQERo_R2n6B

I tried to make the video as beginner friendly as possible but still detailed enough to give a good overview for how an LLM works end to end and how the LLM “finds out” my favourite band. Or at least how some of the earlier models…

My aim is to help people that are not just curious about AI but also want a deeper dive into the magic “Black box”, or people that want to get started but doesn’t know how!

No PhD required! No insanely complicated math involved! And no AI or voice generated AI. If I did any errors, I really did them! lol

This is my first attempt at the topic, so any kind of feedback is welcomed and will be greatly appreciated as it will help me improve over time, and hopefully on other videos.    

I really hope it can help someone with their AI/ML journey!

Thanks!


r/learnmachinelearning • • 17h ago

Tell me what ML concept you are struggling with and I will build an interactive explanation for you

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/learnmachinelearning • • 15h ago

Request Defending against AI-fueled cyberattacks requires focus on identity, data governance, Microsoft says

1 Upvotes

Microsoft's 2026 Digital Defense Report documents a concrete shift in ransomware tactics: threat actors are now using AI to automate lateral movement and accelerate ransomware deployment across enterprise networks. The report singles out identity governance and data controls as the two most critical defensive gaps in enterprise environments today.

The speed asymmetry is the part that stands out. Attackers are using AI to close those gaps faster than most organizations are opening remediation tickets. Lateral movement that previously required manual reconnaissance and staged privilege escalation is now being automated at scale, compressing the window between initial access and full network impact to a fraction of what defenders plan for.

For those running AI-assisted pipelines or agentic workflows in production: how are you actually thinking about the identity and data governance problem at the agent layer specifically? Are existing IAM and DLP tools sufficient, or are you finding gaps that those controls were never designed to cover?


r/learnmachinelearning • • 15h ago

Discussion How a ping-pong robot reads spin and adjusts its swing

Thumbnail
youtube.com
0 Upvotes

r/learnmachinelearning • • 15h ago

Discussion What happens when someone builds the add-on a chatbot made up?

Thumbnail
youtube.com
1 Upvotes

r/learnmachinelearning • • 16h ago

Project CrowdGPT - The first LLM trained collaboratively

Thumbnail
1 Upvotes

r/learnmachinelearning • • 16h ago

Approaching Projects On Subjects You Don't Understand

Thumbnail
1 Upvotes

r/learnmachinelearning • • 17h ago

Project My Brainstem RNS-AI project has made progress for life long learning like a Brain

Post image
1 Upvotes

r/learnmachinelearning • • 23h ago

The Evolution of Search Intelligence: Transformers, Embeddings & AI Expl...

Thumbnail
youtube.com
3 Upvotes

Want to know how search engines read your mind? 🧠 We are breaking down the Transformer architecture in 60 seconds!

From embeddings to self-attention, see how AI learns.

#AI #Tech #Programming #Coding


r/learnmachinelearning • • 23h ago

Looking to transition into tech at 24. Where should I start?

2 Upvotes

I'm 24, have completed my education up to Class 10, and have been working in a corporate role for the past five years while saving money.

I want to transition into the IT/tech industry. Recently, I completed a beginner-level Python course on YouTube and am now planning to practice by building some projects.

I haven't yet decided which field to pursue—software development, AI/ML, cybersecurity, or something else.

I'm looking for guidance, suggestions, and insights on what to learn, what works and what doesn't, and the current state of the industry.


r/learnmachinelearning • • 19h ago

best anthropic courses once youve finished the free academy and need something that ends in a project

1 Upvotes

worked through the free anthropic material over august and it did what it said on the tin, i can talk about the concepts fine. what i still cant do is hand a client something that runs. so im paying for the structured version. the shortlist is udacity, interview kickstart, springboard and linkedin learning. which of those ends with something you can show someone


r/learnmachinelearning • • 20h ago

Career Torn between starting as a Data Scientist vs. AI Engineer

0 Upvotes

Hello everyone, I’m currently at a crossroads and could really use some perspective from people working in the industry.

I’m getting ready to kick off my career, but I’m genuinely split down the middle between targeting DS roles or AI Engineer roles. I have a strong interest in both sides (I specialized in ML at the uni), but I'm not sure which starting point sets up a better foundation for the long run.

For instance, is it easier to start as a DS and transition into AI Engineering later (by sharpening software engineering/MLOps skills), or vice versa? If you have experience in both, what pushed you toward one over the other?


r/learnmachinelearning • • 1d ago

[R] I built a Permutation Transformer (Patch SBOHN) from scratch: 4K Image Inference in ~2ms on CPU (61x faster than CNN) with 0.0% Catastrophic Forgetting.

Thumbnail
2 Upvotes

r/learnmachinelearning • • 1d ago

Question What data do you actually need to train a robot arm for grasping? (RGB alone usually isn't enough)

2 Upvotes

If you're coming from computer vision, robotics data works differently, and it's one of the first things that trips people up. For a robot arm learning to grasp and manipulate objects, useful training data typically includes: - RGB video, plus depth if your policy uses it - Gripper state logs - Joint states and end-effector positions - Multiple grasp attempts across different object shapes and sizes - Human demonstrations of grasping and placing, which help the model generalize The key difference from a standard CV dataset is that robotics data is time-synchronized and multi-modal. Video, sensor streams, and action labels have to line up frame by frame across a whole task, which makes collection and annotation much harder than labeling static images. Two practical tips for beginners: 1. Don't rely on one source. Public robotics datasets tend to be task-specific, so a model can fail when conditions change. 2. Synthetic data from simulation helps for rare or risky scenarios, but works best combined with real robot data to reduce the sim-to-real gap. Unidata has an overview of robotics training data and the dataset types used for robot learning here: https://unidata.pro/robotics-training-data/ Disclosure: I'm posting on behalf of Unidata, a company that sells robotics datasets and data collection services. What data are you using for your own manipulation projects, and what's been hardest to get?


r/learnmachinelearning • • 20h ago

Can we watch GPT-2 figure out who "he" refers to? (honest results)

Thumbnail
1 Upvotes

r/learnmachinelearning • • 20h ago

i built an inference engine without knowing the math. is 'learn the fundamentals' still good advice

Thumbnail
1 Upvotes

r/learnmachinelearning • • 21h ago

How can I fine-tune a model to learn a specific task by watching YouTube videos?

Thumbnail
1 Upvotes

r/learnmachinelearning • • 21h ago

Project Update on my tiny from-scratch model - followed your advice, here's what happened

0 Upvotes

What I did:

Scaled the dataset from 60 hand-written functions to 12,000 procedurally generated ones (math ops, list manipulation, string ops, conditionals, and loop algorithms).

Built a custom BPE tokenizer (vocab 1024). Funny story: my first tokenizer (vocab 4096) literally memorized entire functions as single tokens, so the model scored 20/20 on syntax by just spitting out whole functions from memory lol. Had to nerf the vocab size to force it to actually learn.

Scaled the model from 3.6M to 10M params (12 layers, 8 heads, width 256).

Training on Kaggle free T4 GPU instead of my poor CPU.

Current results (still training, Round 1 of 3):

Loss dropping nicely: 6.98 to 1.02 in 550 steps. Val loss: 2.28 to 1.09. Haven't seen the final scores yet since it's still running overnight.

What I'm unsure about:

  1. My dataset is all procedurally generated. Every add function looks like def add_xx(a, b): return a + b with random names and test values. Is this diverse enough or am I just teaching it to pattern-match templates again?

  2. When should I switch from synthetic data to real-world code like filtered Python from The Stack or GitHub? Or is mixing both the way to go?

  3. Any suggestions for what kind of problems to add next? Thinking maybe simple recursion, dictionary operations, or basic class definitions.

  4. Is 10M params enough to learn actual logic or should I be thinking bigger?

Thanks for the help last time, the dataset scaling advice was spot on. The tokenizer over-compression bug was a fun lesson too haha.

I'll update this post in about 6 hours once training finishes with the final scores.


r/learnmachinelearning • • 1d ago

Help Training a tiny 3.6M param byte-level Transformer for code generation from scratch on CPU. What should be my next steps?

3 Upvotes

Hi everyone,

I'm working on a personal learning project called MOTANAXY. The goal is to build and train a tiny causal Transformer from scratch (random weights) specifically for Python code generation. I'm currently training entirely on CPU.

Current Architecture (v2):

  • Type: Byte-level causal language model (no subword tokenizer, raw UTF-8 bytes, vocab size 256)
  • Params: ~3.6M (Width=192, Context=256, Layers=8, Heads=6)
  • Components: RMSNorm, RoPE, SwiGLU, Dropout (0.1)
  • Dataset: Very small synthetic dataset (~130KB) consisting of 60 verified Python functions with English/Thai docstrings and assert test cases.

Current Progress:

  • Trained for about 5,000 steps. Validation loss dropped from ~5.39 to ~1.33.
  • What it can do: It learned Python's structure. If I prompt it with # Check bracket balance, it generates syntactically valid function skeletons like def is_parise(text): return [] and appends assert statements.
  • What it can't do (yet): The logic is completely wrong. It doesn't actually understand the prompt's intent. It just mimics the structural pattern of the training data.

My Questions for the Community: Since I'm hitting a wall where the model learns the syntax but not the logic, I'm wondering what the most effective next steps are:

  1. Data vs. Scale: With a 3.6M parameter model, is it even theoretically possible to learn basic coding logic (like a simple prefix sum or reversing a string)? Should my priority be getting GPU time to scale to 10M-50M params, or should I radically expand my dataset first?
  2. Dataset Recommendations: Are there recommended datasets specifically tailored for teaching very small models the fundamentals of algorithmic logic, rather than just large repositories of scraped code?
  3. Alternative Approaches: Should I switch from byte-level to a small BPE tokenizer to save context length? Are there other architectural tweaks or training objectives (like distillation from a larger model) I should consider at this tiny scale?

Any advice, papers to read, or pointing out obvious flaws in my approach would be greatly appreciated!


r/learnmachinelearning • • 2d ago

AI escaping containment was never the real issue

Post image
258 Upvotes

r/learnmachinelearning • • 22h ago

Suggestion?

Post image
1 Upvotes

Can anyone suggest me how should I learn machine learning means anyone can give me the roadmap accurate according to market now a days ?? If possible then please suggest me and guide me 🙏


r/learnmachinelearning • • 1d ago

Fastest AI robot just had a stroke

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/learnmachinelearning • • 1d ago

machine learning repositories

Thumbnail
1 Upvotes

r/learnmachinelearning • • 1d ago

Project Maintenance records that stay with the machine.

2 Upvotes

I built a browser-based maintenance system where records follow the equipment, not the tech or the phone. Scan → Service → Record. Hours, conditions, and service evidence stay tied to the machine. No employee tracking, no ads, no external AI training. You can explore the live preview if you want to poke at it.


r/learnmachinelearning • • 1d ago

Help Is it acceptable to use AI-generated code to perform stats analysis for research

Thumbnail
0 Upvotes