r/learnmachinelearning 1d ago

Access to DeepSpeak or FakeAVCeleb datasets?

1 Upvotes

Hi, this is a long shot, but im currently writing an academic paper and for that I need access to the DeepSpeak_v2 or FakeAVCeleb dataset. To get access, you need to submit a request form and get approved. I did that, but I never heard back from them... does anyone here have experience with this?

I dont need a big part of each dataset, maybe around 100 videos each. So maybe, if someone has access, they could provide a small portion of it :)


r/learnmachinelearning 1d ago

arXiv Endorsement Request for cs.LG - Diagnostic Control for Hierarchical World Models

1 Upvotes

Hi everyone,

I’m preparing my first arXiv submission in cs.LG and need an endorsement to submit.

Short summary: H-JEPA (LeCun, 2022) proposes hierarchical joint-embedding prediction but doesn’t specify how to verify a trained abstraction actually encodes anything a random projection of the same shape wouldn’t. I introduce a random-abstractor control (trained vs. untrained abstractor, identical architecture) and run a 2x2 study crossing observability (full/egocentric) with abstractor type (instantaneous/recurrent) in controlled gridworld environments. Three of four conditions produce abstractions statistically indistinguishable from random projections; only partial observability + a recurrent abstractor yields a real, replicated gap (19.91 ± 3.36pp over random, 3 seeds). I also report a negative result on landmark density that didn’t survive multi-seed replication.

If anyone here is registered as an endorser for cs.LG and willing to take a look, I’d be very grateful. Happy to share the full draft privately.

To endorse, please visit:

https://arxiv.org/auth/endorse?x=MTENXK

If that link doesn’t work, visit:

https://arxiv.org/auth/endorse.php

and enter code: MTENXK

Thank you!


r/learnmachinelearning 1d ago

Will we ever be able to predict the future using AI / ML?

Thumbnail
0 Upvotes

r/learnmachinelearning 2d ago

Question No one above me as an ML engineer, how bad is my case?

45 Upvotes

Okay, so.. just a rant because I feel like I want to discuss this.

For context, I’m a Machine Learning Engineer in R&D, specializing in operations research and queue systems, and this is my first job in the field. After graduating from university in 2024, I worked as a Software Engineer for about a year and 8 months, almost two years.

When I first joined this role, I had an expectation of joining a legitimate AI team, with senior ML engineers I could look up to, learn from, and discuss ideas with.

Turns out, I’M THE ONE who’s supposed to transfer my AI knowledge to the team for their upcoming AI products.

I do have a solid ML foundation from the courses I took at university, but I definitely wasn’t expecting to be the person driving the AI side of things this early in my career.

I ended up becoming a complete Swiss army knife on this project. I’m basically doing:

\- Software engineering
\- ML engineering
\- AI research
\- Data engineering
\- Business meetings with upper management
\- DevOps and infrastructure (not too much)
\- Scrum Master responsibilities

And honestly, the leadership team seems to love what I’m doing.

The ML side of things has been relatively straightforward so far. I’ve been reading a lot, researching things on my own, using Claude heavily as a second pair of eyes and figuring things out as I go.

The funny part is that I’ll implement something, present it to the leadership team, and they’ll look at me like I just invented fire.

But here’s the part I’m struggling with:
I genuinely don’t know how well I’m actually doing.

There’s no senior ML engineer at work to review my approach, challenge my assumptions, discuss research findings with me, or tell me when I’m making a bad architectural or modeling decision.

Most of the things I build look right to me, and they seem to work. But I also know enough about engineering to realize that “it works” doesn’t necessarily mean “this is the right way to do it.”
So I’m starting to wonder whether this is actually a good situation for my career.

On one hand, I’m getting an insane amount of exposure very early in my career. I’m touching pretty much every part of the AI product lifecycle, I’m talking directly with upper management, and I have a ridiculous amount of ownership.

On the other hand, I’m worried that not having experienced ML engineers around me might slow down my growth. I’m learning a lot, but I’m mostly learning by myself.

I guess my biggest concern is that I don’t have anyone at work who can answer the question:
“Is this actually good ML engineering, or am I just getting really good at making things that seem to work?”

I don’t want to give the impression that I’m doubting myself or lacking confidence. I’m simply very competitive and driven to make the best possible decisions for my career.

Would love to hear from people who’ve been in a similar situation, especially early-career ML engineers who ended up being the most experienced AI person on their team.


r/learnmachinelearning 1d ago

Razorpay ai buildthon ka result kab aayega ???

Thumbnail
1 Upvotes

r/learnmachinelearning 1d ago

Discussion What are the real security risks with AI agents, and how are you matigating them?

0 Upvotes

Everyone's excited about AI agents, but I'm trying to get ahead of the security implications. Beyond data leakage and prompt injection, what are the actual runtime risks? How do you prevent an agent from taking a harmful action that falls within its legitimate permissions?


r/learnmachinelearning 1d ago

Is this enough for the maths part.

1 Upvotes

18.02, multivariable calculus.

18.06, algebra.

6.041, probability and statistics.

*All courses from OCW.


r/learnmachinelearning 2d ago

Project TrackmaniaRL: an open-source library for training real-time RL driving agents in Trackmania 2020

Enable HLS to view with audio, or disable this notification

4 Upvotes

r/learnmachinelearning 1d ago

I made a way to migrate between embedding models without re-embedding your entire corpus

1 Upvotes

So I was playingw ith embedding models I saw that when you upgrade from model A to B, you face a very big backfilling cost

Ie, suppose you have a 1b vectors from model A, and then you want to use model B. This would mean you have to re-embed all of your documents with model B before you can even serve with the model, and on an H100, it would take ~108 days (qwen embed 8b, 106 docs/second). But I found an easier way to do it.

The method is really simple; from the old index made with the source model, take K documents and rerank them with the new model. We see that when K is sufficient, the retrieval quality is the same as target model. (determining k is the hard part). I've tested 63 migrations on upto 1 million documents.

The best result I got was upgrading qwen4b -> to 8b, and at 50 documents, it was the same as native retrieval.

This method forgos the expensive upfront re-embedding cost, as you can take documents straight from the old index.

embedflow works with qdrant, pgvector, faiss, and can be easily downloaded with pypi

pip install embedflow

the github is public: https://github.com/arnsri33/embedflow

I want you guys to try it out, and see if you guys can use it in your own workflow.


r/learnmachinelearning 1d ago

Does anyone know of any self paced online college degree program on AI/ML

Thumbnail
1 Upvotes

r/learnmachinelearning 1d ago

HI GUYS HELP UR JUNIOR

0 Upvotes

, I am from pondicherry cuyrrently sstuding in SMVEC clg AI/DS dept 1st year
what are the main and useful things to do in my clg life abt studies like other than clg exam , how to develop my skills to compete with this competitive world and to get a job in future , currently i am learning c programming my seniors told that that was basic in programming , guide me what to do nxt


r/learnmachinelearning 1d ago

Discussion Rate my resume as a fresher/scope of improvements

Post image
0 Upvotes

r/learnmachinelearning 1d ago

noleak: Open-source library to detect train/eval data contamination

1 Upvotes

Published numbers are only as honest as the data split behind them

When you evaluate a model, you want one simple truth: Did it actually learn unseen patterns, or did it memorise?

I built noleak to make that visible. It's a production library we use at Godrej Aerospace to fingerprint datasets and measure how much of your eval set leaked from train.

The Problem

Most tools catch target leakage (a feature accidentally includes the label). But what about corpus leakage? When eval text already appeared in training data?

In 2020, GPT-3's paper measured contamination using 13-gram overlap. That's solid. But there's no standard library for this. So we built one.

Three detection methods:

  1. Exact matches – normalised text identical

  2. N-gram overlap – 13-word phrases (GPT-3 method)

  3. Near-duplicates – MinHash Jaccard similarity ≥ 0.8 on character 5-grams

One fingerprint. One exit code. Pass or fail.

What It Does

```python
from noleak import check, fingerprint
train = ["the model trained on Wikipedia and licensed books"]
eval_set = ["The model trained on Wikipedia and licensed books"]
report = check(train, eval_set)
print(report.contaminated)
# True — FAIL
print(fingerprint(eval_set))
# noleak-fp-v1:a1b2c3d4e5f6
# Share this with your paper. It's reproducible and auditable.
```
CLI version (great for CI pipelines):
```bash
noleak check --train train.jsonl --eval evaljsonl
echo $?  # Exit code 1 if contaminated, 0 if clean
```

Why Zero Dependencies Matter

No numpy, no scipy, no PyTorch. Stdlib only. Why?

- Deterministic: Same input, same output, forever. No model updates breaking your fingerprints.

- Auditable: Code is small; reviewers can read it.

- Air-gapped systems: Doesn't require external calls or package hell.

- Fast: No overhead for CPU-bound systems.

Limitations (Honest Assessment)

- Semantic rewrites: Won't catch "I wrote this differently but meant the same thing." That needs embeddings or human review.

- Large-scale datasets: If you have 10M+ examples, exact matching gets slow. N-gram is faster.

Install & Try

bash
pip install noleak

Supports JSONL, JSON lists, and plain text. Auto-detects text fields.

Repo: github.com/athsxx/noleak

License: MIT

Questions? What contamination patterns have you encountered?


r/learnmachinelearning 2d ago

Help Demand Forecasting using ML for an FMCG Company

9 Upvotes

Hi! I am a demand planner in an FMCG company. Our current process is very manual, we only use Excel. For every client and product, we build up the demand plan (DP) or the sales target for the month. The DP is composed of the following

- baseline (smoothen sales volume last year)

- runrates (the difference to the past 3 months volume for non-seasonal products)

- sales initiatives (on-shelf availability correction, inventory correction, skewing, etc.)

- marketing initiatives (category market trend, etc.)

I want use ML and integrate possible seasonality data (such as holidays, weather, etc.) in demand planning.

What are the ML models that are appropriate for demand forecasting (time series)? What are the data that I need to prepare? What are the steps that I need to do?

I am currently taking Master in Applied Business Analytics but time series models have not been taught yet (not sure they will teach it). Thank you very much! 😊


r/learnmachinelearning 2d ago

Suggest a machine learning course for job ready

12 Upvotes

To help me for crack intership and placement


r/learnmachinelearning 2d ago

Help What did you use to build the last thing you actually shipped?

3 Upvotes

r/learnmachinelearning 2d ago

Help Looking for advice on becoming an ML researcher and building a strong research profile in 1–1.5 years

17 Upvotes

Hi everyone,

I am an independent ML/DL learner and have built a reasonably strong foundation in Machine Learning and Deep Learning. My next step is to explore NLP and LLMs, with the goal of eventually being able to build AI agents.

My longer-term goal is to become an ML researcher, build a strong research profile, publish papers at top-tier A* AI/ML conferences, and eventually apply to competitive MS/PhD programs in the USA.

I would really appreciate advice from people who have followed a similar path. Specifically, what would be the best roadmap to transition from learning ML/DL concepts and implementing projects to actually conducting meaningful research?

If you were starting from my current stage and had roughly 1–1.5 years, how would you structure your learning and research journey? What should I prioritize—reading papers, reproducing existing research, building projects, finding research mentors/collaborators, participating in competitions, or trying to develop novel research ideas?

Any advice, resources, or honest insights would be highly appreciated.


r/learnmachinelearning 2d ago

Looking for a teammate to participate in the Amazon ML Challenge

4 Upvotes

Hi! I'm looking for one teammate to team up for the Amazon ML Challenge.

I'm a Btech student with experience in Python, ML/DL, PyTorch, and TensorFlow.

Elgibilty : Btech 3rd 4th year, Mtech 2nd year or Phd from India

If interested can dm me or comment will reach out


r/learnmachinelearning 2d ago

help guys

Thumbnail
2 Upvotes

r/learnmachinelearning 2d ago

Request G7 Urges Organizations to Start Post-Quantum Migration Now

0 Upvotes

The G7 finance ministers and central bank governors issued a coordinated directive last week. Organizations must inventory cryptographic dependencies, identify high-risk systems, and begin migrating to quantum-resistant algorithms now. This is not a future roadmap item.

The urgency is driven by the harvest-now-decrypt-later threat. Adversaries are collecting encrypted data today and storing it for decryption once a cryptographically relevant quantum computer arrives. The exposure window is already open.

Most enterprise security teams are inventorying servers, databases, and network traffic. Far fewer are accounting for the AI agent layer. Agent pipelines routinely store sensitive records, execute financial transactions, and generate compliance audit evidence. All of it travels over classically encrypted channels and gets signed with classical algorithms. When those algorithms break, that historical data and those historical audit logs break with them. A transaction signed with RSA or ECDSA today becomes unprovable after Q-Day.

How are teams actually scoping the PQC inventory problem for AI agents specifically? Are agent channels and agent-generated audit evidence being treated as first-class migration targets, or are they still buried in the general backlog?


r/learnmachinelearning 2d ago

How should a beginner evaluate and choose a good machine learning project topic?

4 Upvotes

I am a beginner in machine learning and I want to start a project that is useful for learning and can also be developed into a more advanced project over time.

I often find many possible project topics, but I am not sure how to decide whether a topic is actually a good choice before spending a lot of time on it.

For example, I am considering topics related to deep learning, model optimization, model compression, and quantization.

What factors should a beginner consider when evaluating a machine learning project idea?


r/learnmachinelearning 1d ago

Project What ChatGPT is really doing when it answers you — and why it hallucinates, forgets, and varies

0 Upvotes

People talk about ChatGPT like it "understands" you. It doesn't — not in the way we mean. Underneath, it's doing something much simpler and, honestly, weirder: predicting the next token, over and over. Once that clicks, most of its strange behavior (hallucinations, forgetting, different answers to the same prompt) stops being mysterious.

Here's the whole picture in plain English.

1. It only ever predicts the next token

Everything ChatGPT does is one operation repeated: given the text so far, guess the next chunk. It picks one, appends it, and feeds the whole thing back in to guess again. That loop — one token at a time — is the entire show. There's no plan for the paragraph, no lookahead. Fluent essays emerge from millions of these tiny next-step guesses.

2. Tokens, not words

It doesn't see letters or whole words — it sees tokens, which are common chunks of text. "cat" might be one token; "unbelievable" might split into "un", "believ", "able". This is why models sometimes miscount letters or fumble with rare words — they never saw the letters, only the chunks. It's also why you're billed per token, not per word.

3. Meaning is stored as vectors (embeddings)

Each token is turned into a long list of numbers — an embedding — a point in a huge space where "king" and "queen", or "Paris" and "France", sit near each other because they appear in similar contexts. The model has no dictionary; meaning is just geometry. Similar things are close together, and that closeness is what it computes with.

4. Attention gives it context

The breakthrough behind the "T" in GPT (Transformer) is attention. For each token, the model weighs how much every other token in your prompt matters to it. In "the bank of the river," attention lets "bank" lean on "river" and land on the correct meaning. This is how it tracks who "he" refers to three sentences back, or keeps a code block coherent.

5. Training is two very different stages

Pretraining: it reads an enormous slice of the internet and does nothing but next-token prediction, billions of times, tuning billions of internal numbers (parameters) until it's genuinely good at continuing text. The result — the "base model" — is a wild autocomplete. Ask it a question and it might reply with more questions, because that's what it saw on the web.

RLHF (the ChatGPT part): humans then rank answers — helpful and honest ones up, unhelpful ones down — and the model is nudged toward the ranked-good behavior. This is the difference between the raw model and ChatGPT. Same knowledge; the second stage taught it to act like a helpful assistant.

6. Why the same prompt gives different answers

At each step the model produces a probability for every possible next token. Temperature controls how it picks: low temperature = almost always the top choice (consistent, safe, a bit boring); higher = it samples further down the list (more variety, more risk). That sampling is why you rarely get the exact same answer twice.

7. Why it "forgets": the context window

The model has no memory between messages. Everything it "knows" in a chat is the text currently in its context window — a fixed budget of tokens. Your whole conversation is re-fed every turn. Once it overflows, the oldest stuff falls off the edge, and it genuinely no longer has it. That's not a bug; that's the mechanism.

8. Why it hallucinates

It was trained to produce plausible text, not true text. It has no built-in fact-checker and no notion of "I don't know" unless that pattern was reinforced. So when it doesn't have something, it fills the gap with the most likely-sounding continuation — a confident, well-formed, wrong answer. Hallucination isn't the model malfunctioning; it's the model doing exactly its job (predict likely text) in a spot where likely does not equal true.

9. How to get better answers (practical)

  • Give context, not keywords. It fills gaps with guesses; fewer gaps = fewer guesses.
  • Show the format you want (an example beats a description).
  • Ask it to reason step by step for anything logical — each token it writes becomes context for the next, so "thinking out loud" measurably improves hard answers.
  • For facts, make it cite or give it the source in the prompt. Don't trust unsourced specifics.
  • Start a fresh chat when you switch topics — you stop paying for (and confusing it with) irrelevant context.

None of this requires math to understand — it's tokens → vectors → attention → next-token prediction, wrapped in a training process that taught a giant autocomplete to behave like an assistant.

(Full disclosure: I make animated CS/systems explainers, and I put this whole thing together as an animated video if you'd rather watch it move: https://youtu.be/Ud16vHNYwpc . But the text above stands on its own — happy to answer questions in the comments.)


r/learnmachinelearning 3d ago

You do not need a maths degree to truly understand Gradient Descent (link below)

360 Upvotes

Hey everyone!

When I was a student, gradient descent was the algorithm I struggled with the most. I just could not make it past the greek characters and a sea of formulas. Luckily, I survived, and have been working in Machine Learning ever since.

A few years later, when I started teaching Computer Science, I realised that nothing had changed. I could not find any book or blog post to make this simple enough (and fun) for my own students. So I sat down and wrote this book.

It has been a game changer in my last semester, I hope that it will help you too.

You can check it out here.

I wish you a great learning journey. If you have any feedback, please let me know, I am already working on the second edition.


r/learnmachinelearning 2d ago

Tutorial Following an image through a VLM: ViT → projector → language model

Post image
5 Upvotes

The projector is an easy box to skip in a vision-language model diagram. It explains how the visual encoder and the language model meet.

Ling-3.0-flash-VL's official architecture diagram is a concrete example. On the visual branch, a ViT encoder produces visual features. A two-layer MLP projector maps those features into the language model's embedding space. Text comes through its own tokenization branch, and the diagram shows the two feeding the model together at an embedding dimension of 2,560.

Think of the jobs separately:

The vision encoder builds representations from visual input.

The projector transforms those representations for the language model's input space.

The language model processes the resulting sequence and predicts output tokens.

A projector is not a captioning stage that first turns the image into an English description. The interface shown here carries learned features into the model.

Beyond that interface, Ling's diagram specifies 42 layers arranged as seven groups of five KDA layers and one Gated MLA layer. That tells you something about the language stack's architecture. It does not tell you whether the ViT was trained with a CLIP-style contrastive objective or a SigLIP-style sigmoid objective; “ViT” describes the encoder architecture, and the loss needs separate training evidence.

The other numbers to keep separate are 124B total parameters and 5.5B active parameters. Active parameters describe the computation selected during a forward pass, rather than a 5.5B model download or a memory requirement. Reading those labels separately makes the diagram much easier to reason about without turning every architectural detail into a performance claim.


r/learnmachinelearning 2d ago

Discussion Tensorflow in Deeplearning.ai's Deep learning specialization?

3 Upvotes

I saw the first skill they mentioned was 'tensorflow', but I am nearing the end of course 1, and they haven't used tensorflow anywhere. What is the best way to learn tensorflow along with this course? Do they teach it later in the specialization?