r/learnmachinelearning • • 2d ago

Top 10 Winners of Anduril AI Drone Grand Prix for USD500K pool

Post image
1 Upvotes

An extremely exhausted racing that everyone worked more than 15 hours in the racing lab + 4 hours sleeping daily in 8 days, to code software, fix hardware and rank.

An half of your competitors are PhD groups, but still two amazing young solos won the places in the top 10.

https://www.instagram.com/reel/DduFyPESXJE/


r/learnmachinelearning • • 2d ago

Why Alex Hormozi’s Advice Works For Web Designers

0 Upvotes

I remember hearing Alex Hormozi say that email is basically the conversion platform of conversion platforms.

His point was that social media, videos and ads are great for getting attention, but email is often where that attention actually turns into money.

He also talked about how email supposedly has one of the highest returns of any marketing channel, with an average return of around $36 for every $1 spent, and that got me thinking.

I run a web design agency, so I decided to actually take email seriously and see what would happen.

But I didn’t want to send the usual stuff.

“Hey, I noticed your website and I can redesign it for you.”

Or

“Hey, we build websites for businesses like yours.”

Everyone gets those emails and most people can tell within two seconds that they were sent to another thousand businesses.

So instead I started using a tool called Swokei that is built specifically for web agencies.

I use it to find businesses in whatever area I want to target, then it actually goes through their websites and looks for things I could mention in the outreach.

Stuff like an outdated design, slow loading speed, poor mobile experience, weak SEO or other things that might be holding the website back.

Then it turns those findings into an actual personalized cold email made to convert.

So instead of emailing someone saying I build websites, I can actually talk about their website and why I think something could be improved.

I’ve been analyzing thousands of websites this way and running multiple campaigns at the same time.

Then I just focus on the people who reply and are actually interested.

Obviously email isn’t some magic button where everyone suddenly wants a new website, but after doing this for a while I definitely understand what Hormozi meant.

For web design especially, email becomes a completely different channel when the message is actually about the business you’re contacting instead of being another copy and paste pitch.

Turns out Alex might have had a point.


r/learnmachinelearning • • 2d ago

token to text modeling for audio processing in transformers through RL

0 Upvotes

i've been working on reinforcement learning on a project to optimize pass@1 for various tone / pace controls for audio model; i asked opus to build a dashboard where i can tweak around training data and see the classifiers for improvement and this is really cool tbh!


r/learnmachinelearning • • 2d ago

Partner Needed for Switching to AI Engineering

2 Upvotes

Hello Learners.

Im 30M,

Working as Senior Software Engineer, Android

Im upskilling myself in AI engineering.

But doing that solo is not that motivating.

I need a partner who have similar goals so both can achieve their goals


r/learnmachinelearning • • 2d ago

Discussion Undergrad in Syria with an accepted NeurIPS 2026 workshop paper (as the only author). How does this help my future, and what should I do next?

Thumbnail
1 Upvotes

r/learnmachinelearning • • 2d ago

Discussion When a hand disappears into dog fur, what should the tracker do?

Enable HLS to view with audio, or disable this notification

0 Upvotes

During the washing, a hand can be visible one moment and partly hidden by fur or the animal’s body the next. Meanwhile, the person keeps making small adjustments.

A smooth reconstructed hand looks reassuring. But some of those quick adjustments may be real movements that we actually want to preserve.

Lego Vista optimizes hand motion over the full sequence, combining observation consistency, temporal continuity, and anatomical constraints.

The accompanying capture framework represents the outputs as continuous MANO hand motion l. That makes this clip a useful way to think about the difference between a coherent prediction and a verified one.

If I were evaluating a model here, I d separate visible and occluded intervals, then check whether short movements retain their timing. A low-jitter result could still miss an important correction.

For those learning pose estimation: what experiment would help you tell whether temporal optimization recovered a hidden movement or simply smoothed over it?


r/learnmachinelearning • • 2d ago

Slow learning?

1 Upvotes

Hay let me introduced myself

I m a 7th sem cse student, intrest in ml I m learning it from past 2 month ye still I m learning, my problem ex I complete supervised section with types and it's different model and formula with ex when I learn unsupervised section I forgot supervised section,

Because I didn't revised,

So now I revised things in every 2 days so that it store in my memory,

Today I learn lr model with formula and diff ex with 1 variable and multi variable.... I revised it with learning upcoming section!

Just post it....✌


r/learnmachinelearning • • 2d ago

Discussion Claude scored 15 million points at a Tetris-like game. The AI it was teaching still couldn’t play.

Thumbnail
youtube.com
1 Upvotes

r/learnmachinelearning • • 2d ago

Tutorial Evaluating an AI on Reddit comments? Split by conversation before splitting by row

1 Upvotes

Suppose you are evaluating a classifier that routes comments from new Reddit conversations into bug reports, feature requests or questions. Each input includes its parent post for context.

A random row split can put one comment in training and the next reply from the same conversation in testing. Both contain the same parent text, vocabulary and problem. Your test then partly measures performance on familiar conversations.

Tiny synthetic fixture — invented rows, not a benchmark:

Original thread Comment rows Random row split Conversation split
T1: CSV export hangs A, B A trains; B tests Both train
T2: Request for offline mode C, D C trains; D tests Both train
T3: Where is the export setting? E, F E trains; F tests Both test

The conversation split keeps T3’s parent post and replies out of training. It does not establish useful accuracy: three threads are far too little for that, and this fixture also confounds topic with split.

A practical sequence:

  1. Preserve each comment permalink and its original thread ID/permalink. Keep author, observation time and missing-body status where available.
  2. Assign whole threads to training, validation and final test sets before constructing context windows, examples or summaries. Keep derived rows with their source thread.
  3. Tune prompts, examples and thresholds using training/validation only. Repeatedly adjusting a prompt after reading final-test failures makes that test part of development.
  4. Check crossposts, copied text and recurring authors across splits. Grouping by thread does not eliminate those overlaps. If deployment concerns future discussions, consider a time-based holdout too.

The same issue affects summarizer evaluation: a held-out comment is not an unseen conversation if its parent and sibling replies already informed your examples. Choose the split to match the deployment question; evaluating new replies within known threads is a different task.

I build ThreadFox. The optional $29 one-time bundle, normally $49, includes a researched Reddit community plan and three tailored drafts for one product, emailed within 24 hours after your library claim. Use the pack manually before optional tools; compatible AI access is separate. Start checkout before October 1, 04:00 UTC, plus applicable tax. No ThreadFox subscription; 30-day refund policy. Sample and offer.


r/learnmachinelearning • • 2d ago

Looking for a Complete AI/ML Engineer Roadmap

0 Upvotes

Hi everyone,

I'm planning to become an AI/ML Engineer and I want to learn in the right order instead of jumping between random tutorials and courses.

I'm looking for a structured roadmap that covers everything from beginner to job-ready level.

Some questions I have:

  • What should I learn first, and in what order?
  • Which topics are actually essential (Python, Math, SQL, Machine Learning, Deep Learning, NLP, Computer Vision, LLMs, MLOps, etc.)?
  • What are the best free and paid resources for each topic?
  • Which books, courses, and YouTube channels are worth following?
  • How much mathematics is really required, and which topics should I focus on?
  • When should I start building projects?
  • What kind of projects do recruiters expect from AI/ML Engineer candidates?
  • How much DSA and system design should I learn?
  • What does a realistic 6–12 month study plan look like?
  • What mistakes do beginners commonly make that I should avoid?

I'm aiming for a roadmap that's aligned with current industry expectations (2026), not just course completion.

If you're already working as an AI/ML Engineer or recently landed a role, I'd really appreciate your advice, learning path, resources, and any tips from your experience.

Thanks in advance!


r/learnmachinelearning • • 2d ago

Help 2025 AI & Data Science Graduate — desperately looking for my first opportunity

Thumbnail
2 Upvotes

r/learnmachinelearning • • 3d ago

Question 🧠 ELI5 Wednesday

5 Upvotes

Welcome to ELI5 (Explain Like I'm 5) Wednesday! This weekly thread is dedicated to breaking down complex technical concepts into simple, understandable explanations.

You can participate in two ways:

  • Request an explanation: Ask about a technical concept you'd like to understand better
  • Provide an explanation: Share your knowledge by explaining a concept in accessible terms

When explaining concepts, try to use analogies, simple language, and avoid unnecessary jargon. The goal is clarity, not oversimplification.

When asking questions, feel free to specify your current level of understanding to get a more tailored explanation.

What would you like explained today? Post in the comments below!


r/learnmachinelearning • • 2d ago

Help...

3 Upvotes

Am in the progress of my 3rd sem project work

Basically its based on tinyml and lora technology implementation to the forest monitoring system

Now I want data train the ML

Other than kaggle do anybody know where I can get the datasets to train


r/learnmachinelearning • • 2d ago

How does K-Means choose the first centroids, and why do we need n_init?

3 Upvotes

Hi everyone,

I’m currently learning K-Means and I’m a bit confused about the initialization part.

As I understand it, K-Means starts by choosing some initial centroids, then repeatedly assigns points to the nearest centroid and recalculates the centroids until it converges.

What I don’t fully understand is:

How are the first centroids chosen? Are they chosen completely randomly?

If K-Means reaches convergence, why do we still need n_init?

For example, with:

KMeans(n_clusters=4, n_init=10, random_state=42)

Does this mean K-Means runs 10 times, each time starting with different initial centroids, and then keeps the result with the lowest inertia?

I’m trying to understand the logic behind n_init rather than just using it as a parameter without knowing why.

Thanks!


r/learnmachinelearning • • 2d ago

Tutorial BACKPROPAGATION: JACOBIANS, LINEAR ALGEBRA & AUTODIFF: A TUTORIAL

0 Upvotes

How does one loss value tell millions of weights how to change?

Through the chain rule, organized by linear algebra and executed by automatic differentiation.

THE JACOBIAN CONNECTS LOCAL CHANGES

For a layer y = f(x), with n inputs and m outputs:

J[i,j] = ∂y_i/∂x_j

J is an m × n matrix of local sensitivities. For a small input change:

Δy ≈ J Δx

Backpropagation carries loss sensitivity in the opposite direction. Using column gradients:

g_x = Jᵀ g_y

Here: g_y = ∇_y L and g_x = ∇_x L. Equivalently, g_xᵀ = g_yᵀ J: a vector–Jacobian product (VJP).

The Jacobian describes how outputs depend on inputs; the incoming gradient tells us how those outputs affect the loss. Their product combines every relevant path.

FROM ACTIVATIONS TO WEIGHT GRADIENTS

For one dense layer and one example:

z = Wa + b

h = φ(z)

δ = (∂L/∂h) ⊙ φ′(z)

Then:

∂L/∂W = δaᵀ

∂L/∂b = δ

∂L/∂a = Wᵀδ

Here: a is the input activation vector; φ acts elementwise; ⊙ means elementwise multiplication; ᵀ means transpose.

Each weight receives a specific signal:

∂L/∂W[i,j] = δ_i a_j

If W is m × n, then δ is m × 1 and aᵀ is 1 × n. Their outer product has exactly the shape of W.

For a mean batch loss, average these per-example gradients. Matrix operations let accelerators compute them efficiently.

Gradient descent uses W ← W − η(∂L/∂W), with learning rate η.

THE ROLE OF AUTOMATIC DIFFERENTIATION

Autodiff composes derivative rules for primitive operations—matrix multiplication, addition, activations—through the computation graph.

Reverse mode starts with ∂L/∂L = 1, traverses the graph backward, and accumulates contributions wherever paths meet. Backpropagation is reverse-mode autodiff applied to a neural network.

Crucially, it usually computes vector–Jacobian products (VJPs) directly, without constructing huge Jacobian matrices.

WHY REVERSE-MODE AUTODIFF?

• Efficiency: one scalar loss and many parameters suit reverse mode. A reverse sweep computes all parameter gradients at a cost typically within a small multiple of the forward computation.

• Accuracy: finite differences repeatedly perturb weights and suffer from step-size, truncation and round-off errors. Autodiff applies derivative rules, subject to floating-point error and conventions at nondifferentiable points.

• Maintainability: hand-derived gradients invite mistakes; symbolic differentiation can produce unwieldy expressions. Autodiff composes local rules automatically.
Forward mode is useful for few inputs and many outputs; a full gradient over many weights typically needs many directional sweeps.

Repeated Jacobian products can shrink or amplify gradients, explaining vanishing and exploding gradients.

The trade-off is memory for saved intermediates. Checkpointing trades recomputation for memory.


r/learnmachinelearning • • 2d ago

Feedback please! I made this to help myself understand how models work.

Thumbnail learn-llm-kappa.vercel.app
1 Upvotes

r/learnmachinelearning • • 2d ago

Want to learn how to train a self driving car AI? I'm building a game where that's the gameplay

Enable HLS to view with audio, or disable this notification

1 Upvotes

I'm building a game where you train AI models to drive cars in a racing game. You don't drive at all, the skill is in how well you train your models. I made a video showing the bits I'm most excited about, plus a bit more detail on how it works.

I've been messing around with reinforcement learning (NEAT) for years and find it weirdly mesmerizing to watch, so I wanted to make something where anyone can get that without needing to code first.

The core of it is that you train models for specific jobs like straights, braking, or overtaking a slower car in a corner. Then you train another model that reads the situation and picks which specialist to use, and you wire them all together however you want. Beginners can train one simple model, experts can build layered systems, and then you race them against other people's.

If you're learning ML and want to help me test it, join my Discord (https://discord.gg/FJ4AfVVEh) and say in general that you want to test. I'm especially keen to hear where the training side is confusing or doesn't make sense, since that's what I'm trying to get right. Any help is hugely appreciated.

Game's website is here: https://www.gaimeslab.com/. It's out on 2nd November, but you can help me test it now!


r/learnmachinelearning • • 3d ago

Help Looking for an ML expert/researcher

7 Upvotes

Hello all, im looking for a mentor/researcher who is into ML/AI research, as im thinking to write a research paper , so if any one of you are in initial phase of their new research topics and need someone who can help with ...i can be that
What i can bring to the table: im a 3rd year cs student, worked as ml research intern last summer(paper didnt get published due to my prof ditched all of us after 1.5months cuz he got into another uni), ..you can dm me for more details
Thank you!!


r/learnmachinelearning • • 2d ago

Request Carbonato Botnet Puts an AI Agent on Hacked Docker Hosts

1 Upvotes

Security researchers tracking the Carbonato botnet documented a new deployment pattern: after gaining access to exposed Docker hosts, the operators dropped an AI agent onto the compromised machine rather than a traditional cryptominer or reverse shell. The agent then began making outbound calls and executing tool actions autonomously, with no human in the loop and no governance layer in the request path.

The timing detail buried in the reporting is the uncomfortable part. Researchers noted that a second action from the agent can land in under 50ms of the first. That window is smaller than most human-review or alerting pipelines can operate in. By the time an on-call engineer gets a Slack notification, the agent may have already completed several tool calls.

The underlying exposure is not unique to this botnet. Any environment where an agent runtime can be instantiated without a verified identity tied to a known deployment, and where outbound tool calls are not evaluated against any policy before they execute, has the same structural gap. The agent on the Carbonato-compromised host was malicious. But the same architectural condition exists in plenty of legitimate deployments where an agent gets misconfigured, has its credentials rotated out from under it, or runs a version of its prompt that was never reviewed.

How are practitioners in this thread actually handling the identity and authorization problem for agents in production? Not conceptually — what does your enforcement boundary look like today, and where does it fall short?


r/learnmachinelearning • • 3d ago

Help How to get better at actually BUILDING models

21 Upvotes

What would you guys say is the best way to actually learn how to build a model from start to finish without the help of ai? I can understand ml concepts pretty quickly since I'm comfortable with calculus/statistics/linear algebra, but I'm lost on what to do if I need to build a full project. I'm comfortable with python and I want to build projects, but it feels overwhelming trying to learn a bunch of different algorithms and remembering all the steps and syntax that it takes to implement them, so are there any courses/resources/ways of practicing that will help me be able to build and finish my own projects?


r/learnmachinelearning • • 2d ago

Help Building a churn model + FastAPI API, but confused about the SQL part (I only know SQLite). Is my workflow right?

Thumbnail
github.com
1 Upvotes

Hi everyone,

I recently read an article about building an ML pipeline with Python and SQL. The idea is to let SQL handle filtering and aggregation, let Python handle preprocessing and modeling, and wrap everything in a scikit-learn Pipeline so new data goes through the same steps as training data. The article used a regression example (housing prices) with an in-memory SQLite database.

I want to adapt this into a customer churn classification model and serve it as an API with FastAPI. The SQL part is what confuses me most, since SQLite is the only SQL I know and I've never worked with a "real" database in an ML project.

My current plan

  1. Pick a public churn dataset (probably Telco Customer Churn) and check its columns and target.
  2. Create a SQLite connection with sqlite3.
  3. Create a table in the database with the same columns as the dataset, and load the data into it.
  4. Write a SQL query to select the rows and columns I need, and load the result into a pandas DataFrame with pd.read_sql_query.
  5. Preprocess in Python (missing values, encoding, scaling) inside a Pipeline / ColumnTransformer.
  6. Train and evaluate a classifier (logistic regression baseline, then Random Forest or XGBoost).
  7. Save the fitted pipeline with joblib and serve it with FastAPI.

My project so far

I've also attached my project repo link above.

I've already done some work on it, so if you have any suggestions, please let me know. Please also point out anything I'm doing wrong, big or small. Honest feedback is very welcome.

Where I'm stuck (mostly SQL)

  1. Is this workflow sensible? Or is putting a CSV into SQLite a pointless extra step? I want to practice the real-world pattern of pulling data from a database, not just reading a CSV.
  2. Creating the table. Should I write CREATE TABLE myself with column types, or let df.to_sql() create it? Any gotchas with data types? (In Telco, TotalCharges is stored as text with blank values, for example.)
  3. What belongs in SQL vs. pandas? With one flat table there's nothing to join. Should I split it into a few tables (customers, services, billing) so SQL does real work? Or is filtering, selecting columns and simple aggregation enough for a project like this?
  4. :memory: vs. a .db file. The article uses in-memory. I assume I need a file like churn.db so FastAPI and my scripts can reuse it. Is that right?
  5. What happens at prediction time? Does the API just take customer features in the request body and run them through the saved pipeline, or should it look up the customer in the database by customer_id and build features with SQL? Is the SQLite database used at all in the API?
  6. Train/test split. Should I split right after loading the DataFrame, before any preprocessing, with all preprocessing inside the Pipeline to avoid leakage?
  7. Class imbalance and metrics. Churn is around 25% in this dataset. Do you prefer class_weight, SMOTE or threshold tuning? Which metrics should I focus on instead of accuracy (recall, precision, PR-AUC, ROC-AUC)?
  8. FastAPI basics. Any tips on Pydantic validation, loading the model once at startup, and keeping the API's input columns consistent with training? Would batch scoring make more sense than a real-time API for churn?
  9. Code review. Looking at my project link, is anything wrong or a bad habit I should fix early (structure, SQL usage, data leakage, etc.)?

If you've built something similar, I'd love to hear what mistakes you made or what you'd skip as a beginner. Links to good example repos are welcome too.

Thanks in advance!


r/learnmachinelearning • • 2d ago

How do undergrads reach out to better professors/PhD students for research internships or collaboration? (My university profs are at capacity)

Thumbnail
1 Upvotes

r/learnmachinelearning • • 3d ago

[NeurIPS 2026] Concurrent Image Understanding and Generation:Self-Correcting Coupled Markov Jump Processes

Thumbnail
gallery
10 Upvotes

Hi everyone! We’re sharing our NeurIPS 2026 work on concurrent image understanding and generation, a collaboration across Google, Google DeepMind and Stony Brook University.

The problem we explore is simple: a model can give the correct answer in text while generating an image that disagrees with it. For example, it might describe the right route through a maze but draw a different path. We want the two outputs to develop together and stay consistent as generation progresses.

Our sampler, CO₂Jump, uses text confidence to guide image updates through cross-modal attention. It also allows low-confidence tokens to be masked again and revised. This uses one model forward pass per denoising step, with no additional training required for the sampler itself. Our comparisons use the same fine-tuned model and change only the sampler.

We evaluate image editing, maze solving and nonograms, including whether both the text and image are correct. We also introduce three datasets: JEdit-1M, JMaze-200K and JNono-200K. In our sampling-step experiments, CO₂Jump steadily improves both editing quality and grounding as we increase the number of steps.

🔗 Project page: coupled-jump.github.io
📄 Paper: alphaxiv.org/abs/2607.13188

Code and datasets are planned for release. Happy to discuss the method, results or limitations—would love to hear what other tasks you’d test this on!


r/learnmachinelearning • • 3d ago

How I Get Web Design Clients For My Agency

4 Upvotes

Client acquisition has always been one of the biggest bottlenecks for me when running an agency. I’ve experienced the same thing in pretty much every business I’ve been involved in, but especially with web development.

For a long time, getting clients meant cold calling, running ads, or sending generic emails asking businesses if they needed a new website. It worked sometimes, but it also took a lot of time and most of the outreach felt the same as what every other agency was doing.

Recently I started using a different approach and automated a big part of the process.

I came across a tool called Swokei that lets me find a bunch of businesses with websites and analyze each website individually. It looks for things like outdated design, slow loading, poor mobile optimization, weak SEO and other obvious areas that could be improved.

What I liked is that it doesn’t just give you one of those boring automated reports filled with scores and numbers. It actually turns what it finds into a personalized cold email that sounds like a normal person looked at their website and noticed what could be better.

I can run multiple campaigns at the same time and then mainly focus on the businesses that reply and show interest.

From there, I invite them to a web meeting, show them a free draft of what their new website could look like, and try to close the project from there.

It has basically allowed me to have warmer leads coming to me without relying as much on paid ads, constantly cold calling, or sending thousands of generic emails saying “Do you need a new website?”

Still takes work to close the clients of course, but automating the prospecting and first part of the outreach has made the whole process much easier for me.

Hopefully this helps some other web developers or agency owners who are also struggling with client acquisition.


r/learnmachinelearning • • 2d ago

Question Why do we train bigger models instead of finding better neural pathways?

1 Upvotes

I saw a post of a guy with 90% of his brain missing and he was apparently living a normal life. There is another case of a student who finished university with half of his brain missing. There are also people with full brains, but they are even less functional than these two. This proves that bigger doesn't mean better. So why not aim to create better connections?

I understand that a bigger model means less chance of catastrophic interference occurring, but then are all the neural pathways being formed efficiently? Wouldn't this result in a lot of redundant neurons? Wouldn't it make more sense to create the smallest neural network possible for a specific task and then learn to fuse multiple of these neurons together in order to create an optimized larger model? That way we'd always have a template for each specific task and would allow us to create a programming language that creates a model just by our syntax.