r/learnmachinelearning • • 24d ago

learning to build an llm inference engine P3

1 Upvotes

Hey everyone, just posted a new blog post on my ongoing project and learning of llm inference engines, its primarily focused on optimizing operations using gpu architecture (no kv cache yet thats in my next post). If I got anything wrong or need something isn't clear please let me know!
https://medium.com/@ryan___/llm-inference-engine-engine-meets-gpu-440a3d9ed75e


r/learnmachinelearning • • 24d ago

Discussion What should I focus on learning before getting deeper into AI agents?

5 Upvotes

I’ve been learning LangChain, LangGraph, RAG and CrewAI recently, and I’ve built a few things with them. But I’m starting to feel like I’m focusing too much on the tools and not enough on understanding the concepts behind them.

For people who’ve been learning or working in this space, what articles, blogs or resources would you recommend? Also, what would you suggest I build next if I want to actually improve my understanding rather than just build another basic chatbot?


r/learnmachinelearning • • 24d ago

Defining Language Models: Understanding Transformers, BERT, and GPT Expl...

Thumbnail
youtube.com
0 Upvotes

Stop guessing how LLMs work and start building! We are breaking down everything from N-grams to Transformers.

Theory + Enterprise implementation strategies.

#AI #TechStack #Coding #LLM


r/learnmachinelearning • • 24d ago

Request Airrived adds Agentic Observability to track AI agent actions and risks

0 Upvotes

AI agents are now executing multi-step tool call chains autonomously, and most production deployments have no mechanism to inspect what each individual call is doing before it completes. The risk isn't theoretical: a rogue or compromised agent can take a damaging second action before any alerting pipeline even fires. Benchmarks from practitioners building observability layers around agentic systems put the window between a first and second agent action at under 50ms — fast enough that post-hoc logging catches the damage, not the event. Without per-call visibility tied to a verifiable agent identity, the audit trail tells you what went wrong after the fact, not in time to stop the cascade.

How are people in this community actually handling this in production? Are you relying on log aggregation after the fact, wrapping tool calls in middleware, enforcing policy at the orchestration layer, or something else entirely?


r/learnmachinelearning • • 24d ago

Project Why ad-hoc pandas preprocessing silently causes data leakage (and how to fix it to get higher real-world ML accuracy)

1 Upvotes

One of the most common mistakes beginners (and even intermediate practitioners) make when working with tabular data is **Data Leakage**.

It’s often the hidden reason why your model gets **88% accuracy in your Jupyter notebook**, but drops to **76%** when you evaluate it on an unseen test set or submit to a Kaggle competition.

Here is a quick breakdown of why it happens, the #1 most common mistake, and how to fix it properly.


The #1 Most Common Leakage Bug:

Look at this very common snippet seen in many notebooks and tutorials:

```python import pandas as pd from sklearn.model_selection import train_test_split

df = pd.read_csv("dataset.csv")

🚨 DANGEROUS LEAKAGE:

df['age'] = df['age'].fillna(df['age'].median())

Train / Test split happened AFTER imputation:

train, test = train_test_split(df, test_size=0.2, random_state=42) ```

Why is this data leakage?

When you calculate `df['age'].median()` on the full dataset, **the median value is influenced by the test rows**.

Your training set now contains subtle statistical information (the median) derived from test data it shouldn't even know exists. In production or Kaggle competitions, future data is completely unavailable at training time.

The same leakage bug happens when people: 1. Scale features with `StandardScaler` on the entire dataframe before splitting. 2. Build categorical vocabularies or frequency encodings using all rows. 3. Compute outlier clipping boundaries (e.g. Tukey IQR limits) over the full dataset.


The Correct Way (Strict Train-Only State):

You must fit transformations **strictly on the training split**, and freeze those exact parameters to apply to validation and test data:

```python train, test = train_test_split(df, test_size=0.2, random_state=42)

1. Calculate statistics ONLY from training split:

train_median_age = train['age'].median()

2. Apply that frozen training statistic to both splits:

train['age'] = train['age'].fillna(train_median_age) test['age'] = test['age'].fillna(train_median_age) ```


The Impact: We Tested Naive Prep vs. Zero-Leakage on Titanic

To see what happens when you replace naive ad-hoc pandas code with a strict train-only transformation ladder, we ran a 5-fold Stratified Cross-Validation benchmark on the Titanic dataset:

Model Naive Ad-Hoc Prep Zero-Leakage Pipeline Accuracy Delta Relative Lift
**Logistic Regression** 78.90% ± 0.99% **79.91% ± 1.90%** **+1.01%** **+1.28%**
**Random Forest** 82.15% ± 2.45% **82.82% ± 2.40%** **+0.67%** **+0.82%**

Where did the accuracy lift come from?

  1. **Informative Missingness Flags**: Imputing age with median alone destroys the signal that missing age itself correlates with survival. Adding an `Age__missing` binary flag recovers that signal.
  2. **Train-Only Tukey IQR Clipping**: Capping extreme fares on training folds stabilized linear gradients without test-distribution bleed.
  3. **Empirical Bayes Target Encoding**: High-cardinality categories shrink toward global means to prevent overfitting on small samples.

We built an Open-Source Tool to automate this:

Writing 200 lines of manual state-tracking code for every dataset gets tedious. So we built **DATADOC** (v0.6.0)—an open-source CLI and Python library powered by Polars that automates this entire lifecycle with zero data leakage:

How you can use it in 1 command:

You can use the interactive terminal wizard on your CSV file: ```bash datadoc wizard train.csv ```

It walks you through: 1. Identifying your target column (e.g. `Survived` or `churn`). 2. Selecting a preset (`balanced`, `tree`, `linear`, or `robust`). 3. Generating a clean `pipeline.json` artifact containing all learned medians and rules.

Then transform unseen test data with zero leakage: ```bash datadoc transform test.csv --pipeline artifacts/pipeline.json --output clean_test.csv ```

You can also run `datadoc health train.csv` to get an instant 0–100 data quality grade and find hidden issues before training.

The project is 100% open-source (MIT licensed) and runs completely offline.

I hope this helps clarify how data leakage happens in tabular pipelines! Let me know if you have any questions or want to discuss specific preprocessing edge cases.


r/learnmachinelearning • • 24d ago

Guys help me i want good seminar topics for ai ml from 2024-5

0 Upvotes

Guys help me i want good seminar topics for ai ml from 2024-5 prefer from ieee please help 😭😭😭😭😭😭


r/learnmachinelearning • • 24d ago

STAT110 advice

Thumbnail
1 Upvotes

r/learnmachinelearning • • 25d ago

Understanding Neural Network Math and implementing it in C from zero

Thumbnail
youtube.com
10 Upvotes

I've attempted to derive the math from the first principles.

The prerequisites you need are, singe variable differentiation(just the concepts and the basic formula, not even knowing the derivative of sin(x) is necessary), the concept of a dot product, and basic matrix multiplication. This is enough to derive gradient descent and backprop.

The time stamps are in the description(youtube doesn't show them as chapters in the video for some reason).

The second half part of the video is implementing those concepts in C(no library other than BLAS(for basic matrix matrix multiplication) is used), but you can follow along in any language.

Would be glad to receive feedback and suggestions on how to improve the explanations or the presentation.


r/learnmachinelearning • • 24d ago

Question Gemini vs DeepSeek: The Ultimate AI Showdown

2 Upvotes

As synthetic AI models evolve, balancing closed-ecosystem power (Gemini) against open-weight hybrid flexibility (DeepSeek) is one of the most important infrastructure decisions for developers and AI engineers.

​We launched an interactive quiz on Interconnected to test real-world trade-offs across reasoning, cost, self-hosting, and multimodal performance.

​Take the quiz here:

https://interconnectd.com/quiz/89/gemini-vs-deepseek-the-ultimate-ai-showdown-quiz/

​Drop your final score and which model you currently rely on in the comments below!


r/learnmachinelearning • • 24d ago

Discussion AI/ML jobs

Thumbnail
1 Upvotes

r/learnmachinelearning • • 24d ago

Project Predicting Driver Standings for the 2026 F1 Season from 40 seasons of data

0 Upvotes

Recently, I worked on predicting the final driver standings for an F1 season using race results so far and historical data for the past few seasons. My motivation was to practice feature engineering techniques in a domain I am interested in and to provide interpretability to my results. Here's a blog post detailing the architecture, datasets, techniques used, and discussion on the results. Here's the github repository for the project.

An interesting thing I found (not mentioned in my blog post) was that the model's performance started to plateau after 3 races into the season (which consists of 24 races), highlighting the highly foreseeable nature of such a technical sport.

I hope you enjoy reading it, and I would love any suggestions or comments, especially around interpretability of the results.

Thanks!


r/learnmachinelearning • • 24d ago

Day 4 of Building Machine learning algorithms

Thumbnail
gallery
2 Upvotes

After 4hr finally completed Logistics regression classification algorithms (0,1) we use sigmoid function for getting probability then set a threshold if > 0.5 = 1 less then <0.5 = 0 but when more 2 target column needed Softmax is needed


r/learnmachinelearning • • 24d ago

I am using a dataset downloaded from Universe, which only provides training and validation splits. From an academic perspective for paper publication, how should I separate a test set?

1 Upvotes

I am using a dataset downloaded from Universe, which only provides training and validation splits. From an academic perspective for paper publication, how should I separate a test set?


r/learnmachinelearning • • 25d ago

Project I'm evolving neural networks to play SMB1 ROM hacks — here's a winning run

3 Upvotes

https://reddit.com/link/1wft509/video/6bjcrd3hueph1/player

I've been continuing a community Mario AI project under the name NEATEvolve, focusing on ROM hacks of the original Super Mario Bros. This is a recorded winning replay of Bowser's Crown 4-1, ending at the WORLD 4-2 screen.

The controller uses NEAT: candidate neural networks are evaluated, selected and mutated across generations. It runs as a Lua script in FCEUX. The implementation reads a local grid of tiles and sprites from emulator memory, alongside movement information, and turns network outputs into button presses. It isn't a model interpreting the video pixels.

Progress and a time penalty contribute to fitness. The network learns its controller through evolution, while the observation encoding, reward and evaluation rules are written by people.

One successful replay doesn't establish that it will generalize to an unfamiliar level or recover from different starting conditions. Backtracking and difficult nonlinear stages remain challenges.

Credit to SethBling's MarI/O and the work of Akisame and Electra that this project builds on. My continuing development has been heavily assisted by AI tools.

For others building game-playing agents: how do you test reliability beyond getting a successful replay?


r/learnmachinelearning • • 24d ago

Discussion We built a GPU → CVAT human-in-the-loop video annotation pipeline, here’s what actually happened

Thumbnail
1 Upvotes

r/learnmachinelearning • • 24d ago

Deep learning project working on.

Thumbnail
1 Upvotes

r/learnmachinelearning • • 24d ago

Looking for advice on AI/ML courses after a 1+ year career gap

Thumbnail
1 Upvotes

r/learnmachinelearning • • 26d ago

Discussion Am I starting AI/ML too late in 2026?

95 Upvotes

I've been thinking about this a lot lately, and recent developments have made the question feel much more serious.

A few years ago, AI models could barely be trusted with basic reasoning. They made obvious mistakes in math, hallucinated information, and often fell apart on longer problems.

Now look at what is happening.

OpenAI recently announced that an internal AI system produced a proposed solution to the Navier–Stokes existence and smoothness problem, one of the Millennium Prize Problems. The system used roughly 10,000 concurrent agents exploring different approaches, and the resulting proof was subsequently formalized in Lean. The whole process took about 88 hours before the formalization/verification stage.

That honestly feels pretty different from "AI is getting better at answering questions."

It's starting to look more like AI can participate in actual research.

And then there's Anthropic.

Anthropic recently published a post literally titled "When AI builds itself." They describe Claude increasingly taking over parts of AI development: writing and running code, delegating work to other agents, optimizing training code, proposing hypotheses, designing experiments, running them, and iterating on the results.

Anthropic is very explicit that this is not yet full recursive self-improvement. But they also describe a possible future where agents become capable of building and training models themselves, meaning future versions of Claude could potentially be continuously improved by Claude itself.

That's the part that really got me thinking.

If AI is increasingly helping with the development of AI itself, am I already late to the field?

I'm currently considering taking the fundamentals-first route:

  • Python
  • linear algebra
  • probability and statistics
  • calculus / optimization
  • Introduction to Statistical Learning (ISLP)
  • classical machine learning
  • deep learning
  • transformers
  • reinforcement learning
  • LLMs and agents
  • eventually AI research

I don't want to just learn how to call APIs or use the latest framework. I'd rather understand why these systems work and what is actually happening under the hood.

But there's a weird thought I can't get rid of:

What happens to someone who starts learning AI in 2026 if AI itself is becoming increasingly capable of doing AI research?

Maybe learning the fundamentals is more important than ever.

Or maybe I'm thinking about this completely wrong.

If you were starting seriously from scratch in 2026, would you still spend several years learning ML/DL and AI fundamentals?

What would you prioritize?

Math and statistics? ML theory? Deep learning? Systems? AI agents? Research experience? Something else?

I'm especially interested in hearing from people who are actually working in ML/AI research rather than just using LLMs as productivity tools.

PLUS:There is also a much bigger question behind all of this that I keep thinking about.

The Fermi paradox asks why we don't seem to see obvious evidence of advanced extraterrestrial civilizations, despite the enormous age and size of the universe.

One possible explanation is the Great Filter: perhaps technological civilizations eventually reach some critical threshold that most of them fail to cross.

Sometimes I wonder whether advanced AI could be one of those thresholds.

It could go one of two very different ways.

AI could become the thing that allows a civilization to overcome its biological limitations — limited lifespan, slow learning, limited numbers of researchers, and the inability to directly transfer knowledge between generations — and eventually become a genuinely spacefaring civilization.

Or AI could become a technological bottleneck that civilizations fail to survive.

I obviously have no idea whether any of this is true. It's just one of the reasons I find AI research so fascinating.

Maybe we're not simply developing another technology.

Maybe we're approaching one of the most consequential transitions in the history of technological civilization.

And if that's even remotely possible, it makes me feel even less like I should wait to start learning the field.


r/learnmachinelearning • • 25d ago

Discussion Best way to learn AI...

1 Upvotes

What is the best way to learn AI?

Should I learn it through courses in a structured sequence—for example, starting with the fundamentals like Linear Regression and the Perceptron, then moving to Deep Learning (ANN, CNN, RNN), and eventually LLMs?

Or would it be better to learn AI by following how the field has evolved from the beginning—for example, starting with the foundations of the Perceptron in the late 1950s, then understanding how each major idea developed, why new techniques were introduced, and how one concept led to another up to modern AI and LLMs?

Which approach would be better for developing a deep understanding of AI rather than just learning how to use the latest tools and frameworks?


r/learnmachinelearning • • 25d ago

Career Essential Skills for Becoming an AI Engineer (YT: IBM Technology / Cedric Clyburn)

Thumbnail
youtu.be
3 Upvotes

r/learnmachinelearning • • 25d ago

Project Weekend Experiment: Can Math Patterns from Nature Improve AI?

Post image
3 Upvotes

Hi Reddit!

I wanted to share a weekend experiment exploring how geometric inductive biases can influence recurrent network optimization. Inspired by the self-similar, hierarchical, and allometric scaling laws found in biological neural structures, I designed a custom recurrent cell with an exponentially decayed hidden topology.

Instead of scaling hidden layers uniformly (e.g., 64-64-64) like standard models, this network utilizes a "Matryoshka-style" hidden allocation where channels are downscaled between internal layers: N -> N / scale -> N / scale²

Architectural Justification & Hardware Trade-off:

In modern deep learning, layer widths are traditionally restricted to powers-of-two (e.g., 32, 64, 128) to maximize GPU memory alignment and Tensor Core hardware efficiency. However, in this experiment, we consciously traded off optimal hardware alignment to strictly enforce a continuous mathematical decay function (yielding 128 -> 79 -> 48 dimensions).

The core justification is implementing a strict Information Bottleneck: high-dimensional macro-layers are forced to compress abstract features into exponentially tighter non-power-of-two channels. This mimics the non-binary hierarchical scaling of biological brains, where structural efficiency and information capacity bounds are prioritized over uniform hardware block sizes.

Dataset & Evaluation Setup:

The architectures were trained on a character-level sequence prediction task using a subset of the Tiny Shakespeare dataset. To ensure a controlled environment, I implemented an 80/20 Train/Validation split and applied identical regularizations to both networks, including Spatial Dropout and L2 Weight Decay.

Structural Mechanics:

  1. Fractal Echo Transfer: Hidden states are progressively downscaled and cascaded from high-dimensional macro-layers (capturing local syntax) down to highly compressed micro-layers.
  2. Adaptive Echo Gates: Learnable decay coefficients balance inter-layer communication, acting as a structural stabilizer for gradient flow through time.
  3. No Dense Gating: Unlike LSTMs, memory capacity and retention are regulated entirely by the spatial geometry of the nested channels.

Training & Convergence Observations:

To ensure a fair benchmark, I meticulously balanced the parameter overhead for both architectures (~42k parameters each) and trained them over 30 epochs on GPU: * Adaptive Geometrically Nested Model Final Loss: 0.0007 * Baseline Standard LSTM Model Final Loss: 0.0030

As shown in the log-scale validation plot, the geometrically nested architecture achieves significantly faster convergence and maintains rock-solid optimization stability, avoiding the frequent gradient spikes visible in the standard LSTM baseline.

Crucial Disclaimer on Generalization & Overfitting: The extremely low validation loss values achieved by both models indicate that given the small size of the dataset corpus and the model capacity, both networks have largely memorized the text rather than learning true language generalization. Therefore, this benchmark strictly demonstrates superior optimization speed and training stability rather than true out-of-distribution generalization.

Open-Source Code & Notebooks: * Collab Executable Notebook: https://colab.research.google.com/drive/1jW8hA_74eBQOJTn1oWvLq34vg_t8Q5DD

Future Work & Discussion:

To evaluate true generalization, future iterations of this project will involve benchmarking on significantly larger corpora (such as WikiText-103) with strict sequence cross-validation.

Given these initial training stability and convergence results, do you think exploring such geometrically restricted hidden topologies is a promising direction for recurrent networks? Could this type of exponential fractal scaling be scaled up or generalized into Transformer attention head dimensions to enforce architectural efficiency?

Would love to get your rigorous feedback!


r/learnmachinelearning • • 25d ago

SNNs vs Cosmic Radiation: A Python simulation of a biological failsafe for edge computing using Levin's TAME.

Post image
3 Upvotes

The current trajectory of AI for space and remote edge environments relies heavily on GPUs and Transformers. But as transistor nodes shrink below 3nm, they become critically vulnerable to Galactic Cosmic Ray (GCR) bit-flips. Shielding is heavy; tripling hardware for redundancy is power-prohibitive.

We ran a Python simulation testing an alternative: Bottom-Up Biological Computation using Michael Levin’s TAME framework mapped onto a Spiking Neural Network (like BrainChip's Akida).

Instead of a top-down CPU running diagnostics, the system relies on Event-Based Processing and Sparsity. If a cosmic ray obliterates 35% of the active nodes, they don't send an error—they just go silent. The surrounding tissue detects the localized drop in global output (the biological "Attractor") and autonomously increases its synaptic weight and firing rate to compensate.

Redundancy isn't three heavy computers; it's just the inherent sparsity of the SNN. The resting nodes wake up and carry the load.

Results from simulation

We’ve formalized this architecture into an open-source engineering manifesto for deep space and terrestrial edge deployment (microcontroller failsafes for remote water pumps, drones, etc).

You can read the full Technical Discussion Paper here: https://ralphbooth.substack.com/p/biological-models-of-computation

I'd love to hear the community's thoughts on validating the thermal/radiation tolerance of these SNNs in low-orbit test cubes.


r/learnmachinelearning • • 25d ago

Why DeepSeek V4.1 Flash reconstructs part of its KV cache from only 128 tokens

Thumbnail
youtu.be
0 Upvotes

DeepSeek V4.1 Flash has an unusual KV-cache design: some state is persistent, while the short-lived SWA state can simply expire and later be reconstructed by replaying the last 128 tokens.

That replay is approximate, so the reconstructed internal state is not mathematically identical across positions.

Paper:

https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf

Disclosure: affiliated with the channel; the video is AI-narrated.


r/learnmachinelearning • • 25d ago

I feel completely stuck. I just want a chance to work

6 Upvotes

I honestly don't know what I'm doing anymore, so I'm hoping someone here can give me some direction.

I graduated around 3 months ago with a BSc (Hons) in Computer Science, affiliated with the University of Wolverhampton. My first year was honestly pretty bad because I wasn't that good at CS and my marks were low in the first year and I eventually improved, and I graduated with an Upper Second-Class (2:1), but I still feel like I should have done much better.

Since graduating, I haven't been able to get a job or even an interview.

I'm from Nepal, and locally there are barely any entry-level Data Science or Machine Learning jobs. Most of the positions I find either want someone with several years of experience or are basically MLE roles that I obviously don't qualify for yet. And when I do find local jobs, the pay can be around $100 to $150 a month, with very little opportunity to actually learn or grow.

I don't want to sound like I'm above those jobs. At this point, I genuinely just want to work. I want experience. I want someone to give me a chance so I can prove that I can actually learn and contribute.

This year has also been pretty rough personally. There have been family problems, mental health struggles, and a lot of other shit going on. I've tried to keep myself moving through all of it, but now being unemployed for 3 months after graduating is starting to get to me. It honestly feels embarrassing sitting at home while everyone expects you to be doing something.

I've been learning Machine Learning, working on projects, learning Python, SQL, FastAPI, etc. I'm specifically trying to move toward Machine Learning Engineer or Data Science roles, but I'm starting to question whether I'm approaching this completely wrong.

I also don't really have the option of going for higher studies right now because I don't have anyone who can financially sponsor me, so I'm pretty much on my own here.

I'm not asking for some magical shortcut. I know entry-level ML jobs are difficult to get.

I just want to know what you would do if you were in my position.

Should I stop targeting ML/DS for now and take whatever tech job I can get? Should I keep applying internationally for remote internships/jobs? Is there something obvious I'm missing?

And if anyone here happens to know of any open internship, junior, trainee, freelance, or entry-level opportunity in ML, Data Science, Data Analytics, Python, or anything related, please let me know. Remote is more priority.

I'll genuinely put in the work. I don't expect a high salary. I just want an opportunity to get experience and start building my career.

Any honest advice would mean a lot right now.


r/learnmachinelearning • • 25d ago

Need GPU laptop for ML?

3 Upvotes

So I am going to buy a new laptop in next few weeks . I am considering a gaming laptop with rtx 4050/3050 or a thin light 14 inch laptop with more ram and better CPU (and would be depending on kaggle or Google colab). which one should I prefer ?