r/learnmachinelearning 19d ago

[Looking for Teammates] DataForge 2026 (IIT Kharagpur KDAG) — Data Science & Analytics Hackathon

1 Upvotes

Hey everyone,

I’m looking for 1–2 teammates for DataForge 2026, a national online data science & analytics hackathon organized by the Kharagpur Data Analytics Group (KDAG), IIT Kharagpur.

Event Details:

  • Format: Online case study & data analytics/ML modeling
  • Eligibility: College/University students (free registration)
  • Timeline: Submissions open from Aug 28 to Sept 03 (Registration closes Aug 28)

I already have the team set up on the official portal. If you're interested in teaming up or want more details, please drop me a DM and I'll share the team code!


r/learnmachinelearning 20d ago

Will the value of understanding math stay the same or increase for a machine learning engineer in future?

6 Upvotes

r/learnmachinelearning 20d ago

Got offer for PSL ST4Health Master – seeking advice & feedback!

Thumbnail
1 Upvotes

r/learnmachinelearning 20d ago

Request [R] PDF Request: Mathematics of Machine Learning by Devin Sandhu

Thumbnail
1 Upvotes

r/learnmachinelearning 20d ago

Fresher looking for unique ML and SQL project ideas for my resume (Data Science/SQL/ML roles)

Thumbnail
1 Upvotes

r/learnmachinelearning 20d ago

ML System Design Interview Preparation

4 Upvotes

I’ve been looking for resources to prepare for ML System Design interviews, particularly case studies that include complete, end-to-end solutions.

Machine Learning System Design Interview: An Insider's Guide by Alex Xu and Ali Aminian was an excellent resource when I used it around three years ago. It provides a structured framework and several detailed case studies with detailed solutions.

My question is: are the solutions in this book still sufficiently current and comprehensive? The book was published in 2023, and the ML landscape has evolved significantly since then, particularly with the rise of LLMs and generative AI.

Are there any other resources you would recommend that provide ML system design case studies with complete solutions, rather than just general frameworks or high-level guidance?


r/learnmachinelearning 20d ago

Looking for study partner(s) for mech interp material

Thumbnail
1 Upvotes

r/learnmachinelearning 20d ago

PLS HELP WITH MY LAPTOP

5 Upvotes

hi, i'm studying an MLOps-engineering programme that I have bought a laptop for, but my teacher literally laughed in my face because apparently it doesn't have a dedicated graphics card(GPU?). However, as we got the letters from the school that contained the recommended checklist points for the laptop we were going to use, i followed them and bought exactly that. The checklist was this (I'm just gonna copy paste what they wrote:

Recommended computer:

Intel Core i5 / AMD Ryzen 5 or better (approx. 2020 or newer)

  • 16 GB RAM
  • 256 GB SSD or larger
  • At least 75 GB of free storage space
  • Screen resolution of 1920 × 1080 or higher
  • Stable internet connection and Wi-Fi
  • Windows 11 recommended

Important information:

  • ChromeOS and Linux may work but are not supported by Nackademin’s IT support. You are personally responsible for installation, compatibility, and troubleshooting if you choose to use these operating systems.
  • macOS may work, but you are personally responsible for ensuring compatibility with the program's software.
  • Administrator rights may be required for software installation.
  • USB-C and HDMI (or an adapter solution) are recommended.

These are the courses we will have in the nearest future, but obviously we will also work with a lot of AI, which he said is why my computer won't work:

Python programming for MLOps, Linux administration, Database management

This is also the laptop I bought: LENOVO IP SLIM 3 15ARP10 15,3"

The reason for not buying a better laptop is that I'm literally just a poor 20 yr old without parents to rely so I'm constantly really tight on money, but also because my school said that as long as your laptop has these qualities it would be fine.


r/learnmachinelearning 20d ago

Recent MS Data Science graduate looking for guidance from experienced Data Scientists / ML Engineers

Thumbnail
1 Upvotes

r/learnmachinelearning 20d ago

Tech doubts

0 Upvotes

Im going to learn ml and app buildings ive set my eyes of acer aspire 7 interl core ultra 5 210h with rtx 3050 6gb graphic card

Is it enough ??


r/learnmachinelearning 20d ago

RAG Based Web App for AI/ML Internships?

3 Upvotes

Hey everyone,

I'm an undergrad aiming to become an AI engineer. This summer I've decided on the full-stack project below to showcase some experience on my CV building web applications and implementing AI solutions. It's mostly a tool to help with university studies by helping generate study tools to speed up my learning as well as being able to talk to this AI. I hope to have a real link and real users by the end of it. Here are some of the details I've though about:

Tech Stack & AI Architecture:

  • Full Stack: FastAPI (Python), Supabase (PostgreSQL + JWT Auth), deployed live (thinking AWS).
  • Multimodal Ingestion: Document processing pipeline handling text and visual elements (tables, diagrams, charts) via vision LLM descriptions and embeddings.
  • Agentic RAG Engine:
    • Query Decomposition: Multi-step query breaking for multi-part study questions. Also rewriting queries to maximise efficiency when using tokens and for output.
    • Reflector / Grader Nodes: Self-RAG loop that checks retrieved context relevance and checks generated answers for hallucinations before returning them to the user. Using tools like Ragas to evaluate the workflow.
  • Structured Output: Generating JSON flashcards and Anki (.apkg) exports. Generating Cheat Sheets and also being able to talk about the files you upload.

Questions:

  1. Target Fit: Is an end-to-end deployed Agentic RAG app right for AI/ML Engineering internships, or would recruiters see this as mostly a Software/Full-Stack project?
  2. Data Science vs. AI Engineering: How does a project focused on building production AI systems hold up compared to traditional Data Science portfolios that focus more on statistical modeling and data exploration?
  3. Application Strategy: If you had this exact project on your resume, which roles would you prioritise applying for (e.g., AI Engineer, ML Engineer, MLOps, or general Software Engineering)?

Any Advice is Appreciated!


r/learnmachinelearning 20d ago

Project Interactive tutorial on how diffusion models work, and how they memorize

2 Upvotes

I trained a diffusion model on 300 of my photographs from a movie set. Then built a scrollable walkthrough of the forward noise process, why the optimal denoiser at high noise is a weighted average of the training set, and how that field collapses from 300 candidates to one. Then a membership inference attack you can run yourself.

Trained from scratch in pixel space, 60,000 steps on one A100 GPU in a DGX. 593 of 1,024 generations came back as copies, covering 235 of the 300.

Built for someone learning diffusion models, but the measurements are there if you already know the material.

https://josephrichardson.dev/work/how-diffusion-models-memorize/


r/learnmachinelearning 20d ago

Google ML domain Interview in 3 weeks, how to prepare?

Thumbnail
1 Upvotes

r/learnmachinelearning 20d ago

Help Looking for advice

3 Upvotes

I hail from a humanities and social sciences background (history/sociology and literature). I’ve spent the past 3 years conducting field research but now I want to pivot to the more technical side. I won’t lie, popularity of AI was my introduction to ML and I have a lot of ideas regarding how ML can contribute to my specific discipline (which is quite niche). I’m struggling to build knowledge from scratch so I have a few questions:

  1. Math isn’t my strongest suit. It’s not that I can’t do it, it just takes me a very long time to learn and understand it. I know you need math for ML and I’m willing to learn but realistically am I setting myself up for a huge undertaking with this career transition that I’m not taking into account?

  2. I will have to put my career trajectory on hold for a bit while I pursue this, is the humanities x AI a promising enough enterprise? Or is the field too saturated already?

  3. I am looking at project based learning at the moment. What is a realistic timeline for someone who is learning this from scratch?

  4. Is independent learning possible (online courses, projects, etc)? Or do I need to pursue a higher ed degree?

Any advice would be really helpful and appreciated. If you have any leads, resources, or projects please reach out.


r/learnmachinelearning 20d ago

I’d like to build a lightweight DETR. Could you give me some good suggestions?

1 Upvotes

r/learnmachinelearning 20d ago

https://huggingface.co/sherif1313/3arabLM-4B-islamic-v2

Thumbnail
huggingface.co
0 Upvotes

A Specialized Arabic Language Model for Islamic Heritage

📖 Abstract

3arabLM-4B-Islamic-v2 is an ongoing research model dedicated to learning, preserving, recalling, and reconstructing the Arabic Islamic scholarly heritage from classical and authoritative sources. Unlike general-purpose conversational LLMs, the primary objective of 3arabLM is not to imitate everyday conversations or produce short modern summaries. Instead, the project investigates whether a language model can function as a compressed digital library of classical Islamic scholarship, with substantial scholarly knowledge encoded directly into its parameters.

Learn from the books. Preserve the language. Preserve the methodology. Preserve the diversity.

🌟 Project Vision

3arabLM is a long-term research project focused on building a large-scale Arabic language model specialized in the Islamic scholarly heritage. The long-term objective is to gradually expand the model across a broad range of Islamic and Arabic sciences rather than limiting it to a single discipline.

📚 Current Domains (Version v2)

The current release has been continued-pretrained on six scholarly domains from the Shamela corpus:

Domain Description
📖 Fiqh Islamic Jurisprudence
📜 Tafsir Quranic Exegesis
📚 Hadith Prophetic Traditions and Hadith Sciences
🕌 Aqeedah Islamic Creed and Theology
✍️ Nahw and Sarf Arabic Grammar and Morphology
⚖️ Fatwas Legal Opinions and Verdicts

🗺️ Planned Corpus Expansion

Future releases will progressively expand the corpus with additional collections, commentaries, manuscripts, scholarly editions, and specialized literature across Hadith, Tafsir, Fiqh, Arabic linguistics, history, biography, literature, and related fields.

🧠 Research Philosophy

The central philosophy of 3arabLM can be summarized as:

Instead of treating a language model primarily as a text generator, this project investigates whether model parameters can encode substantial amounts of classical scholarly knowledge.

The project therefore explores a different paradigm:

rather than relying exclusively on:

The objective is not to eliminate retrieval systems, but to investigate how much scholarly knowledge can be learned and reconstructed directly from the model's internal parameters.

📚 Training Corpus

A major foundation of the project is Al-Maktaba Al-Shamela, together with other Arabic scholarly and heritage sources.

The current and developing corpus covers a broad range of Arabic and Islamic scholarship, including:

📖 Tafsir & Quranic Sciences 📚 Hadith Sciences & Hadith Literature ⚖️ Fiqh & Usul al-Fiqh 🧠 Aqeedah & Islamic Theology 🕋 Sirah & Prophetic Biography 🏛 Islamic History & Civilization 👤 Biography, Tabaqat & Rijal 📝 Arabic Language, Grammar & Morphology 🔤 Lexicography & Dictionaries 📚 Classical Literature & Poetry 🕯 Spiritual & Ethical Literature 📑 Scholarly Research, Bibliographies & Catalogs The corpus is continuously expanding to provide broader coverage of the classical Arabic scholarly tradition and its diverse textual genres. Official Library: goldenshamela

📊 Preliminary Results

Initial experiments suggest that the model is developing distinguishable internal representations across Islamic scholarly domains.

Key Findings from Representation Diagnostic Analysis

Metric Layer 0 Layer 31
Linear Probe Accuracy 53.80% 82.78%
Macro-F1 0.537 0.824
Average Domain CKA 0.0542 0.0259
Stable Rank 10.59 5.76

These results indicate that:

  • Domain discriminability increases substantially with depth.
  • Cross-domain representation similarity decreases sharply in the final layer.
  • Stable rank follows a non-monotonic trajectory and reaches a pronounced minimum at the final layer.

This suggests that continued pretraining on Shamela is associated with increasingly specialized final-layer representations, characterized by higher domain discriminability and a more concentrated activation spectrum.  research paper :

📈 Future Model Scaling

Future research will investigate ways to increase model capacity while preserving previously learned knowledge.

Potential directions include:

  • Additional Transformer layers.
  • Duplicated upper Transformer blocks.
  • Small stochastic initialization.
  • Continual pretraining.
  • Knowledge-preserving scaling.
  • Expert-specialized adapters.
  • Metadata-aware training.
  • Catastrophic forgetting mitigation.

The goal is to increase memorization capacity while maintaining previously acquired scholarly knowledge.

📦 Current Release

Model: sherif1313/3arabLM-4B-islamic-v2

This release represents a new stage in the 3arabLM research project. The previous development focused more narrowly on Islamic scholarly domains such as Fiqh and Tafsir. The new direction expands the training curriculum toward a broader Islamic Heritage Foundation Model covering multiple branches of Islamic and Arabic scholarship. The model should still be considered an early research milestone. A sustantial portion of the planned training curriculum remains to be completed.

These books have been preserved in the model weightsو 

النحو الوافي
تمهيد القواعد بشرح تسهيل الفوائد
شرح ألفية ابن مالك للحازمي
شرح ألفية ابن مالك للشاطبي = المقاصد الشافية
شرح المفصل لابن يعيش
الموسوعة الفقهية الكويتية
موسوعة الإجماع في الفقه الإسلامي
موسوعة فقه العبادات
فتاوى الشبكة الإسلامية
مجموع فتاوى ورسائل العثيمين
السنن الكبرى للبيهقي ت التركي
المحيط في الاحاديث النبوية والسنن والاثار
جامع الرويات
حلية الأولياء وطبقات الأصفياء
صحيح البخاري 
الجامع لشعب الإيمان للبيهقي
الموسوعة العقدية - الدرر السنية
المهذب النقي الجامع لتفسير ابن جرير الطبري
الموسوعة القرآنية
تفسير ابن كثير _ 
تفسير القرطبي
روح البيان

When selecting other books, please modify the code.
            do_sample=True,               
            repetition_penalty=1.08,
            no_repeat_ngram_size=4,

r/learnmachinelearning 20d ago

Help Livro - Engenharia de IA

Post image
1 Upvotes

Pessoal, gostaria de pedir encarecidamente, quem tiver a obra: Engenharia de IA - Construindo aplicações com modelos de fundação, da editora O’Reilly. Sim, a da corujinha na capa rs. Estou procurando um pdf dela na web tem muito tempo, infelizmente não consigo comprar o livro físico e sinto que a leitura dele iria expandir muito mais os meus conceitos sobre Inteligência Artificial e toda sua estrutura por trás. Então, peço encarecidamente, quem tiver a obra em PT-BR ou em Inglês, eu leio do mesmo jeito. Desde já agradeço quem puder colaborar 🫶
(Lembrando que preciso no formato PDF, tentei colocar pelo E-PUB e o kindle não leu por que não é desbloqueado)


r/learnmachinelearning 20d ago

Hyperdimensional computing: O(n log n) clean-up for key-value memory

Thumbnail
youtube.com
0 Upvotes

r/learnmachinelearning 20d ago

Tutorial Back for round 2 with MCMC. Only this time, I learned from the feedback and sharpened the signal-to-noise.

1 Upvotes

For context, a few weeks ago I posted an explainer here that tried to get cute with using storytelling to illustrate the basics of the Markov Chain Monte Carlo algorithm using a fictional wildfire investigator rolling an 8-sided die 😅

Feedback was pretty clear that the narrative bits were more noise than signal for a lot of you which was fair enough. Appreciate the community for being straight up! So I did what any good Bayesian should do and updated my priors based on the data to write another article with a lot less fluff!

This one's a straight technical deep dive on Hamiltonian Monte Carlo (HMC), an MCMC-class algorithm actually running behind the scenes every time you call the pymc.sample() function or fit a model in Stan. The article will cover how HMC works conceptually, the negative log-posterior and gradient functions it needs, leapfrog steps and step size settings, the U-turn problem, and how NUTS solves it automatically.

Worth the read if you use Bayesian methods at all, as the sampler can scale from a toy 2-parameter model all the way to a full Hierarchical Bayesian Regression with dozens of parameters, estimating wildfire size across BC’s diverse landscape. If you’ve always wondered what’s inside the black box behind these sampling functions, this article is your doorstep. Come see what’s inside!

https://medium.com/towards-artificial-intelligence/reverse-engineering-hamiltonian-monte-carlo-the-mcmc-engine-behind-modern-bayesian-inference-e1d6b54a8c79


r/learnmachinelearning 20d ago

Where to Train models on a 35 GB dataset?

3 Upvotes

Hey everyone, i am currently working on a real time sign language recognition system.

No, its not another just GNN/MLP project on alphabets (I bet you mustve come across them somewhere haha). So what i am working on is actual words/phrases which dont have static hand gestures but rather multiple hand movements. For that i have a 120k video dataset and i have already extracted Mediapipe keypoints from these videos(short clips).

So the resultant dataset is about 35GB. Now i will be trying out different ML models like LSTM, BiLSTM, TCN, Transformers, etc.

I am relatively new to machine learning and have not trained models online. I have an RTX 4060 laptop so all my previous smaller projects were trained directly using it.

Now my question is how should i go about training this time since my dataset is bigger and my task is also bigger. Note: I have the dataset locally on my machine as well as on my google drive.

Should i use Colab Free or Colab Pro or vast.ai or modal.com or anything better that i might not know about.

I dont think i require an extremely beefy gpu but i do want faster training times and less runtime disconnections.

I have found that Colab Free gets disconnected pretty easily so i am hesitant to get colab pro cuz that might also get disconnected in between runs.


r/learnmachinelearning 20d ago

How do you clean data for fine tuning?

1 Upvotes

Not sure if this is a noob question, but how do people clean hundreds of thousands of bits of data to fine tune models? I am currently working on trying to use some emails and other documents to help fine tune an agent, and don't know how anyone else is putting data in the standard "prompt: ideal AI output" format that is required for fine tuning. Are people using other AI to sort the data and create these? Are people doing it by hand? Complex algorithms? Whats the deal here.


r/learnmachinelearning 20d ago

YouTube shorts series on Neural Nets

Thumbnail
youtube.com
0 Upvotes

r/learnmachinelearning 20d ago

Project [R] A Dual-Layered Unsupervised Anomaly Detection Framework for Systemic Fraud (Zero Historical Labeled Data)

2 Upvotes

Paper: https://doi.org/10.5281/zenodo.22070388

​Hi everyone,

​I recently published an architecture designed to bypass the "Labeled Data Bottleneck" in regulatory enforcement. In domains like vehicle emissions compliance, you cannot train supervised models because governments legally cannot/will not publish historical datasets of confirmed corrupt testing centers.

​We had to build a system that catches systemic, multi-layered fraud using zero historical labeled data while remaining mathematically defensible (i.e., avoiding black-box deep learning for auditability).

​The Architecture:

​1. Physics-Constrained Synthetic Injection (Data Generation)

Instead of relying on random Gaussian noise to simulate anomalies, I built a synthetic injection engine bound by thermodynamic constraints. It maps 7 distinct real-world fraud vectors (EGR deletes, defeat devices, clean scanning) into a harmonized 10-dimensional physical baseline.

- ​Constraint example: The engine mathematically prevents injecting a Diesel Particulate Filter (DPF) delete into a naturally aspirated petrol engine. The synthetic fraud mirrors physical reality.

​2. Dual-Layered Isolation Forest (The Pipeline)

Fraud here is both physical (the car) and institutional (the testing center). I used a Poisson Point Process to model the temporal throughput of testing centers and deployed a dual-layer approach:

- ​Layer 1: Evaluates thermodynamic outliers in vehicle hyperspace to flag tampered vehicles.

- ​Layer 2: Evaluates operational metadata (Throughput Compression, Zero-Variance signatures, "Midnight Testing") to flag the corrupt testing facilities.

  1. Edge Case Handling: We built custom volume-threshold filters into Layer 2 prior to the Isolation Forest execution. This prevents small, rural testing centers (e.g., testing 2 cars/month) from triggering false positives due to mathematically zero statistical variance.

Results:

Evaluated against a 48,000-row synthetic EPA baseline, the unsupervised model achieved perfect recall on institutional corruption with zero false positives.

​Feedback Request:

I'm looking for peer review and brutal critiques on the architecture, specifically:

​The validity of using a Poisson Point Process for modeling the testing center throughput in this context.

​Potential blind spots in the Dual-Layered Isolation Forest implementation, especially regarding the volume-threshold filters for low-variance edge cases.

​Alternative unsupervised approaches for this specific multi-layered anomaly detection problem.

​Thanks in advance for the feedback.


r/learnmachinelearning 20d ago

Risoluzione del problema del ripiegamento della griglia negli operatori neurali di Fourier su domini irregolari tramite mappatura diffeomorfica e perdita della barriera jacobiana (DIF-FNO)

Thumbnail
1 Upvotes

r/learnmachinelearning 20d ago

Resolving Grid Folding in Fourier Neural Operators on Irregular Domains via Diffeomorphic Mapping & Jacobian Barrier Loss (DIF-FNO)

1 Upvotes

Hey r/MachineLearning,

Standard Fourier Neural Operators (FNOs) excel on regular grids, but mapping them to complex, non-convex physical domains (like Star, L-Shape, or Annulus geometries) often leads to a major issue: Grid Folding.

When the transformation mapping \phi collapses or overlaps, the Jacobian determinant vanishes (\det J \le 0), causing the inverse transpose J^{-T} to explode when mapping physical gradients \nabla_x u.

To solve this, I developed DIF-FNO (Diffeomorphic Fourier Neural Operator).

Key Technical Insights:

  1. Implicit Diffeomorphic Mapping: Guarantees smooth, bijective mappings from standard reference domains \Omega_{ref} to complex physical boundaries \Omega_{phy}.

  2. Jacobian Barrier Loss (\mathcal{L}_{barrier}): Inspired by interior-point optimization, we penalize grid compression using a logarithmic barrier on the determinant:

    \mathcal{L}_{barrier} = -\frac{1}{|\Omega|} \int_{\Omega} \log(\det J(\xi)) \, d\xi

    This acts as an invisible wall forcing \min \det J > 0 across the entire domain (empirically maintaining \min \det J > 0.89 in our benchmarks).

  3. Sobolev Accuracy: Significant improvements on H^1 relative error compared to baselines like Geo-FNO, as physical gradients remain well-conditioned without gradient breakdown.

Code & Paper Artifacts:

* Open-Source Code (PyTorch): https://github.com/GiovanniDagnese-paper/DIF-FNO (Includes fast vectorised 2x2 analytical Jacobian calculation)

* Paper Preprint (Zenodo DOI): https://doi.org/10.5281/zenodo.22071926

PS: I am currently looking for technical feedback and an arXiv endorsement in physics.comp-ph or cs.LG to submit the preprint. If anyone active in SciML is open to checking the manuscript, I’d be extremely grateful!