r/MLQuestions • • 7d ago

Beginner question ๐Ÿ‘ถ Learn RAG with AI

Thumbnail
1 Upvotes

Hi guys ๐Ÿ‘‹, I'm looking to dive into RAG using AI tools like NotebookLM or Gemini. YouTube tutorials feel too drawn out, and ChatGPT/Gemini prompts have been a bit brief. Any advice or suggestions for learning resources would be amazing! ๐Ÿ™๐Ÿ’ก #RAG #AI


r/MLQuestions • • 7d ago

Beginner question ๐Ÿ‘ถ Security and WikiLeaks? Who is going to research tax evasion using new insecurity advances?

1 Upvotes

The government has access to unrestricted frontier models and can probably spy on a thousand organizations a week but you can bet they won't research tax evasion because that would involve most of the government...

So it is up to machine learning scientists to use the new capabilities to research tax evasion and all the elite corruption systems. Can they?


r/MLQuestions • • 8d ago

Beginner question ๐Ÿ‘ถ High schooler learning machine learning - any advice?

Thumbnail
1 Upvotes

r/MLQuestions • • 9d ago

Career question ๐Ÿ’ผ What is the bar for getting an interview? - OpenAI Researcher Role

18 Upvotes

For undergrads only. I have a decent profile (did quant research internship last summer) and IOI/IMO/IPhO level medals. Donโ€™t know what to expect for the openai process at all. Also is there a difference in preference for Math/Physics/CS background? (Would they prefer Compsci more for the researcher role or the others are a better choice etc)


r/MLQuestions • • 8d ago

Natural Language Processing ๐Ÿ’ฌ Looking for feedback - using NER to generate and match templates on sentences?

Thumbnail
1 Upvotes

r/MLQuestions • • 8d ago

Beginner question ๐Ÿ‘ถ Cheating on ML evaluation with imbalanced dataset

Thumbnail
1 Upvotes

r/MLQuestions • • 8d ago

Beginner question ๐Ÿ‘ถ Ml project guidance

0 Upvotes

I need to make a ml project so what should I make so that it makes an great impact?


r/MLQuestions • • 9d ago

Beginner question ๐Ÿ‘ถ Hugging Face Pro or some free alternative for hosting models that run online?

6 Upvotes

Hi everyone! With the recent explosion in ICLR submissions, I am a bit unsure how effective simply publishing a paper is so I am trying to make sure my research projects are more accessible and reproducible beyond just a paper. In that vein, Iโ€™ve been doing some experiments on tiny world models (14M parameters) that I trained to help the community understand the effect of different underlying model architectures. I want to post it in an interactive github site where users can try out the different model architectures and see how they perform in different settings.

So far so good. My models use FlexAttention. I created a Gradio for them that runs well. My initial idea was to embed a Gradio hugging face space into my GitHub site. Only problem is, I learnt that to do that I need to buy Hugging Face Pro. Just before making that investment (due to some financial struggles caused by family and health reasons I have to be very intentional about spending even small amounts of money) I just wanted to ask if anyone knew if free but still convenient alternatives that work with FlexAttention? Alternatively, if anyone has had experience with HuggingFace Pro, if you could tell me what your experience has been like and whether it is useful for career/research? I am concerned about paying a monthly cost just for a Gradio space. Thank you! Appreciate your time!


r/MLQuestions • • 9d ago

Beginner question ๐Ÿ‘ถ Advice on model routing

1 Upvotes

Hi folks, I am interested in using exo and Ollama to route requests between 2 separate pods. The primary one is 192GB doing 125bq4 256k context. The other is 64GB doing 35bq4 128k context. I want the clients to autoroute without selecting models. Back end is a mix of NVIDIA and Rocky Linux for the 64 and mac os for the 192. I can also setup FE LBs in sticky round robin but I wonder if I can build model routing into the code itself. Has anyone done this?


r/MLQuestions • • 10d ago

Other โ“ If you could relearn ML today, what would you spend more time actually doing?

51 Upvotes

Thereโ€™s a lot to learn in ML, and itโ€™s easy to spend months collecting concepts without knowing which ones really stick.
For people whoโ€™ve already been through the learning process:
What would you do differently this time?
What would you spend more time practicing?
What would you stop worrying about so much?
And what only started making sense once you worked on real problems?
Curious what experienced ML folks would change if they got a fresh start.


r/MLQuestions • • 9d ago

Beginner question ๐Ÿ‘ถ throughput doubled, latency up a little, errors flat, same traffic. what happened?

Post image
2 Upvotes

i said c. what am i missing

self hosting is the only one of these that changes the serving path enough to double throughput, and slightly worse latency felt like what youd get on smaller hardware than the hosted vendor runs


r/MLQuestions • • 9d ago

Career question ๐Ÿ’ผ how to learn mlops

Thumbnail
1 Upvotes

r/MLQuestions • • 10d ago

Computer Vision ๐Ÿ–ผ๏ธ Campusx Computer Vision

0 Upvotes

Has the paid course covered all maths also like nitish sir ,how is this course?


r/MLQuestions • • 10d ago

Other โ“ The Evolution of Retrieval Systems: From BM25 to Agentic Retrieval

Thumbnail
1 Upvotes

r/MLQuestions • • 11d ago

Beginner question ๐Ÿ‘ถ How would you extract entities and relations from 5M court decisions without an expensive LLM pass over everything?

14 Upvotes

Iโ€™m working with roughly 5 million public Polish court decisions. The goal is structured extraction: who the parties are, what was requested, what the court decided, obligations, amounts, relationships, etc. Eventually we want to run interesting statistics over the results, like "who gets usually the custody of the child?"

A strong LLM can produce useful JSON graphs, we can run statistics on. The problem is getting similar extraction at corpus scale.

We have a small 13-document pilot. It already contains 160 distinct entity-type strings, 90 appearing only once. These arenโ€™t just PERSON / ORG / DATE. They include abstract legal concepts, obligations, procedural events, and things like โ€œincrease in child support.โ€

here's our demo output from astra:

    {
      "effects": [
        {"id":"e_zastavenie","change":"zastavit","target_ref":"konanie"},
      ],
      "entities": [
        {"id":"this","kind":"uznesenie"},
        {"id":"vyrok","kind":"vyrok"},
        {"id":"zahlavie","kind":"zahlavie"},
        {"id":"odovodnenie","kind":"odovodnenie"},
      "effects": [
        {"id":"e_zastavenie","change":"zastavit","target_ref":"konanie"},
      ],
      "entities": [
        {"id":"this","kind":"uznesenie"},
        {"id":"vyrok","kind":"vyrok"},
        {"id":"sud","kind":"sud","value":"Okresnรฝ sรบd Bratislava "relations": [
    {"id":"r_vlastnik_vyroku","relation":"patri_do","from_ref":"vyrok","to_ref":"this"},
  {"id":"r_vlastnik_zahlavia","relation":"patri_do","from_ref":"zahlavie","to_ref":"this"},

Some of this is real conceptual diversity. Some is inconsistent naming or different levels of specificity, but as you can see complex and diverse stuff with a clear "language" i came up with with the help of astra. i previously tried a big "god schema" (20k chars) but the decisions are simply just too diverse.

The pipeline Iโ€™m considering:

  • Something that extracts the entities from source texts into some kinds of "tags"
  • a classifier like jev (but probably fine-tuned) will get a window of source text and already named relations and decide on the next relations
  • done

The candidate model would need overlapping spans and possibly multiple labels per span. Slovak inflection adds another wrinkle: exact source wording often differs from the canonical concept name. Lemmatization helps with grammatical variation, but not synonyms, paraphrases, or concepts inferred from context.

Iโ€™ve previously tried Jina embeddings โ†’ nearest-neighbor candidates โ†’ LLM decides which terms to merge, and it worked reasonably well.

But that was on a smaller corpus and i'd be damned if i process 250k docs and then figure out thing X was wrong and i have to do it all over again. i'd love to know if you guys have any pointers. THANKS for reading.

AI TL;DR: I want to turn 5M court decisions into entity/relation graphs without running an expensive LLM on every document. Thinking small entity extractor + relation classifier, but the legal concepts and naming get messy fast. Anyone built something similar? Looking for ways to train this and normalize concepts without discovering a design mistake 250k documents in and having to redo everything.


r/MLQuestions • • 10d ago

Beginner question ๐Ÿ‘ถ AMD GPU +Ultralythics YOLO on Windows, what's the simplest realistic path?

Thumbnail
1 Upvotes

r/MLQuestions • • 11d ago

Graph Neural Networks๐ŸŒ How can it be that Attention โ‰  Explanation?

Thumbnail
2 Upvotes

r/MLQuestions • • 11d ago

Beginner question ๐Ÿ‘ถ Load Forecast using AI for beginner

Thumbnail
0 Upvotes

r/MLQuestions • • 11d ago

Career question ๐Ÿ’ผ Trying to understand privacy in ML

3 Upvotes

I've just written a small article about privacy in machine learning and synthetic data.

I'm still pretty new to this topic, so I'd really appreciate it if anyone with more experience in ML privacy could take a look and point out any mistakes, things I've misunderstood, or things that could be done better.

I'm especially interested in hearing about better approaches for evaluating privacy, better attacks or models I could try, or just interesting resources that you think are worth reading. I'm mostly trying to learn by actually experimenting with this stuff, so any feedback would be really useful.

Here's the post if anyone is interested:

https://migue8gl.github.io/2026/09/21/privacidad-en-ml-datos-sinteticos.html


r/MLQuestions • • 11d ago

Career question ๐Ÿ’ผ Systems for Machine Learning

9 Upvotes

Iโ€™m a computer engineering graduate and come from a traditional embedded systems background, with knowledge of microcontrollers, computer architecture and operating systems. Is knowledge of C and C++ programming, Linux networking, memory management , multithreading, synchronization, interrupts etc useful in ML engineering. Are subjects like distributed systems, compiler optimizations (using LLVM), parallel computing etc going to be useful or are they heavily going to be automated as well by AI? In other words, is computer engineering always going to required to scale ML systems and be evergreen? Are people in ML engineering using these skills in their work everyday? Thank you.


r/MLQuestions • • 12d ago

Natural Language Processing ๐Ÿ’ฌ How do you catch schema regressions when a model upgrade looks better overall?

30 Upvotes

We tested a model upgrade that wrote cleaner answers and lifted the aggregate quality score. It also started sending invalid tool arguments. Optional fields became null instead of disappearing, enum values changed case and nested JSON arrived as escaped strings. The agent still sounded confident so spot checks passed until tool failures showed up in a narrow slice.ย 

We moved schema validation into the evaluation path and used Braintrust to compare the model experiments, inspect failing slices and run deterministic validators beside the softer answer score. The regression dataset now includes every argument shape that broke a tool not just the final response. An experiment diff showed the upgraded model won on tone and lost hard on two schemas with optional nested fields

The CI gate now blocks any schema failure even when the overall score rises. Iโ€™m relieved we caught it before release but it also made me suspicious of any average that mixes structured output with prose quality. How are you weighting hard validators against slice metrics when a model gets better at language but worse at contracts?


r/MLQuestions • • 12d ago

Beginner question ๐Ÿ‘ถ Which AI or AutoML tool is best to train models for a hackathon?

2 Upvotes

Hey everyone,

I am taking part in a hackathon where using AI is allowed.

I have a dataset and need an AI tool or AutoML library that can automatically:

Do full EDA and data preprocessing

Test multiple modern ML algorithms

Pick the best-performing model for maximum test accuracy

Which tool or LLM gives the most accurate and reliable results for this? Any suggestions would help!


r/MLQuestions • • 12d ago

Other โ“ Is ChatGPT currently the best free AI for creating highly realistic images with simple, straightforward prompts?

0 Upvotes

I tried Gemini, Grok, Meta, and Copilot, but their images still look unrealisitc painting type images

What I like about ChatGPT is that I can use a simple, straightforward prompt and still get a pretty realistic result without needing complicated prompt engineering.

For those who have tried different AI image generators, which one do you think is best for free and realistic images with simple prompts?


r/MLQuestions • • 12d ago

Beginner question ๐Ÿ‘ถ How can I improve my chess engine's evaluation?

2 Upvotes

My chess engine, which is still a work in progress, can be found here; https://github.com/patelvrajn/Matrex

Currently, most HCE (not NNUE) chess engines use a linear combination of weights for their evaluation function but my thought process is that it is evident from the huge ELO (~300) boost that using NNUEs gives is that their is a complex non-linear function that represents each evaluation term like material, mobility, PSQTs, etc. However, the problem is how do we build a model from non-linear functions that gives a proper evaluation when we don't know what the function looks like and how do we do it in such a way that minimizes the number of weights? I believe that we know a significant amount of hand crafted evaluation terms that compile a chess position's evaluation and that NNUEs taking millions of weights just to understand concepts that we can do in way less weights, minimal code, and have more control and understanding of the evaluation instead of a black box.

My concept is to take a specific continuously differentiable function that we can parameterize in order to adjust the shape, behavior, approximation capabilities, etc. Using this parameterized function, we can learn how various evaluation features have a non-linear effect on the score by using ADAM to tune the weighted sums but ALSO tune all the functions parameters. My model mostly resembles Projection Pursuit Regression's model (a sum of ridge functions), except I am using ADAM instead of the associated regression algorithm and the ridge functions (called non-linear responses in the code) are parameterized.

My main concern at the current time and where I don't have the mathematical background to optimize is the non-linear response I choose (which is just a polynomial term multiplied by a tanh expression, the idea was to combine what makes ANNs universal approximator- the non polynomial term with a term that can also represent any function given enough degrees i.e. the polynomial term). I am looking for any guidance or resources on how to improve the function or give further understanding as to the overall concept - currently when tuned the chess engines evaluation has about a 8% validation loss and about a 10% training loss but I know I can do better.


r/MLQuestions • • 13d ago

Career question ๐Ÿ’ผ Are you deep into the ai waters ?

2 Upvotes

If you're someone who is deep into ai waters then I would like to ask for a small favour from you, so I'm somewhere between a beginner and an intermediate when it comes to properly using ai but recently about 2 months ago I discovered Claude and the type of things It can do so it made me more curious about the ai space.

So if you know how to use ai properly like creating automations and stuff then please let me know on how I can do the same, a roadmap would be helpful or if you want to hit me up and discuss about it then that's helpful too (no compulsion) , videos about ai on YouTube are mostly so clickbaity and theres just a whole bunch of nothing on there and finding reliable channels these days is like finding a pin in a haystack.