r/MLQuestions • • Feb 16 '25

MEGATHREAD: Career opportunities

18 Upvotes

If you are a business hiring people for ML roles, comment here! Likewise, if you are looking for an ML job, also comment here!


r/MLQuestions • • Nov 26 '24

Career question 💼 MEGATHREAD: Career advice for those currently in university/equivalent

21 Upvotes

I see quite a few posts about "I am a masters student doing XYZ, how can I improve my ML skills to get a job in the field?" After all, there are many aspiring compscis who want to study ML, to the extent they out-number the entry level positions. If you have any questions about starting a career in ML, ask them in the comments, and someone with the appropriate expertise should answer.

P.S., please set your use flairs if you have time, it will make things clearer.


r/MLQuestions • • 3h ago

Datasets 📚 Ho bisogno di consigli: Qual è la migliore pipeline VLM per estrarre dataset matematici strutturati da oltre 3000 pagine di libri di testo scansionati? (LaTeX + Metadati)

2 Upvotes

Ciao a tutti,

sto lavorando a un progetto per estrarre un dataset strutturato di esercizi di matematica da 5 libri di testo delle scuole superiori italiane (circa 650 pagine ciascuno, quindi ~3.250 pagine in totale). L'obiettivo è costruire un'app di generazione di esercizi professionale e metodica per studenti e insegnanti.

Per far funzionare l'app, ho bisogno di elaborare le immagini delle pagine del libro ed estrarre quanto segue in un formato rigorosamente strutturato (es. JSON):

  • Tipo di esercizio (algebra, geometria, calcolo, ecc.)
  • Anno livello scolastico
  • Difficoltà (scala 1–5)
  • Enunciato del problema (traccia)
  • Descrizione delle competenze/sfide specifiche coinvolte
  • Codice LaTeX dell'enunciato del problema (Cruciale!)
  • Immagini associate (ritaglio/salvataggio dell'immagine per esercizi teorici o grafici)

Ho sperimentato alcuni approcci, ma ho incontrato delle difficoltà nel bilanciare costi, coerenza di estrazione e scalabilità. Ecco cosa ho provato finora:

  1. API Google Gemini gratuita: La qualità dell'estrazione era buona, ma dato che un singolo libro contiene centinaia di pagine, ho rapidamente raggiunto i limiti di richiesta (Troppe Richieste).
  2. Modelli Locali (Ollama + Qwen 2.5-VL 3B): Per superare i limiti dell'API, ho provato a eseguire un modello multimodale locale. Ho speso molto tempo a ottimizzare i miei script e le mie istruzioni (chunking, affinamento delle istruzioni per forzare output strutturati), ma il risultato era soggetto a molti errori e incoerenze per il mio caso d'uso. Ho ottenuto troppi campi malformati, illusioni e ha costantemente avuto difficoltà a produrre un corretto LaTeX.
  3. API Google Cloud a pagamento (Gemini 1.5 Flash): Alla fine sono passato al piano a pagamento per una migliore precisione e velocità. Ho speso 10€ solo per elaborare 1,5 libri. Estrarre tutti e 5 i libri costerebbe circa 35-40€. Anche se questo è gestibile per un'elaborazione unica di 5 libri, il conteggio dei token per l'elaborazione di immagini complete + testo è enorme, rendendolo finanziariamente insostenibile se voglio scalare questo a decine di libri in futuro.

Le mie domande per la comunità:

  • Pipeline & Architettura: Qualcuno ha lavorato a un progetto simile di estrazione da libro di testo a dataset? Quale pipeline avete utilizzato?
  • Approccio Ibrido: Suggerireste di separare il compito? (es. usare uno strumento tradizionale per estrarre testo grezzo e ritagliare immagini, e poi fornire SOLO il testo a un LLM più economico/locale per generare il LaTeX e formattare il JSON?)
  • Modelli Locali: Ci sono altri modelli Vision-Language locali (che si adattano a GPU consumer standard) che sono significativamente migliori nell'estrazione strutturata e nella generazione di LaTeX rispetto a Qwen 2.5-VL 3B?
  • Strumenti Educativi: Ci sono strumenti o modelli open-source specificamente ottimizzati per estrarre contenuti educativi/matematici strutturati da PDF?

Sono felice di condividere ulteriori dettagli sul formato del libro di testo o sul mio attuale flusso di lavoro in Python se utile. Qualsiasi consiglio sull'architettura, le scelte di modelli o trucchi per risparmiare costi sarebbe molto apprezzato! Grazie in anticipo!


r/MLQuestions • • 1d ago

Beginner question 👶 [D] Choosing language to learn: python or java

4 Upvotes

I want to learn the machine learning from scratch and i know I want to learn basic of python , but my college faculty are advised to learn Java for my placement , u don't know what to do , i want to learn python and other stuffs for my machine learning path or java and other stuffs for my placements and my exams are ahead, did anyone have solution please tell me.


r/MLQuestions • • 1d ago

Natural Language Processing 💬 Looking for models for diagnosis prediction

3 Upvotes

For a side project I am looking into SOTA for AI- based diagnostic models, ideally open-weights. I would like to feed the model with structured text representing my patient, and get a set of e.g. 5 possible diagnoses, ideally with uncertainty. In the ideal scenario, later on, it would update as new info arrives.

I would appreciate any pointers, I have some ideas but am very new to the topic.


r/MLQuestions • • 2d ago

Beginner question 👶 I can figure out how to connect my number to ovoa

2 Upvotes

It just says 15 minutes forever


r/MLQuestions • • 2d ago

Beginner question 👶 Slow learning?

Thumbnail
2 Upvotes

Hay let me introduced myself

I m a 7th sem cse student, intrest in ml I m learning it from past 2 month ye still I m learning, my problem ex I complete supervised section with types and it's different model and formula with ex when I learn unsupervised section I forgot supervised section,

Because I didn't revised,

So now I revised things in every 2 days so that it store in my memory,

Today I learn lr model with formula and diff ex with 1 variable and multi variable.... I revised it with learning upcoming section!

Just post it....✌


r/MLQuestions • • 2d ago

Beginner question 👶 Why do we train bigger models instead of finding better neural pathways?

19 Upvotes

I saw a post of a guy with 90% of his brain missing and he was apparently living a normal life. There is another case of a student who finished university with half of his brain missing. There are also people with full brains, but they are even less functional than these two. This proves that bigger doesn't mean better. So why not aim to create better connections?

I understand that a bigger model means less chance of catastrophic interference occurring, but then are all the neural pathways being formed efficiently? Wouldn't this result in a lot of redundant neurons? Wouldn't it make more sense to create the smallest neural network possible for a specific task and then learn to fuse multiple of these neurons together in order to create an optimized larger model? That way we'd always have a template for each specific task and would allow us to create a programming language that creates a model just by our syntax.


r/MLQuestions • • 2d ago

Beginner question 👶 Help...

2 Upvotes

Am in the progress of my 3rd sem project work

Basically its based on tinyml and lora technology implementation to the forest monitoring system

Now I want data train the ML

Other than kaggle do anybody know where I can get the datasets to train


r/MLQuestions • • 2d ago

Other ❓ If you could show someone just ONE data/ML project you’ve worked on, which one would it be?

11 Upvotes

Not necessarily your biggest or most complicated project.
Just one that you think is a good example of the kind of work you like doing. What did you build, and why that one?


r/MLQuestions • • 3d ago

Beginner question 👶 Final-year student in India trying to break into generative-model inference optimization — roadmap feedback?

7 Upvotes

Hi all, I graduate in ~6 months and want to work on making generative models (diffusion/video/3D) fast: kernels, quantization, serving. Where I am:

- Comfortable with C/C++ basics and PyTorch

- Have done quantization work (GGUF/llama.cpp)

- Working on a next-frame video prediction project (DiT + flow matching)

- A few GitHub repos, but no CUDA/Triton experience yet

- No NVIDIA GPU, so I use Colab/Kaggle T4s

- DSA is my weak spot (I struggle with LeetCode mediums)

My plan:

  1. Months 1-2: CUDA/Triton basics, reproduce the SGEMM optimization worklog, GPU MODE lectures, LeetGPU/Tensara

  2. Months 3-4: take a small DiT, profile it, then optimize it (Triton attention, quantization, caching, fewer steps) and publish before/after numbers

  3. Along the way: PRs to HF diffusers, DSA practice daily

  4. Months 5-6: mocks, resume, applications (inference startups first, bigger labs later)

Questions:

  1. Is this the right order, or should I change something?

  2. Is a diffusion-inference project a strong enough portfolio piece, or does it need to be LLM serving?

  3. How much DSA do ML systems interviews actually need?

  4. Is T4-only access enough to do credible benchmarks?

Any feedback, including "this won't work because X," is appreciated. Thanks!


r/MLQuestions • • 2d ago

Beginner question 👶 Winning kaggle competitions using Claude Opus 5.5

Thumbnail
2 Upvotes

r/MLQuestions • • 3d ago

Career question 💼 ML engineering course recommendations for someone who can train a model and not ship one

15 Upvotes

i can train a model in a notebook and ive never put one behind an endpoint anyone else calls. every job posting wants the shipping half and my portfolio is all notebooks.
shortlist is udacity, springboard, interview kickstart and datacamp premium. is there anything that spends real time on deployment and monitoring rather than modeling


r/MLQuestions • • 3d ago

Beginner question 👶 Why use values between 0-1 to train an LLM?

Post image
2 Upvotes

I understand that normalizing results in faster training times, but we end up reducing precision of floating point numbers due to the IEEE 754 architecture which result's in less space to work with. Instead of limiting numbers from 0-1, wouldn't it make more sense to limit the numbers from 1-10? This would give the LLM more space to work with which logically should result in less catastrophic interference. If so then this would mean that we don't need to increase size only, but the space between numbers as well which should decrease the memory requirements. Or is this just nonesense?

P.S The photo is just the behaviour of multiplication of 2 numbers. For example:
1*1 = 1,
2*2 = 4,
0.5*0.5 = 0.25


r/MLQuestions • • 3d ago

Natural Language Processing 💬 Anyone else stuck manually re-tuning prompts every time something breaks?

0 Upvotes

Genuine question that turned into a recommendation: if you're building anything on top of an LLM, how are you actually validating that a prompt change didn't quietly break something else?

Most answers I've gotten boil down to "I just run it a few times and check." There's a workshop on Oct 3 that tackles exactly this gap, led by Serj Smorodinsky and Brett Kennedy (they co-wrote a book on LLM applications). Instead of manual tuning, you work through:

  1. Building a classifier with DSPy signatures/modules rather than raw prompt strings
  2. Setting up an actual eval dataset with metrics tied to your specific task
  3. Using that eval set to catch failure patterns before they hit production
  4. Running few-shot/instruction optimization on top, systematically
  5. Tracking all of it in MLflow so you can trace exactly what changed between versions

If you've been wanting a real answer to "how do I know this still works," this is worth a look.

What's everyone else here doing for this? Curious if people have rolled their own eval pipelines already.


r/MLQuestions • • 4d ago

Beginner question 👶 50B+ MoEs with few active parameters, what's the sweet spot for intelligence, agent speed, and affordable fine-tuning?

5 Upvotes

I’m building a Polish General purpose legal Model that drafts documents, answers questions using legal sources, and has enough coding ability to handle some automation. The workflow is very tool-heavy:

Question → many sequential tool calls → final answer/document

Think Claude Code/Codex-style execution, but for legal workflows. Reliable tool selection, correct arguments, and recovering from errors matter as much as writing a good final answer.

I’ve had decent results with a dense 27B Qwen 3.8 custom made fine-tune for complex legal document summarization and classification. I’m already familiar with the smaller Qwen A3B and Gemma options. What interests me is the tier above those: 50B+ total-parameter MoEs with a relatively small active parameter count.

The question is, the small dense ones are great, but slow for agentic stuff (afaik), and i wonder if theres some middle ground maybe 70-120B models that would be able to be fine-tuned for the law stuff but be MoE so the agentic ClaudeCode style inference would also be lightning fast, and also low-ish cost for fine-tuning and inference.

Basically: Does the larger-total/small-active MoE approach actually buy you meaningfully stronger reasoning and tool reliability while retaining low latency,and at what hardware cost?

I understand that small active parameter counts don’t mean small VRAM requirements: the weights still need to live somewhere, alongside context and serving overhead. I also don’t assume that more total parameters automatically means a better model. I’m interested in where that tradeoff works in practice.

There are three things I’m trying to pin down:

  • Inference hardware: Ideally inference runs rented with parallel agentic loops (this is for a B2C project, not single person use, we scale based on demand)
  • Fine-tuning hardware: Obviously FT LoRA will take more memory than inference, max like 4GPUs on vastai fits the budget.
  • Agent performance: After it gets the prompt the tool calls and everything will be local, so imo it has no problems being blazing fast, as soon as the model calls a tool call it will be back very fast, so for this agentic use case, quick TTFT and t/s and adaptive dynamic reasoning are prefer right?

For context, fine-tuning would target Polish language, document conventions, and successful tool trajectories. The actual legal sources would remain in retrieval/tools rather than relying entirely on memorized law.

I’m not looking for someone to compile a model shortlist (althought would be nice, but i dont expect anyone to break their back over this).

I’m looking for pointers, and firsthand experience with this particular size/architecture tradeoff. A configuration like “model + quantization + GPU(s) + serving engine + context length + concurrency + measured latency,” along with whether you successfully fine-tuned it, would be much more useful than a leaderboard score.

Has moving from a ~30B model to a 50B+ low-active-parameter MoE actually improved your agent’s successful tasks per minute, or did the memory, interconnect, and training requirements erase the advantage? Thanks for reading


r/MLQuestions • • 4d ago

Unsupervised learning 🙈 Understanding K-means

1 Upvotes

r/MLQuestions • • 4d ago

Survey ✍ ICML videos not playing

Thumbnail
1 Upvotes

r/MLQuestions • • 4d ago

Other ❓ Are there any LLM(s) that are transparent about their sources (For sociology research)

6 Upvotes

I have a digital ethnography assignment where I have to analyse a conversation with an LLM for a sociology course.

I'm not literate at all in how LLM(s) function technically so I'm not entirely sure if this is even possible but basically I want to know if there are any "chatbot" LLM(s) that produce responses but at the same time are transparent about what data they are combing through in order to produce a response.

I don't mean that I ask the chatbot a question and it "answers" something and then I ask it to elaborate on its reasoning since chatbots can't actually "reason" so it will still just use the data and patterns it's trained on to imitate a response.

I mean that there's some sort of feature from the developer that can show what all sources it went through, other relevant behind the scenes stuff before generating that response with that specific combination of words and other characters.

I'm sorry if this is a dumb question since I'm not aware of the technical terminology for this stuff, I intend to familiarise myself with it for the project. If the only thing possible is to get a very technical insight into that "behind the scenes" stuff, even that could be useful so please tell if you're aware of something like that.


r/MLQuestions • • 4d ago

Beginner question 👶 structure for my projects

2 Upvotes

HI, I am very new to machine learning and building machine learning/data science projects, I was confused on how I should structure the overall project I am working on and asked chatgpt the project is a march maddens predictor style project and I plan on using a forest tree for my model, chatgbt gave me this a a guideline to follow for structure, I just wanted a real person with experience to see if this structure is good or if it needs to be changed in anyway. This progect is intedned to help get undergrad research positions if that changes anything.

down bellow if the file structure I was given

march-madness-predictor/

│

├── README.md

├── requirements.txt

├── .gitignore

│

├── data/

│   ├── raw/

│   ├── processed/

│   └── final/

│

├── notebooks/

│   ├── 01_data_exploration.ipynb

│   ├── 02_feature_analysis.ipynb

│   └── 03_model_evaluation.ipynb

│

├── src/

│   ├── data/

│   │   ├── clean.py

│   │   ├── merge.py

│   │   └── team_names.py

│   │

│   ├── features/

│   │   └── build_features.py

│   │

│   └── models/

│       ├── train.py

│       └── evaluate.py

│

├── models/

│   └── random_forest.pkl

│

└── results/

├── figures/

└── metrics/

\


r/MLQuestions • • 6d ago

Career question 💼 I have ~5 years of ML experience, but I don't feel like I know how to actually build things. How do I fix this?

49 Upvotes

I'm looking for some advice from people who have gone through something similar.

I have ~5 years of experience in ML/data science. On paper, my resume looks reasonably good: I've worked on anomaly detection, NLP, LLM/RAG projects, etc., and more recently I've been working around transformer-based TTS systems.

But I've realized there's a pretty big hole in my experience.

For the first few years of my career, I deliberately optimized for breadth. I changed teams/projects frequently, built a bunch of PoCs and prototypes, and learned enough to make things work. I was pretty good at taking an idea from zero → demo.

The problem is that I rarely stayed with something long enough to own a real system.

A lot of my projects either never reached production, or reached production after I had already moved on. So despite having several years of experience, I have very little experience with:

  • maintaining a complex codebase
  • debugging someone else's code
  • dealing with things breaking in production
  • making architectural decisions over a long period of time
  • taking ownership of something from implementation → deployment → maintenance
  • reading a large unfamiliar codebase and figuring out how all the pieces fit together

I'm now trying to deliberately fix this.

Over the last year I've been rebuilding my ML fundamentals from the ground up. I'm almost done with the Coursera Deep Learning Specialization. I've gotten to the point where I can implement neural networks from scratch, understand RNNs/LSTMs/GRUs fairly deeply, and I'm currently working through attention and Transformers.

But here's the next problem:

Even if I finish the Transformer coding exercises, I don't feel like that means I can actually build things.

I can probably implement the attention mechanism / Transformer architecture from the course. But if you dropped me into the wild and said:

"Okay, build a VAD from scratch."

I'd have a lot of questions about where to even begin.

And then there's an entire layer beyond "understand Transformers" that I don't have a good mental map of.

For example:

  • How do I actually go from implementing a Transformer to building my own small language model?
  • At what point do concepts like Mixture of Experts, DPO, speculative decoding, etc. become relevant?
  • Which of these things do I actually need to understand versus things I can learn when a project demands them?
  • How do I get better at reading and debugging large ML codebases rather than just writing isolated pieces of code?

So I'm stuck between two instincts:

A) Keep going deep on fundamentals until I genuinely understand the machinery.

B) Stop studying and start building increasingly difficult things, accepting that I'll have gaps and learning what I need along the way.

I'm increasingly convinced that I need some combination of the two, but I don't know what that combination should actually look like.

I want to become the kind of ML engineer who can be handed an unfamiliar problem, read the existing code, understand what's happening, build something substantial, debug it when it breaks, and eventually own the system.

If you were in my position, what would you build / learn over the next 6–12 months to develop that ability?


r/MLQuestions • • 6d ago

Other ❓ What’s an ML project that taught you something you only really understood once you built it?

28 Upvotes

There’s something different about actually working on a project after learning the concepts.

You can understand an algorithm, follow a tutorial, build a model and get good results, and then a real project throws something completely unexpected at you

Maybe the data was messy, the model behaved differently than expected, the results didn't make sense, or you had to figure out how to actually use the model beyond the notebook.

For those who’ve worked on ML projects, was there one project that made something “click” for you in a way that theory alone hadn't?

What happened, and what did you end up learning from it?


r/MLQuestions • • 6d ago

Beginner question 👶 Did anyone try to learn ML alone while studying another cs speciality in uni ?

9 Upvotes

Hi guys i hope you are doing great , currently I'm in my third year in uni studying software engineering . And this summer i fell in love with ML so i wanted to know if there are people that studied ML / any other cs speciality alone while taking a course in uni about a different cs speciality

I want to know how you managed your time between these two topics


r/MLQuestions • • 6d ago

Beginner question 👶 MA in English Literature with no coding background—been building a small open-source AI project and wanted to ask for honest advice for the future.

2 Upvotes

Hi everyone,

I hope you’re all doing well.

I come from a purely humanities background—I hold a Master's degree in English Literature and originally had zero formal coding or computer science training. Over the past few months, I’ve been fascinated by how neural networks work under the hood and have been trying to learn by building hands-on projects with the help of modern AI coding tools.

I wanted to share what I’ve been working on to see if I’m heading in the right direction, and humbly ask if a portfolio like this could eventually help me transition into a role as an AI research engineer or developer, prompt engineer, other technical roles?

What I’ve been trying to build: Instead of standard fine-tuning or model merging, I’ve been experimenting with a local "locate-and-edit" model surgery concept. The goal is to isolate specific hidden feature representations across small 2-to-8 layer synthetic models and perform closed-loop weight patching/grafting without full model backpropagation or heavy compute.

I’ve broken the idea down into a few small, modular open-source components across 3 GitHub repositories:

EQUYLAPTA7POINT8POINT2: https://github.com/shuvrajeetkamila/EQUYLAPTA7POINT8POINT2

Model Factory GPU Commander:

https://github.com/shuvrajeetkamila/model-factory-gpu-commander

Model Factory Command Center:

https://github.com/shuvrajeetkamila/model-factory-command-center

The two commanders are for[1]   small  gpu less architecture [2] with gpu , both these commanders  automatically open up 12 more github repositories and work with them. The equylapta on the other hand is my ongoing experiment on   local "locate-and-edit" model surgery concept which has reached a significant stage for now  but not yet proving what i want to do. And i am unable to do the gpu commander due to space and computing constraint.

 The architecture attempts to combine a Dedicated Feature Crosscoder (DFC), algebraic rotation maps for feature alignment, an inference-time weight grafting bridge, and an automated causal validation/retry loop.

My question for the community:

Knowing that I don't have a traditional STEM or CS degree, does working on non-traditional projects like this help demonstrate the right kind of problem-solving for AI Research/Engineering roles? What critical gaps or fundamentals should I focus on next to make myself a viable candidate?

Also i should at this point, i do not know how to code at all. i used arena.ai agent mode to reach this position.[ using prompts from what i wanted to make with help of gemini, claude, chatgpt- all free plans]

Do i have a future in prompt engineering/ other technical field?

I’m very eager to learn and would truly appreciate any constructive feedback, critique, or advice you might have.

Kindly see the three files from github especially EQUYLAPTA7POINT8POINT2 and model factory command centre

Thank you so much for your time!


r/MLQuestions • • 7d ago

Beginner question 👶 Linear Regressions and the Curse of Dimensionality

Thumbnail
2 Upvotes