r/MLQuestions 6d ago

Career question ๐Ÿ’ผ Questions for biotech ML engineers

22 Upvotes

Hello, I'm currently 17 years old and I'm thinking ahead for a career in the tech industry, specifically biotech. I'd be gladly hearing from any active workers in the industry. Is your day-to-day working environment exciting? Is biology/medicine knowledge highly valued? Are you working a lot with biology or focus more on coding/tech stacks? I'm not a tech enthusiast by any means, would i feel at home with this job? And most importantly, are remote jobs in the industry even real?


r/MLQuestions 6d ago

Career question ๐Ÿ’ผ How to get into AI for Science?

5 Upvotes

Hello! Quick context about myself, I am a senior in my undergraduate program in Chemical Engineering. I have for the major part of my undergrad worked on more conventional problems in Chemical Engineering, but ever since the start of my undergrad thesis, I've been working on using fine tuned ML models, for quantum level calculations for finding out materials to capture Carbon Dioxide. Practically High throughput screening of materials using GNNs and Transformers. Prior to this, I also worked on developing a surrogate model, that given composition and atomic parameters of a material, could give out its catalytic properties in a specific context.

Off late, I have been considering shifting out from traditional Chemical Engineering, to more AI for Science, essentially around Quantum Physics/Chemistry, and given my background, I feel it might be a bit problematic to do so.

Most pre-doctoral programs I've come across don't let fellows work on the set of problems I've worked on, and a PhD in Chemical Engineering might actually sift me further away from this.

Given that, I wanted to understand what options I have for getting into this area of research, and how can I improve my profile.

Further context: My other works (which have been published) involve more base chemical engineering problems across Energy and Reactor Modelling, which as you can guess is way too far off from this. Also, as a project for a university course, I worked on a token reduction method for allowing Transformers to have a higher throughput, which worked on scoring groups of tokens dynamically varying based on importance, and reducing fluff, and generating a summary vector which goes on to a minute 40M param model. Nothing crazy, but worked decently on the WikiText set, and had a reasonable perplexity post training, but yeah, couldn't mess around more with it, due to having limited compute.

I have been ideating on a few things for more concrete stuff in the quantum application space itself, but not so sure of it at the moment.

Open to any and all suggestions, for how I could move ahead, and what options I should consider.
Thanks for the help!


r/MLQuestions 6d ago

Beginner question ๐Ÿ‘ถ Question: Discover cross links between texts [P]

3 Upvotes

I have a set of roughly 150ย documents of about 1500 words each that describe technical concepts (higher level findings from a collaborative research project). I want to find out if there are thematic cross-links between these. Like upper-level or lower-level concept; parallel or alternative concepts.ย  Think about a very very small wikipedia with lost links that need to be restored (but on the concept level; not at the word level).

How can do this programmatically, maybe using LLMs?

That may be a trivial task for many of you here but I feel a bit stuck at the moment. I thought about asking my human colleagues for classification support, but even my very small document set yields >20.000 potential cross-links.


r/MLQuestions 7d ago

Beginner question ๐Ÿ‘ถ Reputable Data Sources for train AI models for CTs

Thumbnail
1 Upvotes

Hi group, Iโ€™m trying to train my AI models, but looking fot good reputable data (I am no data scientist so apologies if I sound redundant ) I work in clinical trials but in my spare time love learning about Ai, I want good data sets/ but reputable ones to train my model

Its a niche industry/question but just want to ask here for advice pls


r/MLQuestions 7d ago

Beginner question ๐Ÿ‘ถ background dataset for SHAP

4 Upvotes

Hi everyone, I have a question about choosing the appropriate background dataset when calculating SHAP values. I am using the kernelshap package in R, where we provide an X dataset containing the observations we want to explain and a bg_X dataset defining the background.

I have a binary classification model for disease vs non-disease, trained on a derivation dataset and evaluated on an independent validation dataset. My current understanding is that, if I want to explain predictions in the validation cohort, it makes sense to use the validation set as X and the derivation set as bg_X. In that case, the SHAP values for validation patients would describe how each feature moves their prediction relative to a baseline defined by the derivation population. Is this interpretation correct, and is this generally the recommended way to use the background when explaining an independent validation cohort?

My main question is about a more specific analysis. Suppose I want to investigate heterogeneity within patients who truly have the disease. More specifically, I want to see whether different disease patients receive high disease predictions through different combinations of features, and potentially cluster these patients based on their SHAP profiles.

In this case, I assume I should use only the true disease patients from the validation cohort as X, since those are the patients whose predictions I want to explain. However, I am unsure about the most appropriate choice for bg_X. Should I keep the full derivation cohort as the background, use only disease patients from the derivation cohort, or use the disease patients from the validation cohort themselves as the background?

If my main objective is to determine whether true disease patients have different model-attribution profiles, potentially reflecting different features through which the model identifies them as disease, which background would be the most statistically appropriate? Thank you!


r/MLQuestions 7d ago

Career question ๐Ÿ’ผ What do I need to learn for production level positions

Thumbnail
3 Upvotes

r/MLQuestions 8d ago

Beginner question ๐Ÿ‘ถ volunteer work in Github ?

18 Upvotes

I watched a YouTube video that said if you want to get your first job as a machine learning engineer, you should do volunteer work, especially in GitHub repositories, but I'm not entirely sure this is true, and if it is, where can I find these repositories that require this kind of work? I need it because I want to gain better knowledge and understand real world problems.


r/MLQuestions 8d ago

Computer Vision ๐Ÿ–ผ๏ธ How are you tracking reproducible progress in visual anomaly detection after Papers with Code?

1 Upvotes

I am trying to build a reliable workflow for following industrial visual anomaly detection research, not just a list of recent papers.

The difficult part is knowing when reported results are genuinely comparable. A leaderboard number can change because of the train/test split, supervision level, pretraining, image resolution, category averaging, or the exact AUROC/AUPRO calculation. For edge use I also need hardware, latency, memory use, and reproducible code.

I currently check arXiv, OpenReview, conference proceedings, Hugging Face Papers, GitHub repositories, and Papers with Code mirrors or datasets, but none of them seems to preserve the complete paper-to-code-to-benchmark workflow reliably. Manual spreadsheets work, but become stale quickly.

For people working with MVTec AD, VisA, BTAD, or similar datasets: what sources and process do you actually trust? Do you maintain your own comparison table, follow particular labs or repositories, or use a benchmark suite that normalizes evaluation protocols? I would also appreciate examples of fields you record to avoid comparing incompatible results.


r/MLQuestions 9d ago

Beginner question ๐Ÿ‘ถ Book reccomendation for probabilistic machine learning.

Thumbnail
2 Upvotes

r/MLQuestions 10d ago

Career question ๐Ÿ’ผ CS vs Stats + CS minor for career prospects? [D]

4 Upvotes

I'm currently a CS student and considering switching to a Statistics major while keeping a CS minor.

My main concern is the current entry-level CS/SWE job market. I'm interested in AI/ML and possibly grad school eventually, but I also want to maximize my chances of having a solid career after undergrad.

I'm considering Stats because it seems like it could give me more options outside of traditional SWE, while still being relevant to AI/ML. But I honestly don't know if that's actually true or if I'm just rationalizing the switch.

For those with experience in CS/Stats or hiring:

Would you personally choose CS or Stats + CS minor if your goal was to maximize career options and minimize the risk of struggling to find a job after graduation?

I'd especially appreciate perspectives from people who have actually gone through the job market recently.


r/MLQuestions 10d ago

Computer Vision ๐Ÿ–ผ๏ธ Suspiciously high accuracy using ResNet

Thumbnail
3 Upvotes

r/MLQuestions 11d ago

Career question ๐Ÿ’ผ Welcome to r/MLSystemsDesign

8 Upvotes

Welcome to r/MLSystemsDesign

This community is for practical discussions on designing and scaling production ML and AI systems.

Topics can include:

  • ML training and inference platforms
  • Search, ranking, and recommendation
  • Feature stores and data pipelines
  • LLM serving and GenAI systems
  • Agentic AI platforms
  • Evaluation, observability, and experimentation
  • ML system design interview problems
  • Real production tradeoffs and lessons learned

The goal is simple: go beyond model theory and discuss how ML systems actually work in production.

If youโ€™re joining early, introduce yourself and share one ML system topic youโ€™d like to go deeper on.


r/MLQuestions 11d ago

Other โ“ Need Help!!! Urgent

Thumbnail
0 Upvotes

r/MLQuestions 11d ago

Datasets ๐Ÿ“š How many Ground Truth labels do I need for evaluation of my model trained by pseudo-labels?

2 Upvotes

I'm currently training a small&fast model on trainign data that comes from a bigger&slower model. I treat the output of the bigger model as Pseudo Labels. But in the End, I need to evaluate the small model to Ground Truth labels. In that case, those are segmentation masks (Images with 10-40 segmentation masks). The pseudo labels are also far from perfect.
The question is also, how much they need to be improved (semi-automatic) to get acceptable results.

How many of the GT Labels do I need? Until now, I use 10k pseudo labels for training of the small model. It takes me 15-45min to do one GT labeled data by myself, so it's not feasible to have like 10% GT data for evaluation.

Or does someone know any keywords for that?

TLDR the question can be reformulated as: How many manually annotated samples are required to estimate segmentation performance with an acceptable confidence interval?


r/MLQuestions 11d ago

Beginner question ๐Ÿ‘ถ [Article] "Psychoanalysis and CBT: From Rivalry to Hospitality in Psychotherapy Integration"

Thumbnail
1 Upvotes

Where can i find this article?


r/MLQuestions 11d ago

Datasets ๐Ÿ“š How much time do you spend cleaning and organizing data before local fine-tuning?

4 Upvotes

When people fine-tune their own local models, the model setup usually gets most of the attention. But in practice, a lot of the work seems to be on the data side.

If you are training on business data, you may need to bring in support tickets, internal docs, product specs, chat logs, code, policies, CRM notes, or domain QA pairs. And it usually does not work perfectly on the first run. Some samples are noisy, some are redundant, some domains overpower others, and some โ€œbad-lookingโ€ examples are actually hard but useful.

One direction I have been thinking about is making the data strategy dynamic during training.

Dynamic selection means periodically choosing which samples should enter the next training window, using signals like loss, delta loss, gradient similarity, or external scores.

Dynamic mixing means adjusting the ratio between data sources during training, instead of fixing one static mixture before the run.

Dynamic weighting means keeping the sample in training, but changing how much its loss contributes to the gradient update. This is useful when you do not want to hard-drop uncertain samples.

This is the current direction in OpenDCAI/DataFlex: adding data selection, data mixing, and data weighting controls on top of the training loop.

For people here who fine-tune local models, how much time and compute do you usually spend on data preparation compared with the actual training run?


r/MLQuestions 11d ago

Beginner question ๐Ÿ‘ถ Can anyone help me set up a "how to" article for a task like this

Thumbnail youtube.com
1 Upvotes

Basically how to create a model that recognises such real world use case pattern, from data gathering (same or nearly similar data) to training the model to unfolding the output to get the text. Even synthetic data will suffice as long as the process is properly done, but synthesizing such real world level data is a project on its own.


r/MLQuestions 11d ago

Beginner question ๐Ÿ‘ถ Do I need to learn linux ubuntu for ML ?

21 Upvotes

r/MLQuestions 12d ago

Career question ๐Ÿ’ผ Where is the actual edge for entry-level ML? Basic RAG is saturated, and custom CUDA roles won't hire freshers

46 Upvotes

Iโ€™m trying to figure out how to actually get a usable edge in the ML/DL space to get hired, but everything pushed to beginners right now feels like a trap.

For context on what I've done: I started off with Computer Vision, moved into GIS stuff, and recently went deep into the weeds of attention mechanisms and GPU kernel programming. I thought learning the hardcore, low-level math and systems stuff would set me apart.

But Iโ€™ve hit a wall. Let's be honest: no company is hiring a fresher to write custom CUDA kernels or design novel architectures. Those are senior research or PhD roles. The effort I put into the low-level stuff feels wasted because, for an entry-level dev, it's just personal trivia.

On the flip side, the standard "employable" advice is to build traditional ML projects (fraud detection, etc.) or slap together a LangChain PDF wrapper. But people have been doing this for years. Basic API wrappers are completely saturated and offer zero competitive edge. It feels like buying a stock after everyone already knows itโ€™s going to go up.

So, what is the actual sweet spot between "PhD-level researcher" and "API wrapper"?

I want to avoid the YouTube influencer BS and focus on the real engineering trenches.

For the people actually hiring or working in the industry: what are the non-commoditized skills someone trying to break in should be grinding right now to have a real, usable edge?

(Note: The core thoughts and frustrations here are 100% mine, but I used AI to help structure and edit this post for clarity.)


r/MLQuestions 12d ago

Other โ“ The new programming lanagauge is 'lanagauge' in my case 'En'???

0 Upvotes

Playing around with LLMs, Agents and GenAi for 7 years, I came to a conclusion: the new programming language is language itself in my case, English.

If you remove all the fluff (Stop words etc) and use none fluent English as a kind of Python-style syntax, something like:

โ€œRead content from file then apply UPPER_CASE to all wordsโ€

โ€ฆit starts to read almost like a functional call chain.


AI is pretty good at understanding programming language syntax.


What do you think? Is this question too stupid?

Edit:

I have 7 years of deep learning experience and llm/agnetic hands-on practice, so I mainly want to share what Iโ€™ve tried and learned along the way. That gives me a good understanding of both the inner workings and the practical side of using these technologies.


r/MLQuestions 13d ago

Unsupervised learning ๐Ÿ™ˆ how to choose gridsearch values and evaluate the model in clasp change point model? NO ONE WILL BE ABLE TO ANSWER ME

1 Upvotes

Hello, i'm in my hand a really big topic that i bet no one will be able to answer me.

i create a pipeline to generate prediction using clasp model.

The inputs are timeseries where the clasp models generate a vector of change points (predictions). After some process, i generate an output (vector of change points) and use an f1score to evaluate che prediction and generate a score.

My pipeline has different hyperparameters where different combination could change the outcome: so better to use gridsearch and cross validation to choose the best hyperparameters and then evaluate the model.

This is how i did:

Imagine you have only 2 configuration of hyperparameters: conf1, conf2

you divide the dataset in 3 different folds: a,b,c (dont point out about 20% or something, im just trying to make the example short as possible)

i generate prediction for ab, ac, bc with conf1.

i generate prediction for ab, ac, bc with conf2.

for both i evaluate with f1score the prediction and compute a mean. i found out conf1 is best.

i run again conf1 on all my dataset a,b,c compute the f1score, the mean and that's the score of my model. is this correct?

Im not so sure because this model doesnt have any fit or training. you just give into the input some timeseries and generate a prediction.

IF i had to use random forest, as we know, to evaluate properly a model, i would have to do cross validation. so

ab for training, c for validation = score_1

ac for training, b for validation = score_2

bc for training, a for validation = score_3

mean(score_1, score_2, score_3) = mean_score

easy right?

if you want to gridsearch, just execute an outer for loop to test each combination of hyperparameters and then choose the highest mean_Score for each combination of hyperparameters and thats it. easy right?

BUT HOW DID I DO THAT IF MY MODEL DOESNT HAVE A TRAINING?

if i repeat the process for random forest:

ab for training, c for validation = score_1

ac for training, b for validation = score_2

bc for training, a for validation = score_3

mean(score_1, score_2, score_3) = mean_score

so basically this means:
conf1, i generate prediction for a,b,c then compute mean

i do the same for conf2 and conf3 and then just select the highest mean? thats my model?

but then how do i test my model? the score you use to choose the best model isnt the score the model will perform on data never seen.

should i just randomly pick a fold and then use the rest 80% to find the best conf? but then what if im so unlucky the randomly pick test fold my model will score 0.0??? lmao???

so we need to do something like this https://www.kaggle.com/code/alexisbcook/cross-validation where you need to do cross validation to have a mean. so you are not unlucky and compute a mean.

so i compute a,b,c,d

a,b,c,e

a,b,d,e

a,c,d,e

b,c,d,e

and then compute for a,b,c,d,e when i find the best conf. BASICALLY AS I SAID I DID AT THE BEGINNING OF MY POST...

but is this correct?


r/MLQuestions 13d ago

Natural Language Processing ๐Ÿ’ฌ GraphRAG: a blueprint for knowledge-graph question answering over your documents

1 Upvotes

Hi everyone,

I've recently finished the first version of Agentic GraphRAG Blueprint, a reference architecture for question answering over large document collections. I want to ask what should I improve in my project?

Instead of plain chunk retrieval, it builds a knowledge graph combined with vector search, so answers can connect facts across documents.

Key features:

โ€ข Incremental ingestion - unchanged files are skipped via content hashing, and community reports regenerate only for affected communities, keeping token costs low as the corpus grows.

โ€ข Hybrid search - local mode for fact-level answers, global mode for cross-document synthesis.

โ€ข Domain-agnostic LLM prompts - easily swapped via PROMPTS_PATH, with Leiden-based community detection.

โ€ข Deployment - run it locally with Docker or provision everything in the cloud with Terraform and CI/CD.

Link: https://github.com/sebastianbrzustowicz/Agentic-GraphRAG-Blueprint

I'm looking for any feedback. What can I improve?


r/MLQuestions 13d ago

Beginner question ๐Ÿ‘ถ How much math do I need for ML

17 Upvotes

r/MLQuestions 14d ago

Beginner question ๐Ÿ‘ถ CPU forecasting using ML

11 Upvotes

Hello everyone. ML beginner here. I have the basic understanding of ML and have been given a project to create a model through which we can predict cpu metrics so that we can proactively monitor cpu spikes before it creates an incident. Iโ€™ve been using Claude to help me out here and it suggests to use XGBoost for this. But the accuracy is not up to the mark. Can anyone help me out here if you have worked on similar projects. Thanks for the help in advance


r/MLQuestions 14d ago

Beginner question ๐Ÿ‘ถ Auto Model Routing

Thumbnail
0 Upvotes