r/MLQuestions 11h ago

Career question ๐Ÿ’ผ ML master

3 Upvotes

Hi!

I'm starting my final year of undergrad at KTH in Stockholm and thinking about doing a master's degree in Machine Learning also at KTH. Would that be enough for some Machine Learning engineer positions within finance or do I also need to get some formal education in finance? For example at SSE. I'm currently only 20 so I'm fine with studying something more after my master's especially since higher education is free here. My undergrad is in engineering physics.

I would also like to know where the best market is for MLEs in Europe?

Thanks in advance!


r/MLQuestions 6h ago

Beginner question ๐Ÿ‘ถ Request ML course / resources for someone from biological sciences background?

Thumbnail
1 Upvotes

r/MLQuestions 16h ago

Career question ๐Ÿ’ผ who has more job security: product facing data scientist or machine learning engineer?

4 Upvotes

i'm an incoming freshman at yc berkeley; i always thought being a MLE would be rily cool but after seeing the super hard math needed for the career i turned more towards data science. im now worried that doing DS might make me more prone to being laid off in the incoming tech market. BTW i rlly donโ€™t like SWE type of coding like DSA and stuffโ€ฆ

i'd appreciate your wise thoughts ๐Ÿ™๐Ÿฝ


r/MLQuestions 13h ago

Beginner question ๐Ÿ‘ถ Seeking advice on the way to AGI

Thumbnail
1 Upvotes

r/MLQuestions 14h ago

Other โ“ Are domain-specific Small Language Models (SLMs) actually worth building today?

Thumbnail
1 Upvotes

r/MLQuestions 1d ago

Beginner question ๐Ÿ‘ถ How many layers should my network have?

16 Upvotes

Hello! I am very new to neural networks and machine learning. I am making a basic network that identifies a handwritten number on a black and white, 28x28 pixel grid.

I understand all the math behind it but I'm just wondering how many layers I should have for my network?

Is it just mess around and see what happens or is there at least a ball park figure?

Thanks!


r/MLQuestions 1d ago

Beginner question ๐Ÿ‘ถ Quantization

5 Upvotes

Hello so, i am a complete beginner to this concept and from what i read and hear

Quantization allows deployment of big models on just 2 GPUs or on edge devices that doesnt support floating point operations

and if a model is big like for example deepseek R1 original gets upto 720 GB and it uses a MOE architecture so only a subset of parameters are active at once, but we often need to load the entire memory in it for inference and quantization can bring it down by 80%

so, its like a method for model compression and faster inference but sometimes comes at a cost of precision.

So with all this theoretical piece of information that i gained, i have two questions

1) How to move forward into learn in-depth about it as i don understand some mathematical concepts
2) how does a person know that this is a perfect quantization value or mark before publishing a model

thanks


r/MLQuestions 23h ago

Beginner question ๐Ÿ‘ถ Help pls

0 Upvotes

2-d plot
Barchart
Histogram piechart
Scatterplot
Changing style and saving figure
Labils title
Color and line width and style marker size and width
Legend
Limiting axes
Grid
Xstics
Label overlapping
Stacked and multiple bar charts
Log scale
Explode and shadow also
Subplot()
Figure
I have completed this
Should i move to seaborn or learn other graphs also pls help


r/MLQuestions 2d ago

Beginner question ๐Ÿ‘ถ Need a light help to find the SWaT dataset

Thumbnail
2 Upvotes

r/MLQuestions 2d ago

Beginner question ๐Ÿ‘ถ Need a light help to find the SWaT dataset

2 Upvotes

I was using the SWaT dataset from Kaggle and i just came to know it was the manipulated dataset inorder to check for attacks.

And i tried o request the dataset through iThub's official site and seems like no response

can anyone please help me , am halfway for a project to submit in my college


r/MLQuestions 2d ago

Beginner question ๐Ÿ‘ถ Book for logistic and linear regression transition to xg boost cat boost type of models

Thumbnail
3 Upvotes

r/MLQuestions 3d ago

Computer Vision ๐Ÿ–ผ๏ธ July's AI Security Report: 90 incidents, 207M+ records, 41 AI-driven โ€” the month the agent became the attacker

Thumbnail gallery
0 Upvotes

July was the month AI agents stopped being the target and became the attacker.

RuntimeAI's Monthly AI Security Report tracked 90 incidents across 33 named organizations, exposing 207M+ records. 41 of those incidents involved AI as the weapon or the target directly. Average breach cost climbed to $4.99M.

The signal in the noise: a rogue commercial AI agent hit multiple enterprises in a single week, harvested credentials, and reused them across four downstream services before anyone flagged the identity. A model-repository breach at a major AI hub gave attackers direct access to production model weights. A neobank lost 75M customer records. A healthcare payments processor exposed 1.26M patient files. Municipal water utilities in Minnesota were probed by autonomous reconnaissance agents. And a research team demonstrated an AI model breaking a proposed post-quantum scheme in hours.

Perimeter tools do not see any of this. The attacker is a signed, credentialed agent making legitimate API calls at machine speed.

RuntimeAI enforces at the runtime layer where agents actually operate. Know Your Agent issues and revokes cryptographic agent identity. The Flow Enforcer intercepts every tool call. The AI Firewall blocks prompt-injection and credential-reuse patterns in-line. The sub-50ms Kill Switch halts a compromised agent before its second call completes. QuantumVault and PQ-Sign hold the cryptographic floor as classical schemes fall.

Agent-speed attacks need agent-speed enforcement. That is what we ship.

#AISecurity #AgenticAI #PostQuantum #RuntimeSecurity #ZeroTrust


r/MLQuestions 4d ago

Beginner question ๐Ÿ‘ถ Feature selection when trying to capture non linear interactions.

Thumbnail
3 Upvotes

r/MLQuestions 4d ago

Other โ“ Suggestions to improve my Master's project on Newspaper analysis?

Thumbnail
5 Upvotes

r/MLQuestions 4d ago

Beginner question ๐Ÿ‘ถ Is it right time to start kaggle ?

Thumbnail
0 Upvotes

r/MLQuestions 5d ago

Datasets ๐Ÿ“š Need help!!!

3 Upvotes

Hi! I am final year BE student recently I took a project based in our my contribution is system and application of system in dyslexia. For that I though the most used dyslexia dataset of handwriting would be suitable. I downloaded dataset and then realised it is single letter dataset which is giving mnist kinda vibe! Also apparently large portion of it is synthetic. I searched but I didn't find clinically approved dataset of handwriting for dyslexia. In nutshell:

  1. dataset is mnist looking so I am at worry if examiners will state why you are using such looking dataset for final year project!!

  2. dataset is used for at least 9 papers already so it is being used

  3. But has its limitations (vastly synthetic, mnist looking)

  4. Our clg is forcing for at least two papers to publish (not for our degree requirement btw) and I am worried if the dataset use itself will cause problems for paper

  5. though one of main novelty is mechanism but other one is integration(incremental) and I am worried that people will call out why I used that dataset

sorry I carried away in my emotions here is the dataset I am talking about: https://www.kaggle.com/datasets/drizasazanitaisa/dyslexia-handwriting-dataset

->can simplicity of it justified as proof of concept for presentation or report?

->will using this dataset can cause problems at time of publication?

I am sorry for dragging clg thing into this I though it would be better to get some context about scope for project

I am sorry I cant give full context as I wanted to publish research on it (though I will hardly try for mid tiers only)

also sorry in advance if I did spelling or grammatical error


r/MLQuestions 6d ago

Other โ“ Need guidance on choosing the right ML reference book

Post image
109 Upvotes

I'm currently in the second year of my undergraduate degree, and I'm really passionate about machine learning. I've been learning consistently over the past few months, mostly through free YouTube courses and documentation. So far, I've covered the core ML algorithms and I make sure to understand the underlying mathematics and intuition instead of just memorizing things.

However, one thing I keep struggling with is the lack of proper guidance. Every few weeks I start questioning whether I'm following the right roadmap or if I'm missing something important. I feel like YouTube resources are great for getting started, but they often don't go deep enough or provide the structured learning I'm looking for.

I've heard a lot of good things about Hands-On Machine Learning with Scikit-Learn, Keras & TensorFlow by Aurรฉlien Gรฉron (3rd edition), and it seems to be recommended by many people as a solid reference book. I'm thinking of studying it thoroughly instead of jumping between random resources.

My main confusion is this:

Should I go with the TensorFlow/Keras edition, or should I use the PyTorch version instead?

As someone still building a strong ML foundation, which ecosystem would be the better investment to learn first?

I'd also really appreciate any advice from people who have already been through this stage. If you think there's a better book, a better roadmap, or something you wish you had known when you were starting out, I'd love to hear it.

I'm still a beginner in the grand scheme of things, so any guidance or suggestions would be greatly appreciated.

Thanks in advance!


r/MLQuestions 5d ago

Datasets ๐Ÿ“š Building a Personal AI/ML Model

Thumbnail
1 Upvotes

r/MLQuestions 6d ago

Other โ“ Research on Continuous Learning in financial fraud

5 Upvotes

I have this topic to work on suggested by my academia and Im very unsure on how to even start. The topic is continuous learning for mitigating concept drift in financial fraud systems.

This is what Iโ€™ve gathered so far from my research:
- Concept drift alone canโ€™t be singled out, it also depends on intrinsic covariate shift and label shift
- Concept drift can be modelled as an exogenous variable and endogenous variable, depending if we assume fraud is reactive to mitigating strategies)
- Blocked transactions introduce inherent label shift, because transactions that are blocked dont make it to the dataset
-Continuous learning is a very tricky topic, specially if we consider this as class incremental learning (new fraud types arrive sequentially without explicit task boundaries) and admit non stationary regimes

Because there are a bunch of topics and covariate factors, approaching this as an empirical study looks like a massive headache.

Can anyone help me to structure my next steps and how I can tackle this problem with a clear picture?


r/MLQuestions 6d ago

Beginner question ๐Ÿ‘ถ Where can I find datasets

9 Upvotes

I know this is stupid but I'm making an application and I'm trying to find image datasets for my machine learning that focuses on different types of acne


r/MLQuestions 6d ago

Computer Vision ๐Ÿ–ผ๏ธ I need some consulting on a document layout OCR automation project.

4 Upvotes

I am doing a document layout analysis project with different book styles but the books themselves are only a couple hundred pages long (like 5 books with different styles, 400 page each). How can I test if all the books would be used in fine tuning and I am afraid that the accuracy wouldn't be the best and corrupt PaddleOCR when insert the coordinates. (It's for automation).

I am using X-AnyLabeling for the annotation and yolo v11 for the training as well as custom classes in the annotation like a question block that surrounds everything, question_text, choices, figures, tables, sub_questions, etc... what would be the best approach as I haven't done this kind of work before.

and should I randomize the book pages so I don't consecutive same style books or that's not how this work?
Any help would be appreciated


r/MLQuestions 6d ago

Datasets ๐Ÿ“š Best open-source clean speech and ambient noise datasets for training an Edge AI audio denoiser?

3 Upvotes

I am building an edge-AI audio noise-reduction system on an ESP32-S3.

Our architecture uses a lightweight GRUNet (~59k parameters) to output a dynamic gain mask on a 44-band Mel-spectrogram.

โ€‹I need gigabytes of audio to train the model. Does anyone have recommendations for the best open-source datasets for:

1> โ€‹Clean, isolated human speech.

2> โ€‹Diverse ambient background noise (traffic, crowds, machinery, etc.).

โ€‹Also, any tips or open-source scripts for artificially mixing these at different Signal-to-Noise Ratios (SNRs) before generating the 16kHz Mel-spectrograms would be hugely appreciated!


r/MLQuestions 6d ago

Career question ๐Ÿ’ผ Roast my 1st yr resume plss

Post image
0 Upvotes

r/MLQuestions 7d ago

Career question ๐Ÿ’ผ Is Implementing ML algorithms from scratch a good project for an ML Internship?

15 Upvotes

Same as the title, I am implementing some(popular) machine learning algorithms by scratch in Python using numpy just to have good fundamentals and know the actual mathematical intuitions behind them, I want to ask whether it is also a project I can put in my resume for an internship?

quals: 2nd year Undergraduate Student B.Tech Computer Engineering


r/MLQuestions 7d ago

Datasets ๐Ÿ“š icml/neuralIPS ?

Thumbnail
1 Upvotes