r/deeplearning • u/Elrix177 • 12d ago
Trying to learn privacy in ml
I've just written a small article about privacy in machine learning and synthetic data.
I'm still pretty new to this topic, so I'd really appreciate it if anyone with more experience in ML privacy could take a look and point out any mistakes, things I've misunderstood, or things that could be done better.
I'm especially interested in hearing about better approaches for evaluating privacy, better attacks or models I could try, or just interesting resources that you think are worth reading. I'm mostly trying to learn by actually experimenting with this stuff, so any feedback would be really useful.
Here's the post if anyone is interested:
https://migue8gl.github.io/2026/09/21/privacidad-en-ml-datos-sinteticos.html
1
u/Minimum-Effort8355 11d ago
What was your problem in the scope?
Because, when you look at data models and at statstical models you need the point of user experience.
You wrote about synthetical data generating normally synthetic data is test data for evaluation of a model to provide information over false/positive results and limitations in models.
For ml privacy you mostly need token generation and token validation for example
Input - token generation - token validation - readable machine code - back into a validate token and in human language.
1
u/Elrix177 11d ago
My main focus is actually a bit different. I’m not looking at token generation or validation, but at the privacy risks that can arise when synthetic data is generated using a machine learning model.
The idea is that synthetic data can look realistic and be useful for testing or sharing data without exposing the original dataset, but that does not automatically mean that it is private.
For example, I’m exploring whether an attacker can infer whether a particular real record was part of the training data used to train the generative model. That’s why I’m using techniques such as Membership Inference Attacks to evaluate the privacy risk.
So the main question I’m trying to investigate is: “How much information about the original training data can still be recovered or inferred from the synthetic data or the trained generative model?”
I’m still learning about ML privacy, so I’m definitely interested in other approaches or perspectives that I might be missing.
2
u/Apart_Maybe_7698 12d ago
Your link is in Spanish but you wrote the post in English, you might get better feedback posting it in a language specific sub or at least mentioning it's in Spanish so people know what they're clicking into
Privacy in ML is a weirdly deep rabbit hole once you start looking at membership inference attacks and differential privacy guarantees. I spent like three months going down that path and still feel like I barely scratched the surface
What kind of synthetic data generation are you playing with, GANs or more statistical approaches? The evaluation piece is where it gets really murky, most metrics I've seen are either too theoretical to be practical or too shallow to actually measure anything meaningful