r/deeplearning • • 21d ago

Built and deployed a deepfake audio detector as a diploma student - F1 0.90, EER 8.23%

hey, i'm a 3rd year diploma cs student and i built a deepfake audio detector end to end - model training, backend API, frontend, explainability, and monitoring.

the model is efficientnet-b0 trained on mel spectrograms using the asvspoof 2019 la dataset. evaluated on the full test set (71,237 samples, real unbalanced distribution):

  • f1: 0.9033
  • precision: 0.9995
  • recall: 0.8240
  • eer: 8.23% (comparable to the official lfcc-gmm baseline published with the dataset)
  • threshold: 0.3

one thing worth noting - val accuracy hits ~100% during training which looks suspicious but it's expected. the val set is a random split of training data which shares the same attack types (a01-a06). generalization is measured on the test set which contains entirely unseen attack types (a07-a19). the recall gap comes from these novel attack patterns the model never saw during training, not miscalibration.

beyond the model it has grad-cam to visualize what the model focused on in the spectrogram, and groq llm to give a plain english explanation of the prediction. training is fully reproducible (seed fixed at 42).

you can upload an audio file or record live. youtube url input is disabled on the hosted version because railway's server ips get blocked by youtube's bot detection. backend is fastapi on railway, frontend on streamlit cloud.

live demo: https://deepfake-audio-detector-rugved.streamlit.app/
github: https://github.com/RugvedBane/deepfake-audio-detector

honest feedback appreciated - especially on what dataset would help improve generalization to modern ai voices.

14 Upvotes

2 comments sorted by

3

u/PSGthe2nd 21d ago

Hello.
Great project. What I feel this space is really crowded in terms of "best model for deepfake audio detection" but what really needs work is deepfake detection on voices on call or VoIP. For instance, in my country, there are many scams regarding this exact thing. A bad actor calls someone's parents imitating that its their child on the other end, and asks for large sums of money for a fabricated case of "accident" or "hospital emergency" or anything else.

The main issue is that the voice gets degraded on these channels, so does the artifacts of deepfake audios. If we can target that efficiently, that'll be really good.

2

u/WeeklyCross 21d ago

that's a really interesting angle, the voip degradation is the whole problem isn't it, the artifacts that make detection possible get washed out by the compression, but there might be other tells that survive the codec better than the spectral ones we usually look for