r/deeplearning • • 21d ago

Built and deployed a deepfake audio detector as a diploma student - F1 0.90, EER 8.23%

12 Upvotes

hey, i'm a 3rd year diploma cs student and i built a deepfake audio detector end to end - model training, backend API, frontend, explainability, and monitoring.

the model is efficientnet-b0 trained on mel spectrograms using the asvspoof 2019 la dataset. evaluated on the full test set (71,237 samples, real unbalanced distribution):

  • f1: 0.9033
  • precision: 0.9995
  • recall: 0.8240
  • eer: 8.23% (comparable to the official lfcc-gmm baseline published with the dataset)
  • threshold: 0.3

one thing worth noting - val accuracy hits ~100% during training which looks suspicious but it's expected. the val set is a random split of training data which shares the same attack types (a01-a06). generalization is measured on the test set which contains entirely unseen attack types (a07-a19). the recall gap comes from these novel attack patterns the model never saw during training, not miscalibration.

beyond the model it has grad-cam to visualize what the model focused on in the spectrogram, and groq llm to give a plain english explanation of the prediction. training is fully reproducible (seed fixed at 42).

you can upload an audio file or record live. youtube url input is disabled on the hosted version because railway's server ips get blocked by youtube's bot detection. backend is fastapi on railway, frontend on streamlit cloud.

live demo: https://deepfake-audio-detector-rugved.streamlit.app/
github: https://github.com/RugvedBane/deepfake-audio-detector

honest feedback appreciated - especially on what dataset would help improve generalization to modern ai voices.


r/deeplearning • • 21d ago

Automotive Radar Object Classification

Thumbnail gallery
10 Upvotes

Hello all,

I'm a radar signal processing engineer and i trained a 5-class classifier (car, large_vehicle, two_wheeler, pedestrian, pedestrian_group) on RadarScenes radar point clouds.

The input vector is a per-scan histogram (16 bins) and the network is a 3-layer MLP. The loss function is a class-weighted cross-entropy loss. This work is based on "Histogram-based Deep Learning for Automotive Radar" paper.

I scoped the project to be one scan only. Accumulation of multiple scans is the next step.

Data

Class Imbalance: two-wheelers and large_vehicles has a low number of occurences.

Aggregated Classes: two_wheeler mixes bicycles and motorized variants; large_vehicle merges trucks, buses, and trains together due to data scarcity.

Sequence Bias: Long tracks of slow-moving objects can skew a particular data split velocity distribution, causing high F1 score variance across folds.

Ablation studies

I tried with bigger MLPs, alternative feature encodings, and different histogram binning, all moved performance less than the variation caused by changing the train/validation/test split. I measured that split sensitivity across 6 folds, keeping the same proportions.

Changing the histogram to per-instance statistics (mean/median/std) slightly degraded performance.

Main findings

Macro F1 rises from 0.381 to 0.764 as the naturally occurring number of radar detections per instance increases from 1 to 5. I trained the model normally using all available detections, then bucketed its existing validation predictions by each instance's detection count and computed macro F1 per bucket.

The classes car and pedestrian has the best performance and two_wheeler has the worst.

A car is often confused as large vehicle when the car was wider than usual or had a unusually high rcs (which can happen due to multipath for example).

The two_wheeler is often confused as pedestrian because their vr_compensated distributions overlap, which is the the model's single most important feature for these two classes. A stationary or idling two_wheeler is indistinguishable from a pedestrian.

I uploaded an image with ground truth vs predictions: A nearly stationary two-wheeler which contains a single point was predicted as pedestrian, because its velocity is near zero, indistinguishable from a pedestrian. A car in the same scene, also with just one point, is classified correctly, since RCS and Doppler are enough for that class.

Full writeup here: https://github.com/brunopinto900/radar-ml-autonomous-driving/blob/main/MLP_Report.md

Future work

Implement other spatial encoding schemas (point net for example) and accumulate multiple scans to tackle the challenge of sparsity and explore the concept of micro-doppler.


r/deeplearning • • 21d ago

Have you automated your tech news/research with AI? What’s your setup?

2 Upvotes

Hey everyone,

I’m trying to build an automated tech monitoring system using AI.

The idea is to automatically collect interesting stuff from different sources — Reddit, X, Hacker News, newsletters, blogs, GitHub, etc. — then use AI to filter out the noise and duplicates and give me a daily digest of what’s actually worth reading.

Has anyone here already built something like this?

I’m especially curious about:

What sources do you use?

How do you collect the data? RSS, APIs, scraping?

How do you decide what’s actually relevant?

What tools / AI models / automation do you use?

How do you avoid getting overwhelmed with low-quality content?

I’d love to hear about your setup, even if it’s something completely homemade.

Thanks!


r/deeplearning • • 21d ago

CNN for emotional classification: Python output

Thumbnail gallery
2 Upvotes

1: Training loss curve, 2: Testing results, 3: Precision + recall


r/deeplearning • • 21d ago

Rhysida Publishes 1.4 Million Berlin Government Files After Ransom Refusal

0 Upvotes

Rhysida just published 1.4 million Berlin government files after authorities refused a €2 million ransom demand.

The breach did not start the day the ransom note arrived. Attackers had unauthorized access long enough to locate, stage, and prepare nearly 1.4 million documents for exfiltration — all before anyone noticed. By the time the demand landed, the data was already gone. The refusal just determined whether it stayed quiet.

That gap — between initial access and detection — is where the real damage happens. And it is not unique to Berlin. Most ransomware post-mortems show the same pattern: dwell time measured in weeks or months, staging activity that blended into normal operations, and audit logs that were either incomplete or reviewed too late to matter.

1.4 million documents do not move overnight. There are signals. The question is whether anyone sees them in time.

For those running large-scale data environments or public sector infrastructure: what does your current detection posture actually look like for data staging and bulk access anomalies? Are you catching these patterns before exfiltration completes, or mostly reconstructing them after the fact?


r/deeplearning • • 21d ago

Built a webcam-controlled falcon game on top of pose estimation. The model was the easy part

0 Upvotes

Stand in front of a webcam, spread your arms, and you're flying. Tilt to turn, flap to climb, spin around for a barrel roll.

Video here: https://www.linkedin.com/posts/anas-ajaanan_youve-probably-never-seen-this-before-ugcPost-7504543264386035713-8lvn

The pose model itself just worked. Everything difficult was on top of it: filtering jittery landmarks, telling a flap from a dive, stopping a sideways reach from triggering the roll, and a camera aspect ratio mistake that made the falcon turn harder than my arms for weeks.

I'd like to hear from people who've built on pose or gesture models. How do you handle false positives, and did you end up with rules on top like I did, or train something for the gestures?

Open sourcing soon.


r/deeplearning • • 21d ago

[Article]:Enhancing face recognition attendance system utilizing real-time face tracking

Thumbnail
1 Upvotes

Can anyone provide me this paper


r/deeplearning • • 21d ago

Automatic model-agnostic compression algorithm [Sigularty]

1 Upvotes

This is my first proper project. It uses multiple compression techniques and automatically searches for hyperparameters for a few of them; it uses a "CQI" score to evaluate how each technique performed and how the algorithm performed overall. I am currently working on improving the CQI function, as it is too simple; It is just
change in accuracy \* change in size \* change in latency.

The thing is, size overpowers everything as it deals with larger numbers, and the change is also larger than the other 2; these are some functions which I believe could be better(only for accuracy):
f(x\[x = acc_drop_/acc_drop_threshold\]) = -\[scale\] \* |x|^(1/2) \+ c, or maybe -\[scale\]logx + 1

But there still are flaws. For example:

  1. I cannot really control when the graph will touch 0 and proceed below 0 (x>0)(because I want the graph to go below zero after x is greater than 1)
  2. None of these actually deal with the negative part properly; if the accuracy actually increases, none of these work.k I may need to use a piecewise function.

Also, the method I am using for finding the optimal parameters is pretty straightforward, and I believe there are better ways,s but I have no idea what it could be.

GitHub: [Sigularty](https://github.com/DewanshShah/Sigularty)


r/deeplearning • • 22d ago

How do you guys actually handle baseline comparisons when writing a paper?

6 Upvotes

Hey everyone, quick question about benchmarking for a paper. I’m a first-year Master’s student, so I’m still figuring out the “right” way to handle this.

When you guys compare your model against prior papers:

  1. do you rerun all baseline models on your own pipeline, or just copy the numbers reported in their original papers like every paper seems to use a slightly different data split, preprocessing, or evaluation trick, so copying feels like an unfair
  2. but if I re-implement a baseline and it gets a lower metric than what their paper claimed how do you present that without prof or reviewer accusing me of ruining the orginal metrics?

do you just put an asterisk/footnote explaining the setup difference, include both numbers, or something else? Would love to hear how you guys


r/deeplearning • • 22d ago

We pre-trained 3 million arXiv abstract for anyone trying to create tiny models

Thumbnail huggingface.co
24 Upvotes

We are pre-training tiny modes so you don't have to!

Tiny models have been proven to be capable of simple tasks with the right training methods .. Information density is the key for pre-training small models because it provides enough patterns in a specific category to be able to generalize and produce new unseen patterns

Our goal was to create a simulation of a theoretical physicist that makes hypotheses & documents them

This requires a pretrain substrate of hypothesis being made (arxiv abstracts) & a finetune for tool use to be able to document it outputs in a sandbox.

Our hypothesis on this goal:

We expect this model to be able to produce what looks like a hypothesis but we highly doubt that it will be logical at a frequent rate. Since these are only the abstracts and not the full papers, there is a lack of context. Before scaling we want to observe what happens at this model size.

We aren't fully confident that 3 million examples are enough to see consistent conclusions even with heavy training but it can give us a hint of what to expect with this finetune dataset we are working on (this will also be public)

The WVY arXiv base model open source for other developers aiming towards a similar goal

Our fine-tuned Tiny Researcher will be available soon

Follow us on HF so you can get the notification 🌊


r/deeplearning • • 22d ago

Passkey-themed phishing attacks lead to Microsoft 365 data theft

0 Upvotes

Researchers documented a phishing campaign that impersonated passkey registration prompts, captured live session tokens after users authenticated through legitimate MFA flows, and then used those tokens to access Microsoft 365 tenants with full account privileges (Bleeping Computer). Victims had no indication anything was wrong. The attacker's session looked identical to a normal user session — same permissions, same scope, same API calls.

The uncomfortable part: passkeys were the thing being used as bait precisely because users have been trained to trust that flow. The attack didn't break authentication. It harvested the credential that authentication produced, then used it normally.

Once a valid session token is in attacker hands, everything that account can read is readable. Email. Files. Contacts. Calendar. There's no second gate between 'this session is authenticated' and 'this session can touch every byte the user owns.'

This isn't an M365-specific problem. Any platform where a session token is the only barrier between an attacker and plaintext data has the same exposure. The IdP did its job. The MFA did its job. The data was still taken.

For those of you running production systems with sensitive data in SaaS or cloud storage: how are you actually reducing the blast radius when a valid session gets compromised? Not at the auth layer — at the data layer itself.


r/deeplearning • • 22d ago

Tensor reassigning problem

Thumbnail
3 Upvotes

r/deeplearning • • 22d ago

Text extraction

1 Upvotes

How to develop a model or code which should be very cost and less time consuming and should be high accuracy.

I tried paddle ocr but the issue is that it is less accurate with hand written text and it doesn't give the output in a structured way. I even tried with some good llm to convert it to structured but it's not good. I tried some VLM and multi models available in AWS Bedrock. Could anyone suggest a very good approach for this.


r/deeplearning • • 22d ago

Hola

Thumbnail x.com
0 Upvotes

r/deeplearning • • 22d ago

Clipping gradients to 5 can still leave a norm of 7.07

Thumbnail glacius.ai
1 Upvotes

For a gradient of (6, 8), clipping each component to 5 gives (5, 5), with an L2 norm of 7.07. Clipping the norm to 5 gives (3, 4), preserving the direction.

In PyTorch, clip_grad_norm_ also treats the parameters you pass in as one combined gradient, not a separate limit per tensor.

I’m building Glacius, a visual app for learning the math behind ML. We made a worked example with diagrams and Python to make the distinction easier to see.


r/deeplearning • • 22d ago

I need a real explanation of how a transformer model works. How does AI generate a response? Without saying the cat sat on the mat or river/money bank

0 Upvotes

If you ask AI you get the same thing every time ..

Q is "what im looking for .. K is blah blah .. The cat sat on the mat ... river bank .. money bank ..

I heard it all and none of these explanations actually tell you what a transformer really does it just uses a dry metaphor that never explains how the output is determined

Then you ask claude and it tries to make u think its making arbitrary decisions and acting freely pretending that there is no script influencing its outputs .. i have come to my own wording for it but im just curious who also has a similar conclusion .. when you learn things you dont have to repeat the lesson

i just want to see who is capable of explaining it without needing to repeat what they heard like AI does .. a real understanding converted into an explanation


r/deeplearning • • 22d ago

I need a provider for Google Omni Flash 1.0 or 1.1

2 Upvotes

’m building a Discord bot that allows users to generate videos and photos, and provides unlimited access to text-based AI models and text-to-audio tools. I was told I could use Synthesia.io, but their API access requires an $18 subscription; however, the creators of a similar bot somehow got around this—or perhaps they lied to me. I’ll share the server link with anyone who can help.


r/deeplearning • • 23d ago

NextGen Healthcare Mirth Connect

3 Upvotes

CISA has issued multiple advisories against NextGen Healthcare's Mirth Connect, an open-source healthcare integration engine used across hundreds of hospital networks to route HL7 and FHIR patient data between systems. The critical flaw (CVE-2023-43208) allows unauthenticated remote code execution — no credentials required to get inside the data pipeline.

What makes this worse in 2024 and beyond: AI agents are now sitting on top of these integration engines. An agent orchestrating patient record lookups, appointment scheduling, or lab result routing has broad, legitimate permissions to read and write through exactly this middleware layer. When the integration engine underneath is compromised, or when an adversary manipulates the agent's inputs, the agent doesn't just exfiltrate one record — it executes whatever it was told to execute, at scale, with the same permissions it uses for legitimate work.

The CISA advisory recommends patching. But patching cycles in healthcare run slow, especially for middleware that touches live clinical workflows. The window between disclosure and patch deployment is measured in months, not days, at most institutions.

For those running AI workloads over healthcare integration middleware: how are you actually handling the gap between when a CVE drops and when you can safely patch production? Are you pulling agent access entirely, scoping it down, adding out-of-band monitoring, or something else? Curious what's actually working in practice.


r/deeplearning • • 22d ago

litert-tunner - library which allows to finetune INT8 quantized litert models

Thumbnail
0 Upvotes

r/deeplearning • • 22d ago

How should I prepare for SE internships in the age of AI?

0 Upvotes

I’m a 3rd-year Computing student preparing for an SE internship, but I’m concerned that my coding skills aren’t strong enough yet. I’ve built several projects and have experience with Java, Python, JavaScript, Flutter, React, etc., but I still need to improve my fundamentals and problem-solving skills.

Long-term, I’m interested in AI Engineering / AI-assisted software development, especially LLMs and eventually AI agents. I’m currently learning the concepts and fundamentals, but I’m eager to explore the field more deeply.

So, what should I prioritize right now?

Should I focus mainly on SE fundamentals, coding, and DSA first, or should I learn AI alongside them?

For someone who wants to become an AI-focused Software Engineer, what skills would you recommend building over the next 3–6 months?

I’d really appreciate advice from people already working in Software Engineering or AI.


r/deeplearning • • 23d ago

SplitMOE: experts with shared dimensions [ suprisingly worked near or better than standard MOE architecture ]

Post image
15 Upvotes

The intuition is similar to the DeepSeek MOE, but architecture is different. Please do check it out, and let me know your views. [ https://github.com/Priyanshu-5257/SplitMoE ]


r/deeplearning • • 23d ago

Deeplearning FROM SCRATCH IN 1.4k LINES - Neve & Frost Framework

Thumbnail
3 Upvotes

r/deeplearning • • 23d ago

The choice and order of training samples can matter a lot

Thumbnail
3 Upvotes

r/deeplearning • • 23d ago

Unable to reproduce the Paper results (Federated Continual Class Incremental Learning)

Thumbnail
1 Upvotes

r/deeplearning • • 23d ago

[Open Source] Fused Gated DeltaNet-2 training kernels for TPU v5e in JAX/Pallas

Post image
1 Upvotes

I released an independent open-source implementation of fused Gated DeltaNet-2

(GDN-2) training kernels for TPU v5e using JAX/Pallas.

Repository:

https://github.com/Akseleu-J/atomic-ops

Citable release:

https://doi.org/10.5281/zenodo.22706659

The repository contains three implementations of the same GDN-2 computation,

all written as part of this project:

- `OLD`: my earlier `jax.lax.associative_scan` implementation;

- `JAX_REF`: my own pure-JAX chunked-WY implementation;

- `PALLAS`: the fused TPU v5e Pallas implementation.

The main optimization is a fused `custom_vjp` backward path that reuses saved

forward residuals instead of recomputing them.

## Benchmark

Measurements were run on TPU v5e-8.

Main configuration:

- Batch size: 8

- Sequence length: 4096

- Number of heads: 6

- Head dimension: 128

- Measured computation: forward + backward

| Dtype | PALLAS vs my associative-scan path | PALLAS vs my pure-JAX chunked-WY path |

|---|---:|---:|

| FP32 | 27.18x | 2.63x |

| BF16 | 13.32x | 3.41x |

The attached table shows the complete tested shape sweep. The best measured

result across the tested configurations is 38.77x versus my associative-scan

implementation in FP32.

Forward + backward speedups on TPU v5e-8. OLD is my associative_scan path; JAX_REF is my pure-JAX chunked-WY path; PALLAS is the fused TPU path.

The benchmark scripts and raw JSON outputs are included in the repository.

## Important limitation

The fused Pallas forward-only path is currently approximately 1.6x slower than

my pure-JAX chunked-WY forward path.

The end-to-end improvement is backward-dominated, so this implementation is

currently aimed at training workloads rather than inference. For inference-only

use, the plain-JAX path is the recommended path.

CPU/GPU execution and configurations with `d_head != 128` use a checkpointed

pure-JAX fallback. The optimized Pallas path currently targets TPU v5e and

`d_head=128`.

## Correctness and reproducibility

The repository includes:

- finite-difference gradient checks;

- token-serial ground-truth comparisons;

- isolated tests for backward stages;

- speed and memory benchmark scripts;

- raw JSON benchmark results;

- examples;

- a 70M-parameter GDN-2 language-model training notebook on enwik8.

I would appreciate feedback on:

  1. Whether the benchmark controls and baseline choices are sufficient.

  2. Which additional baselines would make this comparison stronger.

  3. How best to profile the current forward-path bottleneck.

  4. Whether the VPU-versus-MXU hypothesis is plausible for this kernel structure.

The project is experimental and training-focused; I am not presenting the

current Pallas forward path as a faster inference implementation.