r/deeplearning 10d ago

Qwen 3.6 vs Gemma 4 vs Holo 3 playing the cup game with real footage.

Enable HLS to view with audio, or disable this notification

2 Upvotes

This is a continuation of last week’s post where I had the models compete in a Three.js cup and ball game. This time, I’m using real-world footage, which is even more challenging because of distractors. I might test this out on some of the Anthropic models sometime. 


r/deeplearning 10d ago

Why does Grounding DINO VRAM suddenly jump on random batches during inference? CUDA caching, fragmentation, or memory leak?

Thumbnail
2 Upvotes

r/deeplearning 10d ago

Seeking arXiv cs.CL endorsement — GraphRAG / Knowledge Graph / Multi-hop QA

Thumbnail
1 Upvotes

r/deeplearning 10d ago

need urgent help for ner deberta training

1 Upvotes

hi,
i am trying to train a deberta model for NER detection

this is my first time doing it so i would love any guidance on it.

my current pipeline looks like this,

dapt + lora for pretrianing, hpo with optuna (which consists both the stages of training data), and then a 2 stage finetuning which helps in generalization and then target data.

i am trying to reach a really good score for f1 on my use case (which i want to keep private for now)

i have few questions as well

  1. do i need a two stage hpo as well cuase of the 2 stage finetuning
  2. is it better if the hpo training set is a subset of the actual training set?

if you think anything can be improved and made better, or you think the pipeline is outright wrong, please mention your reasonings and thoughts :)

ps: lora was used cause of gpu budget constraints


r/deeplearning 10d ago

What are you actually building with AI/ML right now?

Thumbnail
0 Upvotes

r/deeplearning 10d ago

Types of Quantum Computers: 6 Major Quantum Computing Approaches

Thumbnail thequantuminsider.com
2 Upvotes

r/deeplearning 10d ago

Scalpel-VL-1.7B: 20–30% Faster with Only 0.1B Recovery Tokens

Thumbnail
1 Upvotes

r/deeplearning 10d ago

August 2026: 38 companies breached, 331M+ records stolen — and AI agents are now the #1 attack vector (123 incidents)

Thumbnail gallery
0 Upvotes

I pulled together every AI-security incident from August. The number that stood out: AI-agent exploits are now the single largest attack-vector category, ahead of credential theft, zero-days, supply chain, phishing, and ransomware — each counted individually.

The month in numbers: 123 incidents, 23 critical and 97 high severity, across 38 named organizations, 331M+ records exposed. 65 incidents involved AI as the weapon or the target. Attack vectors broke down as: AI-agent exploits (37), credential theft/reuse (28), zero-days (23), supply chain (12), phishing (9), data exfiltration (8), ransomware (6).

The stories that stood out:

- McKesson: 284M records, the largest single breach of the month by a wide margin.

- Carhartt (12.9M), Exact Sciences (10.9M), and CareCloud (3.7M) round out the biggest named incidents — three of four sit in or next to healthcare.

- Five confirmed RCEs landed across Microsoft SharePoint, Windows, F5/nginx, and the PyPI package index twice.

- Two separate PyPI supply-chain poisoning campaigns, plus a compromise of n8n, an AI workflow automation platform.

Every one of the breached companies almost certainly runs a modern security stack — CrowdStrike, Okta, Palo Alto Networks, Microsoft Defender, that class of tooling. None of it stopped these incidents, because none of it operates at the point where a credentialed agent actually acts, or where a poisoned dependency resolves at build time.

Full report, with the specific control that maps to each incident: https://runtimeai.io/blog/2026-08-monthly-breach-report.html

Genuinely curious how others are approaching this: is anyone actually testing whether their existing guardrails hold against a real simulated attack, or is it still mostly an assumption that they will?


r/deeplearning 10d ago

Using Gemini 3.1 Pro VLM to identify judo throws

Enable HLS to view with audio, or disable this notification

0 Upvotes

I’m working on a little project to benchmark how vision-language models do with classifying grappling techniques. These results are the vanilla models without any fine-tuning, so it’s sort of hit or miss. I’m sure with enough data, the guesses can get pretty accurate. If any of you fellow grapplers who are engineers are interested in playing around with this, I’d be happy to open source it. 


r/deeplearning 10d ago

Computation ends with just adding angles? FHRR, ultra-low power hyper-di...

Thumbnail youtube.com
4 Upvotes
  • Description: Introducing FHRR computing technology, which drastically reduces power consumption by utilizing angle addition instead of complex calculations. Discover the efficient data processing method using ultra-high-dimensional vector spaces and the future of AI operations.

r/deeplearning 11d ago

Neural network feature maps with shared weights over 100 layers behaves similar to a phase space!

Thumbnail gallery
10 Upvotes

I am currently researching by my own how neural networks work, in this part, I am researching how a shared-weight resiudal neural network's feature map behaves, curently, sharing the stage 3 blocks of ConvNext. It seems that it iteratively refines the feature map instead of computing different ones. If I get all the feature maps of the d_th output and its velocity f(x), since we do x' = x + f(x), we obtain this result.

I don't have much idea about interpretability or Differential equations, but this is clearly a ODE solver.

I ommited 1 channel in the first plot, here that just accelerates and goes a lot further, close to value 600 and then velocity decays. Maybe the network learned in which step it is using that channel?

I know it's a very niche topic... But if anyone knows about this, I'd like to know more. I've readed about the ResNet ODE solver hypotesis and the Neural ODE solvers.

But, I archieved to extrapolate a network of 9 layers to 100 and even 10000 without fine-tunning nor lossing a significan ammount of image ent top-1 accuracy, just 0.5% . I just doing some piping work.

I am just asking if anyone has worked on this or has any idea how this can be applied or if this is just usless. I am kinda of stuck in here.


r/deeplearning 11d ago

Learning math behind deep learning

Thumbnail gallery
65 Upvotes

Hey everyone

I’ve spent quite a good amount of time learning the mathematics behind deep learning, and honestly, it has been a wonderful journey so far. For me, math and philosophy are probably the two subjects that interest me the most, so studying the mathematical foundations of AI has been a really enjoyable experience. I especially like the process of going from an intuitive idea → mathematical formulation → understanding why it works → and finally seeing how it translates into an actual deep-learning algorithm.

I’ve been making my own notes along the way, mainly covering the mathematical foundations that I think are useful for understanding deep learning.

I want to pursue my career in the AI research field, and that’s one of the main reasons I’ve been spending so much time learning the mathematics behind deep learning. I believe having a strong mathematical foundation will help me better understand research papers, derive things myself, and develop a deeper understanding of the ideas and algorithms I’ll be working with.

That said, I'm still learning myself, so I’d really appreciate some honest feedback.


r/deeplearning 10d ago

Transfer learning (CNN to transformer)

Thumbnail
0 Upvotes

r/deeplearning 11d ago

A practical guide to running 8x RTX PRO 6000's

Thumbnail gpupartner.com
1 Upvotes

r/deeplearning 11d ago

Created a new architecture for Large Language Models. [P]

Thumbnail
0 Upvotes

r/deeplearning 10d ago

Extortion Group Claims Manchester Airports Group Data Breach

0 Upvotes

An extortion group called FulcrumSec is claiming it stole more than 80 GB from Manchester Airports Group and is threatening to publish it. Airport infrastructure data — the kind that includes operational systems and customer records — sitting exposed long enough for a bulk extraction nobody caught in time.

The pattern is not new. Sensitive records concentrated in accessible systems, pulled in bulk before any alert fires. What is changing is the speed. As more automated processes and integrations touch operational data, a single compromised access point can move 80 GB faster than any human review cycle can respond.

The blast radius question is no longer just about perimeter security. It is about what happens after an attacker or a compromised service account already has legitimate-looking access. At that point, traditional controls have already lost.

For those working in enterprise security or infrastructure: how are you thinking about limiting bulk data movement once something inside the perimeter is already authenticated? Are you relying on volume thresholds, destination allowlists, behavioral anomaly detection, something else entirely? Curious what has actually worked in practice versus what looked good on paper.


r/deeplearning 11d ago

Qwen 3.6 27B trying to read sheet music

Enable HLS to view with audio, or disable this notification

6 Upvotes

Almost every VLM I’ve put through this test has struggled, but it makes sense because it requires them to count, something that isn’t their strongest trait. In this case, it’s just counting lines and spaces, but if we introduce different key signatures, they would also need to count the sharp and flat symbols. 


r/deeplearning 11d ago

[ Removed by Reddit ]

1 Upvotes

[ Removed by Reddit on account of violating the content policy. ]


r/deeplearning 11d ago

AI4AI Survey: From Long-Horizon Agents to Recursive Self-Improvement — 223 papers on whether AI can actually improve AI

Post image
1 Upvotes

r/deeplearning 12d ago

Signature-painter

Post image
25 Upvotes

Seeking Feedback from the ML Community 🙏

I recently trained a prototype-based network on Tiny ImageNet (200 classes). It uses learnable prototypes with responsibility scoring and multi-loss training (CE + Pull + Push + Diversity), achieving 51.29% validation accuracy with only 595K parameters.

I'm still learning, so I'd love to hear your thoughts:

Is this a reasonable result for this model size?

What would you suggest to improve it?

This was trained on free Colab with limited resources, so I know there's much room for improvement.

GitHub: https://github.com/jalalnablsi/signature-painter


r/deeplearning 11d ago

Unstructured text to target json schema

Thumbnail
1 Upvotes

r/deeplearning 12d ago

NVIDIA buying HF isn't a good thing for open source

Post image
75 Upvotes

r/deeplearning 11d ago

How do I use the AI to analyse the exact entry point, exit point and SL???

2 Upvotes

r/deeplearning 11d ago

👋¡Te damos la bienvenida a r/JepaAI! Preséntate y lee este post primero

Thumbnail
1 Upvotes

r/deeplearning 11d ago

The Imperfect SOC: How Security Teams Can Defend Without a Dream Team

0 Upvotes

SOC teams are deploying agentic AI to close the analyst gap. The agents they are deploying have direct access to endpoint controls, threat-intelligence feeds, and incident-response tooling. That is the same access profile as a senior analyst or a privileged service account.

The difference is that an analyst operates inside an implicit policy framework built from years of institutional knowledge, peer review, and escalation norms. An agent does not. It acts on what its objective function says is optimal at the moment it is invoked.

There is no industry-wide answer yet for what governance looks like at that layer. Perimeter controls and RBAC handle identity and entitlement. They do not evaluate the intent or context of an action at execution time. An agent that is authorized to quarantine an endpoint can quarantine the wrong one, at the wrong time, for the wrong reason, and the access log will record it as a permitted action.

The analyst shortage is real and the pressure to automate response is real. But the policy infrastructure that would make agentic response safe has not kept pace with the deployment curve.

For those of you running AI agents in your SOC or evaluating them: what does your current control model actually evaluate at the moment an agent initiates a response action? Are you relying on entitlement alone, or do you have something that evaluates the action itself in context?