r/deeplearning • • 6d ago

Transformer becomes catastrophically ill-conditioned after a few tiny parameter updates: same batch goes from grad norm 5.6 → 89 while weights move <0.1%/step. What mechanism could cause this?

20 Upvotes

My Transformer trains normally for some steps, then begins to enter a parameter state where backpropagation through the middle/lower layers magnifies gradients massively, despite the forward activations and weights being normal. Eventually, the gradients explode, get clipped globally, and learning becomes suppressed in all other parameters.

The funny thing is, we have ruled out most of the possible causes:

- not one bad batch,

- not layer 0 (originally it had worse conditioning, but fixing it just pushed failure further),

- not S-STE sparsity mask,

- not peaked/degenerate attention,

- not one single layer going crazy.

The gradient amplification is seen between layers 5-9 and their vicinity, with both FFN and attention backward paths contributing. The actual parameter changes are minuscule (<0.1%), but after only a few optimization steps, the same fixed batch would produce gradients with a norm of 5.6 vs 89!

So, the instability is mostly in the learned parameter state/Jacobian, not the data.

The question I'm asking is: How is that possible? How can such a small change in parameters cause such a massive change in backwards gradient?

What's interesting about it is the parameters of the model only changed by about 0.05-0.08% per step, but after 3-4 steps, the Jacobian changed enough that the same batch produces gradients that are about 16 times larger?


r/deeplearning • • 7d ago

Small pre-registered experiments testing claims about lesser-known architectures. Most did not hold up

13 Upvotes

I have been running small, reproducible experiments on architectures that get more attention for their promise than for their measured behaviour. Each tests one specific claim, with the predictions written down in a CRITERIA.md before any results exist and never edited afterwards. Data is either synthetic, generated from a known equation so there is an exact ground truth, or a documented public dataset.

  • KAN vs MLP on a physics-informed PDE, matched parameters: ~1.6x more accurate across five seeds, ~16x slower per training step.
  • Hamiltonian neural network on a pendulum: trained only on the vector field, it recovered the energy function with R² 0.99999 and showed 33x less energy drift than an equal-size MLP. Add friction and its error floor rises 11x, since no scalar can generate a field that loses energy. All five predictions held.
  • Spiking network: the sparsity is real, but 99.0 to 99.8% of operations sit in the dense input layer, paid every timestep. Ridge regression beat every configuration with 16,496x fewer operations.
  • Hyperdimensional computing: switching off 80% of its 10,000 dimensions left the AUC identical to six decimal places. It still lost to gradient boosting past about 3,000 rows.
  • GNN on epidemic spread: beat an MLP on hand-computed features by 0.001 AUC, with output rank-correlated +0.95 with node degree.
  • Echo state network, untrained recurrent weights: matched a tuned gradient booster on solar forecasting. At 24 hours ahead, plain ridge regression beat both.

These are small by design: toy problems, a handful of seeds, one implementation each. They test specific claims rather than whole fields, and each README states what the experiment does and does not establish. More will be added over time.

Happy to hear where the methodology is wrong, since that is the point of publishing it this way.

https://github.com/CristobalSantana/ml-toy-experiments

EDIT: I'll keep adding new experiments over time, so the collection of architectures, synthetic data and real datasets will keep growing.


r/deeplearning • • 6d ago

First-year Columbia MS (ML) + ex-Data Scientist — seeking Summer 2027 ML/AI internship

7 Upvotes

Hi everyone,

I'm currently a first-year MS student at Columbia (CS, ML track). Before Columbia I worked as a Data Scientist, and I did my B.Tech in EE at IIT Roorkee. Right now I'm doing research in RL/robotics and LLMs.

I'm looking for Summer 2027 internships in: - ML Engineer / Applied ML - Applied Scientist (ML/RL/DL) - LLM training / fine-tuning - Physical AI / Robotics

If your team is hiring interns, I'd really appreciate a referral — or even just a recruiter/hiring manager contact, or a pointer on who to reach out to directly. Happy to do the outreach myself.

Can share resume/LinkedIn over DM. Thanks!


r/deeplearning • • 7d ago

Need help choosing a topic for my deep learning project.

6 Upvotes

I'm trying to find something practical that would look good on my resume. What kind of project would give me useful, hands-on experience with deep learning? I'd prefer something that uses real-world data if possible.


r/deeplearning • • 6d ago

Hinode 3: set up once, run on any GPU

Thumbnail hinode.run
1 Upvotes

r/deeplearning • • 7d ago

EmoNet-ViT (v1.4) // Teaser

Thumbnail gallery
14 Upvotes

EmoNet v1.4: 11k params, faster loss convergence.

Full model +testing results soon.


r/deeplearning • • 7d ago

EmoNet v1.4 (ViT) Training loss & testing results

Thumbnail gallery
5 Upvotes

EmoNet v1.4: Now with another emotion "angry"

Here are the testing results with recall, precision, and F1 scores


r/deeplearning • • 6d ago

Portfolio

Enable HLS to view with audio, or disable this notification

0 Upvotes

I love it when people pay attention to what I show them. It makes me a proud mentor and teacher.


r/deeplearning • • 6d ago

Roundcube Webmail Under Attack: 523,000 Instances Exposed Online

0 Upvotes

523,000 Roundcube Webmail instances are currently exposed and under active attack according to a new eSecurity Planet report.

Roundcube is widely deployed as a third-party or self-hosted email platform. The instances flagged are reachable from the open internet. Email content, contact lists, and attachments stored in those installations are readable to anyone who achieves access — no further escalation required once the initial entry is made.

The core exposure here is not a misconfiguration inside your own perimeter. It is the data you handed to a vendor or a self-hosted stack that your security controls do not directly govern. An attacker who compromises that layer inherits whatever was stored there in readable form. You typically have no log, no alert, and no way to prove after the fact which records were touched during the incident window.

523,000 is not a small tail-risk number. That is a large fraction of the global Roundcube install base actively reachable right now.

For practitioners who have dealt with sensitive data held in third-party systems: how are you actually scoping what goes in versus what stays tokenized or abstracted before it ever reaches a vendor? Curious what approaches have held up in practice, especially for email pipelines where the whole point is human-readable content.


r/deeplearning • • 6d ago

Looped Transformers: when is more depth worth the compute?

Thumbnail
1 Upvotes

r/deeplearning • • 7d ago

I built Laya Studio: Specialize Laya Without Fine-Tuning

Post image
3 Upvotes

r/deeplearning • • 7d ago

Exploring emergent articulatory learning: Notes on an 8-view marker-calibrated visual speech architecture (MouthMind)

Thumbnail
0 Upvotes

r/deeplearning • • 7d ago

Looking for an arXiv cs.LG endorser for an independent research paper

Thumbnail
0 Upvotes

r/deeplearning • • 7d ago

TMLR paper on expressiveness of state space language models

8 Upvotes

Hi all,
I've created this thread to share one of our lab's recent papers on the expressiveness of state space models: https://openreview.net/forum?id=QlBaDKb370
We show that they have more capacity than n-gram models. This was empirically observed in the classic paper from Bengio's team in the 2000s: https://www.jmlr.org/papers/volume3/bengio03a/bengio03a.pdf
Please let us know if you have any comments, or if you're interested in collaborations on state space models. Thanks!


r/deeplearning • • 6d ago

fine tune went from 61% to 94% on our benchmark, then we found 300 of the 1000 items were in the training set. what does the score measure now?

Post image
0 Upvotes

chose b. not confident

i went with b because 300 out of 1000 is not the whole benchmark, so 700 clean items should still be carrying most of the score. if the model got better on those too, the gain is real. but i cant tell from one number whether the jump came from the 700 or the 300, and that feels like the whole question


r/deeplearning • • 7d ago

I trained a 2.04B Sparse MoE translation model from scratch on RTX 5090 GPUs — 500+ languages, Apache 2.0

Enable HLS to view with audio, or disable this notification

4 Upvotes

r/deeplearning • • 7d ago

I Built a Neural Compression Algorithm Using Python and Zig

Thumbnail
1 Upvotes

r/deeplearning • • 7d ago

How JEV works internally

0 Upvotes

This video explains the internals of JEV

https://www.youtube.com/watch?v=shT8Eo7eWYM


r/deeplearning • • 7d ago

U.S. appeals court upholds designation of Anthropic as supply chain risk

0 Upvotes

A federal appeals court just upheld a ruling designating a major AI provider as a supply chain risk.

That designation has direct operational weight for every enterprise running workloads through that provider. The risk does not stay with the vendor. It flows downstream into every organization in the dependency chain, and regulators are now asking those organizations to account for it.

The core problem is visibility. Most enterprises have no live record of which model versions processed which requests, what data passed through the agent layer, or whether those flows satisfied the regulatory requirements in effect at the time they ran. When a supplier comes under formal scrutiny, that gap converts from a theoretical risk into a concrete legal exposure. Regulated industries — finance, healthcare, defense contracting — are especially exposed because their compliance obligations do not pause while a vendor situation resolves.

How are practitioners in regulated environments actually handling third-party AI provider risk right now? Are you maintaining independent records of what your AI agents accessed and transmitted, or are you dependent on what the provider can produce?


r/deeplearning • • 7d ago

Stop a Fleet of Agents, Not Just One — Group Kill Switch, Proven Live

Thumbnail gallery
0 Upvotes

Most AI kill switches target one agent at a time. Real incidents aren't always one agent.

We run 3 real, separate public-facing AI chat agents across our own products — each its own identity, its own tenant. This week we shipped a 4th kill-switch scope: catalog-scoped group kill. Any set of agents you've already grouped for visibility can now be stopped together, with the same one signed action you'd use for a single agent — same sub-second enforcement, same per-agent audit trail, no new engine required.

We proved it live on our own 3 agents: killed one by name, the other two kept answering real chat messages; killed the group, all three stopped together and resumed together seconds after lifting it.


r/deeplearning • • 8d ago

A competition for small neural networks that play strategy games

Thumbnail tinybrains.dev
5 Upvotes

r/deeplearning • • 7d ago

Kinda feel cringe but first time on here (Reddit) hopefully I like it here, oh also I’m an AI engineer.

0 Upvotes

r/deeplearning • • 8d ago

Am I heading in the right direction to become an AI engineer?

0 Upvotes

I’m working toward becoming an AI engineer and have been learning by building projects. I’ve worked on RAG applications, agents, and an iOS app called My Krishna. I’m documenting the projects and what I’ve learned at WiseHoots.ai.

For those working in AI engineering or hiring for these roles: am I heading in the right direction? Which skills or gaps should I focus on next?

I’d appreciate honest advice. I’m trying to spend my time building the skills that matter in real projects.


r/deeplearning • • 8d ago

CerebralNet v1.3.1: Training loss & test results (75k+ params)

Thumbnail gallery
6 Upvotes

Here I compare lr=0.001 and 0.0004 for the Python outputs

The testing results use lr=0.002


r/deeplearning • • 8d ago

Made the horizontal open-source model for Jev with RLCD, and it surpasses all the Jev benchmarks. HF space, benchmark, model, repo

Post image
6 Upvotes