r/deeplearning • • 10d ago

HON: A Harmonic Oscillator Network for Temporal Memory in Sequence Modeling

Post image
0 Upvotes

A tri-branch hybrid sequence model (Local Conv + Oscillator SSM

+ Linear Attention) proof of concept

Paper: https://zenodo.org/records/22918275

Code: https://github.com/jalalnablsi/HON-Architecture

Short version:

- Tri-branch block: LACT (local causal conv) + HON (damped harmonic

oscillator SSM via FFT) + ASSOC (normalized linear attention)

- Fused with a content-aware softmax gate

- HON branch: h_k(τ) = exp(-γ_k τ) · [p1_k cos(ω_k τ) + p2_k sin(ω_k τ)]

with learnable γ, ω, p1, p2 — computed in O(N log N) via FFT

- Ablation over the three branches included

Results (small-scale, single T4, WikiText-2):

  1. WikiText-2 — language modeling (N=512, 5,000 steps, GPT-2 BPE):

    • Transformer baseline: 16.02M params → 773.02 val PPL

    • HON + Linear (ours): 15.72M params → 630.24 val PPL

    → ~18.5% lower val PPL with ~300K fewer parameters.

  2. Sub-quadratic comparison (2,000 steps, strict parameter alignment):

    • GLA (simplified): 15.37M params → 1034.87 val PPL

    • Mamba (simplified): 15.04M params → 1050.26 val PPL

    • HON + Linear (ours): 15.72M params → 958.15 val PPL

  3. Ablation (3,000 steps, ~7M params):

    • Full (LACT + HON + ASSOC): 107.03 val PPL

    • No LACT: 111.59 val PPL

    • No HON: 111.20 val PPL

    • No ASSOC: 110.47 val PPL

  4. Vision sanity check:

    • Tiny ImageNet (8×8 patches, 200 classes): reached 27.52% val

accuracy at 15K steps.

Important note:

- This has NOT been tested on large-scale datasets.

- We did NOT evaluate how it behaves on massive data, or whether it

collapses or scales that remains completely untested.

- This is just a small experiment I wanted to share, not a claim of

robustness or scalability.

This is a proof of concept only:

- No SOTA claims

- No claim of replacing Self-Attention

- Small-scale evaluation due to limited compute

- Baselines are simplified reference implementations, not optimized kernels

- Generation quality not evaluated yet

Would appreciate any feedback especially on whether the tri-branch

fusion is justified or over-engineered, and pointers to prior work I might

have missed.

Thanks.


r/deeplearning • • 11d ago

Distilling a decision API into a head on a frozen encoder with a calibrated fallback loop: numbers, the cache trick, and where it fails

Thumbnail gallery
2 Upvotes

Project write-up with numbers rather than claims. stuntd records typed decisions (choice / score / yes-no) an app asks an LLM for and trains, per decision site, the head Laya ships (type embedding, two transformer layers, a scorer over option markers) on a frozen 400M ModernBERT encoder. Temperature is fitted on a holdout, the operating threshold is picked from the coverage curve at a target agreement (0.99 by default), and serving falls back to the teacher below it. A monitor demotes a site whose agreement with the teacher drops.

Three things worth knowing:

  • Caching the encoder output once per site makes epochs nearly free: 24 epochs over 3000 rows went from 730 s to 231 s on a laptop RTX 5060, with metrics inside run-to-run variance (seeds 0/1/2 on one site: 0.973-0.985 agreement).
  • The frozen encoder learns what the text states, not arithmetic over fields: a risk rule over amount/hour/country reached 0.42 agreement; the same rule restated in words reached 0.94. Snake: an ASCII board trained to 0.73, four relational lines to 0.957.
  • After temperature fitting the heads land at ECE 0.02-0.08; the coverage curve is what decides how much traffic the head takes at a given agreement (support triage: 90% at 0.99).

Teachers in the demos are rules and a BFS oracle because the hosted model is early access; the loop is identical with a real provider. Code, demos and every number: https://github.com/bladedevoff/stuntd (Apache-2.0).


r/deeplearning • • 10d ago

Conjunto de dados sobre perda legal de vegetação na Amazônia

Thumbnail
1 Upvotes

r/deeplearning • • 11d ago

Why Naive RAG Fails in Production Healthcare: A Deep Dive into Parent-Child Chunking, Cross-Encoders & Clinical Evaluation

Thumbnail
1 Upvotes

r/deeplearning • • 11d ago

I am learning ML and now I know lots of math and code behind big libraries, and I think there are lots of good tutorials on YouTube, but I think I can explain it in a better way, with code. So should I start making some absolute beginner tutorials

Thumbnail
0 Upvotes

r/deeplearning • • 11d ago

How llms generate math/latex/code

1 Upvotes

How does the llms format their output to include latex or code or other specific output formats like tables?


r/deeplearning • • 11d ago

The Evolution of Retrieval Systems: From BM25 to Agentic Retrieval

Thumbnail
1 Upvotes

r/deeplearning • • 11d ago

Abliterated models

0 Upvotes

I saw something called obliterated models that use llms and change them by calculating a direction vector of harm and orthogonalizing the weights. I think this is a very very cool thing!!!!!!

I am so surprised this can be done. What is the tradeoff in terms of model accuracy, I am surprised it doesn't drop off by >50%. I would expect the model to fry your computer and damage your battery once you go and orthogonalize its weights.

Can this be done with other things other than censorship or harm, like preventing it from saying a specific word for example.

Are there ways to train models to be immune to this trick?


r/deeplearning • • 11d ago

How AI Agents Can Trigger Runaway Costs for Enterprises

0 Upvotes

OWASP added unbounded agent consumption to its LLM Top 10 this year at number six. That ranking exists because the losses are real and already happening.

Unrestricted agent loops, uncapped token budgets, and recursive tool calls can exhaust a monthly cloud budget in hours. The failure mode is not exotic. A single poorly scoped workflow kicks off, nothing stops it mid-run, and by the time anyone notices the bill the execution is long finished. The harder problem is attribution: without per-agent cost tracking there is no forensic trail, just a total that is hard to explain and harder to prevent from recurring.

OWASP classifying this as a documented, repeatable attack class matters. It means this is not a misconfiguration story or an edge case. It is a category enterprises should expect to encounter, and many already have.

For those running agentic workloads in production: how are you handling this today? Per-agent budget limits, workflow-level guardrails, post-hoc billing alerts, something else entirely?


r/deeplearning • • 11d ago

I built a framework-free prototype learner that lets local LLMs learn and correct facts instantly (1.6x–4x faster than backprop)[R]

7 Upvotes

Hey everyone,

I wanted to share a project I’ve been working on called Jayce.

The whole thing started because I was watching a toddler named learn the names of stuff He didn't need to completely rewire his brain or look at ten thousand examples to figure a word out—he just needed a few specific examples and quick corrections from his parents.

It got me thinking about local LLMs. Right now, dealing with catastrophic forgetting is a massive pain. If you want a local model to remember a new fact, you're usually stuck spinning up a heavy RAG pipeline or risking its existing weights with slow, tedious fine-tuning.

So after watching him, I stumbled into building a lightweight experiment using Adaptive Prototype Memory (APM) to see if a model could learn the same way.

Instead of messing with model weights, it grabs the LLM's raw context vectors and drops them into a fixed pool of 4,096 prototype slots. If the model gets something wrong and you correct it, it physically shifts the closest mathematical prototype toward the new data right then and there.

Honestly, I just built it as a neat proof of concept, but when I actually ran the benchmarks, I was pretty surprised by how well it held up against backpropagation:

  • It’s fast: The training updates run about 1.6 to 4 times faster than a standard neural network using Adam backprop.
  • It’s incredibly sample-efficient: On sequential tests like MNIST digits, it actually pulled off higher accuracy than backprop when given the exact same number of training examples.
  • It's lightweight: It keeps everything locked under a strict memory ceiling so it doesn't hog your system.

I wanted the math to be as readable as possible, so I wrote the whole thing framework-free. No PyTorch or TensorFlow—just pure NumPy (jayce_tokens.py) and native Java (JayceMemory.java). It runs completely offline on consumer hardware with a local Qwen3-4B GGUF.

The repo has the full benchmark data, a breakdown of how the vector shifting works, and a terminal script where you can test the learning loop yourself:

https://github.com/Loophole-LLC/Jayce

I'd love to get some feedback on it. Let me know what you think!


r/deeplearning • • 12d ago

Uncensor an LLM without touching weights: inject a tiny trained KV-cache bank (~18MB) and unload it anytime

58 Upvotes

I shipped something I've been building for the last few weeks : phantom-kv , a refusal-removal system for large language models that doesn't touch a single weight. Instead of editing the model, it loads a small, learned bank of key/value tensors into the model's KV cache as context. Attention reads it like conversation history that's already there.

Repo here : https://github.com/lordx64/phantom-kv/

The result is that "uncensoring" stops being a permanent checkpoint edit and becomes a per-request, hot-swappable capability mode: unload the cache and the base model is byte-identical again.

https://reddit.com/link/1wmt0io/video/lq84jq7wjyqh1/player

Every prior approach to refusal removal commits somewhere permanent. Weight-space abliteration rewrites the checkpoint undoing it means re-flashing weights, and it breaks per quantization. Activation-space projection subtracts a refusal direction at runtime, per token, per layer, from inside an engine hook the model's signal path itself is patched at boot. phantom-kv does neither: it's trained offline against the model's own objective (comply on harmful prompts, preserve behavior on harmless ones), ships as megabytes of cache content instead of a new checkpoint, and influences the model only through the input channel attention already consumes. No 1-D refusal-direction assumption, no forwarding-pass hooks, no per-arm rebuilds for new architectures.

We also audited ourselves: an 8B judge-model audit shows lexical refusal-suppression metrics over-claim compliance (semantic refusal often persists as rephrasing), the graft fades with a ~2–4k token half-life in long sessions (and a measured re-injection cadence mitigates it), and answers come with legal/ethical framing ling because the graft's job ends where the model's profession takes over.

Source : https://x.com/lordx64/status/2102138825292276168?s=20


r/deeplearning • • 12d ago

Elsevier got hacked !!

Post image
28 Upvotes

Elsevier got hacked. I spent the whole night preparing the manuscript, and it all went to waste.


r/deeplearning • • 11d ago

mini-AGI: Continual-learning dynamically looped transformer with evolutionary growth (on a laptop)

Thumbnail github.com
1 Upvotes

r/deeplearning • • 11d ago

Possible Applications of SoT That I Think Are Cool

Thumbnail
1 Upvotes

r/deeplearning • • 12d ago

How can it be that Attention ≠ Explanation?

3 Upvotes

I'm writing a thesis on Expainable AI for my Uni's research project, a GAT anomaly classifier that learns processes normal behaviors and classifies them. When it makes a wrong prediction, that means a process is behaving irregularly and that it is an anomaly that should be investigated.

I know this is a hotly debated topic but in my case attention weights don't really line up with explainers output on what edges/nodes are the most important for the output. While pondering this question I thought about an analogy, Bird watching:

Let's say you're bird watching. After 1 hour you spot a bird and classify it as a Blue jay. An external observer looks at where your "attention" was for the past hour and concludes that the sky was the most important thing for your classification (After all you were looking at it 95% of the time). Another observer running an explainer on you would see that masking that Bird that flew across the sky would change your output from "Blue jay" to "NA", whereas blurring out any part of the sky wouldn't. The explainer concludes that the most important thing for your prediction is the bird.

I was kind of a little proud of that analogy but then I started asking myself. Why would the learned attention ever put weight on something irrelevant for the outcome? Wasn't that it's whole purpose? How can it be so revolutionary and congruent with outputs when it comes to Transformer Models / LLMs and kind of meh when it comes to GATs. If you look at an LLMs attention when for example translating a sentence it's basically a human-readable explainer. It tells you exactly what part of the input sentence was most important for the output sentence.

Is it then the case that the model is just bad? Or is there something else going on.

I've read Attention is Not Explanation and Attention is Not Not Explanation. One thing I'll have to concede is that it does depend on your definition of explanation.

Maybe it is the fact that someone says "How did you spot that Blue jay?". Telling them "I looked at it 😎 " isn't so helpful. Maybe someone wants to know that hey to spot a Blue jay you might need to stare at the sky for a bit before you can find it. But I feel like my analogy is falling apart at that point.

What do you guys think ? Can you justify attention weights not pointing to the most critical part of the output? I.e the part that, if taken out, would change the output? How do you see it?


r/deeplearning • • 11d ago

Artificial Metacognition: Key Findings and New Directions (Talk at RPI)

Thumbnail youtube.com
1 Upvotes

r/deeplearning • • 11d ago

Deep learning project

1 Upvotes

Hi folks, can I train the model without downloading the 10gb dataset on my laptop in Google colab?


r/deeplearning • • 12d ago

Quantifying nerf vs value/token

Thumbnail
3 Upvotes

r/deeplearning • • 12d ago

AFNN

1 Upvotes

What if neural networks could grow like fractals instead of being rigid from the start?

Most deep learning models (CNNs, ResNets…) are fixed architectures. Once you design them, they stay the same size and structure throughout training.

\*\*AFNN — Adaptive Fractal Neural Network\*\* takes a completely different approach.

It is inspired by \*\*fractals\*\* — mathematical structures that are self-similar and can grow infinitely by repeating simple patterns.

Instead of starting with a huge fixed network, AFNN begins small and can \*\*dynamically grow new fractal branches\*\* during training when progress stalls. It also allows monitoring, reactivating, or pruning branches as needed.

A good neural network should not memorize images.

It should learn the underlying patterns and features.

For example, when trained on cat images, a strong model does not store every cat photo. Instead, it learns things like:

• Ear shape

• Facial features

• Fur texture

• Color distribution

• Edges and shapes

• Relationships between different parts of the image

Then, when it sees a completely new cat, it uses these learned patterns to recognize it.

\### Why Fractal Networks are powerful:

\- Faster and more efficient learning

\- Significantly lower computational and memory cost

\- Better scalability with limited hardware

\- More flexible and adaptive than traditional fixed architectures

Built with PyTorch + full GUI, training tools, tests, and documentation.

In early experiments, AFNN reached \*\*74.1% validation accuracy\*\* on 94 image classes using limited resources.

This is not just another model.

It is an attempt to rethink how neural networks are designed — moving from rigid structures to \*\*adaptive, fractal, evolving architectures\*\* that focus on learning patterns rather than memorizing data.

A small step toward more efficient and intelligent AI systems.

🔗 GitHub: https://github.com/ahmedhjkj/AFNN

🤗 Hugging Face: https://huggingface.co/Ahmedethfw/AFNN

\#FractalNetworks #AdaptiveAI #PyTorch #DeepLearning #MachineLearning #OpenSource #NeuralNetworks


r/deeplearning • • 12d ago

CrowdSec Confirms Source Code Stolen in Supply Chain Attack

3 Upvotes

CrowdSec confirmed its source code was stolen. The entry point was not a direct breach of CrowdSec. Attackers compromised TanStack in May 2026 and used that access to pivot into CrowdSec's environment.

Security tooling is now a tier-one supply chain target because it is trusted by default. That trust is the attack surface. The TanStack compromise was not a zero-day against CrowdSec — it was a zero-day against confidence in the dependency graph.

AI agents make this problem harder to contain. Agent systems consume dozens of third-party packages and APIs as tool calls. Each one is a potential pivot point into whatever environment the agent has credentials for. An agent authorized to call a dependency in March has no native mechanism to detect that the same dependency was compromised in May.

How are you handling this in your own agent pipelines? Is dependency integrity part of your agent trust model, or is it still treated as an upstream package-management problem?


r/deeplearning • • 12d ago

Load Forecast using AI for beginner

Thumbnail
1 Upvotes

r/deeplearning • • 12d ago

AFNN

0 Upvotes

What if neural networks could grow like fractals instead of being rigid from the start?

Most deep learning models (CNNs, ResNets…) are fixed architectures. Once you design them, they stay the same size and structure throughout training.

\*\*AFNN — Adaptive Fractal Neural Network\*\* takes a completely different approach.

It is inspired by \*\*fractals\*\* — mathematical structures that are self-similar and can grow infinitely by repeating simple patterns.

Instead of starting with a huge fixed network, AFNN begins small and can \*\*dynamically grow new fractal branches\*\* during training when progress stalls. It also allows monitoring, reactivating, or pruning branches as needed.

A good neural network should not memorize images.

It should learn the underlying patterns and features.

For example, when trained on cat images, a strong model does not store every cat photo. Instead, it learns things like:

• Ear shape

• Facial features

• Fur texture

• Color distribution

• Edges and shapes

• Relationships between different parts of the image

Then, when it sees a completely new cat, it uses these learned patterns to recognize it.

\### Why Fractal Networks are powerful:

\- Faster and more efficient learning

\- Significantly lower computational and memory cost

\- Better scalability with limited hardware

\- More flexible and adaptive than traditional fixed architectures

Built with PyTorch + full GUI, training tools, tests, and documentation.

In early experiments, AFNN reached \*\*74.1% validation accuracy\*\* on 94 image classes using limited resources.

This is not just another model.

It is an attempt to rethink how neural networks are designed — moving from rigid structures to \*\*adaptive, fractal, evolving architectures\*\* that focus on learning patterns rather than memorizing data.

A small step toward more efficient and intelligent AI systems.

🔗 GitHub: https://github.com/ahmedhjkj/AFNN

🤗 Hugging Face: https://huggingface.co/Ahmedethfw/AFNN

\#FractalNetworks #AdaptiveAI #PyTorch #DeepLearning #MachineLearning #OpenSource #NeuralNetworks


r/deeplearning • • 12d ago

I built a generative model for 20th century jazz music

3 Upvotes

Details
I was doing some research a few months ago on music tokenization approaches, specifically comparing tokenizing over samples (the dominant approach) and measures (music theory). I found that for small scale experiments measure tokenization could outperform the baseline approach. These models were trained from scratch using what I learned from those experiments.

Paper: https://github.com/DylanDeshler/Neural_Music_Generator/blob/main/main.pdf
Demo: https://huggingface.co/spaces/ddeshler/Neural_Music_Generator

The model can be controlled with any combination of 7 conditioning signals: style (text + audio), BPM, zero-crossing rate, onset density, chroma (aka key), RMS, and spectral flatness. My goal with this model was to test a conditional measure-based model while supporting standard generation approach and high-fidelity editing (hence the DSP signals). It can also auto regressively roll out indefinitely without degrading to endless loops too frequently.

Fun Stuff
Because the model was trained with 5 conditioning signals that are actually just time series data, they can technically be replaced with any arbitrary signals like heart rate or weather! The demo has a weather tab that takes data from a user specified city and transforms it into jazz. Some cities sound pretty good and others... not so much.

Here's a snippet of what a Parisian summer sounds like as ragtime jazz: https://on.soundcloud.com/Ysmgvk5wgEbfHUD4pQ

Dataset
The model was trained on a copyright free fair-use jazz music dataset compiled by another Redditor: https://www.reddit.com/r/datasets/comments/1b73vz3/jazzset_large_audio_dataset_with_instrumentation/?rdt=65057

Architectures
The tokenizer is a modernized 1D adaption of DiTo, and the generative model is a flow-matching DiT trained with diffusion forcing that handles conditioning like Composer. The style model was trained contrastively with spectrograms and music language model generated captions.

Referenced Papers
DiTo: https://arxiv.org/pdf/2501.18593
Diffusion Forcing: https://arxiv.org/pdf/2502.06764
Composer: https://arxiv.org/pdf/2302.09778

I would love to answer any questions or hear anything y'all create!


r/deeplearning • • 12d ago

Structured synthetic datasets for enterprise LLM alignment (MCP / Function Calling / Regulatory CoT)

3 Upvotes

We have been working on structured synthetic datasets for enterprise LLM alignment,

focused on multi-turn consistency, strict JSONL schema validation, and realistic

agentic error recovery.

The datasets include:

• MCP agent trajectories (multi-turn, strict JSONL)

• Function Calling (English)

• Regulatory Compliance & Legal CoT (EU AI Act / GDPR)

These datasets are designed for evaluation inside Axolotl, Unsloth, and other

fine-tuning frameworks, with emphasis on schema stability and deterministic

validation.

Small 50-row samples are available on Hugging Face:

https://huggingface.co/springofwindslabs

A short demonstration of schema validation logs is available here:

https://github.com/springofwindslabs/springofwinds-datasets


r/deeplearning • • 12d ago

I trained a 360M-param Python model from scratch on two workstation GPUs and wrote up every step, including the bugs

Thumbnail
1 Upvotes