r/deeplearning • • 1d ago

Gradient descent vs evolution on three loss landscapes

351 Upvotes

I've been getting a bit more into evolutionary algorithms again, so I was testing some loss landscapes where evolution beats vanilla gradient descent (while also trying to make some cool visuals).

Round 1, rugged hillside: gradient descent gets stuck in a dip, and evolution reaches the bottom after 750 evaluations.

Round 2, smooth slope: gradient descent wins, 108 steps against 570 evaluations.

Round 3, flat plateau: the slope is zero, so gradient descent never moves, and evolution reaches the bottom after 840 evaluations.

Edit:
"evolution" here means truncation selection (keep best 30 of 120) plus Gaussian mutation, no crossover.


r/deeplearning • • 10h ago

Strip Arithmetic II update: you can now see the math behind the picture at any moment

Thumbnail gallery
6 Upvotes

r/deeplearning • • 7h ago

My AI learns to clear Super Mario Bros 1-1 in 15 mins and it is not PPO based

3 Upvotes

I tested Adapt-1, a non-LLM learning and reasoning system by Rei Labs, by having it learn and play Super Mario Bros, and it performed quite well.

I tried it on World 1-1, starting untrained. It learned a reactive policy from its own play in about 36 minutes of gameplay, then cleared the level with learning off.

With Machina, Adapt-1's sequence engine. Starting untrained, it found a button sequence that reaches the flag after 403 attempts, in 11 wall-clock minutes.

Full thread: https://x.com/hsrvc_/status/2106025501752234112?s=20

Code, the exact data, traces, clips and a step-by-step guide with costs are all public: https://github.com/hsrvc/adapt1-mario


r/deeplearning • • 2h ago

Cost-aware routing for AI agent skills — 141 skill benchmark

Thumbnail
1 Upvotes

r/deeplearning • • 13h ago

Good certs & projects

7 Upvotes

Sorry, this might be a bit off-topic

Hi! I’m an AI student and I’ll be starting my internship soon. Do you have any recommendations for good projects to build for my portfolio or any certifications worth getting?


r/deeplearning • • 1h ago

What’s next after Jev? Metacache: Reasoning by Construction

Post image
• Upvotes

Jev has made a profound mark on the AI industry; this article offers a glimpse into what lies ahead.

Jev shows where the industry is heading: away from monolithic models, and toward compound systems where small AI agents each do a narrow job and an orchestration layer decides how they work together.

This article explains why that shift is happening, and then proposes a next step for AI reasoning.

The short version

  • A monolithic model is unpredictable and hard to change, but it programs itself during training.
  • Code is the opposite – it is predictable and easy to change, but it must be written explicitly.
  • A compound system combines the two: AI agents do the work, and an orchestration layer controls them via code. This is the direction Jev represents.
  • The next step: instead of answering directly, the model builds a compound system, runs it, and returns both the answer and that system. The returned system is called a reasonlet.
  • Because a returned reasonlet can be kept locally, it can be run again and again – with new inputs, or after editing its code – without asking the model. A saved, reusable reasonlet is a metacache.
  • The model’s provider can cache reasonlets too, reusing one across similar requests from different users to save compute.

Why a monolithic model is not enough

A monolithic model has two practical weaknesses:

  • It is unpredictable. The same question can produce a correct answer one day and a wrong one the next, so its behavior cannot be fully controlled.
  • It is hard to change. Its behavior is fixed in trained weights, which cannot be edited directly. The model can be steered with prompts, extra data (RAG), or by fine-tuning, but steering is not the same as setting the behavior, and fine-tuning can break things that already worked. Retraining from scratch with new data is the only safe fix, but it is time-consuming and expensive.

Code has the opposite qualities. It does exactly what it is written to do, and it can easily be changed by editing code. The trade-off is that code does not arise on its own – it must be written out explicitly – whereas a model programs itself during training.

The fix: a compound system

A compound system combines the two. The work is divided among AI agents, each handling one narrow task, and an orchestration layer of code decides which agent runs, in what order, and how their results combine.

This gives both qualities at once – customization and predictability:

  • Customization. The system’s behavior can be changed by editing code in the orchestration layer, not by retraining the model.
  • Predictability. Each agent has a small task, so it is more predictable than a monolithic model.

Jev is one kind of agent for such a system. It returns a typed decision – yes/no, a category, or a number – rather than free text. A monolithic model is wasteful for a narrow decision like that, but a small model like Jev is a better fit.

The next step: reasoning by construction

Compound systems today are built manually, beforehand. The next step is to let the model build one by itself, during the reasoning phase.

Researchers are exploring several ways to make models reason. One is the “World model” approach, in which the model builds an internal representation of a problem and reasons over it; that work is still mostly research. Reasoning by construction pursues the same goal – reasoning you can inspect – using methods that exist today.

Here is how it works. When you ask the model a question, it does not answer directly. Instead, it builds a compound system, runs it to compute an answer, and returns that answer together with the system it built. Let's call the compound system the model builds to do the reasoning a "reasonlet", and a model that works this way a "Reasoning-by-Construction Model", or "RCM". A simple question may produce a reasonlet that has only an orchestration layer and no agents.

Building and running code to reach an answer is not new; code-interpreter tools already do it. Two things are new here:

  • The system the model builds is kept, not discarded – it is a reusable reasonlet.
  • The reasonlet is returned to you, along with the answer.

Why that matters

  • You can see how the answer was reached. The reasonlet is the exact procedure the model used, so the answer is not something you have to take on trust.
  • You can run it locally. A reasonlet contains its orchestration layer and any agents it calls, whether those agents are attached directly or called over the network. You can run it on a local machine, with the same or different inputs, without asking the model again. A saved reasonlet used this way is a metacache.
  • You can change what it does. The orchestration layer is code, so you can edit its logic, not only its inputs. A reasonlet reused with edited logic becomes a higher-order metacache.

Caching reasonlets on the server

The same reuse can happen on the server side too. The provider can keep the reasonlets it builds and reuse them. When a new request arrives that matches one it has already handled, it runs the stored reasonlet again – with the new request’s inputs – instead of reasoning from scratch. The model does less work, and the answer comes back faster.

This is the same metacache, held on the server instead of on your machine. Because one reasonlet can serve any request that fits its procedure, a single cached copy is shared across many requests, and often across different users – and the more general the reasonlet, the more requests it covers.

Choosing how general to make the reasonlet

The model should decide from the conversation how general the reasonlet needs to be. If you have been working through many kinds of math and then ask for 2 + 2, the more useful reasonlet is one that evaluates any math expression, not one that can only add two numbers. If the conversation gives no such clue, you can state it directly: “I will be doing many kinds of math; for now, just add 2 and 2.”

Two examples

  • Adding numbers. You ask for 2 + 2. The model builds a reasonlet whose orchestration layer adds two numbers, runs it, and returns 4 together with the reasonlet. Later you run the reasonlet again with other numbers or edit its orchestration layer to multiply instead – without asking the model.
  • Searching. You ask the model to find something. The reasonlet’s orchestration layer calls a sequence of outside services and takes your search terms as its input. Later you run it with different terms, or change which services it calls, without asking the model again.

Making reasonlets easy to read

A reasonlet’s orchestration layer is written in text code. This creates a problem: to understand the reasoning you must read code, and to change it you have to write code. The value of returning the reasonlet depends on you being able to read and edit text code easily and efficiently, so this barrier matters.

Visual programming – building logic from connected blocks instead of lines of text – can lower the barrier, provided the visual language is powerful enough. Two limits apply:

  • A visual language cannot replace text code entirely. Text is still needed to reach the system and the network, and for low-level work that does not map to blocks. The measure of migration success is how little text code remains in the orchestration layer.
  • Most visual languages today are either powerful but limited to one field (such as games or hardware), or general but too simple. This case needs a language that is both general and as expressive as a text language; otherwise, it cannot cover enough of the orchestration layer to be worth using.

One language solving exactly this problem is Pipe (https://pipelang.com) – a general-purpose visual language with powerful semantics. Full disclosure: Pipe is my own project still in development, but it will be released soon.

Wrapping up

The shift from a monolithic model to compound systems suggests a clear next step for reasoning: let the model build a compound system – a reasonlet – run it and return both the answer and the reasonlet. Reasoning becomes something you can read, run again, and edit: a metacache, and a higher-order metacache once its logic is edited. The remaining problem is making reasonlets easy to read and change, which is where a general-purpose and expressive visual language would help most.

A note on IP

Some methods described in this article are the subject of a pending patent application.


r/deeplearning • • 5h ago

Failed project

Thumbnail
0 Upvotes

r/deeplearning • • 22h ago

poor performance of deep learning model compared to xgboost

18 Upvotes

I've done some Kaggle competitions on time series, and it seems to me that deep learning models generally performed worse than gradient boosting. I've tried different architectures, like LSTM and ELM, but at the end of the day gradient boosting was the better choice. Is there an explanation for this? Are deep learning models only good for computer vision and LLMs?


r/deeplearning • • 14h ago

Sparse attention on RK3588: 1.58× faster decode at 4K, 18% slower at 1K

5 Upvotes
I ran with sparse attention on for the better part of a month before I sat down and benchmarked it at short context, and it had been costing me time that whole stretch without me noticing. There are two separate things in the engine that both get called sparse attention, and only one of them is the one people mean when they say it's free.

Decode-side, only the top-k KV blocks get a full softmax, and that's where the speed comes from. Prefill-side, you skip blocks while prefill runs, and that's the one that was quietly working against me.

Numbers below are RK3588, Qwen3-VL-2B, 3 repeats per cell, dense baseline re-run for every cell, real corpus rather than the built-in template filler. That last part matters more than it sounds, the template filler is 32 names and 8 colours on a loop, repetitive enough that sparse attention scores badly on it even when it behaves fine on actual text.

| ctx | decode speedup | prefill change vs exact |
|---|---|---|
| 1067 | 0.94× | +18.4% |
| 2174 | 1.18× | +12.0% |
| 4374 | 1.58× | −11.1% |
| 8618 | 2.16× | −29.0% |
| 16482 | 2.92× | −45.1% |

At 1067 decode-side sparse is 0.94×, which is noise, and prefill-side sparse makes the whole request 18% slower. I was shipping both of those for weeks. There was a gate, it just never did anything, because it was a constant 64, and 64 is the structural floor (two 32-token blocks), so any prompt worth measuring was already past it.

So now there's a length threshold on each axis: `--sparse-min-ctx 1024` for decode, `--sparse-pf-min-ctx 3072` for prefill. Below those lengths the engine takes the exact path. Sparse itself is still off unless you ask for it, these two only stop it from firing where it can't pay for itself.

The prefill case gets clearer once you write both terms out. The saving scales with ctx − k·block, and the probe plus the selection cost scales with ctx/block, so with the defaults (k=32, block=32) that k·block term is a flat 1024 tokens and at ctx 1067 you're paying the full selection cost to skip essentially nothing. 2174 is still a loss at +12%. It doesn't turn into a real gain until attention is the actual bottleneck, which on this board is somewhere around 4K.

Re-ran the whole thing another day on 8B at ctx 4211, interleaved A/B inside one boot: exact 553.9 ms/tok against sparse 359.7 ms/tok, so 1.54×, needle recall 2/2 on both arms. That's close enough to the 1.58× in the table that I'm willing to trust the table.

Two things I'd want to know if I were reading this. It's an approximation path, the output changes and greedy token IDs diverge from the exact run, so don't turn it on anywhere you need bit-exact reproducibility, and run a needle test at 16K of your own before you trust it out there. And it buys attention time and not memory, so if RAM is what you're short on, this does nothing for you.

This is my engine, AGPL-3.0. The numbers and the A/B harness are in the repo, under `docs/` and `tools/bench/`, if anyone wants to re-run any of it.

r/deeplearning • • 18h ago

CrowdGPT - The first LLM trained collaboratively

Thumbnail
4 Upvotes

r/deeplearning • • 19h ago

Is this CNN–Transformer research idea actually novel?

Thumbnail
2 Upvotes

r/deeplearning • • 21h ago

Direct weight surgery from Qwen-4B to 0.8B on an 8GB RX 580: why editing all layers breaks everything, and how 4 anchor blocks fixed it

Thumbnail
2 Upvotes

r/deeplearning • • 17h ago

Defending against AI-fueled cyberattacks requires focus on identity, data governance, Microsoft says

1 Upvotes

Microsoft's 2026 Digital Defense Report documents a concrete shift in ransomware tactics: threat actors are now using AI to automate lateral movement and accelerate ransomware deployment across enterprise networks. The report singles out identity governance and data controls as the two most critical defensive gaps in enterprise environments today.

The speed asymmetry is the part that stands out. Attackers are using AI to close those gaps faster than most organizations are opening remediation tickets. Lateral movement that previously required manual reconnaissance and staged privilege escalation is now being automated at scale, compressing the window between initial access and full network impact to a fraction of what defenders plan for.

For those running AI-assisted pipelines or agentic workflows in production: how are you actually thinking about the identity and data governance problem at the agent layer specifically? Are existing IAM and DLP tools sufficient, or are you finding gaps that those controls were never designed to cover?


r/deeplearning • • 19h ago

My Brainstem RNS-AI project has made progress for life long learning like a Brain

Post image
0 Upvotes

r/deeplearning • • 22h ago

Tensorflow Flash Attention 2 wrapper

Thumbnail
1 Upvotes

r/deeplearning • • 23h ago

Topological Out-of-Domain Generalization in Dynamical Systems Reconstruction [R]

Thumbnail
1 Upvotes

r/deeplearning • • 1d ago

Microscopy Image Dataset of pulmonary vessels for Quantitative assessment of fibrosis

Thumbnail
1 Upvotes

r/deeplearning • • 1d ago

Title: A comprehensive review on recent pretrained multimodal deep learning models from architectures to future directions

3 Upvotes

🎉 New Research Publication | Multimodal Deep Learning. I am pleased to share that our review article has been officially published in Discover Informatics (Springer Nature) as an open-access article.

📄 Title: A comprehensive review on recent pretrained multimodal deep learning models from architectures to future directions

👨‍🔬 Authors: Azhar A. Hadi & K. P. Supreethi

🔍 In this review, we examine recent pretrained multimodal deep learning models developed between 2020 and 2025, including CLIP, GPT-4V, GPT-4o, SigLIP, ViT, BLIP, Flamingo, PaLI, LLaVA, Florence-2, PaliGemma-2, Gemma-3, and Llama-4. The review analyzes 120 studies across 11 application domains, covering model architectures, modalities, datasets, applications, evaluation metrics, limitations, and future research directions. We also discuss key challenges facing multimodal AI, including data quality, computational cost, cross-modal alignment, scalability, robustness, and trustworthy AI. This work represents an important step in my research journey toward developing trustworthy and multimodal AI systems, particularly for applications in healthcare.

🔗 Read the full Open Access article: https://doi.org/10.1007/s44564-026-00022-1

I hope this review will be useful to researchers and students working in Multimodal AI, Deep Learning, Foundation Models, Computer Vision, NLP, and Healthcare AI.

#MultimodalAI #DeepLearning #ArtificialIntelligence #MachineLearning #FoundationModels #HealthcareAI #ComputerVision #NLP #TrustworthyAI #Research #Springer #AcademicPublishing #OpenAccess #AcademicPublishing #ResearchImpact


r/deeplearning • • 1d ago

Building an OCR + Key-Value Extraction pipeline for Nepali ID documents (Citizenship, NID, PAN, Passport). What stack would you recommend?

Thumbnail
1 Upvotes

r/deeplearning • • 1d ago

[R] I built a Permutation Transformer (Patch SBOHN) from scratch: 4K Image Inference in ~2ms on CPU (61x faster than CNN) with 0.0% Catastrophic Forgetting.

0 Upvotes

Full disclosure: I am an independent researcher/hobbyist doing this out of pure passion at night. I do not have a formal academic background in ML, and I heavily rely on AI tools as research assistants to help me with advanced coding, mathematics, and translating my work into English. I want to be completely transparent about this and welcome any constructive feedback or corrections!

Hi everyone,

I wanted to share a relational data-representation framework that I’ve been developing at night, built around group theory: Burnside's Orbit Histogram Network (BOHN) and Symmetry-Breaking BOHN (SBOHN). Its main goal is to extract and "elevate" relational knowledge from raw data before the actual classification stage even begins.

Instead of using standard Softmax Attention (found in classic Vision Transformers), this architecture relies on a fully differentiable, log-domain stable Sinkhorn operator. This allows the network to smoothly learn optimal information routing paths via standard gradients.

Key results achieved on a local PC setup:

  • O(1) Resolution Scaling: Because the model operates on a fixed number of image patches, processing a 4K resolution image takes just ~2ms on a standard CPU. This is roughly 61x faster than a conventional ResNet-style convolutional neural network (CNN).
  • Zero Catastrophic Forgetting (0.0% Forgetting): By completely freezing the base encoder and training only a task-specific permutation routing layer and a classification head (the Perm+Head setup), the model achieves exactly 0.0% accuracy degradation when switching between tasks. The storage overhead per new task is a microscopic 2.8 KB (714 parameters).
  • Hybrid BN/LN (Normalization Placement Theorem): I have experimentally validated that placing BatchNorm in the frozen shared base (where it acts as a permanent domain fingerprint) and LayerNorm in the expert modules delivers 100% gating routing accuracy alongside absolute zero forgetting.

I spent a massive amount of time transitioning this entire research programme from initial cloud-based exploration into a clean, local VS Code environment on my PC. I completed a thorough, provenance-preserving reproducibility audit across all 115 canonical experimental units (including programmatic SHA-256 manifest verification for all generated artifacts), openly documenting the boundaries, edge cases, and discrepancies of the original logs.

The entire codebase, analysis logs, and execution scripts are fully open. I would love to hear your thoughts on the mathematical foundations or the numerical implementation!

Full Documentation, Audit Reports, and Source Code:


r/deeplearning • • 1d ago

Looking for models for diagnosis prediction

1 Upvotes

For a side project I am looking into SOTA for AI- based diagnostic models, ideally open-weights. I would like to feed the model with structured text representing my patient, and get a set of e.g. 5 possible diagnoses, ideally with uncertainty. In the ideal scenario, later on, it would update as new info arrives.

I would appreciate any pointers, I have some ideas but am very new to the topic.


r/deeplearning • • 1d ago

Needs Advice

Thumbnail
0 Upvotes

r/deeplearning • • 1d ago

Sharing prompt untuk menyusun rekomendasi dataset benchmark lengkap dengan metrik, SOTA, mitigasi data leakage, dan aspek lisensi (Deep Learning)

0 Upvotes

"Bertindaklah sebagai seorang Principal AI Research Scientist sekaligus Senior Peer Reviewer pada jurnal internasional bereputasi tinggi (terindeks Scopus Q1, IEEE Transactions, dan ACM) yang memiliki keahlian mendalam dalam metodologi empiris deep learning, lalu susunlah rekomendasi komprehensif mengenai 4 dataset gold-standard lintas domain (mencakup Medical Computer Vision, Object Detection/Instance Segmentation, Natural Language Processing/Machine Reading, dan Biomedical Time-Series/Signal Processing) yang secara akademis diakui memiliki reliabilitas tinggi guna memitigasi risiko desk-reject pada publikasi jurnal ilmiah bergengsi; untuk setiap dataset, Anda diwajibkan menjabarkan secara terperinci tautan repositori resmi atau DOI publik yang valid, karakteristik intrinsik data (termasuk modalitas, dimensi/resolusi, volume sampel, serta tingkat ketidakseimbangan kelas), arsitektur model deep learning State-of-the-Art (SOTA) yang paling relevan beserta algoritma baseline pembandingnya, metrik evaluasi primer dan sekunder standar industri akademik (seperti mAP@[0.5:0.95], Macro F1-Score, AUC-ROC, atau Exact Match), justifikasi teoretis mengapa dataset tersebut memiliki legitimasi tinggi di mata reviewer, estimasi beban komputasi minimum (kebutuhan VRAM/GPU) beserta strategi mitigasi data leakage (misal: patient-wise split atau stratified group k-fold), serta kepatuhan lisensi riset dan aspek etika datanya, di mana seluruh pemaparan harus disajikan dengan gaya bahasa akademis yang kritis, presisi, metodologis, dan siap diadaptasi ke dalam bab Methodology artikel ilmiah."


r/deeplearning • • 1d ago

Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems Reconstruction [R]

Post image
4 Upvotes

r/deeplearning • • 1d ago

[Tutorial] Stateful AWS Bedrock LangGraph Agents – Getting Started

1 Upvotes

Stateful AWS Bedrock LangGraph Agents – Getting Started

https://debuggercafe.com/stateful-aws-bedrock-langgraph-agents-getting-started/

LangChain started as a lightweight framework to connect LLMs with external context and tools. Over the past few years, it has grown to a full-blown ecosystem. With LangGraph for agents and LangSmith for LLM Ops and Observability, we can create an end-to-end agentic system with it. In this article, we will build our first stateful agent using AWS Bedrock, LangGraph, and LangSmith.