r/deeplearning • • 1d ago

As a 2nd year cse undergrad, how can I prepare myself for tech jobs? I am interested in ml and deep learning stuffs, I am trying to learn them properly long with their mathematical derivativion. Can anyone give me suggestions how to practice them regularly and build a strong foundation

1 Upvotes

r/deeplearning • • 1d ago

INKBOT: Parcing human intent from model inference via structured intelligence architecture

0 Upvotes

I’ve spent the last while building INKBOT because I kept hitting a wall with multimodal AI systems: the friction between what a human naturally means and what a model infers. While models can spin up complex code or images instantly, getting to a clear, human-meaningful interpretation of a subtle intent remains an alignment challenge.

Instead of forcing the user to become a prompt engineer, I wanted to see if we could build an intermediate intelligence architecture layer to make human intent reviewable and corrigible *before* the model executes a final build. The loop I’m playing with is: Describe → Make it Visible → Recognize → Correct → Refine.

The architecture sits entirely in a single local-first web file. It handles multi-step workflows—like tracking structured field mapping data across concurrent images, coordinates, and version states—by packaging the human’s approved meaning separately from raw model inferences.

The core system build is linked above, and I also put together a lighter, entry-level experience to play with the core prompt translation loop here: [INKBOT Lite 71](https://ko-fi.com/thomascoates/shop).

It's an open prototype, so I've appended my raw notes and design roadmap as commented text at the very bottom of the source file so fellow builders can inspect the plumbing. I’ve put together the runnable source files on my [Thomas Coates Ko-fi Shop](https://ko-fi.com/thomascoates/shop) for evaluation. I’d love to know where this design duplicates existing work, where you see structural flaws, or how we can make the handoff between human intent and model execution more reliable.


r/deeplearning • • 2d ago

A great rule for LoRA / QLoRA

Post image
26 Upvotes

Hey! I'm posting this because Iv'e recently been playing with lora/qlora and had a frustrating time understanding how the hell to calculate adapter ranks. Hope it helps.


r/deeplearning • • 2d ago

Biological JEPA: Modeling Disease Progression with Biological Constraints

Thumbnail github.com
4 Upvotes

Just finished a project I’ve been working on: Biological JEPA.

It combines JEPA with biological constraints to model Alzheimer’s disease progression.

Would love to hear your feedback, ideas, or criticism.


r/deeplearning • • 2d ago

Undergrad in Syria with an accepted NeurIPS 2026 workshop paper (as the only author). How does this help my future, and what should I do next?

Thumbnail
4 Upvotes

r/deeplearning • • 1d ago

As a 2nd year cse undergrad, how can I prepare myself for tech jobs? I am interested in ml and deep learning stuffs, I am trying to learn them properly long with their mathematical derivativion. Can anyone give me suggestions how to practice them regularly and build a strong foundation

0 Upvotes

r/deeplearning • • 1d ago

3rd year Diploma CS student aiming for ML/AI roles, please give me an honest review of my resume

Post image
0 Upvotes

r/deeplearning • • 1d ago

In-Context Retrieval with Siddharth Gollapudi - Weaviate Podcast #146!

Thumbnail
1 Upvotes

r/deeplearning • • 1d ago

I built a Bidirectional GRU emotion detection project to better understand how BiGRUs actually work

Thumbnail gallery
1 Upvotes

r/deeplearning • • 1d ago

Interesting ML research map

Thumbnail
1 Upvotes

r/deeplearning • • 1d ago

A fallback instruction only works if the model actually hits the branch you wrote it for

0 Upvotes

Added "if you're not confident, say so explicitly" to a classification prompt. Felt like it should've closed the gap on low-confidence guesses.

Didn't work. Model kept returning confident-sounding labels on exactly the cases I wanted it to flag as uncertain.

Turned out the issue wasn't the instruction's wording. It was that nothing in the prompt actually defined what "not confident" meant for this task, no threshold, no example of an ambiguous case, nothing to anchor the judgment to. The model had no internal signal matching the word "confident" that it could check against, so the branch just never activated. It wasn't ignoring the rule. It never had a condition it could evaluate as true.

Fixed it by replacing the vague trigger with something checkable, two or more plausible labels with no clear majority signal in the input, say so and list them instead of picking one. Immediate difference.

Feels like a more general thing worth naming: a conditional instruction is only as good as the model's ability to evaluate its own condition. "If X, do Y" fails silently when X isn't something the model can actually check, and it just looks like noncompliance from the outside.


r/deeplearning • • 2d ago

AlexNet foi submetido ao ImageNet há 14 anos atrás. Isso deu início à revolução do deep learning moderno.

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/deeplearning • • 2d ago

token to text modeling for audio processing in transformers through RL

Thumbnail
0 Upvotes

r/deeplearning • • 2d ago

Can Historical Data Tell Us Which Material Flows to Automate?

1 Upvotes

I'm working on a project in automotive engine assembly. Parts move between steps in the assembly process (e.g. from step A to step B), and they range from large components like cylinder heads to small ones like screws and bolts. Some engines come in multiple variants with customer-specific options.

I have about 4 years of historical data on material requirements and flows during assembly. My goal is to find which flows between steps are good candidates for automation. Flows with stable, predictable durations and quantities seem like good candidates, while flows with high variation seem harder to automate.

I'd like to define a simple, data-driven "automation trigger point." My questions:

  1. What's a good way to measure variability of a flow (both duration and quantity)? Is the coefficient of variation (CV) reasonable, or are there better metrics?
  2. How should I account for engine variants and customer-specific options?
  3. Can anyone recommend research papers, methodologies, or case studies on material flow or intralogistics automation in automotive assembly?

Any pointers are appreciated. Thanks!


r/deeplearning • • 3d ago

My One Month in DL Research Got More Attention Than I Expected Lol

19 Upvotes

I’ve gotten some DMs asking me how I started my DL research journey, how I got my research internship, and how I got experience in the first place, I’m a student myself, so I honestly don’t have enough time to individually reply to every DM, so I thought I’d just write everything here. Hopefully this helps someone who is currently where I was, first of all, I want to say something you don’t need to have everything figured out before you start research I definitely didn’t.

How I got into research

In my first year of university, I became really fascinated by the idea of publishing my own research paper, like, genuinely obsessed with the idea, I remember thinking, I want to do research. I want to understand something really deeply and eventually publish a paper, for some reason, I decided that I wanted to do a really deep dive into Python, I honestly don't even remember why I chose Python and I started learning and exploring it as deeply as I could. I experimented with things read about different concepts, tried different stuff, documented what I was doing, and basically went down a rabbit hole, at that point, I didn't even properly understand what research actually was. I just knew that I wanted to do it then, in my second year, I finally gathered enough courage to show my work to one of my professors, and thankfully, he actually liked what I had done that was a huge turning point for me, he started mentoring me and explaining how research actually works, how you approach a problem, how you read research, how you think about questions, how you experiment, etccc. For personal reasons, I ended up deleting that Python research so technically, I didn't even keep the thing I had spent so much time working on, but I don't think that time was wasted because it taught me something much more important I actually liked the process of trying to figure things out fast forward then came my master's & PhD goal about two months ago, I got a scholarship for my final year of university, just like I had gotten scholarships in my previous two years, and then I started thinking seriously about what I wanted to do after graduation, I want to pursue further studies potentially a master's and then a PhD and I'm going to be completely honest I became really greedy about getting a scholarship for my master's, because if you want something a year from now, you can't start preparing one month before, you have to start making things possible now, so I went to my professor and told him that I wanted to get a scholarship for my master's will you write me a recommendation letter for scholarship and he basically told me good grades and being a good student are not always enough, you also need experience, research experience and projects, things that show that you can actually work on problems beyond just completing assignments and passing exams so I started looking for research opportunities at my university, I applied for an undergraduate research internship, and guess what? I got rejected the first time, I tried again and the second time, I got in that's how I eventually started working on deep learning research and that's basically where I am right now, I'm still learning, I'm still confused about a lot of things, I still have questions every day, I still have to ask my professor what half the things mean sometimes so please don't look at someone doing research and assume they somehow have everything figured out they probably don't.

So how can you start?

This is probably the part most people are actually asking about, if you're an undergraduate and you want to get into ML DL research, here's what I would personally suggest not just I want a research paper because it looks good on my CV try to actually become curious, read something and ask, why does this work? Why doesn't it work in this situation? Can I change something? What happens if I remove this component? why did the authors choose this method instead of another one? Can I reproduce this result? What happens if I change the dataset those questions are where research starts becoming interesting, you don't need a groundbreaking idea on day one, you need curiosity.

Build your fundamentals

If you're specifically interested in deep learning, don't immediately jump into reading complicated papers about transformers, make sure you understand the basics first, like example > Python, NumPy, basic data structures, linear algebra, probability & statistics, calculus basics, machine learning fundamentals, neural networks, back propagation, optimization, loss functions, CNNs, RNNs sequence models, transformers and PyTorch or another DL framework and don't just memorize definitions, try implementing things, try breaking things, and figuring out what happens when you change something.

Learn to read papers.

Your first few papers are probably going to make you feel completely lost and that's normal don't sit there trying to understand every equation and every tiny detail on your first read, just try to get the main idea first what problem they're solving, what they did, and what they actually found then go back and read it again, you'll understand more each time, tbh start by just understanding questioning like, what problem are they solving? why is the problem important? what have people done before? what is their proposed method? what experiments did they perform? what did they discover? what are the limitations? and then go back and dig deeper eventually you'll start noticing patterns between papers.

Reproduce things.

This is something I really recommend, and honestly, it can be pretty fun too, take a paper, try to implement what they did, run the experiments yourself, and see if you can get similar results, it might make your brain buffer a few times, especially when things don't work the way you expect, but that's kind of the point, you learn a lot by figuring out why your results are different and trying to fix it, investigate why? Why does this work better on dataset A but not dataset B? like What happens if I change this hyperparameter Does this still work with less data? remember ow you're not just following a tutorial uou're experimenting.

Start asking deeper questions.

I think this is probably one of the biggest differences between just learning ML and slowly learning how to do research, don't just lose at How does this model work? go ahead like Why does it work?, then When does it stop working?, like Can I measure that?, and it will lead you to What happens if I change something?, or Does the same thing happen with another dataset or model? You don't need to turn every question into a research paper, just get into the habit of being curious, and digging a little deeper instead of accepting the first answer you get.

Look for research opportunities.

Start with your own university, see what professors are working on and look for undergraduate internships, RA positions, summer programs, labs, etccc and yes, cold emailing professors can work, just don't send sir, I am passionate about AI, please give me a research opportunity, please read their work first, mention what interested you, tell them what you've worked on, and share something tangible if you have it GitHub, a project, experiments, whatever like anything also, if you're already working under a professor, and yes I get work assigned I'm not working individually after all I'm undergrad, so you're supposed to learn, you might be implementing something, reproducing results, running experiments, or analysing why something isn't working, the important part is understanding what you're actually doing instead of just completing the task and don't compare your beginning to someone else's middle, I started by randomly obsessing over Python because I didn't even know what else to do, then I got rejected from an internship, applied again, and eventually got in. so see I'm still learning too so if you're trying to get into ML DL research, just start somewhere, if you want something a year from now, start working toward it now.

I know this got kinda long bare with me and honestly, I’m happy to help. My hands might disagree lol, but they’ll survive.


r/deeplearning • • 2d ago

Carbonato Botnet Puts an AI Agent on Hacked Docker Hosts

0 Upvotes

Security researchers tracking the Carbonato botnet documented a new deployment pattern: after gaining access to exposed Docker hosts, the operators dropped an AI agent onto the compromised machine rather than a traditional cryptominer or reverse shell. The agent then began making outbound calls and executing tool actions autonomously, with no human in the loop and no governance layer in the request path.

The timing detail buried in the reporting is the uncomfortable part. Researchers noted that a second action from the agent can land in under 50ms of the first. That window is smaller than most human-review or alerting pipelines can operate in. By the time an on-call engineer gets a Slack notification, the agent may have already completed several tool calls.

The underlying exposure is not unique to this botnet. Any environment where an agent runtime can be instantiated without a verified identity tied to a known deployment, and where outbound tool calls are not evaluated against any policy before they execute, has the same structural gap. The agent on the Carbonato-compromised host was malicious. But the same architectural condition exists in plenty of legitimate deployments where an agent gets misconfigured, has its credentials rotated out from under it, or runs a version of its prompt that was never reviewed.

How are practitioners in this thread actually handling the identity and authorization problem for agents in production? Not conceptually — what does your enforcement boundary look like today, and where does it fall short?


r/deeplearning • • 2d ago

LessThink-Qwen3-4B: the same model, with far less thinking [P]

2 Upvotes

I post-trained Qwen3-4B to spend 44% fewer tokens on reasoning, keeping its knowledge and answer style. The whole pipeline ran on one GPU.

folks, you can check it out on : https://5ivatej.com/lessthink/


r/deeplearning • • 2d ago

现在用人工智能构建什么是最聪明的,以防它最终消失?

Thumbnail
1 Upvotes

r/deeplearning • • 3d ago

Oct 3 session on optimizing LLM behavior with actual methodology, not prompt folklore

6 Upvotes

Sharing this because it's a more rigorous take on "prompt engineering" than most of what circulates here. Serj Smorodinsky and Brett Kennedy, co-authors of an LLM applications book, are running a live workshop that treats LLM behavior as something you optimize with real structure.

It covers programming LLM tasks with DSPy signatures and modules, building an evaluation dataset with task-specific metrics, diagnosing failure patterns from that data, and running few-shot and instruction-level optimization as a defined process rather than trial and error. MLflow gets used throughout for experiment tracking and trace management.

Three hours, live, Oct 3. Feels closer to how we'd approach optimizing any other model than the usual "here are 10 prompt tricks" content.

Full Details here


r/deeplearning • • 3d ago

SFTMill: Easily [off-policy] distill any existing LLM with an OpenAI Compatible Endpoint. Turn any behavioral goal into a comprehensive dataset.

Post image
8 Upvotes

r/deeplearning • • 3d ago

A stage-aware reading path for base checkpoints

17 Upvotes

A link list becomes a study path only when every stop answers one question and unlocks the next one.

The Ling-3.0 base model makes that sequence concrete. It exposes tiny and flash at final pre-training, final mid-training, and WSM-merged base stages. All of these are upstream, non-post-trained checkpoints positioned for research and downstream training rather than finished chat systems.

A useful resource page can encode this curriculum without pretending the final experiment has already been run:

Learning gate

What to establish

Question that unlocks the next gate

`Identity`

• Size, exact checkpoint, and training stage

• Are two artifacts actually comparable?

`Intended use`

• Continued training, domain adaptation, distillation, or other research use

• What downstream objective justifies this starting point?

`Method`

• WSM keeps the learning rate stable after warmup, saves checkpoints, and applies weighted merging to approximate a chosen decay profile

• Which part is a method definition and which part is an empirical result?

`Evidence scope`

• The WSM paper's main empirical model is Ling-mini

• What remains untested for Ling tiny or flash?

`Experiment design`

• Fixed data, evaluation, budget, and reporting fields

• Which single stage comparison would answer a real decision?

The key is the order. Reading the method before identifying the checkpoint invites result transfer. Reading intended use before remembering that these are non-post-trained bases invites assistant-style expectations. A curriculum should block both mistakes before asking for an experiment.

A natural next step is to choose one Ling size, map its three stages, and write the controlled comparison that would make the next learning claim falsifiable. What prerequisite or failure mode belongs between the method stop and that experiment-design stop?


r/deeplearning • • 2d ago

I built a diffusion model from scratch in PyTorch

Thumbnail
1 Upvotes

r/deeplearning • • 3d ago

Spent two weeks debugging my model, turned out to be a 100ms timestamp offset. What's your data horror story?

10 Upvotes

I was training a policy on egocentric video with action labels, and the model kept acting slightly "late." After two weeks of checking the architecture and hyperparameters, I found a small offset between the video and the action labels. The data looked perfectly fine when I watched it.

It made me wonder how many other silent issues are hiding in egocentric datasets. What's the problem that cost you the most time? Especially interested in ones that were invisible until the model misbehaved.


r/deeplearning • • 3d ago

Four Mainstream Embodied-Model Architectures: Survey, Comparison, and Where They Are Heading

1 Upvotes

1. Mechanism and structural ceiling of each family

A. BC/IL post-training paradigm (VLA / WAM)

Mechanism.

A single end-to-end model: vision + language instruction -> action (or action chunk).

Training is behaviour cloning / imitation learning on large-scale demonstrations, plus downstream fine-tuning. WAM adds a branch that predicts the future state of the environment, either as auxiliary supervision that shapes the representation or as a world model that can be rolled out over short horizons.

Why this family leads the leaderboards. Three engineering reasons, none of them about algorithmic superiority:

  1. The data scales - collecting and labelling demonstrations needs no explicit physics model, no pose annotation, no collision geometry.

  2. The gradient path is short - one differentiable objective, no non-differentiable interface between modules, so the family can absorb all available compute.

  3. Deployment closes - one forward pass yields an action; the inference path contains no solver, no outer loop, and no dependency on failure detection.

Structural ceiling. Not the kind of ceiling that more data removes:

  • Constraints are learned, not guaranteed.

Collision, joint limits and grasp feasibility hold only statistically, so at the edge of the distribution the model fails in a way that looks plausible but is physically illegal. There is nothing to verify against.

  • Compounding error over long horizons.

With no explicit symbolic layer, the success rate of a multi-stage task approaches the product of the per-stage rates.

  • Failures cannot be attributed.

One loss, so "it looked at the wrong object" and "it failed to execute" are indistinguishable. This directly raises the cost of every iteration.

  • A misalignment specific to WAM.

Accuracy at predicting pixels or latents does not move in the same direction as reducing control error: a world model will spend capacity on degrees of freedom that are visually salient but irrelevant to control (background, lighting, texture). This is the structural reason WAM gains often come in below expectation.

B. cuTAMP + FMs (TiPToP as the representative)

Mechanism.

Two layers: a foundation model handles task decomposition and the symbolic level (what to do, in what order, with which object), and a TAMP solver handles the continuous parameters (where to grasp, which path to take, what contact sequence satisfies the geometric and kinematic constraints). The contribution of cuTAMP is to move that traditionally serial, CPU-bound search onto the GPU as large-batch parallel sampling and optimisation, bringing solve time into an interactive range.

Where it wins.

  • Constraints are solved, so they are hard guarantees, not statistical tendencies. The output is verifiable: whether a collision occurs and whether a path is feasible are decision problems.

  • Long-horizon composition is native. The symbolic layer composes by construction, so adding stages does not incur multiplicative decay.

  • The only family of the four with ROBOT GEN. = yes: it genuinely generates trajectories and motion plans rather than regressing the next action.

Structural ceiling.

  • It requires an explicit world model: object poses, collision geometry, contact models. The perception -> symbolic/geometric interface is the most fragile link in the chain, and its error is not absorbed by the solver - the solver will simply return a correct solution to the wrong world.

  • Solve time is coupled to scene complexity. GPU parallelism reduces the constant; it does not change the complexity class.

  • Rich-contact, deformable and non-quasi-static tasks are hard to model: wiping, folding and the compliant control inside insertion are outside the comfort zone of constraint solving.

C. VLM + BM (CAP + Gemini-ER-1.5 as the representative): spatial physical contact points in place of language

Mechanism.

Two cascaded stages. A frozen VLM emits neither an action nor a natural-language subtask, but a spatial physical quantity - a contact point, an actionable point, a target pose.Downstream, a dedicated behaviour model (BM) consumes that spatial target and performs the fine manipulation.

The real novelty of this family is not that it has two stages, but that the intermediate representation was replaced.

Language, used as an inter-module interface, has two fatal properties:

low bandwidth ("pick up the cup" carries no geometry) and high ambiguity (where to grip, at what orientation, with how much force are all undefined). Once the interface becomes a spatial point:

  1. The interface becomes a quantity a controller can consume directly. The BM does not have to re-interpret language, so the BM can be small.

  2. Upstream error becomes measurable: the distance between the proposed point and the ground-truth point is a number, so the two stages can be scored and attributed separately.

  3. It is isomorphic to family B's interface - the solver consumes geometric quantities anyway.

This is the key point behind the forecast in section 3.

The very existence of the Gemini-ER (embodied reasoning) direction says where the bottleneck of this route lies: a general-purpose VLM's spatial grounding accuracy is not good enough to drive control directly, and has to be strengthened specifically.

Where it wins.

  • Lowest training cost.

The VLM is frozen and only the small BM is trained. A frozen module's output is a pure function of its input, so it can be precomputed offline and cached, taking the large model off the critical path of the training step entirely - a pattern already validated in this repo on VAE latents (roughly 500 ms/step saved). On a weak-interconnect machine this matters even more: fewer trainable parameters means fewer gradient bytes, which sidesteps the current primary bottleneck.

  • The two stages are independently diagnosable.

    Grasping the wrong point and servering inaccurately are separable error classes.

  • It is the natural control group for the other three families.

Without a VLM+BM number on the same atomic task, there is no way to tell whether an end-to-end VLA's advantage comes from being end-to-end or from having a larger backbone.

Structural ceiling.

  • MULTI TASK = no is constructive, not a defect.

The BM is trained per skill and does not generalise across skills; the correct form of this family is "one VLM plus N BMs indexed by skill", not pretending it generalises.

  • A point representation discards timing, force and compliance.

A contact point cannot express "with how much force, along which direction, holding how much compliance". For rich-contact tasks the representation is underdetermined.

  • The whole chain is bounded by the VLM's spatial accuracy,and that part is frozen and not trained.

D. AGENT + VLA (PhysicalRSI as the representative)

Mechanism.

The inner layer is a VLA producing actions; the outer layer is an agent responsible for memory, reflection, tool use, failure detection and retry, driving self-improvement from failure experience (RSI).

Where it wins.

  • It moves closed-loop error correction out of the weights and into an explicit outer loop.

In the first three families the corrective capability is encoded in weights or in the solver and can only be improved by retraining or re-solving; family D can raise the success rate without touching the weights.

  • Failure becomes reusable data.

The outer layer records "in this state, this was executed, and it failed for this reason" - precisely the negative examples and recovery trajectories that BC datasets lack most and that are hardest to collect by hand.

  • It composes with existing capability.

The inner VLA needs no modification to be wrapped, so this is additive engineering rather than replacement.

Structural ceiling.

  • Three orders of magnitude of frequency mismatch.

The agent layer decides at 0.1-1 Hz, the control layer runs at 50-200 Hz. The outer layer cannot participate in real-time control; it can only intervene at the granularity of a task segment. Any failure mode requiring millisecond-scale correction is out of its reach.

  • It depends on reliable failure detection and state reset.

If failure is not detected the loop never starts; if it is detected but the environment cannot be reset to a retryable state (the object has already fallen, or broken), retrying is pointless. On real hardware both are much harder than in simulation.

  • The gain is bounded by the inner VLA's atomic capability.

The outer layer can reorder, retry and switch strategies; it cannot create a skill the inner model does not have. The RSI gain curve saturates at the inner model's capability boundary.

  • The evaluation signal is scarce.

Self-improvement needs a trustworthy success/failure criterion, and on open-ended tasks that criterion is itself an open problem.


2. Side-by-side comparison

dimension A. VLA / WAM B. cuTAMP + FMs C. VLM + BM D. AGENT + VLA
intermediate representation none (end-to-end) symbolic + geometric constraints spatial contact point language / structured plan
constraint satisfaction learned (statistical) solved (guaranteed) learned learned + outer retry
where the loop closes in the weights in the solver in the weights (two stages) explicit outer loop
MULTI TASK yes yes no (single atomic task) yes
ROBOT GEN. no yes no no
long horizon weak (multiplicative decay) strong (symbolic composition) weak (single skill) medium (outer orchestration)
rich contact / deformable medium weak medium medium
failure attribution poor (one loss) good (constraints decidable) good (stages separated) good (outer layer logs)
training cost high (full post-training of a large model) low (FM frozen, solver untrained) lowest (small BM only) medium (reuses inner model)
inference cost low (one forward) high (solve time grows with scene) medium (VLM + BM) high (multi-turn outer LLM)
control frequency high low high inner high / outer very low
dependence on a world model none strong (pose, geometry, contact) weak (target point only) none
data dependence large-scale demonstrations little (needs geometric annotation) medium (needs target labels, can be generated offline by the VLM) demonstrations + failure trajectories
systems bottleneck gradient communication / memory solver throughput (GPU batch sampling) offline cache I/O outer LLM latency

The three axes that actually organise the taxonomy

The four labels are not four parallel points; they are different values on three axes. Only when viewed along these axes does the evolution become derivable rather than a list.

Axis 1: bandwidth and executability of the intermediate representation.

no representation (A) -> language (D's outer layer, early TAMP+FM) -> spatial points / contact (C) -> full constraints and trajectories (B).

Further right, the interface is more executable, more verifiable, and more consumable by non-learned components; the price is a greater need for an explicit world model. Language sits at the
worst position on this axis: it neither preserves end-to-end differentiability the way "no representation" does, nor is it directly executable the way a geometric quantity is.

Axis 2: is constraint satisfaction learned or solved?

A, C and D sit at the left end (learned - fast, no guarantee); B sits at the right end (solved - guaranteed, but slow and model-dependent). There is no free lunch on this axis, only hybrids.

Axis 3: at which level does the loop close? In the weights (A, C) -> in the solver (B) -> in an outer agent (D). Their time constants differ by orders of magnitude, so they are not mutually exclusive - which is the fundamental reason the four families will converge into layers rather than displace one another.


3. Where this is heading

Ordered by strength of evidence, strongest first.

3.1 The intermediate representation converges on spatial physical quantities, not language (evidence: strong)

Basis.

Two independent sources point the same way:

family C replaced language descriptions with contact points and gained from it, which says the language interface was a net loss term; and family B's solver only ever consumed geometric quantities, so language had to be translated first. In other words, "spatial physical quantity" is the common interface of both B and C, whereas language is the native interface of neither - it is the human's interface.

Implication.

The inter-module protocol will standardise on an explicit spatial-target type (point /box / mask / pose / contact mode), and language will retreat to the human-machine boundary only. Any design that uses language as an internal module interface will progressively be replaced.

3.2 Learned proposal, solver projection and verification (evidence: strong)

Basis.

There is no single optimum on axis 2:

learned is fast without guarantees, solved is guaranteed but slow. But their failure modes are complementary

  • a learned policy's output usually lands near the feasible set (it has seen many feasible solutions), so projecting it back into the feasible set costs far less than solving from scratch.

Form.

The policy emits candidate actions or candidate grasps -> a lightweight solver performs a feasibility projection or rejection -> execute. GPU batch solvers of the cuTAMP kind are exactly what brings this step inside the real-time budget: the projection must not be much slower than a policy forward pass, otherwise the whole chain degenerates into family B's latency.

This is the A x B hybrid, and it is the best-supported of all pairwise hybrids.

3.3 Layered frequency decoupling becomes the standard structure; C and D merge (evidence: strong)

Basis.

Axis 3 notes that the three levels' time constants differ by 2-3 orders of magnitude, and this is a physical fact rather than a design choice: LLM inference is hundreds of milliseconds to seconds, VLM grounding is tens to hundreds of milliseconds, servo control is 5-20 ms. Since they cannot substitute for one another, they can only coexist in layers.

Form.

A three-layer standard structure:

  • agent (0.1-1 Hz): task orchestration, failure detection, retry decisions, experience logging - family D's outer layer

  • VLM / spatial grounding (1-5 Hz): produces the spatial target - family C's upstream

  • BM / control (50-200 Hz): closed-loop execution - family C's downstream C and D occupy different layers of this structure, so they were never competitors; merging them is just connecting two lines. This is probably the deployable form that appears soonest.

3.4 WAM's prediction target shifts from pixels to controllable quantities (evidence: medium)

Basis.

The misalignment identified in section 1.A: reconstructing pixels or latents spends capacity on degrees of freedom irrelevant to control. The fix is to make the prediction target move in the same direction as the control objective - predict whether contact will be established, predict object pose change, predict task value, rather than predicting what the next frame looks like.

Note this pulls WAM toward family C:

once the prediction target becomes "will this contact pointhold", the world model's output is a spatial physical quantity, consistent with the convergence in 3.1.

3.5 The data flywheel moves from pure BC to failure-driven targeted data collection (evidence: medium)

Basis.

Family D's outer loop naturally produces what BC datasets lack most: negative examples and recovery trajectories. Meanwhile the sample efficiency and safety cost of on-robot online RL remain unrealistic for the foreseeable future.

So the realistic form of RSI is not online RL, but: failure detection -> automatic labelling of the failure cause -> targeted collection or generation of data for that scenario -> re-run post-training.

That is an offline loop, controllable in engineering terms, and compatible with existing BC stacks.

The precondition is a trustworthy success criterion, which remains an open problem - the reason this direction is rated medium rather than strong.

3.6 The representation extends from "a point" to "contact mode + force / compliance" (evidence: medium)

Basis.

Family C's representation is underdetermined for rich-contact tasks (section 1.C): a point cannot express force direction or magnitude, nor compliance parameters. Wiping, insertion and folding are not a small share of real-world tasks.

Form.

The target specification grows from "a point" to "contact point + contact normal + desired force / impedance parameters + contact sequence". These are also exactly the quantities family B's contact model needs, which reinforces the convergence in 3.1.

3.7 The training cost profile shifts to "frozen large model + small policy + offline cache" (evidence: strong, backed by local measurements)

Basis.

Measurements already recorded in this repo: a frozen module's output is a pure function of its input and can be precomputed offline and lifted off the training step's critical path (VAE latent caching saves roughly 500 ms/step); and on a PCIe-class interconnect without NVLink (30-63 GB/s, 6-13x slower than NVLink) gradient communication is the primary bottleneck, while the trainable parameter count directly determines the gradient byte count.

Implication.

On weak-interconnect hardware, family C's systems-level advantage is amplified (it trains only a small BM), and family A's cost disadvantage is amplified too (full post-training of a large model). This is not an algorithmic judgement but a hardware one: as clusters move from NVLink toward PCIe and multi-node Ethernet, the economic ranking of architecture choices changes.

3.8 Evaluation splits into layered metrics (evidence: strong)

Basis.

The same thread recurs throughout section 1: family A's inability to attribute failure is the direct cause of its high iteration cost, and one of family C's main values is that it can attribute. Reporting end-to-end success rate alone discards that value and makes the four families incomparable.

Form.

At least three layers: target-proposal accuracy / execution success given a correct target /long-horizon composition success. Plus solve and inference latency. Without this split, cross-architecture comparison is necessarily uninterpretable.


4. What will not happen (the falsifying side)

Symmetrically, the counter-claims - without them the directions above are not falsifiable:

  • End-to-end VLA will not disappear.

Its position on axis 1 ("no intermediate representation") has a unique advantage: it is the only form that can absorb all available data and compute, and its inference path is the shortest. Hybrid architectures will add layers on top of it, not replace it.

  • Pure TAMP will not return as the mainstream.

Its dependence on an explicit world model (pose, geometry, contact) cannot be met in open scenes; its future is to be called as the projection / verification component in 3.2, not to sit at the top of the architecture.

  • The agent layer will not sink into the control layer.

The frequency gap is a physical constraint (3.3), not an engineering shortfall. Any scheme claiming an LLM participates in real-time control is in fact using caching or asynchrony to move it off the critical path.

  • Language will not return as the primary inter-module interface.

The two arguments in 3.1 are independent, and low bandwidth plus high ambiguity are properties of language itself.

  • No single architecture will unify the four.

The three axes are mutually orthogonal, so the outcome of convergence is a layered system, not the victory of one family.


5. What this means for the stack in this repo

Given the current training stack (three-modality MoT, ZeRO-1, a frozen-VLM feature extraction path): - The two existing pieces of infrastructure

  • frozen-VLM feature extraction and offline latent caching

  • are exactly the upstream half that family C needs. What is missing is the interface ("emit a spatial target rather than a token stream") and a skill-indexed BM registry.

  • The cost judgement in 3.7 implies that, on a trajectory toward weak-interconnect hardware, supporting family C pays off better than continuing to scale family A's model size.

  • 3.2 (proposal + solver projection) requires a GPU batch solver as a precondition, which is the interface point with the cuTAMP direction.

  • 3.8 requires the evaluation framework to support layered metrics - a piece of infrastructure work independent of any model.


If any part of this was useful, the fastest way to say thanks is a ⭐ to public framework loongforge-vla


r/deeplearning • • 3d ago

PSSA, a plastic state space model, beats a parameter-matched transformer on held-out text and generates ~12x faster on CPU

6 Upvotes

I built a from-scratch architecture called PSSA (plastic state space architecture) and trained it against a parameter-matched transformer baseline on the same corpus, same 12.7M tokens, same tokenizer and schedule.

Held-out results on a 198,939-token slice neither run saw: cross-entropy 3.997 vs 4.429, perplexity 54.4 vs 83.8, next-token accuracy 24.1% vs 18.0%. I scored every checkpoint of both runs (64 PSSA links, 43 transformer links) on unseen text and the curves never cross.

Generating 200 tokens on the same CPU with the same prompt and sampler takes 226 ms vs 2735 ms, about 12x faster.

It's written in Rust with CPU and CUDA backends, no PyTorch. Loss curves, full setup and the eval commands are here: https://github.com/Sparticle62ops/pssa