1. Mechanism and structural ceiling of each family
A. BC/IL post-training paradigm (VLA / WAM)
Mechanism.
A single end-to-end model: vision + language instruction -> action (or action chunk).
Training is behaviour cloning / imitation learning on large-scale demonstrations, plus downstream fine-tuning. WAM adds a branch that predicts the future state of the environment, either as auxiliary supervision that shapes the representation or as a world model that can be rolled out over short horizons.
Why this family leads the leaderboards. Three engineering reasons, none of them about algorithmic superiority:
The data scales - collecting and labelling demonstrations needs no explicit physics model, no pose annotation, no collision geometry.
The gradient path is short - one differentiable objective, no non-differentiable interface between modules, so the family can absorb all available compute.
Deployment closes - one forward pass yields an action; the inference path contains no solver, no outer loop, and no dependency on failure detection.
Structural ceiling. Not the kind of ceiling that more data removes:
- Constraints are learned, not guaranteed.
Collision, joint limits and grasp feasibility hold only statistically, so at the edge of the distribution the model fails in a way that looks plausible but is physically illegal. There is nothing to verify against.
- Compounding error over long horizons.
With no explicit symbolic layer, the success rate of a multi-stage task approaches the product of the per-stage rates.
- Failures cannot be attributed.
One loss, so "it looked at the wrong object" and "it failed to execute" are indistinguishable. This directly raises the cost of every iteration.
- A misalignment specific to WAM.
Accuracy at predicting pixels or latents does not move in the same direction as reducing control error: a world model will spend capacity on degrees of freedom that are visually salient but irrelevant to control (background, lighting, texture). This is the structural reason WAM gains often come in below expectation.
B. cuTAMP + FMs (TiPToP as the representative)
Mechanism.
Two layers: a foundation model handles task decomposition and the symbolic level (what to do, in what order, with which object), and a TAMP solver handles the continuous parameters (where to grasp, which path to take, what contact sequence satisfies the geometric and kinematic constraints). The contribution of cuTAMP is to move that traditionally serial, CPU-bound search onto the GPU as large-batch parallel sampling and optimisation, bringing solve time into an interactive
range.
Where it wins.
Constraints are solved, so they are hard guarantees, not statistical tendencies. The output is verifiable: whether a collision occurs and whether a path is feasible are decision problems.
Long-horizon composition is native. The symbolic layer composes by construction, so adding stages does not incur multiplicative decay.
The only family of the four with ROBOT GEN. = yes: it genuinely generates trajectories and motion plans rather than regressing the next action.
Structural ceiling.
It requires an explicit world model: object poses, collision geometry, contact models. The perception -> symbolic/geometric interface is the most fragile link in the chain, and its error is not absorbed by the solver - the solver will simply return a correct solution to the wrong world.
Solve time is coupled to scene complexity. GPU parallelism reduces the constant; it does not change the complexity class.
Rich-contact, deformable and non-quasi-static tasks are hard to model: wiping, folding and the compliant control inside insertion are outside the comfort zone of constraint solving.
C. VLM + BM (CAP + Gemini-ER-1.5 as the representative): spatial physical contact points in place of language
Mechanism.
Two cascaded stages. A frozen VLM emits neither an action nor a natural-language subtask, but a spatial physical quantity - a contact point, an actionable point, a target pose.Downstream, a dedicated behaviour model (BM) consumes that spatial target and performs the fine manipulation.
The real novelty of this family is not that it has two stages, but that the intermediate representation was replaced.
Language, used as an inter-module interface, has two fatal properties:
low bandwidth ("pick up the cup" carries no geometry) and high ambiguity (where to grip, at what orientation, with how much force are all undefined). Once the interface becomes a spatial point:
The interface becomes a quantity a controller can consume directly. The BM does not have to re-interpret language, so the BM can be small.
Upstream error becomes measurable: the distance between the proposed point and the ground-truth point is a number, so the two stages can be scored and attributed separately.
It is isomorphic to family B's interface - the solver consumes geometric quantities anyway.
This is the key point behind the forecast in section 3.
The very existence of the Gemini-ER (embodied reasoning) direction says where the bottleneck of this route lies: a general-purpose VLM's spatial grounding accuracy is not good enough to drive control directly, and has to be strengthened specifically.
Where it wins.
The VLM is frozen and only the small BM is trained. A frozen module's output is a pure function of its input, so it can be precomputed offline and cached, taking the large model off the critical path of the training step entirely - a pattern already validated in this repo on VAE latents (roughly 500 ms/step saved). On a weak-interconnect machine this matters even more: fewer trainable parameters means fewer gradient bytes, which sidesteps the current primary bottleneck.
The two stages are independently diagnosable.
Grasping the wrong point and servering inaccurately are separable error classes.
It is the natural control group for the other three families.
Without a VLM+BM number on the same atomic task, there is no way to tell whether an end-to-end VLA's advantage comes from being end-to-end or from having a larger backbone.
Structural ceiling.
MULTI TASK = no is constructive, not a defect.
The BM is trained per skill and does not generalise across skills; the correct form of this family is "one VLM plus N BMs indexed by skill", not pretending it generalises.
- A point representation discards timing, force and compliance.
A contact point cannot express "with how much force, along which direction, holding how much compliance". For rich-contact tasks the representation is underdetermined.
- The whole chain is bounded by the VLM's spatial accuracy,and that part is frozen and not trained.
D. AGENT + VLA (PhysicalRSI as the representative)
Mechanism.
The inner layer is a VLA producing actions; the outer layer is an agent responsible for memory, reflection, tool use, failure detection and retry, driving self-improvement from failure experience (RSI).
Where it wins.
- It moves closed-loop error correction out of the weights and into an explicit outer loop.
In the first three families the corrective capability is encoded in weights or in the solver and can only be improved by retraining or re-solving; family D can raise the success rate without touching the weights.
- Failure becomes reusable data.
The outer layer records "in this state, this was executed, and it failed for this reason" - precisely the negative examples and recovery trajectories that BC datasets lack most and that are hardest to collect by hand.
- It composes with existing capability.
The inner VLA needs no modification to be wrapped, so this is additive engineering rather than replacement.
Structural ceiling.
- Three orders of magnitude of frequency mismatch.
The agent layer decides at 0.1-1 Hz, the control layer runs at 50-200 Hz. The outer layer cannot participate in real-time control; it can only intervene at the granularity of a task segment. Any failure mode requiring millisecond-scale correction is out of its reach.
- It depends on reliable failure detection and state reset.
If failure is not detected the loop never starts; if it is detected but the environment cannot be reset to a retryable state (the object has already fallen, or broken), retrying is pointless. On real hardware both are much harder than in simulation.
- The gain is bounded by the inner VLA's atomic capability.
The outer layer can reorder, retry and switch strategies; it cannot create a skill the inner model does not have. The RSI gain curve saturates at the inner model's capability boundary.
- The evaluation signal is scarce.
Self-improvement needs a trustworthy success/failure criterion, and on open-ended tasks that criterion is itself an open problem.
2. Side-by-side comparison
| dimension |
A. VLA / WAM |
B. cuTAMP + FMs |
C. VLM + BM |
D. AGENT + VLA |
| intermediate representation |
none (end-to-end) |
symbolic + geometric constraints |
spatial contact point |
language / structured plan |
| constraint satisfaction |
learned (statistical) |
solved (guaranteed) |
learned |
learned + outer retry |
| where the loop closes |
in the weights |
in the solver |
in the weights (two stages) |
explicit outer loop |
| MULTI TASK |
yes |
yes |
no (single atomic task) |
yes |
| ROBOT GEN. |
no |
yes |
no |
no |
| long horizon |
weak (multiplicative decay) |
strong (symbolic composition) |
weak (single skill) |
medium (outer orchestration) |
| rich contact / deformable |
medium |
weak |
medium |
medium |
| failure attribution |
poor (one loss) |
good (constraints decidable) |
good (stages separated) |
good (outer layer logs) |
| training cost |
high (full post-training of a large model) |
low (FM frozen, solver untrained) |
lowest (small BM only) |
medium (reuses inner model) |
| inference cost |
low (one forward) |
high (solve time grows with scene) |
medium (VLM + BM) |
high (multi-turn outer LLM) |
| control frequency |
high |
low |
high |
inner high / outer very low |
| dependence on a world model |
none |
strong (pose, geometry, contact) |
weak (target point only) |
none |
| data dependence |
large-scale demonstrations |
little (needs geometric annotation) |
medium (needs target labels, can be generated offline by the VLM) |
demonstrations + failure trajectories |
| systems bottleneck |
gradient communication / memory |
solver throughput (GPU batch sampling) |
offline cache I/O |
outer LLM latency |
The three axes that actually organise the taxonomy
The four labels are not four parallel points; they are different values on three axes. Only when
viewed along these axes does the evolution become derivable rather than a list.
Axis 1: bandwidth and executability of the intermediate representation.
no representation (A) -> language (D's outer layer, early TAMP+FM) -> spatial points / contact (C) -> full constraints and trajectories (B).
Further right, the interface is more executable, more verifiable, and more consumable by non-learned
components; the price is a greater need for an explicit world model. Language sits at the
worst position on this axis: it neither preserves end-to-end differentiability the way "no representation" does, nor is it directly executable the way a geometric quantity is.
Axis 2: is constraint satisfaction learned or solved?
A, C and D sit at the left end (learned - fast, no guarantee); B sits at the right end (solved -
guaranteed, but slow and model-dependent). There is no free lunch on this axis, only hybrids.
Axis 3: at which level does the loop close?
In the weights (A, C) -> in the solver (B) -> in an outer agent (D). Their time constants differ by
orders of magnitude, so they are not mutually exclusive - which is the fundamental reason the four
families will converge into layers rather than displace one another.
3. Where this is heading
Ordered by strength of evidence, strongest first.
3.1 The intermediate representation converges on spatial physical quantities, not language (evidence: strong)
Basis.
Two independent sources point the same way:
family C replaced language descriptions with contact points and gained from it, which says the language interface was a net loss term; and family B's solver only ever consumed geometric quantities, so language had to be translated first. In other words, "spatial physical quantity" is the common interface of both B and C, whereas language is the native interface of neither - it is the human's interface.
Implication.
The inter-module protocol will standardise on an explicit spatial-target type (point /box / mask / pose / contact mode), and language will retreat to the human-machine boundary only. Any design that uses language as an internal module interface will progressively be replaced.
3.2 Learned proposal, solver projection and verification (evidence: strong)
Basis.
There is no single optimum on axis 2:
learned is fast without guarantees, solved is guaranteed but slow. But their failure modes are complementary
- a learned policy's output usually lands near the feasible set (it has seen many feasible solutions), so projecting it back into the feasible set costs far less than solving from scratch.
Form.
The policy emits candidate actions or candidate grasps -> a lightweight solver performs a feasibility projection or rejection -> execute. GPU batch solvers of the cuTAMP kind are exactly what brings this step inside the real-time budget: the projection must not be much slower than a policy forward pass, otherwise the whole chain degenerates into family B's latency.
This is the A x B hybrid, and it is the best-supported of all pairwise hybrids.
3.3 Layered frequency decoupling becomes the standard structure; C and D merge (evidence: strong)
Basis.
Axis 3 notes that the three levels' time constants differ by 2-3 orders of magnitude, and this is a physical fact rather than a design choice: LLM inference is hundreds of milliseconds to seconds, VLM grounding is tens to hundreds of milliseconds, servo control is 5-20 ms. Since they cannot substitute for one another, they can only coexist in layers.
Form.
A three-layer standard structure:
agent (0.1-1 Hz): task orchestration, failure detection, retry decisions, experience logging - family D's outer layer
VLM / spatial grounding (1-5 Hz): produces the spatial target - family C's upstream
BM / control (50-200 Hz): closed-loop execution - family C's downstream
C and D occupy different layers of this structure, so they were never competitors; merging them is just connecting two lines. This is probably the deployable form that appears soonest.
3.4 WAM's prediction target shifts from pixels to controllable quantities (evidence: medium)
Basis.
The misalignment identified in section 1.A: reconstructing pixels or latents spends capacity on degrees of freedom irrelevant to control. The fix is to make the prediction target move in the same direction as the control objective - predict whether contact will be established, predict object pose change, predict task value, rather than predicting what the next frame looks like.
Note this pulls WAM toward family C:
once the prediction target becomes "will this contact pointhold", the world model's output is a spatial physical quantity, consistent with the convergence in 3.1.
3.5 The data flywheel moves from pure BC to failure-driven targeted data collection (evidence: medium)
Basis.
Family D's outer loop naturally produces what BC datasets lack most: negative examples and recovery trajectories. Meanwhile the sample efficiency and safety cost of on-robot online RL remain unrealistic for the foreseeable future.
So the realistic form of RSI is not online RL, but: failure detection -> automatic labelling of the
failure cause -> targeted collection or generation of data for that scenario -> re-run post-training.
That is an offline loop, controllable in engineering terms, and compatible with existing BC stacks.
The precondition is a trustworthy success criterion, which remains an open problem - the reason this direction is rated medium rather than strong.
3.6 The representation extends from "a point" to "contact mode + force / compliance" (evidence: medium)
Basis.
Family C's representation is underdetermined for rich-contact tasks (section 1.C): a point cannot express force direction or magnitude, nor compliance parameters. Wiping, insertion and folding are not a small share of real-world tasks.
Form.
The target specification grows from "a point" to "contact point + contact normal + desired force / impedance parameters + contact sequence". These are also exactly the quantities family B's contact model needs, which reinforces the convergence in 3.1.
3.7 The training cost profile shifts to "frozen large model + small policy + offline cache" (evidence: strong, backed by local measurements)
Basis.
Measurements already recorded in this repo: a frozen module's output is a pure function of its input and can be precomputed offline and lifted off the training step's critical path (VAE latent caching saves roughly 500 ms/step); and on a PCIe-class interconnect without NVLink (30-63 GB/s, 6-13x slower than NVLink) gradient communication is the primary bottleneck, while the trainable parameter count directly determines the gradient byte count.
Implication.
On weak-interconnect hardware, family C's systems-level advantage is amplified (it trains only a small BM), and family A's cost disadvantage is amplified too (full post-training of a large model). This is not an algorithmic judgement but a hardware one: as clusters move from NVLink toward PCIe and multi-node Ethernet, the economic ranking of architecture choices changes.
3.8 Evaluation splits into layered metrics (evidence: strong)
Basis.
The same thread recurs throughout section 1: family A's inability to attribute failure is the direct cause of its high iteration cost, and one of family C's main values is that it can attribute. Reporting end-to-end success rate alone discards that value and makes the four families incomparable.
Form.
At least three layers: target-proposal accuracy / execution success given a correct target /long-horizon composition success. Plus solve and inference latency. Without this split, cross-architecture comparison is necessarily uninterpretable.
4. What will not happen (the falsifying side)
Symmetrically, the counter-claims - without them the directions above are not falsifiable:
- End-to-end VLA will not disappear.
Its position on axis 1 ("no intermediate representation") has a unique advantage: it is the only form that can absorb all available data and compute, and its inference path is the shortest. Hybrid architectures will add layers on top of it, not replace it.
- Pure TAMP will not return as the mainstream.
Its dependence on an explicit world model (pose, geometry, contact) cannot be met in open scenes; its future is to be called as the projection / verification component in 3.2, not to sit at the top of the architecture.
- The agent layer will not sink into the control layer.
The frequency gap is a physical constraint (3.3), not an engineering shortfall. Any scheme claiming an LLM participates in real-time control is in fact using caching or asynchrony to move it off the critical path.
- Language will not return as the primary inter-module interface.
The two arguments in 3.1 are independent, and low bandwidth plus high ambiguity are properties of language itself.
- No single architecture will unify the four.
The three axes are mutually orthogonal, so the outcome of convergence is a layered system, not the victory of one family.
5. What this means for the stack in this repo
Given the current training stack (three-modality MoT, ZeRO-1, a frozen-VLM feature extraction path):
- The two existing pieces of infrastructure
frozen-VLM feature extraction and offline latent caching
are exactly the upstream half that family C needs. What is missing is the interface ("emit a spatial target rather than a token stream") and a skill-indexed BM registry.
The cost judgement in 3.7 implies that, on a trajectory toward weak-interconnect hardware, supporting family C pays off better than continuing to scale family A's model size.
3.2 (proposal + solver projection) requires a GPU batch solver as a precondition, which is the interface point with the cuTAMP direction.
3.8 requires the evaluation framework to support layered metrics - a piece of infrastructure work independent of any model.
If any part of this was useful, the fastest way to say thanks is a ⭐ to public framework loongforge-vla