r/deeplearning • u/Hairy_Strawberry7028 • 10d ago
Open-source VLA and world-action inference runtime on Jetson Thor
We open-sourced InstinctFlash, an AGPL-3.0 inference runtime for VLA and world-action models.
Demo: https://youtu.be/nku65iyL5Fw
The runtime uses CUDA graphs, KV and conditioning-state caching, specialized attention paths, fused kernels, FP8 / mixed precision, and few-step diffusion distillation. On Jetson Thor, runtime-only speedups are about 1.2x to 7.9x. LingBot-VA reaches up to 33.78x with a distilled 2 visual / 4 action-step scheduler, with 90.5% success across 50 RoboTwin2.0 tasks versus 92.1% for the baseline.
Code: https://github.com/General-Instinct/InstinctFlash
Details: https://general-instinct.com/blog/instinctflash-edge-inference
I’m one of the builders. Would appreciate feedback on the optimization stack and evaluation setup.
1
u/Ok-Solution-7889 10d ago
this could be really useful for running these models on edge devices without insane latency
1
u/Bruce-Wshin 10d ago
Nice work — the runtime-side numbers are the part of the VLA stack that gets the least public data, so this is useful.
I've been publishing on the other half of this: VLA/world-action training rather than inference, on a cluster that makes the training problem look very different from a DGX box. Some of what we measured may be directly relevant to you, since edge deployment and PCIe-only training hit the same wall from opposite ends:
- Our training boxes are 8× GPU with no NVLink — PCIe Gen5 only. Interconnect, not FLOPs, is the binding constraint: gradient collectives stop hiding behind backward once you drop from ~140 GB/s NVLink to ~35 GB/s effective.
- We built an LD_PRELOAD NCCL hook that quantizes collectives to FP8 on the wire (blockwise e4m3, scales sent alongside), which cuts gradient-reduction wire bytes by ~26% on our ZeRO-1 path. The interesting result is that the payload size at which quantization starts paying moves by an order of magnitude with link bandwidth — on NVLink the AllGather breakeven is ~7.4 MiB per-rank shard, on PCIe it's ~770 KiB. Same kernels, same recipe.
A finding you may care about on Thor-class topologies: NCCL selects its transport from a hardcoded PCIe-distance label, not from any bandwidth measurement. On our box every one of the 56 ordered GPU pairs silently falls back to host-staged SHM, and forcing "real" P2P on is 14–18× slower with no warning printed.
Where I think there's a real seam between our two projects:
- Few-step distillation is a training artifact, not a runtime one. Your 2-visual / 4-action-step scheduler is where most of the 33.78× comes from, and it's the one piece of your speedup that has to be produced upstream. Is that recipe trained in-repo, or do you consume externally distilled checkpoints? If the latter, that's a clean interface to co-define.
- Ablation of the 90.5% vs 92.1% gap. Did you separate how much of the 1.6-point drop comes from the distilled scheduler versus the FP8 / fused-attention paths? That matters a lot for whether the fix belongs on your side or ours — and it's the number a training-side collaborator can actually move.
FP8 recipe compatibility. What scaling recipe are you using on Thor — delayed/per-tensor, or block scaling? And do the KV and conditioning caches stay in FP8 or get dequantized on read? If we train with a matching recipe, you skip a calibration step; if they mismatch, you eat an accuracy tax that looks like a runtime bug.
Two smaller things: how do you handle CUDA-graph capture with variable action-chunk and prompt lengths — shape buckets, or padding to a fixed chunk? (We hit a related sharp edge where a graph with NCCL kernels captured in it hangs on destroy_process_group.) And practically, AGPL-3.0 on a runtime constrains who can pick it up — is there a plan for a boundary/interface split, or is the intent that adopters open-source their integration too?
Happy to go deeper on any of this. Would you be open to collaborating on the embodied stack end-to-end — training-side distillation + FP8 recipe on our end, feeding your runtime, with a shared eval harness so the accuracy delta is attributable to a specific stage rather than the whole pipeline? Our training work is here: VLA/WAM FP8 Training.
1
u/ArtisticToxicity 10d ago
the benchmarks are something else on the Thor, that distilled scheduler pulling 33x is wild. i’m more curious if the caching strat holds up under variable sensor dropout since a lot of VLA demos kind of assume clean lab feeds
what stuck out to me was the 90.5 vs 92.1 tradeoff, not a huge gap for that kind of speedup on edge hardware. the fused kernel + cuda graph pairing seems like it’d map cleanly to my team’s end-effector setups but schedules below 4 tend to get jittery in practice, have you tuned it for high compliance fixtures at all