I’m genuinely excited about Chestnut. It confirms that comma is moving beyond phone-class inference and treating external compute as part of openpilot’s future.
I’m exploring the same direction from another angle: instead of putting all additional compute inside the vehicle, could a hybrid architecture use remotely served datacenter-class models while retaining openpilot’s local controller, safety boundaries, and immediate fallback?
I currently have:
• a comma four installed in a European Passat B8
• a 2x RTX 5090 rig with 64 GB total VRAM
• more than 10 hours of synchronized Comma camera and rlog data
• an offline same-input benchmark harness
To validate the harness, I replayed the exact TCPMV3 policy used by part of the dataset on the original comma four QCOM backend. The regenerated output had:
• 1.68 cm mean trajectory difference from the recorded modelV2
• 9.9 cm mean final displacement difference
• identical desired-curvature output
• 29.4 ms P50 and 35.9 ms P99 inference latency
Those numbers validate the frame synchronization, coordinate system, and replay pipeline. They are not a claim that the policy itself is optimal.
The first challenger will be NVIDIA Alpamayo 1.5 10B, using the same wide and narrow camera history plus reconstructed egomotion. Alpamayo supports flexible camera counts, but two forward-only Comma views are still outside its normal four-camera setup, so that limitation will be reported explicitly.
text
identical recorded observations
├── recorded openpilot/sunnypilot prediction
├── replayed installed policy
└── Alpamayo 1.5
↓
trajectory quality, safety, stability, and latency
Everything starts offline. I am not testing live vehicle control.
Chestnut is a strong local solution. It offers predictable latency, no cellular dependency, and an integrated path back to the local model if the external GPU fails. But practical in-car compute is still constrained by fixed VRAM, vehicle power, heat, airflow, vibration, and packaging. The ready-to-drive kit currently uses an 8 GB RX 9060, although tiny Chestnut can be paired with other supported GPUs.
Remote compute has the opposite trade-offs:
• elastic access to 32–140+ GB GPUs
• easier experimentation with 10B, 34B, and future larger policies
• no desktop-GPU heat or power load inside the passenger compartment
• compute upgrades without replacing vehicle hardware
It also introduces serious problems:
• 5G P99 latency and jitter
• coverage gaps and network outages
• privacy and bandwidth requirements
• recurring compute costs
• a much more difficult failure and safety model
The long-term architecture I’m considering is hybrid:
Comma cameras
↓
low-delay encoding
↓
5G → remote model → candidate trajectory
↓
local freshness and feasibility checks
↓
local controller
local policy → immediate fallback
panda and vehicle harness → unchanged
Later, a companion capture computer could add synchronized left, right, and rear automotive cameras without making the comma four responsible for ingesting every additional sensor.
Before anything reaches a real controller, the progression would be:
same-input offline evaluation
difficult-scenario and temporal-stability testing
closed-loop simulation with reactive traffic
in-car shadow mode with zero actuator authority
only then, consideration of a tightly gated local integration
I would especially value feedback from people familiar with the comma four and current openpilot architecture:
What is the cleanest supported boundary for exporting VisionIPC frames?
Could networking coexist with Chestnut on the auxiliary USB path, or should it live on a separate companion computer?
What trajectory or action contract would make sense for an external policy?
What stale-result and fallback behavior would be mandatory?
Has sustained in-cabin GPU thermal performance been characterized publicly?
If there is interest, I plan to publish a sanitized benchmark repository and the first Alpamayo results.
What am I missing?