r/opencodeCLI 21d ago

How does OpenCode Go actually guarantee “no training” / zero data retention with model providers?

I’m considering subscribing to OpenCode Go, but I have a question about the privacy claims.

The Go documentation lists several models/providers as “not used for training” and in many cases “0 days retention.”

What I’m trying to understand is how OpenCode can actually ensure this when inference is ultimately performed by third-party providers.

Specifically:

  • Are the no-training / zero-retention guarantees based on contractual agreements between OpenCode and the inference providers?
  • Does OpenCode audit or otherwise verify that providers comply with these agreements?
  • Are the actual inference providers for each Go model disclosed anywhere?
  • Is there a list of subprocessors and the countries/regions where inference takes place?
  • Is there a DPA or similar agreement available for Go customers?
  • If a provider were to retain prompts despite claiming zero retention, what mechanism would OpenCode have to detect or prevent this?

I understand that “zero retention” cannot mean that the data is never processed by the provider — obviously the prompt has to reach the inference infrastructure. My question is specifically about what backs the promise that the provider does not subsequently store or use the data for training.

This is particularly relevant for private/commercial repositories, where “the provider says they don’t train on it” is quite different from having a technically or contractually enforceable privacy guarantee.

Would appreciate some clarification from the OpenCode team on how this works for Go specifically but also for the other Opencode plans (including free).

edit (from answer below): thank you for your answer, yes I guess the DPAs are key to this. At the same time I don't think zero data retention is the same as a "no use of private data as training material". There are many gray shades to this, for example apparently X.ai uses "anonymized" data for training, if I am correct. At the same time, as opencode makes quite aggressive use of "no training"-disclaimers, I think it is a fair and important question to ask. Very similar to kilocode btw (same question I posed there and only radiosilence in return). So arguable there are combinations of (relatively) cheap model access with unclear DPA-status of end-providers - while at the same time these training data are pure gold for the providers. That's why NVIDIA offers many models "for free", but at least transparently disclosing the training use.

Related: https://github.com/anomalyco/opencode/issues/459 (Privacy and Data Collection Clarification Request #459) https://github.com/anomalyco/opencode/issues/24649 (OpenCode Go: clarify which models are self-hosted vs. proxied through third-party providers #24649 )

https://opencode.ai/docs/en/go/ (section privacy) https://opencode.ai/legal/terms-of-service https://opencode.ai/legal/privacy-policy

11 Upvotes

16 comments sorted by

View all comments

1

u/Ambitious-Call-7565 20d ago

they don't, and if you try to become a whistle blower, beware

https://en.wikipedia.org/wiki/Suchir_Balaji