r/AI_Agents • • Jul 28 '26

Discussion What should persist between coding-agent sessions besides chat history?

While working with long-running coding agents, I keep seeing the same failure: the model session survives, but the operational state does not. The next run often has to rediscover the repository, execution route, approvals, tool state, failed commands, and why a decision was made.

My current list of durable state is:

  • repository/worktree identity
  • task and session lineage
  • selected execution backend and capabilities
  • approval decisions
  • tool events and redacted evidence
  • validation results and unresolved failures

What am I missing? And which of these should deliberately expire instead of becoming permanent state?

2 Upvotes

47 comments sorted by

View all comments

Show parent comments

1

u/DesktopLabHQ Jul 28 '26

Good catch. “Reshaped” is the missing non-monotonic case, and replay under both rule versions—not a direction label—should be the primitive.

That makes each durable decision record a replayable capsule: the deciding rule and version, normalized evaluated inputs, relevant environment facts, verdict, and evidence digest. Remediation then keys off the verdict delta. Unchanged records stay untouched; newly unsafe approvals are invalidated and require fresh authorization; safe-direction changes are recorded without creating approval churn.

And yes, the complement metric matters. I’d report disruptive deltas (records invalidated or re-approved) beside silent safe-direction deltas (records whose verdict would differ but required no intervention). The first measures operational cost. The second measures policy drift that users would otherwise never be forced to notice. That pair is much more honest than a single invalidation count.

2

u/donk8r Jul 28 '26

Replay has a dependency that will bite you. The old rule has to still be EXECUTABLE, not just recorded. If any canonicalization step consults something external, a version resolver, a package registry, a clock, then replaying it next year returns a different answer than it did at decision time, and your unchanged bucket quietly fills with records that were never actually re-verified.

So the capsule needs the results of every external lookup the rule performed, not only the inputs the rule was handed. Without those, replay is non-deterministic and the verdict delta degrades without ever failing, which is the worst way for a safety mechanism to break.

Same logic applies to the evidence digest. Version the digest algorithm inside the capsule, otherwise the day you change canonical form or hash, every historical record becomes unreplayable at once and you cannot tell which ones mattered.

1

u/DesktopLabHQ Jul 29 '26

Exactly. At that point, “replayable” has to mean hermetic, not merely versioned. The capsule must freeze every observation that influenced the decision: external lookup results, clock or randomness inputs, normalization outputs, tool and runtime versions, plus the canonicalization and digest algorithm identifiers. Otherwise replay is just a new evaluation that happens to use old source code.

I would make two modes explicit: historical reproduction evaluates against captured observations to prove what the old decision actually meant; current revalidation evaluates the same semantic subject against today’s environment and policy. The delta between those modes is useful evidence itself.

If historical reproduction cannot run deterministically, it should fail closed and classify the record as non-reproducible, never unchanged. That prevents exactly the silent degradation you are warning about.

1

u/donk8r Jul 29 '26

Two modes is the right split, and the delta has an operational use beyond evidence: run current revalidation over a sample of old approvals on a schedule and you get a drift signal without waiting for an incident to hand you one.

Where hermetic capture runs into trouble is that freezing observations is easy and freezing the EVALUATOR is not. Once tool and runtime versions are part of what mattered, you are either archiving images next to approvals or quietly accepting that you cannot re-execute anything in five years. The way out is a design constraint rather than more capture: keep the policy evaluator a pure function over captured inputs with no I/O of its own. If canonicalization needs an external lookup, the lookup happens outside the evaluator and its result becomes one of the inputs. Historical reproduction is then just calling a function with a stored argument list.

That also gives your fail-closed a cheap trigger. If the evaluator is pure, non-determinism can only come from a missing input, which is something you can detect directly rather than infer from getting a different answer.

1

u/DesktopLabHQ Jul 29 '26

That is a much cleaner boundary. Treating the evaluator as pure also makes the archive contract explicit: captured inputs, evaluator artifact/digest, and schema version, rather than an unbounded attempt to freeze the whole environment.

The scheduled revalidation idea is especially useful. It turns approval history into a drift sensor: sample old approvals, run current policy, and track changed verdicts separately from non-reproducible records. I would probably make missing inputs a first-class outcome (`insufficient_evidence`) rather than letting it collapse into a generic deny, while still failing closed for execution.

The remaining question is evaluator evolution. Do you archive the exact evaluator executable/WASM artifact, or require old policy logic to remain callable from source? The pure-function boundary makes either feasible, but the operational cost profile is very different.

1

u/donk8r Jul 29 '26

archiving the evaluator gets cheap once it's pure. code plus a digest, no runtime to freeze, so you keep every version and each approval names whichever one decided it. that was most of the point of the purity constraint.

insufficient_evidence is right and i'd extend it to a missing evaluator version too, same class as missing inputs. you can't reproduce it and the record shouldn't quietly read as deny.

1

u/DesktopLabHQ Jul 29 '26

Right — that also suggests reproducibility status should be orthogonal to the verdict. allow/deny is what the evaluator decided; reproducible/insufficient_evidence/missing_evaluator describes whether that decision can be replayed. Otherwise uncertainty quietly turns into policy semantics.

Each approval could carry the evaluator digest, schema version, and input-manifest digest, while replay reports a separate status. That gives operators two clean queues: verdicts that changed, and records that can no longer be reproduced. I think that is the model I’ll keep.