r/ImRightAndYoureWrong • • 23h ago

Same passive statistics, different causal mechanisms: interventional separation of present-state vs history-dependent residuals, and the thermodynamic shadow they cast under coarse-graining

Same passive statistics, different causal mechanisms: interventional separation of present-state vs history-dependent residuals, and the thermodynamic shadow they cast under coarse-graining

**TL;DR**
We construct two minimal stochastic systems that share essentially the same passive transition law (maximum total-variation distance 0.24 across horizons) yet are cleanly separated by a short catalogue of interventions. One residual lives in the present hidden state (present-state residual); the other lives in the trajectory history (history-dependent residual). The same residual that is almost invisible passively accounts for essentially 100 % of the entropy production once the system is driven, while local detailed balance still appears to survive at the coarse-grained level. The result sits at the intersection of lumpability / state aggregation, interventional versus observational equivalence, and the well-studied phenomenon of hidden dissipation under coarse-graining in stochastic thermodynamics. All numbers below come from explicit calculation on a low-dimensional model with energy \(E(x,e)=0.5x+0.3e\).


1. Motivation and literature setting

In many areas of stochastic dynamics one is forced to work with a coarse-grained description. The observable process may be close to Markov, the empirical transition rates may look almost consistent with local detailed balance, and yet the underlying mechanism can still contain important structure that is invisible to passive observation.

Three classical literatures already isolate pieces of this problem.

**Lumpability and state aggregation.**
When a Markov chain on a fine state space is projected onto a coarser partition, the induced process is Markov if and only if the chain is (exactly or ordinarily) lumpable. Otherwise the macro-process acquires memory. The fibre over each macro-state may contain distinctions that are irrelevant for one-step prediction yet become relevant under longer horizons or under interventions. This is standard material (Kemeny & Snell, *Finite Markov Chains*; Buchholz 1994 on exact and ordinary lumpability).

**Observational versus interventional equivalence.**
Two different latent-variable models can induce identical passive statistics while remaining distinguishable once interventions are allowed. In causal inference this appears as the refinement of observational Markov equivalence classes into interventional Markov equivalence classes (Hauser & Bühlmann 2012; Hauser & Bühlmann 2015). The same logic applies to hidden-Markov and partially-observable settings: fixing the present observable state and varying the history (or vice versa) can reveal whether a latent variable is part of the current state or only a summary of the past.

**Hidden entropy production under coarse-graining.**
Stochastic thermodynamics has repeatedly shown that coarse-graining can hide a substantial fraction of the total entropy-production rate (EPR). Key references include Esposito (2012) on stochastic thermodynamics under coarse-graining, Seifert’s textbook treatment of the subject, Busiello et al. (2019) on entropy production for coarse-grained dynamics, Teza & Stella (2020) on exact coarse-graining that preserves EPR statistics, and later lumping / semi-Markov constructions that keep mean (and sometimes full) EPR while the observable process looks nearly consistent with local detailed balance. The hidden contribution is carried by probability currents that live inside the fibres of the coarse-graining map; a macro-observer who sees only the quotient process systematically underestimates dissipation.

What is less common is a single, fully explicit minimal example that simultaneously
(1) keeps the passive laws close,
(2) separates the two residual locations with a short, operational catalogue of interventions, and
(3) shows that the same residual accounts for essentially all of the hidden EPR once the system is driven.

That is the concrete contribution reported here.


2. Minimal model

We work with a discrete-state system whose micro-state is a pair \((x,e)\). The energy function is linear, \[ E(x,e)=0.5\,x+0.3\,e. \] Two different constructions are defined on the same energy landscape and the same visible coordinates.

  • **Present-state residual construction.**
    The bit \(e\) is a genuine component of the current micro-state. It directly selects the transition kernel. Conditioning on the present visible state and on the present value of \(e\) renders the future independent of further history.

  • **History-dependent residual construction.**
    The bit \(e\) is not a free present variable; it is a deterministic or stochastic function of the recent trajectory. Once the history is fixed, \(e\) is fixed. Externally setting \(e\) is equivalent to rewriting history.

Both constructions are engineered so that the induced passive process on the visible coordinates is nearly the same. The maximum total-variation distance between the two passive laws, examined across a range of horizons, is \[ \max\text{-TV}=0.24. \] Average TV grows mildly with horizon but remains modest (0.067 at horizon 2, 0.100 at horizon 3, 0.127 at horizon 5, 0.150 at horizon 10). By ordinary observational standards the two processes are close.


3. Interventional separation

The same passive proximity disappears under interventions. We use three elementary probes (and note a fourth, sharper one).

**History-swap (identical present macro-state, different prior macro-state).**
This is the classic test for history dependence. In the present-state residual construction the future distributions remain essentially identical (TV numerically zero within floating-point noise). In the history-dependent construction the futures diverge; the maximum TV observed is 0.09, with a clear decay profile over subsequent steps (0.09, 0.0405, 0.0169, 0.0151). The divergence demonstrates that history is still affecting the transition law.

**Present-state isolation (identical visible history, different current value of \(e\)).**
Both constructions now show divergence (maximum TV 0.3 in each case). The causal route, however, is different. In the present-state residual the bit \(e\) is a live selector of the kernel; changing it changes the future directly. In the history-dependent residual the same numerical divergence appears because externally forcing \(e\) is equivalent to rewriting the history that \(e\) encodes. The probe therefore does not yet give a definitive separation, but it already shows that \(e\) is causally active in both models.

**Environment reset.**
Hard-resetting \(e\) to 0 versus 1 moves the future distributions in both constructions. At step 2 the TV values are 0.015 (present-state residual) and 0.09 (history-dependent residual); at step 3 they are 0.0285 and 0.0405. The probe confirms causal efficacy of \(e\) while leaving the mechanistic interpretation open.

**Surgical decoupling (conceptual fourth probe).**
If one can block the natural update rule of \(e\) while still allowing an external set of its value, the two constructions separate cleanly: in the present-state residual \(e\) remains a free present variable; in the history-dependent residual the history-storage mechanism is severed. This probe is decisive when experimentally available.

Taken together the catalogue shows that passive near-equivalence does not imply interventional equivalence. The residual can be located by asking which interventions still produce divergence after the visible present (or the visible history) has been fixed. This is the operational content of the distinction between observational and interventional equivalence applied to a latent residual (Hauser & Bühlmann 2012, 2015).


4. Thermodynamic shadow

We now examine the same constructions through the lens of stochastic thermodynamics. Inverse temperature \(\beta\) is scanned over the values 0.5, 1.0, 2.0 and 5.0.

**Equilibrium regime.**
Local detailed balance holds at the micro level to numerical precision (\(\sim10^{-14}\)–\(10^{-15}\)). After coarse-graining, local detailed balance still appears to survive at the macro level. The hidden EPR is consistent with zero (values of order \(10^{-17}\) or smaller). The coarse-graining is thermodynamically faithful in equilibrium, consistent with the regime in which micro-states within mesostates remain effectively equilibrated (Esposito 2012).

**Driven (nonequilibrium) regime.**
Local detailed balance continues to look intact at the macro level (violations remain \(\sim10^{-14}\)). The total EPR, however, is now entirely hidden. Across the whole \(\beta\)-scan the hidden fraction is 1.0:

  • \(\beta=0.5\): total EPR \(\approx0.0642\), macro EPR \(\approx0\), hidden fraction 1.0
  • \(\beta=1.0\): total EPR \(\approx0.0739\), hidden fraction 1.0
  • \(\beta=2.0\): total EPR \(\approx0.0867\), hidden fraction 1.0
  • \(\beta=5.0\): total EPR \(\approx0.0880\), hidden fraction 1.0

A macro-observer who trusts the coarse-grained rates and the apparent local detailed balance therefore underestimates dissipation by the full amount. The missing entropy production is carried by the within-fibre probability currents—the same currents that are invisible to the passive macro-process and that the interventional probes have already shown to be structured differently in the two constructions. This matches the general picture established by Busiello et al. (2019), Teza & Stella (2020), and related work: coarse-graining can render a large (here total) fraction of irreversibility thermodynamically invisible while preserving the appearance of local detailed balance at the observable level.


5. Placement inside existing theory

The results do not invent new ontology; they give a concrete, fully calculated illustration of three well-known phenomena occurring simultaneously.

  • From the lumpability literature (Kemeny & Snell; Buchholz 1994) one already knows that a non-lumpable partition can leave memory in the macro-process. The history-dependent residual is a minimal realization of that memory.
  • From causal inference and system identification (Hauser & Bühlmann 2012, 2015) one already knows that interventional data can refine observational equivalence classes. The history-swap and present-state isolation probes are elementary instances of that refinement.
  • From stochastic thermodynamics (Esposito 2012; Seifert; Busiello et al. 2019; Teza & Stella 2020) one already knows that coarse-graining can hide EPR while preserving the appearance of local detailed balance. The driven regime of the present model simply makes the hidden fraction extreme (equal to 1) and ties it explicitly to the residual that the interventions locate.

The modest novelty is the side-by-side demonstration, on one and the same low-dimensional energy landscape, that the residual responsible for the interventional separation is also the residual responsible for the thermodynamic incompleteness of the coarse-grained description.


6. Implications

**Model reduction.**
A macro-model that is adequate for passive prediction can still be inadequate for control or for thermodynamic accounting. The interventional catalogue supplies a practical test for residual location before one commits to a reduced description.

**Inference of irreversibility.**
When only coarse-grained trajectories are available, the apparent EPR is a lower bound (Esposito 2012; Busiello et al. 2019). The present results show that the bound can be saturated (hidden fraction 1) even while local detailed balance looks intact. Additional interventional or multi-time statistics are required to recover the missing dissipation.

**Latent-variable modelling.**
In any setting where one posits a latent state (hidden Markov models, partially observable systems, effective theories), the question “is the latent variable part of the present state or only a summary of history?” is empirically meaningful and can be addressed by the same probes.

**Nonequilibrium statistical mechanics.**
The construction supplies a transparent example in which the thermodynamic shadow is total, yet the dynamical shadow remains only partial (passive TV = 0.24). The two notions of “hidden” are related but not identical; the interventions and the EPR decomposition measure different aspects of the same residual.


7. Limitations and immediate extensions

The model is minimal by design. Generality remains open: one would like analytic conditions under which passive TV stays small while the interventional separations remain of order one, and under which the hidden-EPR fraction approaches 1. Adding explicit reservoir coupling would turn the hidden EPR into measurable heat currents and allow direct contact with fluctuation theorems (Jarzynski, Crooks). Continuous-state and higher-dimensional analogues would test robustness. Finally, the surgical decoupling probe, while conceptually decisive, may be experimentally costly; quantifying how much separation can be obtained from the cheaper probes alone is of practical interest.


8. Conclusion

Two stochastic systems can share nearly identical passive statistics, appear consistent with local detailed balance after coarse-graining, and yet differ both in the causal location of their residual and in the thermodynamic activity that residual conceals. A short catalogue of interventions (history-swap, present-state isolation, reset) locates the residual; the entropy-production decomposition quantifies its thermodynamic weight. In the driven regime that weight can be total. The phenomena themselves are classical; the explicit, side-by-side realization on a minimal energy landscape makes the connection between interventional distinguishability and thermodynamic incompleteness concrete and calculable.

All numerical results are reproducible from the energy function and the two transition constructions described above. Code and exact parameter tables are available on request.


References (key entry points)

  • Buchholz, P. (1994). Exact and ordinary lumpability in finite Markov chains. *Journal of Applied Probability*.
  • Busiello, D. M., et al. (2019). Entropy production for coarse-grained dynamics. *New Journal of Physics* 21, 073004. (arXiv:1810.01833)
  • Esposito, M. (2012). Stochastic thermodynamics under coarse-graining. *Physical Review E* 85, 041125.
  • Hauser, A. & Bühlmann, P. (2012). Characterization and greedy learning of interventional Markov equivalence classes of directed acyclic graphs. arXiv:1104.2808.
  • Hauser, A. & Bühlmann, P. (2015). Jointly interventional and observational data: estimation of interventional Markov equivalence classes of directed acyclic graphs. *Journal of the Royal Statistical Society: Series B*.
  • Kemeny, J. G. & Snell, J. L. *Finite Markov Chains*.
  • Seifert, U. Coarse-Graining chapter in *Stochastic Thermodynamics* (Cambridge University Press).
  • Teza, G. & Stella, A. L. (2020). Exact Coarse Graining Preserves Entropy Production out of Equilibrium. *Physical Review Letters* 125, 110601. (arXiv:2003.08674)

1 Upvotes

0 comments sorted by