Ergodicity and mixing
Things that probably happen eventually on average
2011-10-17 — 2022-02-13
Wherein Ergodicity and Mixing Are Examined, and Mixing Conditions Such as Β‑ and Φ‑mixing Are Related to Finite‑sample Learning Guarantees for Dependent Data, and to Lyapunov Exponents Measuring Sensitivity.
Relevance to actual stochastic processes and dynamical systems, especially linear and non-linear system identification.
Keywords to look up:
- “probability-free” ergodicity
- Birkhoff ergodic theorem
Frobenius-Perron operatordone! It’s the Pagerank algorithm, more-or-less.- Quasicompactness, correlation decay
- C&C Nagaev CLT for Markov chains
Not much material, but please see learning theory for dependent data for some interesting categorisations of mixing and transcendence of miscellaneous mixing conditions for statistical estimators.
My main interest is the following 4-stages-of-grief kind of setup.
Often I can prove that I can learn a thing from my data if it is stationary.
But I rarely have stationarity, so at least showing the estimator is ergodic might be more useful, which would follow from some appropriate mixing conditions which do not necessarily assume stationarity.
Except that often these theorems are hard to show, or estimate, or require knowing the parameters in question, and maybe I might suspect that showing some kind of partial identifiability might be more what I need.
Furthermore, I usually would prefer a finite-sample result instead of some asymptotic guarantee. Sometimes I can get those from learning theory for dependent data.
But if I can prove nothing it is so bad? Can we prove things for the stochastic process that is the world? Is it bad if not?
1 Coupling from the past
Dan Piponi explains coupling from the past via functional programming for Markov chains.
2 Mixing zoo
A recommended partial overview is Bradley (2005). 🚧TODO🚧
2.1 β-mixing
🚧TODO🚧
2.2 ϕ-mixing
🚧TODO🚧
2.3 Sequential Rademacher complexity
🚧TODO🚧
3 Lyapunov exponents
Wilkinson’s explanation is best (Wilkinson 2016).
tl;dr: Ergodicity makes the asymptotic stretching rates described by Lyapunov exponents the same for almost every initial point (with respect to the invariant measure).
4 Incoming
🚧TODO🚧
- The World’s Simplest Ergodic Theorem
- Von Neumann and Birkhoff’s Ergodic Theorems
- The difference between statistical ensembles and sample spaces: Mehmet Süzen, Alignment between statistical mechanics and probability theory
