Maximum entropy is a classical principle for choosing a probability distribution when you have limited information: pick the distribution that is as spread out as possible while still satisfying whatever constraints you know, such as a fixed average value. It works beautifully as a static recipe, but it says nothing about the journey. If a system starts in one probability state and evolves toward the maximum-entropy state over time, how does probability flow from place to place? How much energy or information is consumed along the way? How do you account for the difference between entropy being created internally versus being fed in from outside? This paper builds a mathematical framework, inspired by classical mechanics and thermodynamics, that answers exactly these questions by treating the entire path of a probability distribution as the object of study rather than just its endpoints.
The core idea is to write down a Lagrangian, which is a function physicists use to encode the rules governing how a system moves, but for probability flows instead of physical particles. By tracking the full history of how probability has moved, the authors derive equations that simultaneously enforce conservation of probability, describe how the distribution evolves, and keep a clean accounting ledger separating internally generated entropy from information exchanged with the outside world. They also show that familiar results, including Bayes' theorem and the maximum-entropy principle itself, fall out naturally as special limiting cases when nothing is flowing and the system is at rest. On top of this foundation they introduce modular building blocks called information and structural potentials that can be swapped in to control features like the heaviness of distribution tails, sparsity, or the ability to represent multiple separate clusters of probability, all without disrupting the underlying accounting.
The significance of this work is that it unifies several ideas that have traditionally lived in separate communities: variational inference in machine learning, thermodynamic descriptions of nonequilibrium systems, and information-theoretic optimization. Having a single coherent framework means that algorithms for transporting probability distributions, which appear in tasks like generative modeling, optimal transport, and Bayesian updating, can now be designed with rigorous energy and entropy budgets attached. The authors verify the framework with numerical examples confirming that probability is conserved and that the energy bookkeeping adds up correctly, suggesting this could serve as a principled foundation for building more interpretable and thermodynamically consistent probabilistic models.