A step-by-step explainer
Why a system "wants" to do something is never just about energy. By the end of this page, free energy — and why it, not energy, is the thing that decides what happens — should feel obvious.
The whole idea, in one sentence
Free energy is energy discounted by how many ways there are to arrange it — nature doesn't seek the lowest energy, it seeks the best deal between low energy and lots of room to wiggle.
Step 1 · The puzzle
Drop a ball and it rolls downhill: systems move toward lower energy. That intuition is so reliable we assume it's the whole story. But it can't be — ice melts in a warm room, gases fill their container, proteins partly unfold at body temperature. In every one of those, the system spontaneously moves to higher energy. Something is competing with energy and winning.
That something is counting. A spread-out, jumbled arrangement can be realized in vastly more ways than a neat, compact one. Below, a handful of particles can sit in a tidy low-energy cluster or spread across the box. Raise the temperature and watch which the system prefers — even though spreading out costs energy.
Step 2 · The other half
Give "the number of ways" a name: entropy. Boltzmann's idea is almost embarrassingly simple — entropy is just the (log of the) count of microscopic arrangements consistent with what you can see:
\[ S = k_B \ln \Omega \]
\(\Omega\) is that count of microstates, and \(k_B\) is Boltzmann's constant (the conversion factor between "counting" and "energy units"). The logarithm is what makes entropy add up: two independent systems have \(\Omega_1\Omega_2\) joint arrangements, and \(\ln(\Omega_1\Omega_2)=\ln\Omega_1+\ln\Omega_2\).
Slide the number of particles below and watch \(\Omega\) — the count of ways to place them — explode, while \(\ln\Omega\) grows gently and linearly. That gentle, additive quantity is the real currency.
More generally, for a probability distribution \(p_i\) over microstates, the Gibbs entropy is \( S = -k_B \sum_i p_i \ln p_i \), which reduces to \( k_B\ln\Omega \) when all \(\Omega\) accessible states are equally likely (\(p_i = 1/\Omega\)). This is the same functional form as Shannon entropy up to the constant \(k_B\) and the choice of log base — entropy is the missing information about which microstate the system is in.
Step 3 · What energy is
We've been counting microstates abstractly. Here is where they come from physically. Picture a particle in a potential well made of two stacked Gaussians: a wide, shallow basin with a narrow, steep well sunk into its center. Its energy is split between motion (kinetic) and depth (potential), and the dynamics endlessly trade one for the other while keeping the total fixed. That's all a Hamiltonian does: shuffle energy between kinetic and potential along a path of constant total energy.
We're looking straight down on the well, colored by depth — the bright core is the steep central well, the broad tint is the shallow basin. The orange ring is the energy level \(V=E\): the particle can go anywhere inside it and turns back right at the ring, where all its energy is momentarily potential and none kinetic. Energy is literally how far the particle can roam. Raise it and watch the reachable region widen — and notice how it jumps outward once the energy clears the steep core and the wide basin opens up.
Now discretize. Lay a fine grid over the map: each cell is one single-particle state. A cell is accessible if it lies inside the energy contour (a particle could be there without violating energy conservation). Call the number of accessible cells \(M\). With a single particle there are \(M\) microstates and \(S=\ln M\); the lone particle traces a thin ribbon, so turn up the particle count or just wait to see the \(M\) cells really are all available.
But with \(N\) particles, each free to occupy any of the \(M\) cells, the system's microstates are the combinations of those choices — exactly the "\(N\) particles in \(M\) boxes" counting from earlier:
\[ \Omega = M^{N}, \qquad S = \ln\Omega = N\ln M. \]
Here's the payoff that ties this page together: raising the energy widens the basin, so \(M\) grows; and \(\Omega=M^N\) — hence \(S=N\ln M\) — grows both with the energy and with every particle added. That is "adding energy adds entropy," in continuous phase space, before we abstract it. The next step strips the bowl away and keeps only the count, as labeled quanta dropped into bins. (The grid counter tracks \(M\) and the single-particle cells explored; \(\Omega\) is then \(M^N\).)
Each particle obeys Hamilton's equations for \(H=\tfrac12|\mathbf p|^2 + V(\mathbf q)\) with the two-Gaussian well \(V\); the integrator is velocity-Verlet (symplectic), so the energy shell \(H=E\) is conserved up to \(O(\Delta t^2)\) bounded error. The classically accessible region is \(\{\mathbf q : V(\mathbf q)\le E\}\); discretized, it holds \(M(E)\) single-particle cells. For \(N\) independent, distinguishable particles each on that shell the joint count is \(\Omega = M^{N}\) and \(S = k_B\ln\Omega = N k_B\ln M\) — extensive in \(N\), matching the "\(N\) in \(M\) boxes" counting used earlier (identical particles would carry a Gibbs \(1/N!\); for \(M\gg N\) this shifts \(S\) by a constant and doesn't change the picture). Because \(V\) is confining, \(M(E)\) increases with \(E\), so \(\partial S/\partial E = 1/T > 0\) recovers temperature from a purely mechanical picture. Ergodicity is why a long single trajectory, or many short ones, fill the accessible cells — the assumption underneath equating time averages with the microstate count.
Step 4 · The hinge
We now have the two players — energy, which systems shed, and entropy, which they accumulate. The natural question is what couples them. The answer is the most important and least intuitive idea on this page: temperature is not "how hot something is." Temperature is defined entirely in terms of entropy and energy, as the rate one trades for the other:
\[ \frac{1}{T} \;\equiv\; \frac{\partial S}{\partial E}. \]
Read it literally: \(\partial S/\partial E\) is "how many extra arrangements you unlock per unit of energy you pour in." A system that gains a lot of entropy from a little energy has a large \(\partial S/\partial E\), hence a small \(T\) — it is cold, and energy-hungry: feed it energy and its disorder shoots up. A system that barely gains entropy from more energy has a small slope, hence a large \(T\) — it is hot and energy-indifferent: it already has so many arrangements that more energy hardly adds any.
So temperature is the local steepness of the \(S(E)\) curve, flipped. Drag the point along the entropy curve below and watch the slope — and therefore \(T\) — change. Notice the curve always bends the same way: each extra unit of energy buys fewer new arrangements than the last (the slope flattens). That single fact — \(S(E)\) is concave — is why heating something always raises its temperature.
But before the curve, see why entropy rises at all. Below, the six particles are drawn as bins, and the energy you add as labeled quanta scattered among them. The point isn't where any particular quantum lands — they carry no real identity, the labels are just so you can watch them move. The point is how many ways there are to scatter them. Each quantum can fall into any of the six particles independently, so every quantum you add multiplies the number of configurations by six: \( \Omega = 6^{E} \). At zero energy there is one configuration and \(\Omega = 1\); add energy and \(\Omega\) explodes. Entropy isn't poured in directly — it is the ballooning count of ways to arrange the energy. Hit resample to jump to another equally-likely scattering.
Counting this way, \(\ln\Omega = E\ln N\) grows in a perfectly straight line — each quantum adds the same \(\ln N\) of entropy. The real story has one subtle twist: energy quanta are genuinely identical, so swapping two of them is not a new state at all. That indistinguishability trims the count (from \(N^E\) down to the number of distinct splits, \(\binom{E+N-1}{E}\)) just enough to bend the straight line into a gently curving \(S(E)\) — one whose slope shrinks as energy grows. That curvature is exactly what makes temperature rise as you heat. Drag the point along the curve and watch the slope — and therefore \(T\) — change.
Here's the payoff, and the reason this definition is the spine of everything: that one slope does two jobs at once.
Job one — it governs a system's own response. The slope at a system's current energy is its temperature. Nothing else needed.
Job two — it sets the price of trading with a bath. Put two systems in contact and let energy flow. Total entropy changes by \( dS_\text{tot} = \big(\frac{1}{T_A}-\frac{1}{T_B}\big)\,dE_A \), so energy flows from the shallower-slope system (small \(\partial S/\partial E\), high \(T\)) into the steeper-slope one (low \(T\)) — because that move creates more entropy than it destroys. Energy flows "hot to cold" for no other reason than that it increases the total count of arrangements. It stops exactly when the two slopes match: \(T_A = T_B\). That equality is thermal equilibrium.
Below, two systems with adjustable starting energies are brought into contact. Press play and watch energy flow until their \(S(E)\) slopes — their temperatures — line up, while the total entropy climbs to its max.
The thermodynamic definition is \(\frac{1}{T} = \left(\frac{\partial S}{\partial E}\right)_{N,V}\), with the derivatives of \(S=k_B\ln\Omega\) taken at fixed particle number and volume. For two systems exchanging energy at fixed total \(E=E_A+E_B\), maximizing \(S_A(E_A)+S_B(E_B)\) gives \(\frac{\partial S_A}{\partial E_A}=\frac{\partial S_B}{\partial E_B}\), i.e. \(T_A=T_B\). Concavity of \(S(E)\) (equivalently, positive heat capacity \(C=\partial E/\partial T>0\)) guarantees this is a maximum, not a saddle, and that the equilibrium is stable. The quanta-exchange simulation a couple of steps down is this exact argument realized microscopically: the bath's near-constant slope \(-1/T\) is what makes \(\ln\Omega_\text{bath}\) linear in the system's energy, which is the origin of the Boltzmann factor \(e^{-E/k_BT}\).
Step 5 · The resolution
It's tempting to present \(F = E - TS\) as a clever definition: put energy and entropy on one ledger, weight entropy by temperature, done. But that just relabels the mystery — why that combination, why multiply by \(T\), why subtract? None of it is a choice. Every piece is forced by a single fact we've been ignoring: the system is never alone.
A protein in water, a gas in a room, a spin in a magnet — each sits in contact with a vast bath (a "reservoir") that it freely trades energy with. The second law doesn't say the system's entropy increases. It says the total entropy of everything — system plus bath — increases. So the quantity nature actually maximizes is
\[ S_\text{tot} = S_\text{sys} + S_\text{bath}. \]
This is the whole derivation, and it's three short steps:
1 — The bath's entropy depends only on the energy it has left. Suppose the system grabs energy \(E\) for itself. That energy comes out of the bath, so the bath now has \(E\) less. How much entropy does the bath lose? By the definition of temperature, \( \frac{1}{T} \equiv \frac{\partial S_\text{bath}}{\partial E_\text{bath}} \) — temperature just is "how much bath entropy you gain per unit of energy you hand the bath." Removing \(E\) from the bath therefore costs it
\[ \Delta S_\text{bath} = -\frac{E}{T}. \]
There is your \(T\), and there is your minus sign — neither was a modeling choice. The \(1/T\) is the bath's exchange rate between energy and entropy; the minus sign is because energy the system takes is energy the bath loses.
2 — Add up the total. Substitute into \(S_\text{tot}\):
\[ S_\text{tot} = S_\text{sys} - \frac{E}{T}. \]
3 — Repackage from the system's point of view. The system can't see the bath's internal states; it only knows its own \(E\) and \(S_\text{sys}\). Multiply the line above by \(-T\) (a positive number, so it flips "maximize" into "minimize"):
\[ -T\,S_\text{tot} = E - T\,S_\text{sys} \;\equiv\; F. \]
So \(F = E - TS\) is nothing but the total entropy of the universe, rewritten in energy units and with its sign flipped. Maximizing \(S_\text{tot}\) and minimizing \(F\) are the same statement. The two questions answer themselves:
Why multiply entropy by \(T\)? Because the bath converts energy into entropy at rate \(1/T\). To compare the system's own entropy against the energy it borrows from the bath, you must run that entropy back through the same exchange rate — multiplying \(S\) by \(T\) puts it in energy units. A hot bath (\(T\) large) barely gains entropy from extra energy, so the energy term dominates; a cold bath (\(T\) small) gains a lot, so even tiny energy savings matter enormously.
Why subtract? Because energy the system hoards is entropy the bath loses. The system's gain (\(E\)) and the universe's loss (\(E/T\) of bath entropy) pull in opposite directions, so they enter with opposite signs.
The intuition builder below makes this concrete. It plots the universe's entropy as the system's energy varies, built from its two pieces — and shows that the peak of total entropy sits exactly where \(F\) is smallest. Reveal the layers, then drag the bath temperature and watch the balance point move.
Expand \(S_\text{bath}(E_\text{bath})\) to first order around the system's typical energy: \(S_\text{bath}(E_0 - E) \approx S_\text{bath}(E_0) - E\,\partial S_\text{bath}/\partial E_\text{bath} = \text{const} - E/T\). The linear term is all that survives because the bath is enormous (its second derivative \(\propto 1/C_\text{bath}\) vanishes as the heat capacity \(\to\infty\)) — this is exactly what "reservoir" means. Then \(S_\text{tot}=S_\text{sys}-E/T+\text{const}\), and \(F=-T\,S_\text{tot}\) up to that constant.
Maximizing \(S_\text{tot}\) over the system's energy gives \(\partial S_\text{sys}/\partial E = 1/T\) — i.e. the system's own temperature equals the bath's. That stationarity condition is the equilibrium the widget's peak is tracking. Formally, \(F = U-TS\) is the Legendre transform of \(U\) swapping \(S\) for its conjugate \(T\), and the whole construction is just \(\Delta S_\text{universe}\ge 0\) written from the system's chair: a process is spontaneous iff \(\Delta F \le 0\). At constant pressure the same argument with the bath also exchanging volume gives \(G = H - TS\).
Step 6 · Watch it happen
Everything so far has been curves and formulas. Here is the actual mechanism — no equations driving it, just random energy exchange and counting. We'll use the simplest exactly-countable model: a small system (a few oscillators, on the left) and a large bath (many oscillators, on the right). Energy comes in indivisible quanta. There is a fixed total number of quanta, and at every tick one quantum hops to a uniformly random oscillator anywhere in the combined pool.
That single rule — energy wanders blindly — is the entire engine. Nothing in it "wants" anything. Yet watch what happens to the bars at the bottom: the system doesn't drift to zero energy (lowest energy) and it doesn't hog everything (highest energy). It settles into a steady share, fluctuating around a value, while the total entropy climbs and then plateaus at its maximum. That plateau is equilibrium, and it is nothing but the most probable way to spread the quanta.
The reason is pure counting. For an Einstein solid of \(N\) oscillators holding \(q\) quanta, the number of arrangements is \( \Omega(N,q) = \binom{q+N-1}{q} \). The total number of ways for the universe is \( \Omega_\text{sys}(q_s)\,\Omega_\text{bath}(q-q_s) \), and the system spends time in each split in proportion to that product. The product is overwhelmingly peaked — so the system is found near the peak not by force but by sheer weight of numbers. Press play and watch the microstate count, the entropies, and the running histogram of where the system's energy lives.
The equilibrium split maximizes \( S_\text{tot}=k_B[\ln\Omega_\text{sys}(q_s)+\ln\Omega_\text{bath}(q-q_s)] \), giving \( \partial_{q_s}\ln\Omega_\text{sys} = \partial_{q_s}\ln\Omega_\text{bath} \) — the two temperatures equalize, since \(1/T \equiv k_B\,\partial(\ln\Omega)/\partial E\). For a large bath, \(\ln\Omega_\text{bath}(q-q_s)\approx \text{const} - q_s/ (k_BT)\) is linear in the system's energy with slope \(-1/T\) — exactly the straight bath line from the previous step, now emerging from a real counting argument. The probability the system holds energy \(E_s\) is \(P(E_s)\propto \Omega_\text{sys}(E_s)\,e^{-E_s/k_BT}\); divide out the system's own degeneracy and you have the Boltzmann factor \(e^{-E_s/k_BT}\) that the next step builds on. The histogram the widget accumulates is a Monte Carlo estimate of exactly this \(P(E_s)\).
Step 7 · Where it comes from
Why that formula and not some other combination? Because of one fact about how nature distributes probability over states. A system in contact with a heat bath at temperature \(T\) occupies a microstate of energy \(E_i\) with probability proportional to the Boltzmann factor:
\[ p_i = \frac{e^{-E_i/k_BT}}{Z}, \qquad Z = \sum_i e^{-E_i/k_BT} \]
The normalizer \(Z\) — the partition function — is the whole story in one number: it sums the Boltzmann factor over every microstate, automatically counting low-energy states heavily and high-energy ones faintly. And here is the punchline that ties everything together:
\[ F = -k_B T \ln Z \]
The free energy is just the log of that sum. Below, build the partition function up term by term and watch how \(Z\), and therefore \(F\), emerges from summing over states — and how raising \(T\) lets high-energy states start to count.
From \(F=-k_BT\ln Z\) everything thermodynamic falls out by differentiation: \(S=-\partial F/\partial T\), \(\langle E\rangle = -\partial \ln Z/\partial\beta\) with \(\beta = 1/k_BT\), and pressure \(p=-\partial F/\partial V\). Substituting the Boltzmann \(p_i\) into \(F = \langle E\rangle - TS\) recovers \(-k_BT\ln Z\) exactly — the two definitions of \(F\) are the same object. The relation \(F=-k_BT\ln Z\) is also why \(F\) is the natural target for sampling methods: differences in \(F\) are differences in log-normalizers, i.e. log-ratios of partition functions.
Step 8 · The landscape
In practice we rarely care about the single global \(F\). We care how free energy varies along some coordinate we can interpret — a reaction's progress, a protein's fraction folded, a distance between two atoms. Call it \(x\). Integrate out everything else and you get a free energy surface, or potential of mean force:
\[ F(x) = -k_B T \ln p(x) \]
Read this backwards and it's intuitive: states the system visits often (high \(p\)) are low in free energy; states it almost never visits are high. Valleys are stable states; barriers are what the system must climb to get between them. The depth of a basin sets how populated it is; the height of a barrier sets how slowly the system crosses.
Reshape the landscape below. The cloud of samples is the equilibrium distribution \(p(x)\propto e^{-F(x)/k_BT}\) — deepen a basin and watch samples pile into it; raise the temperature and watch them spill over the barrier.
\(F(x) = -k_BT\ln\!\int \delta(\xi(\mathbf r)-x)\, e^{-E(\mathbf r)/k_BT}\,d\mathbf r + \text{const}\), the marginal of the Boltzmann density along a collective variable \(\xi\). The gradient \(-\nabla F(x)\) is the mean force on \(x\) (hence "potential of mean force"). Free energy differences between basins — \(\Delta F = -k_BT\ln(p_A/p_B)\) — are observable and reproducible even when absolute \(F\) is not, because the unknown constant and \(Z\) cancel. Estimating these differences when barriers make direct sampling of \(p(x)\) intractable is the entire job of enhanced sampling and free energy methods (umbrella sampling, metadynamics, thermodynamic integration, and their learned-sampler descendants).
Start from the puzzle: systems don't just minimize energy, or ice would never melt. The missing ingredient is counting — the number of arrangements, named entropy, \(S=k_B\ln\Omega\). Temperature is the exchange rate that lets us weigh energy against entropy on one ledger, giving the free energy \(F = E - TS\): energy discounted by how many ways there are to arrange it.
That trade-off isn't a guess — it falls straight out of the Boltzmann distribution, where \(F = -k_BT\ln Z\) is just the log of the sum over all states. And projected onto a coordinate we can interpret, \(F(x) = -k_BT\ln p(x)\) turns into a landscape of valleys and barriers that explains both what's stable and how fast things change.
Where it shows up: every protein-folding free energy surface, every binding affinity \(\Delta G\), every phase transition, and — closer to home — every method that estimates \(\Delta F\) by reweighting or biased sampling is, at bottom, trying to compute a ratio of partition functions that brute-force sampling can't reach.