Gauge fields → diffusion → optimization

Diffusion models for SU(N) gauge theory

A ground-up tutorial on the group, the lattice, the probability distribution, what the recent diffusion models learn, and how diffeomorphic optimization could search their base space without confusing a low-action configuration with a quantum ground state.

SU(N) lattice gauge theory is like a matrix-valued cousin of the Ising model: the sampled variables live on edges rather than sites, and the local action is built around squares rather than pairs. A gauge diffusion model learns their equilibrium ensemble. Diffeomorphic optimization then searches the random starting matrices for a generated configuration with low loss—not automatically for the quantum ground state.

Start here · the Ising bridge

Move the spins from sites to edges, then replace signs by matrices

The ordinary Ising model is the right starting point. It separates three ideas that are easy to blur together: what is sampled, where it lives, and which fixed parameter controls its energy.

In the Ising model, each site carries a sampled spin \(s_x=\pm1\). The number \(J\) on a bond is usually a fixed coupling, not a sampled variable: \[ E_{\mathrm{Ising}}[s] =-J\sum_{\langle xy\rangle}s_xs_y. \] The notation \(\langle xy\rangle\) means that \(x\) and \(y\) are nearest neighbors, and the sum counts every bond once. Each local energy term therefore examines two neighboring site spins.

In the standard homogeneous model, we choose the same number \(J\) beforehand for every bond. It is a fixed coupling, not another sampled variable. When \(J>0\), equal neighboring spins lower the energy; when \(J<0\), opposite spins lower it. A more general graph may assign a different fixed value \(J_{xy}\) to each bond: \[ E[s]=-\sum_{\langle xy\rangle}J_{xy}s_xs_y. \] During an ordinary Monte Carlo simulation, the \(J_{xy}\) values remain fixed while the site spins \(s_x\) are sampled.

Minimization and sampling are different questions

If we ask for the zero-temperature ground-state configuration, then yes: the task is \[ s^\star=\arg\min_{\{s_x=\pm1\}}E[s]. \] For a connected graph with every \(J_{xy}>0\), the two minima are all \(+1\) and all \(-1\). Mixed positive and negative couplings can create frustration and make the optimization much harder.

At finite temperature, however, the Ising model does not keep only the minimum. It assigns every spin configuration the probability \[ p_{\beta_T}(s)=\frac{1}{Z}e^{-\beta_T E[s]}, \qquad Z=\sum_{\{s_x\}}e^{-\beta_T E[s]}, \] where \(\beta_T=1/(k_BT)\) is inverse temperature. Low-energy configurations are more likely, but their number also matters. Only as \(T\to0\), or \(\beta_T\to\infty\), does this distribution concentrate on the energy minima. The same distinction between sampling an ensemble and minimizing one configuration will matter for SU(N) gauge theory.

A lattice gauge theory changes the placement of the dynamical variables. A \(\mathbb Z_2\) gauge theory puts a sampled sign \(u_\ell=\pm1\) on every edge \(\ell\). The smallest gauge-invariant energy term must close into a loop, so it multiplies the four edge signs around a plaquette \(p\): \[ E_{\mathbb Z_2\text{ gauge}}[u] =-\beta\sum_p\prod_{\ell\in p}u_\ell. \] Here \(\beta\), not \(u_\ell\), plays the role of the fixed coupling.

What is optimized or sampled in the Z2 theory?

We do not choose a subset of edges. Every edge receives one dynamical sign, and a complete configuration is \[ u=\{u_\ell\in\{-1,+1\}\}_{\text{every edge}}. \] For the zero-temperature optimization problem, we search over all such assignments for the one minimizing \(E[u]\). With positive coupling, each plaquette prefers \[ \prod_{\ell\in p}u_\ell=+1. \] For the ensemble problem, we instead sample complete edge assignments with Boltzmann probability proportional to \(e^{-E[u]}\) in the dimensionless convention used here.

The minimizing edge assignment is not unique. Choose a node and flip every edge touching it. Each adjacent plaquette contains two flipped edges, so its product is unchanged. This local redundancy is the \(\mathbb Z_2\) version of a gauge transformation. The plaquette products are physical; the individual edge signs depend on the chosen gauge.

SU(N) gauge theory keeps this edge-and-plaquette architecture but upgrades each sign to a noncommuting matrix \(U_\ell\in\SU{N}\). The product around a square becomes an ordered matrix product, and the trace turns it into a scalar action: \[ S_W[U] =-\frac{\beta}{N}\sum_p \Re\Tr\!\left(U_1U_2U_3^\dagger U_4^\dagger\right) \quad\text{(up to counting conventions).} \]

Where the sampled variables move staged · Ising to SU(N)
Ising: \(s_x=\pm1\) on sites
\(\mathbb Z_2\): \(u_\ell=\pm1\) on edges
SU(N): \(U_\ell\in\SU{N}\) on edges
With matter: \(\psi(x)\) also on sites
Ising samples one sign on every site; a bond term compares two signs.

Watch two changes separately: first the variables move from sites to edges; then each edge sign is upgraded to an SU(N) matrix. The pure-gauge diffusion papers stop at the third stage.

Ordinary Ising

Sample \(s_x\) on sites. \(J\) is fixed. Local energy lives on bonds: \(-Js_xs_y\).

Pure SU(N) gauge

Sample \(U_\mu(x)\) on edges. \(\beta\) is fixed. Local action lives on plaquettes: \(-(\beta/N)\Re\Tr U_p\).

Gauge plus matter

Sample both site fields \(\psi(x)\) and gauge links \(U_\mu(x)\). Covariant hopping terms couple them.

Step 1 · the object

SU(N) is the space each link is allowed to live in

Start with a complex \(N\times N\) matrix \(U\). Requiring \(U^\dagger U=I\) says it preserves lengths; requiring \(\det U=1\) removes an overall phase. The matrices satisfying both conditions form the special unitary group \(\SU{N}\).

A group is not merely a set: multiplying two allowed matrices gives another allowed matrix, every matrix has an inverse, and multiplication is associative. It is also a smooth curved space. Near the identity, every small motion is generated by a traceless anti-Hermitian matrix \(A\in\su{N}\):

\[ A^\dagger=-A,\qquad \Tr A=0,\qquad U(s)=e^{sA}U(0)\in\SU{N}. \]

There are \(N^2-1\) independent generator directions. This is why an \(\SU{3}\) link has eight continuous degrees of freedom.

An SU(2) matrix as a unit quaternion live · tracked axis projections
unitarity error 0.00e+0 determinant 1.000 + 0.000i quaternion norm 1.000000

The crosshairs track the imaginary quaternion vector in the \(a_1a_2\), \(a_1a_3\), and \(a_2a_3\) planes. Hold the group and polar angles fixed, then rotate the azimuth: the \(a_1a_2\) projection follows the highlighted circle while \((a_0,a_1,a_2,a_3)\) remains on the unit three-sphere.

Why SU(2) is a three-sphere

Every \(U\in\SU{2}\) can be written \[ U=a_0I+i\sum_{k=1}^{3}a_k\sigma_k = \begin{pmatrix} a_0+ia_3 & a_2+ia_1\\ -a_2+ia_1 & a_0-ia_3 \end{pmatrix}, \] where the Pauli matrices are \(\sigma_k\). Unitarity and unit determinant collapse to one condition: \(\sum_{k=0}^3a_k^2=1\). Thus SU(2) is diffeomorphic to \(S^3\), not to an unconstrained Euclidean vector space.

Step 2 · the physics

A lattice gauge field puts one SU(N) transporter on every edge

First: what is a field?

A field is a rule that assigns a value to every position. A temperature field assigns one number \(T(x)\); a wind field assigns a vector \(\mathbf v(x)\); a matter field assigns an internal color vector \(\psi(x)\). The symbol \(x\) asks “at which point?” and the field name tells us what value lives there.

On a continuum there are infinitely many positions. A lattice replaces them by a finite grid, so a field becomes a finite collection of values. A site field stores one value at every node. A gauge field is different: it tells us how to compare internal directions at neighboring nodes, so its lattice values \(U_\mu(x)\in\SU{N}\) live on oriented edges.

The continuum gauge field is usually written \(A_\mu(x)\), a Lie-algebra-valued quantity at every spacetime point and direction. A short lattice link packages its accumulated effect into \[ U_\mu(x)\approx \exp\!\left[iag\,A_\mu\!\left(x+\tfrac{a}{2}\hat\mu\right)\right], \] where \(a\) is the lattice spacing and \(g\) the coupling. So the link matrix is not a different kind of physics: it is the finite-step version of the continuum gauge field.

Analogy: a network of laboratories

Imagine a laboratory at every lattice node. Every lab describes the same kind of \(N\)-channel signal, but each is free to choose its own channel labels and reference phases. A column vector records the signal using one lab's convention. The same signal generally has different components when written down by the neighboring lab.

An edge matrix \(U_\mu(x)\) is the translation dictionary between two neighboring conventions. Multiplying by it converts a vector written by one lab into the coordinates understood by the other. \(\Omega(x)\) means that the lab at \(x\) changed its private convention, so every dictionary touching that lab must be updated.

A closed chain of perfect dictionaries would bring a signal back unchanged. If multiplying the dictionaries around a square produces a net rotation, that mismatch is curvature. Gauge theory uses this same structure, but the \(N\) channels are components of an internal SU(N) vector rather than literal laboratory instruments. In QCD specifically, this internal degree of freedom is called color.

A lattice site is one point in discretized spacetime. Attached to every site \(x\) is an abstract copy of \(\mathbb C^N\): the vector space carrying the fundamental representation of SU(N). In QCD, this is called the local color space.

If the theory contains matter, a field \(\psi(x)\in\mathbb C^N\) picks one vector in that local color space. For SU(3), for example, \(\psi(x)=(\psi_r,\psi_g,\psi_b)^T\) contains the complex amplitudes for the three color basis states. This is one quantum color state with \(N\) components, not \(N\) little colored objects sitting on the node.

Call the neighboring site \(y=x+\hat\mu\). Suppose its field value is the column vector \(\psi(y)\), written using the coordinate convention at \(y\). We cannot compare those numbers directly with \(\psi(x)\), because \(\psi(x)\) uses the convention at \(x\). First translate the neighbor's vector into \(x\)'s coordinates: \[ \underbrace{\psi_{y\to x}}_{\text{neighbor's vector, written at }x} = \underbrace{U_\mu(x)}_{\text{dictionary from }y\text{ to }x} \underbrace{\psi(y)}_{\text{vector written at }y}. \] Now \(\psi_{y\to x}\) and \(\psi(x)\) are written in the same coordinate system and can be compared. To translate the other way, use the inverse dictionary: \[ \psi_{x\to y}=U_\mu^\dagger(x)\psi(x). \] The dagger appears because an SU(N) matrix is unitary, so \(U_\mu^\dagger=U_\mu^{-1}\). Nothing physically travels during this calculation; we are rewriting vectors so both descriptions use the same local coordinates.

What, then, lives on a node? In a theory with matter, a node can carry a dynamical field such as a quark color vector \(\psi(x)\in\mathbb C^N\). But the diffusion papers study pure gauge theory, so no dynamical value is sampled on the nodes. A node supplies only the spacetime label \(x\), an abstract local color space \(\mathbb C^N\), and a freely chosen local basis. The matrix \(\Omega(x)\) describes changing that basis; it is a gauge choice, not an additional physical field. The sampled dynamical data are only the edge matrices \(U_\mu(x)\).

The clean mental model

A node does not contain a randomly chosen vector in the pure-gauge theory. Instead, site \(x\) has an abstract vector space \[ V_x\cong\mathbb C^N, \] meaning “the kind of \(N\)-component vector that could be described here.” It is a vector space, not a group, and no particular element of \(V_x\) is part of a pure-gauge configuration.

The dynamical object on the edge is one matrix \[ U_\mu(x)\in\SU{N}, \qquad U_\mu(x):V_{x+\hat\mu}\longrightarrow V_x. \] Its matrix entries are not the elements of a node vector. Together they specify one allowed linear map between two neighboring copies of \(\mathbb C^N\).

A complete gauge-field configuration is the collection of all edge matrices, \[ U=\{U_\mu(x)\}_{\text{every oriented edge}}. \] HMC or the diffusion model samples this whole collection from \[ p(U)=\frac{1}{Z}e^{-S_W[U]} \prod_{\text{edges}}\dd\mu_{\mathrm{Haar}}\!\left(U_\mu(x)\right). \] The individual links are therefore not independent Haar draws: the Wilson action \(S_W\) correlates neighboring links through plaquettes. Finally, \(\Omega(x)\in\SU{N}\) is a local change of coordinates on \(V_x\), not another physical field applied after sampling.

Now change the color coordinates independently at every site: \(\psi(x)\mapsto\Omega(x)\psi(x)\). The link must change with both endpoint bases: \[ U_\mu(x)\longmapsto U_\mu^\Omega(x)= \Omega(x)U_\mu(x)\Omega^\dagger(x+\hat\mu). \] This rule ensures \[ U_\mu^\Omega(x)\psi^\Omega(x+\hat\mu) =\Omega(x)U_\mu(x)\psi(x+\hat\mu), \] so everyone agrees on the transported color state even if they use different coordinates. Individual link matrices depend on those coordinates; physics must be built from combinations in which the basis changes cancel.

The smallest closed loop is a square called a plaquette, hence the letter \(P\). A square needs two directions: \(\mu\) names its first axis and \(\nu\) its second. Thus \(P_{\mu\nu}(x)\) means “the plaquette based at site \(x\), lying in the \(\mu\)-\(\nu\) plane, with the orientation fixed by the ordered pair \((\mu,\nu)\).” Its four factors represent the edges in the \(+\mu,+\nu,-\mu,-\nu\) directions. For example, \(P_{12}(x)\) is the square in lattice directions 1 and 2 that starts at \(x\): \[ P_{\mu\nu}(x)=U_\mu(x)U_\nu(x+\hat\mu) U_\mu^\dagger(x+\hat\nu)U_\nu^\dagger(x). \] The ordered subscripts matter because SU(N) matrices generally do not commute. Reversing the orientation produces the inverse loop matrix (up to the corresponding shift of base point). It transforms as \(P\mapsto\Omega(x)P\Omega^\dagger(x)\), so its trace is unchanged. The trace measures the curvature trapped in that square.

At a site

A local color space \(\mathbb C^N\). With matter, \(\psi(x)\) is an \(N\)-component vector in it. In pure gauge theory, no matter vector is stored.

On a link

\(U_\mu(x)\in\SU{N}\), an \(N\times N\) matrix that translates color components between neighboring local bases.

Around a loop

Multiply the link translators in order. Returning rotated reveals gauge curvature; the traced product is independent of all chosen bases.

What lives where on a gauge lattice staged · build the notation
site \(x\)
neighbor \(x+\hat\mu\)
link \(U_\mu(x)\in\SU{N}\)
\(U_\mu(x)\psi(x+\hat\mu)\)
\(\Omega(x)U_\mu(x)\Omega^\dagger(x+\hat\mu)\)
Pick one lattice site and call it \(x\).

Each equation is a label for one visible part of the same picture. Advance one stage at a time; the highlighted equation names what has just appeared.

Notation decoder

Yes: matrices written next to one another are simply multiplied. Thus \(AB\) means ordinary matrix multiplication—but order matters, because generally \(AB\neq BA\). The remaining symbols are compact bookkeeping:

\(U_\mu(x)\) is the matrix assigned to the oriented link from \(x\) toward direction \(\mu\); \(x+\hat\mu\) is the neighboring site one lattice step away. With the convention above, \(U_\mu\) translates the neighbor's components back to the basis at \(x\), while \(U_\mu^\dagger\) translates forward. \(\Omega(x)\) is a change of color basis at site \(x\). The dagger \(U^\dagger\) means transpose and complex-conjugate; for a unitary matrix it is also the inverse, so it traverses the same link backward.

Finally, \(\Tr U\) adds the diagonal entries of \(U\), \(\Re\) keeps the real part, and \(\mapsto\) means “changes into under a gauge transformation.” There is no new exotic multiplication hiding in the plaquette formula: it is four ordered matrix multiplications around a closed square.

Local bases change; the plaquette does not click · gauge transform
before: Re Tr(P)/2 = 0.000000 after: Re Tr(P)/2 = 0.000000 difference = 0.00e+0

The four link descriptions are basis-dependent. Multiply around the closed loop and every intermediate basis rotation cancels.

Bridge · from curvature to probability

The Wilson action turns every plaquette mismatch into a local cost

The symbol \(U_p\) means the plaquette matrix: the ordered product of the four link matrices around one square \(p\). For the plaquette based at \(x\) in the \(\mu\)-\(\nu\) plane, \[ U_p\equiv P_{\mu\nu}(x) = U_\mu(x)\, U_\nu(x+\hat\mu)\, U_\mu^\dagger(x+\hat\nu)\, U_\nu^\dagger(x). \] The daggered matrices traverse the final two edges backward. Thus \(U_p\) is one SU(N) matrix answering: “after transporting around this closed square, what net transformation remains?” If every internal vector returns unchanged, then \(U_p=I\): the plaquette is locally flat. The action needs to measure how far every \(U_p\) is from that ideal.

First compress the matrix to one real, gauge-invariant number: \[ w_p=\frac{1}{N}\Re\Tr U_p. \] The normalized trace reaches its maximum \(w_p=1\) when \(U_p=I\). Therefore \(1-w_p\) is zero for a flat plaquette and positive when going around the square produces a nontrivial rotation.

Add that local cost over every plaquette: \[ \boxed{ S_W[U] =\beta\sum_p\left(1-\frac{1}{N}\Re\Tr U_p\right) =\frac{\beta}{N}\sum_p\left(N-\Re\Tr U_p\right). } \] This is the Wilson gauge action. The lattice and all link matrices \(U_\mu(x)\) are the variables; \(\beta\) is one fixed coupling. In a common four-dimensional SU(N) convention, \(\beta=2N/g_0^2\), where \(g_0\) is the bare gauge coupling.

Build the plaquette matrix on a 2D grid staged · one edge matrix at a time
\(U_\mu(x)\)
\(U_\nu(x+\hat\mu)\)
\(U_\mu^\dagger(x+\hat\nu)\)
\(U_\nu^\dagger(x)\)
Start at \(x\) and follow the stored \(+\mu\) link to \(x+\hat\mu\).

Thin gray arrows show how link matrices are stored in the positive coordinate directions. Thick colored arrows show how this loop is traversed. Moving against a stored arrow uses its inverse, \(U^{-1}=U^\dagger\).

A Wilson action assembled one square at a time click a plaquette · bend its holonomy
selected \(\Re\Tr U_p/2\) 0.4536 selected cost 1.0928 total \(S_W\) 1.0928

This SU(2) slice uses \(U_p(\theta)=\mathrm{diag}(e^{i\theta},e^{-i\theta})\), so \(\Re\Tr U_p/2=\cos\theta\) and the local Wilson cost is \(\beta(1-\cos\theta)\). Click any square, then change its loop angle.

Why does this reproduce Yang–Mills theory? For a small square of side \(a\), the loop matrix is approximately \[ U_{\mu\nu}(x) \approx\exp\!\left[i\,a^2g_0F_{\mu\nu}(x)\right], \] where \(F_{\mu\nu}\) is the continuum field strength. Expanding the exponential gives \[ \Re\Tr U_{\mu\nu} = N-\frac{a^4g_0^2}{2}\Tr(F_{\mu\nu}^2)+O(a^6). \] Thus \(N-\Re\Tr U_p\) is proportional to curvature squared times the plaquette volume. Summing over squares becomes the continuum integral \[ S_{\mathrm{YM}}\propto \int\dd^dx\;\Tr(F_{\mu\nu}F_{\mu\nu}) \] as \(a\to0\).

Why the trace is essential

A gauge transformation conjugates the loop matrix: \(U_p\mapsto\Omega(x)U_p\Omega^\dagger(x)\). Conjugation changes the matrix entries but not its trace. The Wilson action therefore assigns the same cost to every gauge-equivalent description. It measures curvature, not the arbitrary coordinate convention at a node.

Equivalent conventions in papers and code

Authors may write \(S_W=-(\beta/N)\sum_p\Re\Tr U_p\), dropping the additive constant \(\beta\sum_p1\), because constants cancel from normalized expectation values. Some sum over ordered pairs \(\mu\ne\nu\), counting each unoriented plaquette twice, and compensate with a factor of \(1/2\). These conventions generate the same probability distribution when their prefactors are matched.

Step 3 · the target distribution

The calculation needs an ensemble, not one best configuration

After Wick rotating time, quantum expectation values become averages over classical-looking Euclidean field configurations: \[ \langle\mathcal O\rangle= \frac{1}{Z}\int\!\mathcal DU\;\mathcal O[U]e^{-S_W[U]}, \qquad Z=\int\!\mathcal DU\;e^{-S_W[U]}. \]

Low-action fields are more probable, but they are not the only fields that matter. The enormous number of less-than-minimal configurations contributes entropy. Standard lattice calculations use Markov-chain Monte Carlo, especially HMC, to draw correlated samples from this Boltzmann weight. At fine lattice spacing, long autocorrelation times and frozen topological sectors can make those samples expensive.

One SU(2) plaquette: minimum versus ensemble sample · change coupling
samples 0 sample mean cos(theta) -- exact ensemble mean 0.4331 action-minimizing value 1.0000

The action minimum is the single point \(\theta=0\), but Haar measure contributes a \(\sin^2\theta\) volume factor. The quantum path integral averages the whole shaded distribution.

The ground-state trap

Minimizing \(S_W[U]\) gives a classical minimum-action field (for the pure Wilson action, flat plaquettes \(U_p=I\), modulo gauge and boundary degeneracies). A quantum ground state is instead a lowest-energy wavefunctional \(\Psi_0[U]\), or the state projected out by long Euclidean time. It is not one link configuration obtained by minimizing the Euclidean action.

Step 4 · the generative mechanism

Forward diffusion forgets the field by taking a random walk on SU(N)

A diffusion model creates a bridge between hard samples \(U_0\sim p_{\rm data}\) and an easy reference distribution. On a compact group, the natural easy distribution is Haar measure: uniform over the group.

The intrinsic construction in Komijani, Marinkovic, and Turgut multiplies each link by many tiny random group elements: \[ U_{t+\Delta t} = \exp\!\left[\sigma(t)\sqrt{\Delta t}\,\eta_t\right]U_t, \qquad \eta_t\in\su{N}. \] Because the exponential of a traceless anti-Hermitian matrix is in SU(N), every intermediate state remains a valid gauge field. As accumulated variance grows, memory of \(U_0\) disappears and \(U_t\) approaches Haar.

Brownian motion spreading over SU(2) real group walk · advance noise
steps 0 mean Re Tr(U)/2 1.000 mean quaternion norm 1.000000

Every dot is an SU(2) matrix. The histogram tracks \(x=\Re\Tr(U)/2=a_0\). It begins at \(x=1\) and approaches the Haar shape proportional to \(\sqrt{1-x^2}\), while every quaternion norm stays one.

The continuous group diffusion

With generators normalized by \(\Tr(T^aT^b)=-\delta^{ab}\), write \(dW_t=dW_t^aT^a\). The time-ordered process is \[ U_t=K_{t,0}U_0,\qquad K_{t,0}=\mathcal T\exp\!\left(\int_0^t\sigma(\tau)\,dW_\tau\right). \] Time ordering matters because SU(N) is non-Abelian: increments at different times generally do not commute. The cited SU(N) paper uses \(\sigma(t)=\sigma_0/\sqrt{1-t+\varepsilon}\), whose diverging accumulated variance drives the endpoint toward exact Haar measure.

Step 5 · learning the way back

The score points toward configurations that are more probable at each noise level

At time \(t\), noisy fields have density \(\rho_t(U)\). Its score is the tangent vector \[ s(U,t)=T^a\partial^a\log\rho_t(U)\in\su{N}. \] It answers a local question: “which infinitesimal group motion most increases log probability?”

A neural network \(s_\theta(U,t)\) learns this vector from noised training fields. Reversing the diffusion uses the score to oppose the entropy-producing random walk. The deterministic member of the reverse family is the probability-flow ODE \[ \frac{dU_t}{dt} = -\frac12\sigma^2(t)s_\theta(U_t,t)U_t, \] integrated from noisy \(t=1\) toward data \(t=0\).

A score field on one SU(2) subgroup live · drag noise level
mean log-density gain 0.000 probe steps 0

This is a one-parameter SU(2) cross-section, not the full lattice. Arrows show \(\partial_\theta\log\rho_t\): broad and weak at high noise, sharp near learned modes at low noise. The network must produce the analogous tangent matrix for every lattice link.

The SU(N) paper avoids needing the unknown marginal score directly. For algebra increments \(d\Gamma_t=\log(U_{t+dt}U_t^\dagger)\), the conditional increment score is Gaussian:

\[ \nabla_{\Gamma_t}\log q_t(\Gamma_t\mid\Gamma_0) = -\frac{\Gamma_t-\Gamma_0}{\sigma_{\rm cum}^2(0,t)}. \]

It trains \(s_\theta\) by weighted implicit score matching against this known target. For non-Abelian groups, translating the algebra-space result to the exact group reverse drift is subtle; the paper gives a heuristic stochastic-equivalence argument rather than a full proof.

Why equivariance belongs in the architecture

If \(U\mapsto U^\Omega\), a gauge-equivariant score must transform in the corresponding link tangent representation. GaugeLinkConv layers build features from Wilson staples and parallel transport them before combining them. The network therefore cannot mistake a local basis change for new physics. Augmentation can encourage this behavior; equivariant layers guarantee it by construction.

Step 6 · what the models actually solve

Two recent papers make two importantly different choices

Both papers learn to generate finite-lattice equilibrium gauge configurations, but only one treats SU(N) as the curved group it is throughout diffusion. Use the staged comparison below; each click adds one design decision.

Flat DDPM versus intrinsic Lie-group diffusion staged · compare papers

Why use ML?

After training, produce new approximately independent equilibrium fields without walking one long Markov chain from its previous state.

What is learned?

A noise predictor or group score at every noise time—not the action minimum, not a wavefunction, and not an exact sampler guarantee.

What validates it?

Agreement of Wilson-loop distributions and other observables with HMC or exact results, plus scaling, topology, and autocorrelation tests.

Results and limitations from the two papers

Alharazin, Panteleeva, and Sun (arXiv:2602.09045v1). SU(2) links are unit quaternions, but ordinary Gaussian DDPM noise moves intermediate states into \(\mathbb R^8\); links are renormalized only at the end. Gauge symmetry is learned through random gauge augmentation, not enforced. On \(8^2\) at \(\beta=2\), the reported plaquette \(0.4329\pm0.0001\) agrees with the exact \(0.4331\), but transfer to distant couplings and large unseen volumes degrades substantially.

Komijani, Marinkovic, and Turgut (arXiv:2605.06134). Multiplicative Brownian motion stays on SU(N), the score is Lie-algebra valued, and GaugeLinkConv is equivariant. Tests use SU(3) in 2D and 4D. A plain RK4 reverse ODE works in easier settings; at 4D \(\beta=6\), they require 20 predictor steps with HMD correctors. Those correctors are unadjusted, the estimated cost exceeds their HMC thermalization baseline, and only a small set of Wilson loops is tested. The exact GaugeLinkUNet used in the paper is not publicly released.

Neither reported sampler includes a final Metropolis correction, so finite model and integration error can bias observables. Neither paper claims a quantum ground-state calculation.

Step 7 · your proposed connection

Diffeomorphic optimization turns the frozen reverse flow into coordinates for search

Set stochastic reverse noise to zero and freeze the trained score model. Integrating its probability-flow ODE gives a deterministic map \(g:\mathcal Z\to\mathcal C\) from easy base fields \(Z\) to generated gauge fields \(U=g(Z)\).

Given any differentiable gauge-invariant loss \(\mathcal L(U)\), optimize its pullback: \[ \widetilde{\mathcal L}(Z)=\mathcal L(g(Z)). \] Backpropagation through the full ODE gives the gradient at the base field. Every outer update is made intrinsically on each SU(N) base link: \[ Z_{\ell}^{(k+1)} = \exp\!\left[-\eta A_{\ell}^{(k)}\right]Z_{\ell}^{(k)}, \qquad A_{\ell}^{(k)} = \operatorname{grad}_{Z_\ell} \mathcal L(g(Z^{(k)}))\in\su{N}. \]

To first order, mapping this step through \(g\) is Riemannian gradient descent on the generated manifold. In Euclidean notation, \[ \Delta U\approx-\eta\,J_gJ_g^\top\nabla_U\mathcal L. \] The learned generator supplies a geometry-aware preconditioner \(J_gJ_g^\top\): base-space motion is pushed into directions the model knows how to generate.

Optimize through a diffeomorphic SU(2) flow iterate · backpropagate to base
base z 2.450 generated angle x=g(z) 2.801 Jacobian g'(z) 0.576 loss 1-cos(x) 1.942 iteration 0

The toy flow \(x=g(z)=z+0.55\sin z\) is a genuine orientation-preserving diffeomorphism of a one-parameter SU(2) subgroup because \(g'(z)=1+0.55\cos z>0\). The gradient is pulled back as \(d\mathcal L/dz=\sin(g(z))g'(z)\), then the group-valued base is updated.

So: can it find the “ground state SU(N) diffusion model”?

Yes, under a precise reinterpretation: it can search for a low-loss SU(N) gauge configuration inside the image of a frozen, deterministic, gauge-equivariant diffusion flow. Choosing \(\mathcal L=S_W\) searches for a low classical lattice action. Choosing a differentiable Wilson-loop target, smooth topological surrogate, or variational energy gives a different constrained design problem.

No, under the literal quantum interpretation: minimizing \(S_W(g(Z))\) does not produce the quantum Yang–Mills ground-state wavefunctional or an unbiased path-integral ensemble. It returns \(\arg\min_Z\mathcal L(g(Z))\), which may differ from the global \(\arg\min_U\mathcal L(U)\) and intentionally collapses sampling diversity.

A technically sound experiment

  1. Start with the intrinsic SU(N) model, not the flat quaternion DDPM. Use \(\mathcal Z=\SU{N}^{|E|}\) with independent Haar base links so the base and configuration spaces have compatible topology.
  2. Use the deterministic probability-flow ODE only. Predictor–corrector noise is stochastic and does not define the invertible map required by the diffeomorphic-optimization argument.
  3. Enforce gauge equivariance in \(g_\theta\) and choose a gauge-invariant differentiable objective. Otherwise the optimizer can exploit gauge artifacts.
  4. Backpropagate through the group-valued ODE using stored solver states, checkpointing, or a matrix-Lie-group adjoint. The adjoint contains the noncommutative commutator term; a Euclidean adjoint copied verbatim is not enough.
  5. Project each base gradient to its traceless anti-Hermitian part and update with the matrix exponential. Monitor unitarity, determinant, gradient norms, reversibility error, and ODE tolerance.
  6. Compare against direct Riemannian optimization of \(U\), cooling/gradient flow, HMC-derived best-of-\(K\) samples, and random base restarts. Report both the optimized objective and physical observables not used as the loss.
  7. If the actual target is a quantum ground state, formulate a variational wavefunctional \(\Psi_\phi[U]\) and minimize \(\langle\Psi_\phi|H|\Psi_\phi\rangle/\langle\Psi_\phi|\Psi_\phi\rangle\). Diffeomorphic optimization might optimize samples or variational parameters, but the Hamiltonian expectation—not \(S_W\) alone—defines the task.
Geometric and physical caveats

Topology. There is no global diffeomorphism \(\mathbb R^d\to\SU{N}\): the former is noncompact and the latter compact. A Euclidean Gaussian latent plus one global Lie-algebra chart cannot satisfy the theorem globally. Use group-valued base variables, intrinsic flows, or an explicitly charted construction.

Gauge zero modes. A gauge-invariant loss is flat along gauge orbits. That degeneracy is physical. Equivariance keeps it from turning into a spurious learning signal, but the Hessian remains singular along pure-gauge directions.

Topological sectors. A smooth invertible flow cannot cross disconnected components. At finite lattice spacing sectors may not be exactly disconnected, but continuum-like admissibility makes sector transitions increasingly difficult. Multiple base initializations or sector-conditional models may be necessary.

Approximate invertibility. The theorem assumes a smooth diffeomorphism. A finite-step ODE solver, learned-score error, projection back to SU(N), and stiff dynamics all make the implemented map only approximately invertible. Large optimization steps also invalidate the local first-order equivalence.

Putting it together

SU(N) supplies the curved link space. Gauge symmetry says local color bases are redundant, so observables come from closed loops. The Euclidean path integral needs a whole Boltzmann ensemble. Intrinsic diffusion destroys that ensemble by a group-valued random walk and learns the score needed to reverse it. Diffeomorphic optimization then freezes the deterministic reverse flow and uses its base space as learned coordinates for a separate optimization problem.

Defensible claim

“We optimize a differentiable gauge-invariant objective over configurations generated by an SU(N)-equivariant probability-flow ODE, using intrinsic base-space gradient updates.”

Claim to avoid

“Minimizing the Wilson action through a diffusion model finds the quantum ground state or preserves the equilibrium sampling distribution.”

Primary sources and code map

  1. H. Alharazin, J. Yu. Panteleeva, and B.-D. Sun, Diffusion Models for SU(2) Lattice Gauge Theory in Two Dimensions (2026).
  2. J. Komijani, M. K. Marinkovic, and L. Turgut, Diffusion model for SU(N) gauge theories (2026).
  3. L. Winkler, A. Leaver-Fay, J. Kleinhenz, and P. Kessel, Diffeomorphic Optimization (2026).
  4. jkomijani/lattice_ml: active implementation. Relevant modules are diffusion, gauge_tools, and integrate. The Wilson action is in gauge_action.py, the intrinsic process in _lie_diffusion_process.py, and HMD/Langevin correctors in gauge/_lie_corrector.py.
  5. jkomijani/diffusion: archived predecessor; development moved to lattice_ml.diffusion.

Conventions for signs and plaquette counting differ across sources. This page uses the positive-shifted Wilson action when discussing minimization; the code repository implements the equivalent negative-trace convention.

Interactive tutorial · equations rendered with MathJax · all simulations run locally in this page