By the end you will see the sphere’s covariant derivative, Christoffel symbols, and parallel transport as three views of one two-word act — with one-forms as an optional extension.
For the sphere’s canonical, length-preserving derivative: differentiate a tangent vector field in the surrounding space, then delete the part of the result that sticks out of the surface. This Levi-Civita covariant derivative is “differentiate, then project.”
The \(\Gamma\)-symbols and transport equation are intrinsic bookkeeping for that deletion, written in coordinates that live on the surface and cannot see the outside. More general connections exist; this page builds the canonical one selected by the sphere’s metric.
A derivative along a curve \(\gamma(t)\) tries to compare \(V(\gamma(t+\varepsilon))\) with \(V(\gamma(t))\) as \(\varepsilon\to0\). For a scalar function this is innocent — two numbers, subtract. For a vector field on a sphere it is quietly illegal: \(V(p)\) lives in the tangent plane at \(p\), and \(V(q)\) lives in the tangent plane at \(q\). These are two different two-dimensional vector spaces, tilted against each other in \(\mathbb{R}^3\). There is no God-given way to subtract an arrow in one from an arrow in the other.
The tempting cheat is to subtract components: write both vectors in local coordinates and difference the numbers. The plate below shows why that fails. Take the arrow at \(P\), copy its components \((V^\theta, V^\varphi)\) to \(Q\), and look at what you get: a genuinely different arrow pointing a genuinely different way. Component-subtraction silently assumes an identification between the two planes — and the choice of that identification is the missing structure. Supplying it, consistently, is exactly what a connection does. The covariant derivative is the derivative you get once it’s supplied.
The sphere has one great advantage: it sits inside \(\mathbb{R}^3\), where differentiation is unproblematic. So differentiate there. Take a vector field along a curve — say the field \(\hat e_\varphi\), “always point due east,” along a circle of latitude — and compute the honest ambient derivative \(\frac{dV}{dt}\in\mathbb{R}^3\). The result generally pokes out of the surface.
Now notice the poking-out part isn’t the field’s fault. It’s the surface’s fault — the sphere curves away underneath you, so even a field doing its best to stay constant gets dragged. That component carries no information about how the field changes within the sphere’s own world. So delete it:
\[ \nabla_X V \;=\; \underbrace{P_{\text{tangent}}}_{\text{project}}\big(\underbrace{D_X V}_{\text{ordinary derivative}}\big) \]
That’s the sphere’s Levi-Civita covariant derivative. For the eastward field at polar angle \(\theta\) you can watch the split happen below: the ambient derivative \(\partial_\varphi \hat e_\varphi\) decomposes as \(-\sin\theta\,\hat r \;-\;\cos\theta\,\hat e_\theta\). The \(\hat r\) piece is normal — deleted. What survives is \(\nabla_{\partial_\varphi}\hat e_\varphi=-\cos\theta\,\hat e_\theta\): “due east” rotates poleward at rate \(\cos\theta\) per radian of longitude. Per unit distance, the direction is \(\hat e_\varphi=(\sin\theta)^{-1}\partial_\varphi\), so \(\nabla_{\hat e_\varphi}\hat e_\varphi=-\cot\theta\,\hat e_\theta\).
For a surface \(S\subset\mathbb{R}^3\) with tangent projection \(P_p:\mathbb{R}^3\to T_pS\), define \(\nabla_X Y := P\,(D_X Y)\) where \(D\) is the flat directional derivative of any smooth extension. One checks: \(\nabla\) is \(\mathbb{R}\)-bilinear, tensorial in \(X\), satisfies the Leibniz rule \(\nabla_X(fY)=(Xf)Y+f\nabla_XY\), is metric-compatible \(\big(X\langle Y,Z\rangle = \langle\nabla_XY,Z\rangle+\langle Y,\nabla_XZ\rangle\), because projection is self-adjoint and the discarded part is orthogonal to all tangents\(\big)\), and is torsion-free \((\nabla_XY-\nabla_YX=[X,Y])\). By the fundamental theorem of Riemannian geometry those two properties pin the connection down uniquely — so “differentiate, then project” isn’t a choice of connection, it is the canonical one, and the intrinsic formula in Step 3 recovers it with no reference to the embedding.
Now go intrinsic. On the sphere use \((\theta,\varphi)\) and the coordinate basis \(e_\theta=\partial_\theta\mathbf r\), \(e_\varphi=\partial_\varphi\mathbf r\). Watch the notation change: Step 2 used the unit east arrow \(\hat e_\varphi\), whereas the coordinate arrow is \(e_\varphi=\sin\theta\,\hat e_\varphi\). It shrinks near the pole because one radian of longitude covers less distance there. Write \(V=V^k e_k\) and differentiate with the product rule:
\[ \partial_i V \;=\; (\partial_i V^k)\,e_k \;+\; V^k\,\partial_i e_k. \]
The first term is what a naive component-derivative sees. The second is what it misses: the basis arrows themselves swing around as you move. On a sphere, “the \(e_\varphi\) direction” at one longitude is not “the \(e_\varphi\) direction” at the next. Project that basis-swing onto the surface and record its tangential components — those recorded numbers are the Christoffel symbols:
\[ P\big(\partial_i e_j\big) \;=\; \Gamma^k_{\;ij}\, e_k \qquad\Longrightarrow\qquad \nabla_i V^k \;=\; \partial_i V^k \;+\; \Gamma^k_{\;ij} V^j . \]
A Christoffel symbol is not an extra force or a piece of curvature by itself. It is one coefficient in the expansion \[ \nabla_{e_i}e_j=\Gamma^k_{\;ij}e_k. \] In words: move in coordinate direction \(i\), watch basis arrow \(j\), and record how much of its tangential change points along basis direction \(k\). Repeated \(k\) is summed, so the full change is assembled from all output directions.
Where do you move?
The first lower index selects the direction of differentiation.
What do you watch?
The second lower index selects the basis vector being differentiated.
Where does the change point?
The upper index selects one component of the resulting tangent vector.
For example, \(\Gamma^\theta_{\;\varphi\varphi}\) asks: move east in the \(\varphi\) direction; watch the east-pointing coordinate arrow \(e_\varphi\); how much does its change point south along \(e_\theta\)? On the unit sphere the answer is \(-\sin\theta\cos\theta\).
1. Nudge the coordinate. Replace \(x^i\) by \(x^i+\delta\) and compare the same basis field \(e_j\) at the two nearby points in the surrounding space: \[ \frac{e_j(x^1,\ldots,x^i+\delta,\ldots)-e_j(x^1,\ldots,x^i,\ldots)}{\delta} \;\longrightarrow\; \partial_i e_j. \]
2. Remove the normal change. Only the tangent part describes change within the surface: \[ P_{\mathrm{tangent}}(\partial_i e_j)=\nabla_{e_i}e_j. \]
3. Expand what remains. Every tangent vector is a combination of the local basis arrows, so there must be coefficients \(\Gamma^k_{\;ij}\) satisfying \[ P_{\mathrm{tangent}}(\partial_i e_j)=\Gamma^k_{\;ij}e_k. \] Those coefficients are the Christoffel symbols. Because a coordinate basis need not be orthonormal, extracting them uses the inverse metric: \[ \boxed{\Gamma^k_{\;ij}=g^{k\ell}\left\langle\partial_i e_j,e_\ell\right\rangle}. \]
The metric entries \(g_{j\ell}=\langle e_j,e_\ell\rangle\) record the lengths and mutual angles of the coordinate basis. Differentiate one entry:
\[ \partial_i g_{j\ell} =\left\langle\nabla_i e_j,e_\ell\right\rangle +\left\langle e_j,\nabla_i e_\ell\right\rangle. \]
Write this equation again with \(i,j,\ell\) cyclically exchanged. Add the first two versions and subtract the third. Metric compatibility preserves the inner products, while torsion-freeness gives \(\nabla_i e_j=\nabla_j e_i\) for a coordinate basis. The unwanted terms cancel:
\[ 2\left\langle\nabla_i e_j,e_\ell\right\rangle =\partial_i g_{j\ell}+\partial_j g_{i\ell}-\partial_\ell g_{ij}. \]
Multiplying by \(g^{k\ell}\) raises the remaining index and isolates the coefficient:
\[ \boxed{\Gamma^k_{\;ij} =\frac12 g^{k\ell} \left(\partial_i g_{j\ell}+\partial_j g_{i\ell}-\partial_\ell g_{ij}\right)}. \]
For the unit sphere, \(g_{\theta\theta}=1\), \(g_{\varphi\varphi}=\sin^2\theta\), and \(g^{\varphi\varphi}=1/\sin^2\theta\). Therefore
\[ \Gamma^\theta_{\;\varphi\varphi} =-\frac12\,\partial_\theta(\sin^2\theta) =-\sin\theta\cos\theta, \qquad \Gamma^\varphi_{\;\theta\varphi} =\frac12\frac{1}{\sin^2\theta}\, \partial_\theta(\sin^2\theta) =\cot\theta. \]
The direct basis measurement and the metric-only calculation agree. Plate III performs the first derivation numerically: its finite quotient converges to these two coefficients as \(\delta\to0\).
Read the corrected formula as a sentence: total change = the components changed + the rulers changed underneath the components. Nothing about \(\Gamma\) is mysterious or intrinsically “curvy” — polar coordinates on a flat sheet of paper have nonzero \(\Gamma\)s too, because their basis arrows also swivel. \(\Gamma\) is a property of your grid, not (directly) of the geometry. On the sphere: \(\Gamma^\theta_{\;\varphi\varphi}=-\sin\theta\cos\theta\) and \(\Gamma^\varphi_{\;\theta\varphi}=\cot\theta\), which the plate below lets you read directly off the swinging arrows.
The collection \(\Gamma^k_{\;ij}\) is not a tensor: changing coordinates changes both the basis arrows and their derivatives, producing an extra second-derivative term in the transformation law. At any chosen point one can use normal coordinates to make every Christoffel symbol vanish there, even on a curved sphere. What cannot generally be removed throughout a neighborhood is the path-dependence encoded by the curvature tensor, which combines derivatives and quadratic products of \(\Gamma\). Thus nonzero \(\Gamma\) can arise from coordinates alone, while nonzero curvature is geometric.
Once you can differentiate, you can say what “a constant vector along a curve” means: one whose covariant derivative vanishes, \(\nabla_{\dot\gamma} V = 0\). Componentwise this is a linear ODE, \(\dot V^k = -\Gamma^k_{\;ij}\,\dot x^i V^j\) — at every instant, counter-rotate the components by exactly the amount the basis swings, so the arrow itself never turns. This is parallel transport: the connection’s answer to Step 1’s missing identification between tangent planes, delivered curve by curve.
And here is the punchline of the whole subject. On flat paper, transport a vector around any closed loop and it returns identical — the identification doesn’t care about the path. On the sphere, carry a vector around a circle of latitude, never once turning it, and it comes back rotated. If \(\theta_0\) is the colatitude measured down from the north pole, the mismatch angle is \(2\pi(1-\cos\theta_0)\) — precisely the solid angle the loop encloses, \(\iint K\,dA\). Path-dependence of transport is curvature; the sphere confesses through an angle you can measure below. The plate integrates the actual transport ODE (RK4), so the arrow’s return angle and unchanged length are computed, not drawn.
Along \(\theta=\theta_0,\ \varphi=t\): \(\dot V^\theta = \sin\theta_0\cos\theta_0\, V^\varphi\), \(\dot V^\varphi = -\cot\theta_0\, V^\theta\). In the orthonormal frame \((\hat e_\theta, \hat e_\varphi)\) this is rigid rotation at rate \(-\cos\theta_0\) per unit \(\varphi\); after a full loop the frame-relative rotation is \(-2\pi\cos\theta_0\), a mismatch of \(2\pi(1-\cos\theta_0)\bmod 2\pi\). For the unit sphere \(K\equiv 1\), and the spherical cap above the loop has area \(2\pi(1-\cos\theta_0)\) — so the holonomy equals \(\int_{\text{cap}} K\,dA\), the baby case of the general fact that curvature is the infinitesimal holonomy: \(R(X,Y)Z = (\nabla_X\nabla_Y - \nabla_Y\nabla_X - \nabla_{[X,Y]})Z\), the failure of transports along the two edges of a small parallelogram to commute.
The main story ends with holonomy in Plate IV. Open this bonus section to see how the same connection differentiates measuring devices as well as arrows.
A one-form \(\omega\) at a point is not an arrow; it is a measuring device for arrows — a linear machine that eats a tangent vector and returns a number. The right mental picture is a ruled stack of evenly spaced level lines laid across the tangent plane: feed it a vector \(v\), and \(\omega(v)\) is the number of lines the arrow pierces. Denser ruling = stronger form. The prototype is the differential \(df\) of a function: its level lines are literally the level sets of \(f\), and \(df(v)\) is how many contour lines you cross moving along \(v\). At the starting point on the sphere, take \(f=z\) (height): its ruling is the circles of latitude, spaced \(1/\sin\theta\) apart in the tangent plane — dense near the equator, sparse near the poles, exactly matching \(df = -\sin\theta\, d\theta\). Once transported, the sheet is a carried copy of that initial \(df\), not generally the local height differential at its new point.
How should \(\nabla\) act on \(\omega\)? You don’t get to choose — it’s forced. A number like \(\omega(V)\) is a scalar, and scalars must obey the plain product rule: \(\partial_\mu\big(\omega_\nu V^\nu\big) = (\nabla_\mu\omega)(V) + \omega(\nabla_\mu V)\). Solving for \(\nabla\omega\):
\[ \nabla_\mu\, \omega_\nu \;=\; \partial_\mu \omega_\nu \;-\; \Gamma^\lambda_{\;\mu\nu}\,\omega_\lambda \]
— one \(+\Gamma\) correction for an upper vector index and one \(-\Gamma\) correction for a lower covector index. The plate makes their cancellation intuitive: parallel-transport the arrow and the ruled sheet together around the latitude circle. Their coordinate components obey different index rules, but in the orthonormal tangent frame the physical arrow and measuring sheet co-rotate. The reading \(\omega(V)\), the pierced-line count, therefore stays frozen to the last decimal: measurements do not drift when both objects are parallel.
Demand that \(\nabla\) obey Leibniz over the pairing and reduce to \(\partial\) on scalars: \[ \partial_\mu(\omega_\nu V^\nu) = (\nabla_\mu\omega_\nu)V^\nu + \omega_\nu\big(\partial_\mu V^\nu + \Gamma^\nu_{\;\mu\lambda}V^\lambda\big). \] Expand the left side with the ordinary product rule, cancel \((\partial_\mu\omega_\nu)V^\nu + \omega_\nu\partial_\mu V^\nu\) from both sides, and relabel dummy indices: \((\nabla_\mu\omega_\nu)V^\nu = -\,\Gamma^\lambda_{\;\mu\nu}\omega_\lambda V^\nu\) for all \(V\), hence the boxed formula. The same argument propagates \(\nabla\) to every tensor rank: one \(+\Gamma\) per upper index, one \(-\Gamma\) per lower index. Transport-invariance of the pairing — what Plate V shows numerically — is the coordinate-free statement of the same fact.
For the sphere’s Levi-Civita connection, one act — differentiate, then project — produces four core ideas:
Optional dual extension (Plate V): one-forms are ruled measuring sheets; the lower-index \(-\Gamma\) rule keeps pairings \(\omega(V)\) honest under transport.
General relativity. A geodesic is a curve that parallel-transports its own velocity: \(\ddot x^k + \Gamma^k_{\;ij}\dot x^i \dot x^j = 0\). “Gravity” in GR is nothing added to this equation — it is this equation, with \(\Gamma\) determined by the metric of spacetime. Free fall is Plate IV applied to your own four-velocity.
Natural gradient & information geometry. The loss gradient \(\partial_\theta \mathcal L\) is a one-form on parameter space — a measuring device (“how fast does loss change per unit move this way?”), not a direction. Turning it into a step requires a metric; choosing the Fisher information gives the natural gradient \(F^{-1}\partial_\theta\mathcal L\), and the statistical manifold’s Christoffel symbols govern how estimators and flows must be corrected under reparametrization — the exact “rulers turning underneath you” story of Step 3, with coordinate charts replaced by model parametrizations.
Riemannian optimization. Momentum methods on manifolds (SGD on the Stiefel or SPD manifold, hyperbolic embeddings) must parallel-transport the momentum vector from the previous iterate’s tangent space to the current one before adding — Plate I’s illegality, resolved by Plate IV’s transport, inside an optimizer loop.