Skip to content

Classical Field Theory and the Lagrangian: From a Variational Principle for Infinitely Many Degrees of Freedom to Noether's Theorem

Prerequisite:Why We Need Quantum Field Theory: One-Particle Quantum Mechanics Is Incompatible with Relativity

Raw
  • A field is a quantity that takes a value at every point of space at every instant of time. Mechanically it is a system with infinitely many degrees of freedom, labelled by the points of space; taking the lattice spacing of a chain of coupled oscillators to zero makes this picture concrete.
  • The action of a field can always be written as S[ϕ]=d4xL(ϕa,μϕa)S[\phi] = \int d^4x\, \mathcal{L}(\phi_a, \partial_\mu \phi_a). If the integrand — the Lagrangian density L\mathcal{L} — is a Lorentz scalar, the equations of motion that follow from it are automatically Lorentz covariant. This is the single strongest reason to start from the Lagrangian formalism.
  • The variational principle δS=0\delta S = 0 yields the Euler–Lagrange equations for fields, μ(L/(μϕa))L/ϕa=0\partial_\mu\big(\partial\mathcal{L}/\partial(\partial_\mu\phi_a)\big) - \partial\mathcal{L}/\partial\phi_a = 0 (Theorem 4.3).
  • The simplest Lagrangian density allowed for a real scalar field is that of the Klein–Gordon field, and the plane-wave solutions of its equation of motion satisfy E2=p2+m2E^2 = \boldsymbol{p}^2 + m^2. This is why mm is called a mass.
  • Every continuous symmetry of the action corresponds to one conserved current, μjμ=0\partial_\mu j^\mu = 0 (Noether’s theorem, Theorem 6.2). Spacetime translations give the energy–momentum tensor; a global U(1)U(1) phase rotation gives a conserved charge.
  • The classical field theory assembled here is the starting point for canonical quantization and path-integral quantization in the chapters that follow. Quantization means laying quantum rules on top of the space of classical field configurations.

1. Motivation: what is it that we quantize?

Section titled “1. Motivation: what is it that we quantize?”

In the previous chapter we saw that making relativity and quantum mechanics coexist forces the particle number to change, so that the fundamental object must be a field rather than a particle wave function. But before we can speak of quantizing a field, we need the classical field theory that is to be quantized. Building that foundation is the purpose of this chapter.

For a system with finitely many degrees of freedom, classical mechanics offers three formulations: Newton’s equations, the Lagrangian formalism and the Hamiltonian formalism. All three are available in field theory as well, but as a starting point the Lagrangian formalism is overwhelmingly the most convenient. There are four reasons.

First, relativistic invariance stays manifest. The Hamiltonian is the generator of time evolution, so it cannot help but single out the time coordinate from the three spatial ones. The action S=d4xLS = \int d^4x\,\mathcal{L}, by contrast, is an integral over four-dimensional spacetime, and if L\mathcal{L} is chosen to be a Lorentz scalar then SS itself is Lorentz invariant. Since the equations of motion are fixed as the stationarity condition of SS, Lorentz covariance is guaranteed the moment the expression is written down. For the requirements of special relativity see Lorentz transformations.

Second, it is the natural language for symmetries. Internal symmetries and gauge symmetries alike can be stated in a single line as the invariance of L\mathcal{L}. And by Noether’s theorem, described below, conservation laws can then be read off mechanically from the symmetries.

Third, the path integral uses the action itself. Transition amplitudes in quantum theory take the form DϕeiS[ϕ]\int \mathcal{D}\phi\, e^{iS[\phi]} (path-integral quantization). What appears here is the action, not the Hamiltonian.

Fourth, it serves as the blueprint of a theory. The modern recipe for building a field theory is: specify the field content and the symmetries, then list every term compatible with them in order of increasing mass dimension. Rather than postulating interactions out of thin air, one narrows the candidates down by symmetry and dimensional counting. This methodology works only because all the information in the theory is concentrated in a single function L\mathcal{L}.

Historically, too, the Lagrangian formalism for fields is older than quantum theory. Nineteenth-century elasticity already treated the vibrations of continuous media in Lagrangian form. The theory of the electromagnetic field, begun with Faraday’s lines of force and completed in Maxwell’s equations, had by the early twentieth century arrived at the recognition that the electromagnetic field is an independent mechanical degree of freedom carrying energy and momentum. Noether’s 1918 paper established the relation between symmetries and conservation laws in general form, and from the late 1920s Dirac, Jordan, Heisenberg and Pauli began attempting to quantize these classical fields. What we are about to retrace is that first step.


2. Preliminaries: notation and conventions

Section titled “2. Preliminaries: notation and conventions”

Throughout we use natural units =c=1\hbar = c = 1. In these units length, time and inverse mass all carry the same dimension, so the dimension of any quantity can be expressed by a single number, its mass dimension. The notational conventions are collected in the table below.

SymbolMeaning and convention
ημν=diag(1,1,1,1)\eta_{\mu\nu} = \mathrm{diag}(1,-1,-1,-1)Minkowski metric. ημν\eta^{\mu\nu} is its inverse, with the same components
xμ=(x0,x1,x2,x3)=(t,x)x^\mu = (x^0, x^1, x^2, x^3) = (t, \boldsymbol{x})Spacetime coordinates. Greek indices run over 0,1,2,30,1,2,3, Latin indices i,ji,j over 1,2,31,2,3
μ=/xμ\partial_\mu = \partial/\partial x^\muμ=ημνν=(t,)\partial^\mu = \eta^{\mu\nu}\partial_\nu = (\partial_t, -\nabla)
=μμ=t22\Box = \partial_\mu \partial^\mu = \partial_t^2 - \nabla^2d’Alembert operator
a,ba, bIndices distinguishing the components (internal degrees of freedom) of a field. Repeated indices are always summed
d4x=dtd3xd^4x = dt\, d^3xSpacetime volume element. Invariant under Lorentz transformations (their determinant is 11)

Write [][\,\cdot\,] for mass dimension. For eiSe^{iS} to make sense, SS must be dimensionless, and since [d4x]=4[d^4x] = -4 we need

[L]=4.[\mathcal{L}] = 4 .

This one line is the yardstick we shall use later to classify interaction terms.


3. The continuum limit: from point masses to fields

Section titled “3. The continuum limit: from point masses to fields”

Let us verify the statement that a field is a mechanical system with infinitely many degrees of freedom on a concrete example. Consider a one-dimensional chain of point masses joined by springs and let the lattice spacing go to zero. This example has been standard since Goldstein’s textbook, but it is worth once following the mechanical ancestry of the concept of a field with one’s own hands.

Discrete: N degrees of freedom y_1(t), …, y_N(t)aa → 0 (continuum limit)Continuum: one value at each point x (uncountably many)phi(x, t)x
In the limit where the lattice spacing a of the discrete spring-and-mass system (top) goes to zero, a continuum field emerges. In the discrete system the degrees of freedom number as many as the point masses; in the limit there is one degree of freedom for each point of space.

Example 3.1Continuum limit of a spring chain and the wave equation

Let NN point masses of mass mm be arranged along a line with spacing aa, neighbouring masses joined by springs of spring constant kk. Writing yi(t)y_i(t) for the displacement of the ii-th mass from its equilibrium position, the Lagrangian is

L=i[12my˙i212k(yi+1yi)2].L = \sum_{i} \left[ \frac{1}{2} m\, \dot{y}_i^{\,2} - \frac{1}{2} k\, (y_{i+1} - y_i)^2 \right] .

Apply the finite-dimensional Lagrange equation ddtLy˙i=Lyi\dfrac{d}{dt}\dfrac{\partial L}{\partial \dot{y}_i} = \dfrac{\partial L}{\partial y_i}. The left-hand side is my¨im\ddot{y}_i; the right-hand side receives contributions from the two terms containing yiy_i, namely 12k(yiyi1)2-\frac{1}{2}k(y_i - y_{i-1})^2 and 12k(yi+1yi)2-\frac{1}{2}k(y_{i+1} - y_i)^2, giving

Lyi=k(yiyi1)+k(yi+1yi)=k(yi+12yi+yi1).\frac{\partial L}{\partial y_i} = -k(y_i - y_{i-1}) + k(y_{i+1} - y_i) = k\,(y_{i+1} - 2y_i + y_{i-1}) .

The equation of motion is therefore my¨i=k(yi+12yi+yi1)m\ddot{y}_i = k(y_{i+1} - 2y_i + y_{i-1}).

Now divide both sides by aa and take a0a \to 0, NN \to \infty (with the total length NaNa held fixed) while keeping the linear density μ:=m/a\mu := m/a and the Young modulus Y:=kaY := ka fixed. Writing the position of the ii-th mass as x=iax = ia and reading yi(t)ϕ(x,t)y_i(t) \to \phi(x,t), we get

k(yi+12yi+yi1)a=(ka)yi+12yi+yi1a2Y2ϕx2\frac{k(y_{i+1} - 2y_i + y_{i-1})}{a} = (ka)\,\frac{y_{i+1} - 2y_i + y_{i-1}}{a^2} \longrightarrow Y\,\frac{\partial^2 \phi}{\partial x^2}

(using the convergence of the central difference quotient to the second derivative). Hence

μ2ϕt2=Y2ϕx2,\mu\, \frac{\partial^2 \phi}{\partial t^2} = Y\, \frac{\partial^2 \phi}{\partial x^2} ,

a wave equation with wave speed v=Y/μv = \sqrt{Y/\mu}.

Taking the same limit in the Lagrangian gives

L=ia[12may˙i212(ka)(yi+1yia)2]dx[12μ(tϕ)212Y(xϕ)2]=: L.L = \sum_i a \left[ \frac{1}{2}\frac{m}{a}\dot{y}_i^{\,2} - \frac{1}{2}(ka)\left(\frac{y_{i+1}-y_i}{a}\right)^2 \right] \longrightarrow \int dx\, \underbrace{\left[ \frac{1}{2}\mu\,(\partial_t \phi)^2 - \frac{1}{2}Y (\partial_x \phi)^2 \right]}_{=:\ \mathcal{L}} .

The sum ia\sum_i a has become an integral dx\int dx, and the integrand that appears is the Lagrangian density L\mathcal{L}.

From this computation we can read off a dictionary between finitely many degrees of freedom and field theory.

Finite-dimensional systemField theory
Index i=1,,Ni = 1,\dots,NSpatial coordinate xR3\boldsymbol{x} \in \mathbb{R}^3 (a continuous “index”)
Generalized coordinate qi(t)q_i(t)Field ϕ(x,t)\phi(\boldsymbol{x}, t)
i\sum_id3x\int d^3x
Lagrangian L(qi,q˙i)L(q_i, \dot{q}_i)L=d3xL(ϕ,μϕ)L = \int d^3x\, \mathcal{L}(\phi, \partial_\mu\phi)
Action S=dtLS = \int dt\, LAction S=d4xLS = \int d^4x\, \mathcal{L}
L/qi\partial L/\partial q_iFunctional derivative δS/δϕ(x)\delta S/\delta \phi(x)

Remark 3.2A field need not be the displacement of anything

In the example above, ϕ\phi had the concrete meaning of the displacement of an elastic medium. The fields of field theory, however, are not necessarily displacements of any medium. Nineteenth-century physicists tried to understand the electromagnetic field as a strain in the ether, but the Michelson–Morley experiment and special relativity denied that picture. The electromagnetic field, and the scalar fields treated from the next section onward, are themselves fundamental mechanical degrees of freedom; there is no vibrating medium behind them.

The continuum-limit argument supplies a mathematical template for handling systems with infinitely many degrees of freedom; it does not supply an ontology of fields. Note in particular that we adjusted μ\mu and YY precisely so that L\mathcal{L} would remain finite in the limit a0a \to 0. This “tuning of how the limit is taken” is the prototype of an operation that reappears later in renormalization theory (introduction to renormalization).


4. The variational principle for fields and the Euler–Lagrange equations

Section titled “4. The variational principle for fields and the Euler–Lagrange equations”

Definition 4.1Action functional of a field and Lagrangian density

Let ΩR4\Omega \subset \mathbb{R}^4 be a bounded region of spacetime with smooth boundary Ω\partial\Omega, and let ϕa:ΩR\phi_a : \Omega \to \mathbb{R} (a=1,,na = 1,\dots,n) be a collection of fields of class C2C^2.

Let L=L(ϕa,μϕa)\mathcal{L} = \mathcal{L}(\phi_a, \partial_\mu \phi_a) be a C2C^2 function of the nn real variables ϕa\phi_a and the 4n4n real variables μϕa\partial_\mu\phi_a; we call it the Lagrangian density. Then

S[ϕ]:=Ωd4x L(ϕa(x),μϕa(x))S[\phi] := \int_\Omega d^4x\ \mathcal{L}\big(\phi_a(x), \partial_\mu \phi_a(x)\big)

is called the action functional.

We are assuming here that L\mathcal{L} contains no derivatives of ϕa\phi_a beyond the first, and that it does not depend explicitly on the coordinate xx. The former is required in order to keep the equations of motion second order; the latter expresses translation invariance of spacetime, that is, the absence of any externally imposed background.

Before stating the variational principle we need the fundamental lemma of the calculus of variations. It says that if the integral of δS\delta S vanishes for every test function, then the integrand vanishes; without it the equations of motion would emerge only in integrated form.

Lemma 4.2Fundamental lemma of the calculus of variations

Let ΩR4\Omega \subset \mathbb{R}^4 be a nonempty open set and f:ΩRf : \Omega \to \mathbb{R} a continuous function. If

Ωf(x)η(x)d4x=0\int_\Omega f(x)\, \eta(x)\, d^4x = 0

holds for every CC^\infty function η\eta with compact support contained in Ω\Omega, then f0f \equiv 0 on Ω\Omega.

Proof(Lemma 4.2)

We prove the contrapositive. Suppose f(x0)0f(x_0) \ne 0 at some point x0Ωx_0 \in \Omega; without loss of generality c:=f(x0)>0c := f(x_0) > 0 (if f(x0)<0f(x_0) < 0, replace ff by f-f).

Since ff is continuous, taking an open ball B:={x:xx0<r}B := \{ x : |x - x_0| < r \} contained in Ω\Omega small enough (here |\cdot| is the Euclidean norm on the coordinates of R4\mathbb{R}^4 — a purely topological tool, unrelated to the metric ημν\eta_{\mu\nu}), we have f(x)>c/2f(x) > c/2 on BB.

As a test function supported on this BB take the standard bump function

η(x):={exp ⁣(1r2xx02)(xx0<r)0(xx0r).\eta(x) := \begin{cases} \exp\!\left(\dfrac{-1}{r^2 - |x - x_0|^2}\right) & (|x - x_0| < r) \\[2mm] 0 & (|x-x_0| \ge r) . \end{cases}

It is of class CC^\infty, positive on BB and zero outside BB. Therefore

Ωfηd4x=Bfηd4x>c2Bηd4x>0,\int_\Omega f \eta\, d^4x = \int_B f \eta\, d^4x > \frac{c}{2}\int_B \eta\, d^4x > 0 ,

contradicting the hypothesis. Hence f=0f = 0 at every point of Ω\Omega.

Theorem 4.3Euler–Lagrange equations for fields

In the setting of Definition 4.1, suppose a field configuration ϕ=(ϕa)\phi = (\phi_a) is a stationary point of SS in the following sense: for every collection η=(ηa)\eta = (\eta_a) of CC^\infty functions with compact support contained in Ω\Omega,

ddεε=0S[ϕ+εη]=0.\left.\frac{d}{d\varepsilon}\right|_{\varepsilon = 0} S[\phi + \varepsilon \eta] = 0 .

Then on Ω\Omega

μ(L(μϕa))Lϕa=0(a=1,,n).\partial_\mu \left( \frac{\partial \mathcal{L}}{\partial (\partial_\mu \phi_a)} \right) - \frac{\partial \mathcal{L}}{\partial \phi_a} = 0 \qquad (a = 1, \dots, n) .

Conversely, any ϕ\phi satisfying this system is a stationary point in the above sense. Here μ\partial_\mu denotes the total derivative with respect to xx, including the dependence through ϕa(x)\phi_a(x).

Proof(Theorem 4.3)

Put ϕaε:=ϕa+εηa\phi_a^\varepsilon := \phi_a + \varepsilon \eta_a. Since the support of ηa\eta_a is compact and contained in Ω\Omega, we have ηa=0\eta_a = 0, and hence μηa=0\partial_\mu \eta_a = 0, in a neighbourhood of Ω\partial\Omega. As L\mathcal{L} is of class C2C^2 we may interchange the ε\varepsilon-derivative with the integral, and the chain rule gives

ddε0S[ϕε]=Ωd4x[Lϕaηa+L(μϕa)μηa]\left.\frac{d}{d\varepsilon}\right|_{0} S[\phi^\varepsilon] = \int_\Omega d^4x \left[ \frac{\partial \mathcal{L}}{\partial \phi_a}\, \eta_a + \frac{\partial \mathcal{L}}{\partial (\partial_\mu \phi_a)}\, \partial_\mu \eta_a \right]

(where we used μ(ϕa+εηa)=μϕa+εμηa\partial_\mu(\phi_a + \varepsilon\eta_a) = \partial_\mu\phi_a + \varepsilon\,\partial_\mu\eta_a).

In the second term use the Leibniz rule to move the derivative off ηa\eta_a:

L(μϕa)μηa=μ ⁣(L(μϕa)ηa)μ ⁣(L(μϕa))ηa.\frac{\partial \mathcal{L}}{\partial (\partial_\mu \phi_a)}\, \partial_\mu \eta_a = \partial_\mu\!\left( \frac{\partial \mathcal{L}}{\partial (\partial_\mu \phi_a)}\, \eta_a \right) - \partial_\mu\!\left( \frac{\partial \mathcal{L}}{\partial (\partial_\mu \phi_a)} \right) \eta_a .

The first term is a divergence, so by Gauss’s theorem it becomes a boundary integral:

Ωd4x μ ⁣(L(μϕa)ηa)=ΩdΣμ L(μϕa)ηa=0.\int_\Omega d^4x\ \partial_\mu\!\left( \frac{\partial \mathcal{L}}{\partial (\partial_\mu \phi_a)}\, \eta_a \right) = \oint_{\partial\Omega} d\Sigma_\mu\ \frac{\partial \mathcal{L}}{\partial (\partial_\mu \phi_a)}\, \eta_a = 0 .

The last equality holds because ηa=0\eta_a = 0 on Ω\partial\Omega. The boundary term drops out only thanks to this assumption; a variation with no boundary condition imposed yields no equations of motion. Therefore

ddε0S[ϕε]=Ωd4x [Lϕaμ ⁣(L(μϕa))]ηa.\left.\frac{d}{d\varepsilon}\right|_{0} S[\phi^\varepsilon] = \int_\Omega d^4x\ \left[ \frac{\partial \mathcal{L}}{\partial \phi_a} - \partial_\mu\!\left( \frac{\partial \mathcal{L}}{\partial (\partial_\mu \phi_a)} \right) \right] \eta_a .

By hypothesis this vanishes for every ηa\eta_a, so for each aa separately (take the η\eta of the other components to be 00) we may apply Lemma 4.2 to the expression in square brackets and obtain the conclusion. Since ϕa\phi_a and L\mathcal{L} are of class C2C^2, that expression is continuous and the hypothesis of the lemma is met.

The converse follows immediately by reading the identity above from right to left.

This is a system of nn coupled second-order partial differential equations: the equations of motion of the field. Compared with the finite-dimensional Lagrange equation ddtLq˙iLqi=0\frac{d}{dt}\frac{\partial L}{\partial \dot{q}_i} - \frac{\partial L}{\partial q_i} = 0, the only changes are that the time derivative d/dtd/dt has become the spacetime derivative μ\partial_\mu, and the one equation per index ii has become one equation per field component aa. Checking it on the elastic string of Example 3.1: for L=12μ(tϕ)212Y(xϕ)2\mathcal{L} = \frac{1}{2}\mu(\partial_t\phi)^2 - \frac{1}{2}Y(\partial_x\phi)^2 we have L/(tϕ)=μtϕ\partial \mathcal{L}/\partial(\partial_t\phi) = \mu\,\partial_t\phi, L/(xϕ)=Yxϕ\partial\mathcal{L}/\partial(\partial_x \phi) = -Y \partial_x\phi and L/ϕ=0\partial\mathcal{L}/\partial\phi = 0, so t(μtϕ)+x(Yxϕ)=0\partial_t(\mu\partial_t\phi) + \partial_x(-Y\partial_x\phi) = 0, which is exactly the wave equation obtained above by taking the continuum limit.

Remark 4.4Why it is good for the Lagrangian density to be a scalar

Suppose L\mathcal{L} is a Lorentz scalar, that is, L(x)=L(x)\mathcal{L}'(x') = \mathcal{L}(x) under a Lorentz transformation xΛxx \to \Lambda x. Since d4xd^4x is invariant (detΛ=1|\det\Lambda| = 1), SS is invariant too. The equations of motion are determined by stationarity of SS, so whatever is a solution in one inertial frame remains a solution when transported to another. In other words, Lorentz covariance of the equations of motion is guaranteed the instant L\mathcal{L} is written down.

This is an advantage the Hamiltonian formalism does not offer. The Hamiltonian H=d3xHH = \int d^3x\,\mathcal{H} is an integral over a spatial slice at a fixed time, so its behaviour under Lorentz transformations is not manifest.

4.3. The ambiguity in the Lagrangian density

Section titled “4.3. The ambiguity in the Lagrangian density”

The equations of motion do not determine L\mathcal{L} uniquely. The ambiguity by a total derivative(Proposition 4.4)[Lagrangian Mechanics] seen in the finite-dimensional case carries over verbatim to field theory. The following proposition is what will later justify extending Noether’s theorem to quasi-invariance.

Proposition 4.5A total-derivative term does not change the equations of motion

Let Kμ=Kμ(ϕa,x)K^\mu = K^\mu(\phi_a, x) be a C2C^2 function of the field values and the coordinates (with no dependence on derivatives of the fields). Then the equations of motion obtained from

L:=L+μKμ\mathcal{L}' := \mathcal{L} + \partial_\mu K^\mu

via Theorem 4.3 are identical to those obtained from L\mathcal{L}.

Proof(Proposition 4.5)

We have S[ϕ]=S[ϕ]+Ωd4xμKμS'[\phi] = S[\phi] + \int_\Omega d^4x\, \partial_\mu K^\mu. Applying Gauss’s theorem to the second term,

Ωd4x μKμ=ΩdΣμ Kμ(ϕa(x),x),\int_\Omega d^4x\ \partial_\mu K^\mu = \oint_{\partial\Omega} d\Sigma_\mu\ K^\mu\big(\phi_a(x), x\big) ,

which is determined by the values of ϕa\phi_a on the boundary Ω\partial\Omega alone.

In the variation of Theorem 4.3, ηa\eta_a vanishes in a neighbourhood of Ω\partial\Omega, so ϕaε=ϕa\phi_a^\varepsilon = \phi_a on Ω\partial\Omega and this boundary integral does not depend on ε\varepsilon. Hence

ddε0S[ϕε]=ddε0S[ϕε],\left.\frac{d}{d\varepsilon}\right|_0 S'[\phi^\varepsilon] = \left.\frac{d}{d\varepsilon}\right|_0 S[\phi^\varepsilon] ,

so the stationarity conditions coincide exactly, and so do the equations of motion obtained through Lemma 4.2.

(If KμK^\mu also depended on derivatives of the fields, the argument would not go through as it stands, since μηa\partial_\mu\eta_a need not vanish on the boundary. Only the above form is needed in what follows.)


Suppose we have a single real-valued field ϕ(x)\phi(x) which transforms as a scalar under Lorentz transformations, ϕ(x)=ϕ(x)\phi'(x') = \phi(x). We narrow down the Lagrangian densities allowed for this field by three requirements.

  1. L\mathcal{L} is a Lorentz scalar (Remark 4.4).
  2. L\mathcal{L} contains no derivatives of ϕ\phi beyond the first (so that the equations of motion are second order).
  3. L\mathcal{L} is at most quadratic in ϕ\phi (we want a free field, for which the superposition principle holds).

Requirements 1 and 2 leave μϕμϕ\partial_\mu\phi\,\partial^\mu\phi as the only scalar that can be built with derivatives (no scalar containing exactly one μϕ\partial_\mu\phi exists, since there is nothing to contract the index with). Together with requirement 3, the most general form is, with constants A,B,C,DA, B, C, D,

L=Aμϕμϕ+Bϕ2+Cϕ+D.\mathcal{L} = A\, \partial_\mu\phi\,\partial^\mu\phi + B\, \phi^2 + C\,\phi + D .

Here DD is a constant that does not contribute to the equations of motion, and CϕC\phi can be absorbed into Bϕ2B\phi^2 by a shift ϕϕ+const\phi \to \phi + \text{const} (when B0B \ne 0). The constant AA can be normalized to 1/21/2 by rescaling ϕ\phi by a constant factor. Writing the remaining BB as B=12m2B = -\frac{1}{2}m^2, we arrive at

LKG=12μϕμϕ12m2ϕ2=12ϕ˙212(ϕ)212m2ϕ2,\mathcal{L}_{\mathrm{KG}} = \frac{1}{2}\, \partial_\mu\phi\, \partial^\mu\phi - \frac{1}{2}\, m^2 \phi^2 = \frac{1}{2}\dot\phi^2 - \frac{1}{2}(\nabla\phi)^2 - \frac{1}{2}m^2\phi^2 ,

the Lagrangian density of the Klein–Gordon field.

The choice of signs has physical content. Unless A>0A > 0 (positive kinetic term), the energy density fails to be bounded below, as we shall see. Unless m2>0m^2 > 0, the configuration ϕ=0\phi = 0 is a maximum of the potential V(ϕ)=12m2ϕ2V(\phi) = \frac{1}{2}m^2\phi^2 and is unstable. What happens when m2<0m^2 < 0 belongs to the subject of spontaneous symmetry breaking and is treated in gauge theory and spontaneous symmetry breaking.

Example 5.1The Klein–Gordon equation and its dispersion relation

Apply Theorem 4.3 to LKG\mathcal{L}_{\mathrm{KG}}. First,

LKGϕ=m2ϕ.\frac{\partial \mathcal{L}_{\mathrm{KG}}}{\partial \phi} = -m^2\phi .

Next the partial derivative with respect to the derivatives. Writing μϕμϕ=ηαβαϕβϕ\partial_\mu\phi\,\partial^\mu\phi = \eta^{\alpha\beta}\,\partial_\alpha\phi\,\partial_\beta\phi and differentiating with respect to ρϕ\partial_\rho\phi, both factors contribute and

(ρϕ)(12ηαβαϕβϕ)=12(ηρββϕ+ηαραϕ)=ρϕ\frac{\partial}{\partial(\partial_\rho\phi)}\Big( \tfrac{1}{2}\eta^{\alpha\beta}\partial_\alpha\phi\,\partial_\beta\phi \Big) = \tfrac{1}{2}\big( \eta^{\rho\beta}\partial_\beta\phi + \eta^{\alpha\rho}\partial_\alpha\phi \big) = \partial^\rho \phi

(using the symmetry of η\eta). The Euler–Lagrange equation is therefore

(+m2)ϕ=μμϕ+m2ϕ=0,(\Box + m^2)\,\phi = \partial_\mu\partial^\mu\phi + m^2\phi = 0 ,

the Klein–Gordon equation.

Substitute a plane wave ϕ(x)=eikx\phi(x) = e^{-i k\cdot x} (with kx=kμxμ=k0tkxk\cdot x = k_\mu x^\mu = k^0 t - \boldsymbol{k}\cdot\boldsymbol{x}). Since μeikx=ikμeikx\partial_\mu e^{-ik\cdot x} = -i k_\mu e^{-ik\cdot x}, we have eikx=(ikμ)(ikμ)eikx=kμkμeikx\Box e^{-ik\cdot x} = (-ik_\mu)(-ik^\mu) e^{-ik\cdot x} = -k_\mu k^\mu\, e^{-ik\cdot x}, so

(kμkμ+m2)eikx=0    kμkμ=m2    (k0)2=k2+m2.(-k_\mu k^\mu + m^2)\, e^{-ik\cdot x} = 0 \iff k_\mu k^\mu = m^2 \iff (k^0)^2 = |\boldsymbol{k}|^2 + m^2 .

Reading k0=Ek^0 = E and k=p\boldsymbol{k} = \boldsymbol{p} in units =c=1\hbar = c = 1, this is nothing but the relativistic energy–momentum relation

E2=p2+m2E^2 = \boldsymbol{p}^2 + m^2

(see the energy–momentum relation(Theorem 4.4)[Relativistic Mechanics] in relativistic mechanics). This is why the parameter mm is called a mass. Since ϕ\phi is real, the general solution can be written, with ωk:=k2+m2\omega_{\boldsymbol{k}} := \sqrt{|\boldsymbol{k}|^2 + m^2}, as

ϕ(x)=d3k(2π)32ωk[α(k)eikx+α(k)eikx]k0=ωk.\phi(x) = \int \frac{d^3k}{(2\pi)^3\, 2\omega_{\boldsymbol{k}}} \left[ \alpha(\boldsymbol{k})\, e^{-ik\cdot x} + \alpha(\boldsymbol{k})^{*}\, e^{ik\cdot x} \right]_{k^0 = \omega_{\boldsymbol{k}}} .

Canonical quantization is precisely the step in which α(k)\alpha(\boldsymbol{k}) turns into an operator a^(k)\hat{a}(\boldsymbol{k}) (canonical quantization of the scalar field).

5.2. A static source and the Yukawa potential

Section titled “5.2. A static source and the Yukawa potential”

Let us confirm from another angle that mm really acts as a mass in the physics. Placing a static point source in a Klein–Gordon field makes the range of the resulting force 1/m1/m. This is exactly the calculation by which Yukawa predicted the meson in 1935.

Example 5.2The field of a point source and the range of the nuclear force

Represent a point source at rest at the origin, coupled to the field, by

L=12μϕμϕ12m2ϕ2+gϕδ3(x).\mathcal{L} = \frac{1}{2}\partial_\mu\phi\,\partial^\mu\phi - \frac{1}{2}m^2\phi^2 + g\,\phi\,\delta^3(\boldsymbol{x}) .

By Theorem 4.3, with L/ϕ=m2ϕ+gδ3(x)\partial\mathcal{L}/\partial\phi = -m^2\phi + g\delta^3(\boldsymbol{x}) and μ(L/(μϕ))=ϕ\partial_\mu(\partial\mathcal{L}/\partial(\partial_\mu\phi)) = \Box\phi,

(+m2)ϕ=gδ3(x).(\Box + m^2)\phi = g\,\delta^3(\boldsymbol{x}) .

We look for a static solution, so tϕ=0\partial_t\phi = 0 and ϕ=2ϕ\Box\phi = -\nabla^2\phi, giving

(2+m2)ϕ(x)=gδ3(x).(-\nabla^2 + m^2)\,\phi(\boldsymbol{x}) = g\,\delta^3(\boldsymbol{x}) .

Inserting the Fourier transform ϕ(x)=d3k(2π)3ϕ~(k)eikx\phi(\boldsymbol{x}) = \int \frac{d^3k}{(2\pi)^3} \tilde\phi(\boldsymbol{k}) e^{i\boldsymbol{k}\cdot\boldsymbol{x}} gives (k2+m2)ϕ~=g(|\boldsymbol{k}|^2 + m^2)\tilde\phi = g, that is, ϕ~(k)=g/(k2+m2)\tilde\phi(\boldsymbol{k}) = g/(|\boldsymbol{k}|^2+m^2). Carry out the inverse transform in polar coordinates: with r:=xr := |\boldsymbol{x}|, k:=kk := |\boldsymbol{k}| and x\boldsymbol{x} taken as polar axis,

ϕ(r)=g(2π)302π ⁣dφ0π ⁣dθsinθ0 ⁣dk k2eikrcosθk2+m2=g(2π)20dk k2k2+m22sin(kr)kr=g2π2r0dk ksin(kr)k2+m2.\begin{aligned} \phi(r) &= \frac{g}{(2\pi)^3}\int_0^{2\pi}\! d\varphi \int_0^\pi\! d\theta\,\sin\theta \int_0^\infty\! dk\ \frac{k^2\, e^{ikr\cos\theta}}{k^2+m^2} \\ &= \frac{g}{(2\pi)^2}\int_0^\infty dk\ \frac{k^2}{k^2+m^2}\cdot \frac{2\sin(kr)}{kr} = \frac{g}{2\pi^2 r}\int_0^\infty dk\ \frac{k\,\sin(kr)}{k^2+m^2} . \end{aligned}

In the second line we used 0πdθsinθeikrcosθ=[eikrcosθ/(ikr)]θ=0π=2sin(kr)/(kr)\int_0^\pi d\theta \sin\theta\, e^{ikr\cos\theta} = \big[ e^{ikr\cos\theta}/(-ikr) \big]_{\theta=0}^{\pi} = 2\sin(kr)/(kr).

Call the remaining integral II. The integrand is even in kk, so I=12ksin(kr)k2+m2dkI = \frac{1}{2}\int_{-\infty}^{\infty} \frac{k\sin(kr)}{k^2+m^2}dk. Moreover kcos(kr)/(k2+m2)k\cos(kr)/(k^2+m^2) is odd and integrates to 00, whence

keikrk2+m2dk=0+i2I.\int_{-\infty}^{\infty} \frac{k\,e^{ikr}}{k^2+m^2}\,dk = 0 + i\cdot 2I .

Evaluate the left-hand side by residues. Since r>0r > 0, close the contour with a semicircle in the upper half plane; the only pole of the integrand there is k=imk = im, with residue imei(im)r2im=emr2\dfrac{im\, e^{i(im)r}}{2im} = \dfrac{e^{-mr}}{2}. The left-hand side is thus 2πiemr2=iπemr2\pi i \cdot \frac{e^{-mr}}{2} = i\pi e^{-mr}, giving I=π2emrI = \frac{\pi}{2}e^{-mr}. Therefore

ϕ(r)=g2π2rπ2emr=g4πemrr.\phi(r) = \frac{g}{2\pi^2 r}\cdot\frac{\pi}{2}e^{-mr} = \frac{g}{4\pi}\,\frac{e^{-mr}}{r} .

This is the Yukawa potential. As m0m \to 0 it becomes the Coulomb-type long-range force g/(4πr)g/(4\pi r), while for m0m \ne 0 the factor emre^{-mr} makes it decay rapidly over a distance of order 1/m1/m.

Let us put in numbers. The range of the nuclear force is about 1.4 fm1.4\ \mathrm{fm}. Using c=197.3 MeVfm\hbar c = 197.3\ \mathrm{MeV\cdot fm},

mc1.4 fm=197.3 MeVfm1.4 fm1.4×102 MeV.m \approx \frac{\hbar c}{1.4\ \mathrm{fm}} = \frac{197.3\ \mathrm{MeV\cdot fm}}{1.4\ \mathrm{fm}} \approx 1.4 \times 10^2\ \mathrm{MeV} .

The pion, actually discovered in 1947, has a mass of about 135135140 MeV140\ \mathrm{MeV}, in good agreement with this estimate.

Nothing happens in a free field alone: scattering and decay require interactions. Drop requirement 3 (quadratic in ϕ\phi) and add potential terms:

L=12μϕμϕ12m2ϕ2n3λnn!ϕn.\mathcal{L} = \frac{1}{2}\partial_\mu\phi\,\partial^\mu\phi - \frac{1}{2}m^2\phi^2 - \sum_{n \ge 3}\frac{\lambda_n}{n!}\phi^n .

Remark 5.3Mass dimension classifies interactions

In §2 we established [L]=4[\mathcal{L}] = 4. The kinetic term μϕμϕ\partial_\mu\phi\,\partial^\mu\phi gives 2+2[ϕ]=42 + 2[\phi] = 4, that is,

[ϕ]=1.[\phi] = 1 .

Hence [λn]=4n[\lambda_n] = 4 - n, and the classification is as follows.

TermMass dimension of the couplingName
ϕ3\phi^3[λ3]=1[\lambda_3] = 1super-renormalizable
ϕ4\phi^4[λ4]=0[\lambda_4] = 0renormalizable
ϕ5,ϕ6,\phi^5, \phi^6, \dots[λn]<0[\lambda_n] < 0non-renormalizable (meaningful only as an effective theory)

The convention that “the first interaction to write down for a real scalar field in four-dimensional spacetime is ϕ4\phi^4” comes from this dimensional counting. Why a negative mass dimension is problematic is treated in the classification of renormalizability(Definition 5.4)[くりこみ理論入門] in introduction to renormalization. What matters here is that writing down L\mathcal{L} is not a mere formality: it is simultaneously an enumeration of the interactions the theory is allowed to have.


6. Noether’s theorem (field-theoretic version)

Section titled “6. Noether’s theorem (field-theoretic version)”

The theorem Emmy Noether proved in 1918 states that every continuous symmetry of the action corresponds to one conserved quantity. In the finite-dimensional version (symmetries and conservation laws) the conserved object was a single function of time; in field theory we obtain the stronger statement of a local conservation law, the continuity equation μjμ=0\partial_\mu j^\mu = 0. This asserts that charge cannot disappear here and reappear there, and it is consistent with relativistic causality.

flowchart TD
L["Lagrangian density L"] --> S["Action S = ∫ d⁴x L"]
S --> V["Variational principle δS = 0"]
V --> EL["Euler-Lagrange equations"]
S --> SY["Continuous symmetry of the action"]
SY --> N["Noether's theorem"]
EL --> N
N --> J["Conserved current ∂μ jμ = 0"]
J --> Q["Conserved charge Q = ∫ d³x j⁰"]
The logical structure of this chapter. A single function, the Lagrangian density, yields both the equations of motion and the conservation laws.

The definition of a symmetry requires care. Demanding invariance of the action alone would throw away the freedom of Proposition 4.5, so we adopt quasi-invariance, allowing L\mathcal{L} to shift by a total derivative.

Definition 6.1Continuous symmetry of the action (quasi-invariance)

Consider a smooth one-parameter family of transformations with real parameter ε\varepsilon,

xμxμ=xμ+εXμ(x)+O(ε2),ϕa(x)ϕa(x)=ϕa(x)+εFa(x)+O(ε2),x^\mu \longmapsto x'^\mu = x^\mu + \varepsilon\, X^\mu(x) + O(\varepsilon^2), \qquad \phi_a(x) \longmapsto \phi'_a(x') = \phi_a(x) + \varepsilon\, F_a(x) + O(\varepsilon^2) ,

where XμX^\mu is a smooth function of xx and FaF_a is a smooth function of ϕb(x)\phi_b(x), μϕb(x)\partial_\mu\phi_b(x) and xx.

If there exists a KμK^\mu (a smooth function of ϕb\phi_b, μϕb\partial_\mu\phi_b and xx) such that the identity

L(ϕ(x),ϕ(x)) det ⁣(xx)=L(ϕ(x),ϕ(x))+εμKμ+O(ε2)\mathcal{L}\big(\phi'(x'),\, \partial' \phi'(x')\big)\ \det\!\left( \frac{\partial x'}{\partial x} \right) = \mathcal{L}\big(\phi(x),\, \partial\phi(x)\big) + \varepsilon\, \partial_\mu K^\mu + O(\varepsilon^2)

holds for every field configuration, whether or not it satisfies the equations of motion, then the family is called a continuous symmetry of the action. The case Kμ=0K^\mu = 0 is called strict invariance, the general case quasi-invariance.

The factor det(x/x)\det(\partial x'/\partial x) on the left represents the change of the integration measure, d4x=det(x/x)d4xd^4x' = \det(\partial x'/\partial x)\, d^4x; including it makes the left-hand side “the action density of the transformed theory measured in the original coordinates”.

The essential point is that the condition holds identically, independently of the equations of motion. An identity valid only on solutions of the equations of motion yields no new information.

Theorem 6.2Noether's theorem

Let L=L(ϕa,μϕa)\mathcal{L} = \mathcal{L}(\phi_a, \partial_\mu\phi_a) have no explicit dependence on the coordinate xx, and let a continuous symmetry in the sense of Definition 6.1 be given. Put

πμa:=L(μϕa)\pi^\mu{}_a := \frac{\partial \mathcal{L}}{\partial(\partial_\mu \phi_a)}

and define the Noether current by

jμ:=πμa(FaXννϕa)+LXμKμ.j^\mu := \pi^\mu{}_a \big( F_a - X^\nu \partial_\nu \phi_a \big) + \mathcal{L}\, X^\mu - K^\mu .

Then on any solution of the Euler–Lagrange equations (Theorem 4.3),

μjμ=0.\partial_\mu j^\mu = 0 .
Proof(Theorem 6.2)

Step 1: write the symmetry condition to order ε\varepsilon.

The Jacobian matrix is xμ/xν=δνμ+ενXμ+O(ε2)\partial x'^\mu/\partial x^\nu = \delta^\mu_\nu + \varepsilon\,\partial_\nu X^\mu + O(\varepsilon^2), so its determinant is

det ⁣(xx)=1+εμXμ+O(ε2)\det\!\left( \frac{\partial x'}{\partial x} \right) = 1 + \varepsilon\, \partial_\mu X^\mu + O(\varepsilon^2)

(using det(I+εM)=1+εtrM+O(ε2)\det(I + \varepsilon M) = 1 + \varepsilon\,\mathrm{tr}\,M + O(\varepsilon^2)).

Next the transformation of derivatives. The inverse map gives xν/xμ=δμνεμXν+O(ε2)\partial x^\nu/\partial x'^\mu = \delta^\nu_\mu - \varepsilon\,\partial_\mu X^\nu + O(\varepsilon^2), so by the chain rule

μϕa(x)=xνxμν[ϕa(x)+εFa(x)]=μϕa+ε(μFa(μXν)νϕa)+O(ε2).\partial'_\mu \phi'_a(x') = \frac{\partial x^\nu}{\partial x'^\mu}\,\partial_\nu\big[\phi_a(x) + \varepsilon F_a(x)\big] = \partial_\mu \phi_a + \varepsilon\big( \partial_\mu F_a - (\partial_\mu X^\nu)\, \partial_\nu\phi_a \big) + O(\varepsilon^2) .

Since L\mathcal{L} has no explicit xx-dependence, the chain rule gives

L(ϕ(x),ϕ(x))=L+ε[LϕaFa+πμa(μFa(μXν)νϕa)]+O(ε2).\mathcal{L}\big(\phi'(x'), \partial'\phi'(x')\big) = \mathcal{L} + \varepsilon\left[ \frac{\partial\mathcal{L}}{\partial\phi_a}F_a + \pi^\mu{}_a\big( \partial_\mu F_a - (\partial_\mu X^\nu)\partial_\nu\phi_a \big) \right] + O(\varepsilon^2) .

Multiplying by the determinant and comparing with the right-hand side of Definition 6.1 at order ε\varepsilon, we obtain the identity

LμXμ+LϕaFa+πμaμFaπμa(μXν)νϕa=μKμ,\mathcal{L}\,\partial_\mu X^\mu + \frac{\partial\mathcal{L}}{\partial\phi_a}F_a + \pi^\mu{}_a\,\partial_\mu F_a - \pi^\mu{}_a\,(\partial_\mu X^\nu)\,\partial_\nu\phi_a = \partial_\mu K^\mu ,

which we shall call the symmetry identity. Note that no equation of motion has been used so far, and that it is an identity precisely because Definition 6.1 demanded validity for arbitrary field configurations.

Step 2: use the equations of motion to assemble the left-hand side into a divergence.

From now on let ϕ\phi be a solution of the Euler–Lagrange equations, that is, Lϕa=μπμa\dfrac{\partial\mathcal{L}}{\partial\phi_a} = \partial_\mu \pi^\mu{}_a.

Combine the second and third terms. By the Leibniz rule,

LϕaFa+πμaμFa=(μπμa)Fa+πμaμFa=μ(πμaFa).\frac{\partial\mathcal{L}}{\partial\phi_a}F_a + \pi^\mu{}_a\partial_\mu F_a = (\partial_\mu \pi^\mu{}_a) F_a + \pi^\mu{}_a \partial_\mu F_a = \partial_\mu\big( \pi^\mu{}_a F_a \big) .

Now the first term. Since L\mathcal{L} has no explicit xx-dependence, its total derivative with respect to xx is

μL=Lϕaμϕa+πνaνμϕa=(νπνa)μϕa+πνaνμϕa=ν(πνaμϕa)\partial_\mu \mathcal{L} = \frac{\partial\mathcal{L}}{\partial\phi_a}\partial_\mu\phi_a + \pi^\nu{}_a\, \partial_\nu\partial_\mu \phi_a = (\partial_\nu \pi^\nu{}_a)\,\partial_\mu\phi_a + \pi^\nu{}_a\,\partial_\nu\partial_\mu\phi_a = \partial_\nu\big( \pi^\nu{}_a\, \partial_\mu \phi_a \big)

(the second equality uses the equations of motion again). Therefore

LμXμ=μ(LXμ)XμμL=μ(LXμ)Xμν(πνaμϕa).\mathcal{L}\,\partial_\mu X^\mu = \partial_\mu\big(\mathcal{L}X^\mu\big) - X^\mu\,\partial_\mu\mathcal{L} = \partial_\mu\big(\mathcal{L}X^\mu\big) - X^\mu\, \partial_\nu\big( \pi^\nu{}_a \partial_\mu\phi_a \big) .

Split the fourth term likewise by the Leibniz rule:

πμa(μXν)νϕa=μ(Xνπμaνϕa)+Xνμ(πμaνϕa).-\pi^\mu{}_a (\partial_\mu X^\nu)\partial_\nu\phi_a = -\partial_\mu\big( X^\nu \pi^\mu{}_a \partial_\nu\phi_a \big) + X^\nu \partial_\mu\big( \pi^\mu{}_a \partial_\nu\phi_a \big) .

Step 3: the extra terms containing XX cancel.

The terms carrying XX that appeared in the last two equations are

Xμν(πνaμϕa)+Xνμ(πμaνϕa),- X^\mu \partial_\nu\big( \pi^\nu{}_a \partial_\mu\phi_a \big) + X^\nu \partial_\mu\big( \pi^\mu{}_a \partial_\nu\phi_a \big) ,

and relabelling the summation indices μν\mu \leftrightarrow \nu in the second term makes it the negative of the first, so the sum is 00.

Substituting all of this into the symmetry identity gives

μ[LXμ+πμaFaXνπμaνϕa]=μKμ,\partial_\mu\Big[ \mathcal{L}X^\mu + \pi^\mu{}_a F_a - X^\nu \pi^\mu{}_a \partial_\nu\phi_a \Big] = \partial_\mu K^\mu ,

that is, μ[πμa(FaXννϕa)+LXμKμ]=0\partial_\mu\big[ \pi^\mu{}_a(F_a - X^\nu\partial_\nu\phi_a) + \mathcal{L}X^\mu - K^\mu \big] = 0, which is the required identity.

From the local conservation law μjμ=0\partial_\mu j^\mu = 0 a globally conserved quantity follows. Note, however, that a physical assumption is needed: the current must fall off fast enough at infinity.

Corollary 6.3Conservation of the Noether charge

In the situation of Theorem 6.2, suppose the current jμj^\mu of a solution ϕ\phi satisfies, at each time tt, on the sphere SRS_R of radius RR,

limRSRjinidS=0\lim_{R\to\infty} \oint_{S_R} j^i\, n^i\, dS = 0

(nin^i the outward unit normal; for instance ji=o(R2)|j^i| = o(R^{-2}) suffices). Suppose also that j0j^0 is integrable over R3\mathbb{R}^3 at each time and that the time derivative may be interchanged with the integral. Then

Q(t):=R3d3x j0(t,x)Q(t) := \int_{\mathbb{R}^3} d^3x\ j^0(t, \boldsymbol{x})

is a constant independent of time.

Proof(Corollary 6.3)

Since μjμ=0j0+iji=0\partial_\mu j^\mu = \partial_0 j^0 + \partial_i j^i = 0,

dQdt=R3d3x 0j0=R3d3x iji=limRxRd3x j=limRSRjinidS=0.\frac{dQ}{dt} = \int_{\mathbb{R}^3} d^3x\ \partial_0 j^0 = -\int_{\mathbb{R}^3} d^3x\ \partial_i j^i = -\lim_{R\to\infty}\int_{|\boldsymbol{x}| \le R} d^3x\ \nabla\cdot \boldsymbol{j} = -\lim_{R\to\infty}\oint_{S_R} j^i n^i\, dS = 0 .

The first equality uses the interchange of derivative and integral, the third Gauss’s theorem, and the last the assumption.

Remark 6.4The current is ambiguous

Adding νΣνμ\partial_\nu \Sigma^{\nu\mu} to jμj^\mu, where Σνμ=Σμν\Sigma^{\nu\mu} = -\Sigma^{\mu\nu} is any antisymmetric tensor, preserves the conservation law, since μνΣνμ=0\partial_\mu\partial_\nu\Sigma^{\nu\mu} = 0 (a symmetric pair of derivatives contracted with an antisymmetric tensor). If moreover Σ\Sigma falls off fast enough at infinity, the charge QQ is unchanged as well. The Noether current, in other words, is not unique. We shall use this freedom in §6.5.

6.3. Spacetime translations and the energy–momentum tensor

Section titled “6.3. Spacetime translations and the energy–momentum tensor”

Definition 6.5Canonical energy–momentum tensor

When L\mathcal{L} has no explicit xx-dependence,

Tμν:=πμaνϕaδνμL,Tμν=ηνρTμρ=πμaνϕaημνLT^\mu{}_\nu := \pi^\mu{}_a\, \partial_\nu \phi_a - \delta^\mu_\nu\, \mathcal{L}, \qquad T^{\mu\nu} = \eta^{\nu\rho}\,T^\mu{}_\rho = \pi^\mu{}_a\, \partial^\nu\phi_a - \eta^{\mu\nu}\mathcal{L}

is called the canonical energy–momentum tensor.

Consider a spacetime translation xμ=xμ+εξμx'^\mu = x^\mu + \varepsilon\,\xi^\mu (with ξμ\xi^\mu a constant vector; we write ξ\xi to avoid confusion with the field-component index aa). The field is merely carried along, ϕa(x)=ϕa(x)\phi'_a(x') = \phi_a(x), so Xμ=ξμX^\mu = \xi^\mu and Fa=0F_a = 0. The Jacobian matrix is the identity, with det=1\det = 1, and since L\mathcal{L} has no explicit xx-dependence the left-hand side of Definition 6.1 is just L(ϕ(x),ϕ(x))\mathcal{L}(\phi(x),\partial\phi(x)), so we may take Kμ=0K^\mu = 0. Applying Theorem 6.2,

jμ=πμa(0ξννϕa)+Lξμ=ξν(πμaνϕaδνμL)=ξνTμν.j^\mu = \pi^\mu{}_a\big( 0 - \xi^\nu \partial_\nu\phi_a \big) + \mathcal{L}\,\xi^\mu = -\xi^\nu\Big( \pi^\mu{}_a \partial_\nu\phi_a - \delta^\mu_\nu \mathcal{L} \Big) = -\xi^\nu\, T^\mu{}_\nu .

Since ξν\xi^\nu is arbitrary, we obtain four conservation laws,

μTμν=0(ν=0,1,2,3).\partial_\mu T^\mu{}_\nu = 0 \qquad (\nu = 0,1,2,3) .

By Corollary 6.3 the corresponding conserved quantities are

Pν=d3x T0ν,P^\nu = \int d^3x\ T^{0\nu} ,

with P0P^0 the energy and PiP^i the momentum. Indeed T00=π0aϕ˙aLT^0{}_0 = \pi^0{}_a\dot\phi_a - \mathcal{L} has the same form as the finite-dimensional Legendre transform H=pq˙LH = p\dot q - L (definition of the Hamiltonian(Definition 3.6)[ハミルトン形式の力学]).

Example 6.6Energy and momentum of the Klein–Gordon field

For LKG\mathcal{L}_{\mathrm{KG}} we computed πμ=μϕ\pi^\mu = \partial^\mu\phi in Example 5.1, so

Tμν=μϕνϕημν[12αϕαϕ12m2ϕ2].T^{\mu\nu} = \partial^\mu\phi\,\partial^\nu\phi - \eta^{\mu\nu}\left[ \frac{1}{2}\partial_\alpha\phi\,\partial^\alpha\phi - \frac{1}{2}m^2\phi^2 \right] .

Noting αϕαϕ=ϕ˙2(ϕ)2\partial_\alpha\phi\,\partial^\alpha\phi = \dot\phi^2 - (\nabla\phi)^2, compute the μ=ν=0\mu = \nu = 0 component:

T00=ϕ˙2[12ϕ˙212(ϕ)212m2ϕ2]=12ϕ˙2+12(ϕ)2+12m2ϕ2.T^{00} = \dot\phi^2 - \left[ \frac{1}{2}\dot\phi^2 - \frac{1}{2}(\nabla\phi)^2 - \frac{1}{2}m^2\phi^2 \right] = \frac{1}{2}\dot\phi^2 + \frac{1}{2}(\nabla\phi)^2 + \frac{1}{2}m^2\phi^2 .

All three terms are nonnegative, so the energy density is bounded below. This is why we required a positive coefficient for the kinetic term and m2>0m^2 > 0. The momentum density is

T0i=0ϕiϕ=ϕ˙ϕxi,P=d3x ϕ˙ϕ.T^{0i} = \partial^0\phi\,\partial^i\phi = -\dot\phi\, \frac{\partial\phi}{\partial x^i}, \qquad \boldsymbol{P} = -\int d^3x\ \dot\phi\, \nabla\phi .

Let us check the sign. For a wave travelling in the +x+x direction, ϕ=Acos(ωtkx)\phi = A\cos(\omega t - k x) with k>0k > 0, we have ϕ˙=Aωsin(ωtkx)\dot\phi = -A\omega\sin(\omega t - kx) and ϕ/x=Aksin(ωtkx)\partial\phi/\partial x = A k \sin(\omega t - kx), so ϕ˙xϕ=A2ωksin2(ωtkx)0-\dot\phi\,\partial_x\phi = A^2\omega k \sin^2(\omega t - kx) \ge 0: the momentum indeed points in the +x+x direction.

Note also that TμνT^{\mu\nu} is symmetric (both μϕνϕ\partial^\mu\phi\,\partial^\nu\phi and ημν\eta^{\mu\nu} are). This will matter in §6.5.

6.4. Internal symmetry and conserved charge

Section titled “6.4. Internal symmetry and conserved charge”

Next come transformations that leave spacetime alone and rotate only within field space. For a complex scalar field ϕ\phi (with its complex conjugate ϕˉ\bar\phi regarded as an independent variable), consider

L=μϕˉμϕm2ϕˉϕ.\mathcal{L} = \partial_\mu \bar\phi\, \partial^\mu\phi - m^2 \bar\phi \phi .

One may equally think of this as two real fields ϕ1,ϕ2\phi_1, \phi_2 packaged as ϕ=(ϕ1+iϕ2)/2\phi = (\phi_1 + i\phi_2)/\sqrt{2}.

Example 6.7Global U(1) symmetry and its conserved current

Consider the transformation

ϕ(x)=eiεϕ(x)=ϕiεϕ+O(ε2),ϕˉ(x)=e+iεϕˉ(x)=ϕˉ+iεϕˉ+O(ε2).\phi'(x) = e^{-i\varepsilon}\phi(x) = \phi - i\varepsilon\phi + O(\varepsilon^2), \qquad \bar\phi'(x) = e^{+i\varepsilon}\bar\phi(x) = \bar\phi + i\varepsilon\bar\phi + O(\varepsilon^2) .

Since ε\varepsilon is a constant independent of spacetime, this is a “global” symmetry. In the notation of Definition 6.1, Xμ=0X^\mu = 0, Fϕ=iϕF_\phi = -i\phi and Fϕˉ=+iϕˉF_{\bar\phi} = +i\bar\phi. Because L\mathcal{L} is built solely out of the combinations ϕˉϕ\bar\phi\phi and μϕˉμϕ\partial_\mu\bar\phi\,\partial^\mu\phi, the phase factors cancel as e+iεeiε=1e^{+i\varepsilon}e^{-i\varepsilon} = 1 and L\mathcal{L} is strictly invariant. We may thus take Kμ=0K^\mu = 0.

The conjugate momenta are

πμϕ=L(μϕ)=μϕˉ,πμϕˉ=L(μϕˉ)=μϕ,\pi^\mu{}_{\phi} = \frac{\partial\mathcal{L}}{\partial(\partial_\mu\phi)} = \partial^\mu\bar\phi, \qquad \pi^\mu{}_{\bar\phi} = \frac{\partial\mathcal{L}}{\partial(\partial_\mu\bar\phi)} = \partial^\mu\phi ,

so Theorem 6.2 gives

jμ=πμϕFϕ+πμϕˉFϕˉ=μϕˉ(iϕ)+μϕ(iϕˉ)=i(ϕˉμϕϕμϕˉ).j^\mu = \pi^\mu{}_\phi F_\phi + \pi^\mu{}_{\bar\phi} F_{\bar\phi} = \partial^\mu\bar\phi\,(-i\phi) + \partial^\mu\phi\,(i\bar\phi) = i\big( \bar\phi\,\partial^\mu\phi - \phi\,\partial^\mu\bar\phi \big) .

Let us verify the conservation directly. The Euler–Lagrange equations of L\mathcal{L} are (+m2)ϕ=0(\Box + m^2)\phi = 0 from varying ϕˉ\bar\phi, and (+m2)ϕˉ=0(\Box+m^2)\bar\phi = 0 from varying ϕ\phi. Hence

μjμ=i(μϕˉμϕ+ϕˉϕμϕμϕˉϕϕˉ)=i(ϕˉ(m2ϕ)ϕ(m2ϕˉ))=0.\partial_\mu j^\mu = i\big( \partial_\mu\bar\phi\,\partial^\mu\phi + \bar\phi\,\Box\phi - \partial_\mu\phi\,\partial^\mu\bar\phi - \phi\,\Box\bar\phi \big) = i\big( \bar\phi(-m^2\phi) - \phi(-m^2\bar\phi) \big) = 0 .

The first and third terms cancel, and the equations of motion kill what remains. The conserved charge is

Q=d3x j0=id3x (ϕˉϕ˙ϕϕˉ˙).Q = \int d^3x\ j^0 = i\int d^3x\ \big( \bar\phi\,\dot\phi - \phi\,\dot{\bar\phi} \big) .

Making this U(1)U(1) local (letting ε\varepsilon be a function of xx) forces a coupling to the electromagnetic field to appear, and QQ becomes the electric charge itself (gauge theory and spontaneous symmetry breaking).

Remark 6.8j⁰ is not a probability density

The quantity j0=i(ϕˉϕ˙ϕϕˉ˙)j^0 = i(\bar\phi\dot\phi - \phi\dot{\bar\phi}) above has no definite sign. Indeed, for a negative-frequency solution ϕ=e+iωt\phi = e^{+i\omega t} one finds j0<0j^0 < 0. The early attempts to read the Klein–Gordon equation as a relativistic one-particle Schrödinger equation broke down precisely because this j0j^0 could not be interpreted as a probability density (see the Klein–Gordon current and its lack of positivity(Proposition 3.2)[Why We Need Quantum Field Theory] in why we need quantum field theory).

Within classical field theory this causes no trouble at all. QQ is a charge, not a probability, and it is perfectly natural for a charge to take either sign. After quantization one finds that the eigenvalues of QQ count the number of particles minus the number of antiparticles.

6.5. Lorentz transformations and angular momentum

Section titled “6.5. Lorentz transformations and angular momentum”

Proposition 6.9Angular momentum tensor of a scalar field

Let L\mathcal{L} be the Lagrangian density of a real scalar field and a Lorentz scalar. As the Noether conserved quantity associated with the infinitesimal Lorentz transformation xμ=xμ+εωμνxνx'^\mu = x^\mu + \varepsilon\,\omega^\mu{}_\nu x^\nu (with ωμν=ωνμ\omega_{\mu\nu} = -\omega_{\nu\mu}) together with ϕ(x)=ϕ(x)\phi'(x') = \phi(x), the quantity

Mμαρ:=xαTμρxρTμαM^{\mu\alpha\rho} := x^\alpha\, T^{\mu\rho} - x^\rho\, T^{\mu\alpha}

satisfies μMμαρ=0\partial_\mu M^{\mu\alpha\rho} = 0. Moreover, it follows from this that the canonical energy–momentum tensor is symmetric, Tαρ=TραT^{\alpha\rho} = T^{\rho\alpha}.

Proof(Proposition 6.9)

In the notation of Definition 6.1, Xμ=ωμνxνX^\mu = \omega^\mu{}_\nu x^\nu and F=0F = 0. The Jacobian is 1+εμXμ=1+εωμμ1 + \varepsilon\,\partial_\mu X^\mu = 1 + \varepsilon\,\omega^\mu{}_\mu, but ωμμ=ημαωαμ\omega^\mu{}_\mu = \eta^{\mu\alpha}\omega_{\alpha\mu} is the contraction of the symmetric tensor η\eta with the antisymmetric tensor ω\omega and therefore vanishes. Since L\mathcal{L} is a scalar, L(ϕ(x),ϕ(x))=L(ϕ(x),ϕ(x))\mathcal{L}(\phi'(x'),\partial'\phi'(x')) = \mathcal{L}(\phi(x),\partial\phi(x)), and we may take Kμ=0K^\mu = 0. Hence

jμ=TμνXν=Tμνωνρxρ=Tμαωαρxρj^\mu = -T^\mu{}_\nu X^\nu = -T^\mu{}_\nu\, \omega^\nu{}_\rho\, x^\rho = -T^{\mu\alpha}\,\omega_{\alpha\rho}\, x^\rho

(using Tμνηνα=TμαT^\mu{}_\nu \eta^{\nu\alpha} = T^{\mu\alpha}). Antisymmetrizing explicitly with the help of the antisymmetry of ω\omega,

jμ=12ωαρ(xρTμαxαTμρ)=12ωαρMμαρ.j^\mu = -\frac{1}{2}\omega_{\alpha\rho}\big( x^\rho T^{\mu\alpha} - x^\alpha T^{\mu\rho} \big) = \frac{1}{2}\,\omega_{\alpha\rho}\, M^{\mu\alpha\rho} .

Since ωαρ\omega_{\alpha\rho} is an arbitrary constant antisymmetric matrix, the conclusion μjμ=0\partial_\mu j^\mu = 0 of Theorem 6.2 means μMμαρ=0\partial_\mu M^{\mu\alpha\rho} = 0 for each pair (α,ρ)(\alpha,\rho).

Now the last assertion. From μxα=δμα\partial_\mu x^\alpha = \delta^\alpha_\mu and the Leibniz rule,

μMμαρ=TαρTρα+xαμTμρxρμTμα.\partial_\mu M^{\mu\alpha\rho} = T^{\alpha\rho} - T^{\rho\alpha} + x^\alpha\,\partial_\mu T^{\mu\rho} - x^\rho\,\partial_\mu T^{\mu\alpha} .

The last two terms vanish by μTμν=0\partial_\mu T^{\mu\nu} = 0, shown in §6.3, so μMμαρ=TαρTρα=0\partial_\mu M^{\mu\alpha\rho} = T^{\alpha\rho} - T^{\rho\alpha} = 0.

Forming Ji=12ϵijkd3xM0jkJ^i = \frac{1}{2}\epsilon^{ijk}\int d^3x\, M^{0jk} from the spatial components of MμαρM^{\mu\alpha\rho} gives the ordinary angular momentum. That this is the field-theoretic version of r×p\boldsymbol{r} \times \boldsymbol{p} can be read off from M0jk=xjT0kxkT0jM^{0jk} = x^j T^{0k} - x^k T^{0j}.

Remark 6.10For fields with spin the canonical tensor is not symmetric

The argument of Proposition 6.9 is specific to scalar fields. For vector or spinor fields the field components mix under Lorentz transformations, so Fa0F_a \ne 0, and the current acquires a “spin part” SμαρS^{\mu\alpha\rho}, giving Mμαρ=xαTμρxρTμα+SμαρM^{\mu\alpha\rho} = x^\alpha T^{\mu\rho} - x^\rho T^{\mu\alpha} + S^{\mu\alpha\rho}. What μM=0\partial_\mu M = 0 then requires is TαρTρα=μSμαρT^{\alpha\rho} - T^{\rho\alpha} = -\partial_\mu S^{\mu\alpha\rho}, not the symmetry of TT.

The canonical tensor of the electromagnetic field is in fact neither symmetric nor gauge invariant. Using the freedom of Remark 6.4 to absorb the spin part and thereby produce a symmetric, gauge-invariant tensor is the Belinfante–Rosenfeld improvement, and the result coincides with the TμνT^{\mu\nu} that appears as the source of gravity in general relativity. For details see, for example, Weinberg’s textbook.


7. Toward the Hamiltonian formalism: preparing for quantization

Section titled “7. Toward the Hamiltonian formalism: preparing for quantization”

To proceed to canonical quantization we must pass from the Lagrangian to the Hamiltonian formalism. This is the stage at which time is singled out, so Lorentz covariance ceases to be manifest (it is still present in the theory).

Definition 7.1Conjugate momentum density and Hamiltonian density

Define the momentum density conjugate to the field ϕa\phi_a by

πa(x):=Lϕ˙a=π0a(x).\pi_a(x) := \frac{\partial \mathcal{L}}{\partial \dot\phi_a} = \pi^0{}_a(x) .

Assuming that ϕ˙a\dot\phi_a can be solved for in terms of πa\pi_a (that is, that the Legendre transform exists), define the Hamiltonian density by

H:=πaϕ˙aL\mathcal{H} := \pi_a \dot\phi_a - \mathcal{L}

and call H:=d3x HH := \int d^3x\ \mathcal{H} the Hamiltonian.

Comparing with Definition 6.5 shows that H=T00=T00\mathcal{H} = T^0{}_0 = T^{00} and H=P0H = P^0. The Hamiltonian is thus nothing but the time component of the Noether charge of spacetime translation symmetry. The statement that energy is conserved because of time-translation symmetry holds in field theory exactly as it does for finitely many degrees of freedom.

For the Klein–Gordon field π=ϕ˙\pi = \dot\phi, so

H=ϕ˙2LKG=12π2+12(ϕ)2+12m2ϕ2,\mathcal{H} = \dot\phi^2 - \mathcal{L}_{\mathrm{KG}} = \frac{1}{2}\pi^2 + \frac{1}{2}(\nabla\phi)^2 + \frac{1}{2}m^2\phi^2 ,

in agreement with the T00T^{00} of Example 6.6. The ϕ\nabla\phi term penalizes differences between the values at neighbouring points; it can be read as the elastic energy of the springs of Example 3.1, surviving intact.

Equal-time Poisson brackets are defined by reading the finite-dimensional {qi,pj}=δij\{q_i, p_j\} = \delta_{ij} with a continuous index:

{ϕa(t,x),πb(t,y)}=δabδ3(xy),{ϕa(t,x),ϕb(t,y)}={πa(t,x),πb(t,y)}=0.\{\phi_a(t,\boldsymbol{x}),\, \pi_b(t,\boldsymbol{y})\} = \delta_{ab}\,\delta^3(\boldsymbol{x} - \boldsymbol{y}), \qquad \{\phi_a(t,\boldsymbol{x}),\, \phi_b(t,\boldsymbol{y})\} = \{\pi_a(t,\boldsymbol{x}),\, \pi_b(t,\boldsymbol{y})\} = 0 .

The only change is that the Kronecker delta has become a Dirac delta function (Hamiltonian mechanics, canonical transformations and Poisson brackets).

In the next chapter we replace this bracket by a commutator via { , }i[ , ]\{\ ,\ \} \to -i[\ ,\ ]. That is, imposing

[ϕ^(t,x),π^(t,y)]=iδ3(xy)[\hat\phi(t,\boldsymbol{x}),\, \hat\pi(t,\boldsymbol{y})] = i\,\delta^3(\boldsymbol{x}-\boldsymbol{y})

is canonical quantization of the scalar field (canonical quantization of the real scalar field(Definition 3.1)[スカラー場の正準量子化]). The alternative route, which keeps time on the same footing as space and uses the action S[ϕ]S[\phi] directly, is path-integral quantization. Whichever road one takes, the starting point is the L\mathcal{L} built in this chapter, and the symmetries encoded in it are inherited by the quantum theory — or else broken by quantum effects, which is what an anomaly is.


Exercise 8.1Hard

Consider the Lagrangian density of the electromagnetic field,

L=14FμνFμνAμJμ,Fμν:=μAννAμ,\mathcal{L} = -\frac{1}{4}F_{\mu\nu}F^{\mu\nu} - A_\mu J^\mu, \qquad F_{\mu\nu} := \partial_\mu A_\nu - \partial_\nu A_\mu ,

where JμJ^\mu is a given external current independent of AμA_\mu. The fundamental fields are the four components AνA_\nu.

  1. Apply Theorem 4.3 and derive the equations of motion.
  2. Verify that the ν=0\nu = 0 component of the resulting equation is Gauss’s law.
Solution

1. First, since FμνF_{\mu\nu} contains only derivatives of AA, the derivative with respect to the field itself is

LAν=Jν.\frac{\partial \mathcal{L}}{\partial A_\nu} = -J^\nu .

Next the derivative with respect to μAν\partial_\mu A_\nu. From Fαβ=αAββAαF_{\alpha\beta} = \partial_\alpha A_\beta - \partial_\beta A_\alpha,

Fαβ(μAν)=δαμδβνδανδβμ.\frac{\partial F_{\alpha\beta}}{\partial(\partial_\mu A_\nu)} = \delta^\mu_\alpha \delta^\nu_\beta - \delta^\nu_\alpha \delta^\mu_\beta .

The chain rule then gives

(μAν)(FαβFαβ)=2Fαβ(δαμδβνδανδβμ)=2(FμνFνμ)=4Fμν\frac{\partial}{\partial(\partial_\mu A_\nu)}\big( F_{\alpha\beta}F^{\alpha\beta} \big) = 2F^{\alpha\beta}\big( \delta^\mu_\alpha\delta^\nu_\beta - \delta^\nu_\alpha\delta^\mu_\beta \big) = 2\big( F^{\mu\nu} - F^{\nu\mu} \big) = 4F^{\mu\nu}

(the last step uses the antisymmetry of FF). Hence

πμ(ν)=L(μAν)=144Fμν=Fμν.\pi^\mu{}_{(\nu)} = \frac{\partial\mathcal{L}}{\partial(\partial_\mu A_\nu)} = -\frac{1}{4}\cdot 4 F^{\mu\nu} = -F^{\mu\nu} .

Substituting into the Euler–Lagrange equation μ(L/(μAν))L/Aν=0\partial_\mu\big(\partial\mathcal{L}/\partial(\partial_\mu A_\nu)\big) - \partial\mathcal{L}/\partial A_\nu = 0,

μFμν+Jν=0μFμν=Jν.-\partial_\mu F^{\mu\nu} + J^\nu = 0 \qquad\Longleftrightarrow\qquad \partial_\mu F^{\mu\nu} = J^\nu .

These are the two inhomogeneous Maxwell equations (Gauss’s law and the Ampère–Maxwell law). The two homogeneous ones (absence of magnetic flux and Faraday’s law) follow as identities from the definition F=dAF = dA, so they do not come out of the variational principle.

2. With Aμ=(φ,A)A^\mu = (\varphi, \boldsymbol{A}) we have Aμ=(φ,A)A_\mu = (\varphi, -\boldsymbol{A}), and F0i=0AiiA0=tAiiφ=EiF_{0i} = \partial_0 A_i - \partial_i A_0 = -\partial_t A^i - \partial_i\varphi = E^i (since E=φtA\boldsymbol{E} = -\nabla\varphi - \partial_t\boldsymbol{A}). Raising indices, F0i=η00ηiiF0i=EiF^{0i} = \eta^{00}\eta^{ii}F_{0i} = -E^i, that is, Fi0=EiF^{i0} = E^i. Noting F00=0F^{00} = 0 by antisymmetry, the ν=0\nu = 0 component reads

μFμ0=0F00+iFi0=E=J0=ρ,\partial_\mu F^{\mu 0} = \partial_0 F^{00} + \partial_i F^{i0} = \nabla\cdot\boldsymbol{E} = J^0 = \rho ,

which is Gauss’s law.

Exercise 8.2Standard

Consider a massless real scalar field, L0=12μϕμϕ\mathcal{L}_0 = \frac{1}{2}\partial_\mu\phi\,\partial^\mu\phi.

  1. Verify that the shift ϕ(x)=ϕ(x)+εc\phi'(x) = \phi(x) + \varepsilon c by a constant cc (with the spacetime coordinates untouched) is a symmetry in the sense of Definition 6.1, and find the Noether current. Confirm directly from the equation of motion that it is conserved.
  2. Show that adding a mass term 12m2ϕ2-\frac{1}{2}m^2\phi^2 (with m0m \ne 0) makes this transformation no longer a symmetry.
Solution

1. Here Xμ=0X^\mu = 0 and F=cF = c. The density L0\mathcal{L}_0 contains no ϕ\phi itself, only μϕ\partial_\mu\phi, and μ(ϕ+εc)=μϕ\partial_\mu(\phi + \varepsilon c) = \partial_\mu\phi, so L0\mathcal{L}_0 is strictly invariant and we may take Kμ=0K^\mu = 0. Since the coordinates are not moved, the Jacobian is 11.

With πμ=μϕ\pi^\mu = \partial^\mu\phi, Theorem 6.2 gives

jμ=πμF=cμϕ.j^\mu = \pi^\mu F = c\,\partial^\mu\phi .

Conservation is checked directly: μjμ=cμμϕ=cϕ=0\partial_\mu j^\mu = c\,\partial_\mu\partial^\mu\phi = c\,\Box\phi = 0, the last equality being the Klein–Gordon equation with m=0m = 0. The corresponding charge is Q=cd3xϕ˙Q = c\int d^3x\,\dot\phi, which states that the “average velocity” of the field does not change.

2. Applying the transformation to L=L012m2ϕ2\mathcal{L} = \mathcal{L}_0 - \frac{1}{2}m^2\phi^2,

LL=12m2[(ϕ+εc)2ϕ2]=εm2cϕ+O(ε2).\mathcal{L}' - \mathcal{L} = -\frac{1}{2}m^2\big[ (\phi + \varepsilon c)^2 - \phi^2 \big] = -\varepsilon\, m^2 c\, \phi + O(\varepsilon^2) .

The term at order ε\varepsilon is m2cϕ-m^2 c\,\phi, and this can never equal the total derivative μKμ\partial_\mu K^\mu of any Kμ(ϕ,x)K^\mu(\phi, x). Indeed, the chain rule decomposes the total derivative of KμK^\mu as

μKμ=Kμϕμϕ+(μKμ)expl,\partial_\mu K^\mu = \frac{\partial K^\mu}{\partial \phi}\,\partial_\mu\phi + \big( \partial_\mu K^\mu \big)_{\mathrm{expl}} ,

the second term differentiating only the explicit xx-dependence at fixed ϕ\phi. Suppose the identity m2cϕ=μKμ-m^2c\,\phi = \partial_\mu K^\mu held for every field configuration. The left-hand side contains no μϕ\partial_\mu\phi, whereas the first term on the right is proportional to μϕ\partial_\mu\phi; hence Kμ/ϕ=0\partial K^\mu/\partial\phi = 0, so KμK^\mu must be a function of xx alone. But then the right-hand side does not depend on ϕ\phi while the left-hand side is proportional to ϕ\phi, so m2c=0m^2 c = 0, that is, m=0m = 0, on pain of contradiction.

Thus for m0m \ne 0 the shift symmetry is broken. Incidentally, the fact that a massless scalar field enjoys a shift symmetry is deeply related to the reason why Nambu–Goldstone bosons cannot have a mass.

Exercise 8.3Standard

For the energy–momentum tensor of the Klein–Gordon field,

Tμν=μϕνϕημν[12αϕαϕ12m2ϕ2],T^{\mu\nu} = \partial^\mu\phi\,\partial^\nu\phi - \eta^{\mu\nu}\left[ \frac{1}{2}\partial_\alpha\phi\,\partial^\alpha\phi - \frac{1}{2}m^2\phi^2 \right] ,

verify μTμν=0\partial_\mu T^{\mu\nu} = 0 by direct computation, without going through Noether’s theorem (take ϕ\phi to be a solution of the Klein–Gordon equation).

Solution

Differentiate each term by the Leibniz rule:

μ(μϕνϕ)=(ϕ)νϕ+μϕμνϕ.\partial_\mu\big( \partial^\mu\phi\,\partial^\nu\phi \big) = (\Box\phi)\,\partial^\nu\phi + \partial^\mu\phi\,\partial_\mu\partial^\nu\phi .

For the second term, ημνμ=ν\eta^{\mu\nu}\partial_\mu = \partial^\nu, so

μ(ημν[12αϕαϕ12m2ϕ2])=ν[12αϕαϕ]+12m2ν(ϕ2)=αϕναϕ+m2ϕνϕ.-\partial_\mu\left( \eta^{\mu\nu}\left[ \frac{1}{2}\partial_\alpha\phi\,\partial^\alpha\phi - \frac{1}{2}m^2\phi^2 \right] \right) = -\partial^\nu\left[ \frac{1}{2}\partial_\alpha\phi\,\partial^\alpha\phi \right] + \frac{1}{2}m^2\,\partial^\nu(\phi^2) = -\partial_\alpha\phi\,\partial^\nu\partial^\alpha\phi + m^2\phi\,\partial^\nu\phi .

Now μϕμνϕ\partial^\mu\phi\,\partial_\mu\partial^\nu\phi and αϕναϕ\partial_\alpha\phi\,\partial^\nu\partial^\alpha\phi are the same object: relabel the summation index μα\mu \to \alpha and use that partial derivatives commute, since ϕ\phi is of class C2C^2. These two therefore cancel, leaving

μTμν=(ϕ)νϕ+m2ϕνϕ=(ϕ+m2ϕ)νϕ=0.\partial_\mu T^{\mu\nu} = (\Box\phi)\,\partial^\nu\phi + m^2\phi\,\partial^\nu\phi = \big( \Box\phi + m^2\phi \big)\,\partial^\nu\phi = 0 .

The final equality is the Klein–Gordon equation. Note that the equation of motion was used only in that last line; everything before it was an identity.

Exercise 8.4Easy

Work in natural units in dd-dimensional spacetime (one time dimension plus d1d-1 spatial dimensions).

  1. Find the mass dimension [ϕ][\phi] of a real scalar field ϕ\phi.
  2. Express the mass dimension [λn][\lambda_n] of the coupling in the interaction term λnn!ϕn-\dfrac{\lambda_n}{n!}\phi^n in terms of dd and nn, and describe what happens for d=4d = 4 and for d=2d = 2.
Solution

1. The action S=ddxLS = \int d^dx\,\mathcal{L} is dimensionless and [ddx]=d[d^dx] = -d, so [L]=d[\mathcal{L}] = d. The kinetic term μϕμϕ\partial_\mu\phi\,\partial^\mu\phi has dimension 2[]+2[ϕ]=2+2[ϕ]2[\partial] + 2[\phi] = 2 + 2[\phi], hence

2+2[ϕ]=d    [ϕ]=d22.2 + 2[\phi] = d \iff [\phi] = \frac{d-2}{2} .

2. From [λn]+n[ϕ]=d[\lambda_n] + n[\phi] = d,

[λn]=dnd22.[\lambda_n] = d - n\,\frac{d-2}{2} .

For d=4d = 4 this is [λn]=4n[\lambda_n] = 4 - n: the value is 11 for n=3n = 3 (a coupling with the dimension of a mass), 00 for n=4n = 4 (dimensionless), and negative for n5n \ge 5. Couplings of negative dimension grow more important at high energies and are non-renormalizable (Remark 5.3).

For d=2d = 2 we get [ϕ]=0[\phi] = 0: the field itself is dimensionless. Consequently [λn]=2[\lambda_n] = 2 for every nn, and an arbitrary function V(ϕ)V(\phi) of ϕ\phi is allowed as a potential. This is why theories containing non-polynomial functions of ϕ\phi, such as the sine-Gordon model with V(ϕ)cos(βϕ)V(\phi) \propto \cos(\beta\phi), are studied in two dimensions.


  • M. E. Peskin and D. V. Schroeder, An Introduction to Quantum Field Theory, Westview Press, 1995 — the first half of Chapter 2, “The Klein-Gordon Field”, gives a concise account of almost exactly the path taken here (Lagrangian density, Euler–Lagrange equations, Noether’s theorem, the Klein–Gordon field).
  • H. Goldstein, C. Poole and J. Safko, Classical Mechanics, 3rd edition, Addison-Wesley, 2002 — Chapter 13. Treats the continuum limit of coupled oscillators and the Lagrangian and Hamiltonian formalisms for fields independently of quantum theory. Example 3.1 follows the discussion of that chapter.
  • S. Weinberg, The Quantum Theory of Fields, Volume I: Foundations, Cambridge University Press, 1995 — Chapter 7, “The Canonical Formalism”. Goes further than this article on Noether’s theorem and the improvement (Belinfante–Rosenfeld) of the energy–momentum tensor.
  • T. Kugo, Gēji-ba no Ryōshiron I, Baifūkan, 1989 (in Japanese) — Chapter 1. A standard reference giving a careful treatment of classical field theory and Noether’s theorem.
  • E. Noether, “Invariante Variationsprobleme”, Nachrichten von der Gesellschaft der Wissenschaften zu Göttingen, Mathematisch-Physikalische Klasse (1918), 235–257. English translation: M. A. Tavel, “Invariant Variation Problems”, arXiv:physics/0503066.
  • H. Yukawa, “On the Interaction of Elementary Particles. I”, Proceedings of the Physico-Mathematical Society of Japan 17 (1935), 48–57. The original source of the calculation in Example 5.2.

Appendix: Rewriting in terms of functional derivatives

Section titled “Appendix: Rewriting in terms of functional derivatives”

Definition of the functional derivative. In the main text we treated variations as derivatives with respect to ε\varepsilon, but the literature on field theory makes wide use of functional-derivative notation. Given a functional S[ϕ]S[\phi], if

ddε0S[ϕ+εη]=Ωd4x δSδϕa(x)ηa(x)\left.\frac{d}{d\varepsilon}\right|_{0} S[\phi + \varepsilon\eta] = \int_\Omega d^4x\ \frac{\delta S}{\delta \phi_a(x)}\,\eta_a(x)

holds for every smooth ηa\eta_a with compact support contained in Ω\Omega, the coefficient δSδϕa(x)\dfrac{\delta S}{\delta\phi_a(x)} in the integrand is called the functional derivative. By Lemma 4.2, this coefficient is uniquely determined (within continuous functions).

Rewriting the Euler–Lagrange equations. In this notation, the expression obtained in the proof of Theorem 4.3 reads

δSδϕa(x)=Lϕaμ ⁣(L(μϕa)),\frac{\delta S}{\delta \phi_a(x)} = \frac{\partial\mathcal{L}}{\partial\phi_a} - \partial_\mu\!\left( \frac{\partial\mathcal{L}}{\partial(\partial_\mu\phi_a)} \right) ,

so the equations of motion become the single line δSδϕa(x)=0\dfrac{\delta S}{\delta\phi_a(x)} = 0. This is the finite-dimensional Sqi=0\dfrac{\partial S}{\partial q_i} = 0 with the index ii replaced by the continuous variable xx, corresponding to the last row of the dictionary in §3.

A basic formula. The definition immediately yields

δϕa(x)δϕb(y)=δabδ4(xy).\frac{\delta \phi_a(x)}{\delta \phi_b(y)} = \delta_{ab}\,\delta^4(x - y) .

Indeed, applying the defining relation above to S[ϕ]=ϕa(x)S[\phi] = \phi_a(x) (the functional that returns the field value at a fixed xx), the left-hand side is ηa(x)\eta_a(x) while the right-hand side is d4yδϕa(x)δϕb(y)ηb(y)\int d^4y\,\frac{\delta\phi_a(x)}{\delta\phi_b(y)}\eta_b(y); for these to agree for every η\eta, the kernel must be a delta function. This formula is used repeatedly in deriving the Schwinger–Dyson equations from the path integral.

Report an error in this article ・Operated by: Mugen Giken LLCPricingTermsLegal notice

© 2026 夢現技研合同会社 ・Feeding the text to an LLM is welcome. Code samples are MIT licensed.