Skip to content

Why We Need Quantum Field Theory: One-Particle Quantum Mechanics Is Incompatible with Relativity

Prerequisite:The Birth of Quantum Mechanics: How Black-Body Radiation, the Photoelectric Effect and Matter Waves Broke the Classical PictureThe Principles of Special Relativity: From Galilean Relativity to Einstein's Two PostulatesFoundations of Newtonian Mechanics: From the Three Laws to Momentum and Energy Conservation

Raw
  • Nonrelativistic quantum mechanics is built on the Hilbert space L2(R3N)L^2(\mathbb{R}^{3N}) with the particle number NN held fixed. The framework contains no place in which to write the creation or annihilation of a particle.
  • Substituting operators naively into the relativistic dispersion relation E2=p2c2+m2c4E^2 = \boldsymbol{p}^2c^2 + m^2c^4 produces the Klein–Gordon equation, but negative-energy solutions inevitably appear and the time component of the conserved current can be negative. It cannot be read as a probability density.
  • The Dirac equation restores positivity of the probability density, yet the negative-energy solutions remain. The “Dirac sea” is in effect a many-body theory carrying infinitely many particles, and it is unusable for bosons.
  • Causality is the more serious failure. One-particle time evolution generated by H=2+m2H = \sqrt{-\nabla^2 + m^2} destroys compact support instantaneously, and the amplitude to propagate to a spacelike-separated point does not vanish either.
  • There is exactly one way out: quantize the field ϕ(x)\phi(x). A free field becomes a collection of independent harmonic oscillators, one per mode, and their excitation quanta appear as particles obeying the dispersion relation E=p2+m2E = \sqrt{\boldsymbol{p}^2 + m^2}.
  • Causality is implemented not as “the amplitude vanishes” but as “the commutator vanishes at spacelike separation”. This cancellation requires antiparticles, so the existence of antiparticles is a consequence of relativity together with quantum mechanics.

1. Motivation: where does the photon come from?

Section titled “1. Motivation: where does the photon come from?”

Let us begin with a naive question. When an excited atom drops to its ground state, one photon is emitted. Before the emission no photon exists; afterwards one does. In the framework of the Schrödinger equation and the wave function, however, a state is a function ψ(x1,,xN,t)\psi(\boldsymbol{x}_1,\ldots,\boldsymbol{x}_N,t) of the coordinates of NN particles, and NN is fixed the moment the equation is written down. There is nowhere at all to write a process in which the number of particles increases by one.

The same thing happens in the annihilation of an electron and a positron, ee+γγe^- e^+ \to \gamma\gamma. The initial state consists of two massive particles, the final state of two massless ones. Not only the number of particles changes, but their species as well.

Historically, this difficulty was recognized almost as soon as quantum mechanics was completed. The first wave equation Schrödinger wrote down, at the end of 1925, was a relativistic one — what we now call the Klein–Gordon equation — but it disagreed with experiment on the fine structure of hydrogen, so he abandoned it and published the nonrelativistic equation instead. The same equation, rediscovered independently by Klein, Gordon, Fock and others in 1926, carried with it the disease that its probability density can go negative.

Dirac, meanwhile, quantized the electromagnetic field itself in a paper of 1927 and computed the emission and absorption of light by atoms from first principles. This was the first quantum field theory. The electromagnetic field is a field already classically, and quantizing it produced a particle: the photon. A natural question follows.

Might the electron, too, be the quantum of some field?

The answer is yes. To reach that conclusion, however, we need to know precisely how one-particle relativistic quantum mechanics fails. Below we exhibit the failure in three layers — negative energy, the probabilistic interpretation, and causality — and see that quantizing the field cures all of them at once.

flowchart TD
A["Nonrelativistic quantum mechanics<br/>particle number is conserved"] --> B["Substitute the relativistic dispersion relation"]
B --> C["Klein-Gordon equation"]
B --> D["Dirac equation"]
C --> E["Negative-energy solutions<br/>probability density goes negative"]
D --> F["Negative-energy solutions<br/>the Dirac sea fails for bosons"]
E --> G["Propagation amplitude nonzero even at spacelike separation"]
F --> G
G --> H["Quantize the field itself"]
H --> I["Particle = excitation quantum of a field<br/>antiparticles, creation and annihilation, microcausality"]
How one-particle relativistic quantum mechanics breaks down, and the route to field quantization

2. Preliminaries: two assumptions hidden in nonrelativistic quantum mechanics

Section titled “2. Preliminaries: two assumptions hidden in nonrelativistic quantum mechanics”

Unless stated otherwise we use natural units =c=1\hbar = c = 1 and the metric ημν=diag(+1,1,1,1)\eta_{\mu\nu} = \mathrm{diag}(+1,-1,-1,-1). A spacetime point is x=(t,x)x = (t,\boldsymbol{x}), the inner product is px=p0tpxp\cdot x = p^0 t - \boldsymbol{p}\cdot\boldsymbol{x}, and the on-shell energy is written ωp=p2+m2\omega_{\boldsymbol{p}} = \sqrt{\boldsymbol{p}^2 + m^2}. We write =μμ=t22\Box = \partial_\mu\partial^\mu = \partial_t^2 - \nabla^2.

Nonrelativistic quantum mechanics tacitly assumes the following two things. Both breakdowns below originate here.

(A1) The particle number is fixed. The state space is HN=L2(R3N)\mathcal{H}_N = L^2(\mathbb{R}^{3N}) (for identical particles, its symmetrized or antisymmetrized subspace). The Hamiltonian maps HN\mathcal{H}_N into HN\mathcal{H}_N, so the particle number NN is conserved by definition. We have not proved a conservation law; we have merely chosen a space in which a situation with non-conservation cannot be written down.

(A2) Time and space are treated asymmetrically. The position x^\hat{\boldsymbol{x}} is an operator, whereas the time tt is not an operator but a mere parameter. The canonical commutation relation reads [x^i,p^j]=iδij[\hat{x}^i, \hat{p}^j] = i\delta^{ij}, and tt does not appear in it. But a Lorentz transformation mixes tt with x\boldsymbol{x}. Since a quantity that is an operator and a quantity that is a parameter get mixed by the transformation, this distinction is not Lorentz invariant.

Remark 2.1The route through a time operator is blocked

One might think that (A2) can be repaired by promoting tt to an operator. Suppose, however, that a self-adjoint T^\hat T with [T^,H^]=i[\hat T,\hat H] = i exists. Since [T^,H^][\hat T,\hat H] is a c-number, the Baker–Campbell–Hausdorff expansion terminates at first order:

eiϵT^H^eiϵT^=H^+iϵ[T^,H^]=H^ϵ.e^{i\epsilon \hat T}\hat H e^{-i\epsilon \hat T} = \hat H + i\epsilon[\hat T,\hat H] = \hat H - \epsilon .

The left-hand side is a unitary transformation and therefore leaves the spectrum unchanged. Hence the spectrum of H^\hat H is invariant under translation by an arbitrary real number ϵ\epsilon, which forces it to be all of R\mathbb{R}. For a system whose energy is bounded below this cannot hold (Pauli’s argument). The direction to take is therefore the opposite one: demote x^\hat{\boldsymbol{x}} from being an operator, and let both entries of x=(t,x)x = (t,\boldsymbol{x}) be mere labels. What becomes an operator instead is the field ϕ(x)\phi(x), which takes a value at each spacetime point.

3. The naive relativistic upgrade: two diseases of the Klein–Gordon equation

Section titled “3. The naive relativistic upgrade: two diseases of the Klein–Gordon equation”

In the nonrelativistic case we obtained the Schrödinger equation by substituting EitE \to i\partial_t and pi\boldsymbol{p}\to -i\nabla into E=p2/2mE = \boldsymbol{p}^2/2m. Let us apply the same substitution to the relativistic relation E2=p2+m2E^2 = \boldsymbol{p}^2 + m^2 (see relativistic mechanics).

Definition 3.1The Klein–Gordon equation

For a complex scalar field of mass m>0m > 0 — or a candidate one-particle wave function — ϕ(t,x)\phi(t,\boldsymbol{x}), the equation

(+m2)ϕ=0,=t22\left(\Box + m^2\right)\phi = 0, \qquad \Box = \partial_t^2 - \nabla^2

is called the Klein–Gordon equation.

Disease 1: the negative-energy solutions cannot be discarded

Section titled “Disease 1: the negative-energy solutions cannot be discarded”

Substituting the plane wave ϕ=Nei(Etpx)\phi = N e^{-i(Et - \boldsymbol{p}\cdot\boldsymbol{x})} gives E2+p2+m2=0-E^2 + \boldsymbol{p}^2 + m^2 = 0, that is, the two branches

E=±ωp=±p2+m2.E = \pm\,\omega_{\boldsymbol{p}} = \pm\sqrt{\boldsymbol{p}^2 + m^2} .

The negative branch means that the energy is not bounded below: perturb the system and it can keep falling into ever lower states.

One is tempted to say that the negative-energy solutions are unphysical and may simply be thrown away. They cannot be. The reason lies in the structure of the initial value problem. The Klein–Gordon equation is second order in time, so the Cauchy data consist of two arbitrary functions, ϕ(0,x)\phi(0,\boldsymbol{x}) and ϕ˙(0,x)\dot\phi(0,\boldsymbol{x}). A spatial Fourier transform gives ϕ~¨=ωp2ϕ~\ddot{\tilde\phi} = -\omega_{\boldsymbol p}^2\tilde\phi for each p\boldsymbol{p}, whose general solution is

ϕ~(t,p)=a(p)eiωpt+b(p)e+iωpt.\tilde\phi(t,\boldsymbol{p}) = a(\boldsymbol{p})e^{-i\omega_{\boldsymbol p}t} + b(\boldsymbol{p})e^{+i\omega_{\boldsymbol p}t} .

Restricting to positive frequencies — setting b0b \equiv 0 — is nothing other than imposing the constraint ϕ~˙(0,p)=iωpϕ~(0,p)\dot{\tilde\phi}(0,\boldsymbol{p}) = -i\omega_{\boldsymbol{p}}\tilde\phi(0,\boldsymbol{p}). In position space this is the nonlocal relation ϕ˙(0,)=i2+m2ϕ(0,)\dot\phi(0,\cdot) = -i\sqrt{-\nabla^2+m^2}\,\phi(0,\cdot), so we have surrendered the right to prescribe the initial data freely and locally. What this nonlocality destroys is the subject of Proposition 5.1.

Disease 2: the time component of the conserved current goes negative

Section titled “Disease 2: the time component of the conserved current goes negative”

For the Schrödinger equation, ρ=ψ20\rho = \lvert\psi\rvert^2 \ge 0 satisfies a continuity equation and can be read as a probability density. The Klein–Gordon equation also admits a conserved current, but positivity is lost.

Proposition 3.2The Klein–Gordon current and its failure of positivity

Let ϕ\phi be a solution of Definition 3.1. Then

jμ=i2m(ϕμϕ(μϕ)ϕ)j^\mu = \frac{i}{2m}\left(\phi^{*}\partial^\mu\phi - (\partial^\mu\phi^{*})\phi\right)

satisfies μjμ=0\partial_\mu j^\mu = 0. Moreover, for a plane-wave solution ϕ=Nei(Etpx)\phi = Ne^{-i(Et-\boldsymbol{p}\cdot\boldsymbol{x})} with E=±ωpE = \pm\omega_{\boldsymbol p},

j0=EmN2.j^0 = \frac{E}{m}\lvert N\rvert^2 .

Hence j0<0j^0 < 0 for the negative-energy solutions, and j0j^0 cannot be interpreted as a probability density.

Proof(Proposition 3.2)

Let us first verify conservation. By the product rule,

μjμ=i2m(μϕμϕ+ϕϕ(ϕ)ϕμϕμϕ)=i2m(ϕϕ(ϕ)ϕ).\partial_\mu j^\mu = \frac{i}{2m}\left(\partial_\mu\phi^{*}\partial^\mu\phi + \phi^{*}\Box\phi - (\Box\phi^{*})\phi - \partial_\mu\phi^{*}\partial^\mu\phi\right) = \frac{i}{2m}\left(\phi^{*}\Box\phi - (\Box\phi^{*})\phi\right).

The first and fourth terms have cancelled. Now Definition 3.1 gives ϕ=m2ϕ\Box\phi = -m^2\phi and, on taking complex conjugates, ϕ=m2ϕ\Box\phi^{*} = -m^2\phi^{*}, so

μjμ=i2m(m2ϕϕ+m2ϕϕ)=0.\partial_\mu j^\mu = \frac{i}{2m}\left(-m^2\phi^{*}\phi + m^2\phi^{*}\phi\right) = 0 .

Next we compute j0j^0. With our metric convention 0=η000=t\partial^0 = \eta^{00}\partial_0 = \partial_t, so j0=i2m(ϕϕ˙ϕ˙ϕ)j^0 = \frac{i}{2m}(\phi^{*}\dot\phi - \dot\phi^{*}\phi). For a plane wave ϕ˙=iEϕ\dot\phi = -iE\phi and ϕ˙=+iEϕ\dot\phi^{*} = +iE\phi^{*}, whence

j0=i2m(N2(iE)(iE)N2)=i2m(2iEN2)=EmN2.j^0 = \frac{i}{2m}\left(\lvert N\rvert^2(-iE) - (iE)\lvert N\rvert^2\right) = \frac{i}{2m}\left(-2iE\lvert N\rvert^2\right) = \frac{E}{m}\lvert N\rvert^2 .

On the branch E=ωp<0E = -\omega_{\boldsymbol p} < 0 the right-hand side is negative. A density that goes negative is not a probability density.

Example 3.3The nonrelativistic limit recovers the correct probability density

It is not that j0j^0 is entirely meaningless. Extract the rest energy by writing ϕ(t,x)=eimtψ(t,x)\phi(t,\boldsymbol{x}) = e^{-imt}\psi(t,\boldsymbol{x}) and assume the nonrelativistic limit ψ˙mψ\lvert\dot\psi\rvert \ll m\lvert\psi\rvert. Then

ϕ˙=eimt(ψ˙imψ)imeimtψ,\dot\phi = e^{-imt}\left(\dot\psi - im\psi\right) \simeq -im\,e^{-imt}\psi ,

so substituting into the expression for j0j^0 in Proposition 3.2 gives

j0i2m(ψˉ(imψ)(imψˉ)ψ)=i2m(2imψ2)=ψ2,j^0 \simeq \frac{i}{2m}\left(\bar\psi\,(-im\psi) - (im\bar\psi)\,\psi\right) = \frac{i}{2m}\left(-2im\lvert\psi\rvert^2\right) = \lvert\psi\rvert^2 ,

which is exactly the probability density of the Schrödinger equation. In other words, j0j^0 is a quantity that looks like a probability density in the nonrelativistic limit and changes sign in the relativistic regime.

Remark 3.4The disease was a misreading

Let us give the game away in advance: jμj^\mu itself is a perfectly good conserved quantity. What was wrong was the reading of it as a probability density. In quantum field theory ejμe\,j^\mu is the electric charge current, and charge can be positive or negative. A solution with j0<0j^0 < 0 describes not a particle of negative probability but a particle carrying charge of the opposite sign — an antiparticle. Pauli and Weisskopf established this rereading in 1934 by quantizing the Klein–Gordon equation as a field. We take the matter up after Theorem 7.1.

In 1928 Dirac traced the failure of positivity of the probability density to the equation being second order in time. For a first-order equation, ρ=ψψ0\rho = \psi^\dagger\psi \ge 0 ought to be a conserved density, just as for the Schrödinger equation.

Definition 4.1The Dirac equation

The first-order equation

itψ=H^ψ,H^=αp^+βmi\,\partial_t \psi = \hat H\psi, \qquad \hat H = \boldsymbol{\alpha}\cdot\hat{\boldsymbol{p}} + \beta m

for an NN-component wave function ψ(t,x)CN\psi(t,\boldsymbol{x}) \in \mathbb{C}^N is called the Dirac equation. Here α1,α2,α3,β\alpha^1,\alpha^2,\alpha^3,\beta are Hermitian N×NN\times N matrices, required to satisfy

{αi,αj}=2δij1,{αi,β}=0,β2=1\{\alpha^i,\alpha^j\} = 2\delta^{ij}\mathbb{1}, \qquad \{\alpha^i,\beta\} = 0, \qquad \beta^2 = \mathbb{1}

so that H^2=p^2+m2\hat H^2 = \hat{\boldsymbol{p}}^2 + m^2 holds.

That these relations are equivalent to H^2=p^2+m2\hat H^2 = \hat{\boldsymbol{p}}^2 + m^2 becomes visible on expanding:

H^2=αiαjp^ip^j+m(αiβ+βαi)p^i+β2m2=12{αi,αj}p^ip^j+m{αi,β}p^i+β2m2.\hat H^2 = \alpha^i\alpha^j \hat p_i\hat p_j + m\left(\alpha^i\beta + \beta\alpha^i\right)\hat p_i + \beta^2 m^2 = \tfrac{1}{2}\{\alpha^i,\alpha^j\}\hat p_i\hat p_j + m\{\alpha^i,\beta\}\hat p_i + \beta^2m^2 .

In the first term p^ip^j\hat p_i \hat p_j is symmetric under iji \leftrightarrow j, so only the symmetric part of αiαj\alpha^i\alpha^j contributes, and that part is half the anticommutator. Under the three conditions above we therefore obtain H^2=δijp^ip^j+m2=p^2+m2\hat H^2 = \delta^{ij}\hat p_i\hat p_j + m^2 = \hat{\boldsymbol{p}}^2 + m^2.

Since H^\hat H is Hermitian, ρ=ψψ\rho = \psi^\dagger\psi and j=ψαψ\boldsymbol{j} = \psi^\dagger\boldsymbol{\alpha}\psi satisfy the continuity equation tρ+j=0\partial_t\rho + \nabla\cdot\boldsymbol{j} = 0, and moreover ρ0\rho \ge 0. Disease 2 of Proposition 3.2 has been cured. Disease 1, however, survives.

Theorem 4.2The Dirac Hamiltonian is unbounded below

Suppose there exist Hermitian N×NN\times N matrices αi,β\alpha^i,\beta satisfying the conditions of Definition 4.1. For the matrix H(p)=αp+βmH(\boldsymbol{p}) = \boldsymbol{\alpha}\cdot\boldsymbol{p} + \beta m attached to a momentum eigenvalue p\boldsymbol{p}, the following hold.

  1. trαi=0\operatorname{tr}\alpha^i = 0 for i=1,2,3i=1,2,3, and trβ=0\operatorname{tr}\beta = 0. Hence trH(p)=0\operatorname{tr}H(\boldsymbol{p}) = 0.
  2. The eigenvalues of H(p)H(\boldsymbol{p}) are +ωp+\omega_{\boldsymbol p} and ωp-\omega_{\boldsymbol p} only, and the two have equal multiplicity N/2N/2. In particular NN is even, the spectrum of H^\hat H is (,m][m,)(-\infty,-m]\cup[m,\infty), and it is unbounded below.
Proof(Theorem 4.2)

Proof of 1. From {αi,β}=0\{\alpha^i,\beta\}=0 we have βαi=αiβ\beta\alpha^i = -\alpha^i\beta; multiplying on the right by β\beta and using β2=1\beta^2=\mathbb{1} gives

βαiβ=αiβ2=αi.\beta\alpha^i\beta = -\alpha^i\beta^2 = -\alpha^i .

Take traces of both sides. By cyclicity, the left-hand side is tr(βαiβ)=tr(αiβ2)=trαi\operatorname{tr}(\beta\alpha^i\beta) = \operatorname{tr}(\alpha^i\beta^2) = \operatorname{tr}\alpha^i. Hence trαi=trαi\operatorname{tr}\alpha^i = -\operatorname{tr}\alpha^i, that is, trαi=0\operatorname{tr}\alpha^i = 0. Similarly, from (α1)2=1(\alpha^1)^2 = \mathbb{1} and {α1,β}=0\{\alpha^1,\beta\}=0 we get α1βα1=β(α1)2=β\alpha^1\beta\alpha^1 = -\beta(\alpha^1)^2 = -\beta, and taking traces yields trβ=trβ=0\operatorname{tr}\beta = -\operatorname{tr}\beta = 0. By linearity, trH(p)=pitrαi+mtrβ=0\operatorname{tr}H(\boldsymbol{p}) = p_i\operatorname{tr}\alpha^i + m\operatorname{tr}\beta = 0.

Proof of 2. As checked immediately after Definition 4.1, H(p)2=(p2+m2)1=ωp21H(\boldsymbol{p})^2 = (\boldsymbol{p}^2+m^2)\mathbb{1} = \omega_{\boldsymbol p}^2\mathbb{1}. Since H(p)H(\boldsymbol p) is Hermitian it is diagonalizable, and its eigenvalues λ\lambda obey λ2=ωp2\lambda^2 = \omega_{\boldsymbol p}^2, so λ=±ωp\lambda = \pm\omega_{\boldsymbol p}. Writing n+,nn_+,n_- for the multiplicities we have n++n=Nn_++n_-=N, and part 1 gives

trH(p)=n+ωpnωp=0.\operatorname{tr}H(\boldsymbol p) = n_+\omega_{\boldsymbol p} - n_-\omega_{\boldsymbol p} = 0 .

Since m>0m > 0 we have ωpm>0\omega_{\boldsymbol p} \ge m > 0, whence n+=n=N/2n_+ = n_- = N/2. As N1N \ge 1, we get n1n_- \ge 1: negative eigenvalues necessarily exist. Letting p\boldsymbol{p} range over all of space, ωp-\omega_{\boldsymbol p} takes every value continuously from m-m down to -\infty, so the spectrum is unbounded below.

Thus making the equation first order does not remove the negative-energy solutions. Theorem 4.2 says more than that they survive: exactly half of the states have negative energy.

Example 4.3The Dirac sea and its price

Dirac’s prescription runs as follows. The electron is a fermion, so the Pauli exclusion principle applies. Define the vacuum as the state in which every negative-energy level is occupied by an electron; then a positive-energy electron cannot fall into a negative-energy level, because exclusion forbids it. A hole in the sea behaves like a particle of charge +e+e and positive energy. This is the positron, discovered by Anderson in a cloud chamber in 1932 — a spectacular success as a prediction.

The price, however, is high.

  1. The vacuum contains infinitely many electrons, and both its charge density and its energy density diverge. Observables have to be redefined as differences from the vacuum.
  2. The construction relies on exclusion, so it is unusable for bosons. That spin-0 pions have antiparticles (π+\pi^+ and π\pi^-) is an experimental fact.
  3. Decisively, this is no longer a one-particle theory. The moment infinitely many particles are introduced the description has become a many-body one, and the prescription itself concedes that one-particle relativistic quantum mechanics does not hold up.

Remark 4.4Spin emerges from relativity

Theorem 4.2 also shows that NN is even. In fact N=2N = 2 is impossible, and the minimal dimension is 44 (Exercise 9.2). Merely attempting to write a relativistic first-order equation forces the wave function to have several components, and that internal degree of freedom corresponds to the spin 1/21/2 introduced in angular momentum and spin. Spin is a consequence of combining relativity with quantum mechanics.

5. A deeper disease: a one-particle theory cannot respect causality

Section titled “5. A deeper disease: a one-particle theory cannot respect causality”

One escape route from the negative-energy problem still seems open: keep only the positive-frequency part. That is, restrict the Hilbert space to positive-energy solutions and adopt

H=2+m2H = \sqrt{-\nabla^2 + m^2}

as the Hamiltonian. This is a self-adjoint operator on L2(R3)L^2(\mathbb{R}^3) with spectrum [m,)[m,\infty); it is indeed bounded below, and the difficulty of Theorem 4.2 is formally evaded. But at this point a deeper disease appears: causality, the very heart of relativity, is violated.

5.1. Compact support is destroyed instantaneously

Section titled “5.1. Compact support is destroyed instantaneously”

Proposition 5.1A positive-energy particle cannot stay localized

Let m>0m > 0, H=2+m2H = \sqrt{-\nabla^2+m^2} and ψt=eitHψ0\psi_t = e^{-itH}\psi_0. If ψ0L2(R3)\psi_0 \in L^2(\mathbb{R}^3) has compact support and ψ00\psi_0 \ne 0, then ψt\psi_t fails to have compact support at every time t0t \ne 0.

Proof(Proposition 5.1)

Under a spatial Fourier transform, ψ^t(p)=eitωpψ^0(p)\hat\psi_t(\boldsymbol{p}) = e^{-it\omega_{\boldsymbol p}}\hat\psi_0(\boldsymbol{p}) with ωp=p2+m2\omega_{\boldsymbol p} = \sqrt{\boldsymbol{p}^2+m^2}.

We argue by contradiction. Suppose that for some t0t\ne 0 the function ψt\psi_t also has compact support. By the Paley–Wiener–Schwartz theorem, the Fourier transform of a compactly supported L2L^2 function continues analytically to an entire function (of exponential type) on C3\mathbb{C}^3. Hence ψ^0\hat\psi_0 and ψ^t\hat\psi_t both extend to entire functions F0,FtF_0, F_t.

Since ψ00\psi_0 \ne 0, the function F0F_0 is not identically zero, and as F0F_0 is continuous at real points there is a real vector p0\boldsymbol{p}_0 with F0(p0)0F_0(\boldsymbol{p}_0)\ne 0. Fix a real unit vector e\boldsymbol{e}, restrict to the complex line p=p0+ze\boldsymbol{p} = \boldsymbol{p}_0 + z\boldsymbol{e} with zCz\in\mathbb{C}, and set f0(z)=F0(p0+ze)f_0(z) = F_0(\boldsymbol{p}_0+z\boldsymbol{e}), ft(z)=Ft(p0+ze)f_t(z) = F_t(\boldsymbol{p}_0+z\boldsymbol{e}). Both are entire in zz, and f0(0)0f_0(0)\ne 0.

Along this line,

Q(z):=(p0+ze)2+m2=z2+2(p0e)z+p02+m2.Q(z) := (\boldsymbol{p}_0+z\boldsymbol{e})^2 + m^2 = z^2 + 2(\boldsymbol{p}_0\cdot\boldsymbol{e})z + \boldsymbol{p}_0^2 + m^2 .

Its discriminant is 4(p0e)24(p02+m2)4(\boldsymbol{p}_0\cdot\boldsymbol{e})^2 - 4(\boldsymbol{p}_0^2+m^2), which is negative because Cauchy–Schwarz gives (p0e)2p02<p02+m2(\boldsymbol{p}_0\cdot\boldsymbol{e})^2 \le \boldsymbol{p}_0^2 < \boldsymbol{p}_0^2+m^2 (here we used m>0m>0). Hence QQ has two distinct non-real roots z±z_\pm, and Q(z)=(zz+)(zz)\sqrt{Q(z)} = \sqrt{(z-z_+)(z-z_-)} has genuine branch points at z±z_\pm.

On the real axis, ft(z)=g(z)f0(z)f_t(z) = g(z)f_0(z) with g(z)=eitQ(z)g(z) = e^{-it\sqrt{Q(z)}}. Since f0,ftf_0, f_t are entire and f0≢0f_0\not\equiv 0, the quotient ft/f0f_t/f_0 is a single-valued meromorphic function on C\mathbb{C}. On the other hand, continuing gg once around z+z_+ reverses the sign of Q\sqrt{Q} and carries gg into e+itQe^{+it\sqrt{Q}}. For t0t\ne 0, the equation e2itQ(z)=1e^{-2it\sqrt{Q(z)}} = 1 holds only at the isolated points where Q(z)(π/t)Z\sqrt{Q(z)} \in (\pi/t)\mathbb{Z}, so gg is not single-valued near z+z_+. A single-valued meromorphic function cannot agree on the real axis with a function that is not single-valued, and we have a contradiction.

That the support ceases to be compact means that a particle confined to a bounded region at the initial time has nonzero amplitude arbitrarily far away after an arbitrarily short time t>0t > 0. The limit imposed by the speed of light is not in force. Let us next estimate the size of that amplitude.

Example 5.2The propagation amplitude outside the light cone

The amplitude for a particle located at the point x\boldsymbol{x} at time 00 to be found at the point y\boldsymbol{y} at time tt is

U(t,r)=yeiHtx=d3p(2π)3  eip(yx)eiωpt,r=yx.U(t,r) = \langle \boldsymbol{y}\,\lvert\, e^{-iHt}\,\rvert\,\boldsymbol{x}\rangle = \int\frac{d^3p}{(2\pi)^3}\;e^{i\boldsymbol{p}\cdot(\boldsymbol{y}-\boldsymbol{x})}\,e^{-i\omega_{\boldsymbol p}t}, \qquad r = \lvert\boldsymbol{y}-\boldsymbol{x}\rvert .

Performing the angular integral (with dΩeiprcosθ=4πsin(pr)/(pr)\int d\Omega\, e^{ipr\cos\theta} = 4\pi\sin(pr)/(pr)) gives

U(t,r)=12π2r0dp  psin(pr)eiωpt.U(t,r) = \frac{1}{2\pi^2 r}\int_0^\infty dp\;p\,\sin(pr)\,e^{-i\omega_p t} .

Deforming the contour around the branch points p=±imp = \pm im of p2+m2\sqrt{p^2+m^2} (the computation is in the Appendix), one obtains, for a spacelike separation r>t>0r > t > 0, the expression

U(t,r)=i2π2rmdρ  ρeρrsinh ⁣(tρ2m2).U(t,r) = \frac{i}{2\pi^2 r}\int_m^\infty d\rho\;\rho\,e^{-\rho r}\,\sinh\!\left(t\sqrt{\rho^2-m^2}\right) .

Inside the range of integration the integrand is strictly positive (for ρ>m\rho > m one has sinh(tρ2m2)>0\sinh(t\sqrt{\rho^2-m^2}) > 0). Hence U(t,r)0U(t,r) \ne 0: the amplitude does not vanish outside the light cone.

Its magnitude can be read off from the fact that for mr1mr \gg 1 and rtr \gg t the lower endpoint ρm\rho\simeq m dominates. Setting ρ=m+u\rho = m+u, using ρ2m22mu\sqrt{\rho^2-m^2}\simeq\sqrt{2mu} and sinhxx\sinh x\simeq x, and using 0duueur=π/(2r3/2)\int_0^\infty du\,\sqrt{u}\,e^{-ur} = \sqrt{\pi}/(2r^{3/2}), we find

U2π  m3/2t4π2r5/2emr.\lvert U\rvert \sim \frac{\sqrt{2\pi}\;m^{3/2}\,t}{4\pi^2\,r^{5/2}}\,e^{-mr} .

A more careful saddle-point evaluation turns the exponent into the Lorentz-invariant form emr2t2e^{-m\sqrt{r^2-t^2}}. The amplitude is exponentially small at distances beyond the Compton wavelength 1/m1/m, but it is never zero.

txOlight conePspacelike regionamplitude ≠ 0
Even at a point P spacelike separated from the event at the origin, the one-particle propagation amplitude is merely exponentially small, never zero. The shaded areas are the spacelike region.

5.2. Localization and causal propagation are incompatible

Section titled “5.2. Localization and causal propagation are incompatible”

Example 5.2 was an explicit computation, but the phenomenon is in fact a general one that follows from the single condition H0H \ge 0. The following theorem is due to Hegerfeldt (1974).

Theorem 5.3Hegerfeldt's theorem

Let H\mathcal{H} be a Hilbert space and let HH be a self-adjoint operator on H\mathcal{H} with H0H \ge 0 (its spectrum contained in [0,)[0,\infty)). Let ψ0H\psi_0\in\mathcal{H}, ψt=eiHtψ0\psi_t = e^{-iHt}\psi_0, and let AA be a bounded positive operator (A0A \ge 0). If

ψt,Aψt=0(tI)\langle \psi_t, A\psi_t\rangle = 0 \quad (\forall t\in I)

holds on some open interval IRI\subset\mathbb{R}, then ψt,Aψt=0\langle\psi_t,A\psi_t\rangle = 0 for every tRt\in\mathbb{R}.

Proof(Theorem 5.3)

Since A0A\ge 0 and AA is bounded, a bounded square root A1/20A^{1/2}\ge 0 exists and ψt,Aψt=A1/2ψt2\langle\psi_t,A\psi_t\rangle = \lVert A^{1/2}\psi_t\rVert^2. The hypothesis is therefore equivalent to A1/2ψt=0A^{1/2}\psi_t = 0 for tIt\in I.

Fix an arbitrary χH\chi\in\mathcal{H} and consider, on the lower half plane C={z:Imz<0}\mathbb{C}_- = \{z : \operatorname{Im}z < 0\},

g(z)=χ,  A1/2eizHψ0.g(z) = \left\langle \chi,\; A^{1/2}e^{-izH}\psi_0\right\rangle .

Let E()E(\cdot) be the spectral measure of HH and introduce the complex measure μ()=χ,A1/2E()ψ0\mu(\cdot) = \langle \chi, A^{1/2}E(\cdot)\psi_0\rangle, whose total variation is finite and at most χA1/2ψ0\lVert\chi\rVert\,\lVert A^{1/2}\rVert\,\lVert\psi_0\rVert. Then

g(z)=[0,)eizλdμ(λ).g(z) = \int_{[0,\infty)} e^{-iz\lambda}\,d\mu(\lambda) .

Because H0H\ge 0 the spectrum is confined to λ0\lambda\ge 0, and for z=tisz = t-is with s>0s>0 we have eizλ=esλ1\lvert e^{-iz\lambda}\rvert = e^{-s\lambda}\le 1. Hence gg is bounded on C\mathbb{C}_-, and since the integrand is holomorphic and admits a uniform integrable majorant, Fubini’s theorem together with Morera’s theorem shows that gg is holomorphic there. By dominated convergence it has the boundary value g(t)=χ,A1/2ψtg(t) = \langle\chi, A^{1/2}\psi_t\rangle.

By hypothesis the boundary value vanishes on II, a set of positive Lebesgue measure. A bounded holomorphic function on a half plane — an element of the Hardy space HH^\infty — whose boundary values vanish on a set of positive measure is identically zero, by the boundary uniqueness theorem (the theorem of F. and M. Riesz–Privalov). Hence g0g\equiv 0.

As χ\chi was arbitrary, A1/2ψt=0A^{1/2}\psi_t = 0 for every real tt, that is, ψt,Aψt=A1/2ψt2=0\langle\psi_t,A\psi_t\rangle = \lVert A^{1/2}\psi_t\rVert^2 = 0.

Corollary 5.4Strict localization and causal propagation are incompatible

Let H0H\ge 0, and suppose that for every bounded open set WR3W\subset\mathbb{R}^3 we are given a projection PWP_W expressing that the particle is found inside WW, with PWPWP_W \le P_{W'} whenever WWW\subset W'. Suppose a state ψ0\psi_0 is strictly localized in a bounded region VV and that propagation is causal, that is,

dist(W,V)>t    ψt,PWψt=0\operatorname{dist}(W,V) > \lvert t\rvert \;\Longrightarrow\; \langle\psi_t, P_W\psi_t\rangle = 0

holds (with c=1c=1). Then for every bounded open set WW lying at positive distance from VV we have ψt,PWψt=0\langle\psi_t,P_W\psi_t\rangle = 0 at all times tRt\in\mathbb{R}.

Proof(Corollary 5.4)

Put d=dist(W,V)>0d = \operatorname{dist}(W,V) > 0. The hypothesis of causal propagation gives ψt,PWψt=0\langle\psi_t,P_W\psi_t\rangle = 0 on the open interval I=(d,d)I = (-d,d). The projection PWP_W is bounded and satisfies PW=PW2=PW0P_W = P_W^2 = P_W^\dagger \ge 0, so Theorem 5.3 applies with A=PWA = P_W, and the conclusion extends to all tt.

The particle would therefore never be found outside Vˉ\bar V, for all time. But the spectrum of H=2+m2H=\sqrt{-\nabla^2+m^2} is absolutely continuous, so the RAGE theorem gives PKψt0\lVert P_K\psi_t\rVert \to 0 as tt\to\infty for every bounded region KK: the particle must spread. This is a contradiction.

The conclusion is this. The three requirements — H0H\ge 0, a notion of strict localization, and causal propagation — cannot hold simultaneously. Abandon H0H\ge 0 and we are back to the difficulty of Theorem 4.2; abandon causality and we abandon relativity; abandon localization and the very question of where the particle is loses its meaning.

6. Particle number is not conserved: the Compton wavelength as a scale

Section titled “6. Particle number is not conserved: the Compton wavelength as a scale”

So far the breakdowns have been logical ones. Physically, too, the reason a one-particle theory cannot hold is plain. In relativity, mass and energy are interchangeable (relativistic mechanics), so with enough energy available particles simply get created.

Definition 6.1The reduced Compton wavelength

For a particle of mass mm, the quantity

λˉC=mc\bar\lambda_C = \frac{\hbar}{mc}

is called the reduced Compton wavelength (multiplying it by 2π2\pi gives the Compton wavelength λC=h/mc\lambda_C = h/mc).

The meaning of this scale can be read off from the uncertainty relation. Confining a particle to a region of size Δx\Delta x creates a momentum uncertainty Δp/Δx\Delta p \gtrsim \hbar/\Delta x, which relativistically amounts to an energy uncertainty

ΔEcΔpcΔx.\Delta E \gtrsim c\,\Delta p \gtrsim \frac{\hbar c}{\Delta x} .

Once ΔxλˉC\Delta x \lesssim \bar\lambda_C we have ΔEmc2\Delta E \gtrsim mc^2, which is enough energy to be borrowed for the creation of a particle–antiparticle pair.

Example 6.2Numbers for the electron

We use =1.055×1034 Js\hbar = 1.055\times10^{-34}\ \mathrm{J\,s}, me=9.109×1031 kgm_e = 9.109\times10^{-31}\ \mathrm{kg} and c=2.998×108 m/sc = 2.998\times10^{8}\ \mathrm{m/s}. Since mec=2.731×1022 kgm/sm_e c = 2.731\times10^{-22}\ \mathrm{kg\,m/s},

λˉC=1.055×10342.731×1022 m=3.86×1013 m=386 fm.\bar\lambda_C = \frac{1.055\times10^{-34}}{2.731\times10^{-22}}\ \mathrm{m} = 3.86\times10^{-13}\ \mathrm{m} = 386\ \mathrm{fm} .

The threshold for pair creation, on the other hand, is 2mec2=1.022 MeV2m_ec^2 = 1.022\ \mathrm{MeV}. So if one tries to squeeze an electron into a region narrower than about 400 fm400\ \mathrm{fm}, the energy poured in for that purpose creates a new electron–positron pair. It ceases to be meaningful to speak of “the wave function of one electron” at a resolution finer than λˉC\bar\lambda_C.

This is the same scale as the decay length 1/m1/m of the amplitude found in Example 5.2 (in natural units). That the distance at which the causality-violating amplitude matters coincides with the distance at which pair creation sets in is no accident. Both are faces of the same fact: the one-particle picture breaks down.

There is a single prescription that resolves all three breakdowns above (negative energy, probability density, causality) together with the non-conservation of particle number: treat the classical field as the dynamical variable and quantize it canonically.

7.1. A free field is a collection of independent harmonic oscillators

Section titled “7.1. A free field is a collection of independent harmonic oscillators”

The starting point is Lagrangian mechanics extended to fields (treated in detail in classical field theory and the Lagrangian). The simplest Lagrangian for a real scalar field ϕ(t,x)\phi(t,\boldsymbol{x}) is

L=d3x  L,L=12ϕ˙212(ϕ)212m2ϕ2,L = \int d^3x\;\mathcal{L}, \qquad \mathcal{L} = \tfrac12\dot\phi^2 - \tfrac12(\nabla\phi)^2 - \tfrac12 m^2\phi^2 ,

and its Euler–Lagrange equation is precisely the Klein–Gordon equation of Definition 3.1. Here lies the decisive change of viewpoint: we reread the Klein–Gordon equation not as the equation for a one-particle wave function but as the equation of motion of a classical field.

Theorem 7.1Mode decomposition and quantization of the free scalar field

Let a real scalar field ϕ\phi inside a cube of side LL (volume V=L3V = L^3, periodic boundary conditions) obey the Lagrangian above. Substituting the Fourier expansion

ϕ(t,x)=1Vkqk(t)eikx,k2πLZ3,qk=qk\phi(t,\boldsymbol{x}) = \frac{1}{\sqrt{V}}\sum_{\boldsymbol{k}} q_{\boldsymbol{k}}(t)\,e^{i\boldsymbol{k}\cdot\boldsymbol{x}}, \qquad \boldsymbol{k}\in\frac{2\pi}{L}\mathbb{Z}^3, \qquad q_{-\boldsymbol{k}} = \overline{q_{\boldsymbol{k}}}

gives

L=k12(q˙k2ωk2qk2),ωk=k2+m2.L = \sum_{\boldsymbol{k}}\frac12\left(\lvert\dot q_{\boldsymbol{k}}\rvert^2 - \omega_{\boldsymbol{k}}^2\,\lvert q_{\boldsymbol{k}}\rvert^2\right), \qquad \omega_{\boldsymbol{k}} = \sqrt{\boldsymbol{k}^2+m^2} .

That is, a free field is equivalent to a collection of independent harmonic oscillators, one per mode, with angular frequency ωk\omega_{\boldsymbol{k}}. Quantizing canonically, with [a^k,a^k]=δkk[\hat a_{\boldsymbol k},\hat a^\dagger_{\boldsymbol k'}] = \delta_{\boldsymbol{k}\boldsymbol{k}'} and all other commutators zero, one obtains

H^=kωk(a^ka^k+12),P^=kka^ka^k,\hat H = \sum_{\boldsymbol{k}}\omega_{\boldsymbol{k}}\left(\hat a^\dagger_{\boldsymbol{k}}\hat a_{\boldsymbol{k}} + \tfrac12\right), \qquad \hat{\boldsymbol{P}} = \sum_{\boldsymbol{k}}\boldsymbol{k}\,\hat a^\dagger_{\boldsymbol{k}}\hat a_{\boldsymbol{k}} ,

with energy eigenvalues (apart from the zero-point energy) knkωk\sum_{\boldsymbol{k}} n_{\boldsymbol{k}}\,\omega_{\boldsymbol{k}}, where nk{0,1,2,}n_{\boldsymbol{k}}\in\{0,1,2,\ldots\}.

Proof(Theorem 7.1)

We repeatedly use the fact that under periodic boundary conditions 1VVd3x  ei(k+k)x=δk,k\frac{1}{V}\int_V d^3x\;e^{i(\boldsymbol{k}+\boldsymbol{k}')\cdot\boldsymbol{x}} = \delta_{\boldsymbol{k},-\boldsymbol{k}'}. First,

d3x  ϕ˙2=k,kq˙kq˙kδk,k=kq˙kq˙k=kq˙k2\int d^3x\;\dot\phi^2 = \sum_{\boldsymbol{k},\boldsymbol{k}'}\dot q_{\boldsymbol{k}}\dot q_{\boldsymbol{k}'}\,\delta_{\boldsymbol{k},-\boldsymbol{k}'} = \sum_{\boldsymbol{k}}\dot q_{\boldsymbol{k}}\dot q_{-\boldsymbol{k}} = \sum_{\boldsymbol{k}}\lvert\dot q_{\boldsymbol{k}}\rvert^2

(the last step uses the reality condition qk=qkq_{-\boldsymbol{k}}=\overline{q_{\boldsymbol{k}}}). Similarly, from eikx=ikeikx\nabla e^{i\boldsymbol{k}\cdot\boldsymbol{x}} = i\boldsymbol{k}\,e^{i\boldsymbol{k}\cdot\boldsymbol{x}},

d3x  (ϕ)2=k(ikqk)(ikqk)=kk2qk2,d3x  ϕ2=kqk2.\int d^3x\;(\nabla\phi)^2 = \sum_{\boldsymbol{k}}(i\boldsymbol{k}q_{\boldsymbol{k}})\cdot(-i\boldsymbol{k}q_{-\boldsymbol{k}}) = \sum_{\boldsymbol{k}}\boldsymbol{k}^2\lvert q_{\boldsymbol{k}}\rvert^2, \qquad \int d^3x\;\phi^2 = \sum_{\boldsymbol{k}}\lvert q_{\boldsymbol{k}}\rvert^2 .

Combining the three gives the asserted expression for LL, with k2+m2=ωk2\boldsymbol{k}^2 + m^2 = \omega_{\boldsymbol{k}}^2.

Next, the complex variables qkq_{\boldsymbol{k}} are not independent, so we pass to real degrees of freedom. Choose a half set K+K_+ containing exactly one of each pair k\boldsymbol{k}, k-\boldsymbol{k}, and for kK+\boldsymbol{k}\in K_+ write qk=(xk+iyk)/2q_{\boldsymbol{k}} = (x_{\boldsymbol{k}} + i y_{\boldsymbol{k}})/\sqrt2. Since qk=qk\lvert q_{-\boldsymbol{k}}\rvert = \lvert q_{\boldsymbol{k}}\rvert, the contributions of ±k\pm\boldsymbol{k} are equal, and

L=kK+12[(x˙k2+y˙k2)ωk2(xk2+yk2)]+12(q˙02m2q02).L = \sum_{\boldsymbol{k}\in K_+}\frac12\left[(\dot x_{\boldsymbol{k}}^2 + \dot y_{\boldsymbol{k}}^2) - \omega_{\boldsymbol{k}}^2(x_{\boldsymbol{k}}^2+y_{\boldsymbol{k}}^2)\right] + \frac12\left(\dot q_{\boldsymbol{0}}^2 - m^2 q_{\boldsymbol{0}}^2\right) .

This is an array of real harmonic oscillators of unit mass and angular frequency ωk\omega_{\boldsymbol{k}}, one per k\boldsymbol{k} (two for each element of K+K_+, and one for k=0\boldsymbol{k}=\boldsymbol{0}).

Passing each oscillator over to Hamiltonian mechanics, imposing [q^,p^]=i[\hat q,\hat p]=i and introducing the standard creation and annihilation operators yields H^osc=ω(a^a^+12)\hat H_{\text{osc}} = \omega(\hat a^\dagger\hat a + \tfrac12). Recombining the pairs xk,ykx_{\boldsymbol{k}},y_{\boldsymbol{k}} into complex combinations produces operators a^k\hat a_{\boldsymbol{k}} and a^k\hat a_{-\boldsymbol{k}} labelled by momenta ±k\pm\boldsymbol{k}, and the total Hamiltonian assembles into the asserted form. As for the momentum, rewriting P=d3xϕ˙ϕ\boldsymbol{P} = -\int d^3x\,\dot\phi\,\nabla\phi — obtained from Noether’s theorem (symmetries and conservation laws) — by the same procedure gives kka^ka^k\sum_{\boldsymbol{k}}\boldsymbol{k}\,\hat a^\dagger_{\boldsymbol{k}}\hat a_{\boldsymbol{k}} (the zero-point terms cancel by the symmetry kk\boldsymbol{k}\to-\boldsymbol{k}).

The eigenstates are products of number states of the individual modes, with eigenvalues kωk(nk+12)\sum_{\boldsymbol{k}}\omega_{\boldsymbol{k}}(n_{\boldsymbol{k}}+\tfrac12).

How this result is to be read is quantum field theory itself. By Theorem 7.1, the state a^k0\hat a^\dagger_{\boldsymbol{k}}\lvert 0\rangle obtained by acting once with a^k\hat a^\dagger_{\boldsymbol{k}} on the vacuum 0\lvert 0\rangle carries

E=ωk=k2+m2,P=k.E = \omega_{\boldsymbol{k}} = \sqrt{\boldsymbol{k}^2+m^2}, \qquad \boldsymbol{P} = \boldsymbol{k} .

These are exactly the energy and momentum of one relativistic particle of mass mm. Acting nn times gives nn particles. The picture that a particle is an excitation quantum of a field has emerged as a consequence of the computation.

Definition 7.2Fock space

The space spanned by all states obtained by acting with creation operators on the vacuum 0\lvert 0\rangle (which satisfies a^k0=0\hat a_{\boldsymbol{k}}\lvert 0\rangle = 0 for every k\boldsymbol{k}),

F=n=0H(n),H(n)=span{a^k1a^kn0},\mathcal{F} = \bigoplus_{n=0}^{\infty}\mathcal{H}^{(n)}, \qquad \mathcal{H}^{(n)} = \operatorname{span}\left\{\hat a^\dagger_{\boldsymbol{k}_1}\cdots \hat a^\dagger_{\boldsymbol{k}_n}\lvert 0\rangle\right\} ,

is called Fock space. The number operator is N^=ka^ka^k\hat N = \sum_{\boldsymbol{k}}\hat a^\dagger_{\boldsymbol{k}}\hat a_{\boldsymbol{k}}.

Fock space is a direct sum of subspaces with different particle numbers, so operators that change the particle number can be written down. The restriction (A1) has been lifted. Moreover a state is determined by the occupation numbers {nk}\{n_{\boldsymbol{k}}\} alone, so symmetry under exchange of particles holds automatically. There is no need to impose symmetrization for identical particles by hand.

Example 7.3Interactions do not conserve particle number

Add λ4!ϕ4-\frac{\lambda}{4!}\phi^4 to the Lagrangian. The corresponding interaction Hamiltonian is H^int=λ4!d3x  ϕ^4\hat H_{\text{int}} = \frac{\lambda}{4!}\int d^3x\;\hat\phi^4. The field operator is a sum of an annihilation part and a creation part, ϕ^a^+a^\hat\phi \sim \hat a + \hat a^\dagger, so expanding ϕ^4\hat\phi^4 produces 24=162^4 = 16 terms, which we may classify by the number of creation operators they contain.

ϕ^4a^a^a^a^ΔN=+4+a^a^a^a^ΔN=+2+a^a^a^a^ΔN=0+a^a^a^a^ΔN=2+a^a^a^a^ΔN=4\hat\phi^4 \sim \underbrace{\hat a^\dagger\hat a^\dagger\hat a^\dagger\hat a^\dagger}_{\Delta N = +4} + \underbrace{\hat a^\dagger\hat a^\dagger\hat a^\dagger\hat a}_{\Delta N=+2} + \underbrace{\hat a^\dagger\hat a^\dagger\hat a\hat a}_{\Delta N = 0} + \underbrace{\hat a^\dagger\hat a\hat a\hat a}_{\Delta N=-2} + \underbrace{\hat a\hat a\hat a\hat a}_{\Delta N=-4}

(each brace stands for all combinations, including rearrangements of the momenta). Since terms with ΔN=±4,±2\Delta N = \pm 4, \pm 2 are present, [H^,N^]0[\hat H,\hat N]\ne 0 and particle number is not conserved. A process in which four particles are born out of the vacuum and a process in which two particles scatter into two both come from the same single term. Only within the framework of Definition 7.2 do such processes become describable.

7.2. How negative energy and negative probability are resolved

Section titled “7.2. How negative energy and negative probability are resolved”

Subtracting the constant zero-point energy kωk/2\sum_{\boldsymbol{k}}\omega_{\boldsymbol{k}}/2 from the H^\hat H of Theorem 7.1 leaves H^=kωka^ka^k0\hat H = \sum_{\boldsymbol{k}}\omega_{\boldsymbol{k}}\hat a^\dagger_{\boldsymbol{k}}\hat a_{\boldsymbol{k}} \ge 0. Where has the negative energy gone? The answer is that the negative-frequency part e+iωte^{+i\omega t} of the solutions became not a negative-energy state but the coefficient of a creation operator. In the mode expansion both appear,

ϕ^(x)    a^keiωktpositive frequency  +  a^ke+iωktnegative frequency,\hat\phi(x) \;\sim\; \underbrace{\hat a_{\boldsymbol{k}}\,e^{-i\omega_{\boldsymbol{k}}t}}_{\text{positive frequency}} \;+\; \underbrace{\hat a^\dagger_{\boldsymbol{k}}\,e^{+i\omega_{\boldsymbol{k}}t}}_{\text{negative frequency}} ,

but a^k\hat a^\dagger_{\boldsymbol{k}} is the operator that raises the energy by ωk\omega_{\boldsymbol{k}}. The very same formula, read differently, yields a spectrum bounded below.

For a complex scalar field the current jμj^\mu of Proposition 3.2 becomes the charge current, and the conserved charge takes the form

Q^=k(a^ka^kb^kb^k).\hat Q = \sum_{\boldsymbol{k}}\left(\hat a^\dagger_{\boldsymbol{k}}\hat a_{\boldsymbol{k}} - \hat b^\dagger_{\boldsymbol{k}}\hat b_{\boldsymbol{k}}\right) .

The particles created by a^\hat a^\dagger and those created by b^\hat b^\dagger have the same mass and opposite charge; the latter are the antiparticles. So j0<0j^0 < 0 meant not negative probability but the existence of antiparticles, exactly as announced in Remark 3.4.

We turn finally to causality. In quantum field theory causality is implemented not as the vanishing of an amplitude but as the statement that field operators at spacelike-separated points commute. If they commute, a measurement at one point cannot influence the outcome of a measurement at the other.

Definition 7.4The free real scalar field operator

In the continuum limit (VV\to\infty) we write

ϕ^(x)=d3p(2π)3  12ωp(a^peipx+a^pe+ipx),p0=ωp.\hat\phi(x) = \int\frac{d^3p}{(2\pi)^3}\;\frac{1}{\sqrt{2\omega_{\boldsymbol{p}}}}\left(\hat a_{\boldsymbol{p}}\,e^{-ip\cdot x} + \hat a^\dagger_{\boldsymbol{p}}\,e^{+ip\cdot x}\right), \qquad p^0 = \omega_{\boldsymbol{p}} .

The commutation relations are [a^p,a^q]=(2π)3δ3(pq)[\hat a_{\boldsymbol{p}},\hat a^\dagger_{\boldsymbol{q}}] = (2\pi)^3\delta^3(\boldsymbol{p}-\boldsymbol{q}), all others being zero.

Theorem 7.5Microcausality

For the field of Definition 7.4, setting D(x):=d3p(2π)32ωpeipxD(x) := \displaystyle\int\frac{d^3p}{(2\pi)^3\,2\omega_{\boldsymbol{p}}}e^{-ip\cdot x}, we have

[ϕ^(x),ϕ^(y)]=D(xy)D(yx),[\hat\phi(x),\hat\phi(y)] = D(x-y) - D(y-x) ,

and this vanishes whenever (xy)2<0(x-y)^2 < 0, that is, at spacelike separation.

Proof(Theorem 7.5)

Step 1: the form of the commutator. Substituting Definition 7.4, the only surviving terms are those containing [a^p,a^q][\hat a_{\boldsymbol{p}},\hat a^\dagger_{\boldsymbol{q}}] and [a^p,a^q]=(2π)3δ3(pq)[\hat a^\dagger_{\boldsymbol{p}},\hat a_{\boldsymbol{q}}] = -(2\pi)^3\delta^3(\boldsymbol{p}-\boldsymbol{q}).

[ϕ^(x),ϕ^(y)]=d3pd3q(2π)64ωpωq([a^p,a^q]eipx+iqy+[a^p,a^q]eipxiqy)[\hat\phi(x),\hat\phi(y)] = \int\frac{d^3p\,d^3q}{(2\pi)^6\sqrt{4\omega_{\boldsymbol{p}}\omega_{\boldsymbol{q}}}} \left([\hat a_{\boldsymbol{p}},\hat a^\dagger_{\boldsymbol{q}}]e^{-ip\cdot x + iq\cdot y} + [\hat a^\dagger_{\boldsymbol{p}},\hat a_{\boldsymbol{q}}]e^{ip\cdot x - iq\cdot y}\right)=d3p(2π)32ωp(eip(xy)eip(xy))=D(xy)D(yx).= \int\frac{d^3p}{(2\pi)^3\,2\omega_{\boldsymbol{p}}}\left(e^{-ip\cdot(x-y)} - e^{ip\cdot(x-y)}\right) = D(x-y) - D(y-x).

Step 2: Lorentz invariance of the measure. The zeros of δ(p2m2)=δ ⁣((p0)2ωp2)\delta(p^2-m^2) = \delta\!\left((p^0)^2 - \omega_{\boldsymbol{p}}^2\right) are at p0=±ωpp^0 = \pm\omega_{\boldsymbol{p}}, where the absolute value of the derivative is 2ωp2\omega_{\boldsymbol{p}}, so

δ(p2m2)=12ωp[δ(p0ωp)+δ(p0+ωp)].\delta(p^2-m^2) = \frac{1}{2\omega_{\boldsymbol{p}}}\left[\delta(p^0-\omega_{\boldsymbol{p}}) + \delta(p^0+\omega_{\boldsymbol{p}})\right].

Consequently, for any function ff,

d4p(2π)3δ(p2m2)θ(p0)f(p)=d3p(2π)32ωpf(ωp,p).\int\frac{d^4p}{(2\pi)^3}\,\delta(p^2-m^2)\,\theta(p^0)\,f(p) = \int\frac{d^3p}{(2\pi)^3\,2\omega_{\boldsymbol{p}}}f(\omega_{\boldsymbol{p}},\boldsymbol{p}) .

On the left-hand side, d4pd^4p and δ(p2m2)\delta(p^2-m^2) are Lorentz invariant, and on the positive mass shell θ(p0)\theta(p^0) is invariant under proper orthochronous Lorentz transformations ΛSO+(1,3)\Lambda\in SO^+(1,3) as well, since such Λ\Lambda preserve the sign of p0p^0. Hence D(Λx)=D(x)D(\Lambda x) = D(x) for every ΛSO+(1,3)\Lambda\in SO^+(1,3).

Step 3: reversing a spacelike vector. Let z=xyz = x-y be spacelike, that is, z2=(z0)2z2<0z^2 = (z^0)^2 - \lvert\boldsymbol{z}\rvert^2 < 0. Then z>z0\lvert\boldsymbol{z}\rvert > \lvert z^0\rvert, so v=z0/zv = z^0/\lvert\boldsymbol{z}\rvert satisfies v<1\lvert v\rvert < 1. Applying a boost with velocity vv along the direction z^\hat{\boldsymbol{z}} gives

z0=γ(z0vz)=0,z'^0 = \gamma\left(z^0 - v\lvert\boldsymbol{z}\rvert\right) = 0 ,

so z=(0,z)z' = (0,\boldsymbol{z}') is an equal-time vector. Next, a rotation by angle π\pi about an axis orthogonal to z\boldsymbol{z}' sends zz\boldsymbol{z}'\to-\boldsymbol{z}', that is, zzz'\to -z'. Both the boost and the rotation lie in SO+(1,3)SO^+(1,3), so their composition Λ\Lambda satisfies Λz=z\Lambda z = -z.

By Step 2, D(z)=D(Λz)=D(z)D(z) = D(\Lambda z) = D(-z). Returning to Step 1 we obtain [ϕ^(x),ϕ^(y)]=D(z)D(z)=0[\hat\phi(x),\hat\phi(y)] = D(z)-D(-z) = 0.

Remark 7.6Causality demands antiparticles

What was essential in the proof of Theorem 7.5 is that two terms, D(xy)D(x-y) and D(yx)D(y-x), appeared with a relative minus sign. Indeed DD by itself does not vanish at spacelike separation. At equal times, xy=(0,r)x-y = (0,\boldsymbol{r}),

D(0,r)=14π2r0dp  psin(pr)p2+m2=m4π2rK1(mr)>0D(0,\boldsymbol{r}) = \frac{1}{4\pi^2 r}\int_0^\infty dp\;\frac{p\sin(pr)}{\sqrt{p^2+m^2}} = \frac{m}{4\pi^2 r}K_1(mr) > 0

(here K1K_1 is the modified Bessel function of the second kind, with K1(u)π/2ueuK_1(u) \sim \sqrt{\pi/2u}\,e^{-u}), so it has the same emre^{-mr} tail found in Example 5.2. It is only the difference that vanishes.

For a complex scalar field the meaning of that difference becomes clear. D(xy)D(x-y) is the amplitude for a particle created at yy to be annihilated at xx, while D(yx)D(y-x) is the amplitude for an antiparticle created at xx to be annihilated at yy. At spacelike separation the two events carry no invariant time ordering, so the two amplitudes have equal magnitude and cancel exactly. Without antiparticles there would be no cancellation and causality would fail. The existence of antiparticles is forced upon us by relativity, quantum mechanics and causality together (Exercise 9.4).

For fermions the same cancellation occurs in the anticommutator. That integer spin must be quantized with commutators and half-integer spin with anticommutators, on pain of losing either causality or positivity of the energy, is the spin-statistics theorem. Even the Pauli exclusion principle that the Dirac sea required becomes a theorem here.

IssueOne-particle relativistic quantum mechanicsQuantum field theory
Negative-energy solutionsThe Hamiltonian is unbounded below (Theorem 4.2)Negative-frequency coefficients become creation operators, so H^0\hat H\ge 0
Sign of the densityj0j^0 goes negative (Proposition 3.2)j0j^0 is a charge density; the two signs are particle and antiparticle
CausalityThe amplitude is nonzero even at spacelike separation (Example 5.2)The commutator vanishes at spacelike separation (Theorem 7.5)
Particle numberNo operator changing it can be writtenHandled naturally in Fock space (Example 7.3)
Identical particlesSymmetrization or antisymmetrization imposed by handAutomatically identical, being quanta of one and the same field
Spin and statisticsAn independent assumptionDerived, as the spin-statistics theorem

Remark 7.7On the name 'second quantization'

For historical reasons this procedure is called second quantization. The term originally referred to rewriting many-body quantum mechanics in terms of creation and annihilation operators, and only later came to be applied to the quantization of fields. The former is no more than a change of notation, available for nonrelativistic many-body systems as well; new physics enters when a relativistic field is quantized.

The name is also misleading. What is quantized is the classical field ϕ(x)\phi(x); it is not that an already quantized wave function is quantized a second time. Quantization happens once and once only (Weinberg emphasizes this point). Accept the term as a historical label, but understand its content as canonical quantization of a classical field.

In the chapters that follow we walk through the door opened here.

The computational techniques of perturbation theory carry over as the foundation. Treating the interaction term of Example 7.3 as a perturbation is where Feynman diagrams begin.

Exercise 9.1Standard

Use natural units =c=1\hbar=c=1 and the metric ημν=diag(+1,1,1,1)\eta_{\mu\nu} = \mathrm{diag}(+1,-1,-1,-1), and let ϕ\phi be a complex solution of the Klein–Gordon equation (+m2)ϕ=0(\Box+m^2)\phi=0.

  1. Write the spatial components j\boldsymbol{j} of the current jμ=i2m(ϕμϕ(μϕ)ϕ)j^\mu = \frac{i}{2m}\left(\phi^{*}\partial^\mu\phi - (\partial^\mu\phi^{*})\phi\right) in terms of ϕ\phi and ϕ\nabla\phi, and verify that they take the same form as the nonrelativistic probability current 12mi(ψψψψ)\frac{1}{2mi}(\psi^{*}\nabla\psi - \psi\nabla\psi^{*}).
  2. Compute j0j^0 for the superposition ϕ=N+eiωt+Ne+iωt\phi = N_+e^{-i\omega t} + N_-e^{+i\omega t} (with p=0\boldsymbol{p}=\boldsymbol{0}, ω=m\omega=m) and show that j0=0j^0 = 0 when N+=N\lvert N_+\rvert = \lvert N_-\rvert. What does this mean?
Solution

1. Since i=ηiii=/xi\partial^i = \eta^{ii}\partial_i = -\partial/\partial x^i (no sum over ii), the spatial components are

j=i2m(ϕ(ϕ)(ϕ)ϕ)=i2m(ϕϕ(ϕ)ϕ).\boldsymbol{j} = \frac{i}{2m}\left(\phi^{*}(-\nabla\phi) - (-\nabla\phi^{*})\phi\right) = -\frac{i}{2m}\left(\phi^{*}\nabla\phi - (\nabla\phi^{*})\phi\right) .

Because i/(2m)=1/(2mi)-i/(2m) = 1/(2mi),

j=12mi(ϕϕϕϕ),\boldsymbol{j} = \frac{1}{2mi}\left(\phi^{*}\nabla\phi - \phi\nabla\phi^{*}\right) ,

which has the same form as the nonrelativistic probability current. So the relativistic upgrade changed the form of the time component only.

2. Set ω=m\omega = m. Multiplying ϕ˙=im(N+eimtNe+imt)\dot\phi = -im\left(N_+e^{-imt} - N_-e^{+imt}\right) by ϕ=N+e+imt+Neimt\phi^{*} = \overline{N_+}e^{+imt} + \overline{N_-}e^{-imt} and writing w:=NN+e2imtw := \overline{N_-}N_+e^{-2imt} (so that N+Ne2imt=wˉ\overline{N_+}N_-e^{2imt} = \bar w), we get

ϕϕ˙=im[(N+2N2)+(wwˉ)]=im(N+2N2)+2mImw\phi^{*}\dot\phi = -im\left[\left(\lvert N_+\rvert^2-\lvert N_-\rvert^2\right) + (w-\bar w)\right] = -im\left(\lvert N_+\rvert^2-\lvert N_-\rvert^2\right) + 2m\operatorname{Im}w

(using wwˉ=2iImww-\bar w = 2i\operatorname{Im}w). The second term is real, so it cancels when we subtract ϕ˙ϕ=ϕϕ˙\dot\phi^{*}\phi = \overline{\phi^{*}\dot\phi}, leaving

j0=i2m(ϕϕ˙ϕ˙ϕ)=i2m(2im)(N+2N2)=N+2N2.j^0 = \frac{i}{2m}\left(\phi^{*}\dot\phi - \dot\phi^{*}\phi\right) = \frac{i}{2m}\cdot\left(-2im\right)\left(\lvert N_+\rvert^2-\lvert N_-\rvert^2\right) = \lvert N_+\rvert^2 - \lvert N_-\rvert^2 .

Hence j0=0j^0 = 0 when N+=N\lvert N_+\rvert=\lvert N_-\rvert.

The meaning: in a state containing equal amounts of positive- and negative-energy components, reading j0j^0 as the probability density for finding the particle there gives the absurd answer zero identically. As explained in Remark 3.4, j0j^0 is a charge density, and the correct reading is that equal amounts of positive and negative charge are present, making the state neutral.

Exercise 9.2Standard

Let α1,α2,α3,β\alpha^1,\alpha^2,\alpha^3,\beta be Hermitian N×NN\times N matrices satisfying {αi,αj}=2δij1\{\alpha^i,\alpha^j\}=2\delta^{ij}\mathbb{1}, {αi,β}=0\{\alpha^i,\beta\}=0 and β2=1\beta^2=\mathbb{1}.

  1. Show that NN is even.
  2. Show that N=2N=2 is impossible, and conclude that the minimal dimension is 44.
Solution

1. As shown in Theorem 4.2, trβ=0\operatorname{tr}\beta = 0. On the other hand β2=1\beta^2=\mathbb{1} and β\beta is Hermitian, so β\beta is diagonalizable with eigenvalues ±1\pm1 only. Writing n±n_\pm for the multiplicities, n++n=Nn_++n_-=N and trβ=n+n=0\operatorname{tr}\beta = n_+-n_-=0, whence n+=n=N/2n_+=n_-=N/2. Since n±n_\pm are integers, NN is even.

2. Suppose N=2N=2. Every traceless Hermitian 2×22\times2 matrix can be written uniquely as a real linear combination of σ1,σ2,σ3\sigma^1,\sigma^2,\sigma^3 (this space is real three-dimensional). By Theorem 4.2 both αi\alpha^i and β\beta are traceless, so there are real vectors a1,a2,a3,bR3\boldsymbol{a}_1,\boldsymbol{a}_2,\boldsymbol{a}_3,\boldsymbol{b}\in\mathbb{R}^3 with

αi=aiσ,β=bσ.\alpha^i = \boldsymbol{a}_i\cdot\boldsymbol{\sigma},\qquad \beta = \boldsymbol{b}\cdot\boldsymbol{\sigma} .

Using the Pauli identity (uσ)(vσ)=(uv)1+i(u×v)σ(\boldsymbol{u}\cdot\boldsymbol{\sigma})(\boldsymbol{v}\cdot\boldsymbol{\sigma}) = (\boldsymbol{u}\cdot\boldsymbol{v})\mathbb{1} + i(\boldsymbol{u}\times\boldsymbol{v})\cdot\boldsymbol{\sigma} and the antisymmetry of u×v\boldsymbol{u}\times\boldsymbol{v}, we get

{uσ,vσ}=2(uv)1.\{\boldsymbol{u}\cdot\boldsymbol{\sigma},\,\boldsymbol{v}\cdot\boldsymbol{\sigma}\} = 2(\boldsymbol{u}\cdot\boldsymbol{v})\,\mathbb{1} .

The three hypotheses are therefore equivalent to

aiaj=δij,aib=0,bb=1,\boldsymbol{a}_i\cdot\boldsymbol{a}_j = \delta_{ij},\qquad \boldsymbol{a}_i\cdot\boldsymbol{b} = 0,\qquad \boldsymbol{b}\cdot\boldsymbol{b} = 1 ,

that is, to the existence of four mutually orthogonal unit vectors inside R3\mathbb{R}^3. Since R3\mathbb{R}^3 has dimension 33, at most three mutually orthogonal nonzero vectors exist, a contradiction. Hence N2N\ne 2.

By part 1, NN is even; N=2N=2 has just been excluded, and N=0N=0 is meaningless for matrices, so N4N\ge 4. And N=4N=4 is actually realized: the Dirac representation αi=(0σiσi0)\alpha^i = \begin{pmatrix}0&\sigma^i\\\sigma^i&0\end{pmatrix}, β=(1001)\beta=\begin{pmatrix}\mathbb{1}&0\\0&-\mathbb{1}\end{pmatrix} satisfies the three conditions, as one checks block by block using {σi,σj}=2δij\{\sigma^i,\sigma^j\}=2\delta^{ij}. The minimal dimension is therefore 44.

Exercise 9.3Easy

  1. Find the ratio of the electron’s reduced Compton wavelength λˉC=/(mec)\bar\lambda_C = \hbar/(m_ec) to the Bohr radius a0=/(mecα)a_0 = \hbar/(m_ec\alpha), and give the numbers (take α1=137.04\alpha^{-1}=137.04).
  2. Express the typical speed and kinetic energy of an electron in a hydrogen atom in terms of α\alpha and mec2=511 keVm_ec^2 = 511\ \mathrm{keV}, and explain why nonrelativistic quantum mechanics is a good approximation.
  3. What happens if one tries to confine an electron to a region of size about λˉC\bar\lambda_C?
Solution

1. Straight from the definitions, a0/λˉC=1/α=137.04a_0/\bar\lambda_C = 1/\alpha = 137.04. Using λˉC=3.86×1013 m\bar\lambda_C = 3.86\times10^{-13}\ \mathrm{m} from Example 6.2,

a0=137.04×3.86×1013 m=5.29×1011 m,a_0 = 137.04\times 3.86\times10^{-13}\ \mathrm{m} = 5.29\times10^{-11}\ \mathrm{m} ,

which agrees with the known Bohr radius.

2. The electron is confined to a spread of order a0a_0, so by the uncertainty relation its typical momentum is p/a0=αmecp\sim\hbar/a_0 = \alpha\, m_ec. Hence

vcα1137,Ekinp22me=12α2mec2=5.325×1052×511 keV=13.6 eV.\frac{v}{c}\sim\alpha\simeq\frac{1}{137}, \qquad E_{\text{kin}}\sim\frac{p^2}{2m_e} = \frac{1}{2}\alpha^2 m_ec^2 = \frac{5.325\times10^{-5}}{2}\times 511\ \mathrm{keV} = 13.6\ \mathrm{eV}.

This is exactly the Rydberg energy of hydrogen. The kinetic energy is only 2.7×1052.7\times10^{-5} times the rest energy, and relativistic corrections enter at relative accuracy (v/c)2α2(v/c)^2\sim\alpha^2, an effect of order 10510^{-5} (this is the size of the fine structure). Nonrelativistic quantum mechanics is therefore a good approximation. That quantum mechanics alone almost suffices for atomic physics is thanks to the smallness of the fine-structure constant.

3. With ΔxλˉC\Delta x\sim\bar\lambda_C we get Δp/λˉC=mec\Delta p\gtrsim\hbar/\bar\lambda_C = m_ec, hence ΔEmec2=511 keV\Delta E\gtrsim m_ec^2 = 511\ \mathrm{keV}. This is the same order as the pair-creation threshold 2mec2=1.022 MeV2m_ec^2 = 1.022\ \mathrm{MeV}, so the energy used for the confinement creates electron–positron pairs. Since the original electron and the newly created one cannot be told apart, the very description in terms of a one-particle wave function breaks down. One needs a space containing states of different particle number, as in Definition 7.2.

Exercise 9.4Hard

Take the complex scalar field operator to be

ϕ^(x)=d3p(2π)312ωp(a^peipx+b^pe+ipx)\hat\phi(x) = \int\frac{d^3p}{(2\pi)^3}\frac{1}{\sqrt{2\omega_{\boldsymbol{p}}}}\left(\hat a_{\boldsymbol{p}}e^{-ip\cdot x} + \hat b^\dagger_{\boldsymbol{p}}e^{+ip\cdot x}\right)

with [a^p,a^q]=[b^p,b^q]=(2π)3δ3(pq)[\hat a_{\boldsymbol{p}},\hat a^\dagger_{\boldsymbol{q}}] = [\hat b_{\boldsymbol{p}},\hat b^\dagger_{\boldsymbol{q}}] = (2\pi)^3\delta^3(\boldsymbol{p}-\boldsymbol{q}) and all commutators between a^\hat a-type and b^\hat b-type operators vanishing.

  1. Show that [ϕ^(x),ϕ^(y)]=D(xy)D(yx)[\hat\phi(x),\hat\phi^\dagger(y)] = D(x-y) - D(y-x) and conclude that it vanishes at spacelike separation (here DD is as in Theorem 7.5).
  2. State what each of D(xy)D(x-y) and D(yx)D(y-x) is the amplitude for, physically, and explain the meaning of the cancellation.
  3. What breaks if antiparticles are assumed not to exist, that is, if the b^\hat b^\dagger term is absent?
Solution

1. We have ϕ^(y)=d3q(2π)312ωq(a^qe+iqy+b^qeiqy)\hat\phi^\dagger(y) = \int\frac{d^3q}{(2\pi)^3}\frac{1}{\sqrt{2\omega_{\boldsymbol{q}}}}\left(\hat a^\dagger_{\boldsymbol{q}}e^{+iq\cdot y} + \hat b_{\boldsymbol{q}}e^{-iq\cdot y}\right). Expanding the commutator, the cross terms between a^\hat a-type and b^\hat b-type operators vanish, so

[ϕ^(x),ϕ^(y)]=d3pd3q(2π)64ωpωq([a^p,a^q]eipx+iqy+[b^p,b^q]e+ipxiqy)[\hat\phi(x),\hat\phi^\dagger(y)] = \int\frac{d^3p\,d^3q}{(2\pi)^6\sqrt{4\omega_{\boldsymbol{p}}\omega_{\boldsymbol{q}}}}\left([\hat a_{\boldsymbol{p}},\hat a^\dagger_{\boldsymbol{q}}]e^{-ip\cdot x+iq\cdot y} + [\hat b^\dagger_{\boldsymbol{p}},\hat b_{\boldsymbol{q}}]e^{+ip\cdot x - iq\cdot y}\right)=d3p(2π)32ωp(eip(xy)e+ip(xy))=D(xy)D(yx).= \int\frac{d^3p}{(2\pi)^3 2\omega_{\boldsymbol{p}}}\left(e^{-ip\cdot(x-y)} - e^{+ip\cdot(x-y)}\right) = D(x-y)-D(y-x).

The sign of the second term comes from [b^p,b^q]=(2π)3δ3(pq)[\hat b^\dagger_{\boldsymbol{p}},\hat b_{\boldsymbol{q}}] = -(2\pi)^3\delta^3(\boldsymbol{p}-\boldsymbol{q}). From here Steps 2 and 3 of Theorem 7.5 apply verbatim: if (xy)2<0(x-y)^2<0 there is a ΛSO+(1,3)\Lambda\in SO^+(1,3) with Λ(xy)=(xy)\Lambda(x-y) = -(x-y), and the invariance of DD gives D(xy)=D(yx)D(x-y)=D(y-x), so the commutator vanishes.

2. Taking vacuum expectation values, 0ϕ^(x)ϕ^(y)0=D(xy)\langle 0\rvert\hat\phi(x)\hat\phi^\dagger(y)\lvert 0\rangle = D(x-y). Here ϕ^(y)\hat\phi^\dagger(y) creates one particle (of a^\hat a type) at yy and ϕ^(x)\hat\phi(x) annihilates it at xx, so D(xy)D(x-y) is the amplitude for a particle born at yy to disappear at xx. On the other hand 0ϕ^(y)ϕ^(x)0=D(yx)\langle 0\rvert\hat\phi^\dagger(y)\hat\phi(x)\lvert 0\rangle = D(y-x), where ϕ^(x)\hat\phi(x) creates an antiparticle (of b^\hat b type) at xx and ϕ^(y)\hat\phi^\dagger(y) annihilates it at yy. That is the amplitude for an antiparticle born at xx to disappear at yy.

Two spacelike-separated points admit no Lorentz-invariant time ordering. An observer who sees yy first sees a particle propagating from yy to xx; an observer who sees xx first sees an antiparticle propagating from xx to yy. The two amplitudes are equal and are subtracted in the commutator, so they cancel. This is the mechanism by which no observable causal influence survives.

3. Without the b^\hat b^\dagger term we would have ϕ^(x)=d3p(2π)32ωpa^peipx\hat\phi(x) = \int\frac{d^3p}{(2\pi)^3\sqrt{2\omega_{\boldsymbol p}}}\hat a_{\boldsymbol{p}}e^{-ip\cdot x} and hence [ϕ^(x),ϕ^(y)]=D(xy)[\hat\phi(x),\hat\phi^\dagger(y)] = D(x-y). As noted in Remark 7.6, at spacelike separation DD equals m4π2rK1(mr)>0\frac{m}{4\pi^2r}K_1(mr)>0 and does not vanish. Microcausality would therefore fail, and two spacelike-separated measurements would influence one another.

In short, building a relativistic local field requires both annihilation and creation operators inside the same ϕ^(x)\hat\phi(x). When the field carries charge (ϕ^ϕ^\hat\phi\ne\hat\phi^\dagger), the latter is necessarily the creation operator for a particle of the opposite charge, that is, for an antiparticle. For a neutral real scalar field (Definition 7.4) one has b^=a^\hat b = \hat a, and the particle is its own antiparticle.

  • M. E. Peskin and D. V. Schroeder, An Introduction to Quantum Field Theory, Addison-Wesley, 1995 — Chapter 1 and §2.1 (the failure of one-particle relativistic quantum mechanics and propagation outside the light cone).
  • S. Weinberg, The Quantum Theory of Fields, Volume I: Foundations, Cambridge University Press, 1995 — Chapter 1 (historical introduction), Chapter 5 (the necessity of fields and spin-statistics).
  • M. Sakamoto, Ba no Ryōshiron: Fuhensei to Jiyūba o Chūshin ni shite, Shōkabō, 2014 (in Japanese) — Chapters 1–3 (difficulties of the one-particle theory, canonical quantization).
  • P. A. M. Dirac, “The Quantum Theory of the Electron”, Proceedings of the Royal Society A 117 (1928), 610–624. DOI: 10.1098/rspa.1928.0023
  • C. D. Anderson, “The Positive Electron”, Physical Review 43 (1933), 491–494. DOI: 10.1103/PhysRev.43.491
  • G. C. Hegerfeldt, “Remark on causality and particle localization”, Physical Review D 10 (1974), 3320–3321. DOI: 10.1103/PhysRevD.10.3320

Appendix: Contour deformation for the propagation amplitude outside the light cone

Section titled “Appendix: Contour deformation for the propagation amplitude outside the light cone”

Let us derive the expression used in Example 5.2. The starting point is

U(t,r)=12π2r0dp  psin(pr)eiωpt,ωp=p2+m2.U(t,r) = \frac{1}{2\pi^2 r}\int_0^\infty dp\;p\,\sin(pr)\,e^{-i\omega_p t}, \qquad \omega_p = \sqrt{p^2+m^2} .

Both psin(pr)p\sin(pr) and ωp\omega_p are even functions of pp, so we may extend the range of integration to all of R\mathbb{R} and multiply by 12\frac12. Writing sin(pr)=(eipreipr)/(2i)\sin(pr) = (e^{ipr}-e^{-ipr})/(2i) and substituting ppp\to-p in the eipre^{-ipr} term, which turns it into a copy of the eipre^{ipr} term, we obtain

U(t,r)=14π2irdp  peipreip2+m2t.U(t,r) = \frac{1}{4\pi^2 i\,r}\int_{-\infty}^{\infty}dp\;p\,e^{ipr}\,e^{-i\sqrt{p^2+m^2}\,t} .

Now continue p2+m2\sqrt{p^2+m^2} analytically into the complex pp plane. The branch points are p=±imp=\pm im; take the cuts along [im,i)[im, i\infty) and (i,im](-i\infty,-im], and choose the branch with p2+m2>0\sqrt{p^2+m^2}>0 on the real axis.

Since r>0r>0, the factor eipre^{ipr} decays in the upper half plane. On a large upper semicircle of radius RR we have p2+m2±p\sqrt{p^2+m^2}\simeq \pm p as p\lvert p\rvert\to\infty, so the exponential part of the integrand behaves like eip(rt)=e(rt)Imp\lvert e^{ip(r-t)}\rvert = e^{-(r-t)\operatorname{Im}p}. Only for a spacelike separation r>t>0r>t>0 does this decay and the arc contribution vanish. This is the decisive point: for a timelike separation r<tr<t the deformation below is not available.

Pushing the contour upward, it catches on the cut [im,i)[im,i\infty). What remains is a contour that descends the left side of the cut from ii\infty to imim, rounds imim, and ascends the right side back to ii\infty. Setting p=iρp = i\rho with ρ>m\rho>m, we have dp=idρdp = i\,d\rho and eipr=eρre^{ipr} = e^{-\rho r}. Writing p=ϵ+iρp=\epsilon+i\rho and inspecting p2+m2(m2ρ2)+2iϵρp^2+m^2 \simeq (m^2-\rho^2) + 2i\epsilon\rho, we see that for ρ>m\rho>m the real part is negative while the sign of the imaginary part matches that of ϵ\epsilon. Hence, with s:=ρ2m2s:=\sqrt{\rho^2-m^2}, on the right side of the cut (ϵ>0\epsilon>0) we get p2+m2=+is\sqrt{p^2+m^2} = +is and on the left side (ϵ<0\epsilon<0) we get is-is. Therefore

eip2+m2t={e+ts(right side)ets(left side)e^{-i\sqrt{p^2+m^2}\,t} = \begin{cases} e^{+ts} & (\text{right side}) \\ e^{-ts} & (\text{left side}) \end{cases}

Adding the contributions of the left side (ρ:m\rho:\infty\to m) and the right side (ρ:m\rho: m\to\infty),

dp  peipreiωpt=m(iρ)eρretsidρ+m(iρ)eρre+tsidρ\int_{-\infty}^{\infty}dp\;p\,e^{ipr}e^{-i\omega_pt} = \int_{\infty}^{m}(i\rho)e^{-\rho r}e^{-ts}\,i\,d\rho + \int_{m}^{\infty}(i\rho)e^{-\rho r}e^{+ts}\,i\,d\rho =mdρ  ρeρr(etse+ts)=2mdρ  ρeρrsinh(ts).= \int_m^\infty d\rho\;\rho\,e^{-\rho r}\left(e^{-ts} - e^{+ts}\right) = -2\int_m^\infty d\rho\;\rho\,e^{-\rho r}\sinh(ts).

Substituting this back into the expression for UU and using 2/(4i)=i/2-2/(4i) = i/2, we obtain

U(t,r)=i2π2rmdρ  ρeρrsinh ⁣(tρ2m2).U(t,r) = \frac{i}{2\pi^2 r}\int_m^\infty d\rho\;\rho\,e^{-\rho r}\,\sinh\!\left(t\sqrt{\rho^2-m^2}\right) .

As ρ\rho\to\infty the integrand behaves like ρe(rt)ρ/2\rho\,e^{-(r-t)\rho}/2, so the integral converges for r>tr>t. For t>0t>0 the integrand is strictly positive throughout the range, so the integral is positive, that is, U0U\ne 0. As t0t\to 0 we have sinh0\sinh\to0 and hence U0U\to0, consistent with the value of δ3(r)\delta^3(\boldsymbol{r}) for r>0r>0.

Report an error in this article ・Operated by: Mugen Giken LLCPricingTermsLegal notice

© 2026 夢現技研合同会社 ・Feeding the text to an LLM is welcome. Code samples are MIT licensed.