# Why We Need Quantum Field Theory: One-Particle Quantum Mechanics Is Incompatible with Relativity

> Relativity grafted onto one-particle quantum mechanics yields negative energies, a negative probability density and acausal propagation. Quantizing the field is the only way out.
> https://rikai.mugen-giken.com/en/physics/qft/why-quantum-field-theory

## 0. Key points

- Nonrelativistic quantum mechanics is built on the Hilbert space $L^2(\mathbb{R}^{3N})$ with the particle number $N$ held fixed. The framework contains no place in which to write the creation or annihilation of a particle.
- Substituting operators naively into the relativistic dispersion relation $E^2 = \boldsymbol{p}^2c^2 + m^2c^4$ produces the Klein–Gordon equation, but negative-energy solutions inevitably appear and the time component of the conserved current can be negative. It cannot be read as a probability density.
- The Dirac equation restores positivity of the probability density, yet the negative-energy solutions remain. The "Dirac sea" is in effect a many-body theory carrying infinitely many particles, and it is unusable for bosons.
- Causality is the more serious failure. One-particle time evolution generated by $H = \sqrt{-\nabla^2 + m^2}$ destroys compact support instantaneously, and the amplitude to propagate to a spacelike-separated point does not vanish either.
- There is exactly one way out: quantize the field $\phi(x)$. A free field becomes a collection of independent harmonic oscillators, one per mode, and their excitation quanta appear as particles obeying the dispersion relation $E = \sqrt{\boldsymbol{p}^2 + m^2}$.
- Causality is implemented not as "the amplitude vanishes" but as "the commutator vanishes at spacelike separation". This cancellation requires antiparticles, so the existence of antiparticles is a consequence of relativity together with quantum mechanics.

## 1. Motivation: where does the photon come from?

Let us begin with a naive question. When an excited atom drops to its ground state, one photon is emitted. Before the emission no photon exists; afterwards one does. In the framework of [the Schrödinger equation and the wave function](/en/physics/quantum-mechanics/schrodinger-equation), however, a state is a function $\psi(\boldsymbol{x}_1,\ldots,\boldsymbol{x}_N,t)$ of the coordinates of $N$ particles, and $N$ is fixed the moment the equation is written down. There is nowhere at all to write a process in which the number of particles increases by one.

The same thing happens in the annihilation of an electron and a positron, $e^- e^+ \to \gamma\gamma$. The initial state consists of two massive particles, the final state of two massless ones. Not only the number of particles changes, but their species as well.

Historically, this difficulty was recognized almost as soon as quantum mechanics was completed. The first wave equation Schrödinger wrote down, at the end of 1925, was a relativistic one — what we now call the Klein–Gordon equation — but it disagreed with experiment on the fine structure of hydrogen, so he abandoned it and published the nonrelativistic equation instead. The same equation, rediscovered independently by Klein, Gordon, Fock and others in 1926, carried with it the disease that its probability density can go negative.

Dirac, meanwhile, quantized the electromagnetic field itself in a paper of 1927 and computed the emission and absorption of light by atoms from first principles. This was the first quantum field theory. The electromagnetic field is a field already classically, and quantizing it produced a particle: the photon. A natural question follows.

> Might the electron, too, be the quantum of some field?

The answer is yes. To reach that conclusion, however, we need to know precisely **how** one-particle relativistic quantum mechanics fails. Below we exhibit the failure in three layers — negative energy, the probabilistic interpretation, and causality — and see that quantizing the field cures all of them at once.

<Figure caption="How one-particle relativistic quantum mechanics breaks down, and the route to field quantization">
<Mermaid code={`flowchart TD
  A["Nonrelativistic quantum mechanics<br/>particle number is conserved"] --> B["Substitute the relativistic dispersion relation"]
  B --> C["Klein-Gordon equation"]
  B --> D["Dirac equation"]
  C --> E["Negative-energy solutions<br/>probability density goes negative"]
  D --> F["Negative-energy solutions<br/>the Dirac sea fails for bosons"]
  E --> G["Propagation amplitude nonzero even at spacelike separation"]
  F --> G
  G --> H["Quantize the field itself"]
  H --> I["Particle = excitation quantum of a field<br/>antiparticles, creation and annihilation, microcausality"]`} />
</Figure>

## 2. Preliminaries: two assumptions hidden in nonrelativistic quantum mechanics

Unless stated otherwise we use natural units $\hbar = c = 1$ and the metric $\eta_{\mu\nu} = \mathrm{diag}(+1,-1,-1,-1)$. A spacetime point is $x = (t,\boldsymbol{x})$, the inner product is $p\cdot x = p^0 t - \boldsymbol{p}\cdot\boldsymbol{x}$, and the on-shell energy is written $\omega_{\boldsymbol{p}} = \sqrt{\boldsymbol{p}^2 + m^2}$. We write $\Box = \partial_\mu\partial^\mu = \partial_t^2 - \nabla^2$.

Nonrelativistic quantum mechanics tacitly assumes the following two things. Both breakdowns below originate here.

**(A1) The particle number is fixed.** The state space is $\mathcal{H}_N = L^2(\mathbb{R}^{3N})$ (for identical particles, its symmetrized or antisymmetrized subspace). The Hamiltonian maps $\mathcal{H}_N$ into $\mathcal{H}_N$, so the particle number $N$ is conserved **by definition**. We have not proved a conservation law; we have merely chosen a space in which a situation with non-conservation cannot be written down.

**(A2) Time and space are treated asymmetrically.** The position $\hat{\boldsymbol{x}}$ is an operator, whereas the time $t$ is not an operator but a mere parameter. The canonical commutation relation reads $[\hat{x}^i, \hat{p}^j] = i\delta^{ij}$, and $t$ does not appear in it. But a [Lorentz transformation](/en/physics/relativity/lorentz-transformations) mixes $t$ with $\boldsymbol{x}$. Since a quantity that is an operator and a quantity that is a parameter get mixed by the transformation, this distinction is not Lorentz invariant.

<Remark id="rem-pauli-time-operator" title="The route through a time operator is blocked">
One might think that (A2) can be repaired by promoting $t$ to an operator. Suppose, however, that a self-adjoint $\hat T$ with $[\hat T,\hat H] = i$ exists. Since $[\hat T,\hat H]$ is a c-number, the Baker–Campbell–Hausdorff expansion terminates at first order:

$$
e^{i\epsilon \hat T}\hat H e^{-i\epsilon \hat T} = \hat H + i\epsilon[\hat T,\hat H] = \hat H - \epsilon .
$$

The left-hand side is a unitary transformation and therefore leaves the spectrum unchanged. Hence the spectrum of $\hat H$ is invariant under translation by an arbitrary real number $\epsilon$, which forces it to be all of $\mathbb{R}$. For a system whose energy is bounded below this cannot hold (Pauli's argument). The direction to take is therefore the opposite one: **demote** $\hat{\boldsymbol{x}}$ from being an operator, and let both entries of $x = (t,\boldsymbol{x})$ be mere labels. What becomes an operator instead is the **field** $\phi(x)$, which takes a value at each spacetime point.
</Remark>

## 3. The naive relativistic upgrade: two diseases of the Klein–Gordon equation

In the nonrelativistic case we obtained the Schrödinger equation by substituting $E \to i\partial_t$ and $\boldsymbol{p}\to -i\nabla$ into $E = \boldsymbol{p}^2/2m$. Let us apply the same substitution to the relativistic relation $E^2 = \boldsymbol{p}^2 + m^2$ (see [relativistic mechanics](/en/physics/relativity/relativistic-mechanics)).

<Definition id="def-klein-gordon" title="The Klein–Gordon equation">
For a complex scalar field of mass $m > 0$ — or a candidate one-particle wave function — $\phi(t,\boldsymbol{x})$, the equation

$$
\left(\Box + m^2\right)\phi = 0,
\qquad \Box = \partial_t^2 - \nabla^2
$$

is called the Klein–Gordon equation.
</Definition>

### Disease 1: the negative-energy solutions cannot be discarded

Substituting the plane wave $\phi = N e^{-i(Et - \boldsymbol{p}\cdot\boldsymbol{x})}$ gives $-E^2 + \boldsymbol{p}^2 + m^2 = 0$, that is, the two branches

$$
E = \pm\,\omega_{\boldsymbol{p}} = \pm\sqrt{\boldsymbol{p}^2 + m^2} .
$$

The negative branch means that the energy is not bounded below: perturb the system and it can keep falling into ever lower states.

One is tempted to say that the negative-energy solutions are unphysical and may simply be thrown away. They cannot be. The reason lies in the structure of the initial value problem. The Klein–Gordon equation is second order in time, so the Cauchy data consist of **two** arbitrary functions, $\phi(0,\boldsymbol{x})$ and $\dot\phi(0,\boldsymbol{x})$. A spatial Fourier transform gives $\ddot{\tilde\phi} = -\omega_{\boldsymbol p}^2\tilde\phi$ for each $\boldsymbol{p}$, whose general solution is

$$
\tilde\phi(t,\boldsymbol{p}) = a(\boldsymbol{p})e^{-i\omega_{\boldsymbol p}t} + b(\boldsymbol{p})e^{+i\omega_{\boldsymbol p}t} .
$$

Restricting to positive frequencies — setting $b \equiv 0$ — is nothing other than imposing the **constraint** $\dot{\tilde\phi}(0,\boldsymbol{p}) = -i\omega_{\boldsymbol{p}}\tilde\phi(0,\boldsymbol{p})$. In position space this is the nonlocal relation $\dot\phi(0,\cdot) = -i\sqrt{-\nabla^2+m^2}\,\phi(0,\cdot)$, so we have surrendered the right to prescribe the initial data freely and locally. What this nonlocality destroys is the subject of <Ref to="prop-instant-delocalization" />.

### Disease 2: the time component of the conserved current goes negative

For the Schrödinger equation, $\rho = \lvert\psi\rvert^2 \ge 0$ satisfies a continuity equation and can be read as a probability density. The Klein–Gordon equation also admits a conserved current, but positivity is lost.

<Proposition id="prop-kg-current" title="The Klein–Gordon current and its failure of positivity">
Let $\phi$ be a solution of <Ref to="def-klein-gordon" />. Then

$$
j^\mu = \frac{i}{2m}\left(\phi^{*}\partial^\mu\phi - (\partial^\mu\phi^{*})\phi\right)
$$

satisfies $\partial_\mu j^\mu = 0$. Moreover, for a plane-wave solution $\phi = Ne^{-i(Et-\boldsymbol{p}\cdot\boldsymbol{x})}$ with $E = \pm\omega_{\boldsymbol p}$,

$$
j^0 = \frac{E}{m}\lvert N\rvert^2 .
$$

Hence $j^0 < 0$ for the negative-energy solutions, and $j^0$ cannot be interpreted as a probability density.
</Proposition>

<Proof of="prop-kg-current">
Let us first verify conservation. By the product rule,

$$
\partial_\mu j^\mu = \frac{i}{2m}\left(\partial_\mu\phi^{*}\partial^\mu\phi + \phi^{*}\Box\phi - (\Box\phi^{*})\phi - \partial_\mu\phi^{*}\partial^\mu\phi\right)
= \frac{i}{2m}\left(\phi^{*}\Box\phi - (\Box\phi^{*})\phi\right).
$$

The first and fourth terms have cancelled. Now <Ref to="def-klein-gordon" /> gives $\Box\phi = -m^2\phi$ and, on taking complex conjugates, $\Box\phi^{*} = -m^2\phi^{*}$, so

$$
\partial_\mu j^\mu = \frac{i}{2m}\left(-m^2\phi^{*}\phi + m^2\phi^{*}\phi\right) = 0 .
$$

Next we compute $j^0$. With our metric convention $\partial^0 = \eta^{00}\partial_0 = \partial_t$, so $j^0 = \frac{i}{2m}(\phi^{*}\dot\phi - \dot\phi^{*}\phi)$. For a plane wave $\dot\phi = -iE\phi$ and $\dot\phi^{*} = +iE\phi^{*}$, whence

$$
j^0 = \frac{i}{2m}\left(\lvert N\rvert^2(-iE) - (iE)\lvert N\rvert^2\right) = \frac{i}{2m}\left(-2iE\lvert N\rvert^2\right) = \frac{E}{m}\lvert N\rvert^2 .
$$

On the branch $E = -\omega_{\boldsymbol p} < 0$ the right-hand side is negative. A density that goes negative is not a probability density.
</Proof>

<Example id="ex-nonrelativistic-limit" title="The nonrelativistic limit recovers the correct probability density">
It is not that $j^0$ is entirely meaningless. Extract the rest energy by writing $\phi(t,\boldsymbol{x}) = e^{-imt}\psi(t,\boldsymbol{x})$ and assume the nonrelativistic limit $\lvert\dot\psi\rvert \ll m\lvert\psi\rvert$. Then

$$
\dot\phi = e^{-imt}\left(\dot\psi - im\psi\right) \simeq -im\,e^{-imt}\psi ,
$$

so substituting into the expression for $j^0$ in <Ref to="prop-kg-current" /> gives

$$
j^0 \simeq \frac{i}{2m}\left(\bar\psi\,(-im\psi) - (im\bar\psi)\,\psi\right) = \frac{i}{2m}\left(-2im\lvert\psi\rvert^2\right) = \lvert\psi\rvert^2 ,
$$

which is exactly the probability density of the Schrödinger equation. In other words, $j^0$ is a quantity that *looks* like a probability density in the nonrelativistic limit and changes sign in the relativistic regime.
</Example>

<Remark id="rem-charge-not-probability" title="The disease was a misreading">
Let us give the game away in advance: $j^\mu$ itself is a perfectly good conserved quantity. What was wrong was the reading of it as a probability density. In quantum field theory $e\,j^\mu$ is the **electric charge current**, and charge can be positive or negative. A solution with $j^0 < 0$ describes not a particle of negative probability but a particle carrying charge of the opposite sign — an **antiparticle**. Pauli and Weisskopf established this rereading in 1934 by quantizing the Klein–Gordon equation as a field. We take the matter up after <Ref to="thm-modes-oscillators" />.
</Remark>

## 4. Dirac's resolution and its price

In 1928 Dirac traced the failure of positivity of the probability density to the equation being second order in time. For a first-order equation, $\rho = \psi^\dagger\psi \ge 0$ ought to be a conserved density, just as for the Schrödinger equation.

<Definition id="def-dirac-equation" title="The Dirac equation">
The first-order equation

$$
i\,\partial_t \psi = \hat H\psi,
\qquad
\hat H = \boldsymbol{\alpha}\cdot\hat{\boldsymbol{p}} + \beta m
$$

for an $N$-component wave function $\psi(t,\boldsymbol{x}) \in \mathbb{C}^N$ is called the Dirac equation. Here $\alpha^1,\alpha^2,\alpha^3,\beta$ are Hermitian $N\times N$ matrices, required to satisfy

$$
\{\alpha^i,\alpha^j\} = 2\delta^{ij}\mathbb{1},
\qquad
\{\alpha^i,\beta\} = 0,
\qquad
\beta^2 = \mathbb{1}
$$

so that $\hat H^2 = \hat{\boldsymbol{p}}^2 + m^2$ holds.
</Definition>

That these relations are equivalent to $\hat H^2 = \hat{\boldsymbol{p}}^2 + m^2$ becomes visible on expanding:

$$
\hat H^2 = \alpha^i\alpha^j \hat p_i\hat p_j + m\left(\alpha^i\beta + \beta\alpha^i\right)\hat p_i + \beta^2 m^2
= \tfrac{1}{2}\{\alpha^i,\alpha^j\}\hat p_i\hat p_j + m\{\alpha^i,\beta\}\hat p_i + \beta^2m^2 .
$$

In the first term $\hat p_i \hat p_j$ is symmetric under $i \leftrightarrow j$, so only the symmetric part of $\alpha^i\alpha^j$ contributes, and that part is half the anticommutator. Under the three conditions above we therefore obtain $\hat H^2 = \delta^{ij}\hat p_i\hat p_j + m^2 = \hat{\boldsymbol{p}}^2 + m^2$.

Since $\hat H$ is Hermitian, $\rho = \psi^\dagger\psi$ and $\boldsymbol{j} = \psi^\dagger\boldsymbol{\alpha}\psi$ satisfy the continuity equation $\partial_t\rho + \nabla\cdot\boldsymbol{j} = 0$, and moreover $\rho \ge 0$. Disease 2 of <Ref to="prop-kg-current" /> has been cured. Disease 1, however, survives.

<Theorem id="thm-dirac-spectrum" title="The Dirac Hamiltonian is unbounded below">
Suppose there exist Hermitian $N\times N$ matrices $\alpha^i,\beta$ satisfying the conditions of <Ref to="def-dirac-equation" />. For the matrix $H(\boldsymbol{p}) = \boldsymbol{\alpha}\cdot\boldsymbol{p} + \beta m$ attached to a momentum eigenvalue $\boldsymbol{p}$, the following hold.

1. $\operatorname{tr}\alpha^i = 0$ for $i=1,2,3$, and $\operatorname{tr}\beta = 0$. Hence $\operatorname{tr}H(\boldsymbol{p}) = 0$.
2. The eigenvalues of $H(\boldsymbol{p})$ are $+\omega_{\boldsymbol p}$ and $-\omega_{\boldsymbol p}$ only, and the two have equal multiplicity $N/2$. In particular $N$ is even, the spectrum of $\hat H$ is $(-\infty,-m]\cup[m,\infty)$, and it is unbounded below.
</Theorem>

<Proof of="thm-dirac-spectrum">
**Proof of 1.** From $\{\alpha^i,\beta\}=0$ we have $\beta\alpha^i = -\alpha^i\beta$; multiplying on the right by $\beta$ and using $\beta^2=\mathbb{1}$ gives

$$
\beta\alpha^i\beta = -\alpha^i\beta^2 = -\alpha^i .
$$

Take traces of both sides. By cyclicity, the left-hand side is $\operatorname{tr}(\beta\alpha^i\beta) = \operatorname{tr}(\alpha^i\beta^2) = \operatorname{tr}\alpha^i$. Hence $\operatorname{tr}\alpha^i = -\operatorname{tr}\alpha^i$, that is, $\operatorname{tr}\alpha^i = 0$. Similarly, from $(\alpha^1)^2 = \mathbb{1}$ and $\{\alpha^1,\beta\}=0$ we get $\alpha^1\beta\alpha^1 = -\beta(\alpha^1)^2 = -\beta$, and taking traces yields $\operatorname{tr}\beta = -\operatorname{tr}\beta = 0$. By linearity, $\operatorname{tr}H(\boldsymbol{p}) = p_i\operatorname{tr}\alpha^i + m\operatorname{tr}\beta = 0$.

**Proof of 2.** As checked immediately after <Ref to="def-dirac-equation" />, $H(\boldsymbol{p})^2 = (\boldsymbol{p}^2+m^2)\mathbb{1} = \omega_{\boldsymbol p}^2\mathbb{1}$. Since $H(\boldsymbol p)$ is Hermitian it is diagonalizable, and its eigenvalues $\lambda$ obey $\lambda^2 = \omega_{\boldsymbol p}^2$, so $\lambda = \pm\omega_{\boldsymbol p}$. Writing $n_+,n_-$ for the multiplicities we have $n_++n_-=N$, and part 1 gives

$$
\operatorname{tr}H(\boldsymbol p) = n_+\omega_{\boldsymbol p} - n_-\omega_{\boldsymbol p} = 0 .
$$

Since $m > 0$ we have $\omega_{\boldsymbol p} \ge m > 0$, whence $n_+ = n_- = N/2$. As $N \ge 1$, we get $n_- \ge 1$: negative eigenvalues necessarily exist. Letting $\boldsymbol{p}$ range over all of space, $-\omega_{\boldsymbol p}$ takes every value continuously from $-m$ down to $-\infty$, so the spectrum is unbounded below.
</Proof>

Thus making the equation first order does not remove the negative-energy solutions. <Ref to="thm-dirac-spectrum" /> says more than that they survive: **exactly half** of the states have negative energy.

<Example id="ex-dirac-sea" title="The Dirac sea and its price">
Dirac's prescription runs as follows. The electron is a fermion, so the Pauli exclusion principle applies. Define the vacuum as the state in which every negative-energy level is occupied by an electron; then a positive-energy electron cannot fall into a negative-energy level, because exclusion forbids it. A hole in the sea behaves like a particle of charge $+e$ and positive energy. This is the positron, discovered by Anderson in a cloud chamber in 1932 — a spectacular success as a prediction.

The price, however, is high.

1. The vacuum contains infinitely many electrons, and both its charge density and its energy density diverge. Observables have to be redefined as differences from the vacuum.
2. The construction relies on exclusion, so **it is unusable for bosons**. That spin-0 pions have antiparticles ($\pi^+$ and $\pi^-$) is an experimental fact.
3. Decisively, this is no longer a one-particle theory. The moment infinitely many particles are introduced the description has become a many-body one, and the prescription itself concedes that one-particle relativistic quantum mechanics does not hold up.
</Example>

<Remark id="rem-spin-from-relativity" title="Spin emerges from relativity">
<Ref to="thm-dirac-spectrum" /> also shows that $N$ is even. In fact $N = 2$ is impossible, and the minimal dimension is $4$ (<Ref to="exr-dirac-dimension" />). Merely attempting to write a relativistic first-order equation forces the wave function to have several components, and that internal degree of freedom corresponds to the spin $1/2$ introduced in [angular momentum and spin](/physics/quantum-mechanics/angular-momentum-and-spin). Spin is a consequence of combining relativity with quantum mechanics.
</Remark>

## 5. A deeper disease: a one-particle theory cannot respect causality

One escape route from the negative-energy problem still seems open: keep only the positive-frequency part. That is, restrict the Hilbert space to positive-energy solutions and adopt

$$
H = \sqrt{-\nabla^2 + m^2}
$$

as the Hamiltonian. This is a self-adjoint operator on $L^2(\mathbb{R}^3)$ with spectrum $[m,\infty)$; it is indeed bounded below, and the difficulty of <Ref to="thm-dirac-spectrum" /> is formally evaded. But at this point a deeper disease appears: **causality**, the very heart of relativity, is violated.

### 5.1. Compact support is destroyed instantaneously

<Proposition id="prop-instant-delocalization" title="A positive-energy particle cannot stay localized">
Let $m > 0$, $H = \sqrt{-\nabla^2+m^2}$ and $\psi_t = e^{-itH}\psi_0$. If $\psi_0 \in L^2(\mathbb{R}^3)$ has compact support and $\psi_0 \ne 0$, then $\psi_t$ fails to have compact support at every time $t \ne 0$.
</Proposition>

<Proof of="prop-instant-delocalization">
Under a spatial Fourier transform, $\hat\psi_t(\boldsymbol{p}) = e^{-it\omega_{\boldsymbol p}}\hat\psi_0(\boldsymbol{p})$ with $\omega_{\boldsymbol p} = \sqrt{\boldsymbol{p}^2+m^2}$.

We argue by contradiction. Suppose that for some $t\ne 0$ the function $\psi_t$ also has compact support. By the Paley–Wiener–Schwartz theorem, the Fourier transform of a compactly supported $L^2$ function continues analytically to an entire function (of exponential type) on $\mathbb{C}^3$. Hence $\hat\psi_0$ and $\hat\psi_t$ both extend to entire functions $F_0, F_t$.

Since $\psi_0 \ne 0$, the function $F_0$ is not identically zero, and as $F_0$ is continuous at real points there is a real vector $\boldsymbol{p}_0$ with $F_0(\boldsymbol{p}_0)\ne 0$. Fix a real unit vector $\boldsymbol{e}$, restrict to the complex line $\boldsymbol{p} = \boldsymbol{p}_0 + z\boldsymbol{e}$ with $z\in\mathbb{C}$, and set $f_0(z) = F_0(\boldsymbol{p}_0+z\boldsymbol{e})$, $f_t(z) = F_t(\boldsymbol{p}_0+z\boldsymbol{e})$. Both are entire in $z$, and $f_0(0)\ne 0$.

Along this line,

$$
Q(z) := (\boldsymbol{p}_0+z\boldsymbol{e})^2 + m^2 = z^2 + 2(\boldsymbol{p}_0\cdot\boldsymbol{e})z + \boldsymbol{p}_0^2 + m^2 .
$$

Its discriminant is $4(\boldsymbol{p}_0\cdot\boldsymbol{e})^2 - 4(\boldsymbol{p}_0^2+m^2)$, which is negative because Cauchy–Schwarz gives $(\boldsymbol{p}_0\cdot\boldsymbol{e})^2 \le \boldsymbol{p}_0^2 < \boldsymbol{p}_0^2+m^2$ (here we used $m>0$). Hence $Q$ has two distinct non-real roots $z_\pm$, and $\sqrt{Q(z)} = \sqrt{(z-z_+)(z-z_-)}$ has genuine branch points at $z_\pm$.

On the real axis, $f_t(z) = g(z)f_0(z)$ with $g(z) = e^{-it\sqrt{Q(z)}}$. Since $f_0, f_t$ are entire and $f_0\not\equiv 0$, the quotient $f_t/f_0$ is a **single-valued** meromorphic function on $\mathbb{C}$. On the other hand, continuing $g$ once around $z_+$ reverses the sign of $\sqrt{Q}$ and carries $g$ into $e^{+it\sqrt{Q}}$. For $t\ne 0$, the equation $e^{-2it\sqrt{Q(z)}} = 1$ holds only at the isolated points where $\sqrt{Q(z)} \in (\pi/t)\mathbb{Z}$, so $g$ is not single-valued near $z_+$. A single-valued meromorphic function cannot agree on the real axis with a function that is not single-valued, and we have a contradiction.
</Proof>

That the support ceases to be compact means that a particle confined to a bounded region at the initial time has nonzero amplitude arbitrarily far away after an arbitrarily short time $t > 0$. The limit imposed by the speed of light is not in force. Let us next estimate the size of that amplitude.

<Example id="ex-propagator-tail" title="The propagation amplitude outside the light cone">
The amplitude for a particle located at the point $\boldsymbol{x}$ at time $0$ to be found at the point $\boldsymbol{y}$ at time $t$ is

$$
U(t,r) = \langle \boldsymbol{y}\,\lvert\, e^{-iHt}\,\rvert\,\boldsymbol{x}\rangle
= \int\frac{d^3p}{(2\pi)^3}\;e^{i\boldsymbol{p}\cdot(\boldsymbol{y}-\boldsymbol{x})}\,e^{-i\omega_{\boldsymbol p}t},
\qquad r = \lvert\boldsymbol{y}-\boldsymbol{x}\rvert .
$$

Performing the angular integral (with $\int d\Omega\, e^{ipr\cos\theta} = 4\pi\sin(pr)/(pr)$) gives

$$
U(t,r) = \frac{1}{2\pi^2 r}\int_0^\infty dp\;p\,\sin(pr)\,e^{-i\omega_p t} .
$$

Deforming the contour around the branch points $p = \pm im$ of $\sqrt{p^2+m^2}$ (the computation is in the Appendix), one obtains, for a spacelike separation $r > t > 0$, the expression

$$
U(t,r) = \frac{i}{2\pi^2 r}\int_m^\infty d\rho\;\rho\,e^{-\rho r}\,\sinh\!\left(t\sqrt{\rho^2-m^2}\right) .
$$

Inside the range of integration the integrand is **strictly positive** (for $\rho > m$ one has $\sinh(t\sqrt{\rho^2-m^2}) > 0$). Hence $U(t,r) \ne 0$: the amplitude does not vanish outside the light cone.

Its magnitude can be read off from the fact that for $mr \gg 1$ and $r \gg t$ the lower endpoint $\rho\simeq m$ dominates. Setting $\rho = m+u$, using $\sqrt{\rho^2-m^2}\simeq\sqrt{2mu}$ and $\sinh x\simeq x$, and using $\int_0^\infty du\,\sqrt{u}\,e^{-ur} = \sqrt{\pi}/(2r^{3/2})$, we find

$$
\lvert U\rvert \sim \frac{\sqrt{2\pi}\;m^{3/2}\,t}{4\pi^2\,r^{5/2}}\,e^{-mr} .
$$

A more careful saddle-point evaluation turns the exponent into the Lorentz-invariant form $e^{-m\sqrt{r^2-t^2}}$. The amplitude is exponentially small at distances beyond the Compton wavelength $1/m$, but it is never zero.
</Example>

<Figure caption="Even at a point P spacelike separated from the event at the origin, the one-particle propagation amplitude is merely exponentially small, never zero. The shaded areas are the spacelike region.">
<svg viewBox="0 0 520 320" width="100%" role="img" aria-label="Light cone and the propagation amplitude at spacelike separation">
  <polygon points="240,240 420,60 500,60 500,310 310,310" fill="currentColor" opacity="0.08" />
  <polygon points="240,240 60,60 20,60 20,310 170,310" fill="currentColor" opacity="0.08" />
  <g stroke="currentColor" fill="none" stroke-width="1.1">
    <line x1="20" y1="240" x2="500" y2="240" />
    <line x1="240" y1="310" x2="240" y2="20" />
  </g>
  <g stroke="currentColor" fill="none" stroke-width="1.8">
    <line x1="240" y1="240" x2="60" y2="60" />
    <line x1="240" y1="240" x2="420" y2="60" />
    <line x1="240" y1="240" x2="170" y2="310" />
    <line x1="240" y1="240" x2="310" y2="310" />
  </g>
  <line x1="240" y1="240" x2="450" y2="170" stroke="var(--sl-color-accent)" stroke-width="1.8" stroke-dasharray="6 4" fill="none" />
  <circle cx="450" cy="170" r="4.5" fill="var(--sl-color-accent)" />
  <g fill="currentColor" font-size="13">
    <text x="226" y="32" font-style="italic">t</text>
    <text x="490" y="262" font-style="italic">x</text>
    <text x="222" y="260">O</text>
    <text x="428" y="78">light cone</text>
    <text x="460" y="164">P</text>
    <text x="400" y="228" text-anchor="middle">spacelike region</text>
    <text x="330" y="196" text-anchor="middle" font-size="12">amplitude ≠ 0</text>
  </g>
</svg>
</Figure>

### 5.2. Localization and causal propagation are incompatible

<Ref to="ex-propagator-tail" /> was an explicit computation, but the phenomenon is in fact a general one that follows from the single condition $H \ge 0$. The following theorem is due to Hegerfeldt (1974).

<Theorem id="thm-hegerfeldt" title="Hegerfeldt's theorem">
Let $\mathcal{H}$ be a Hilbert space and let $H$ be a self-adjoint operator on $\mathcal{H}$ with $H \ge 0$ (its spectrum contained in $[0,\infty)$). Let $\psi_0\in\mathcal{H}$, $\psi_t = e^{-iHt}\psi_0$, and let $A$ be a bounded positive operator ($A \ge 0$). If

$$
\langle \psi_t, A\psi_t\rangle = 0 \quad (\forall t\in I)
$$

holds on some open interval $I\subset\mathbb{R}$, then $\langle\psi_t,A\psi_t\rangle = 0$ for every $t\in\mathbb{R}$.
</Theorem>

<Proof of="thm-hegerfeldt">
Since $A\ge 0$ and $A$ is bounded, a bounded square root $A^{1/2}\ge 0$ exists and $\langle\psi_t,A\psi_t\rangle = \lVert A^{1/2}\psi_t\rVert^2$. The hypothesis is therefore equivalent to $A^{1/2}\psi_t = 0$ for $t\in I$.

Fix an arbitrary $\chi\in\mathcal{H}$ and consider, on the lower half plane $\mathbb{C}_- = \{z : \operatorname{Im}z < 0\}$,

$$
g(z) = \left\langle \chi,\; A^{1/2}e^{-izH}\psi_0\right\rangle .
$$

Let $E(\cdot)$ be the spectral measure of $H$ and introduce the complex measure $\mu(\cdot) = \langle \chi, A^{1/2}E(\cdot)\psi_0\rangle$, whose total variation is finite and at most $\lVert\chi\rVert\,\lVert A^{1/2}\rVert\,\lVert\psi_0\rVert$. Then

$$
g(z) = \int_{[0,\infty)} e^{-iz\lambda}\,d\mu(\lambda) .
$$

Because $H\ge 0$ the spectrum is confined to $\lambda\ge 0$, and for $z = t-is$ with $s>0$ we have $\lvert e^{-iz\lambda}\rvert = e^{-s\lambda}\le 1$. Hence $g$ is bounded on $\mathbb{C}_-$, and since the integrand is holomorphic and admits a uniform integrable majorant, Fubini's theorem together with Morera's theorem shows that $g$ is holomorphic there. By dominated convergence it has the boundary value $g(t) = \langle\chi, A^{1/2}\psi_t\rangle$.

By hypothesis the boundary value vanishes on $I$, a set of positive Lebesgue measure. A bounded holomorphic function on a half plane — an element of the Hardy space $H^\infty$ — whose boundary values vanish on a set of positive measure is identically zero, by the boundary uniqueness theorem (the theorem of F. and M. Riesz–Privalov). Hence $g\equiv 0$.

As $\chi$ was arbitrary, $A^{1/2}\psi_t = 0$ for every real $t$, that is, $\langle\psi_t,A\psi_t\rangle = \lVert A^{1/2}\psi_t\rVert^2 = 0$.
</Proof>

<Corollary id="cor-no-causal-localization" title="Strict localization and causal propagation are incompatible">
Let $H\ge 0$, and suppose that for every bounded open set $W\subset\mathbb{R}^3$ we are given a projection $P_W$ expressing that the particle is found inside $W$, with $P_W \le P_{W'}$ whenever $W\subset W'$. Suppose a state $\psi_0$ is strictly localized in a bounded region $V$ and that propagation is causal, that is,

$$
\operatorname{dist}(W,V) > \lvert t\rvert \;\Longrightarrow\; \langle\psi_t, P_W\psi_t\rangle = 0
$$

holds (with $c=1$). Then for every bounded open set $W$ lying at positive distance from $V$ we have $\langle\psi_t,P_W\psi_t\rangle = 0$ at **all** times $t\in\mathbb{R}$.
</Corollary>

<Proof of="cor-no-causal-localization">
Put $d = \operatorname{dist}(W,V) > 0$. The hypothesis of causal propagation gives $\langle\psi_t,P_W\psi_t\rangle = 0$ on the open interval $I = (-d,d)$. The projection $P_W$ is bounded and satisfies $P_W = P_W^2 = P_W^\dagger \ge 0$, so <Ref to="thm-hegerfeldt" /> applies with $A = P_W$, and the conclusion extends to all $t$.
</Proof>

The particle would therefore never be found outside $\bar V$, for all time. But the spectrum of $H=\sqrt{-\nabla^2+m^2}$ is absolutely continuous, so the RAGE theorem gives $\lVert P_K\psi_t\rVert \to 0$ as $t\to\infty$ for every bounded region $K$: the particle must spread. This is a contradiction.

The conclusion is this. **The three requirements — $H\ge 0$, a notion of strict localization, and causal propagation — cannot hold simultaneously.** Abandon $H\ge 0$ and we are back to the difficulty of <Ref to="thm-dirac-spectrum" />; abandon causality and we abandon relativity; abandon localization and the very question of where the particle is loses its meaning.

<Aside type="caution">
The last option is in fact the doorway to the right answer. Quantum field theory removes the position of a particle from the list of fundamental observables. The position operator $\hat{\boldsymbol{x}}$ disappears from the theory, and $\boldsymbol{x}$ becomes a label on the field — the demotion announced in <Ref to="rem-pauli-time-operator" />.
</Aside>

## 6. Particle number is not conserved: the Compton wavelength as a scale

So far the breakdowns have been logical ones. Physically, too, the reason a one-particle theory cannot hold is plain. In relativity, mass and energy are interchangeable ([relativistic mechanics](/en/physics/relativity/relativistic-mechanics)), so with enough energy available particles simply get created.

<Definition id="def-compton-wavelength" title="The reduced Compton wavelength">
For a particle of mass $m$, the quantity

$$
\bar\lambda_C = \frac{\hbar}{mc}
$$

is called the reduced Compton wavelength (multiplying it by $2\pi$ gives the Compton wavelength $\lambda_C = h/mc$).
</Definition>

The meaning of this scale can be read off from the uncertainty relation. Confining a particle to a region of size $\Delta x$ creates a momentum uncertainty $\Delta p \gtrsim \hbar/\Delta x$, which relativistically amounts to an energy uncertainty

$$
\Delta E \gtrsim c\,\Delta p \gtrsim \frac{\hbar c}{\Delta x} .
$$

Once $\Delta x \lesssim \bar\lambda_C$ we have $\Delta E \gtrsim mc^2$, which is enough energy to be borrowed for the creation of a particle–antiparticle pair.

<Example id="ex-compton-numbers" title="Numbers for the electron">
We use $\hbar = 1.055\times10^{-34}\ \mathrm{J\,s}$, $m_e = 9.109\times10^{-31}\ \mathrm{kg}$ and $c = 2.998\times10^{8}\ \mathrm{m/s}$. Since $m_e c = 2.731\times10^{-22}\ \mathrm{kg\,m/s}$,

$$
\bar\lambda_C = \frac{1.055\times10^{-34}}{2.731\times10^{-22}}\ \mathrm{m} = 3.86\times10^{-13}\ \mathrm{m} = 386\ \mathrm{fm} .
$$

The threshold for pair creation, on the other hand, is $2m_ec^2 = 1.022\ \mathrm{MeV}$. So if one tries to squeeze an electron into a region narrower than about $400\ \mathrm{fm}$, the energy poured in for that purpose creates a new electron–positron pair. It ceases to be meaningful to speak of "the wave function of one electron" at a resolution finer than $\bar\lambda_C$.

This is the same scale as the decay length $1/m$ of the amplitude found in <Ref to="ex-propagator-tail" /> (in natural units). That the distance at which the causality-violating amplitude matters coincides with the distance at which pair creation sets in is no accident. Both are faces of the same fact: the one-particle picture breaks down.
</Example>

<Aside type="note">
Conversely, the same computation shows why nonrelativistic quantum mechanics works so well in atomic physics. The Bohr radius is $a_0 = \bar\lambda_C/\alpha \simeq 137\,\bar\lambda_C$, so the region in which the electron moves is two orders of magnitude larger than the Compton wavelength. The numbers, and the fact that the typical speed satisfies $v/c\sim\alpha$, are verified in <Ref to="exr-compton-scale" />.
</Aside>

## 7. The way out: quantize the field

There is a single prescription that resolves all three breakdowns above (negative energy, probability density, causality) together with the non-conservation of particle number: **treat the classical field as the dynamical variable and quantize it canonically.**

### 7.1. A free field is a collection of independent harmonic oscillators

The starting point is [Lagrangian mechanics](/en/physics/mechanics/lagrangian-mechanics) extended to fields (treated in detail in [classical field theory and the Lagrangian](/en/physics/qft/classical-field-theory)). The simplest Lagrangian for a real scalar field $\phi(t,\boldsymbol{x})$ is

$$
L = \int d^3x\;\mathcal{L},
\qquad
\mathcal{L} = \tfrac12\dot\phi^2 - \tfrac12(\nabla\phi)^2 - \tfrac12 m^2\phi^2 ,
$$

and its Euler–Lagrange equation is precisely the Klein–Gordon equation of <Ref to="def-klein-gordon" />. Here lies the decisive change of viewpoint: we reread the Klein–Gordon equation not as the equation for a one-particle wave function but as the **equation of motion of a classical field**.

<Theorem id="thm-modes-oscillators" title="Mode decomposition and quantization of the free scalar field">
Let a real scalar field $\phi$ inside a cube of side $L$ (volume $V = L^3$, periodic boundary conditions) obey the Lagrangian above. Substituting the Fourier expansion

$$
\phi(t,\boldsymbol{x}) = \frac{1}{\sqrt{V}}\sum_{\boldsymbol{k}} q_{\boldsymbol{k}}(t)\,e^{i\boldsymbol{k}\cdot\boldsymbol{x}},
\qquad
\boldsymbol{k}\in\frac{2\pi}{L}\mathbb{Z}^3,
\qquad
q_{-\boldsymbol{k}} = \overline{q_{\boldsymbol{k}}}
$$

gives

$$
L = \sum_{\boldsymbol{k}}\frac12\left(\lvert\dot q_{\boldsymbol{k}}\rvert^2 - \omega_{\boldsymbol{k}}^2\,\lvert q_{\boldsymbol{k}}\rvert^2\right),
\qquad
\omega_{\boldsymbol{k}} = \sqrt{\boldsymbol{k}^2+m^2} .
$$

That is, a free field is equivalent to a collection of independent harmonic oscillators, one per mode, with angular frequency $\omega_{\boldsymbol{k}}$. Quantizing canonically, with $[\hat a_{\boldsymbol k},\hat a^\dagger_{\boldsymbol k'}] = \delta_{\boldsymbol{k}\boldsymbol{k}'}$ and all other commutators zero, one obtains

$$
\hat H = \sum_{\boldsymbol{k}}\omega_{\boldsymbol{k}}\left(\hat a^\dagger_{\boldsymbol{k}}\hat a_{\boldsymbol{k}} + \tfrac12\right),
\qquad
\hat{\boldsymbol{P}} = \sum_{\boldsymbol{k}}\boldsymbol{k}\,\hat a^\dagger_{\boldsymbol{k}}\hat a_{\boldsymbol{k}} ,
$$

with energy eigenvalues (apart from the zero-point energy) $\sum_{\boldsymbol{k}} n_{\boldsymbol{k}}\,\omega_{\boldsymbol{k}}$, where $n_{\boldsymbol{k}}\in\{0,1,2,\ldots\}$.
</Theorem>

<Proof of="thm-modes-oscillators">
We repeatedly use the fact that under periodic boundary conditions $\frac{1}{V}\int_V d^3x\;e^{i(\boldsymbol{k}+\boldsymbol{k}')\cdot\boldsymbol{x}} = \delta_{\boldsymbol{k},-\boldsymbol{k}'}$. First,

$$
\int d^3x\;\dot\phi^2 = \sum_{\boldsymbol{k},\boldsymbol{k}'}\dot q_{\boldsymbol{k}}\dot q_{\boldsymbol{k}'}\,\delta_{\boldsymbol{k},-\boldsymbol{k}'} = \sum_{\boldsymbol{k}}\dot q_{\boldsymbol{k}}\dot q_{-\boldsymbol{k}} = \sum_{\boldsymbol{k}}\lvert\dot q_{\boldsymbol{k}}\rvert^2
$$

(the last step uses the reality condition $q_{-\boldsymbol{k}}=\overline{q_{\boldsymbol{k}}}$). Similarly, from $\nabla e^{i\boldsymbol{k}\cdot\boldsymbol{x}} = i\boldsymbol{k}\,e^{i\boldsymbol{k}\cdot\boldsymbol{x}}$,

$$
\int d^3x\;(\nabla\phi)^2 = \sum_{\boldsymbol{k}}(i\boldsymbol{k}q_{\boldsymbol{k}})\cdot(-i\boldsymbol{k}q_{-\boldsymbol{k}}) = \sum_{\boldsymbol{k}}\boldsymbol{k}^2\lvert q_{\boldsymbol{k}}\rvert^2,
\qquad
\int d^3x\;\phi^2 = \sum_{\boldsymbol{k}}\lvert q_{\boldsymbol{k}}\rvert^2 .
$$

Combining the three gives the asserted expression for $L$, with $\boldsymbol{k}^2 + m^2 = \omega_{\boldsymbol{k}}^2$.

Next, the complex variables $q_{\boldsymbol{k}}$ are not independent, so we pass to real degrees of freedom. Choose a half set $K_+$ containing exactly one of each pair $\boldsymbol{k}$, $-\boldsymbol{k}$, and for $\boldsymbol{k}\in K_+$ write $q_{\boldsymbol{k}} = (x_{\boldsymbol{k}} + i y_{\boldsymbol{k}})/\sqrt2$. Since $\lvert q_{-\boldsymbol{k}}\rvert = \lvert q_{\boldsymbol{k}}\rvert$, the contributions of $\pm\boldsymbol{k}$ are equal, and

$$
L = \sum_{\boldsymbol{k}\in K_+}\frac12\left[(\dot x_{\boldsymbol{k}}^2 + \dot y_{\boldsymbol{k}}^2) - \omega_{\boldsymbol{k}}^2(x_{\boldsymbol{k}}^2+y_{\boldsymbol{k}}^2)\right] + \frac12\left(\dot q_{\boldsymbol{0}}^2 - m^2 q_{\boldsymbol{0}}^2\right) .
$$

This is an array of real harmonic oscillators of unit mass and angular frequency $\omega_{\boldsymbol{k}}$, one per $\boldsymbol{k}$ (two for each element of $K_+$, and one for $\boldsymbol{k}=\boldsymbol{0}$).

Passing each oscillator over to [Hamiltonian mechanics](/physics/mechanics/hamiltonian-mechanics), imposing $[\hat q,\hat p]=i$ and introducing the standard creation and annihilation operators yields $\hat H_{\text{osc}} = \omega(\hat a^\dagger\hat a + \tfrac12)$. Recombining the pairs $x_{\boldsymbol{k}},y_{\boldsymbol{k}}$ into complex combinations produces operators $\hat a_{\boldsymbol{k}}$ and $\hat a_{-\boldsymbol{k}}$ labelled by momenta $\pm\boldsymbol{k}$, and the total Hamiltonian assembles into the asserted form. As for the momentum, rewriting $\boldsymbol{P} = -\int d^3x\,\dot\phi\,\nabla\phi$ — obtained from Noether's theorem ([symmetries and conservation laws](/physics/mechanics/noethers-theorem)) — by the same procedure gives $\sum_{\boldsymbol{k}}\boldsymbol{k}\,\hat a^\dagger_{\boldsymbol{k}}\hat a_{\boldsymbol{k}}$ (the zero-point terms cancel by the symmetry $\boldsymbol{k}\to-\boldsymbol{k}$).

The eigenstates are products of number states of the individual modes, with eigenvalues $\sum_{\boldsymbol{k}}\omega_{\boldsymbol{k}}(n_{\boldsymbol{k}}+\tfrac12)$.
</Proof>

How this result is to be read is quantum field theory itself. By <Ref to="thm-modes-oscillators" />, the state $\hat a^\dagger_{\boldsymbol{k}}\lvert 0\rangle$ obtained by acting once with $\hat a^\dagger_{\boldsymbol{k}}$ on the vacuum $\lvert 0\rangle$ carries

$$
E = \omega_{\boldsymbol{k}} = \sqrt{\boldsymbol{k}^2+m^2},
\qquad
\boldsymbol{P} = \boldsymbol{k} .
$$

These are exactly the energy and momentum of one relativistic particle of mass $m$. Acting $n$ times gives $n$ particles. The picture that **a particle is an excitation quantum of a field** has emerged as a consequence of the computation.

<Definition id="def-fock-space" title="Fock space">
The space spanned by all states obtained by acting with creation operators on the vacuum $\lvert 0\rangle$ (which satisfies $\hat a_{\boldsymbol{k}}\lvert 0\rangle = 0$ for every $\boldsymbol{k}$),

$$
\mathcal{F} = \bigoplus_{n=0}^{\infty}\mathcal{H}^{(n)},
\qquad
\mathcal{H}^{(n)} = \operatorname{span}\left\{\hat a^\dagger_{\boldsymbol{k}_1}\cdots \hat a^\dagger_{\boldsymbol{k}_n}\lvert 0\rangle\right\} ,
$$

is called Fock space. The number operator is $\hat N = \sum_{\boldsymbol{k}}\hat a^\dagger_{\boldsymbol{k}}\hat a_{\boldsymbol{k}}$.
</Definition>

Fock space is a direct sum of subspaces with different particle numbers, so operators that change the particle number **can be written down**. The restriction (A1) has been lifted. Moreover a state is determined by the occupation numbers $\{n_{\boldsymbol{k}}\}$ alone, so symmetry under exchange of particles holds automatically. There is no need to impose symmetrization for identical particles by hand.

<Example id="ex-phi4-number-violation" title="Interactions do not conserve particle number">
Add $-\frac{\lambda}{4!}\phi^4$ to the Lagrangian. The corresponding interaction Hamiltonian is $\hat H_{\text{int}} = \frac{\lambda}{4!}\int d^3x\;\hat\phi^4$. The field operator is a sum of an annihilation part and a creation part, $\hat\phi \sim \hat a + \hat a^\dagger$, so expanding $\hat\phi^4$ produces $2^4 = 16$ terms, which we may classify by the number of creation operators they contain.

$$
\hat\phi^4 \sim \underbrace{\hat a^\dagger\hat a^\dagger\hat a^\dagger\hat a^\dagger}_{\Delta N = +4} + \underbrace{\hat a^\dagger\hat a^\dagger\hat a^\dagger\hat a}_{\Delta N=+2} + \underbrace{\hat a^\dagger\hat a^\dagger\hat a\hat a}_{\Delta N = 0} + \underbrace{\hat a^\dagger\hat a\hat a\hat a}_{\Delta N=-2} + \underbrace{\hat a\hat a\hat a\hat a}_{\Delta N=-4}
$$

(each brace stands for all combinations, including rearrangements of the momenta). Since terms with $\Delta N = \pm 4, \pm 2$ are present, $[\hat H,\hat N]\ne 0$ and particle number is not conserved. A process in which four particles are born out of the vacuum and a process in which two particles scatter into two both come from the same single term. Only within the framework of <Ref to="def-fock-space" /> do such processes become describable.
</Example>

### 7.2. How negative energy and negative probability are resolved

Subtracting the constant zero-point energy $\sum_{\boldsymbol{k}}\omega_{\boldsymbol{k}}/2$ from the $\hat H$ of <Ref to="thm-modes-oscillators" /> leaves $\hat H = \sum_{\boldsymbol{k}}\omega_{\boldsymbol{k}}\hat a^\dagger_{\boldsymbol{k}}\hat a_{\boldsymbol{k}} \ge 0$. Where has the negative energy gone? The answer is that the negative-frequency part $e^{+i\omega t}$ of the solutions became not a negative-energy state but the coefficient of a **creation operator**. In the mode expansion both appear,

$$
\hat\phi(x) \;\sim\; \underbrace{\hat a_{\boldsymbol{k}}\,e^{-i\omega_{\boldsymbol{k}}t}}_{\text{positive frequency}} \;+\; \underbrace{\hat a^\dagger_{\boldsymbol{k}}\,e^{+i\omega_{\boldsymbol{k}}t}}_{\text{negative frequency}} ,
$$

but $\hat a^\dagger_{\boldsymbol{k}}$ is the operator that **raises** the energy by $\omega_{\boldsymbol{k}}$. The very same formula, read differently, yields a spectrum bounded below.

For a complex scalar field the current $j^\mu$ of <Ref to="prop-kg-current" /> becomes the charge current, and the conserved charge takes the form

$$
\hat Q = \sum_{\boldsymbol{k}}\left(\hat a^\dagger_{\boldsymbol{k}}\hat a_{\boldsymbol{k}} - \hat b^\dagger_{\boldsymbol{k}}\hat b_{\boldsymbol{k}}\right) .
$$

The particles created by $\hat a^\dagger$ and those created by $\hat b^\dagger$ have the same mass and opposite charge; the latter are the **antiparticles**. So $j^0 < 0$ meant not negative probability but the existence of antiparticles, exactly as announced in <Ref to="rem-charge-not-probability" />.

### 7.3. Causality returns as microcausality

We turn finally to causality. In quantum field theory causality is implemented not as the vanishing of an amplitude but as the statement that **field operators at spacelike-separated points commute**. If they commute, a measurement at one point cannot influence the outcome of a measurement at the other.

<Definition id="def-free-scalar-field" title="The free real scalar field operator">
In the continuum limit ($V\to\infty$) we write

$$
\hat\phi(x) = \int\frac{d^3p}{(2\pi)^3}\;\frac{1}{\sqrt{2\omega_{\boldsymbol{p}}}}\left(\hat a_{\boldsymbol{p}}\,e^{-ip\cdot x} + \hat a^\dagger_{\boldsymbol{p}}\,e^{+ip\cdot x}\right),
\qquad p^0 = \omega_{\boldsymbol{p}} .
$$

The commutation relations are $[\hat a_{\boldsymbol{p}},\hat a^\dagger_{\boldsymbol{q}}] = (2\pi)^3\delta^3(\boldsymbol{p}-\boldsymbol{q})$, all others being zero.
</Definition>

<Theorem id="thm-microcausality" title="Microcausality">
For the field of <Ref to="def-free-scalar-field" />, setting $D(x) := \displaystyle\int\frac{d^3p}{(2\pi)^3\,2\omega_{\boldsymbol{p}}}e^{-ip\cdot x}$, we have

$$
[\hat\phi(x),\hat\phi(y)] = D(x-y) - D(y-x) ,
$$

and this vanishes whenever $(x-y)^2 < 0$, that is, at spacelike separation.
</Theorem>

<Proof of="thm-microcausality">
**Step 1: the form of the commutator.** Substituting <Ref to="def-free-scalar-field" />, the only surviving terms are those containing $[\hat a_{\boldsymbol{p}},\hat a^\dagger_{\boldsymbol{q}}]$ and $[\hat a^\dagger_{\boldsymbol{p}},\hat a_{\boldsymbol{q}}] = -(2\pi)^3\delta^3(\boldsymbol{p}-\boldsymbol{q})$.

$$
[\hat\phi(x),\hat\phi(y)] = \int\frac{d^3p\,d^3q}{(2\pi)^6\sqrt{4\omega_{\boldsymbol{p}}\omega_{\boldsymbol{q}}}}
\left([\hat a_{\boldsymbol{p}},\hat a^\dagger_{\boldsymbol{q}}]e^{-ip\cdot x + iq\cdot y} + [\hat a^\dagger_{\boldsymbol{p}},\hat a_{\boldsymbol{q}}]e^{ip\cdot x - iq\cdot y}\right)
$$

$$
= \int\frac{d^3p}{(2\pi)^3\,2\omega_{\boldsymbol{p}}}\left(e^{-ip\cdot(x-y)} - e^{ip\cdot(x-y)}\right) = D(x-y) - D(y-x).
$$

**Step 2: Lorentz invariance of the measure.** The zeros of $\delta(p^2-m^2) = \delta\!\left((p^0)^2 - \omega_{\boldsymbol{p}}^2\right)$ are at $p^0 = \pm\omega_{\boldsymbol{p}}$, where the absolute value of the derivative is $2\omega_{\boldsymbol{p}}$, so

$$
\delta(p^2-m^2) = \frac{1}{2\omega_{\boldsymbol{p}}}\left[\delta(p^0-\omega_{\boldsymbol{p}}) + \delta(p^0+\omega_{\boldsymbol{p}})\right].
$$

Consequently, for any function $f$,

$$
\int\frac{d^4p}{(2\pi)^3}\,\delta(p^2-m^2)\,\theta(p^0)\,f(p) = \int\frac{d^3p}{(2\pi)^3\,2\omega_{\boldsymbol{p}}}f(\omega_{\boldsymbol{p}},\boldsymbol{p}) .
$$

On the left-hand side, $d^4p$ and $\delta(p^2-m^2)$ are Lorentz invariant, and on the positive mass shell $\theta(p^0)$ is invariant under proper orthochronous Lorentz transformations $\Lambda\in SO^+(1,3)$ as well, since such $\Lambda$ preserve the sign of $p^0$. Hence $D(\Lambda x) = D(x)$ for every $\Lambda\in SO^+(1,3)$.

**Step 3: reversing a spacelike vector.** Let $z = x-y$ be spacelike, that is, $z^2 = (z^0)^2 - \lvert\boldsymbol{z}\rvert^2 < 0$. Then $\lvert\boldsymbol{z}\rvert > \lvert z^0\rvert$, so $v = z^0/\lvert\boldsymbol{z}\rvert$ satisfies $\lvert v\rvert < 1$. Applying a boost with velocity $v$ along the direction $\hat{\boldsymbol{z}}$ gives

$$
z'^0 = \gamma\left(z^0 - v\lvert\boldsymbol{z}\rvert\right) = 0 ,
$$

so $z' = (0,\boldsymbol{z}')$ is an equal-time vector. Next, a rotation by angle $\pi$ about an axis orthogonal to $\boldsymbol{z}'$ sends $\boldsymbol{z}'\to-\boldsymbol{z}'$, that is, $z'\to -z'$. Both the boost and the rotation lie in $SO^+(1,3)$, so their composition $\Lambda$ satisfies $\Lambda z = -z$.

By Step 2, $D(z) = D(\Lambda z) = D(-z)$. Returning to Step 1 we obtain $[\hat\phi(x),\hat\phi(y)] = D(z)-D(-z) = 0$.
</Proof>

<Remark id="rem-antiparticle-necessity" title="Causality demands antiparticles">
What was essential in the proof of <Ref to="thm-microcausality" /> is that **two terms, $D(x-y)$ and $D(y-x)$, appeared with a relative minus sign**. Indeed $D$ by itself does not vanish at spacelike separation. At equal times, $x-y = (0,\boldsymbol{r})$,

$$
D(0,\boldsymbol{r}) = \frac{1}{4\pi^2 r}\int_0^\infty dp\;\frac{p\sin(pr)}{\sqrt{p^2+m^2}} = \frac{m}{4\pi^2 r}K_1(mr) > 0
$$

(here $K_1$ is the modified Bessel function of the second kind, with $K_1(u) \sim \sqrt{\pi/2u}\,e^{-u}$), so it has the same $e^{-mr}$ tail found in <Ref to="ex-propagator-tail" />. It is only the difference that vanishes.

For a complex scalar field the meaning of that difference becomes clear. $D(x-y)$ is the amplitude for a **particle** created at $y$ to be annihilated at $x$, while $D(y-x)$ is the amplitude for an **antiparticle** created at $x$ to be annihilated at $y$. At spacelike separation the two events carry no invariant time ordering, so the two amplitudes have equal magnitude and cancel exactly. Without antiparticles there would be no cancellation and causality would fail. **The existence of antiparticles is forced upon us by relativity, quantum mechanics and causality together** (<Ref to="exr-microcausality-complex" />).

For fermions the same cancellation occurs in the anticommutator. That integer spin must be quantized with commutators and half-integer spin with anticommutators, on pain of losing either causality or positivity of the energy, is the spin-statistics theorem. Even the Pauli exclusion principle that the Dirac sea required becomes a theorem here.
</Remark>

### 7.4. Correspondence table

| Issue | One-particle relativistic quantum mechanics | Quantum field theory |
|---|---|---|
| Negative-energy solutions | The Hamiltonian is unbounded below (<Ref to="thm-dirac-spectrum" />) | Negative-frequency coefficients become creation operators, so $\hat H\ge 0$ |
| Sign of the density | $j^0$ goes negative (<Ref to="prop-kg-current" />) | $j^0$ is a charge density; the two signs are particle and antiparticle |
| Causality | The amplitude is nonzero even at spacelike separation (<Ref to="ex-propagator-tail" />) | The commutator vanishes at spacelike separation (<Ref to="thm-microcausality" />) |
| Particle number | No operator changing it can be written | Handled naturally in Fock space (<Ref to="ex-phi4-number-violation" />) |
| Identical particles | Symmetrization or antisymmetrization imposed by hand | Automatically identical, being quanta of one and the same field |
| Spin and statistics | An independent assumption | Derived, as the spin-statistics theorem |

<Remark id="rem-second-quantization-name" title="On the name 'second quantization'">
For historical reasons this procedure is called second quantization. The term originally referred to rewriting many-body quantum mechanics in terms of creation and annihilation operators, and only later came to be applied to the quantization of fields. The former is no more than a change of notation, available for nonrelativistic many-body systems as well; new physics enters when a relativistic field is quantized.

The name is also misleading. What is quantized is the classical field $\phi(x)$; it is not that an already quantized wave function is quantized a second time. Quantization happens once and once only (Weinberg emphasizes this point). Accept the term as a historical label, but understand its content as canonical quantization of a classical field.
</Remark>

## 8. The road ahead

In the chapters that follow we walk through the door opened here.

- [Classical field theory and the Lagrangian](/en/physics/qft/classical-field-theory): the Lagrangian formalism for fields and Noether's theorem, placing the starting point of <Ref to="thm-modes-oscillators" /> on a rigorous footing.
- [Canonical quantization of the scalar field](/physics/qft/canonical-quantization): from the equal-time commutation relation $[\hat\phi(t,\boldsymbol{x}),\hat\pi(t,\boldsymbol{y})] = i\delta^3(\boldsymbol{x}-\boldsymbol{y})$ to Fock space and the propagator.
- [Path integral quantization](/physics/qft/path-integral-quantization): a formulation without operators, which comes into its own for gauge theories.
- [Introduction to renormalization](/physics/qft/renormalization): how the divergences appearing in interacting fields point to the scale at which the theory is valid.
- [Gauge theory and spontaneous symmetry breaking](/physics/qft/gauge-theory-and-symmetry-breaking): the skeleton of the Standard Model.

The computational techniques of [perturbation theory](/physics/quantum-mechanics/perturbation-theory) carry over as the foundation. Treating the interaction term of <Ref to="ex-phi4-number-violation" /> as a perturbation is where Feynman diagrams begin.

## 9. Exercises

<Exercise id="exr-kg-current" difficulty="Standard">
Use natural units $\hbar=c=1$ and the metric $\eta_{\mu\nu} = \mathrm{diag}(+1,-1,-1,-1)$, and let $\phi$ be a complex solution of the Klein–Gordon equation $(\Box+m^2)\phi=0$.

1. Write the spatial components $\boldsymbol{j}$ of the current $j^\mu = \frac{i}{2m}\left(\phi^{*}\partial^\mu\phi - (\partial^\mu\phi^{*})\phi\right)$ in terms of $\phi$ and $\nabla\phi$, and verify that they take the same form as the nonrelativistic probability current $\frac{1}{2mi}(\psi^{*}\nabla\psi - \psi\nabla\psi^{*})$.
2. Compute $j^0$ for the superposition $\phi = N_+e^{-i\omega t} + N_-e^{+i\omega t}$ (with $\boldsymbol{p}=\boldsymbol{0}$, $\omega=m$) and show that $j^0 = 0$ when $\lvert N_+\rvert = \lvert N_-\rvert$. What does this mean?

<Solution>
**1.** Since $\partial^i = \eta^{ii}\partial_i = -\partial/\partial x^i$ (no sum over $i$), the spatial components are

$$
\boldsymbol{j} = \frac{i}{2m}\left(\phi^{*}(-\nabla\phi) - (-\nabla\phi^{*})\phi\right) = -\frac{i}{2m}\left(\phi^{*}\nabla\phi - (\nabla\phi^{*})\phi\right) .
$$

Because $-i/(2m) = 1/(2mi)$,

$$
\boldsymbol{j} = \frac{1}{2mi}\left(\phi^{*}\nabla\phi - \phi\nabla\phi^{*}\right) ,
$$

which has the same form as the nonrelativistic probability current. So the relativistic upgrade changed the form of the time component only.

**2.** Set $\omega = m$. Multiplying $\dot\phi = -im\left(N_+e^{-imt} - N_-e^{+imt}\right)$ by $\phi^{*} = \overline{N_+}e^{+imt} + \overline{N_-}e^{-imt}$ and writing $w := \overline{N_-}N_+e^{-2imt}$ (so that $\overline{N_+}N_-e^{2imt} = \bar w$), we get

$$
\phi^{*}\dot\phi = -im\left[\left(\lvert N_+\rvert^2-\lvert N_-\rvert^2\right) + (w-\bar w)\right]
= -im\left(\lvert N_+\rvert^2-\lvert N_-\rvert^2\right) + 2m\operatorname{Im}w
$$

(using $w-\bar w = 2i\operatorname{Im}w$). The second term is real, so it cancels when we subtract $\dot\phi^{*}\phi = \overline{\phi^{*}\dot\phi}$, leaving

$$
j^0 = \frac{i}{2m}\left(\phi^{*}\dot\phi - \dot\phi^{*}\phi\right) = \frac{i}{2m}\cdot\left(-2im\right)\left(\lvert N_+\rvert^2-\lvert N_-\rvert^2\right) = \lvert N_+\rvert^2 - \lvert N_-\rvert^2 .
$$

Hence $j^0 = 0$ when $\lvert N_+\rvert=\lvert N_-\rvert$.

The meaning: in a state containing equal amounts of positive- and negative-energy components, reading $j^0$ as the probability density for finding the particle there gives the absurd answer zero identically. As explained in <Ref to="rem-charge-not-probability" />, $j^0$ is a charge density, and the correct reading is that equal amounts of positive and negative charge are present, making the state neutral.
</Solution>
</Exercise>

<Exercise id="exr-dirac-dimension" difficulty="Standard">
Let $\alpha^1,\alpha^2,\alpha^3,\beta$ be Hermitian $N\times N$ matrices satisfying $\{\alpha^i,\alpha^j\}=2\delta^{ij}\mathbb{1}$, $\{\alpha^i,\beta\}=0$ and $\beta^2=\mathbb{1}$.

1. Show that $N$ is even.
2. Show that $N=2$ is impossible, and conclude that the minimal dimension is $4$.

<Solution>
**1.** As shown in <Ref to="thm-dirac-spectrum" />, $\operatorname{tr}\beta = 0$. On the other hand $\beta^2=\mathbb{1}$ and $\beta$ is Hermitian, so $\beta$ is diagonalizable with eigenvalues $\pm1$ only. Writing $n_\pm$ for the multiplicities, $n_++n_-=N$ and $\operatorname{tr}\beta = n_+-n_-=0$, whence $n_+=n_-=N/2$. Since $n_\pm$ are integers, $N$ is even.

**2.** Suppose $N=2$. Every traceless Hermitian $2\times2$ matrix can be written uniquely as a real linear combination of $\sigma^1,\sigma^2,\sigma^3$ (this space is real three-dimensional). By <Ref to="thm-dirac-spectrum" /> both $\alpha^i$ and $\beta$ are traceless, so there are real vectors $\boldsymbol{a}_1,\boldsymbol{a}_2,\boldsymbol{a}_3,\boldsymbol{b}\in\mathbb{R}^3$ with

$$
\alpha^i = \boldsymbol{a}_i\cdot\boldsymbol{\sigma},\qquad \beta = \boldsymbol{b}\cdot\boldsymbol{\sigma} .
$$

Using the Pauli identity $(\boldsymbol{u}\cdot\boldsymbol{\sigma})(\boldsymbol{v}\cdot\boldsymbol{\sigma}) = (\boldsymbol{u}\cdot\boldsymbol{v})\mathbb{1} + i(\boldsymbol{u}\times\boldsymbol{v})\cdot\boldsymbol{\sigma}$ and the antisymmetry of $\boldsymbol{u}\times\boldsymbol{v}$, we get

$$
\{\boldsymbol{u}\cdot\boldsymbol{\sigma},\,\boldsymbol{v}\cdot\boldsymbol{\sigma}\} = 2(\boldsymbol{u}\cdot\boldsymbol{v})\,\mathbb{1} .
$$

The three hypotheses are therefore equivalent to

$$
\boldsymbol{a}_i\cdot\boldsymbol{a}_j = \delta_{ij},\qquad \boldsymbol{a}_i\cdot\boldsymbol{b} = 0,\qquad \boldsymbol{b}\cdot\boldsymbol{b} = 1 ,
$$

that is, to the existence of four mutually orthogonal unit vectors inside $\mathbb{R}^3$. Since $\mathbb{R}^3$ has dimension $3$, at most three mutually orthogonal nonzero vectors exist, a contradiction. Hence $N\ne 2$.

By part 1, $N$ is even; $N=2$ has just been excluded, and $N=0$ is meaningless for matrices, so $N\ge 4$. And $N=4$ is actually realized: the Dirac representation $\alpha^i = \begin{pmatrix}0&\sigma^i\\\sigma^i&0\end{pmatrix}$, $\beta=\begin{pmatrix}\mathbb{1}&0\\0&-\mathbb{1}\end{pmatrix}$ satisfies the three conditions, as one checks block by block using $\{\sigma^i,\sigma^j\}=2\delta^{ij}$. The minimal dimension is therefore $4$.
</Solution>
</Exercise>

<Exercise id="exr-compton-scale" difficulty="Easy">
1. Find the ratio of the electron's reduced Compton wavelength $\bar\lambda_C = \hbar/(m_ec)$ to the Bohr radius $a_0 = \hbar/(m_ec\alpha)$, and give the numbers (take $\alpha^{-1}=137.04$).
2. Express the typical speed and kinetic energy of an electron in a hydrogen atom in terms of $\alpha$ and $m_ec^2 = 511\ \mathrm{keV}$, and explain why nonrelativistic quantum mechanics is a good approximation.
3. What happens if one tries to confine an electron to a region of size about $\bar\lambda_C$?

<Solution>
**1.** Straight from the definitions, $a_0/\bar\lambda_C = 1/\alpha = 137.04$. Using $\bar\lambda_C = 3.86\times10^{-13}\ \mathrm{m}$ from <Ref to="ex-compton-numbers" />,

$$
a_0 = 137.04\times 3.86\times10^{-13}\ \mathrm{m} = 5.29\times10^{-11}\ \mathrm{m} ,
$$

which agrees with the known Bohr radius.

**2.** The electron is confined to a spread of order $a_0$, so by the uncertainty relation its typical momentum is $p\sim\hbar/a_0 = \alpha\, m_ec$. Hence

$$
\frac{v}{c}\sim\alpha\simeq\frac{1}{137},
\qquad
E_{\text{kin}}\sim\frac{p^2}{2m_e} = \frac{1}{2}\alpha^2 m_ec^2 = \frac{5.325\times10^{-5}}{2}\times 511\ \mathrm{keV} = 13.6\ \mathrm{eV}.
$$

This is exactly the Rydberg energy of hydrogen. The kinetic energy is only $2.7\times10^{-5}$ times the rest energy, and relativistic corrections enter at relative accuracy $(v/c)^2\sim\alpha^2$, an effect of order $10^{-5}$ (this is the size of the fine structure). Nonrelativistic quantum mechanics is therefore a good approximation. That quantum mechanics alone almost suffices for atomic physics is thanks to the smallness of the fine-structure constant.

**3.** With $\Delta x\sim\bar\lambda_C$ we get $\Delta p\gtrsim\hbar/\bar\lambda_C = m_ec$, hence $\Delta E\gtrsim m_ec^2 = 511\ \mathrm{keV}$. This is the same order as the pair-creation threshold $2m_ec^2 = 1.022\ \mathrm{MeV}$, so the energy used for the confinement creates electron–positron pairs. Since the original electron and the newly created one cannot be told apart, the very description in terms of a one-particle wave function breaks down. One needs a space containing states of different particle number, as in <Ref to="def-fock-space" />.
</Solution>
</Exercise>

<Exercise id="exr-microcausality-complex" difficulty="Hard">
Take the complex scalar field operator to be

$$
\hat\phi(x) = \int\frac{d^3p}{(2\pi)^3}\frac{1}{\sqrt{2\omega_{\boldsymbol{p}}}}\left(\hat a_{\boldsymbol{p}}e^{-ip\cdot x} + \hat b^\dagger_{\boldsymbol{p}}e^{+ip\cdot x}\right)
$$

with $[\hat a_{\boldsymbol{p}},\hat a^\dagger_{\boldsymbol{q}}] = [\hat b_{\boldsymbol{p}},\hat b^\dagger_{\boldsymbol{q}}] = (2\pi)^3\delta^3(\boldsymbol{p}-\boldsymbol{q})$ and all commutators between $\hat a$-type and $\hat b$-type operators vanishing.

1. Show that $[\hat\phi(x),\hat\phi^\dagger(y)] = D(x-y) - D(y-x)$ and conclude that it vanishes at spacelike separation (here $D$ is as in <Ref to="thm-microcausality" />).
2. State what each of $D(x-y)$ and $D(y-x)$ is the amplitude for, physically, and explain the meaning of the cancellation.
3. What breaks if antiparticles are assumed not to exist, that is, if the $\hat b^\dagger$ term is absent?

<Solution>
**1.** We have $\hat\phi^\dagger(y) = \int\frac{d^3q}{(2\pi)^3}\frac{1}{\sqrt{2\omega_{\boldsymbol{q}}}}\left(\hat a^\dagger_{\boldsymbol{q}}e^{+iq\cdot y} + \hat b_{\boldsymbol{q}}e^{-iq\cdot y}\right)$. Expanding the commutator, the cross terms between $\hat a$-type and $\hat b$-type operators vanish, so

$$
[\hat\phi(x),\hat\phi^\dagger(y)] = \int\frac{d^3p\,d^3q}{(2\pi)^6\sqrt{4\omega_{\boldsymbol{p}}\omega_{\boldsymbol{q}}}}\left([\hat a_{\boldsymbol{p}},\hat a^\dagger_{\boldsymbol{q}}]e^{-ip\cdot x+iq\cdot y} + [\hat b^\dagger_{\boldsymbol{p}},\hat b_{\boldsymbol{q}}]e^{+ip\cdot x - iq\cdot y}\right)
$$

$$
= \int\frac{d^3p}{(2\pi)^3 2\omega_{\boldsymbol{p}}}\left(e^{-ip\cdot(x-y)} - e^{+ip\cdot(x-y)}\right) = D(x-y)-D(y-x).
$$

The sign of the second term comes from $[\hat b^\dagger_{\boldsymbol{p}},\hat b_{\boldsymbol{q}}] = -(2\pi)^3\delta^3(\boldsymbol{p}-\boldsymbol{q})$. From here Steps 2 and 3 of <Ref to="thm-microcausality" /> apply verbatim: if $(x-y)^2<0$ there is a $\Lambda\in SO^+(1,3)$ with $\Lambda(x-y) = -(x-y)$, and the invariance of $D$ gives $D(x-y)=D(y-x)$, so the commutator vanishes.

**2.** Taking vacuum expectation values, $\langle 0\rvert\hat\phi(x)\hat\phi^\dagger(y)\lvert 0\rangle = D(x-y)$. Here $\hat\phi^\dagger(y)$ creates one particle (of $\hat a$ type) at $y$ and $\hat\phi(x)$ annihilates it at $x$, so $D(x-y)$ is the amplitude for a particle born at $y$ to disappear at $x$. On the other hand $\langle 0\rvert\hat\phi^\dagger(y)\hat\phi(x)\lvert 0\rangle = D(y-x)$, where $\hat\phi(x)$ creates an antiparticle (of $\hat b$ type) at $x$ and $\hat\phi^\dagger(y)$ annihilates it at $y$. That is the amplitude for an antiparticle born at $x$ to disappear at $y$.

Two spacelike-separated points admit no Lorentz-invariant time ordering. An observer who sees $y$ first sees a particle propagating from $y$ to $x$; an observer who sees $x$ first sees an antiparticle propagating from $x$ to $y$. The two amplitudes are equal and are subtracted in the commutator, so they cancel. This is the mechanism by which no observable causal influence survives.

**3.** Without the $\hat b^\dagger$ term we would have $\hat\phi(x) = \int\frac{d^3p}{(2\pi)^3\sqrt{2\omega_{\boldsymbol p}}}\hat a_{\boldsymbol{p}}e^{-ip\cdot x}$ and hence $[\hat\phi(x),\hat\phi^\dagger(y)] = D(x-y)$. As noted in <Ref to="rem-antiparticle-necessity" />, at spacelike separation $D$ equals $\frac{m}{4\pi^2r}K_1(mr)>0$ and does not vanish. Microcausality would therefore fail, and two spacelike-separated measurements would influence one another.

In short, building a relativistic local field requires both annihilation and creation operators inside the same $\hat\phi(x)$. When the field carries charge ($\hat\phi\ne\hat\phi^\dagger$), the latter is necessarily the creation operator for a particle of the opposite charge, that is, for an antiparticle. For a neutral real scalar field (<Ref to="def-free-scalar-field" />) one has $\hat b = \hat a$, and the particle is its own antiparticle.
</Solution>
</Exercise>

## References

- M. E. Peskin and D. V. Schroeder, *An Introduction to Quantum Field Theory*, Addison-Wesley, 1995 — Chapter 1 and §2.1 (the failure of one-particle relativistic quantum mechanics and propagation outside the light cone).
- S. Weinberg, *The Quantum Theory of Fields, Volume I: Foundations*, Cambridge University Press, 1995 — Chapter 1 (historical introduction), Chapter 5 (the necessity of fields and spin-statistics).
- M. Sakamoto, *Ba no Ryōshiron: Fuhensei to Jiyūba o Chūshin ni shite*, Shōkabō, 2014 (in Japanese) — Chapters 1–3 (difficulties of the one-particle theory, canonical quantization).
- P. A. M. Dirac, "The Quantum Theory of the Electron", *Proceedings of the Royal Society A* **117** (1928), 610–624. [DOI: 10.1098/rspa.1928.0023](https://doi.org/10.1098/rspa.1928.0023)
- C. D. Anderson, "The Positive Electron", *Physical Review* **43** (1933), 491–494. [DOI: 10.1103/PhysRev.43.491](https://doi.org/10.1103/PhysRev.43.491)
- G. C. Hegerfeldt, "Remark on causality and particle localization", *Physical Review D* **10** (1974), 3320–3321. [DOI: 10.1103/PhysRevD.10.3320](https://doi.org/10.1103/PhysRevD.10.3320)

## Appendix: Contour deformation for the propagation amplitude outside the light cone

Let us derive the expression used in <Ref to="ex-propagator-tail" />. The starting point is

$$
U(t,r) = \frac{1}{2\pi^2 r}\int_0^\infty dp\;p\,\sin(pr)\,e^{-i\omega_p t},
\qquad \omega_p = \sqrt{p^2+m^2} .
$$

Both $p\sin(pr)$ and $\omega_p$ are even functions of $p$, so we may extend the range of integration to all of $\mathbb{R}$ and multiply by $\frac12$. Writing $\sin(pr) = (e^{ipr}-e^{-ipr})/(2i)$ and substituting $p\to-p$ in the $e^{-ipr}$ term, which turns it into a copy of the $e^{ipr}$ term, we obtain

$$
U(t,r) = \frac{1}{4\pi^2 i\,r}\int_{-\infty}^{\infty}dp\;p\,e^{ipr}\,e^{-i\sqrt{p^2+m^2}\,t} .
$$

Now continue $\sqrt{p^2+m^2}$ analytically into the complex $p$ plane. The branch points are $p=\pm im$; take the cuts along $[im, i\infty)$ and $(-i\infty,-im]$, and choose the branch with $\sqrt{p^2+m^2}>0$ on the real axis.

Since $r>0$, the factor $e^{ipr}$ decays in the upper half plane. On a large upper semicircle of radius $R$ we have $\sqrt{p^2+m^2}\simeq \pm p$ as $\lvert p\rvert\to\infty$, so the exponential part of the integrand behaves like $\lvert e^{ip(r-t)}\rvert = e^{-(r-t)\operatorname{Im}p}$. **Only for a spacelike separation $r>t>0$** does this decay and the arc contribution vanish. This is the decisive point: for a timelike separation $r<t$ the deformation below is not available.

Pushing the contour upward, it catches on the cut $[im,i\infty)$. What remains is a contour that descends the left side of the cut from $i\infty$ to $im$, rounds $im$, and ascends the right side back to $i\infty$. Setting $p = i\rho$ with $\rho>m$, we have $dp = i\,d\rho$ and $e^{ipr} = e^{-\rho r}$. Writing $p=\epsilon+i\rho$ and inspecting $p^2+m^2 \simeq (m^2-\rho^2) + 2i\epsilon\rho$, we see that for $\rho>m$ the real part is negative while the sign of the imaginary part matches that of $\epsilon$. Hence, with $s:=\sqrt{\rho^2-m^2}$, on the right side of the cut ($\epsilon>0$) we get $\sqrt{p^2+m^2} = +is$ and on the left side ($\epsilon<0$) we get $-is$. Therefore

$$
e^{-i\sqrt{p^2+m^2}\,t} = \begin{cases} e^{+ts} & (\text{right side}) \\ e^{-ts} & (\text{left side}) \end{cases}
$$

Adding the contributions of the left side ($\rho:\infty\to m$) and the right side ($\rho: m\to\infty$),

$$
\int_{-\infty}^{\infty}dp\;p\,e^{ipr}e^{-i\omega_pt}
= \int_{\infty}^{m}(i\rho)e^{-\rho r}e^{-ts}\,i\,d\rho + \int_{m}^{\infty}(i\rho)e^{-\rho r}e^{+ts}\,i\,d\rho
$$

$$
= \int_m^\infty d\rho\;\rho\,e^{-\rho r}\left(e^{-ts} - e^{+ts}\right)
= -2\int_m^\infty d\rho\;\rho\,e^{-\rho r}\sinh(ts).
$$

Substituting this back into the expression for $U$ and using $-2/(4i) = i/2$, we obtain

$$
U(t,r) = \frac{i}{2\pi^2 r}\int_m^\infty d\rho\;\rho\,e^{-\rho r}\,\sinh\!\left(t\sqrt{\rho^2-m^2}\right) .
$$

As $\rho\to\infty$ the integrand behaves like $\rho\,e^{-(r-t)\rho}/2$, so the integral converges for $r>t$. For $t>0$ the integrand is strictly positive throughout the range, so the integral is positive, that is, $U\ne 0$. As $t\to 0$ we have $\sinh\to0$ and hence $U\to0$, consistent with the value of $\delta^3(\boldsymbol{r})$ for $r>0$.
