Skip to content

Relativistic Mechanics: What the Four-Momentum Says About E = mc²

Prerequisite:Lorentz Transformations: Time Dilation and Length Contraction from the Light Clock

Raw
  • The conservation law for the Newtonian momentum mvm\boldsymbol v may hold in one inertial frame and fail in another. The reason is that velocities compose by the Lorentz transformation, not by the Galilean one. If we want to keep the conservation law, we have no choice but to rebuild the definition of momentum.
  • The guiding principle for the rebuilding is: use only quantities that transform under Lorentz transformations by the same rule as the spacetime coordinates xμx^\mu — that is, four-vectors. The leading roles go to the four-velocity uμu^\mu, obtained by differentiating with respect to proper time τ\tau, and to the four-momentum pμ=muμp^\mu = m u^\mu.
  • The spatial part of pμp^\mu is p=γmv\boldsymbol p = \gamma m \boldsymbol v and its time component is p0=γmcp^0 = \gamma m c. Multiplying the latter by cc gives E=γmc2E = \gamma m c^2, the relativistic energy, which at low speeds reduces to mc2+12mv2mc^2 + \frac{1}{2}mv^2.
  • If we demand conservation of momentum in every inertial frame, conservation of energy follows automatically. The two are merely different components of a single four-vector conservation law.
  • A body at rest still carries the energy E=mc2E = mc^2. Unlike in Newtonian mechanics, this “constant” is observable: when a reaction changes which particles are present, m\sum m changes, and the difference is exactly the energy that flows in or out.
  • Mass is not additive. The invariant mass of a system is fixed by M2c4=(E)2p2c2M^2c^4 = \left(\sum E\right)^2 - \left|\sum \boldsymbol p\right|^2 c^2, and it is always greater than or equal to the sum of the masses of the constituents. The binding energy of deuterium, and the fact that the Sun shines, are consequences of this one sentence.

1. Motivation: where Newtonian mechanics breaks down

Section titled “1. Motivation: where Newtonian mechanics breaks down”

In the previous article we saw that the coordinate transformation between inertial frames is the Lorentz transformation(Theorem 6.1)[Lorentz Transformations] rather than the Galilean one. Time dilation and length contraction both follow from it. But if the transformation law has changed, then the laws of mechanics that were written down for the old law must be re-examined as well.

The backbone of Newtonian mechanics is the momentum pN=mv\boldsymbol p_{\text{N}} = m\boldsymbol v together with its conservation law. This is a law confirmed by experiment to an extraordinary degree, and as we saw in conservation of momentum(Theorem 6.1)[Foundations of Newtonian Mechanics] in Foundations of Newtonian mechanics, it is central enough to be equivalent to the law of action and reaction. Yet this law does not keep its form under Lorentz transformations. The following computation shows it.

Example 1.1Newtonian momentum conservation is frame-dependent

Consider two lumps of clay A and B of equal mass. Viewed from the inertial frame S, A moves with velocity +u+u and B with velocity u-u; they collide head-on and stick together into a single lump. The Newtonian momentum before the collision is mu+m(u)=0mu + m(-u) = 0, so if we accept the conservation law the merged lump is at rest in S.

Now change to the inertial frame S′, which moves with velocity u-u along the xx axis relative to S. By the velocity-addition rule (Corollary 6.3[Lorentz Transformations] of the previous article), a body with velocity ww in S has velocity

w=w+u1+wu/c2w' = \frac{w + u}{1 + wu/c^2}

in S′. Apply this to the three bodies. Writing βu/c\beta \equiv u/c,

  • A (w=uw = u): wA=2u1+β2\displaystyle w'_{\mathrm A} = \frac{2u}{1+\beta^2}
  • B (w=uw = -u): wB=u+u1β2=0\displaystyle w'_{\mathrm B} = \frac{-u+u}{1-\beta^2} = 0
  • the merged lump (w=0w = 0): w=uw' = u

Compute the Newtonian momentum in S′, assuming that masses add (so that the merged lump has mass 2m2m).

before=m2u1+β2+m0=2mu1+β2,after=2mu=2mu.\begin{aligned} \text{before} &= m\cdot\frac{2u}{1+\beta^2} + m\cdot 0 = \frac{2mu}{1+\beta^2},\\ \text{after} &= 2m\cdot u = 2mu. \end{aligned}

Their ratio is 1/(1+β2)1/(1+\beta^2), which is never 11 as long as u0u \neq 0. For instance with u=0.6cu = 0.6c we have β2=0.36\beta^2 = 0.36, so the value before the collision is 2mu/1.36=1.47mu2mu/1.36 = 1.47\,mu while after it is 2mu2mu — a discrepancy of 26%26\,\%.

In other words, Newtonian momentum is conserved in S but not in S′.

This is serious. The principle of relativity (the principle of special relativity(Axiom 5.2)[The Principles of Special Relativity] in Principles of special relativity) demands that the laws of physics take the same form in every inertial frame, so a conservation law that holds only in one frame forfeits its status as a law of physics. There are only two ways out.

  1. Abandon conservation of momentum.
  2. Rebuild the definition of momentum (or the assumption that masses add).

Experiment forbids the first. That a conserved quantity exists in collisions has been confirmed on every scale from elementary particles to astronomical bodies. Moreover, by Noether’s theorem, the existence of the conserved quantity is a restatement of the homogeneity of space, so discarding it would amount to discarding a symmetry. Hence only the second road is open.

The same conclusion is visible from another direction. In Newtonian mechanics the kinetic energy K=12mv2K = \frac{1}{2}mv^2 has no upper bound, so quadrupling KK should double vv. Yet however much energy we pour into accelerating an electron, the speed measured by time of flight merely sticks to cc and never exceeds it. Bertozzi’s 1962 experiment (reference [5]) showed directly that raising the kinetic energy of electrons from 0.5 MeV0.5\ \mathrm{MeV} to 15 MeV15\ \mathrm{MeV} — a factor of 3030 — moves the time-of-flight value of v2/c2v^2/c^2 only from 0.740.74 to 0.9990.999, never past 11. The very relation between KK and vv is different.

In this article we follow the second road to its end. Its destination is E=mc2E = mc^2.

2. Preliminaries: Lorentz transformations and four-vectors

Section titled “2. Preliminaries: Lorentz transformations and four-vectors”

Let cc denote the speed of light and write γ(w)(1w2/c2)1/2\gamma(w) \equiv \left(1 - w^2/c^2\right)^{-1/2}. We use VV for the relative velocity between inertial frames and v\boldsymbol v (with magnitude vv) for the velocity of a particle, abbreviating γγ(v)\gamma \equiv \gamma(v) for particles. We set βV/c\beta \equiv V/c.

When S′ moves with velocity VV in the positive xx direction relative to S, the Lorentz transformation (boost) reads

ct=γ(V)(ctβx),x=γ(V)(xβct),y=y,z=z\begin{aligned} ct' &= \gamma(V)\left(ct - \beta x\right), \qquad & x' &= \gamma(V)\left(x - \beta\, ct\right),\\ y' &= y, \qquad & z' &= z \end{aligned}

as before. Note that the time coordinate has been rescaled to ctct, giving it the dimension of a length. All four coordinates then carry the same dimension, and we may write them together as xμ=(x0,x1,x2,x3)=(ct,x,y,z)x^\mu = (x^0, x^1, x^2, x^3) = (ct, x, y, z) with μ=0,1,2,3\mu = 0,1,2,3.

The Lorentz transformation is linear in xμx^\mu, and it mixes x0x^0 with x1x^1. From this the notion of “a quadruple that mixes in the same way” arises naturally.

Definition 2.1Four-vectors and the Minkowski product

Suppose that in each inertial frame a quadruple of numbers Aμ=(A0,A1,A2,A3)A^\mu = (A^0, A^1, A^2, A^3) is given, and that under the boost above it transforms as

A0=γ(V)(A0βA1),A1=γ(V)(A1βA0),A2=A2,A3=A3.\begin{aligned} A'^0 &= \gamma(V)\left(A^0 - \beta A^1\right), \qquad & A'^1 &= \gamma(V)\left(A^1 - \beta A^0\right),\\ A'^2 &= A^2, \qquad & A'^3 &= A^3. \end{aligned}

Then AμA^\mu is called a four-vector. For two four-vectors AμA^\mu and BμB^\mu we call

ABA0B0A1B1A2B2A3B3A\cdot B \equiv A^0B^0 - A^1B^1 - A^2B^2 - A^3B^3

the Minkowski product, and AAA\cdot A the squared norm of AμA^\mu.

It follows from the definition that the difference of spacetime coordinates Δxμ=(cΔt,Δx,Δy,Δz)\Delta x^\mu = (c\Delta t, \Delta x, \Delta y, \Delta z) is a four-vector (the Lorentz transformation is linear, so differences obey the same formulas). Its squared norm

ΔxΔx=c2Δt2Δx2Δy2Δz2\Delta x \cdot \Delta x = c^2\Delta t^2 - \Delta x^2 - \Delta y^2 - \Delta z^2

is the invariant interval of the previous article (invariance of the spacetime interval(Theorem 7.2)[Lorentz Transformations]). In fact it is not only the norm of a coordinate difference that is invariant.

Proposition 2.2Lorentz invariance of the Minkowski product

If AμA^\mu and BμB^\mu are four-vectors, then AB=ABA'\cdot B' = A\cdot B for every boost.

Proof(Proposition 2.2)

Since A2A^2 and A3A^3 are unchanged by the transformation, A2B2+A3B3=A2B2+A3B3A'^2B'^2 + A'^3B'^3 = A^2B^2 + A^3B^3. It therefore suffices to examine the 00 and 11 components. Substituting the transformation law of Definition 2.1,

A0B0A1B1=γ2[(A0βA1)(B0βB1)(A1βA0)(B1βB0)]=γ2[A0B0βA0B1βA1B0+β2A1B1A1B1+βA1B0+βA0B1β2A0B0]=γ2[(1β2)A0B0(1β2)A1B1].\begin{aligned} A'^0B'^0 - A'^1B'^1 &= \gamma^2\Big[\left(A^0-\beta A^1\right)\left(B^0-\beta B^1\right) - \left(A^1-\beta A^0\right)\left(B^1-\beta B^0\right)\Big]\\ &= \gamma^2\Big[A^0B^0 - \beta A^0B^1 - \beta A^1B^0 + \beta^2A^1B^1\\ &\qquad\quad - A^1B^1 + \beta A^1B^0 + \beta A^0B^1 - \beta^2A^0B^0\Big]\\ &= \gamma^2\Big[(1-\beta^2)A^0B^0 - (1-\beta^2)A^1B^1\Big]. \end{aligned}

The essential point is that the cross terms cancel in pairs: βA0B1-\beta A^0B^1 against +βA0B1+\beta A^0B^1, and βA1B0-\beta A^1B^0 against +βA1B0+\beta A^1B^0. Finally, γ2=(1β2)1\gamma^2 = (1-\beta^2)^{-1} gives γ2(1β2)=1\gamma^2(1-\beta^2) = 1, so

A0B0A1B1=A0B0A1B1A'^0B'^0 - A'^1B'^1 = A^0B^0 - A^1B^1

as claimed. Boosts along directions other than the xx axis, and spatial rotations, reduce to the same computation after a suitable choice of axes.

Proposition 2.2 is the tool used throughout this article. The product of two four-vectors is the same number no matter which inertial frame computes it. That is why a quantity built from such a product carries meaning as a quantity intrinsic to the object. We shall use this single fact again and again.

To construct a momentum we must first construct a velocity. The naive attempt dxμ/dtdx^\mu/dt is not a four-vector, and the reason is plain: the numerator dxμdx^\mu is a four-vector, but the denominator dtdt is a frame-dependent quantity (one component, x0/cx^0/c). Dividing a four-vector by something that is not invariant destroys the transformation law.

What we need is a time on whose value everyone agrees. That is proper time.

Definition 3.1Proper time

Let dxμ=(cdt,dx)dx^\mu = (c\,dt, d\boldsymbol x) be an infinitesimal displacement along the world line of a particle, with dxdx>0dx\cdot dx > 0 (timelike). The quantity dτ>0d\tau > 0 defined by

c2dτ2dxdx=c2dt2dx2c^2\,d\tau^2 \equiv dx\cdot dx = c^2dt^2 - |d\boldsymbol x|^2

is called the infinitesimal change of proper time, and its integral τ=dτ\tau = \int d\tau along the world line is the proper time of the particle.

Since dτd\tau is built out of a product of four-vectors, it is Lorentz invariant by Proposition 2.2. Proper time thus has the same value in every inertial frame, and physically it is nothing other than the time read by a clock attached to the particle. It is the quantity proper time(Definition 3.2)[Lorentz Transformations] of the previous article, rewritten in terms of an infinitesimal displacement along the world line.

Proposition 3.2Proper time versus coordinate time

If a particle has velocity v\boldsymbol v (of magnitude v<cv < c) in the inertial frame S, then the coordinate time tt of S and the proper time τ\tau satisfy

dtdτ=γ(v)=11v2/c2.\frac{dt}{d\tau} = \gamma(v) = \frac{1}{\sqrt{1 - v^2/c^2}}.
Proof(Proposition 3.2)

Factor dt2dt^2 out of the right-hand side of Definition 3.1. By the definition of velocity, v=dx/dt\boldsymbol v = d\boldsymbol x/dt, so dx=vdt|d\boldsymbol x| = v\,dt and

c2dτ2=c2dt2dx2=c2dt2v2dt2=c2dt2(1v2c2).c^2 d\tau^2 = c^2dt^2 - |d\boldsymbol x|^2 = c^2dt^2 - v^2dt^2 = c^2dt^2\left(1 - \frac{v^2}{c^2}\right).

Dividing both sides by c2dt2c^2dt^2 gives (dτ/dt)2=1v2/c2\left(d\tau/dt\right)^2 = 1 - v^2/c^2. Since v<cv < c, the right-hand side is positive, and as we have taken dτ>0d\tau > 0 and dt>0dt > 0 the sign of the square root is fixed to be positive:

dτdt=1v2c2=1γ(v),\frac{d\tau}{dt} = \sqrt{1 - \frac{v^2}{c^2}} = \frac{1}{\gamma(v)},

that is, dt/dτ=γ(v)dt/d\tau = \gamma(v).

Since γ(v)1\gamma(v) \ge 1, we get dtdτdt \ge d\tau: the conclusion of the previous article that a moving clock runs slow (time dilation(Theorem 3.3)[Lorentz Transformations]) is reproduced here by the very same formula.

Definition 3.3Four-velocity

For the world line xμ(τ)x^\mu(\tau) of a particle of non-zero mass, the quantity

uμdxμdτu^\mu \equiv \frac{dx^\mu}{d\tau}

is called the four-velocity.

Theorem 3.4Properties of the four-velocity

For a particle of non-zero mass the following hold.

  1. uμu^\mu is a four-vector.
  2. If the particle has velocity v\boldsymbol v in the inertial frame S, its components are uμ=γ(v)(c, v)u^\mu = \gamma(v)\,(c,\ \boldsymbol v).
  3. Its squared norm is uu=c2u\cdot u = c^2, independently of the velocity.
Proof(Theorem 3.4)

(1) The numerator dxμdx^\mu is a four-vector (as noted just after Definition 2.1, coordinate differences obey the boost formulas directly, by linearity of the Lorentz transformation). The denominator dτd\tau is a Lorentz-invariant scalar by Proposition 2.2. Dividing a four-vector by a scalar multiplies every component by the same constant, so the transformation law is unaffected. Hence uμu^\mu is a four-vector.

(2) Rewrite in terms of coordinate time by the chain rule. Since dt/dτ=γ(v)dt/d\tau = \gamma(v) by Proposition 3.2,

uμ=dxμdτ=dtdτdxμdt=γ(v)(d(ct)dt, dxdt)=γ(v)(c, v).u^\mu = \frac{dx^\mu}{d\tau} = \frac{dt}{d\tau}\cdot\frac{dx^\mu}{dt} = \gamma(v)\left(\frac{d(ct)}{dt},\ \frac{d\boldsymbol x}{dt}\right) = \gamma(v)\,(c,\ \boldsymbol v).

(3) Insert the components from (2) into the product of Definition 2.1:

uu=γ(v)2(c2v2)=c2v21v2/c2=c2(1v2/c2)1v2/c2=c2.u\cdot u = \gamma(v)^2\left(c^2 - |\boldsymbol v|^2\right) = \frac{c^2 - v^2}{1 - v^2/c^2} = \frac{c^2\left(1 - v^2/c^2\right)}{1 - v^2/c^2} = c^2.

The key step is factoring c2v2=c2(1v2/c2)c^2 - v^2 = c^2(1 - v^2/c^2): the velocity cancels and the constant c2c^2 survives.

Alternatively, (3) follows without any computation: in the rest frame of the particle (v=0\boldsymbol v = \boldsymbol 0) we have uμ=(c,0)u^\mu = (c, \boldsymbol 0) and hence uu=c2u\cdot u = c^2, and the value is the same in every frame by Proposition 2.2.

Remark 3.5

The relation uu=c2u\cdot u = c^2 may be read as saying that every object moves through spacetime at the constant rate cc. A body at rest advances at rate cc purely in the time direction, while a fast-moving body diverts part of that cc into the spatial directions. Time dilation corresponds to the resulting decrease of the time component.

In Newtonian mechanics momentum was “mass times velocity”. In relativity we put the four-velocity in place of the velocity. For the mass we use the value measured in the rest frame of the particle (the rest mass, from now on simply the mass). It is a constant attached to each particle, a number independent of the frame.

Definition 4.1Four-momentum

The four-momentum of a particle of mass m>0m > 0 is defined by

pμmuμ=mγ(v)(c, v).p^\mu \equiv m\,u^\mu = m\gamma(v)\,(c,\ \boldsymbol v).

Its spatial part is called the relativistic momentum pγ(v)mv\boldsymbol p \equiv \gamma(v)\,m\boldsymbol v, and its time component multiplied by cc is called the energy Ecp0=γ(v)mc2E \equiv c\,p^0 = \gamma(v)\,mc^2, so that pμ=(E/c, p)p^\mu = (E/c,\ \boldsymbol p).

Since mm is an invariant scalar and uμu^\mu is a four-vector (Theorem 3.4 (1)), pμp^\mu is a four-vector as well. For vcv \ll c we have γ1\gamma \approx 1 and hence pmv\boldsymbol p \approx m\boldsymbol v, recovering the Newtonian momentum. So this is the Newtonian momentum, modified by the least amount needed to make it a four-vector.

We have not yet explained why the name EE is deserved; that will become clear in Theorem 5.1. First let us see why this definition rescues the conservation law.

4.2. Conservation independent of the frame

Section titled “4.2. Conservation independent of the frame”

Theorem 4.2Frame independence of four-momentum conservation

In a process such as a collision, a decay or a creation event, write inpμ\sum_{\text{in}} p^\mu for the sum of the four-momenta of the incoming particles and outpμ\sum_{\text{out}} p^\mu for the sum for the outgoing ones.

  1. If inpμ=outpμ\sum_{\text{in}} p^\mu = \sum_{\text{out}} p^\mu (all four components) holds in one inertial frame, then it holds in every inertial frame.
  2. Moreover, if conservation of the spatial part inp=outp\sum_{\text{in}} \boldsymbol p = \sum_{\text{out}} \boldsymbol p holds in every inertial frame, then conservation of the time component inE=outE\sum_{\text{in}} E = \sum_{\text{out}} E holds automatically.
Proof(Theorem 4.2)

Consider the difference Δμinpμoutpμ\Delta^\mu \equiv \sum_{\text{in}} p^\mu - \sum_{\text{out}} p^\mu. The transformation law of Definition 2.1 is linear, so sums and differences of four-vectors are again four-vectors. Hence Δμ\Delta^\mu is a four-vector.

(1) The hypothesis is that Δμ=0\Delta^\mu = 0 (all four components vanish) in some frame S. For an arbitrary boost,

Δ0=γ(V)(Δ0βΔ1)=0,Δ1=γ(V)(Δ1βΔ0)=0,\Delta'^0 = \gamma(V)\left(\Delta^0 - \beta\Delta^1\right) = 0,\qquad \Delta'^1 = \gamma(V)\left(\Delta^1 - \beta\Delta^0\right) = 0,

and Δ2=Δ2=0\Delta'^2 = \Delta^2 = 0, Δ3=Δ3=0\Delta'^3 = \Delta^3 = 0. What matters here is that the transformation is a homogeneous linear map, with no constant term. Hence Δμ=0\Delta'^\mu = 0 in every frame, which is the conservation law.

(2) The hypothesis is that the spatial components vanish in every frame, that is, Δ1=Δ2=Δ3=0\Delta^1 = \Delta^2 = \Delta^3 = 0 and Δ1=Δ2=Δ3=0\Delta'^1 = \Delta'^2 = \Delta'^3 = 0. Taking a boost with velocity V0V \neq 0 along xx, the transformation law gives

0=Δ1=γ(V)(Δ1βΔ0)=γ(V)(0βΔ0)=γ(V)βΔ0.0 = \Delta'^1 = \gamma(V)\left(\Delta^1 - \beta\Delta^0\right) = \gamma(V)\left(0 - \beta\Delta^0\right) = -\gamma(V)\,\beta\,\Delta^0 .

Since γ(V)1>0\gamma(V) \ge 1 > 0 and β=V/c0\beta = V/c \neq 0, we must have Δ0=0\Delta^0 = 0. As E=cp0E = cp^0, this is conservation of energy, inE=outE\sum_{\text{in}} E = \sum_{\text{out}} E.

Part (1) of Theorem 4.2 repairs exactly what broke in Example 1.1. The Newtonian momentum mvm\boldsymbol v is not part of a four-vector, so conservation failed upon changing frames. The relativistic p=γmv\boldsymbol p = \gamma m\boldsymbol v is the spatial part of the four-vector pμp^\mu, so as long as it is conserved together with the time component, it is conserved in every frame.

Part (2) is one of the most beautiful consequences of the theory. In relativity one cannot postulate momentum conservation and energy conservation as separate laws. Demanding either one in all inertial frames brings the other with it. Two laws that were independent in Newtonian mechanics are here unified into a single four-vector conservation law.

Remark 4.3

In the language of Noether’s theorem, momentum conservation was the consequence of invariance under spatial translations (conservation of total momentum(Example 4.3)[対称性と保存則]) and energy conservation the consequence of invariance under time translations (homogeneity of time and conservation of energy(Corollary 5.4)[対称性と保存則]). In relativity a boost mixes time with space, so the two symmetries mix as well. The merging of the conservation laws into one is a reflection of this.

Theorem 4.4Energy–momentum relation

For a free particle of mass mm, the energy EE and momentum p\boldsymbol p satisfy

E2=p2c2+m2c4E^2 = |\boldsymbol p|^2c^2 + m^2c^4

in every inertial frame. In particular, at rest (p=0\boldsymbol p = \boldsymbol 0) we have E=mc2E = mc^2.

Proof(Theorem 4.4)

By Definition 4.1 we have pμ=muμp^\mu = m u^\mu, so the product of Definition 2.1 together with Theorem 3.4 (3) gives

pp=m2(uu)=m2c2.p\cdot p = m^2\,(u\cdot u) = m^2c^2 .

On the other hand, computing the product directly from the components pμ=(E/c,p)p^\mu = (E/c, \boldsymbol p),

pp=(Ec)2p2.p\cdot p = \left(\frac{E}{c}\right)^2 - |\boldsymbol p|^2 .

Equating the two gives (E/c)2p2=m2c2\left(E/c\right)^2 - |\boldsymbol p|^2 = m^2c^2, and multiplying by c2c^2 yields E2p2c2=m2c4E^2 - |\boldsymbol p|^2c^2 = m^2c^4.

The left-hand side is Lorentz invariant by Proposition 2.2, so the identity holds in the same form whichever frame measures the components. Setting p=0\boldsymbol p = \boldsymbol 0 gives E2=m2c4E^2 = m^2c^4, and E>0E > 0 gives E=mc2E = mc^2.

In practice it is convenient to remember this relation as a right triangle with EE as the hypotenuse.

p cm c²θEp → 0 gives E → mc²m → 0 gives E → pcsin θ = pc / E = v / c
The energy–momentum relation E² = (pc)² + (mc²)² drawn as a right triangle. The vertical leg is the rest energy mc², the horizontal leg is pc, and the hypotenuse is the total energy E. The angle θ satisfies sin θ = pc/E = v/c. For a particle at rest the horizontal leg vanishes and E = mc²; for a massless particle the vertical leg vanishes and E = pc.

Corollary 4.5Recovering the velocity, and massless particles

  1. For a particle of mass m>0m > 0 we have v=pc2E\displaystyle \boldsymbol v = \frac{\boldsymbol p\,c^2}{E}.
  2. Taking the limit m0m \to 0 while keeping p0\boldsymbol p \neq \boldsymbol 0 fixed, we get E=pcE = |\boldsymbol p|c and v=cv = c. A particle in this limit (a photon, for instance) has zero mass yet non-zero energy and momentum.
Proof(Corollary 4.5)

(1) Simply divide p=γmv\boldsymbol p = \gamma m\boldsymbol v by E=γmc2E = \gamma mc^2 from Definition 4.1:

pc2E=γmvc2γmc2=v.\frac{\boldsymbol p\,c^2}{E} = \frac{\gamma m \boldsymbol v \, c^2}{\gamma m c^2} = \boldsymbol v .

The factors γ\gamma, mm and c2c^2 simply cancel.

(2) Setting m=0m = 0 in Theorem 4.4 gives E2=p2c2E^2 = |\boldsymbol p|^2c^2, and E>0E > 0 gives E=pcE = |\boldsymbol p|c. Substituting this into (1),

v=pc2E=pc2pc=c.v = \frac{|\boldsymbol p|c^2}{E} = \frac{|\boldsymbol p|c^2}{|\boldsymbol p|c} = c .

Conversely, for a particle with v=cv = c the factor γ(v)\gamma(v) diverges, so E=γmc2E = \gamma mc^2 can be finite only if m=0m = 0. Massless particles can travel only at the speed of light, and particles travelling at the speed of light cannot have mass.

Remark 4.6

Older textbooks speak of a “mass that grows with speed”, mrel=γmm_{\text{rel}} = \gamma m (the relativistic mass). It looks convenient, since one can then write p=mrelv\boldsymbol p = m_{\text{rel}}\boldsymbol v and E=mrelc2E = m_{\text{rel}}c^2; but F=mrela\boldsymbol F = m_{\text{rel}}\boldsymbol a fails (force and acceleration need not be parallel), so the notion causes more confusion than it removes. The modern standard is that “mass” always means the invariant rest mass mm, with all velocity dependence pushed into γ\gamma. We follow that convention here. The history is discussed in detail in Okun’s article (reference [6]).

5. Relativistic kinetic energy and rest energy

Section titled “5. Relativistic kinetic energy and rest energy”

In Definition 4.1 we called E=γmc2E = \gamma mc^2 the “energy”, but so far that is only a name. The operational definition that fixes what energy is reads “the work done by an external force on a body initially at rest”, so let us compute it. As the equation of motion we adopt Newton’s second law in the form

F=dpdt,p=γ(v)mv.\boldsymbol F = \frac{d\boldsymbol p}{dt},\qquad \boldsymbol p = \gamma(v)\,m\boldsymbol v .

The point is to keep the form dp/dtd\boldsymbol p/dt rather than F=ma\boldsymbol F = m\boldsymbol a; in this form the law is directly compatible with conservation of four-momentum.

Theorem 5.1Relativistic kinetic energy

Suppose a particle of mass mm starts from rest and, under an external force F=dp/dt\boldsymbol F = d\boldsymbol p/dt, reaches the speed vv (with v<cv < c). Then the work KK done by the force equals

K=(γ(v)1)mc2=mc21v2/c2mc2,K = (\gamma(v) - 1)\,mc^2 = \frac{mc^2}{\sqrt{1-v^2/c^2}} - mc^2 ,

and this quantity is called the relativistic kinetic energy. Equivalently, E=γmc2=mc2+KE = \gamma mc^2 = mc^2 + K.

Proof(Theorem 5.1)

To keep the notation simple we treat one-dimensional motion, with force and motion along the xx axis; the general case is described at the end. From the definition of work and the equation of motion,

K=0xFdx=0xdpdtdx.K = \int_0^x F\,dx' = \int_0^x \frac{dp}{dt}\,dx' .

Change the integration variable from xx' to pp. Since dx=vdtdx' = v\,dt, we have dpdtdx=dpdtvdt=vdp\dfrac{dp}{dt}\,dx' = \dfrac{dp}{dt}\,v\,dt = v\,dp, and the particle starts from rest (p=0p = 0) and reaches momentum p=γmvp = \gamma m v, so

K=0γmvvdp.K = \int_0^{\gamma m v} v\,dp .

Integrate by parts. Since vdp=[vp]pdv\int v\,dp = \left[vp\right] - \int p\,dv,

K=[vγ(v)mv]v=0v0vγ(w)mwdw=γ(v)mv2m0vwdw1w2/c2.K = \Big[v\cdot \gamma(v) m v\Big]_{v=0}^{v} - \int_0^{v} \gamma(w)\,m w\,dw = \gamma(v)\,mv^2 - m\int_0^v \frac{w\,dw}{\sqrt{1-w^2/c^2}} .

Now evaluate the remaining integral. Substituting s=1w2/c2s = 1 - w^2/c^2 gives ds=2wdw/c2ds = -2w\,dw/c^2, that is, wdw=c22dsw\,dw = -\tfrac{c^2}{2}ds, and the range w:0vw: 0 \to v becomes s:11v2/c2s: 1 \to 1 - v^2/c^2:

0vwdw1w2/c2=11v2/c2c22dss=c22[2s]11v2/c2=c2(11v2c2)=c2(11γ(v)).\int_0^v \frac{w\,dw}{\sqrt{1-w^2/c^2}} = \int_1^{1-v^2/c^2} \frac{-\tfrac{c^2}{2}\,ds}{\sqrt{s}} = -\frac{c^2}{2}\Big[2\sqrt{s}\Big]_1^{1-v^2/c^2} = c^2\left(1 - \sqrt{1-\frac{v^2}{c^2}}\right) = c^2\left(1 - \frac{1}{\gamma(v)}\right).

Putting this back,

K=γmv2mc2(11γ)=mc2(γv2c21+1γ).K = \gamma m v^2 - mc^2\left(1 - \frac{1}{\gamma}\right) = mc^2\left(\gamma\,\frac{v^2}{c^2} - 1 + \frac{1}{\gamma}\right).

Simplify the bracket. Since 1γ=1v2c2=γ(1v2c2)\dfrac{1}{\gamma} = \sqrt{1-\dfrac{v^2}{c^2}} = \gamma\left(1-\dfrac{v^2}{c^2}\right) (the last step uses γ2=1v2/c2\gamma^{-2} = 1 - v^2/c^2),

γv2c2+1γ=γv2c2+γ(1v2c2)=γ.\gamma\,\frac{v^2}{c^2} + \frac{1}{\gamma} = \gamma\,\frac{v^2}{c^2} + \gamma\left(1 - \frac{v^2}{c^2}\right) = \gamma .

Hence K=mc2(γ1)K = mc^2(\gamma - 1).

The conclusion is the same for general three-dimensional motion. Differentiating both sides of Theorem 4.4 gives 2EdE=2c2pdp2E\,dE = 2c^2\,\boldsymbol p\cdot d\boldsymbol p (as mm is constant), so using Corollary 4.5 (1),

dE=c2pEdp=vdp=vFdt,dE = \frac{c^2\boldsymbol p}{E}\cdot d\boldsymbol p = \boldsymbol v\cdot d\boldsymbol p = \boldsymbol v\cdot\boldsymbol F\,dt ,

and the right-hand side is precisely the work done by the external force during dtdt. So the total work equals the increment of EE, and the increment of E=γmc2E = \gamma mc^2 from γ=1\gamma = 1 is (γ1)mc2(\gamma-1)mc^2.

012300.51β = v / cK / mc²v = crelativistic K = (γ - 1) mc²Newtonian K = mv² / 2
Comparison of kinetic energies. The solid curve is the relativistic K = (γ - 1)mc², the dashed one is the Newtonian K = mv²/2. Up to about β = v/c ≈ 0.2 the two coincide, but as the speed approaches c the relativistic curve diverges: no amount of energy brings v up to c.

Proposition 5.2The Newtonian limit and its error

With β=v/c\beta = v/c, we have for β<1|\beta| < 1 the expansion

K=(γ1)mc2=12mv2(1+34β2+58β4+).K = (\gamma - 1)mc^2 = \frac{1}{2}mv^2\left(1 + \frac{3}{4}\beta^2 + \frac{5}{8}\beta^4 + \cdots\right).

In particular K12mv212mv2=34β2+O(β4)\dfrac{K - \frac{1}{2}mv^2}{\frac{1}{2}mv^2} = \dfrac{3}{4}\beta^2 + O(\beta^4), so the relative error of the Newtonian expression is of order β2\beta^2.

Proof(Proposition 5.2)

Substitute x=β2x = \beta^2 into the binomial series (1x)1/2=1+12x+38x2+516x3+(1-x)^{-1/2} = 1 + \frac{1}{2}x + \frac{3}{8}x^2 + \frac{5}{16}x^3 + \cdots (valid for x<1|x| < 1):

γ=(1β2)1/2=1+12β2+38β4+516β6+.\gamma = \left(1-\beta^2\right)^{-1/2} = 1 + \frac{1}{2}\beta^2 + \frac{3}{8}\beta^4 + \frac{5}{16}\beta^6 + \cdots .

Therefore

K=mc2(γ1)=mc2(12β2+38β4+516β6+)=12mc2β2(1+34β2+58β4+).K = mc^2(\gamma - 1) = mc^2\left(\frac{1}{2}\beta^2 + \frac{3}{8}\beta^4 + \frac{5}{16}\beta^6 + \cdots\right) = \frac{1}{2}mc^2\beta^2\left(1 + \frac{3}{4}\beta^2 + \frac{5}{8}\beta^4 + \cdots\right).

Since mc2β2=mv2mc^2\beta^2 = mv^2, the leading term is 12mv2\frac{1}{2}mv^2. Rearranging the bracket gives the formula for the relative error.

This justifies the name. The velocity-dependent part of E=γmc2E = \gamma mc^2 is exactly the Newtonian kinetic energy at low speeds, so we are entitled to call EE an energy. And by Theorem 4.2 the sum of these EE is conserved across a reaction.

Example 5.3How close to the speed of light is an LHC proton?

The rest energy of the proton is mpc2=938.272 MeVm_pc^2 = 938.272\ \mathrm{MeV}. In Run 3 of the LHC, the protons in a single beam carry the energy E=6.8 TeV=6.8×106 MeVE = 6.8\ \mathrm{TeV} = 6.8\times 10^6\ \mathrm{MeV}. By Definition 4.1,

γ=Empc2=6.8×106938.272=7.25×103.\gamma = \frac{E}{m_pc^2} = \frac{6.8\times10^6}{938.272} = 7.25\times10^3 .

Now find the speed. From γ2=1β2\gamma^{-2} = 1-\beta^2 we get β=1γ2\beta = \sqrt{1-\gamma^{-2}}, and since γ2=1/(7.248×103)2=1.904×108\gamma^{-2} = 1/(7.248\times10^3)^2 = 1.904\times10^{-8} is very small we may use 1x1x/2\sqrt{1-x} \approx 1 - x/2:

β112γ2=19.52×109.\beta \approx 1 - \frac{1}{2\gamma^2} = 1 - 9.52\times10^{-9}.

The shortfall relative to the speed of light is cv=9.52×109×2.998×108 m/s=2.85 m/sc - v = 9.52\times10^{-9}\times 2.998\times10^8\ \mathrm{m/s} = 2.85\ \mathrm{m/s} — barely faster than walking pace. The LHC ring is 26.66 km26.66\ \mathrm{km} around, so one revolution takes 26.66×103/(2.998×108)=8.89×105 s26.66\times10^3 / (2.998\times10^8) = 8.89\times10^{-5}\ \mathrm{s}, and per revolution light gets ahead of the proton by 2.85×8.89×105=2.5×104 m2.85 \times 8.89\times10^{-5} = 2.5\times10^{-4}\ \mathrm{m}, that is, by 0.25 mm0.25\ \mathrm{mm}.

What would Newtonian mechanics predict for the same kinetic energy? Solving 12mpv2=6.8 TeV\frac{1}{2}m_pv^2 = 6.8\ \mathrm{TeV} gives v=c2×7.25×103=120cv = c\sqrt{2\times 7.25\times10^3} = 120\,c, a hundred and twenty times the speed of light. The measured speed of course never exceeds cc. It is the divergence of γ1\gamma - 1 as vcv \to c in Theorem 5.1 that guarantees this bound.

Setting p=0\boldsymbol p = \boldsymbol 0 in Theorem 4.4, or v=0v = 0 in Theorem 5.1, gives the famous formula

E0=mc2.E_0 = mc^2 .

This is called the rest energy. The formula itself follows almost trivially from Definition 4.1, but its content is not trivial at all. The objection to answer is: “Have we not merely shifted the zero point of energy?” In Newtonian mechanics the zero of potential energy is arbitrary, and adding a constant changes no physics.

The answer is this. mc2mc^2 is indeed a constant, but imic2\sum_i m_i c^2 is not. If a reaction changes which particles are present, the total mass changes. By Theorem 4.2 the total E\sum E is conserved, so if mic2\sum m_ic^2 decreases, exactly that amount must appear as kinetic energy or radiated energy. In Newtonian mechanics the number and kind of particles were assumed fixed, which is the only reason the zero point could be removed; in a world where particles are created and annihilated it can no longer be removed.

The size of the coefficient c2=8.99×1016 m2/s2c^2 = 8.99\times10^{16}\ \mathrm{m^2/s^2} is what gives the formula its bite. A mass of 1 g1\ \mathrm{g} corresponds to

E0=1.0×103 kg×8.99×1016 m2/s2=9.0×1013 J.E_0 = 1.0\times10^{-3}\ \mathrm{kg}\times 8.99\times10^{16}\ \mathrm{m^2/s^2} = 9.0\times10^{13}\ \mathrm{J}.

Since 1 kt1\ \mathrm{kt} of TNT is 4.184×1012 J4.184\times10^{12}\ \mathrm{J}, this is 9.0×1013/4.184×1012=21.5 kt9.0\times10^{13}/4.184\times10^{12} = 21.5\ \mathrm{kt} worth. The rest energy of an everyday object is orders of magnitude away from its everyday kinetic energy.

For a single particle the mass was determined by pp=m2c2p\cdot p = m^2c^2. The same construction works for a system of several particles.

Definition 6.1Invariant mass of a system

For a system of particles i=1,,ni = 1,\dots,n, let the total four-momentum be

Pμi=1npiμ=(Ec, P),E=iEi,P=ipi.P^\mu \equiv \sum_{i=1}^n p_i^\mu = \left(\frac{\mathcal E}{c},\ \boldsymbol P\right), \qquad \mathcal E = \sum_i E_i,\quad \boldsymbol P = \sum_i \boldsymbol p_i .

Then

M1c2E2P2c2M \equiv \frac{1}{c^2}\sqrt{\mathcal E^2 - |\boldsymbol P|^2c^2}

is called the invariant mass of the system.

Being a sum of four-vectors, PμP^\mu is a four-vector, and M2c4=PPc2M^2c^4 = P\cdot P\,c^2 is Lorentz invariant by Proposition 2.2. So MM takes the same value in every inertial frame. If an inertial frame with P=0\boldsymbol P = \boldsymbol 0 exists (the centre-of-momentum frame), then there Mc2=EM c^2 = \mathcal E: the invariant mass is the total energy in the centre-of-momentum frame divided by c2c^2.

Theorem 6.2Non-additivity of mass

For a system of nn particles of masses mi0m_i \ge 0 (each satisfying Ei>0E_i > 0 and Ei2=pi2c2+mi2c4E_i^2 = |\boldsymbol p_i|^2c^2 + m_i^2c^4), the invariant mass MM obeys

M  i=1nmi.M \ \ge\ \sum_{i=1}^n m_i .

Equality holds precisely when all the four-momenta are parallel to one another, that is (for particles of positive mass) when all the velocities coincide.

Proof(Theorem 6.2)

We first treat the case n=2n = 2. From Definition 6.1 and the product of Definition 2.1,

M2c4=(p1+p2)(p1+p2)c2=m12c4+m22c4+2(E1E2c2p1p2)M^2c^4 = (p_1+p_2)\cdot(p_1+p_2)\,c^2 = m_1^2c^4 + m_2^2c^4 + 2\left(E_1E_2 - c^2\,\boldsymbol p_1\cdot\boldsymbol p_2\right)

(where pipic2=mi2c4p_i\cdot p_i\,c^2 = m_i^2c^4 is the identity established in the proof of Theorem 4.4). On the other hand,

(m1c2+m2c2)2=m12c4+m22c4+2m1m2c4.(m_1c^2+m_2c^2)^2 = m_1^2c^4+m_2^2c^4+2m_1m_2c^4 .

Hence, to prove Mm1+m2M \ge m_1+m_2 it suffices to prove the inequality

E1E2c2p1p2  m1m2c4,E_1E_2 - c^2\,\boldsymbol p_1\cdot\boldsymbol p_2 \ \ge\ m_1m_2c^4 ,

which we shall call inequality (A). Writing qipic0q_i \equiv |\boldsymbol p_i|c \ge 0, the Cauchy–Schwarz inequality gives p1p2p1p2\boldsymbol p_1\cdot\boldsymbol p_2 \le |\boldsymbol p_1||\boldsymbol p_2|, so

E1E2c2p1p2  E1E2q1q2.E_1E_2 - c^2\,\boldsymbol p_1\cdot\boldsymbol p_2 \ \ge\ E_1E_2 - q_1q_2 .

It therefore suffices to show E1E2m1m2c4+q1q2E_1E_2 \ge m_1m_2c^4 + q_1q_2. Since Ei=mi2c4+qi2E_i = \sqrt{m_i^2c^4+q_i^2} and both sides are non-negative, it is enough to compare their squares.

(E1E2)2(m1m2c4+q1q2)2=(m12c4+q12)(m22c4+q22)(m1m2c4+q1q2)2=m12m22c8+m12c4q22+m22c4q12+q12q22m12m22c82m1m2c4q1q2q12q22=c4(m1q2m2q1)2  0.\begin{aligned} (E_1E_2)^2 - \left(m_1m_2c^4+q_1q_2\right)^2 &= \left(m_1^2c^4+q_1^2\right)\left(m_2^2c^4+q_2^2\right) - \left(m_1m_2c^4+q_1q_2\right)^2\\ &= m_1^2m_2^2c^8 + m_1^2c^4q_2^2 + m_2^2c^4q_1^2 + q_1^2q_2^2\\ &\qquad - m_1^2m_2^2c^8 - 2m_1m_2c^4q_1q_2 - q_1^2q_2^2\\ &= c^4\left(m_1q_2 - m_2q_1\right)^2 \ \ge\ 0 . \end{aligned}

This proves inequality (A). Equality requires both steps to be equalities simultaneously, that is, p1p2\boldsymbol p_1 \parallel \boldsymbol p_2 (same direction) and m1q2=m2q1m_1q_2 = m_2q_1. When mi>0m_i > 0, substituting qi=γimivicq_i = \gamma_i m_i v_i c turns the second condition into m1γ2m2v2=m2γ1m1v1m_1\gamma_2m_2v_2 = m_2\gamma_1m_1v_1, that is, γ2v2=γ1v1\gamma_2v_2 = \gamma_1v_1; since wγ(w)ww \mapsto \gamma(w)w is strictly increasing on [0,c)[0,c), this forces v1=v2v_1 = v_2, and with the directions agreeing we get v1=v2\boldsymbol v_1 = \boldsymbol v_2.

For n3n \ge 3 we induct. The sum p1++pn1p_1+\cdots+p_{n-1} is a four-vector with E>0E > 0 and E2=p2c2+(Mn1c2)2E^2 = |\boldsymbol p|^2c^2 + (M_{n-1}c^2)^2, so the argument above applies verbatim to “one particle of mass Mn1M_{n-1}” together with the nn-th particle, giving MnMn1+mn(in1mi)+mnM_n \ge M_{n-1} + m_n \ge \left(\sum_{i \le n-1} m_i\right) + m_n.

What Theorem 6.2 says is that mass is not additive, and moreover that the discrepancy always goes in the direction of an excess. Where does the excess come from? From the internal kinetic energy of the system and from the energy of the interactions. The next example is the most direct illustration.

Example 6.3Colliding lumps of clay: heat becomes mass

Let us redo the setting of Example 1.1, this time with the correct conservation law. In the frame S, two lumps of clay of mass mm collide head-on at speed uu and merge into a single lump of mass MM. Write γuγ(u)\gamma_u \equiv \gamma(u).

Conservation in S. Momentum: γumu+γum(u)=0\gamma_u m u + \gamma_u m(-u) = 0, so the momentum after the merger is 00 as well and the lump is at rest. Energy: by Theorem 4.2,

γumc2+γumc2=Mc2M=2γum.\gamma_u mc^2 + \gamma_u mc^2 = Mc^2 \quad\Longrightarrow\quad M = 2\gamma_u m .

Since γu>1\gamma_u > 1 we get M>2mM > 2m: the inequality of Theorem 6.2 is realised with strict inequality. The increase is

(M2m)c2=2(γu1)mc2=2K,(M - 2m)c^2 = 2(\gamma_u-1)mc^2 = 2K,

that is, exactly the total kinetic energy that was lost. The energy that turned into heat inside the clay shows up as mass of the merged lump. For u=0.6cu = 0.6c we have γu=1/10.36=1/0.8=1.25\gamma_u = 1/\sqrt{1-0.36} = 1/0.8 = 1.25 and hence M=2.5mM = 2.5\,m: the mass has grown by 25%25\,\%.

Is it conserved in S′ too? Let us redo the check that failed in Example 1.1. The speed of A in S′ was uA=2u/(1+β2)u'_{\mathrm A} = 2u/(1+\beta^2) with β=u/c\beta = u/c. Compute the corresponding γ\gamma:

1uA2c2=(1+β2)24β2(1+β2)2=(1β2)2(1+β2)2γ(uA)=1+β21β2=γu2(1+β2).1 - \frac{u'^2_{\mathrm A}}{c^2} = \frac{(1+\beta^2)^2 - 4\beta^2}{(1+\beta^2)^2} = \frac{(1-\beta^2)^2}{(1+\beta^2)^2} \quad\Longrightarrow\quad \gamma(u'_{\mathrm A}) = \frac{1+\beta^2}{1-\beta^2} = \gamma_u^2\left(1+\beta^2\right).

Hence, remembering that B is at rest, the momentum before the collision is

γ(uA)muA+0=γu2(1+β2)m2u1+β2=2γu2mu.\gamma(u'_{\mathrm A})\,m\,u'_{\mathrm A} + 0 = \gamma_u^2\left(1+\beta^2\right) m \cdot \frac{2u}{1+\beta^2} = 2\gamma_u^2\,m u .

After the collision, the lump of mass M=2γumM = 2\gamma_u m moves with speed uu, so

γ(u)Mu=γu2γumu=2γu2mu.\gamma(u)\,M\,u = \gamma_u\cdot 2\gamma_u m\cdot u = 2\gamma_u^2\,mu .

The two agree. The check that was off by 26%26\,\% in Example 1.1 comes out exactly right the moment we replace the momentum by γmv\gamma m\boldsymbol v and give up the additivity of mass.

Changes of mass are far too small to measure at everyday energy scales. Taking u=10 m/su = 10\ \mathrm{m/s} in Example 6.3 gives γu112(u/c)2=5.6×1016\gamma_u - 1 \approx \frac{1}{2}(u/c)^2 = 5.6\times10^{-16}, so a 1 kg1\ \mathrm{kg} lump of clay gains about 1015 kg10^{-15}\ \mathrm{kg}. On nuclear scales, however, mass differences are routine measurable quantities.

Example 6.4The binding energy of the deuteron

The deuterium nucleus (the deuteron) is a bound state of one proton and one neutron. Their rest energies are

mpc2=938.272 MeV,mnc2=939.565 MeV,mdc2=1875.613 MeV.m_pc^2 = 938.272\ \mathrm{MeV},\qquad m_nc^2 = 939.565\ \mathrm{MeV},\qquad m_dc^2 = 1875.613\ \mathrm{MeV}.

The sum of the masses of a free proton and a free neutron is 938.272+939.565=1877.837 MeV/c2938.272+939.565 = 1877.837\ \mathrm{MeV}/c^2. The mass of the deuteron is smaller than that, the difference being

B=(mp+mnmd)c2=1877.8371875.613=2.224 MeV.B = (m_p + m_n - m_d)c^2 = 1877.837 - 1875.613 = 2.224\ \mathrm{MeV}.

This is the binding energy, equal to the energy needed to pull the deuteron apart into a proton and a neutron. In the language of Theorem 6.2, the invariant mass of a system consisting of a proton and a neutron at rest is mp+mnm_p+m_n (equality holds, since the velocities coincide); upon binding, the surplus energy BB is discarded as a photon, and the invariant mass of what remains drops to md=mp+mnB/c2m_d = m_p+m_n-B/c^2.

As a fraction of mass this is 2.224/1877.837=1.18×1032.224/1877.837 = 1.18\times10^{-3}, a decrease of 0.12%0.12\,\%. Mass spectrometers in nuclear physics resolve this magnitude effortlessly, so here E=mc2E = mc^2 is an everyday working formula. Indeed the inverse reaction d+γp+nd + \gamma \to p + n (photodisintegration of the deuteron) does not occur for photons of energy below 2.224 MeV2.224\ \mathrm{MeV}.

Example 6.5How much lighter does the Sun get each second?

The luminosity of the Sun is L=3.828×1026 WL_\odot = 3.828\times10^{26}\ \mathrm{W}. All the radiated energy comes out of the depletion of rest energy, so by Theorem 4.4 the mass loss per unit time is

dmdt=Lc2=3.828×1026 J/s8.988×1016 m2/s2=4.26×109 kg/s.\frac{dm}{dt} = \frac{L_\odot}{c^2} = \frac{3.828\times10^{26}\ \mathrm{J/s}}{8.988\times10^{16}\ \mathrm{m^2/s^2}} = 4.26\times10^{9}\ \mathrm{kg/s}.

That is 4.264.26 million tonnes per second. Over the roughly 4.64.6 billion years (1.45×1017 s1.45\times10^{17}\ \mathrm{s}) since the Sun was born, the mass lost is 4.26×109×1.45×1017=6.2×1026 kg4.26\times10^9 \times 1.45\times10^{17} = 6.2\times10^{26}\ \mathrm{kg}, a mere 0.031%0.031\,\% of the solar mass 1.989×1030 kg1.989\times10^{30}\ \mathrm{kg}.

The energy source is the fusion of hydrogen. In the net process by which four protons (4×938.272=3753.09 MeV4\times938.272 = 3753.09\ \mathrm{MeV}) turn into a helium-4 nucleus (3727.379 MeV3727.379\ \mathrm{MeV}) together with two positrons and other products, about 26.7 MeV26.7\ \mathrm{MeV} is released. As a fraction this is 26.7/3753.09=0.71%26.7/3753.09 = 0.71\,\%. The number shows that fusion is a technology which taps 0.7%0.7\,\% of the rest energy.

Remark 6.6

Chemical reactions lose mass too. The heat of combustion of carbon, C+O2CO2\mathrm{C} + \mathrm{O_2} \to \mathrm{CO_2}, is 393.5 kJ/mol393.5\ \mathrm{kJ/mol}, that is 393500/(6.022×1023)=6.53×1019 J=4.08 eV393500/(6.022\times10^{23}) = 6.53\times10^{-19}\ \mathrm{J} = 4.08\ \mathrm{eV} per molecule. The rest energy of the reactants is 44 u×931.494 MeV/u=4.10×1010 eV44\ \mathrm{u} \times 931.494\ \mathrm{MeV/u} = 4.10\times10^{10}\ \mathrm{eV}, so the fractional mass decrease is

4.084.10×1010=1.0×1010.\frac{4.08}{4.10\times10^{10}} = 1.0\times10^{-10}.

That is seven orders of magnitude smaller than the 7×1037\times10^{-3} of fusion. Those seven orders of magnitude are what entitle chemistry textbooks to write that mass is conserved in a reaction. The law of conservation of mass was not wrong; it was an approximation good enough to be invisible at chemical precision. That is the accurate way to put it.

Remark 6.7

Theorem 6.2 can also be read backwards: even a system consisting solely of massless particles has non-zero invariant mass. A system of two photons of energy ε\varepsilon travelling in opposite directions has E=2ε\mathcal E = 2\varepsilon and P=0\boldsymbol P = \boldsymbol 0, hence M=2ε/c2M = 2\varepsilon/c^2. Trap light inside a mirrored box and the box gets heavier.

This is no fantasy. Of the proton mass 938 MeV/c2938\ \mathrm{MeV}/c^2, the contribution of the rest masses of its constituent up and down quarks (2mu+md9 MeV/c22m_u + m_d \approx 9\ \mathrm{MeV}/c^2) is only about 1%1\,\%; the rest is the kinetic energy of the quarks and the energy of the gluon field. Most of our body weight is made of confined energy.

Exercise 7.1Easy

Let β=v/c\beta = v/c.

  1. For v=0.1cv = 0.1c and for v=0.5cv = 0.5c, compute numerically K=(γ1)mc2K = (\gamma-1)mc^2 and KN=12mv2K_{\mathrm N} = \frac{1}{2}mv^2 in units of mc2mc^2, and evaluate (KNK)/K\left(K_{\mathrm N}-K\right)/K.
  2. How small must vv be kept for the relative error of the Newtonian expression to stay below 1%1\,\%? Use the first-order approximation of Proposition 5.2.
Solution

1. Compute γ=(1β2)1/2\gamma = (1-\beta^2)^{-1/2}.

For v=0.1cv = 0.1c we have β2=0.01\beta^2 = 0.01 and γ=1/0.99=1.0050378\gamma = 1/\sqrt{0.99} = 1.0050378, so

Kmc2=5.0378×103,KNmc2=12(0.1)2=5.0000×103,\frac{K}{mc^2} = 5.0378\times10^{-3},\qquad \frac{K_{\mathrm N}}{mc^2} = \frac{1}{2}(0.1)^2 = 5.0000\times10^{-3},KNKK=5.00005.03785.0378=7.5×103.\frac{K_{\mathrm N}-K}{K} = \frac{5.0000-5.0378}{5.0378} = -7.5\times10^{-3}.

The Newtonian value underestimates by 0.75%0.75\,\%.

For v=0.5cv = 0.5c we have β2=0.25\beta^2 = 0.25 and γ=1/0.75=1.1547005\gamma = 1/\sqrt{0.75} = 1.1547005, so

Kmc2=0.154701,KNmc2=12(0.5)2=0.125,\frac{K}{mc^2} = 0.154701,\qquad \frac{K_{\mathrm N}}{mc^2} = \frac{1}{2}(0.5)^2 = 0.125,KNKK=0.1250.1547010.154701=0.192.\frac{K_{\mathrm N}-K}{K} = \frac{0.125-0.154701}{0.154701} = -0.192 .

An underestimate of 19.2%19.2\,\%: no longer usable as an approximation.

2. By Proposition 5.2, (KKN)/KN=34β2+O(β4)\left(K-K_{\mathrm N}\right)/K_{\mathrm N} = \frac{3}{4}\beta^2 + O(\beta^4), so it suffices to require 34β2<0.01\frac{3}{4}\beta^2 < 0.01, that is,

β<4300=0.01333=0.1155.\beta < \sqrt{\frac{4}{300}} = \sqrt{0.01333} = 0.1155 .

Thus the Newtonian expression keeps 1%1\,\% accuracy up to about v=0.115c3.5×107 m/sv = 0.115c \approx 3.5\times10^{7}\ \mathrm{m/s}. Since the first cosmic velocity is 7.9×103 m/s7.9\times10^3\ \mathrm{m/s}, relativistic corrections enter astronautics only at the level of 101010^{-10}.

Exercise 7.2Standard

An atom of mass mm at rest absorbs a single photon of energy ε\varepsilon and becomes an (excited) atom of mass MM. Take the four-momentum of the photon to be (ε/c, ε/c, 0, 0)\left(\varepsilon/c,\ \varepsilon/c,\ 0,\ 0\right).

  1. Express the mass MM and the speed vv of the atom after absorption in terms of mm, ε\varepsilon and cc.
  2. For εmc2\varepsilon \ll mc^2, compute MmM - m to second order, including the deviation from ε/c2\varepsilon/c^2, and explain the physical meaning of that deviation.
Solution

1. By Theorem 4.2 all four components of the four-momentum are conserved. Writing EE for the energy and pp for the magnitude of the momentum of the atom after absorption,

E=mc2+ε,pc=εE = mc^2 + \varepsilon,\qquad pc = \varepsilon

(the atom moves in the xx direction). Applying Theorem 4.4 to the atom after absorption,

M2c4=E2p2c2=(mc2+ε)2ε2=m2c4+2mc2ε.M^2c^4 = E^2 - p^2c^2 = \left(mc^2+\varepsilon\right)^2 - \varepsilon^2 = m^2c^4 + 2mc^2\varepsilon .

Hence

M=m1+2εmc2.M = m\sqrt{1 + \frac{2\varepsilon}{mc^2}} .

The speed follows from Corollary 4.5 (1):

v=pc2E=εcmc2+ε.v = \frac{pc^2}{E} = \frac{\varepsilon c}{mc^2+\varepsilon} .

For εmc2\varepsilon \ll mc^2 this is vε/(mc)v \approx \varepsilon/(mc), which agrees with the non-relativistic recoil speed v=p/mv = p/m.

2. Put xε/(mc2)1x \equiv \varepsilon/(mc^2) \ll 1 and use 1+2x=1+x12x2+O(x3)\sqrt{1+2x} = 1 + x - \frac{1}{2}x^2 + O(x^3):

M=m(1+xx22+)=m+εc2ε22mc4+.M = m\left(1 + x - \frac{x^2}{2} + \cdots\right) = m + \frac{\varepsilon}{c^2} - \frac{\varepsilon^2}{2mc^4} + \cdots .

Naively one expects the mass to grow by exactly the energy absorbed, M=m+ε/c2M = m + \varepsilon/c^2; in fact it grows slightly less. The difference ε2/(2mc4)\varepsilon^2/(2mc^4) coincides with 12mv2/c212m(ε/mc)2/c2=ε2/(2mc4)\frac{1}{2}mv^2/c^2 \approx \frac{1}{2}m\left(\varepsilon/mc\right)^2/c^2 = \varepsilon^2/(2mc^4). In other words, part of the absorbed energy goes into the recoil kinetic energy of the atom and does not become internal energy (that is, mass). This difference becomes essential in the discussion of the Mössbauer effect.

Exercise 7.3Standard

Two photons, each of energy ε\varepsilon, travel with an opening angle θ\theta between them (0θπ0 \le \theta \le \pi).

  1. Find the invariant mass MM of this two-photon system.
  2. The neutral pion π0\pi^0 decays into two photons. When both photons are measured in the laboratory frame to have energy 100 MeV100\ \mathrm{MeV} with an opening angle θ\theta, find the value of θ\theta that gives Mc2=134.98 MeVM c^2 = 134.98\ \mathrm{MeV}.
Solution

1. By Corollary 4.5 (2), each photon has momentum of magnitude pi=ε/c|\boldsymbol p_i| = \varepsilon/c. Substitute into Definition 6.1:

E=2ε,P2=p12+p22+2p1p2=ε2c2(1+1+2cosθ).\mathcal E = 2\varepsilon,\qquad |\boldsymbol P|^2 = |\boldsymbol p_1|^2 + |\boldsymbol p_2|^2 + 2\boldsymbol p_1\cdot\boldsymbol p_2 = \frac{\varepsilon^2}{c^2}\left(1 + 1 + 2\cos\theta\right).

Hence

M2c4=E2P2c2=4ε22ε2(1+cosθ)=2ε2(1cosθ).M^2c^4 = \mathcal E^2 - |\boldsymbol P|^2c^2 = 4\varepsilon^2 - 2\varepsilon^2\left(1+\cos\theta\right) = 2\varepsilon^2\left(1-\cos\theta\right).

Using the half-angle identity 1cosθ=2sin2(θ/2)1-\cos\theta = 2\sin^2(\theta/2) gives M2c4=4ε2sin2(θ/2)M^2c^4 = 4\varepsilon^2\sin^2(\theta/2), and since sin(θ/2)0\sin(\theta/2) \ge 0,

M=2εc2sinθ2.M = \frac{2\varepsilon}{c^2}\,\sin\frac{\theta}{2} .

For θ=0\theta = 0 (same direction) we get M=0M = 0, and for θ=π\theta = \pi (opposite directions) M=2ε/c2M = 2\varepsilon/c^2, agreeing with the value in Remark 6.7. This confirms that a system made only of massless particles can have non-zero mass.

2. Substitute ε=100 MeV\varepsilon = 100\ \mathrm{MeV} and Mc2=134.98 MeVMc^2 = 134.98\ \mathrm{MeV} into the formula of part 1:

sinθ2=Mc22ε=134.98200=0.6749θ2=42.5,θ=85.0.\sin\frac{\theta}{2} = \frac{Mc^2}{2\varepsilon} = \frac{134.98}{200} = 0.6749 \quad\Longrightarrow\quad \frac{\theta}{2} = 42.5^\circ,\qquad \theta = 85.0^\circ .

Experiments run this backwards: computing MM for many photon pairs and histogramming the results produces a peak at 134.98 MeV/c2134.98\ \mathrm{MeV}/c^2. This is particle identification by invariant mass, a basic technique of accelerator experiments.

Exercise 7.4Hard

Antiprotons are produced by firing protons at protons at rest (a liquid hydrogen target). The reaction is

p+p  p+p+p+pˉp + p \ \longrightarrow\ p + p + p + \bar p

and the mass of pˉ\bar p equals that of pp, namely mpm_p. Take mpc2=938.272 MeVm_pc^2 = 938.272\ \mathrm{MeV}.

  1. Find the minimum (threshold) kinetic energy KK of the incident proton.
  2. Find the minimum kinetic energy required in each beam when the same reaction is produced by two beams colliding head-on (a collider), and compare with part 1.
Solution

1. The invariant mass MM of the system is Lorentz invariant by the argument of Theorem 6.2 and unchanged across the reaction by Theorem 4.2. The reaction can occur provided the invariant mass of the final state is at least the sum of the masses of the final-state particles, that is, by Theorem 6.2,

Mc2  4mpc2.Mc^2 \ \ge\ 4m_pc^2 .

Equality holds when all four final-state particles move with the same velocity (all at rest in the centre-of-momentum frame), and this is the threshold.

Compute MM from the initial state. Let the incident proton have energy EE and momentum p1\boldsymbol p_1, while the target is at rest (energy mpc2m_pc^2, momentum 0\boldsymbol 0). By Definition 6.1,

M2c4=(E+mpc2)2p12c2=E2p12c2+2Empc2+mp2c4=mp2c4+2Empc2+mp2c4=2mp2c4+2Empc2,\begin{aligned} M^2c^4 &= \left(E + m_pc^2\right)^2 - |\boldsymbol p_1|^2c^2\\ &= E^2 - |\boldsymbol p_1|^2c^2 + 2Em_pc^2 + m_p^2c^4\\ &= m_p^2c^4 + 2Em_pc^2 + m_p^2c^4 = 2m_p^2c^4 + 2Em_pc^2 , \end{aligned}

where we used E2p12c2=mp2c4E^2 - |\boldsymbol p_1|^2c^2 = m_p^2c^4 from Theorem 4.4. Inserting the threshold condition M2c4=16mp2c4M^2c^4 = 16m_p^2c^4,

16mp2c4=2mp2c4+2Empc2E=7mpc2.16m_p^2c^4 = 2m_p^2c^4 + 2Em_pc^2 \quad\Longrightarrow\quad E = 7m_pc^2 .

The kinetic energy is, by Theorem 5.1, K=Empc2=6mpc2=6×938.272=5629 MeV5.63 GeVK = E - m_pc^2 = 6m_pc^2 = 6\times938.272 = 5629\ \mathrm{MeV} \approx 5.63\ \mathrm{GeV}. The reason that 5.63 GeV5.63\ \mathrm{GeV} is needed to make a single antiproton of 0.94 GeV0.94\ \mathrm{GeV} is that the four final-state particles cannot come to rest in the laboratory frame and are forced to carry off surplus kinetic energy. The Bevatron at Berkeley, which discovered the antiproton in 1955, was designed for 6.2 GeV6.2\ \mathrm{GeV} — this estimate with a margin.

2. In a head-on collision the laboratory frame is itself the centre-of-momentum frame. With energy EE in each beam we have P=0\boldsymbol P = \boldsymbol 0 and E=2E\mathcal E = 2E, so Mc2=2EMc^2 = 2E. The threshold condition 2E=4mpc22E = 4m_pc^2 gives E=2mpc2E = 2m_pc^2, that is,

K=Empc2=mpc2=938 MeV.K = E - m_pc^2 = m_pc^2 = 938\ \mathrm{MeV} .

Even taking both beams together this is 1.88 GeV1.88\ \mathrm{GeV}, one third of the 5.63 GeV5.63\ \mathrm{GeV} needed with a fixed target. The gap widens as the energy rises: with a fixed target Mc2EMc^2 \propto \sqrt{E}, whereas with a collider Mc2EMc^2 \propto E. This one-line comparison is the reason nearly all modern accelerators are colliders.

  1. Shigenobu Sunakawa, Sotaiseiriron no Kangaekata (in Japanese), Iwanami Shoten (Butsuri no Kangaekata 4), 1993 — treats relativistic mechanics and the four-vector formalism carefully, organised around thought experiments.
  2. E. F. Taylor and J. A. Wheeler, Spacetime Physics, 2nd ed., W. H. Freeman, 1992 — Chapter 7 (“Momenergy”) covers four-momentum and Chapter 8 (“Collide. Conserve. Create!”) covers collisions and particle production. This is the textbook closest in spirit to the present article.
  3. W. Rindler, Relativity: Special, General, and Cosmological, 2nd ed., Oxford University Press, 2006 — the chapter on relativistic particle mechanics gives a systematic treatment of four-velocity, four-momentum and four-force.
  4. A. Einstein, “Ist die Trägheit eines Körpers von seinem Energieinhalt abhängig?”, Annalen der Physik 18 (1905), 639–641 — the three-page paper that first stated E=mc2E = mc^2. It is the original source for the argument presented in the Appendix.
  5. W. Bertozzi, “Speed and Kinetic Energy of Relativistic Electrons”, American Journal of Physics 32 (1964), 551–555. DOI: 10.1119/1.1970770 — the experiment that demonstrated directly, by time of flight, that electron speeds have an upper bound.
  6. L. B. Okun, “The Concept of Mass”, Physics Today 42, no. 6 (1989), 31–36. DOI: 10.1063/1.881171 — an article sorting out the confusion surrounding the term “relativistic mass”.

In the main text we defined E=γmc2E = \gamma mc^2 in Definition 4.1 and then verified in Theorem 5.1 that it deserves the name energy. Historically the order was reversed: without any four-vector formalism, Einstein derived Δm=E0/c2\Delta m = E_0/c^2 from a single thought experiment about the emission of light (reference [4]). The argument is still worth reading.

Suppose a body at rest in the inertial frame S emits, simultaneously, two pulses of light of energy E0/2E_0/2 each, one in the +x+x direction and one in the x-x direction. The momenta of the two pulses are equal in magnitude and opposite in direction, so they cancel, and the body remains at rest in S after the emission. The decrease of energy as seen in S is, by definition, E0E_0.

Next move to the inertial frame S′, which travels with speed VV in the x-x direction relative to S. Seen from S′, the body moves with speed VV in the +x+x direction. By Corollary 4.5, the four-momentum of a light pulse is (ε/c)(1,±1,0,0)\left(\varepsilon/c\right)(1,\pm1,0,0), with ε=E0/2\varepsilon = E_0/2 and the sign given by the direction of travel. Since S′ is obtained by a boost of velocity V-V, the transformation law of Definition 2.1 with ββ\beta \to -\beta (where β=V/c\beta = V/c) gives

ε±c=γ(V)(εc+β(±εc))ε±=γ(V)ε(1±β).\frac{\varepsilon'_{\pm}}{c} = \gamma(V)\left(\frac{\varepsilon}{c} + \beta\cdot\left(\pm\frac{\varepsilon}{c}\right)\right) \quad\Longrightarrow\quad \varepsilon'_{\pm} = \gamma(V)\,\varepsilon\,(1 \pm \beta).

This is precisely the Doppler effect for light. Adding the two, the terms in β\beta cancel:

ε++ε=γ(V)ε[(1+β)+(1β)]=2γ(V)ε=γ(V)E0.\varepsilon'_{+} + \varepsilon'_{-} = \gamma(V)\,\varepsilon\,\big[(1+\beta) + (1-\beta)\big] = 2\gamma(V)\,\varepsilon = \gamma(V)\,E_0 .

So from S′ the body has lost the energy γ(V)E0\gamma(V)E_0, which exceeds the loss E0E_0 measured in S by (γ(V)1)E0\left(\gamma(V)-1\right)E_0.

Here is the crux of the argument. The velocity of the body is unchanged by the emission: if it stays at rest in S, it keeps the speed VV in S′. Nevertheless the energy measured in S′ decreases by more than in S. Splitting the energy into “internal energy plus kinetic energy”, the surplus (γ(V)1)E0\left(\gamma(V)-1\right)E_0 must be a decrease of kinetic energy. For the kinetic energy to decrease while the speed stays the same, the mm in K=(γ(V)1)mc2K = \left(\gamma(V)-1\right)mc^2 (Theorem 5.1) must decrease. Writing Δm\Delta m for the decrease of mass,

(γ(V)1)E0=(γ(V)1)Δmc2Δm=E0c2.\left(\gamma(V)-1\right)E_0 = \left(\gamma(V)-1\right)\,\Delta m\,c^2 \quad\Longrightarrow\quad \Delta m = \frac{E_0}{c^2}.

Einstein himself expanded γ\gamma to second order in V/cV/c and read the result as a decrease 12(E0/c2)V2\frac{1}{2}\left(E_0/c^2\right)V^2 of kinetic energy, reaching the same conclusion. The closing sentence of the 1905 paper states, in effect, that the mass of a body is a measure of its energy content, and he proposed testing this with the decay of radium salts. It would be some thirty years before the test became feasible.

Note that what has been obtained here is a relation between differences — “a body that loses the energy E0E_0 loses the mass E0/c2E_0/c^2” — and not the assertion that the whole rest energy equals mc2mc^2. To reach the statement about the whole, one must start from the four-momentum, as we did in the main text. The division of labour between the historical argument and the systematic one is visible right here.

The four-momentum assembled in this article will be the foundation of the chapters that follow. In An invitation to general relativity: the equivalence principle we start from the fact that gravitational and inertial mass are equal (inertial mass and gravitational mass(Definition 2.1)[一般相対性理論への招待]) and move towards the statement that everything possessing energy feels gravity. For if E=mc2E = mc^2 is correct, then even light, which has no mass, ought to be bent by gravity — and that expectation arises quite naturally.

Report an error in this article ・Operated by: Mugen Giken LLCPricingTermsLegal notice

© 2026 夢現技研合同会社 ・Feeding the text to an LLM is welcome. Code samples are MIT licensed.