# Relativistic Mechanics: What the Four-Momentum Says About E = mc²

> Newtonian momentum conservation fails between inertial frames. Building proper time, four-velocity and four-momentum repairs it, yielding E = γmc², E² = (pc)² + (mc²)², and non-additive mass.
> https://rikai.mugen-giken.com/en/physics/relativity/relativistic-mechanics

## 0. Key points

- The conservation law for the Newtonian momentum $m\boldsymbol v$ may hold in one inertial frame and fail in another. The reason is that velocities compose by the Lorentz transformation, not by the Galilean one. If we want to keep the conservation law, we have no choice but to rebuild the definition of momentum.
- The guiding principle for the rebuilding is: use only quantities that transform under Lorentz transformations by the same rule as the spacetime coordinates $x^\mu$ — that is, four-vectors. The leading roles go to the four-velocity $u^\mu$, obtained by differentiating with respect to proper time $\tau$, and to the four-momentum $p^\mu = m u^\mu$.
- The spatial part of $p^\mu$ is $\boldsymbol p = \gamma m \boldsymbol v$ and its time component is $p^0 = \gamma m c$. Multiplying the latter by $c$ gives $E = \gamma m c^2$, the relativistic energy, which at low speeds reduces to $mc^2 + \frac{1}{2}mv^2$.
- If we demand conservation of momentum *in every* inertial frame, conservation of energy follows automatically. The two are merely different components of a single four-vector conservation law.
- A body at rest still carries the energy $E = mc^2$. Unlike in Newtonian mechanics, this "constant" is observable: when a reaction changes which particles are present, $\sum m$ changes, and the difference is exactly the energy that flows in or out.
- Mass is not additive. The invariant mass of a system is fixed by $M^2c^4 = \left(\sum E\right)^2 - \left|\sum \boldsymbol p\right|^2 c^2$, and it is always greater than or equal to the sum of the masses of the constituents. The binding energy of deuterium, and the fact that the Sun shines, are consequences of this one sentence.

## 1. Motivation: where Newtonian mechanics breaks down

In the [previous article](/en/physics/relativity/lorentz-transformations) we saw that the coordinate transformation between inertial frames is the <Ref to="physics/relativity/lorentz-transformations#thm-lorentz" text="Lorentz transformation" /> rather than the Galilean one. Time dilation and length contraction both follow from it. But if the transformation law has changed, then **the laws of mechanics that were written down for the old law must be re-examined as well**.

The backbone of Newtonian mechanics is the momentum $\boldsymbol p_{\text{N}} = m\boldsymbol v$ together with its conservation law. This is a law confirmed by experiment to an extraordinary degree, and as we saw in <Ref to="physics/mechanics/newtonian-mechanics#thm-momentum-conservation" text="conservation of momentum" /> in [Foundations of Newtonian mechanics](/en/physics/mechanics/newtonian-mechanics), it is central enough to be equivalent to the law of action and reaction. Yet this law does not keep its form under Lorentz transformations. The following computation shows it.

<Example id="ex-newtonian-failure" title="Newtonian momentum conservation is frame-dependent">
Consider two lumps of clay A and B of equal mass. Viewed from the inertial frame S, A moves with velocity $+u$ and B with velocity $-u$; they collide head-on and stick together into a single lump. The Newtonian momentum before the collision is $mu + m(-u) = 0$, so if we accept the conservation law the merged lump is at rest in S.

Now change to the inertial frame S′, which moves with velocity $-u$ along the $x$ axis relative to S. By the velocity-addition rule (<Ref to="physics/relativity/lorentz-transformations#cor-velocity-addition" /> of the [previous article](/en/physics/relativity/lorentz-transformations)), a body with velocity $w$ in S has velocity

$$
w' = \frac{w + u}{1 + wu/c^2}
$$

in S′. Apply this to the three bodies. Writing $\beta \equiv u/c$,

- A ($w = u$): $\displaystyle w'_{\mathrm A} = \frac{2u}{1+\beta^2}$
- B ($w = -u$): $\displaystyle w'_{\mathrm B} = \frac{-u+u}{1-\beta^2} = 0$
- the merged lump ($w = 0$): $w' = u$

Compute the Newtonian momentum in S′, assuming that masses add (so that the merged lump has mass $2m$).

$$
\begin{aligned}
\text{before} &= m\cdot\frac{2u}{1+\beta^2} + m\cdot 0 = \frac{2mu}{1+\beta^2},\\
\text{after} &= 2m\cdot u = 2mu.
\end{aligned}
$$

Their ratio is $1/(1+\beta^2)$, which is never $1$ as long as $u \neq 0$. For instance with $u = 0.6c$ we have $\beta^2 = 0.36$, so the value before the collision is $2mu/1.36 = 1.47\,mu$ while after it is $2mu$ — a discrepancy of $26\,\%$.

In other words, Newtonian momentum is conserved in S but not in S′.
</Example>

This is serious. The principle of relativity (<Ref to="physics/relativity/principles-of-special-relativity#ax-relativity-principle" text="the principle of special relativity" /> in [Principles of special relativity](/en/physics/relativity/principles-of-special-relativity)) demands that the laws of physics take the same form in every inertial frame, so a conservation law that holds only in one frame forfeits its status as a law of physics. There are only two ways out.

1. Abandon conservation of momentum.
2. Rebuild the definition of momentum (or the assumption that masses add).

Experiment forbids the first. That a conserved quantity exists in collisions has been confirmed on every scale from elementary particles to astronomical bodies. Moreover, by [Noether's theorem](/physics/mechanics/noethers-theorem), the existence of the conserved quantity is a restatement of the homogeneity of space, so discarding it would amount to discarding a symmetry. Hence only the second road is open.

The same conclusion is visible from another direction. In Newtonian mechanics the kinetic energy $K = \frac{1}{2}mv^2$ has no upper bound, so quadrupling $K$ should double $v$. Yet however much energy we pour into accelerating an electron, the speed measured by time of flight merely sticks to $c$ and never exceeds it. Bertozzi's 1962 experiment (reference [5]) showed directly that raising the kinetic energy of electrons from $0.5\ \mathrm{MeV}$ to $15\ \mathrm{MeV}$ — a factor of $30$ — moves the time-of-flight value of $v^2/c^2$ only from $0.74$ to $0.999$, never past $1$. The very relation between $K$ and $v$ is different.

In this article we follow the second road to its end. Its destination is $E = mc^2$.

## 2. Preliminaries: Lorentz transformations and four-vectors

### 2.1. Notation

Let $c$ denote the speed of light and write $\gamma(w) \equiv \left(1 - w^2/c^2\right)^{-1/2}$. We use $V$ for the relative velocity between inertial frames and $\boldsymbol v$ (with magnitude $v$) for the velocity of a particle, abbreviating $\gamma \equiv \gamma(v)$ for particles. We set $\beta \equiv V/c$.

When S′ moves with velocity $V$ in the positive $x$ direction relative to S, the Lorentz transformation (boost) reads

$$
\begin{aligned}
ct' &= \gamma(V)\left(ct - \beta x\right), \qquad & x' &= \gamma(V)\left(x - \beta\, ct\right),\\
y' &= y, \qquad & z' &= z
\end{aligned}
$$

as before. Note that the time coordinate has been rescaled to $ct$, giving it the dimension of a length. All four coordinates then carry the same dimension, and we may write them together as $x^\mu = (x^0, x^1, x^2, x^3) = (ct, x, y, z)$ with $\mu = 0,1,2,3$.

### 2.2. Four-vectors

The Lorentz transformation is **linear** in $x^\mu$, and it mixes $x^0$ with $x^1$. From this the notion of "a quadruple that mixes in the same way" arises naturally.

<Definition id="def-four-vector" title="Four-vectors and the Minkowski product">
Suppose that in each inertial frame a quadruple of numbers $A^\mu = (A^0, A^1, A^2, A^3)$ is given, and that under the boost above it transforms as

$$
\begin{aligned}
A'^0 &= \gamma(V)\left(A^0 - \beta A^1\right), \qquad & A'^1 &= \gamma(V)\left(A^1 - \beta A^0\right),\\
A'^2 &= A^2, \qquad & A'^3 &= A^3.
\end{aligned}
$$

Then $A^\mu$ is called a **four-vector**. For two four-vectors $A^\mu$ and $B^\mu$ we call

$$
A\cdot B \equiv A^0B^0 - A^1B^1 - A^2B^2 - A^3B^3
$$

the **Minkowski product**, and $A\cdot A$ the **squared norm** of $A^\mu$.
</Definition>

<Aside type="note">
There are two conventions for the signature of the metric, $(+,-,-,-)$ and $(-,+,+,+)$, and both describe the same physics. In this article we adopt $(+,-,-,-)$, the one common in particle physics, taking the time component positive. With this convention the product $p\cdot p$ appearing below comes out as the positive number $m^2c^2$, which saves us some sign-chasing.
</Aside>

It follows from the definition that the **difference** of spacetime coordinates $\Delta x^\mu = (c\Delta t, \Delta x, \Delta y, \Delta z)$ is a four-vector (the Lorentz transformation is linear, so differences obey the same formulas). Its squared norm

$$
\Delta x \cdot \Delta x = c^2\Delta t^2 - \Delta x^2 - \Delta y^2 - \Delta z^2
$$

is the invariant interval of the previous article (<Ref to="physics/relativity/lorentz-transformations#thm-invariant-interval" text="invariance of the spacetime interval" />). In fact it is not only the norm of a coordinate difference that is invariant.

<Proposition id="prop-invariance" title="Lorentz invariance of the Minkowski product">
If $A^\mu$ and $B^\mu$ are four-vectors, then $A'\cdot B' = A\cdot B$ for every boost.
</Proposition>

<Proof of="prop-invariance">
Since $A^2$ and $A^3$ are unchanged by the transformation, $A'^2B'^2 + A'^3B'^3 = A^2B^2 + A^3B^3$. It therefore suffices to examine the $0$ and $1$ components. Substituting the transformation law of <Ref to="def-four-vector" />,

$$
\begin{aligned}
A'^0B'^0 - A'^1B'^1
&= \gamma^2\Big[\left(A^0-\beta A^1\right)\left(B^0-\beta B^1\right) - \left(A^1-\beta A^0\right)\left(B^1-\beta B^0\right)\Big]\\
&= \gamma^2\Big[A^0B^0 - \beta A^0B^1 - \beta A^1B^0 + \beta^2A^1B^1\\
&\qquad\quad - A^1B^1 + \beta A^1B^0 + \beta A^0B^1 - \beta^2A^0B^0\Big]\\
&= \gamma^2\Big[(1-\beta^2)A^0B^0 - (1-\beta^2)A^1B^1\Big].
\end{aligned}
$$

The essential point is that the cross terms cancel in pairs: $-\beta A^0B^1$ against $+\beta A^0B^1$, and $-\beta A^1B^0$ against $+\beta A^1B^0$. Finally, $\gamma^2 = (1-\beta^2)^{-1}$ gives $\gamma^2(1-\beta^2) = 1$, so

$$
A'^0B'^0 - A'^1B'^1 = A^0B^0 - A^1B^1
$$

as claimed. Boosts along directions other than the $x$ axis, and spatial rotations, reduce to the same computation after a suitable choice of axes.
</Proof>

<Ref to="prop-invariance" /> is the tool used throughout this article. **The product of two four-vectors is the same number no matter which inertial frame computes it.** That is why a quantity built from such a product carries meaning as a quantity intrinsic to the object. We shall use this single fact again and again.

## 3. Proper time and four-velocity

### 3.1. Making velocity into a four-vector

To construct a momentum we must first construct a velocity. The naive attempt $dx^\mu/dt$ is not a four-vector, and the reason is plain: the numerator $dx^\mu$ is a four-vector, but the denominator $dt$ is a frame-dependent quantity (one component, $x^0/c$). Dividing a four-vector by something that is not invariant destroys the transformation law.

What we need is **a time on whose value everyone agrees**. That is proper time.

<Definition id="def-proper-time" title="Proper time">
Let $dx^\mu = (c\,dt, d\boldsymbol x)$ be an infinitesimal displacement along the world line of a particle, with $dx\cdot dx > 0$ (timelike). The quantity $d\tau > 0$ defined by

$$
c^2\,d\tau^2 \equiv dx\cdot dx = c^2dt^2 - |d\boldsymbol x|^2
$$

is called the infinitesimal change of **proper time**, and its integral $\tau = \int d\tau$ along the world line is the proper time of the particle.
</Definition>

Since $d\tau$ is built out of a product of four-vectors, it is Lorentz invariant by <Ref to="prop-invariance" />. Proper time thus has the same value in every inertial frame, and physically it is nothing other than **the time read by a clock attached to the particle**. It is the quantity <Ref to="physics/relativity/lorentz-transformations#def-proper-time" text="proper time" /> of the previous article, rewritten in terms of an infinitesimal displacement along the world line.

<Proposition id="prop-proper-time" title="Proper time versus coordinate time">
If a particle has velocity $\boldsymbol v$ (of magnitude $v < c$) in the inertial frame S, then the coordinate time $t$ of S and the proper time $\tau$ satisfy

$$
\frac{dt}{d\tau} = \gamma(v) = \frac{1}{\sqrt{1 - v^2/c^2}}.
$$
</Proposition>

<Proof of="prop-proper-time">
Factor $dt^2$ out of the right-hand side of <Ref to="def-proper-time" />. By the definition of velocity, $\boldsymbol v = d\boldsymbol x/dt$, so $|d\boldsymbol x| = v\,dt$ and

$$
c^2 d\tau^2 = c^2dt^2 - |d\boldsymbol x|^2 = c^2dt^2 - v^2dt^2 = c^2dt^2\left(1 - \frac{v^2}{c^2}\right).
$$

Dividing both sides by $c^2dt^2$ gives $\left(d\tau/dt\right)^2 = 1 - v^2/c^2$. Since $v < c$, the right-hand side is positive, and as we have taken $d\tau > 0$ and $dt > 0$ the sign of the square root is fixed to be positive:

$$
\frac{d\tau}{dt} = \sqrt{1 - \frac{v^2}{c^2}} = \frac{1}{\gamma(v)},
$$

that is, $dt/d\tau = \gamma(v)$.
</Proof>

Since $\gamma(v) \ge 1$, we get $dt \ge d\tau$: the conclusion of the previous article that a moving clock runs slow (<Ref to="physics/relativity/lorentz-transformations#thm-time-dilation" text="time dilation" />) is reproduced here by the very same formula.

### 3.2. Four-velocity

<Definition id="def-four-velocity" title="Four-velocity">
For the world line $x^\mu(\tau)$ of a particle of non-zero mass, the quantity

$$
u^\mu \equiv \frac{dx^\mu}{d\tau}
$$

is called the **four-velocity**.
</Definition>

<Theorem id="thm-four-velocity" title="Properties of the four-velocity">
For a particle of non-zero mass the following hold.

1. $u^\mu$ is a four-vector.
2. If the particle has velocity $\boldsymbol v$ in the inertial frame S, its components are $u^\mu = \gamma(v)\,(c,\ \boldsymbol v)$.
3. Its squared norm is $u\cdot u = c^2$, independently of the velocity.
</Theorem>

<Proof of="thm-four-velocity">
**(1)** The numerator $dx^\mu$ is a four-vector (as noted just after <Ref to="def-four-vector" />, coordinate differences obey the boost formulas directly, by linearity of the Lorentz transformation). The denominator $d\tau$ is a Lorentz-invariant scalar by <Ref to="prop-invariance" />. Dividing a four-vector by a scalar multiplies every component by the same constant, so the transformation law is unaffected. Hence $u^\mu$ is a four-vector.

**(2)** Rewrite in terms of coordinate time by the chain rule. Since $dt/d\tau = \gamma(v)$ by <Ref to="prop-proper-time" />,

$$
u^\mu = \frac{dx^\mu}{d\tau} = \frac{dt}{d\tau}\cdot\frac{dx^\mu}{dt} = \gamma(v)\left(\frac{d(ct)}{dt},\ \frac{d\boldsymbol x}{dt}\right) = \gamma(v)\,(c,\ \boldsymbol v).
$$

**(3)** Insert the components from (2) into the product of <Ref to="def-four-vector" />:

$$
u\cdot u = \gamma(v)^2\left(c^2 - |\boldsymbol v|^2\right) = \frac{c^2 - v^2}{1 - v^2/c^2} = \frac{c^2\left(1 - v^2/c^2\right)}{1 - v^2/c^2} = c^2.
$$

The key step is factoring $c^2 - v^2 = c^2(1 - v^2/c^2)$: the velocity cancels and the constant $c^2$ survives.

Alternatively, (3) follows without any computation: in the rest frame of the particle ($\boldsymbol v = \boldsymbol 0$) we have $u^\mu = (c, \boldsymbol 0)$ and hence $u\cdot u = c^2$, and the value is the same in every frame by <Ref to="prop-invariance" />.
</Proof>

<Remark id="rem-four-velocity-meaning">
The relation $u\cdot u = c^2$ may be read as saying that every object moves through spacetime at the constant rate $c$. A body at rest advances at rate $c$ purely in the time direction, while a fast-moving body diverts part of that $c$ into the spatial directions. Time dilation corresponds to the resulting decrease of the time component.
</Remark>

## 4. Four-momentum and energy

### 4.1. Definition

In Newtonian mechanics momentum was "mass times velocity". In relativity we put the four-velocity in place of the velocity. For the mass we use the value measured in the rest frame of the particle (the **rest mass**, from now on simply the mass). It is a constant attached to each particle, a number independent of the frame.

<Definition id="def-four-momentum" title="Four-momentum">
The **four-momentum** of a particle of mass $m > 0$ is defined by

$$
p^\mu \equiv m\,u^\mu = m\gamma(v)\,(c,\ \boldsymbol v).
$$

Its spatial part is called the **relativistic momentum** $\boldsymbol p \equiv \gamma(v)\,m\boldsymbol v$, and its time component multiplied by $c$ is called the **energy** $E \equiv c\,p^0 = \gamma(v)\,mc^2$, so that $p^\mu = (E/c,\ \boldsymbol p)$.
</Definition>

Since $m$ is an invariant scalar and $u^\mu$ is a four-vector (<Ref to="thm-four-velocity" /> (1)), $p^\mu$ is a four-vector as well. For $v \ll c$ we have $\gamma \approx 1$ and hence $\boldsymbol p \approx m\boldsymbol v$, recovering the Newtonian momentum. So this is the Newtonian momentum, modified by the least amount needed to make it a four-vector.

We have not yet explained why the name $E$ is deserved; that will become clear in <Ref to="thm-kinetic-energy" />. First let us see why this definition rescues the conservation law.

### 4.2. Conservation independent of the frame

<Theorem id="thm-covariant-conservation" title="Frame independence of four-momentum conservation">
In a process such as a collision, a decay or a creation event, write $\sum_{\text{in}} p^\mu$ for the sum of the four-momenta of the incoming particles and $\sum_{\text{out}} p^\mu$ for the sum for the outgoing ones.

1. If $\sum_{\text{in}} p^\mu = \sum_{\text{out}} p^\mu$ (all four components) holds in one inertial frame, then it holds in every inertial frame.
2. Moreover, if conservation of the spatial part $\sum_{\text{in}} \boldsymbol p = \sum_{\text{out}} \boldsymbol p$ holds in every inertial frame, then conservation of the time component $\sum_{\text{in}} E = \sum_{\text{out}} E$ holds automatically.
</Theorem>

<Proof of="thm-covariant-conservation">
Consider the difference $\Delta^\mu \equiv \sum_{\text{in}} p^\mu - \sum_{\text{out}} p^\mu$. The transformation law of <Ref to="def-four-vector" /> is linear, so sums and differences of four-vectors are again four-vectors. Hence $\Delta^\mu$ is a four-vector.

**(1)** The hypothesis is that $\Delta^\mu = 0$ (all four components vanish) in some frame S. For an arbitrary boost,

$$
\Delta'^0 = \gamma(V)\left(\Delta^0 - \beta\Delta^1\right) = 0,\qquad
\Delta'^1 = \gamma(V)\left(\Delta^1 - \beta\Delta^0\right) = 0,
$$

and $\Delta'^2 = \Delta^2 = 0$, $\Delta'^3 = \Delta^3 = 0$. What matters here is that the transformation is a **homogeneous** linear map, with no constant term. Hence $\Delta'^\mu = 0$ in every frame, which is the conservation law.

**(2)** The hypothesis is that the spatial components vanish in every frame, that is, $\Delta^1 = \Delta^2 = \Delta^3 = 0$ and $\Delta'^1 = \Delta'^2 = \Delta'^3 = 0$. Taking a boost with velocity $V \neq 0$ along $x$, the transformation law gives

$$
0 = \Delta'^1 = \gamma(V)\left(\Delta^1 - \beta\Delta^0\right) = \gamma(V)\left(0 - \beta\Delta^0\right) = -\gamma(V)\,\beta\,\Delta^0 .
$$

Since $\gamma(V) \ge 1 > 0$ and $\beta = V/c \neq 0$, we must have $\Delta^0 = 0$. As $E = cp^0$, this is conservation of energy, $\sum_{\text{in}} E = \sum_{\text{out}} E$.
</Proof>

Part (1) of <Ref to="thm-covariant-conservation" /> repairs exactly what broke in <Ref to="ex-newtonian-failure" />. The Newtonian momentum $m\boldsymbol v$ is not part of a four-vector, so conservation failed upon changing frames. The relativistic $\boldsymbol p = \gamma m\boldsymbol v$ is the spatial part of the four-vector $p^\mu$, so as long as it is conserved together with the time component, it is conserved in every frame.

Part (2) is one of the most beautiful consequences of the theory. **In relativity one cannot postulate momentum conservation and energy conservation as separate laws.** Demanding either one in all inertial frames brings the other with it. Two laws that were independent in Newtonian mechanics are here unified into a single four-vector conservation law.

<Remark id="rem-noether">
In the language of [Noether's theorem](/physics/mechanics/noethers-theorem), momentum conservation was the consequence of invariance under spatial translations (<Ref to="physics/mechanics/noethers-theorem#ex-translation" text="conservation of total momentum" />) and energy conservation the consequence of invariance under time translations (<Ref to="physics/mechanics/noethers-theorem#cor-energy" text="homogeneity of time and conservation of energy" />). In relativity a boost mixes time with space, so the two symmetries mix as well. The merging of the conservation laws into one is a reflection of this.
</Remark>

### 4.3. The energy–momentum relation

<Theorem id="thm-energy-momentum-relation" title="Energy–momentum relation">
For a free particle of mass $m$, the energy $E$ and momentum $\boldsymbol p$ satisfy

$$
E^2 = |\boldsymbol p|^2c^2 + m^2c^4
$$

in every inertial frame. In particular, at rest ($\boldsymbol p = \boldsymbol 0$) we have $E = mc^2$.
</Theorem>

<Proof of="thm-energy-momentum-relation">
By <Ref to="def-four-momentum" /> we have $p^\mu = m u^\mu$, so the product of <Ref to="def-four-vector" /> together with <Ref to="thm-four-velocity" /> (3) gives

$$
p\cdot p = m^2\,(u\cdot u) = m^2c^2 .
$$

On the other hand, computing the product directly from the components $p^\mu = (E/c, \boldsymbol p)$,

$$
p\cdot p = \left(\frac{E}{c}\right)^2 - |\boldsymbol p|^2 .
$$

Equating the two gives $\left(E/c\right)^2 - |\boldsymbol p|^2 = m^2c^2$, and multiplying by $c^2$ yields $E^2 - |\boldsymbol p|^2c^2 = m^2c^4$.

The left-hand side is Lorentz invariant by <Ref to="prop-invariance" />, so the identity holds in the same form whichever frame measures the components. Setting $\boldsymbol p = \boldsymbol 0$ gives $E^2 = m^2c^4$, and $E > 0$ gives $E = mc^2$.
</Proof>

In practice it is convenient to remember this relation as a right triangle with $E$ as the hypotenuse.

<Figure caption="The energy–momentum relation E² = (pc)² + (mc²)² drawn as a right triangle. The vertical leg is the rest energy mc², the horizontal leg is pc, and the hypotenuse is the total energy E. The angle θ satisfies sin θ = pc/E = v/c. For a particle at rest the horizontal leg vanishes and E = mc²; for a massless particle the vertical leg vanishes and E = pc.">
<svg viewBox="0 0 760 330" width="100%" role="img" aria-label="Right triangle representing the relation between energy and momentum">
  <path d="M 120 250 L 120 100" fill="none" stroke="currentColor" stroke-width="2.4" />
  <path d="M 120 250 L 460 250" fill="none" stroke="currentColor" stroke-width="2.4" />
  <path d="M 120 100 L 460 250" fill="none" stroke="var(--sl-color-accent)" stroke-width="3" />
  <path d="M 120 230 L 140 230 L 140 250" fill="none" stroke="currentColor" stroke-width="1.4" />
  <path d="M 120 140 A 40 40 0 0 0 156.6 116.1" fill="none" stroke="currentColor" stroke-width="1.4" />
  <g fill="currentColor" font-size="19">
    <text x="290" y="280" text-anchor="middle">p c</text>
    <text x="106" y="182" text-anchor="end">m c²</text>
    <text x="154" y="148" font-size="17">θ</text>
  </g>
  <text x="298" y="156" fill="var(--sl-color-accent)" font-size="19">E</text>
  <g fill="currentColor" font-size="15">
    <text x="500" y="120">p → 0 gives E → mc²</text>
    <text x="500" y="150">m → 0 gives E → pc</text>
    <text x="500" y="180">sin θ = pc / E = v / c</text>
  </g>
</svg>
</Figure>

<Corollary id="cor-velocity-and-massless" title="Recovering the velocity, and massless particles">
1. For a particle of mass $m > 0$ we have $\displaystyle \boldsymbol v = \frac{\boldsymbol p\,c^2}{E}$.
2. Taking the limit $m \to 0$ while keeping $\boldsymbol p \neq \boldsymbol 0$ fixed, we get $E = |\boldsymbol p|c$ and $v = c$. A particle in this limit (a photon, for instance) has zero mass yet non-zero energy and momentum.
</Corollary>

<Proof of="cor-velocity-and-massless">
**(1)** Simply divide $\boldsymbol p = \gamma m\boldsymbol v$ by $E = \gamma mc^2$ from <Ref to="def-four-momentum" />:

$$
\frac{\boldsymbol p\,c^2}{E} = \frac{\gamma m \boldsymbol v \, c^2}{\gamma m c^2} = \boldsymbol v .
$$

The factors $\gamma$, $m$ and $c^2$ simply cancel.

**(2)** Setting $m = 0$ in <Ref to="thm-energy-momentum-relation" /> gives $E^2 = |\boldsymbol p|^2c^2$, and $E > 0$ gives $E = |\boldsymbol p|c$. Substituting this into (1),

$$
v = \frac{|\boldsymbol p|c^2}{E} = \frac{|\boldsymbol p|c^2}{|\boldsymbol p|c} = c .
$$

Conversely, for a particle with $v = c$ the factor $\gamma(v)$ diverges, so $E = \gamma mc^2$ can be finite only if $m = 0$. Massless particles can travel only at the speed of light, and particles travelling at the speed of light cannot have mass.
</Proof>

<Remark id="rem-relativistic-mass">
Older textbooks speak of a "mass that grows with speed", $m_{\text{rel}} = \gamma m$ (the relativistic mass). It looks convenient, since one can then write $\boldsymbol p = m_{\text{rel}}\boldsymbol v$ and $E = m_{\text{rel}}c^2$; but $\boldsymbol F = m_{\text{rel}}\boldsymbol a$ fails (force and acceleration need not be parallel), so the notion causes more confusion than it removes. The modern standard is that "mass" always means the invariant rest mass $m$, with all velocity dependence pushed into $\gamma$. We follow that convention here. The history is discussed in detail in Okun's article (reference [6]).
</Remark>

## 5. Relativistic kinetic energy and rest energy

### 5.1. Kinetic energy from the work done

In <Ref to="def-four-momentum" /> we called $E = \gamma mc^2$ the "energy", but so far that is only a name. The operational definition that fixes what energy is reads "the work done by an external force on a body initially at rest", so let us compute it. As the equation of motion we adopt Newton's second law in the form

$$
\boldsymbol F = \frac{d\boldsymbol p}{dt},\qquad \boldsymbol p = \gamma(v)\,m\boldsymbol v .
$$

The point is to keep the form $d\boldsymbol p/dt$ rather than $\boldsymbol F = m\boldsymbol a$; in this form the law is directly compatible with conservation of four-momentum.

<Theorem id="thm-kinetic-energy" title="Relativistic kinetic energy">
Suppose a particle of mass $m$ starts from rest and, under an external force $\boldsymbol F = d\boldsymbol p/dt$, reaches the speed $v$ (with $v < c$). Then the work $K$ done by the force equals

$$
K = (\gamma(v) - 1)\,mc^2 = \frac{mc^2}{\sqrt{1-v^2/c^2}} - mc^2 ,
$$

and this quantity is called the **relativistic kinetic energy**. Equivalently, $E = \gamma mc^2 = mc^2 + K$.
</Theorem>

<Proof of="thm-kinetic-energy">
To keep the notation simple we treat one-dimensional motion, with force and motion along the $x$ axis; the general case is described at the end. From the definition of work and the equation of motion,

$$
K = \int_0^x F\,dx' = \int_0^x \frac{dp}{dt}\,dx' .
$$

Change the integration variable from $x'$ to $p$. Since $dx' = v\,dt$, we have $\dfrac{dp}{dt}\,dx' = \dfrac{dp}{dt}\,v\,dt = v\,dp$, and the particle starts from rest ($p = 0$) and reaches momentum $p = \gamma m v$, so

$$
K = \int_0^{\gamma m v} v\,dp .
$$

Integrate by parts. Since $\int v\,dp = \left[vp\right] - \int p\,dv$,

$$
K = \Big[v\cdot \gamma(v) m v\Big]_{v=0}^{v} - \int_0^{v} \gamma(w)\,m w\,dw
= \gamma(v)\,mv^2 - m\int_0^v \frac{w\,dw}{\sqrt{1-w^2/c^2}} .
$$

Now evaluate the remaining integral. Substituting $s = 1 - w^2/c^2$ gives $ds = -2w\,dw/c^2$, that is, $w\,dw = -\tfrac{c^2}{2}ds$, and the range $w: 0 \to v$ becomes $s: 1 \to 1 - v^2/c^2$:

$$
\int_0^v \frac{w\,dw}{\sqrt{1-w^2/c^2}}
= \int_1^{1-v^2/c^2} \frac{-\tfrac{c^2}{2}\,ds}{\sqrt{s}}
= -\frac{c^2}{2}\Big[2\sqrt{s}\Big]_1^{1-v^2/c^2}
= c^2\left(1 - \sqrt{1-\frac{v^2}{c^2}}\right) = c^2\left(1 - \frac{1}{\gamma(v)}\right).
$$

Putting this back,

$$
K = \gamma m v^2 - mc^2\left(1 - \frac{1}{\gamma}\right) = mc^2\left(\gamma\,\frac{v^2}{c^2} - 1 + \frac{1}{\gamma}\right).
$$

Simplify the bracket. Since $\dfrac{1}{\gamma} = \sqrt{1-\dfrac{v^2}{c^2}} = \gamma\left(1-\dfrac{v^2}{c^2}\right)$ (the last step uses $\gamma^{-2} = 1 - v^2/c^2$),

$$
\gamma\,\frac{v^2}{c^2} + \frac{1}{\gamma} = \gamma\,\frac{v^2}{c^2} + \gamma\left(1 - \frac{v^2}{c^2}\right) = \gamma .
$$

Hence $K = mc^2(\gamma - 1)$.

The conclusion is the same for general three-dimensional motion. Differentiating both sides of <Ref to="thm-energy-momentum-relation" /> gives $2E\,dE = 2c^2\,\boldsymbol p\cdot d\boldsymbol p$ (as $m$ is constant), so using <Ref to="cor-velocity-and-massless" /> (1),

$$
dE = \frac{c^2\boldsymbol p}{E}\cdot d\boldsymbol p = \boldsymbol v\cdot d\boldsymbol p = \boldsymbol v\cdot\boldsymbol F\,dt ,
$$

and the right-hand side is precisely the work done by the external force during $dt$. So the total work equals the increment of $E$, and the increment of $E = \gamma mc^2$ from $\gamma = 1$ is $(\gamma-1)mc^2$.
</Proof>

<Figure caption="Comparison of kinetic energies. The solid curve is the relativistic K = (γ - 1)mc², the dashed one is the Newtonian K = mv²/2. Up to about β = v/c ≈ 0.2 the two coincide, but as the speed approaches c the relativistic curve diverges: no amount of energy brings v up to c.">
<svg viewBox="0 0 640 380" width="100%" role="img" aria-label="Graph comparing relativistic kinetic energy with its Newtonian approximation">
  <path d="M 70 35 L 70 330 L 620 330" fill="none" stroke="currentColor" stroke-width="1.5" />
  <path d="M 70 330 L 64 330 M 70 233 L 64 233 M 70 137 L 64 137 M 70 40 L 64 40 M 70 330 L 70 336 M 335 330 L 335 336 M 600 330 L 600 336" fill="none" stroke="currentColor" stroke-width="1.5" />
  <g fill="currentColor" font-size="15" text-anchor="end">
    <text x="56" y="335">0</text>
    <text x="56" y="238">1</text>
    <text x="56" y="142">2</text>
    <text x="56" y="45">3</text>
  </g>
  <g fill="currentColor" font-size="15" text-anchor="middle">
    <text x="70" y="356">0</text>
    <text x="335" y="356">0.5</text>
    <text x="600" y="356">1</text>
    <text x="340" y="376">β = v / c</text>
    <text x="74" y="22">K / mc²</text>
  </g>
  <path d="M 600 40 L 600 330" fill="none" stroke="currentColor" stroke-width="1.2" stroke-dasharray="4 4" opacity="0.6" />
  <text x="606" y="52" fill="currentColor" font-size="14">v = c</text>
  <polyline points="70,330 123,329.5 176,328.1 229,325.7 282,322.3 335,317.9 388,312.6 441,306.3 494,299.1 547,290.9 600,281.7" fill="none" stroke="currentColor" stroke-width="2" stroke-dasharray="6 5" />
  <polyline points="70,330 123,329.5 176,328.0 229,325.3 282,321.2 335,315.0 388,305.8 441,291.3 494,265.6 520,243.2 547,204.9 560,172.3 573,117.1 579,81.4 583,40" fill="none" stroke="var(--sl-color-accent)" stroke-width="2.6" />
  <g font-size="15" fill="currentColor">
    <path d="M 110 62 L 150 62" fill="none" stroke="var(--sl-color-accent)" stroke-width="2.6" />
    <text x="158" y="67">relativistic K = (γ - 1) mc²</text>
    <path d="M 110 92 L 150 92" fill="none" stroke="currentColor" stroke-width="2" stroke-dasharray="6 5" />
    <text x="158" y="97">Newtonian K = mv² / 2</text>
  </g>
</svg>
</Figure>

### 5.2. What it looks like at low speed

<Proposition id="prop-newtonian-limit" title="The Newtonian limit and its error">
With $\beta = v/c$, we have for $|\beta| < 1$ the expansion

$$
K = (\gamma - 1)mc^2 = \frac{1}{2}mv^2\left(1 + \frac{3}{4}\beta^2 + \frac{5}{8}\beta^4 + \cdots\right).
$$

In particular $\dfrac{K - \frac{1}{2}mv^2}{\frac{1}{2}mv^2} = \dfrac{3}{4}\beta^2 + O(\beta^4)$, so the relative error of the Newtonian expression is of order $\beta^2$.
</Proposition>

<Proof of="prop-newtonian-limit">
Substitute $x = \beta^2$ into the binomial series $(1-x)^{-1/2} = 1 + \frac{1}{2}x + \frac{3}{8}x^2 + \frac{5}{16}x^3 + \cdots$ (valid for $|x| < 1$):

$$
\gamma = \left(1-\beta^2\right)^{-1/2} = 1 + \frac{1}{2}\beta^2 + \frac{3}{8}\beta^4 + \frac{5}{16}\beta^6 + \cdots .
$$

Therefore

$$
K = mc^2(\gamma - 1) = mc^2\left(\frac{1}{2}\beta^2 + \frac{3}{8}\beta^4 + \frac{5}{16}\beta^6 + \cdots\right)
= \frac{1}{2}mc^2\beta^2\left(1 + \frac{3}{4}\beta^2 + \frac{5}{8}\beta^4 + \cdots\right).
$$

Since $mc^2\beta^2 = mv^2$, the leading term is $\frac{1}{2}mv^2$. Rearranging the bracket gives the formula for the relative error.
</Proof>

This justifies the name. The velocity-dependent part of $E = \gamma mc^2$ is exactly the Newtonian kinetic energy at low speeds, so we are entitled to call $E$ an energy. And by <Ref to="thm-covariant-conservation" /> the sum of these $E$ is conserved across a reaction.

<Example id="ex-lhc-proton" title="How close to the speed of light is an LHC proton?">
The rest energy of the proton is $m_pc^2 = 938.272\ \mathrm{MeV}$. In Run 3 of the LHC, the protons in a single beam carry the energy $E = 6.8\ \mathrm{TeV} = 6.8\times 10^6\ \mathrm{MeV}$. By <Ref to="def-four-momentum" />,

$$
\gamma = \frac{E}{m_pc^2} = \frac{6.8\times10^6}{938.272} = 7.25\times10^3 .
$$

Now find the speed. From $\gamma^{-2} = 1-\beta^2$ we get $\beta = \sqrt{1-\gamma^{-2}}$, and since $\gamma^{-2} = 1/(7.248\times10^3)^2 = 1.904\times10^{-8}$ is very small we may use $\sqrt{1-x} \approx 1 - x/2$:

$$
\beta \approx 1 - \frac{1}{2\gamma^2} = 1 - 9.52\times10^{-9}.
$$

The shortfall relative to the speed of light is $c - v = 9.52\times10^{-9}\times 2.998\times10^8\ \mathrm{m/s} = 2.85\ \mathrm{m/s}$ — barely faster than walking pace. The LHC ring is $26.66\ \mathrm{km}$ around, so one revolution takes $26.66\times10^3 / (2.998\times10^8) = 8.89\times10^{-5}\ \mathrm{s}$, and per revolution light gets ahead of the proton by $2.85 \times 8.89\times10^{-5} = 2.5\times10^{-4}\ \mathrm{m}$, that is, by $0.25\ \mathrm{mm}$.

What would Newtonian mechanics predict for the same kinetic energy? Solving $\frac{1}{2}m_pv^2 = 6.8\ \mathrm{TeV}$ gives $v = c\sqrt{2\times 7.25\times10^3} = 120\,c$, a hundred and twenty times the speed of light. The measured speed of course never exceeds $c$. It is the divergence of $\gamma - 1$ as $v \to c$ in <Ref to="thm-kinetic-energy" /> that guarantees this bound.
</Example>

### 5.3. Rest energy

Setting $\boldsymbol p = \boldsymbol 0$ in <Ref to="thm-energy-momentum-relation" />, or $v = 0$ in <Ref to="thm-kinetic-energy" />, gives the famous formula

$$
E_0 = mc^2 .
$$

This is called the **rest energy**. The formula itself follows almost trivially from <Ref to="def-four-momentum" />, but its content is not trivial at all. The objection to answer is: "Have we not merely shifted the zero point of energy?" In Newtonian mechanics the zero of potential energy is arbitrary, and adding a constant changes no physics.

The answer is this. **$mc^2$ is indeed a constant, but $\sum_i m_i c^2$ is not.** If a reaction changes which particles are present, the total mass changes. By <Ref to="thm-covariant-conservation" /> the total $\sum E$ is conserved, so if $\sum m_ic^2$ decreases, exactly that amount must appear as kinetic energy or radiated energy. In Newtonian mechanics the number and kind of particles were assumed fixed, which is the only reason the zero point could be removed; in a world where particles are created and annihilated it can no longer be removed.

The size of the coefficient $c^2 = 8.99\times10^{16}\ \mathrm{m^2/s^2}$ is what gives the formula its bite. A mass of $1\ \mathrm{g}$ corresponds to

$$
E_0 = 1.0\times10^{-3}\ \mathrm{kg}\times 8.99\times10^{16}\ \mathrm{m^2/s^2} = 9.0\times10^{13}\ \mathrm{J}.
$$

Since $1\ \mathrm{kt}$ of TNT is $4.184\times10^{12}\ \mathrm{J}$, this is $9.0\times10^{13}/4.184\times10^{12} = 21.5\ \mathrm{kt}$ worth. The rest energy of an everyday object is orders of magnitude away from its everyday kinetic energy.

## 6. Mass is a form of energy

### 6.1. The invariant mass of a system

For a single particle the mass was determined by $p\cdot p = m^2c^2$. The same construction works for a system of several particles.

<Definition id="def-invariant-mass" title="Invariant mass of a system">
For a system of particles $i = 1,\dots,n$, let the total four-momentum be

$$
P^\mu \equiv \sum_{i=1}^n p_i^\mu = \left(\frac{\mathcal E}{c},\ \boldsymbol P\right),
\qquad \mathcal E = \sum_i E_i,\quad \boldsymbol P = \sum_i \boldsymbol p_i .
$$

Then

$$
M \equiv \frac{1}{c^2}\sqrt{\mathcal E^2 - |\boldsymbol P|^2c^2}
$$

is called the **invariant mass** of the system.
</Definition>

Being a sum of four-vectors, $P^\mu$ is a four-vector, and $M^2c^4 = P\cdot P\,c^2$ is Lorentz invariant by <Ref to="prop-invariance" />. So $M$ takes the same value in every inertial frame. If an inertial frame with $\boldsymbol P = \boldsymbol 0$ exists (the centre-of-momentum frame), then there $M c^2 = \mathcal E$: **the invariant mass is the total energy in the centre-of-momentum frame divided by $c^2$**.

<Theorem id="thm-invariant-mass" title="Non-additivity of mass">
For a system of $n$ particles of masses $m_i \ge 0$ (each satisfying $E_i > 0$ and $E_i^2 = |\boldsymbol p_i|^2c^2 + m_i^2c^4$), the invariant mass $M$ obeys

$$
M \ \ge\ \sum_{i=1}^n m_i .
$$

Equality holds precisely when all the four-momenta are parallel to one another, that is (for particles of positive mass) when all the velocities coincide.
</Theorem>

<Proof of="thm-invariant-mass">
We first treat the case $n = 2$. From <Ref to="def-invariant-mass" /> and the product of <Ref to="def-four-vector" />,

$$
M^2c^4 = (p_1+p_2)\cdot(p_1+p_2)\,c^2 = m_1^2c^4 + m_2^2c^4 + 2\left(E_1E_2 - c^2\,\boldsymbol p_1\cdot\boldsymbol p_2\right)
$$

(where $p_i\cdot p_i\,c^2 = m_i^2c^4$ is the identity established in the proof of <Ref to="thm-energy-momentum-relation" />). On the other hand,

$$
(m_1c^2+m_2c^2)^2 = m_1^2c^4+m_2^2c^4+2m_1m_2c^4 .
$$

Hence, to prove $M \ge m_1+m_2$ it suffices to prove the inequality

$$
E_1E_2 - c^2\,\boldsymbol p_1\cdot\boldsymbol p_2 \ \ge\ m_1m_2c^4 ,
$$

which we shall call inequality (A). Writing $q_i \equiv |\boldsymbol p_i|c \ge 0$, the Cauchy–Schwarz inequality gives $\boldsymbol p_1\cdot\boldsymbol p_2 \le |\boldsymbol p_1||\boldsymbol p_2|$, so

$$
E_1E_2 - c^2\,\boldsymbol p_1\cdot\boldsymbol p_2 \ \ge\ E_1E_2 - q_1q_2 .
$$

It therefore suffices to show $E_1E_2 \ge m_1m_2c^4 + q_1q_2$. Since $E_i = \sqrt{m_i^2c^4+q_i^2}$ and both sides are non-negative, it is enough to compare their squares.

$$
\begin{aligned}
(E_1E_2)^2 - \left(m_1m_2c^4+q_1q_2\right)^2
&= \left(m_1^2c^4+q_1^2\right)\left(m_2^2c^4+q_2^2\right) - \left(m_1m_2c^4+q_1q_2\right)^2\\
&= m_1^2m_2^2c^8 + m_1^2c^4q_2^2 + m_2^2c^4q_1^2 + q_1^2q_2^2\\
&\qquad - m_1^2m_2^2c^8 - 2m_1m_2c^4q_1q_2 - q_1^2q_2^2\\
&= c^4\left(m_1q_2 - m_2q_1\right)^2 \ \ge\ 0 .
\end{aligned}
$$

This proves inequality (A). Equality requires both steps to be equalities simultaneously, that is, $\boldsymbol p_1 \parallel \boldsymbol p_2$ (same direction) and $m_1q_2 = m_2q_1$. When $m_i > 0$, substituting $q_i = \gamma_i m_i v_i c$ turns the second condition into $m_1\gamma_2m_2v_2 = m_2\gamma_1m_1v_1$, that is, $\gamma_2v_2 = \gamma_1v_1$; since $w \mapsto \gamma(w)w$ is strictly increasing on $[0,c)$, this forces $v_1 = v_2$, and with the directions agreeing we get $\boldsymbol v_1 = \boldsymbol v_2$.

For $n \ge 3$ we induct. The sum $p_1+\cdots+p_{n-1}$ is a four-vector with $E > 0$ and $E^2 = |\boldsymbol p|^2c^2 + (M_{n-1}c^2)^2$, so the argument above applies verbatim to "one particle of mass $M_{n-1}$" together with the $n$-th particle, giving $M_n \ge M_{n-1} + m_n \ge \left(\sum_{i \le n-1} m_i\right) + m_n$.
</Proof>

What <Ref to="thm-invariant-mass" /> says is that **mass is not additive**, and moreover that the discrepancy always goes in the direction of an excess. Where does the excess come from? From the internal kinetic energy of the system and from the energy of the interactions. The next example is the most direct illustration.

<Example id="ex-inelastic-mass-gain" title="Colliding lumps of clay: heat becomes mass">
Let us redo the setting of <Ref to="ex-newtonian-failure" />, this time with the correct conservation law. In the frame S, two lumps of clay of mass $m$ collide head-on at speed $u$ and merge into a single lump of mass $M$. Write $\gamma_u \equiv \gamma(u)$.

**Conservation in S.** Momentum: $\gamma_u m u + \gamma_u m(-u) = 0$, so the momentum after the merger is $0$ as well and the lump is at rest. Energy: by <Ref to="thm-covariant-conservation" />,

$$
\gamma_u mc^2 + \gamma_u mc^2 = Mc^2
\quad\Longrightarrow\quad
M = 2\gamma_u m .
$$

Since $\gamma_u > 1$ we get $M > 2m$: the inequality of <Ref to="thm-invariant-mass" /> is realised with strict inequality. The increase is

$$
(M - 2m)c^2 = 2(\gamma_u-1)mc^2 = 2K,
$$

that is, **exactly the total kinetic energy that was lost**. The energy that turned into heat inside the clay shows up as mass of the merged lump. For $u = 0.6c$ we have $\gamma_u = 1/\sqrt{1-0.36} = 1/0.8 = 1.25$ and hence $M = 2.5\,m$: the mass has grown by $25\,\%$.

**Is it conserved in S′ too?** Let us redo the check that failed in <Ref to="ex-newtonian-failure" />. The speed of A in S′ was $u'_{\mathrm A} = 2u/(1+\beta^2)$ with $\beta = u/c$. Compute the corresponding $\gamma$:

$$
1 - \frac{u'^2_{\mathrm A}}{c^2} = \frac{(1+\beta^2)^2 - 4\beta^2}{(1+\beta^2)^2} = \frac{(1-\beta^2)^2}{(1+\beta^2)^2}
\quad\Longrightarrow\quad
\gamma(u'_{\mathrm A}) = \frac{1+\beta^2}{1-\beta^2} = \gamma_u^2\left(1+\beta^2\right).
$$

Hence, remembering that B is at rest, the momentum before the collision is

$$
\gamma(u'_{\mathrm A})\,m\,u'_{\mathrm A} + 0 = \gamma_u^2\left(1+\beta^2\right) m \cdot \frac{2u}{1+\beta^2} = 2\gamma_u^2\,m u .
$$

After the collision, the lump of mass $M = 2\gamma_u m$ moves with speed $u$, so

$$
\gamma(u)\,M\,u = \gamma_u\cdot 2\gamma_u m\cdot u = 2\gamma_u^2\,mu .
$$

The two agree. The check that was off by $26\,\%$ in <Ref to="ex-newtonian-failure" /> comes out exactly right the moment we replace the momentum by $\gamma m\boldsymbol v$ and give up the additivity of mass.
</Example>

### 6.2. What the experiments show

Changes of mass are far too small to measure at everyday energy scales. Taking $u = 10\ \mathrm{m/s}$ in <Ref to="ex-inelastic-mass-gain" /> gives $\gamma_u - 1 \approx \frac{1}{2}(u/c)^2 = 5.6\times10^{-16}$, so a $1\ \mathrm{kg}$ lump of clay gains about $10^{-15}\ \mathrm{kg}$. On nuclear scales, however, mass differences are routine measurable quantities.

<Example id="ex-deuteron" title="The binding energy of the deuteron">
The deuterium nucleus (the deuteron) is a bound state of one proton and one neutron. Their rest energies are

$$
m_pc^2 = 938.272\ \mathrm{MeV},\qquad m_nc^2 = 939.565\ \mathrm{MeV},\qquad m_dc^2 = 1875.613\ \mathrm{MeV}.
$$

The sum of the masses of a free proton and a free neutron is $938.272+939.565 = 1877.837\ \mathrm{MeV}/c^2$. The mass of the deuteron is smaller than that, the difference being

$$
B = (m_p + m_n - m_d)c^2 = 1877.837 - 1875.613 = 2.224\ \mathrm{MeV}.
$$

This is the **binding energy**, equal to the energy needed to pull the deuteron apart into a proton and a neutron. In the language of <Ref to="thm-invariant-mass" />, the invariant mass of a system consisting of a proton and a neutron at rest is $m_p+m_n$ (equality holds, since the velocities coincide); upon binding, the surplus energy $B$ is discarded as a photon, and the invariant mass of what remains drops to $m_d = m_p+m_n-B/c^2$.

As a fraction of mass this is $2.224/1877.837 = 1.18\times10^{-3}$, a decrease of $0.12\,\%$. Mass spectrometers in nuclear physics resolve this magnitude effortlessly, so here $E = mc^2$ is an everyday working formula. Indeed the inverse reaction $d + \gamma \to p + n$ (photodisintegration of the deuteron) does not occur for photons of energy below $2.224\ \mathrm{MeV}$.
</Example>

<Example id="ex-sun" title="How much lighter does the Sun get each second?">
The luminosity of the Sun is $L_\odot = 3.828\times10^{26}\ \mathrm{W}$. All the radiated energy comes out of the depletion of rest energy, so by <Ref to="thm-energy-momentum-relation" /> the mass loss per unit time is

$$
\frac{dm}{dt} = \frac{L_\odot}{c^2} = \frac{3.828\times10^{26}\ \mathrm{J/s}}{8.988\times10^{16}\ \mathrm{m^2/s^2}} = 4.26\times10^{9}\ \mathrm{kg/s}.
$$

That is $4.26$ million tonnes per second. Over the roughly $4.6$ billion years ($1.45\times10^{17}\ \mathrm{s}$) since the Sun was born, the mass lost is $4.26\times10^9 \times 1.45\times10^{17} = 6.2\times10^{26}\ \mathrm{kg}$, a mere $0.031\,\%$ of the solar mass $1.989\times10^{30}\ \mathrm{kg}$.

The energy source is the fusion of hydrogen. In the net process by which four protons ($4\times938.272 = 3753.09\ \mathrm{MeV}$) turn into a helium-4 nucleus ($3727.379\ \mathrm{MeV}$) together with two positrons and other products, about $26.7\ \mathrm{MeV}$ is released. As a fraction this is $26.7/3753.09 = 0.71\,\%$. The number shows that fusion is a technology which taps $0.7\,\%$ of the rest energy.
</Example>

<Remark id="rem-chemical-vs-nuclear">
Chemical reactions lose mass too. The heat of combustion of carbon, $\mathrm{C} + \mathrm{O_2} \to \mathrm{CO_2}$, is $393.5\ \mathrm{kJ/mol}$, that is $393500/(6.022\times10^{23}) = 6.53\times10^{-19}\ \mathrm{J} = 4.08\ \mathrm{eV}$ per molecule. The rest energy of the reactants is $44\ \mathrm{u} \times 931.494\ \mathrm{MeV/u} = 4.10\times10^{10}\ \mathrm{eV}$, so the fractional mass decrease is

$$
\frac{4.08}{4.10\times10^{10}} = 1.0\times10^{-10}.
$$

That is seven orders of magnitude smaller than the $7\times10^{-3}$ of fusion. Those seven orders of magnitude are what entitle chemistry textbooks to write that mass is conserved in a reaction. The law of conservation of mass was not wrong; it was an approximation good enough to be invisible at chemical precision. That is the accurate way to put it.
</Remark>

<Remark id="rem-proton-mass">
<Ref to="thm-invariant-mass" /> can also be read backwards: even a system consisting solely of massless particles has non-zero invariant mass. A system of two photons of energy $\varepsilon$ travelling in opposite directions has $\mathcal E = 2\varepsilon$ and $\boldsymbol P = \boldsymbol 0$, hence $M = 2\varepsilon/c^2$. Trap light inside a mirrored box and the box gets heavier.

This is no fantasy. Of the proton mass $938\ \mathrm{MeV}/c^2$, the contribution of the rest masses of its constituent up and down quarks ($2m_u + m_d \approx 9\ \mathrm{MeV}/c^2$) is only about $1\,\%$; the rest is the kinetic energy of the quarks and the energy of the gluon field. Most of our body weight is made of confined energy.
</Remark>

<Aside type="tip">
Reading $E = mc^2$ as "mass turns into energy" is slightly inaccurate. By <Ref to="thm-covariant-conservation" /> the total energy is always conserved, so nothing is being converted. The correct statement is that energy stored in the form of rest energy moves into another form, such as kinetic energy or radiation, and that when it does so the **invariant mass** of the system decreases. Mass is a form of energy — that is the theme of this article.
</Aside>

## 7. Exercises

<Exercise id="exr-newtonian-error" difficulty="Easy">
Let $\beta = v/c$.

1. For $v = 0.1c$ and for $v = 0.5c$, compute numerically $K = (\gamma-1)mc^2$ and $K_{\mathrm N} = \frac{1}{2}mv^2$ in units of $mc^2$, and evaluate $\left(K_{\mathrm N}-K\right)/K$.
2. How small must $v$ be kept for the relative error of the Newtonian expression to stay below $1\,\%$? Use the first-order approximation of <Ref to="prop-newtonian-limit" />.

<Solution>
**1.** Compute $\gamma = (1-\beta^2)^{-1/2}$.

For $v = 0.1c$ we have $\beta^2 = 0.01$ and $\gamma = 1/\sqrt{0.99} = 1.0050378$, so

$$
\frac{K}{mc^2} = 5.0378\times10^{-3},\qquad \frac{K_{\mathrm N}}{mc^2} = \frac{1}{2}(0.1)^2 = 5.0000\times10^{-3},
$$

$$
\frac{K_{\mathrm N}-K}{K} = \frac{5.0000-5.0378}{5.0378} = -7.5\times10^{-3}.
$$

The Newtonian value underestimates by $0.75\,\%$.

For $v = 0.5c$ we have $\beta^2 = 0.25$ and $\gamma = 1/\sqrt{0.75} = 1.1547005$, so

$$
\frac{K}{mc^2} = 0.154701,\qquad \frac{K_{\mathrm N}}{mc^2} = \frac{1}{2}(0.5)^2 = 0.125,
$$

$$
\frac{K_{\mathrm N}-K}{K} = \frac{0.125-0.154701}{0.154701} = -0.192 .
$$

An underestimate of $19.2\,\%$: no longer usable as an approximation.

**2.** By <Ref to="prop-newtonian-limit" />, $\left(K-K_{\mathrm N}\right)/K_{\mathrm N} = \frac{3}{4}\beta^2 + O(\beta^4)$, so it suffices to require $\frac{3}{4}\beta^2 < 0.01$, that is,

$$
\beta < \sqrt{\frac{4}{300}} = \sqrt{0.01333} = 0.1155 .
$$

Thus the Newtonian expression keeps $1\,\%$ accuracy up to about $v = 0.115c \approx 3.5\times10^{7}\ \mathrm{m/s}$. Since the first cosmic velocity is $7.9\times10^3\ \mathrm{m/s}$, relativistic corrections enter astronautics only at the level of $10^{-10}$.
</Solution>
</Exercise>

<Exercise id="exr-photon-absorption" difficulty="Standard">
An atom of mass $m$ at rest absorbs a single photon of energy $\varepsilon$ and becomes an (excited) atom of mass $M$. Take the four-momentum of the photon to be $\left(\varepsilon/c,\ \varepsilon/c,\ 0,\ 0\right)$.

1. Express the mass $M$ and the speed $v$ of the atom after absorption in terms of $m$, $\varepsilon$ and $c$.
2. For $\varepsilon \ll mc^2$, compute $M - m$ to second order, including the deviation from $\varepsilon/c^2$, and explain the physical meaning of that deviation.

<Solution>
**1.** By <Ref to="thm-covariant-conservation" /> all four components of the four-momentum are conserved. Writing $E$ for the energy and $p$ for the magnitude of the momentum of the atom after absorption,

$$
E = mc^2 + \varepsilon,\qquad pc = \varepsilon
$$

(the atom moves in the $x$ direction). Applying <Ref to="thm-energy-momentum-relation" /> to the atom after absorption,

$$
M^2c^4 = E^2 - p^2c^2 = \left(mc^2+\varepsilon\right)^2 - \varepsilon^2 = m^2c^4 + 2mc^2\varepsilon .
$$

Hence

$$
M = m\sqrt{1 + \frac{2\varepsilon}{mc^2}} .
$$

The speed follows from <Ref to="cor-velocity-and-massless" /> (1):

$$
v = \frac{pc^2}{E} = \frac{\varepsilon c}{mc^2+\varepsilon} .
$$

For $\varepsilon \ll mc^2$ this is $v \approx \varepsilon/(mc)$, which agrees with the non-relativistic recoil speed $v = p/m$.

**2.** Put $x \equiv \varepsilon/(mc^2) \ll 1$ and use $\sqrt{1+2x} = 1 + x - \frac{1}{2}x^2 + O(x^3)$:

$$
M = m\left(1 + x - \frac{x^2}{2} + \cdots\right) = m + \frac{\varepsilon}{c^2} - \frac{\varepsilon^2}{2mc^4} + \cdots .
$$

Naively one expects the mass to grow by exactly the energy absorbed, $M = m + \varepsilon/c^2$; in fact it grows slightly less. The difference $\varepsilon^2/(2mc^4)$ coincides with $\frac{1}{2}mv^2/c^2 \approx \frac{1}{2}m\left(\varepsilon/mc\right)^2/c^2 = \varepsilon^2/(2mc^4)$. In other words, part of the absorbed energy goes into the recoil kinetic energy of the atom and does not become internal energy (that is, mass). This difference becomes essential in the discussion of the Mössbauer effect.
</Solution>
</Exercise>

<Exercise id="exr-two-photon-mass" difficulty="Standard">
Two photons, each of energy $\varepsilon$, travel with an opening angle $\theta$ between them ($0 \le \theta \le \pi$).

1. Find the invariant mass $M$ of this two-photon system.
2. The neutral pion $\pi^0$ decays into two photons. When both photons are measured in the laboratory frame to have energy $100\ \mathrm{MeV}$ with an opening angle $\theta$, find the value of $\theta$ that gives $M c^2 = 134.98\ \mathrm{MeV}$.

<Solution>
**1.** By <Ref to="cor-velocity-and-massless" /> (2), each photon has momentum of magnitude $|\boldsymbol p_i| = \varepsilon/c$. Substitute into <Ref to="def-invariant-mass" />:

$$
\mathcal E = 2\varepsilon,\qquad
|\boldsymbol P|^2 = |\boldsymbol p_1|^2 + |\boldsymbol p_2|^2 + 2\boldsymbol p_1\cdot\boldsymbol p_2
= \frac{\varepsilon^2}{c^2}\left(1 + 1 + 2\cos\theta\right).
$$

Hence

$$
M^2c^4 = \mathcal E^2 - |\boldsymbol P|^2c^2 = 4\varepsilon^2 - 2\varepsilon^2\left(1+\cos\theta\right) = 2\varepsilon^2\left(1-\cos\theta\right).
$$

Using the half-angle identity $1-\cos\theta = 2\sin^2(\theta/2)$ gives $M^2c^4 = 4\varepsilon^2\sin^2(\theta/2)$, and since $\sin(\theta/2) \ge 0$,

$$
M = \frac{2\varepsilon}{c^2}\,\sin\frac{\theta}{2} .
$$

For $\theta = 0$ (same direction) we get $M = 0$, and for $\theta = \pi$ (opposite directions) $M = 2\varepsilon/c^2$, agreeing with the value in <Ref to="rem-proton-mass" />. This confirms that a system made only of massless particles can have non-zero mass.

**2.** Substitute $\varepsilon = 100\ \mathrm{MeV}$ and $Mc^2 = 134.98\ \mathrm{MeV}$ into the formula of part 1:

$$
\sin\frac{\theta}{2} = \frac{Mc^2}{2\varepsilon} = \frac{134.98}{200} = 0.6749
\quad\Longrightarrow\quad
\frac{\theta}{2} = 42.5^\circ,\qquad \theta = 85.0^\circ .
$$

Experiments run this backwards: computing $M$ for many photon pairs and histogramming the results produces a peak at $134.98\ \mathrm{MeV}/c^2$. This is particle identification by invariant mass, a basic technique of accelerator experiments.
</Solution>
</Exercise>

<Exercise id="exr-antiproton-threshold" difficulty="Hard">
Antiprotons are produced by firing protons at protons at rest (a liquid hydrogen target). The reaction is

$$
p + p \ \longrightarrow\ p + p + p + \bar p
$$

and the mass of $\bar p$ equals that of $p$, namely $m_p$. Take $m_pc^2 = 938.272\ \mathrm{MeV}$.

1. Find the minimum (threshold) kinetic energy $K$ of the incident proton.
2. Find the minimum kinetic energy required in each beam when the same reaction is produced by two beams colliding head-on (a collider), and compare with part 1.

<Solution>
**1.** The invariant mass $M$ of the system is Lorentz invariant by the argument of <Ref to="thm-invariant-mass" /> and unchanged across the reaction by <Ref to="thm-covariant-conservation" />. The reaction can occur provided the invariant mass of the final state is at least the sum of the masses of the final-state particles, that is, by <Ref to="thm-invariant-mass" />,

$$
Mc^2 \ \ge\ 4m_pc^2 .
$$

Equality holds when all four final-state particles move with the same velocity (all at rest in the centre-of-momentum frame), and this is the threshold.

Compute $M$ from the initial state. Let the incident proton have energy $E$ and momentum $\boldsymbol p_1$, while the target is at rest (energy $m_pc^2$, momentum $\boldsymbol 0$). By <Ref to="def-invariant-mass" />,

$$
\begin{aligned}
M^2c^4 &= \left(E + m_pc^2\right)^2 - |\boldsymbol p_1|^2c^2\\
&= E^2 - |\boldsymbol p_1|^2c^2 + 2Em_pc^2 + m_p^2c^4\\
&= m_p^2c^4 + 2Em_pc^2 + m_p^2c^4 = 2m_p^2c^4 + 2Em_pc^2 ,
\end{aligned}
$$

where we used $E^2 - |\boldsymbol p_1|^2c^2 = m_p^2c^4$ from <Ref to="thm-energy-momentum-relation" />. Inserting the threshold condition $M^2c^4 = 16m_p^2c^4$,

$$
16m_p^2c^4 = 2m_p^2c^4 + 2Em_pc^2
\quad\Longrightarrow\quad
E = 7m_pc^2 .
$$

The kinetic energy is, by <Ref to="thm-kinetic-energy" />, $K = E - m_pc^2 = 6m_pc^2 = 6\times938.272 = 5629\ \mathrm{MeV} \approx 5.63\ \mathrm{GeV}$. The reason that $5.63\ \mathrm{GeV}$ is needed to make a single antiproton of $0.94\ \mathrm{GeV}$ is that the four final-state particles cannot come to rest in the laboratory frame and are forced to carry off surplus kinetic energy. The Bevatron at Berkeley, which discovered the antiproton in 1955, was designed for $6.2\ \mathrm{GeV}$ — this estimate with a margin.

**2.** In a head-on collision the laboratory frame is itself the centre-of-momentum frame. With energy $E$ in each beam we have $\boldsymbol P = \boldsymbol 0$ and $\mathcal E = 2E$, so $Mc^2 = 2E$. The threshold condition $2E = 4m_pc^2$ gives $E = 2m_pc^2$, that is,

$$
K = E - m_pc^2 = m_pc^2 = 938\ \mathrm{MeV} .
$$

Even taking both beams together this is $1.88\ \mathrm{GeV}$, one third of the $5.63\ \mathrm{GeV}$ needed with a fixed target. The gap widens as the energy rises: with a fixed target $Mc^2 \propto \sqrt{E}$, whereas with a collider $Mc^2 \propto E$. This one-line comparison is the reason nearly all modern accelerators are colliders.
</Solution>
</Exercise>

## References

1. Shigenobu Sunakawa, *Sotaiseiriron no Kangaekata* (in Japanese), Iwanami Shoten (Butsuri no Kangaekata 4), 1993 — treats relativistic mechanics and the four-vector formalism carefully, organised around thought experiments.
2. E. F. Taylor and J. A. Wheeler, *Spacetime Physics*, 2nd ed., W. H. Freeman, 1992 — Chapter 7 ("Momenergy") covers four-momentum and Chapter 8 ("Collide. Conserve. Create!") covers collisions and particle production. This is the textbook closest in spirit to the present article.
3. W. Rindler, *Relativity: Special, General, and Cosmological*, 2nd ed., Oxford University Press, 2006 — the chapter on relativistic particle mechanics gives a systematic treatment of four-velocity, four-momentum and four-force.
4. A. Einstein, "Ist die Trägheit eines Körpers von seinem Energieinhalt abhängig?", *Annalen der Physik* 18 (1905), 639–641 — the three-page paper that first stated $E = mc^2$. It is the original source for the argument presented in the Appendix.
5. W. Bertozzi, "Speed and Kinetic Energy of Relativistic Electrons", *American Journal of Physics* 32 (1964), 551–555. [DOI: 10.1119/1.1970770](https://doi.org/10.1119/1.1970770) — the experiment that demonstrated directly, by time of flight, that electron speeds have an upper bound.
6. L. B. Okun, "The Concept of Mass", *Physics Today* 42, no. 6 (1989), 31–36. [DOI: 10.1063/1.881171](https://doi.org/10.1063/1.881171) — an article sorting out the confusion surrounding the term "relativistic mass".

## Appendix: Einstein's argument of 1905

In the main text we defined $E = \gamma mc^2$ in <Ref to="def-four-momentum" /> and then verified in <Ref to="thm-kinetic-energy" /> that it deserves the name energy. Historically the order was reversed: without any four-vector formalism, Einstein derived $\Delta m = E_0/c^2$ from a single thought experiment about the emission of light (reference [4]). The argument is still worth reading.

Suppose a body at rest in the inertial frame S emits, simultaneously, two pulses of light of energy $E_0/2$ each, one in the $+x$ direction and one in the $-x$ direction. The momenta of the two pulses are equal in magnitude and opposite in direction, so they cancel, and the body remains at rest in S after the emission. The decrease of energy as seen in S is, by definition, $E_0$.

Next move to the inertial frame S′, which travels with speed $V$ in the $-x$ direction relative to S. Seen from S′, the body moves with speed $V$ in the $+x$ direction. By <Ref to="cor-velocity-and-massless" />, the four-momentum of a light pulse is $\left(\varepsilon/c\right)(1,\pm1,0,0)$, with $\varepsilon = E_0/2$ and the sign given by the direction of travel. Since S′ is obtained by a boost of velocity $-V$, the transformation law of <Ref to="def-four-vector" /> with $\beta \to -\beta$ (where $\beta = V/c$) gives

$$
\frac{\varepsilon'_{\pm}}{c} = \gamma(V)\left(\frac{\varepsilon}{c} + \beta\cdot\left(\pm\frac{\varepsilon}{c}\right)\right)
\quad\Longrightarrow\quad
\varepsilon'_{\pm} = \gamma(V)\,\varepsilon\,(1 \pm \beta).
$$

This is precisely the Doppler effect for light. Adding the two, the terms in $\beta$ cancel:

$$
\varepsilon'_{+} + \varepsilon'_{-} = \gamma(V)\,\varepsilon\,\big[(1+\beta) + (1-\beta)\big] = 2\gamma(V)\,\varepsilon = \gamma(V)\,E_0 .
$$

So from S′ the body has lost the energy $\gamma(V)E_0$, which exceeds the loss $E_0$ measured in S by $\left(\gamma(V)-1\right)E_0$.

Here is the crux of the argument. The velocity of the body is unchanged by the emission: if it stays at rest in S, it keeps the speed $V$ in S′. Nevertheless the energy measured in S′ decreases by more than in S. Splitting the energy into "internal energy plus kinetic energy", the surplus $\left(\gamma(V)-1\right)E_0$ must be a decrease of kinetic energy. For the kinetic energy to decrease while the speed stays the same, the $m$ in $K = \left(\gamma(V)-1\right)mc^2$ (<Ref to="thm-kinetic-energy" />) must decrease. Writing $\Delta m$ for the decrease of mass,

$$
\left(\gamma(V)-1\right)E_0 = \left(\gamma(V)-1\right)\,\Delta m\,c^2
\quad\Longrightarrow\quad
\Delta m = \frac{E_0}{c^2}.
$$

Einstein himself expanded $\gamma$ to second order in $V/c$ and read the result as a decrease $\frac{1}{2}\left(E_0/c^2\right)V^2$ of kinetic energy, reaching the same conclusion. The closing sentence of the 1905 paper states, in effect, that the mass of a body is a measure of its energy content, and he proposed testing this with the decay of radium salts. It would be some thirty years before the test became feasible.

Note that what has been obtained here is a **relation between differences** — "a body that loses the energy $E_0$ loses the mass $E_0/c^2$" — and not the assertion that the whole rest energy equals $mc^2$. To reach the statement about the whole, one must start from the four-momentum, as we did in the main text. The division of labour between the historical argument and the systematic one is visible right here.

The four-momentum assembled in this article will be the foundation of the chapters that follow. In [An invitation to general relativity: the equivalence principle](/physics/relativity/equivalence-principle) we start from the fact that gravitational and inertial mass are equal (<Ref to="physics/relativity/equivalence-principle#def-two-masses" text="inertial mass and gravitational mass" />) and move towards the statement that everything possessing energy feels gravity. For if $E = mc^2$ is correct, then even light, which has no mass, ought to be bent by gravity — and that expectation arises quite naturally.
