Skip to content

The Principles of Special Relativity: From Galilean Relativity to Einstein's Two Postulates

Prerequisite:Foundations of Newtonian Mechanics: From the Three Laws to Momentum and Energy Conservation

Raw
  • The Galilean principle of relativity rests on a computation: Newton’s equation of motion does not change form under a Galilean transformation. The principle does not assert that “absolute rest does not exist” but that “no physical law detects absolute rest”.
  • The wave equation that follows from Maxwell’s equations does change form under a Galilean transformation. Electromagnetism would therefore possess a preferred inertial frame, and that frame was taken to be the rest frame of the ether.
  • The Michelson-Morley experiment failed to detect the fringe shift (about 0.4 fringes) predicted by the ether hypothesis. The motion of the Earth relative to the ether is not found even at second order.
  • Einstein laid down two postulates: the special principle of relativity (all physical laws take the same form in every inertial frame) and the invariance of the speed of light (the speed of light in vacuum depends neither on the source nor on the inertial frame of the observer).
  • Granting both at once, absolute time t=tt' = t cannot be maintained mathematically. What had to go was the Galilean transformation, not the principle of relativity.
  • The immediate consequence is the relativity of simultaneity. The truth of the statement “two events happened at the same time” depends on the inertial frame in which it is judged.

1. Motivation: two successful theories that would not fit together

Section titled “1. Motivation: two successful theories that would not fit together”

At the end of the nineteenth century physics possessed two finished systems: Newtonian mechanics and Maxwell’s electromagnetism. Both agreed well with experiment, and both had the kind of completeness that invites the thought that nothing essential remains to be done. The trouble was that one could not believe in both at once.

Newtonian mechanics has a deep property going back to Galileo. Inside the cabin of a ship moving at constant velocity, no mechanical experiment can detect that the ship is moving. The laws of physics take exactly the same form in a laboratory at rest and in one in uniform rectilinear motion.

Maxwell’s equations, on the other hand, contain a speed cc as a constant, a value fixed by the permittivity and permeability of the vacuum alone and thus intrinsic to the theory. But a “speed” is supposed to be a quantity whose value depends on the coordinate system in which it is measured. In which coordinate system, then, is the cc appearing in Maxwell’s equations measured?

The natural answer was: in the frame in which the medium carrying light (the ether) is at rest. Just as sound travels at a fixed speed relative to the air, light travels at a fixed speed relative to the ether. But adopting this answer means that a preferred inertial frame — the rest frame of the ether — exists, and that the principle of relativity handed down from Galileo does not extend to electromagnetism. Since the Earth should be orbiting through the ether at 30 km per second, an “ether wind” should be blowing at the surface of the Earth, and it should be measurable.

In an autobiographical note written much later, Einstein records a question he had asked himself at the age of sixteen: what would one see if one chased a light beam at the speed of light? According to Maxwell’s equations light is an electromagnetic field that oscillates as it advances. If one could travel alongside it, one would see a field that oscillates but does not advance — something that is not a solution of Maxwell’s equations at all. Nobody has ever seen such a thing.

In this article we verify by computation how serious the contradiction was, and see why the answer Einstein gave in 1905 (his two postulates) forces us to abandon absolute time. Time dilation and length contraction themselves are treated in the next chapter, Lorentz transformations (time dilation(Theorem 3.3)[Lorentz Transformations], Lorentz contraction(Theorem 4.2)[Lorentz Transformations]).

2. Preliminaries: inertial frames and the Galilean principle of relativity

Section titled “2. Preliminaries: inertial frames and the Galilean principle of relativity”

Relativity takes as its basic object something that packages “when” and “where” together.

Definition 2.1Event

A point of spacetime, that is, a pair “where and when”, is called an event. Once an inertial frame is chosen, an event is represented by four numbers (t,x,y,z)(t, x, y, z). “A bulb flashed”, “two particles collided” — occurrences with neither extension nor duration — are events.

Definition 2.2Inertial frame

A coordinate system (a set of measuring rods mutually at rest together with synchronised clocks) in which a point mass free of external forces always appears to move uniformly in a straight line is called an inertial frame.

The notion of an inertial frame is the same one introduced in Foundations of Newtonian mechanics (definition of an inertial frame(Definition 3.1)[Foundations of Newtonian Mechanics]). What is new in relativity is the transformation law connecting inertial frames.

2.2. Galilean transformations and the principle of relativity

Section titled “2.2. Galilean transformations and the principle of relativity”

Definition 2.3Galilean transformation

Let (t,x,y,z)(t, x, y, z) be the coordinates of an inertial frame SS, and let (t,x,y,z)(t', x', y', z') be those of a frame SS' moving with constant speed vv in the positive xx direction relative to SS. Assume the origins coincide at t=0t = 0 and the axes are parallel. If the coordinates of one and the same event are related by

t=t,x=xvt,y=y,z=zt' = t,\qquad x' = x - vt,\qquad y' = y,\qquad z' = z

then this correspondence is called a Galilean transformation. The first equation asserts that “time is common to all frames”; this is the assumption of absolute time.

Axiom 2.4Galilean principle of relativity

All inertial frames are equivalent with respect to the laws of mechanics. That is, if a law of mechanics holds in one inertial frame, then a law of exactly the same form holds in every inertial frame in uniform rectilinear motion with respect to it.

Example 2.5Galileo's ship cabin

In his Dialogue of 1632 Galileo painted the following scene. Shut yourself in the hold of a large ship; release butterflies and flies, set out a bowl of goldfish, hang a bottle from the ceiling dripping water into a vessel below. While the ship stands still, the insects fly equally in all directions of the room, the drops fall straight into the vessel below, and when you jump along the floor you cover the same distance in either direction. Now let the ship move straight ahead at constant speed. Not one of these phenomena changes. Hence no observation made inside the cabin can decide whether the ship is moving or standing still.

What does the work here is the condition “straight ahead at constant speed”. If the ship accelerates or turns, the insects and the drops immediately betray it. What the principle of relativity declares equivalent is inertial frames, not arbitrary coordinate systems.

Why does nothing change inside the cabin? The reason is nothing more than the following computation concerning Newton’s equation of motion (the second law(Axiom 3.3)[Foundations of Newtonian Mechanics]).

Theorem 2.6Galilean invariance of Newton's equation of motion

Consider a system of NN point masses, and let mim_i be the mass of particle ii and ri(t)\boldsymbol{r}_i(t) its position in an inertial frame SS. Suppose the force on particle ii is a function of the relative position vectors alone,

Fi=Fi(r1r2, r1r3, , rN1rN).\boldsymbol{F}_i = \boldsymbol{F}_i(\boldsymbol{r}_1 - \boldsymbol{r}_2,\ \boldsymbol{r}_1 - \boldsymbol{r}_3,\ \ldots,\ \boldsymbol{r}_{N-1} - \boldsymbol{r}_N).

Then, if the equation of motion mir¨i=Fim_i \ddot{\boldsymbol{r}}_i = \boldsymbol{F}_i holds in SS, in any inertial frame SS' related to SS by a Galilean transformation the equation mir¨i=Fi(r1r2,)m_i \ddot{\boldsymbol{r}}\,'_i = \boldsymbol{F}_i(\boldsymbol{r}\,'_1 - \boldsymbol{r}\,'_2, \ldots) holds with the same function Fi\boldsymbol{F}_i.

Proof(Theorem 2.6)

Writing Definition 2.3 in vector form, with v\boldsymbol{v} the relative velocity vector, we have ri=rivt\boldsymbol{r}\,'_i = \boldsymbol{r}_i - \boldsymbol{v}t and t=tt' = t.

Consider first the time derivative. Since t=tt' = t we have ddt=ddt\dfrac{d}{dt'} = \dfrac{d}{dt}, so there is no need to relabel the derivative. Hence

r˙i=r˙iv,r¨i=r¨i.\dot{\boldsymbol{r}}\,'_i = \dot{\boldsymbol{r}}_i - \boldsymbol{v},\qquad \ddot{\boldsymbol{r}}\,'_i = \ddot{\boldsymbol{r}}_i .

Because v\boldsymbol{v} is a constant vector it disappears on the second differentiation; that is, acceleration is invariant under a Galilean transformation.

Consider next the force. The relative position vectors satisfy

rirj=(rivt)(rjvt)=rirj,\boldsymbol{r}\,'_i - \boldsymbol{r}\,'_j = (\boldsymbol{r}_i - \boldsymbol{v}t) - (\boldsymbol{r}_j - \boldsymbol{v}t) = \boldsymbol{r}_i - \boldsymbol{r}_j,

the terms vt\boldsymbol{v}t cancelling, so they too are invariant. By hypothesis the force is a function of the relative positions alone, so the value of Fi\boldsymbol{F}_i is unchanged.

The mass mim_i is, in Newtonian mechanics, a constant independent of the coordinate system. Putting all this together,

mir¨i=mir¨i=Fi(r1r2,)=Fi(r1r2,),m_i \ddot{\boldsymbol{r}}\,'_i = m_i \ddot{\boldsymbol{r}}_i = \boldsymbol{F}_i(\boldsymbol{r}_1 - \boldsymbol{r}_2, \ldots) = \boldsymbol{F}_i(\boldsymbol{r}\,'_1 - \boldsymbol{r}\,'_2, \ldots),

so the equation of motion holds in SS' with the same form.

Gravitation, the Coulomb force and the elastic force of a spring are all functions of relative position alone. Every practical situation in Newtonian mechanics therefore falls within the scope of this theorem, and the velocity of the ship never appears in experiments performed inside its cabin.

Corollary 2.7Galilean addition of velocities

If u\boldsymbol{u} is the velocity of a particle in SS and u\boldsymbol{u}' its velocity in SS', then u=uv\boldsymbol{u}' = \boldsymbol{u} - \boldsymbol{v}.

Proof(Corollary 2.7)

Differentiating r=rvt\boldsymbol{r}\,' = \boldsymbol{r} - \boldsymbol{v}t from Definition 2.3 with respect to t=tt' = t gives

u=drdt=ddt(rvt)=uv.\boldsymbol{u}' = \frac{d\boldsymbol{r}\,'}{dt'} = \frac{d}{dt}(\boldsymbol{r} - \boldsymbol{v}t) = \boldsymbol{u} - \boldsymbol{v}.

The assumption of absolute time, t=tt' = t, is used precisely where the variable of differentiation is exchanged.

Remark 2.8What the principle of relativity actually asserts

The principle of relativity is often summarised as “absolute rest does not exist”, but this is not accurate. What the principle says is that “no physical law detects absolute rest”. Even if a frame of absolute rest did exist, if there is in principle no way to measure one’s velocity relative to it, then the concept may be banished from physics. The controversy over the ether was meaningful exactly because people believed the ether wind could be measured.

3. Maxwell’s equations change form under a Galilean transformation

Section titled “3. Maxwell’s equations change form under a Galilean transformation”

In vacuum (a region free of charge and current) Maxwell’s equations read

E=0,B=0,×E=Bt,×B=ε0μ0Et.\nabla \cdot \boldsymbol{E} = 0,\qquad \nabla \cdot \boldsymbol{B} = 0,\qquad \nabla \times \boldsymbol{E} = -\frac{\partial \boldsymbol{B}}{\partial t},\qquad \nabla \times \boldsymbol{B} = \varepsilon_0 \mu_0 \frac{\partial \boldsymbol{E}}{\partial t}.

Let us confirm that a wave equation follows from them.

Proposition 3.1The wave equation for electromagnetic waves

An electric field E\boldsymbol{E} satisfying Maxwell’s equations in vacuum satisfies

2Et2=c22E,c=1ε0μ0.\frac{\partial^2 \boldsymbol{E}}{\partial t^2} = c^2 \nabla^2 \boldsymbol{E},\qquad c = \frac{1}{\sqrt{\varepsilon_0 \mu_0}}.

The magnetic field B\boldsymbol{B} satisfies the same equation.

Proof(Proposition 3.1)

Take the curl of both sides of the third equation. Applying the vector identity ×(×A)=(A)2A\nabla \times (\nabla \times \boldsymbol{A}) = \nabla(\nabla \cdot \boldsymbol{A}) - \nabla^2 \boldsymbol{A} to the left-hand side gives

×(×E)=(E)2E=2E,\nabla \times (\nabla \times \boldsymbol{E}) = \nabla(\nabla \cdot \boldsymbol{E}) - \nabla^2 \boldsymbol{E} = -\nabla^2 \boldsymbol{E},

where the last equality uses the first equation E=0\nabla \cdot \boldsymbol{E} = 0.

For the right-hand side, exchanging the order of the time and space derivatives,

×(Bt)=t(×B)=ε0μ02Et2,\nabla \times \left(-\frac{\partial \boldsymbol{B}}{\partial t}\right) = -\frac{\partial}{\partial t}(\nabla \times \boldsymbol{B}) = -\varepsilon_0\mu_0 \frac{\partial^2 \boldsymbol{E}}{\partial t^2},

where the last equality uses the fourth equation. Equating the two sides and cancelling the signs gives the asserted equation. For B\boldsymbol{B} take the curl of the fourth equation and use the second and third in the same way.

Example 3.2A combination of constants turns out to be the speed of light

Let us substitute ε0=8.8542×1012 F/m\varepsilon_0 = 8.8542 \times 10^{-12}\ \mathrm{F/m} and μ0=1.25664×106 N/A2\mu_0 = 1.25664 \times 10^{-6}\ \mathrm{N/A^2}.

ε0μ0=8.8542×1.25664×1018=1.11265×1017,\varepsilon_0 \mu_0 = 8.8542 \times 1.25664 \times 10^{-18} = 1.11265 \times 10^{-17},ε0μ0=3.3356×109,c=13.3356×109=2.998×108 m/s.\sqrt{\varepsilon_0\mu_0} = 3.3356 \times 10^{-9},\qquad c = \frac{1}{3.3356\times 10^{-9}} = 2.998 \times 10^{8}\ \mathrm{m/s}.

Combining two constants determined by purely electric and magnetic experiments produces the speed of light. When Maxwell himself performed this computation in 1862 he obtained about 3.11×1083.11\times10^8 m/s, in good agreement with Fizeau’s then-known measurement of the speed of light, 3.15×1083.15\times10^8 m/s. From this agreement the conclusion was drawn that light is an electromagnetic wave.

At the same time the computation shows that nothing whatsoever specifies “in which coordinate system” cc is the measured speed, because ε0\varepsilon_0 and μ0\mu_0 are mere numbers with no direction attached.

Let us now see what happens to this wave equation under a Galilean transformation. To keep things visible we consider the case of variation in the xx direction alone.

Theorem 3.3The wave equation is not Galilean invariant

Rewriting the one-dimensional wave equation

2ut2=c22ux2\frac{\partial^2 u}{\partial t^2} = c^2 \frac{\partial^2 u}{\partial x^2}

in the variables (t,x)(t', x') by means of Definition 2.3 yields

2ut22v2uxt=(c2v2)2ux2.\frac{\partial^2 u}{\partial t'^2} - 2v\,\frac{\partial^2 u}{\partial x' \partial t'} = (c^2 - v^2)\,\frac{\partial^2 u}{\partial x'^2}.

As long as v0v \ne 0, this is not of the same form as the original equation.

Proof(Theorem 3.3)

Assume uu is of class C2C^2 and use the chain rule. Since t=tt' = t and x=xvtx' = x - vt, differentiating partially with respect to xx at fixed tt and using x/x=1\partial x'/\partial x = 1, t/x=0\partial t'/\partial x = 0 gives

x=x.\frac{\partial}{\partial x} = \frac{\partial}{\partial x'} .

Differentiating partially with respect to tt at fixed xx, and using x/t=v\partial x'/\partial t = -v, t/t=1\partial t'/\partial t = 1, gives

t=tvx.\frac{\partial}{\partial t} = \frac{\partial}{\partial t'} - v\frac{\partial}{\partial x'} .

Apply this operator twice. Since uu is C2C^2 the mixed partial derivatives commute, so

2t2=(tvx)2=2t22v2xt+v22x2.\frac{\partial^2}{\partial t^2} = \left(\frac{\partial}{\partial t'} - v\frac{\partial}{\partial x'}\right)^2 = \frac{\partial^2}{\partial t'^2} - 2v\frac{\partial^2}{\partial x' \partial t'} + v^2 \frac{\partial^2}{\partial x'^2} .

Substituting these into the wave equation gives

2ut22v2uxt+v22ux2=c22ux2,\frac{\partial^2 u}{\partial t'^2} - 2v\frac{\partial^2 u}{\partial x' \partial t'} + v^2 \frac{\partial^2 u}{\partial x'^2} = c^2 \frac{\partial^2 u}{\partial x'^2},

and moving v22u/x2v^2 \partial^2 u/\partial x'^2 to the right-hand side yields the asserted equation.

Compared with the original form, a cross term 2v2u/xt-2v\,\partial^2 u/\partial x'\partial t' has appeared and the coefficient on the right has changed from c2c^2 to c2v2c^2 - v^2. Neither can be removed when v0v \ne 0. Indeed, if the same form t2u=c2x2u\partial_{t'}^2 u = c^2 \partial_{x'}^2 u held in SS' as well, subtracting the two would force 2v2u/xt=v22u/x22v\,\partial^2 u/\partial x'\partial t' = v^2 \partial^2 u /\partial x'^2 for every solution, a condition that severely restricts the set of solutions and does not hold in general.

Corollary 3.4The speed of light in a moving frame

A wave travelling in the +x+x direction with speed cc in SS travels with speed cvc - v in SS'. A wave travelling in the x-x direction travels with speed c+vc + v in SS'.

Proof(Corollary 3.4)

The general solution of the one-dimensional wave equation is d’Alembert’s u=f(xct)+g(x+ct)u = f(x - ct) + g(x + ct). Inverting Definition 2.3 gives x=x+vtx = x' + vt' and t=tt = t', so

xct=x+vtct=x(cv)t,x+ct=x+(c+v)t.x - ct = x' + vt' - ct' = x' - (c - v)t',\qquad x + ct = x' + (c + v)t'.

Hence, as seen from SS', the waveform represented by ff advances in the positive direction with speed cvc - v and the one represented by gg advances in the negative direction with speed c+vc + v. This agrees with what one obtains by applying Corollary 2.7 directly to light.

So as long as we accept the Galilean transformation, the speed of light must change to c±vc \pm v with the motion of the observer. Since Maxwell’s equations single out one value cc, there can be only one preferred inertial frame in which the equations hold in their given form. That frame is the rest frame of the ether.

Example 3.5How large is the ether wind?

The orbital speed of the Earth is about v=2.98×104v = 2.98\times10^4 m/s. Compared with c=3.00×108c = 3.00\times10^8 m/s,

βvc=2.98×1043.00×108=9.9×105,β2=9.9×109108.\beta \equiv \frac{v}{c} = \frac{2.98\times10^4}{3.00\times10^8} = 9.9 \times 10^{-5},\qquad \beta^2 = 9.9\times10^{-9} \approx 10^{-8}.

An effect of first order in β\beta, that is, a relative difference of order 10410^{-4}, was comfortably within reach of optical measurements at the end of the nineteenth century. But as we shall see, in a measurement that sends light out and back the first-order effect cancels, and what remains is the second-order effect β2108\beta^2 \sim 10^{-8}, one part in a hundred million. Capturing that requires an instrument called an interferometer.

flowchart TD
A["Newtonian mechanics: form unchanged by Galilean transformations"] --> C{"Irreconcilable"}
B["Maxwell's equations: the light speed c is a constant of the equations"] --> C
C --> D["Response 1: there is a preferred inertial frame, the ether"]
C --> E["Response 2: Maxwell's equations are the incomplete ones"]
C --> F["Response 3: the Galilean transformation is the mistaken one"]
D --> G["Refuted by the Michelson-Morley experiment"]
E --> H["Conflicts with stellar aberration and binary-star observations"]
F --> I["Special relativity"]
The logical impasse at the end of the nineteenth century, and the three roads leading out of it

Response 2 came in several variants. The proposal that the ether is completely dragged along by the Earth conflicts with stellar aberration (the annual variation in the apparent position of a star with the Earth’s orbital motion, discovered by Bradley in 1728). Emission theories, in which the speed of light adds to the velocity of the source, discard Maxwell’s equations and, as de Sitter pointed out in 1913, make the apparent motion of binary stars disagree with observation. Fizeau’s flowing-water experiment of 1851 showed that light travelling through water is dragged by exactly Fresnel’s coefficient 11/n21 - 1/n^2, supporting neither complete dragging nor a simple emission theory. Thus response 1 — that an ether exists which is not at rest relative to the Earth — remained the leading candidate.

If an ether wind is real, it should show up as a difference in the round-trip times of light. The interferometer devised by Michelson sends light out and back along two perpendicular arms and reads the time difference off as an interference pattern.

SourceDetectorMirror M1Mirror M2Beam splitterArm 1 (length L)Arm 2 (length L)Ether wind v
The Michelson interferometer. Light from the source is split in two by a half-silvered mirror, reflected back by the respective mirrors, and recombined to form interference fringes

Proposition 4.1The fringe shift predicted by the ether hypothesis

Let an interferometer with two arms, both of length LL, move with speed vv relative to the ether in the direction of arm 1. Assume that the speed of light is cc in the rest frame of the ether and that the speed of light as seen from the interferometer obeys Corollary 2.7. Writing β=v/c\beta = v/c, the round-trip times along the two arms are

t1=2Lc11β2,t2=2Lc11β2t_1 = \frac{2L}{c}\cdot\frac{1}{1 - \beta^2},\qquad t_2 = \frac{2L}{c}\cdot\frac{1}{\sqrt{1 - \beta^2}}

and for β1\beta \ll 1

t1t2Lβ2c=Lv2c3.t_1 - t_2 \simeq \frac{L\beta^2}{c} = \frac{Lv^2}{c^3}.

Rotating the apparatus through 9090^\circ interchanges the roles of the two arms, so with light of wavelength λ\lambda the fringe shift observed during the rotation is

ΔN=2c(t1t2)λ2Lv2λc2fringes.\Delta N = \frac{2c\,(t_1 - t_2)}{\lambda} \simeq \frac{2Lv^2}{\lambda c^2} \quad \text{fringes}.
Proof(Proposition 4.1)

Arm 1 (parallel to the ether wind). As seen from the interferometer, the outward light travels against the wind with speed cvc - v and the return light travels with the wind at speed c+vc + v (Corollary 2.7). Hence

t1=Lcv+Lc+v=L(c+v)+L(cv)(cv)(c+v)=2Lcc2v2=2Lc11β2.t_1 = \frac{L}{c - v} + \frac{L}{c + v} = \frac{L(c+v) + L(c-v)}{(c-v)(c+v)} = \frac{2Lc}{c^2 - v^2} = \frac{2L}{c}\cdot\frac{1}{1-\beta^2}.

Arm 2 (perpendicular to the ether wind). Here it is quicker to work in the rest frame of the ether. If τ\tau is the time for one leg, the light travels a distance cτc\tau during it. Meanwhile the mirror moves a distance vτv\tau in the direction of travel of the interferometer (opposite to the wind), and since the light must reach the tip of an arm of length LL, its path is the hypotenuse of a right triangle. By the Pythagorean theorem,

(cτ)2=(vτ)2+L2τ=Lc2v2.(c\tau)^2 = (v\tau)^2 + L^2 \quad\Longrightarrow\quad \tau = \frac{L}{\sqrt{c^2 - v^2}} .

For the round trip,

t2=2Lc2v2=2Lc11β2.t_2 = \frac{2L}{\sqrt{c^2 - v^2}} = \frac{2L}{c}\cdot\frac{1}{\sqrt{1-\beta^2}} .

Estimating the difference. For β1\beta \ll 1, using (1β2)1=1+β2+O(β4)(1-\beta^2)^{-1} = 1 + \beta^2 + O(\beta^4) and (1β2)1/2=1+12β2+O(β4)(1-\beta^2)^{-1/2} = 1 + \tfrac{1}{2}\beta^2 + O(\beta^4),

t1t2=2Lc[(1+β2)(1+12β2)]+O(β4)=Lβ2c+O(β4).t_1 - t_2 = \frac{2L}{c}\left[\left(1 + \beta^2\right) - \left(1 + \tfrac{1}{2}\beta^2\right)\right] + O(\beta^4) = \frac{L\beta^2}{c} + O(\beta^4).

Note that no term of first order in β\beta appears: it cancelled because the light was sent out and back (Example 3.5).

The effect of rotation. Rotating the apparatus by 9090^\circ makes arm 1 perpendicular and arm 2 parallel to the wind, so the time difference changes sign and becomes (t1t2)-(t_1 - t_2). The optical path difference therefore changes by c(t1t2)c((t1t2))=2c(t1t2)c\,(t_1-t_2) - c\,(-(t_1-t_2)) = 2c(t_1-t_2) across the rotation. Since the fringes shift by one for every change of λ\lambda in the path difference, the shift is ΔN=2c(t1t2)/λ=2Lβ2/λ=2Lv2/(λc2)\Delta N = 2c(t_1-t_2)/\lambda = 2L\beta^2/\lambda = 2Lv^2/(\lambda c^2) fringes.

Example 4.2The numbers the 1887 experiment aimed at

By reflecting the light back and forth many times, Michelson and Morley extended the effective arm length to L=11L = 11 m. Taking for vv the orbital speed of the Earth, 3.0×1043.0\times10^4 m/s, and for λ\lambda the value 5.9×1075.9\times10^{-7} m (the yellow sodium line), Proposition 4.1 gives

ΔN=2×11×(3.0×104)2(5.9×107)×(3.0×108)2=1.98×10105.31×1010=0.37\Delta N = \frac{2 \times 11 \times (3.0\times10^4)^2}{(5.9\times10^{-7}) \times (3.0\times10^8)^2} = \frac{1.98\times10^{10}}{5.31\times10^{10}} = 0.37

fringes. The numerator is 2×11×9.0×108=1.98×10102\times11\times9.0\times10^8 = 1.98\times10^{10} and the denominator is 5.9×107×9.0×1016=5.31×10105.9\times10^{-7}\times9.0\times10^{16} = 5.31\times10^{10}.

The sensitivity of the apparatus had been pushed to the point of reading a shift of 0.01 fringes. The expected 0.4 fringes is forty times that. Yet the shift actually observed fell below 0.01 fringes. The two authors concluded that the velocity of the Earth relative to the ether is probably less than one sixth of the orbital speed. Since ΔN\Delta N is proportional to v2v^2, a speed one sixth as large gives a fringe shift one thirty-sixth as large, that is, 0.01 fringes.

Remark 4.3What did this experiment refute?

The Michelson-Morley experiment did not directly prove the invariance of the speed of light. What it established is a single negative fact: no ether wind is detected at second order.

There were, in fact, other ways to account for the same result. Lorentz and FitzGerald proposed the hypothesis that a body moving through the ether with speed vv contracts along the direction of motion by a factor 1β2\sqrt{1-\beta^2}. Inserting this into Proposition 4.1 reproduces the null result exactly (see the Appendix). The ether could survive.

Einstein’s paper of 1905 mentions at the outset “the unsuccessful attempts to detect a motion of the Earth relative to the light medium”, but does not name the Michelson-Morley experiment. The motivation he puts first is the asymmetry between magnet and conductor described in the next example.

Example 5.1Magnet and conductor: two explanations for one phenomenon

Move a magnet and a coil relative to one another and a current flows in the coil. The electromagnetic explanation is entirely different according to which of the two is said to be moving.

If the magnet is held fixed and the coil is moved with speed vv, the charges inside the coil move through the magnetic field B\boldsymbol{B} and are set in motion by the Lorentz force qv×Bq\boldsymbol{v}\times\boldsymbol{B}. No electric field appears.

If instead the coil is held fixed and the magnet is moved with speed vv, the magnetic field at the position of the coil varies in time, so by ×E=B/t\nabla \times \boldsymbol{E} = -\partial \boldsymbol{B}/\partial t an electric field E\boldsymbol{E} arises and pushes the charges. This time no magnetic force appears.

The fields invoked as the cause are different, yet the resulting current depends only on the relative velocity, and no experiment can distinguish the two cases. Einstein cited this asymmetry in the opening paragraph of his 1905 paper and said it is a symptom that the concept of “absolute rest” has no counterpart in nature. There is an asymmetry in the theory but none in the phenomena; the verdict is that it is the theory that must be repaired.

The two postulates he then laid down are as follows.

Axiom 5.2Special principle of relativity

All inertial frames are equivalent with respect to every physical law, not only the laws of mechanics but those of electromagnetism as well. That is, the equations expressing the laws of physics take the same form — including the values of the constants appearing in them — in all inertial frames in uniform rectilinear motion with respect to one another.

Axiom 5.3Invariance of the speed of light

The speed of light propagating in vacuum always has the same value cc, independently of whether the source is moving or at rest, and independently of the inertial frame in which it is measured.

Remark 5.4The differing roles of the two postulates

The first postulate merely extends the scope of Axiom 2.4 to electromagnetism; in content it is conservative. It is the second postulate that is radical.

Looked at carefully, the second postulate contains two assertions: (a) within one inertial frame, the speed of light does not depend on the motion of the source; (b) the value of that speed is the same in every inertial frame. In fact (b) follows from (a) together with the first postulate. Suppose (a) holds as a law of physics in an inertial frame SS. By Axiom 5.2, the same law holds in the same form in any inertial frame SS'. The constant cc appearing in the law is, by Proposition 3.1, a value determined by the universal constants ε0,μ0\varepsilon_0, \mu_0, and “the same form” means that even the values of the constants agree; hence the speed of light in SS' is the same cc.

As for (a), it is a rather natural property from the point of view of waves. It is the same as the fact that the speed of sound in air does not depend on whether the source is moving. The second postulate is therefore not a wholly unmotivated leap but the result of joining “light is a wave” to “the principle of relativity” in a straightforward way.

Once the two postulates are granted, the Galilean transformation collapses immediately. And what collapses is not the part x=xvtx' = x - vt but the part t=tt' = t.

Theorem 5.5Absolute time and the invariance of light speed are incompatible

Suppose inertial frames SS and SS' are related, with constants a,ba, b, by the transformation

t=t,x=ax+bt.t' = t,\qquad x' = ax + bt.

If light in vacuum travels with the same speed cc in both the +x+x and the x-x directions in SS and in SS' alike, then a=1a = 1 and b=0b = 0; hence x=xx' = x, that is, SS' is at rest relative to SS.

Proof(Theorem 5.5)

Consider light travelling in the +x+x direction. Its trajectory in SS can be written, with a constant x0x_0, as x=x0+ctx = x_0 + ct. Substituting into the transformation,

x=a(x0+ct)+bt=ax0+(ac+b)t=ax0+(ac+b)t,x' = a(x_0 + ct) + bt = ax_0 + (ac + b)\,t = ax_0 + (ac + b)\,t',

the last equality using the hypothesis t=tt' = t. Hence the speed of this light as seen from SS' is ac+bac + b. Requiring this to equal cc gives

ac+b=c.(i)ac + b = c. \tag{i}

Similarly, substituting the trajectory x=x0ctx = x_0 - ct of light travelling in the x-x direction gives

x=ax0+(ac+b)t,x' = ax_0 + (-ac + b)\,t',

so the speed in SS' is ac+b|-ac+b|. Requiring it to travel in the negative direction with speed cc gives

ac+b=c.(ii)-ac + b = -c. \tag{ii}

Adding (i) and (ii) gives 2b=02b = 0, that is, b=0b = 0; subtracting them gives 2ac=2c2ac = 2c, and since c0c \ne 0, a=1a = 1. Hence x=xx' = x.

The origin of SS' is defined by x=0x' = 0, which now means x=0x = 0 in SS. The origin of SS' does not move in SS, so the relative velocity of the two frames is 00.

Remark 5.6Why we may restrict to linear transformations

In Theorem 5.5 we assumed the transformation to be linear in xx and tt. There is a reason for this assumption. By the definition of an inertial frame (Definition 2.2), a particle moving uniformly in a straight line in one frame must do so in the other as well. The coordinate transformation is therefore a map that “carries straight lines of spacetime to straight lines”. Granting in addition that spacetime is homogeneous (no point is special), such a map must be affine, and hence linear once the origins are made to coincide. A detailed discussion of this point, and the procedure that actually derives the Lorentz transformation from it, is given in Lorentz transformations (the Lorentz transformation (boost)(Theorem 6.1)[Lorentz Transformations]).

Corollary 5.7Abandoning absolute time

If Axiom 5.2 and Axiom 5.3 are accepted, the coordinate transformation connecting two inertial frames with nonzero relative velocity does not satisfy t=tt' = t. In other words, the assumption that the rate of time is independent of the inertial frame must be discarded.

Proof(Corollary 5.7)

By Remark 5.6 we may take the transformation connecting inertial frames to be linear. If t=tt' = t held, all the hypotheses of Theorem 5.5 would be satisfied: indeed, that the speed of light is cc in both frames is guaranteed by Axiom 5.3. The conclusion would then be that the relative velocity is 00, contrary to hypothesis.

This is the branch point. The physicists of the nineteenth century never doubted t=tt' = t. Time was one single thing in the world, shared by all observers — that had been the premise since Newton. Einstein saw that this very premise was the true culprit setting the two theories at odds.

What, concretely, does abandoning absolute time amount to abandoning? The first casualty is the absoluteness of “simultaneous”.

Theorem 5.8Relativity of simultaneity

A carriage whose length measured in the ground frame is \ell runs straight ahead with constant speed vv (0<v<c0 < v < c) relative to the ground. A source at the centre of the carriage emits light signals towards the front and the rear simultaneously. Then:

  1. In the inertial frame moving with the carriage, the two signals reach the front and the rear ends at the same time.
  2. In the ground frame, the arrival at the rear end precedes the arrival at the front end by
Δt=vc2v2.\Delta t = \frac{\ell v}{c^2 - v^2}.

Hence the truth of the statement “two events happened at the same time” depends on the inertial frame in which it is judged.

Proof(Theorem 5.8)

Proof of 1. In the frame moving with the carriage, the distance from the source to the front end equals that to the rear end (the source was placed at the centre). In this frame too, by Axiom 5.3, light travels at the same speed cc both forwards and backwards. Equal distances at equal speeds take equal times, so the signals arrive simultaneously.

Proof of 2. Work in the ground frame. Let t=0t = 0 be the instant of emission and x=0x = 0 the position of the source at that instant. By Axiom 5.3, the speed of light in the ground frame is cc even though the source is moving. This is the decisive difference from the classical theory: it is not c+vc + v or cvc - v as in Corollary 2.7.

The front end is at x=/2x = \ell/2 at t=0t = 0 and moves with speed vv, so its position is /2+vt\ell/2 + vt. The forward light is at x=ctx = ct, so the arrival time t+t_+ satisfies

ct+=2+vt+t+=2(cv).c\,t_+ = \frac{\ell}{2} + v\,t_+ \quad\Longrightarrow\quad t_+ = \frac{\ell}{2(c - v)} .

The rear end is at /2+vt-\ell/2 + vt and the backward light is at x=ctx = -ct, so the arrival time tt_- satisfies

ct=2+vtt=2(c+v).-c\,t_- = -\frac{\ell}{2} + v\,t_- \quad\Longrightarrow\quad t_- = \frac{\ell}{2(c + v)} .

Since 0<v<c0 < v < c we have cv<c+vc - v < c + v, hence t+>tt_+ > t_-: the arrival at the rear end comes first. The difference is

t+t=2(1cv1c+v)=2(c+v)(cv)c2v2=vc2v2,t_+ - t_- = \frac{\ell}{2}\left(\frac{1}{c-v} - \frac{1}{c+v}\right) = \frac{\ell}{2}\cdot\frac{(c+v)-(c-v)}{c^2 - v^2} = \frac{\ell v}{c^2 - v^2},

which is the asserted Δt\Delta t.

By 1 and 2, the two events “the light reached the front end” and “the light reached the rear end” occur at the same time in the frame of the carriage and at different times in the frame of the ground.

Remark 5.9What follows from here

The relativity of simultaneity is not merely a curious phenomenon; it is the source of almost every consequence of relativity.

To measure “the length of a rod”, for instance, one must read off the positions of its two ends at the same time. Since the meaning of “at the same time” differs from frame to frame, it is only to be expected that length does too. The meaning of “two separated clocks agree” likewise differs from frame to frame, so the same happens when the rates of clocks are compared. The quantitative derivation of length contraction and time dilation is carried out in Lorentz transformations (the general form of the offset of simultaneity is written down there as the relativity of simultaneity(Theorem 5.1)[Lorentz Transformations]).

It is equally important that the relativity of simultaneity does not destroy causality. The formula for Δt\Delta t never reverses the order of an event and an event that follows the arrival of light from it. The order can be reversed only for two events too far apart to be joined even by light, that is, events that cannot influence each other (the absoluteness of causal structure(Proposition 7.3)[Lorentz Transformations]).

Let us take stock.

ConceptIts fate in special relativity
Principle of relativityExtended from mechanics to all physical laws; strengthened, if anything
Absolute time t=tt' = tAbandoned (Corollary 5.7)
Galilean transformationAbandoned, though it survives as a good approximation for β1\beta \ll 1
EtherUnnecessary, since there is no means of detecting it (Remark 2.8)
Maxwell’s equationsUntouched; no modification needed
Newton’s equation of motionRequires modification; treated two chapters later
CausalityPreserved

The conclusion of this chapter is that what was discarded was not electromagnetism but the framework on the side of mechanics. How the equation of motion is to be repaired is the subject of Relativistic mechanics; the starting point is the frame dependence of Newtonian momentum conservation(Example 1.1)[Relativistic Mechanics], and the keys are the four-momentum(Definition 4.1)[Relativistic Mechanics] and the frame independence of its conservation law(Theorem 4.2)[Relativistic Mechanics].

Modern tests reach a precision orders of magnitude beyond that of the original. Michelson-Morley type experiments are repeated in the form of rotating optical cavities, and any directional dependence of the speed of light is excluded at the level Δc/c1017\Delta c/c \lesssim 10^{-17}.

Finally, let us record the point of view that sees the principle of relativity as a symmetry. That the laws are invariant under the transformations passing from one inertial frame to another is a symmetry of spacetime itself. And symmetries are tied to conservation laws (Noether’s theorem, Theorem 4.1[対称性と保存則]). Special relativity ends up rewriting the definitions of momentum and energy precisely because the group of symmetries is replaced, from the Galilean group to the Lorentz group.

Exercise 7.1Easy

Sound waves travel with speed csc_s relative to the air. Consequently, in a wind of speed ww, an observer fixed to the ground measures the speed of sound to be cs+wc_s + w downwind and cswc_s - w upwind, so that it differs with direction. Does this contradict the principle of relativity? State what is different in the case of light.

Solution

There is no contradiction. The principle of relativity asserts that the form of the physical laws is the same in every inertial frame, not that every measured value is the same in every inertial frame. Air is a real material substance, and its distribution and motion form part of the initial and boundary conditions in each frame. The fact that a wind is blowing is a circumstance, not a law, so the form of the laws is unaffected even though the speed of sound depends on direction. Indeed, the equations of motion of air (the laws of fluid mechanics) are invariant under Galilean transformations.

In the case of light there are two problems. First, the ether was a medium “postulated solely to carry light and having no other effect whatsoever”. Air can be measured independently with an anemometer, but the ether has no means of detection other than light. Second, one might still hope to measure it using the propagation of light itself — this was precisely the Michelson-Morley experiment — and the velocity was not detected (Example 4.2). A quantity for which no means of detection exists in principle may be banished from physics; that is the position stated in Remark 2.8.

Exercise 7.2Standard

Suppose the two arms of the interferometer are of unequal lengths L1L_1 (initially parallel to the ether wind) and L2L_2 (initially perpendicular). Under the same assumptions as in Proposition 4.1, show that the fringe shift on rotating the apparatus through 9090^\circ is

ΔN(L1+L2)v2λc2.\Delta N \simeq \frac{(L_1 + L_2)v^2}{\lambda c^2}.

Check also that for L1=L2=LL_1 = L_2 = L this reproduces the result of Proposition 4.1.

Solution

Before the rotation, the same computation as in the proof of Proposition 4.1 gives the round-trip times

t1=2L1c11β2,t2=2L2c11β2.t_1 = \frac{2L_1}{c}\cdot\frac{1}{1-\beta^2},\qquad t_2 = \frac{2L_2}{c}\cdot\frac{1}{\sqrt{1-\beta^2}} .

Expanding to first order in β2\beta^2,

t1t22c[L1(1+β2)L2(1+12β2)]=2(L1L2)c+β2c(2L1L2).t_1 - t_2 \simeq \frac{2}{c}\left[L_1(1+\beta^2) - L_2\left(1+\tfrac{1}{2}\beta^2\right)\right] = \frac{2(L_1 - L_2)}{c} + \frac{\beta^2}{c}\left(2L_1 - L_2\right).

After the rotation arm 1 is perpendicular and arm 2 parallel, so interchanging the roles in the same formula,

t1t22c[L1(1+12β2)L2(1+β2)]=2(L1L2)c+β2c(L12L2).t'_1 - t'_2 \simeq \frac{2}{c}\left[L_1\left(1+\tfrac{1}{2}\beta^2\right) - L_2(1+\beta^2)\right] = \frac{2(L_1-L_2)}{c} + \frac{\beta^2}{c}\left(L_1 - 2L_2\right).

The term 2(L1L2)/c2(L_1-L_2)/c coming from the difference of the arm lengths is common to both and cancels in the difference. What remains is

(t1t2)(t1t2)=β2c[(2L1L2)(L12L2)]=β2c(L1+L2).(t_1 - t_2) - (t'_1 - t'_2) = \frac{\beta^2}{c}\left[(2L_1 - L_2) - (L_1 - 2L_2)\right] = \frac{\beta^2}{c}(L_1 + L_2).

Converting to an optical path difference and dividing by the wavelength,

ΔN=cλ[(t1t2)(t1t2)]=β2(L1+L2)λ=(L1+L2)v2λc2.\Delta N = \frac{c}{\lambda}\left[(t_1-t_2)-(t'_1-t'_2)\right] = \frac{\beta^2 (L_1+L_2)}{\lambda} = \frac{(L_1+L_2)v^2}{\lambda c^2}.

Setting L1=L2=LL_1 = L_2 = L gives 2Lv2/(λc2)2Lv^2/(\lambda c^2), in agreement with Proposition 4.1. The point of the computation is that the effect does not disappear even when the arms are of unequal length.

Exercise 7.3Standard

A flash is emitted at the origin of an inertial frame SS at time t=0t = 0. By Axiom 5.3, the wavefront at time tt is the sphere

x2+y2+z2=c2t2.x^2 + y^2 + z^2 = c^2 t^2 .

In an inertial frame SS' whose origin coincides with that of SS at t=0t' = 0, the wavefront must for the same reason satisfy x2+y2+z2=c2t2x'^2 + y'^2 + z'^2 = c^2 t'^2. Verify by direct substitution that the Galilean transformation (Definition 2.3) fails to meet this requirement.

Solution

Inverting Definition 2.3 gives x=x+vtx = x' + vt', y=yy = y', z=zz = z', t=tt = t'. Substitute these into the quantity obtained by subtracting the right-hand side of the wavefront equation in SS from its left-hand side.

x2+y2+z2c2t2=(x+vt)2+y2+z2c2t2=x2+2vxt+v2t2+y2+z2c2t2=(x2+y2+z2c2t2)+2vxt+v2t2.\begin{aligned} x^2 + y^2 + z^2 - c^2t^2 &= (x' + vt')^2 + y'^2 + z'^2 - c^2 t'^2\\ &= x'^2 + 2vx't' + v^2t'^2 + y'^2 + z'^2 - c^2t'^2\\ &= \left(x'^2 + y'^2 + z'^2 - c^2t'^2\right) + 2vx't' + v^2 t'^2 . \end{aligned}

The left-hand side vanishes on the wavefront in SS. Hence, in order for the wavefront in SS' to satisfy x2+y2+z2c2t2=0x'^2+y'^2+z'^2-c^2t'^2 = 0 as well, the extra terms must vanish:

2vxt+v2t2=vt(2x+vt)=0.2vx't' + v^2t'^2 = vt'\left(2x' + vt'\right) = 0 .

This would have to hold at every point (t,x,y,z)(t', x', y', z') of the wavefront; but a wavefront with t>0t' > 0 contains many points with different values of xx', so 2x+vt=02x' + vt' = 0 cannot hold at all of them. Therefore v=0v = 0, and for inertial frames in relative motion the Galilean transformation fails to meet the requirement. This is a three-dimensional restatement of Theorem 5.5.

The computation is also a bridge to the next chapter. Asking which linear transformations leave the combination x2+y2+z2c2t2x^2+y^2+z^2-c^2t^2 (the spacetime interval(Definition 7.1)[Lorentz Transformations]) invariant produces the Lorentz transformation (invariance of the spacetime interval(Theorem 7.2)[Lorentz Transformations]).

Exercise 7.4Standard

For Δt=v/(c2v2)\Delta t = \ell v/(c^2-v^2) of Theorem 5.8, compute the numerical value in the following two cases, then compare and discuss the results. Take c=3.0×108c = 3.0\times10^8 m/s.

  1. A spacecraft of length =100\ell = 100 m flying at v=0.50cv = 0.50c.
  2. A train of length =400\ell = 400 m running at 300 km/h.
Solution

1. Since v=0.50cv = 0.50c, we have c2v2=c2(10.25)=0.75c2c^2 - v^2 = c^2(1 - 0.25) = 0.75\,c^2, so

Δt=100×0.50c0.75c2=100×0.500.75c=66.73.0×108=2.2×107 s\Delta t = \frac{100 \times 0.50\,c}{0.75\,c^2} = \frac{100 \times 0.50}{0.75\,c} = \frac{66.7}{3.0\times10^8} = 2.2\times10^{-7}\ \text{s}

that is, about 0.22 microseconds. This is the time for light to travel 67 m, of the same order as the length of the spacecraft, so it cannot be neglected when discussing the order of events on board.

2. A speed of 300 km/h is v=300×103/3600=83.3v = 300\times10^3/3600 = 83.3 m/s. Since vcv \ll c we may put c2v2c2=9.0×1016c^2 - v^2 \approx c^2 = 9.0\times10^{16}, so

Δt=400×83.39.0×1016=3.33×1049.0×1016=3.7×1013 s\Delta t = \frac{400 \times 83.3}{9.0\times10^{16}} = \frac{3.33\times10^{4}}{9.0\times10^{16}} = 3.7\times10^{-13}\ \text{s}

that is, about 0.37 picoseconds.

Discussion. Since Δt\Delta t is proportional to the first power of vv, it is strictly nonzero even in case 2. The relativity of simultaneity is not a new phenomenon that sets in only at high speed; it is always occurring, and has merely been too small to notice. Nobody could ever detect 0.37 picoseconds by human perception, but measured against the period of an optical oscillation (roughly 2 femtoseconds for visible light) it amounts to about 180 periods, an amount that optical methods can measure comfortably. The only reason absolute time went unquestioned in everyday life is that cc is large, which makes v/c2\ell v/c^2 small.

  • A. Einstein, “Zur Elektrodynamik bewegter Körper”, Annalen der Physik 17 (1905), 891–921. — The original paper on special relativity. The introduction and Part I §1–§2 correspond to Section 5 of this article.
  • A. A. Michelson and E. W. Morley, “On the Relative Motion of the Earth and the Luminiferous Ether”, American Journal of Science 34 (1887), 333–345. — The construction of the interferometer, the expected fringe shift and the observed result are all recorded here.
  • E. F. Taylor and J. A. Wheeler, Spacetime Physics, 2nd ed., W. H. Freeman, 1992 — Chapter 1. Its plan of putting the relativity of simultaneity first is close to the one adopted here.
  • Katsuhiko Sato, Sotaiseiriron (Iwanami Kiso Butsuri Series 9), Iwanami Shoten, 1996 — Chapter 1. A standard introduction available in Japanese (in Japanese).
  • R. P. Feynman, R. B. Leighton, M. Sands, The Feynman Lectures on Physics, Vol. I, Ch. 15 “The Special Theory of Relativity”. The full text is available on the official site.
  • S. Herrmann et al., “Rotating optical cavity experiment testing Lorentz invariance at the 101710^{-17} level”, Physical Review D 80 (2009), 105011. — A modern version of the Michelson-Morley experiment.

Appendix: The Lorentz-FitzGerald contraction as “another explanation”

Section titled “Appendix: The Lorentz-FitzGerald contraction as “another explanation””

Let us look at the route, mentioned in Remark 4.3, that explains the null result while keeping the ether.

FitzGerald (1889) and Lorentz (1892) proposed the hypothesis that a body moving through the ether with speed vv has its length along the direction of motion multiplied by 1β2\sqrt{1-\beta^2}. Arm 1 of the interferometer is parallel to the wind and so contracts, LL1β2L \to L\sqrt{1-\beta^2}, while arm 2 is perpendicular and stays at LL. Substituting into t1t_1 of Proposition 4.1,

t1=2L1β2c11β2=2Lc11β2=t2,t_1 = \frac{2L\sqrt{1-\beta^2}}{c}\cdot\frac{1}{1-\beta^2} = \frac{2L}{c}\cdot\frac{1}{\sqrt{1-\beta^2}} = t_2,

so the time difference vanishes exactly. However the apparatus is turned, the fringes do not move. The explanation is that the ether is real, but the wind cannot be measured because the measuring instrument contracts.

This hypothesis had two weaknesses. One is that it was an assumption introduced solely to rescue the observations, offering no explanation of why bodies should contract (Lorentz later attempted a mechanical justification, arguing that this is what one should expect if intermolecular forces are electromagnetic in origin). The other is that the contraction alone does not account for all the experiments.

Indeed, in 1932 Kennedy and Thorndike used an interferometer with deliberately very unequal arms and confirmed that the fringes do not move even as the velocity of the Earth changes with the seasons. As we saw in Exercise 7.2, with unequal arms the effect does not cancel through contraction alone; a null result is obtained only once one also grants that the rate of time itself is multiplied by 1β2\sqrt{1-\beta^2} (time dilation(Theorem 3.3)[Lorentz Transformations]). A direct measurement of time dilation was achieved by Ives and Stilwell in 1938, who measured the Doppler shift of light emitted by moving atoms both in the forward and in the backward direction.

Taken together, these three experiments (Michelson-Morley, Kennedy-Thorndike, Ives-Stilwell) determine the transformation connecting inertial frames uniquely to be the Lorentz transformation. Lorentz’s own theory of 1904 gave the same predictions as Einstein’s for all observable quantities. The difference lies in the physical interpretation. For Lorentz, both the contraction and the slowing of time were mechanical effects caused by motion relative to an undetectable ether at rest. For Einstein they were properties of the structure of time and space themselves, emerging as kinematics from the two postulates. The judgement at work here is to prefer the option that does not postulate undetectable structure. The same judgement recurs later, when the theory of gravity is built from the equivalence principle (An invitation to general relativity, Einstein's equivalence principle(Definition 3.3)[一般相対性理論への招待]).

Report an error in this article ・Operated by: Mugen Giken LLCPricingTermsLegal notice

© 2026 夢現技研合同会社 ・Feeding the text to an LLM is welcome. Code samples are MIT licensed.