Phyz
Home
  • Classical Mechanics
  • First Quantization
Special Relativity
Relativistic QM
Dirac Equation
QFT: Quantum Field Theory
Home
  • Classical Mechanics
  • First Quantization
Special Relativity
Relativistic QM
Dirac Equation
QFT: Quantum Field Theory
  • Special Relativity

Special Relativity

1. The Two Postulates

Special relativity rests on exactly two assumptions. The first is the principle of relativity: the laws of physics are the same in every inertial frame. No experiment performed entirely within a uniformly moving laboratory can detect that motion — there is no preferred rest frame. This was already implicit in Newtonian mechanics, where Newton's laws take the same form after a Galilean boost, but Einstein elevated it to a universal principle covering electromagnetism as well.

The second postulate is the one that breaks with Newtonian intuition: the invariance of the speed of light. The speed of light in vacuum, c≈3×108 m/sc \approx 3 \times 10^8\ \text{m/s}c≈3×108 m/s, is the same for all inertial observers regardless of the motion of the source or the observer. This is not obvious — it contradicts the Galilean addition of velocities — but it is what Maxwell's equations require, and every precision experiment since Michelson–Morley has confirmed it. Taken together, the two postulates force a revision of how time and space relate across different frames.

2. Natural Coordinates

Setting c=1c = 1c=1 — measuring time in the same units as distance, so one second equals 3×1083 \times 10^83×108 metres — strips the factors of ccc from every formula and makes the underlying geometry legible. The Lorentz transformation becomes t′=γ(t−vx)t' = \gamma(t - vx)t′=γ(t−vx), x′=γ(x−vt)x' = \gamma(x - vt)x′=γ(x−vt); the invariant interval becomes s2=Δt2−Δx2s^2 = \Delta t^2 - \Delta x^2s2=Δt2−Δx2; the energy–momentum relation becomes E2=m2+p2E^2 = m^2 + p^2E2=m2+p2. Velocities are dimensionless numbers between 0 and 1. Mass, energy, and momentum all share the same unit. The only price is that restoring SI units at the end requires inserting factors of ccc by dimensional analysis, which is straightforward once the physics is clear. Natural units are used throughout the rest of this page.

3. Galilean Transformation

Before Einstein, frames in relative motion were related by the Galilean transformation: t′=tt' = tt′=t and x′=x−vtx' = x - vtx′=x−vt, with yyy and zzz unchanged. Time is universal, and velocities add without limit — if a train moves at vvv and a passenger walks at uuu relative to the train, a platform observer sees them at u+vu + vu+v, with no ceiling. There is nothing in the Galilean rules that forbids speeds larger than ccc; light has no special status.

The geometric picture makes the problem vivid. Draw a spacetime diagram with xxx on the vertical axis and ttt on the horizontal (with c=1c = 1c=1, a light ray travels one unit of distance per unit of time). A light ray therefore traces a 45° line — it bisects the angle between the ttt- and xxx-axes. Under a Galilean boost the time axis stays horizontal (t′=tt' = tt′=t) while the x′x'x′-axis tilts upward toward it, so in the new frame the light ray no longer sits at 45° between t′t't′ and x′x'x′: different observers assign different speeds to light.

The Lorentz transformation fixes this by tilting both axes symmetrically toward the 45° light ray — the x′x'x′-axis rotates up toward the light ray and the t′t't′-axis rotates up toward the xxx-axis, both by the same hyperbolic angle — so the light ray always bisects them regardless of vvv. Keeping that bisector fixed at 45° is exactly what it means to preserve the speed of light in every frame. Galilean vs Lorentz boosts

Figure: the same boost in both frameworks — Galilean (left): t′=tt' = tt′=t stays horizontal and only x′x'x′ tilts up, so light no longer bisects the axes; Lorentz (right): both axes tilt symmetrically, so light always bisects t′t't′ and x′x'x′.

4. Lorentz Transformations

If two inertial frames SSS and S′S'S′ are aligned along the xxx-axis with S′S'S′ moving at velocity vvv relative to SSS, the Lorentz transformation relating their coordinates is

t′=γ ⁣(t−vxc2),x′=γ(x−vt),y′=y,z′=zt' = \gamma\!\left(t - \frac{vx}{c^2}\right), \qquad x' = \gamma(x - vt), \qquad y' = y, \qquad z' = z t′=γ(t−c2vx​),x′=γ(x−vt),y′=y,z′=z

where γ=1/1−v2/c2\gamma = 1/\sqrt{1 - v^2/c^2}γ=1/1−v2/c2​ is the Lorentz factor, always ≥1\geq 1≥1 and diverging as v→cv \to cv→c. In the limit v≪cv \ll cv≪c, γ→1\gamma \to 1γ→1 and these reduce to the Galilean transformation t′=tt' = tt′=t, x′=x−vtx' = x - vtx′=x−vt — Newton's kinematics is recovered as a low-velocity approximation.

The crucial novelty is the mixing of ttt and xxx: time is no longer universal. What one observer calls "simultaneous" (t1=t2t_1 = t_2t1​=t2​ at different xxx) another observer in relative motion generally does not. Simultaneity is frame-dependent, and this is not a failure of perception but a structural feature of spacetime.

Velocity addition is also modified. If an object moves at speed uuu in SSS, its speed in S′S'S′ is

u′=u−v1−uv/c2u' = \frac{u - v}{1 - uv/c^2} u′=1−uv/c2u−v​

Setting u=cu = cu=c gives u′=cu' = cu′=c for any vvv — light speed is the same in every frame, as required.

5. Spacetime and the Invariant Interval

Minkowski's insight was that the Lorentz transformations are rotations in a four-dimensional spacetime, but with a metric that mixes a spatial sign with a temporal one. Define the spacetime interval between two events as

s2=c2Δt2−Δx2−Δy2−Δz2s^2 = c^2\Delta t^2 - \Delta x^2 - \Delta y^2 - \Delta z^2 s2=c2Δt2−Δx2−Δy2−Δz2

This quantity is the same in every inertial frame — it is the Lorentz-invariant analog of distance. Depending on its sign, the separation is classified:

  • s2>0s^2 > 0s2>0: timelike — a signal traveling slower than ccc can connect the events; one can always find a frame where they occur at the same place at different times
  • s2=0s^2 = 0s2=0: lightlike (null) — only light connects them
  • s2<0s^2 < 0s2<0: spacelike — no causal influence can connect them; one can find a frame where they are simultaneous

The invariant interval replaces the Euclidean notion of absolute distance. Just as a rotation in space changes xxx and yyy individually while preserving x2+y2x^2 + y^2x2+y2, a Lorentz boost changes ttt and xxx individually while preserving c2t2−x2c^2 t^2 - x^2c2t2−x2.

6. Four-Vectors and Covariant Notation

The natural objects in special relativity are four-vectors, which transform under Lorentz boosts the same way (ct,x,y,z)(ct, x, y, z)(ct,x,y,z) does. The prototype is the spacetime four-position:

xμ=(ct, x, y, z),μ=0,1,2,3x^\mu = (ct,\, x,\, y,\, z), \qquad \mu = 0, 1, 2, 3 xμ=(ct,x,y,z),μ=0,1,2,3

The Minkowski metric ημν=diag(+1,−1,−1,−1)\eta_{\mu\nu} = \text{diag}(+1,-1,-1,-1)ημν​=diag(+1,−1,−1,−1) defines the inner product:

xμxμ=ημνxμxν=c2t2−x2−y2−z2=s2x^\mu x_\mu = \eta_{\mu\nu}x^\mu x^\nu = c^2t^2 - x^2 - y^2 - z^2 = s^2 xμxμ​=ημν​xμxν=c2t2−x2−y2−z2=s2

Any combination of four-vectors contracted with ημν\eta_{\mu\nu}ημν​ is a Lorentz scalar — frame-independent by construction. This is the systematic way to write physical laws that are automatically consistent with special relativity: build them out of four-vector contractions, and they hold in every inertial frame.

The four-velocity uμ=dxμ/dτu^\mu = dx^\mu/d\tauuμ=dxμ/dτ (derivative with respect to proper time) satisfies uμuμ=c2u^\mu u_\mu = c^2uμuμ​=c2 identically, and in the rest frame reduces to (c,0,0,0)(c, 0, 0, 0)(c,0,0,0). The four-momentum pμ=muμp^\mu = m u^\mupμ=muμ has components

pμ=(Ec, px, py, pz)p^\mu = \left(\frac{E}{c},\, p_x,\, p_y,\, p_z\right) pμ=(cE​,px​,py​,pz​)

where EEE is the relativistic energy and p=γmv\mathbf{p} = \gamma m \mathbf{v}p=γmv is the relativistic three-momentum.

7. Metric Tensor, Covariance, and Contravariance

Four-vectors come in two flavors distinguished by where their index sits. A contravariant vector AμA^\muAμ (index up) transforms the same way the coordinate displacement dxμdx^\mudxμ does under a Lorentz transformation Λμν\Lambda^\mu{}_\nuΛμν​:

A′μ=Λμν AνA'^\mu = \Lambda^\mu{}_\nu\, A^\nu A′μ=Λμν​Aν

A covariant vector AμA_\muAμ​ (index down) transforms by the inverse transpose, which for Lorentz transformations is (Λ−1)νμ(\Lambda^{-1})^\nu{}_\mu(Λ−1)νμ​:

Aμ′=(Λ−1)νμ AνA'_\mu = (\Lambda^{-1})^\nu{}_\mu\, A_\nu Aμ′​=(Λ−1)νμ​Aν​

The names come from how each type behaves under a change of coordinates: contravariant components transform inversely to the basis vectors (they "go against" the basis), covariant components transform the same way as the basis (they "go with" it). Derivatives ∂/∂xμ\partial/\partial x^\mu∂/∂xμ are the prototype covariant object; coordinate increments dxμdx^\mudxμ are the prototype contravariant one.

The Minkowski metric ημν=diag(+1,−1,−1,−1)\eta_{\mu\nu} = \text{diag}(+1,-1,-1,-1)ημν​=diag(+1,−1,−1,−1) is the machine that converts between them:

Aμ=ημνAν,Aμ=ημνAνA_\mu = \eta_{\mu\nu} A^\nu, \qquad A^\mu = \eta^{\mu\nu} A_\nu Aμ​=ημν​Aν,Aμ=ημνAν​

where ημν=diag(+1,−1,−1,−1)\eta^{\mu\nu} = \text{diag}(+1,-1,-1,-1)ημν=diag(+1,−1,−1,−1) is the inverse metric (numerically identical here, though that is special to flat spacetime). Lowering the index on the four-position gives xμ=(t,−x,−y,−z)x_\mu = (t, -x, -y, -z)xμ​=(t,−x,−y,−z): the time component is unchanged, the spatial components flip sign.

A contraction pairs one upper index with one lower index and sums over it, producing an object with two fewer indices:

AμBμ=A0B0+A1B1+A2B2+A3B3=A0B0−A⋅BA^\mu B_\mu = A^0 B_0 + A^1 B_1 + A^2 B_2 + A^3 B_3 = A^0 B_0 - \mathbf{A}\cdot\mathbf{B} AμBμ​=A0B0​+A1B1​+A2B2​+A3B3​=A0B0​−A⋅B

This Einstein summation convention — repeated index up/down means sum — is in force throughout. A fully contracted object has no free indices and is a Lorentz scalar: it takes the same numerical value in every inertial frame. The invariant interval s2=xμxμs^2 = x^\mu x_\mus2=xμxμ​, the rest mass m2=pμpμm^2 = p^\mu p_\mum2=pμpμ​, and the phase of a plane wave ϕ=kμxμ\phi = k^\mu x_\muϕ=kμxμ​ are all scalars.

A tensor of type (r,s)(r, s)(r,s) carries rrr contravariant and sss covariant indices, each transforming with its own Λ\LambdaΛ or Λ−1\Lambda^{-1}Λ−1:

T′μ1⋯μrν1⋯νs=Λμ1α1⋯Λμrαr (Λ−1)β1ν1⋯(Λ−1)βsνs  Tα1⋯αrβ1⋯βsT'^{\mu_1\cdots\mu_r}{}_{\nu_1\cdots\nu_s} = \Lambda^{\mu_1}{}_{\alpha_1}\cdots\Lambda^{\mu_r}{}_{\alpha_r}\,(\Lambda^{-1})^{\beta_1}{}_{\nu_1}\cdots(\Lambda^{-1})^{\beta_s}{}_{\nu_s}\; T^{\alpha_1\cdots\alpha_r}{}_{\beta_1\cdots\beta_s} T′μ1​⋯μr​ν1​⋯νs​​=Λμ1​α1​​⋯Λμr​αr​​(Λ−1)β1​ν1​​⋯(Λ−1)βs​νs​​Tα1​⋯αr​β1​⋯βs​​

The metric itself is a (0,2)(0,2)(0,2) tensor. Any equation written as a tensor equality — same index structure on both sides, all free indices consistent — is automatically valid in every Lorentz frame. This is the practical content of covariance: write physics as tensor equations and relativistic invariance is built in.

8. Mass, Energy

The Lorentz-invariant norm of the four-momentum gives the energy–momentum relation:

pμpμ=E2c2−∣p∣2=m2c2p^\mu p_\mu = \frac{E^2}{c^2} - \lvert\mathbf{p}\rvert^2 = m^2 c^2 pμpμ​=c2E2​−∣p∣2=m2c2

Rearranged:

E2=(mc2)2+(pc)2E^2 = (mc^2)^2 + (pc)^2 E2=(mc2)2+(pc)2

For a particle at rest (p=0\mathbf{p} = 0p=0) this collapses to E=mc2E = mc^2E=mc2 — rest mass is a form of energy. For a massless particle (m=0m = 0m=0, e.g. a photon) it gives E=pcE = pcE=pc, and from the four-velocity construction one can show such a particle must always travel at exactly ccc.

The total relativistic energy E=γmc2E = \gamma mc^2E=γmc2 splits into rest energy mc2mc^2mc2 and kinetic energy (γ−1)mc2(\gamma - 1)mc^2(γ−1)mc2. In the limit v≪cv \ll cv≪c, γ−1≈v2/2c2\gamma - 1 \approx v^2/2c^2γ−1≈v2/2c2, so kinetic energy →12mv2\to \tfrac{1}{2}mv^2→21​mv2 — again, Newtonian mechanics is recovered as the low-velocity limit.

The energy–momentum relation is the starting point for relativistic quantum mechanics: replacing E→iℏ ∂/∂tE \to i\hbar\,\partial/\partial tE→iℏ∂/∂t and p→−iℏ∇\mathbf{p} \to -i\hbar\nablap→−iℏ∇ in E2=(mc2)2+(pc)2E^2 = (mc^2)^2 + (pc)^2E2=(mc2)2+(pc)2 gives the Klein–Gordon equation, the first attempt at a relativistic wave equation, and the road that eventually leads to the Dirac equation and quantum field theory.

9. Maxwell Equations

Section 1 left a claim dangling: the two postulates are "what Maxwell's equations require." The reason is that Maxwell's equations are already Lorentz-covariant — they assemble out of four-vector and tensor objects exactly as §7 prescribes, so they keep the same form in every inertial frame. Special relativity did not fix Maxwell; it was built to accommodate it.

In natural units (c=1c = 1c=1), the four equations are

∇⋅E=ρ,∇⋅B=0,∇×E=−∂B∂t,∇×B=∂E∂t+j.\nabla \cdot \mathbf{E} = \rho, \qquad \nabla \cdot \mathbf{B} = 0, \qquad \nabla \times \mathbf{E} = -\frac{\partial \mathbf{B}}{\partial t}, \qquad \nabla \times \mathbf{B} = \frac{\partial \mathbf{E}}{\partial t} + \mathbf{j}. ∇⋅E=ρ,∇⋅B=0,∇×E=−∂t∂B​,∇×B=∂t∂E​+j.

Gauss and Ampère involve the sources; Faraday and "no magnetic monopoles" involve only the fields.

The covariant form packages charge density and current into the four-current Jμ=(ρ,j)J^\mu = (\rho, \mathbf{j})Jμ=(ρ,j), a four-vector whose divergence vanishes, ∂μJμ=0\partial_\mu J^\mu = 0∂μ​Jμ=0 — the continuity equation. The fields assemble into the electromagnetic field tensor, the antisymmetric derivative of the four-potential Aμ=(φ,A)A^\mu = (\varphi, \mathbf{A})Aμ=(φ,A):

Fμν=∂μAν−∂νAμ=(0−Ex−Ey−EzEx0−BzByEyBz0−BxEz−ByBx0).F^{\mu\nu} = \partial^\mu A^\nu - \partial^\nu A^\mu = \begin{pmatrix} 0 & -E_x & -E_y & -E_z \\ E_x & 0 & -B_z & B_y \\ E_y & B_z & 0 & -B_x \\ E_z & -B_y & B_x & 0 \end{pmatrix}. Fμν=∂μAν−∂νAμ=​0Ex​Ey​Ez​​−Ex​0Bz​−By​​−Ey​−Bz​0Bx​​−Ez​By​−Bx​0​​.

The electric and magnetic fields are not separate invariant objects — under a boost they mix into each other; FμνF^{\mu\nu}Fμν is the invariant. Maxwell's equations are then two tensor equations:

∂μFμν=Jν,∂λFμν+∂μFνλ+∂νFλμ=0.\partial_\mu F^{\mu\nu} = J^\nu, \qquad \partial_\lambda F_{\mu\nu} + \partial_\mu F_{\nu\lambda} + \partial_\nu F_{\lambda\mu} = 0. ∂μ​Fμν=Jν,∂λ​Fμν​+∂μ​Fνλ​+∂ν​Fλμ​=0.

E\mathbf{E}E and B\mathbf{B}B are therefore not distinct physical quantities but frame-dependent manifestations of the single object FμνF^{\mu\nu}Fμν — different observers slice the same antisymmetric tensor into different electric and magnetic parts. For a boost along x^\hat{\mathbf{x}}x^ (with c=1c = 1c=1),

E∥′=E∥,E⊥′=γ(E⊥+v×B),B∥′=B∥,B⊥′=γ(B⊥−v×E).\mathbf{E}'_\parallel = \mathbf{E}_\parallel, \quad \mathbf{E}'_\perp = \gamma(\mathbf{E}_\perp + \mathbf{v}\times\mathbf{B}), \qquad \mathbf{B}'_\parallel = \mathbf{B}_\parallel, \quad \mathbf{B}'_\perp = \gamma(\mathbf{B}_\perp - \mathbf{v}\times\mathbf{E}). E∥′​=E∥​,E⊥′​=γ(E⊥​+v×B),B∥′​=B∥​,B⊥′​=γ(B⊥​−v×E).

The classic illustration: a point charge at rest produces a purely electric field; an observer moving past it sees the same FμνF^{\mu\nu}Fμν sliced differently and reports a magnetic field as well — which is why moving charges experience magnetic forces in the first place. There is no "real" versus "apparent" field; the frame-invariant reality is the field tensor.

Conversely, the four familiar equations are the component expansion of the two tensor equations. In ∂μFμν=Jν\partial_\mu F^{\mu\nu} = J^\nu∂μ​Fμν=Jν, the ν=0\nu = 0ν=0 component is Gauss's law ∇⋅E=ρ\nabla\cdot\mathbf{E} = \rho∇⋅E=ρ and the ν=1,2,3\nu = 1, 2, 3ν=1,2,3 components are the three space components of Ampère's law; in the cyclic identity, the purely spatial index choice (λμν)=(123)(\lambda\mu\nu) = (123)(λμν)=(123) is ∇⋅B=0\nabla\cdot\mathbf{B} = 0∇⋅B=0 and the choices with one temporal index give Faraday's law. Maxwell's equations are not four independent postulates — they are the components of two tensor equations, unpacked in a chosen frame. That is why a boost merely rotates the components into one another (the mixing above) while the equations themselves never change form. The homogeneous pair is automatic: it is an identity once Fμν=∂μAν−∂νAμF^{\mu\nu} = \partial^\mu A^\nu - \partial^\nu A^\muFμν=∂μAν−∂νAμ is substituted — the field tensor is the exterior derivative of the four-potential.

Both tensor equations are contractions of four-vector/tensor objects, hence manifestly Lorentz-covariant: by §7 they hold unchanged in every inertial frame, and in particular the wave equation they imply propagates disturbances at exactly ccc for every observer. That is the precise sense in which Maxwell is compatible with the Lorentz transformation — the equations are the same in every frame, not merely similar.

10. Schrödinger Equation

The Schrödinger equation is what you get by quantizing the non-relativistic energy–momentum relation, and it is not Lorentz-covariant: time and space enter at different orders, so the equation singles out one frame.

Restoring ℏ\hbarℏ (still c=1c = 1c=1), the non-relativistic energy is E=p2/2mE = \mathbf{p}^2/2mE=p2/2m. Applying the quantization prescription of §8, E→iℏ ∂/∂tE \to i\hbar\,\partial/\partial tE→iℏ∂/∂t and p→−iℏ∇\mathbf{p} \to -i\hbar\nablap→−iℏ∇, gives the free Schrödinger equation

iℏ∂ψ∂t=−ℏ22m∇2ψ,i\hbar \frac{\partial \psi}{\partial t} = -\frac{\hbar^2}{2m}\nabla^2 \psi, iℏ∂t∂ψ​=−2mℏ2​∇2ψ,

with a potential term V(x)ψV(\mathbf{x})\psiV(x)ψ added by hand for interacting particles. The equation is first order in time but second order in space.

That asymmetry is exactly why it cannot be Lorentz-invariant. The Lorentz scalar built from two derivatives is the d'Alembertian ∂μ∂μ=∂t2−∇2\partial_\mu\partial^\mu = \partial_t^2 - \nabla^2∂μ​∂μ=∂t2​−∇2, which treats time and space democratically; the Schrödinger operator iℏ∂t+ℏ22m∇2i\hbar\partial_t + \tfrac{\hbar^2}{2m}\nabla^2iℏ∂t​+2mℏ2​∇2 has no such four-vector form, so a Lorentz boost does not preserve the equation — observers in relative motion would not agree that it holds.

Equivalently, look at plane waves ψ∝ei(p⋅x−Et)/ℏ\psi \propto e^{i(\mathbf{p}\cdot\mathbf{x} - Et)/\hbar}ψ∝ei(p⋅x−Et)/ℏ. The equation enforces the dispersion relation E=p2/2mE = \mathbf{p}^2/2mE=p2/2m, which is only approximate: it is the low-velocity limit of the exact relation E2=m2+p2E^2 = m^2 + \mathbf{p}^2E2=m2+p2 (§8), valid when ∣p∣≪m|\mathbf{p}| \ll m∣p∣≪m. Special relativity demands the second-order Klein–Gordon equation instead — at the price of negative-energy solutions, which point the way to the Dirac equation and antiparticles. The Schrödinger equation is the v≪cv \ll cv≪c limit of that story, and the starting point of First Quantization.

Last Updated: 8/31/26, 3:25 AM
Contributors: Hanh Huynh Huu