The Dirac Equation
This page continues from Relativistic QM §4: the Dirac equation with its four-component spinor. The questions that section had to defer — what the extra components are, and what the negative-energy branch means — are taken up here.
1. Spin
Relativistic QM §4 deferred the meaning of the spinor's extra components. The identification is now forced by the equation itself. Begin with the standard definition: the total angular momentum of a particle is the sum of its orbital and intrinsic parts,
where is the orbital angular momentum of the motion and the intrinsic angular momentum. The Dirac equation shows that the intrinsic part is not optional. With the free Hamiltonian ,
so orbital angular momentum alone is not conserved. Conservation is a commutator statement: an observable with no explicit time dependence evolves by the Heisenberg equation of motion (§2)
so it is a constant of the motion exactly when — the commutator measures how fast the observable changes. That is the criterion used throughout this section: rules out as a conserved quantity, and the intrinsic term must restore .
The matrices act on the spinor's internal components, coupling the motion to degrees of freedom that — a purely spatial operator — cannot see. What must the intrinsic term be? Two requirements constrain it: it must cancel the deficit, and it must be an angular momentum — its components must obey the angular-momentum commutation table, a property to be checked once the candidate is found. The first requirement fixes the size:
Where does the candidate come from? It is already inside the equation. From Relativistic QM §4, the Dirac matrices are built from Pauli matrices, , so a product of two of them is block-diagonal: . Antisymmetrizing the product, the Pauli commutator reappears in each block,
so the commutators of the equation's own matrices generate the block-diagonal Pauli matrix — the spin structure latent in the equation:
Its commutator with the Hamiltonian follows from :
This is twice the needed deficit — the factor 2 is the one already sitting in the Pauli commutator . Scaling down by exactly that factor produces the operator whose commutator cancels the deficit:
The coefficient is fixed, not chosen: is the ratio of the deficit to what provides, — the from the canonical commutator inside , the 2 from the Pauli matrices. The check:
so the total angular momentum commutes with . Neither orbital angular momentum nor spin is separately conserved — the equation's structure mixes them, exactly as a relativistic theory should — but their sum is.
What kind of operator is this ? Because the satisfy , the components of satisfy
the angular-momentum algebra — the defining commutation relations of angular momentum. "Algebra" in the sense of a set closed under the commutator: the commutator of any two components is again a component (, cyclically), so the three operators form a self-contained structure. The identification carries the physics: any three operators satisfying this table are angular momentum — orbital obeys the same relations, and the Pauli matrices are one particular realization of it, the smallest, not its source — and the table is the infinitesimal statement that rotations about different axes do not commute. The eigenvalues follow at once: has eigenvalues , so has eigenvalues . The extra components are spin, and the equation describes a spin-½ particle; the two components of each pair are spin-up and spin-down.
The equation also fixes the magnetic moment. Coupling to an electromagnetic field (minimal substitution ) gives, in the non-relativistic limit, the Pauli equation with
Dirac's famous prediction: the electron's gyromagnetic ratio is twice the classical value. This answers the first deferred question of Relativistic QM §4 — the extra components are the two spin states of a spin-½ particle, and their transformation properties are the subject of §5 below.
2. Conservation and Commutators
The criterion used in §1 — an observable is conserved exactly when it commutes with the Hamiltonian — deserves a proof. An observable with no explicit time dependence evolves by the Heisenberg equation of motion,
so it is a constant of the motion exactly when : the commutator measures how fast the observable changes. The equation is proved by carrying the time evolution in the operator itself. In the Heisenberg picture,
with . Differentiating — using , valid because is time-independent and so commutes with its own exponential —
Taking expectation values in any state gives : the expectation value is constant exactly when the commutator vanishes. An explicit time dependence in would add a term ; none of the operators on this page has one.
This is the criterion applied in §1: orbital angular momentum fails it — — and the intrinsic term is exactly the correction that restores it, . The same criterion is used, without further proof, on the pages that follow.
3. Antiparticles
The second deferred question: the negative-energy branch. The fastest way in is to solve the equation: plane waves separate into two cases, rest and moving, and the solutions at rest are the whole structure in embryo.
At rest (). The Hamiltonian is simply , so a plane wave with constant four-vector satisfies
Where do these two values come from? The equation is an eigenvalue problem: has eigenvalues (twice) and (twice) — the content of its diagonal form — so multiplying by fixes to exactly the two values and , the two roots of . These are the two branches of the dispersion relation of Relativistic QM §2 seen at : iterating the Dirac equation reproduces , and at rest that is simply the rest energy with either sign, . The equation does not choose between them — both are realized, one in each pair of components. Each eigenspace of being two-dimensional, there are exactly four independent solutions, the basis vectors of the four-component space: two with , supported on the upper pair, and two with , supported on the lower pair,
The four components are forced, not chosen: the equation needs four mutually anticommuting matrices, and in the three Pauli matrices cannot be extended by a fourth — each anticommutes with the others but not with itself — so the matrices, and with them , live in (Relativistic QM §4): two pairs, the particle and antiparticle components of the previous page.
The two members of each pair are the two spin states along — the eigenstates of the spin projection . Call the four rest solutions . Since ,
the two eigenvalues, each occurring once per pair: the first member of each pair is spin up (), the second spin down (). On each pair the operator acts as , the factor being the spin quantum number of §1 — not a consequence of the block structure — and the total spin is , computed directly from the matrices (, since each ): a magnitude , larger than the largest projection — the spin vector never lies along a single axis.
The axis is a choice: for any direction the operator has the same two eigenvalues (the Pauli matrices have eigenvalues along every axis), so every direction defines spin states of its own, the eigenstates of ; the -axis is singled out here only by the basis we wrote down. At rest the equation is therefore completely solved — four states, two energies, two spins each.
Moving (). For a plane wave , the equation becomes algebraic,
and writing for the upper and lower pairs of Relativistic QM §4 recovers the coupled equations found there,
For (with ) the lower pair is determined by the upper, , so there are again two independent solutions, fixed by the choice of spin state in the upper pair. This is where the spinors enter — the momentum-dependent four-vectors, in the standard normalization, with and :
which reduce to the two positive rest solutions at . For the roles reverse — the upper pair is now determined by the lower, , the minus sign forced by the coupled equations — giving the two negative-energy spinors
with the lower pair large, exactly as Relativistic QM §4 found.
The two branches are not independent: they are each other's charge conjugates. Taking the complex conjugate of the Dirac equation and multiplying by a suitable matrix (in the standard representation ) gives a solution of the same form with opposite charge — the charge-conjugated spinor
and on the plane-wave solutions it maps the particle branch onto the antiparticle branch, up to a phase, with momentum reversed and the two spin states interchanged. If describes a particle of charge , then describes one of charge : the Dirac equation is invariant under charge conjugation. The negative-frequency solutions are therefore not redundant — they are the wave functions of a particle with the same mass and opposite charge, the antiparticle. Together with the Stückelberg–Feynman reading of Relativistic QM §3 — negative-frequency waves propagating backward in time — this is the physics of the positron, predicted by the equation and discovered in 1932, six years later.
What cannot be done at this level: making this precise requires particles to be created and destroyed, which a single-particle wave function cannot describe. The statement that survives at this level is that the Dirac equation has room in its mathematics for both particles and antiparticles — the positive-frequency and negative-frequency parts of its solutions — and that both are physical.
4. Negative Energy Solutions
The two families of §3 — the positive-energy -branch and the negative-energy -branch — pose the problem that Relativistic QM §4 sharpened: both carry positive density , so nothing marks a negative-energy state as unphysical, and the energy is unbounded below — an interacting electron could radiate energy forever, falling through negative levels. Dirac's resolution was the hole theory. The vacuum is not empty; every negative-energy state is occupied — a filled sea of electrons, which the Pauli exclusion principle protects from further occupancy. A missing electron in the sea then behaves as a positive-energy particle of positive charge: a hole, the positron. When a positive-energy electron falls into a hole, both disappear — the energy released is radiated away: pair annihilation; the reverse process, lifting an electron out of the sea into a positive level and leaving a hole behind, is pair creation.
Two honest caveats, in the spirit of the previous pages: the sea picture leans on fermionic statistics (the exclusion principle), and on a many-particle vacuum — both belong to the quantized theory, where the hole picture is replaced by creation and annihilation operators acting on the vacuum. What the single-particle equation contributes is definitive: negative-energy solutions exist, they carry opposite charge (§3), and their interpretation is the doorway to field theory.
5. Spinors (Transformations)
The four-component object transforms differently from anything encountered so far. Under a Lorentz transformation , the spinor transforms as
where are the boost/rotation parameters. Three features set apart from the transformations of vectors and scalars:
- Finite-dimensional. is a matrix — a finite-dimensional representation of the Lorentz group — whereas the familiar transformations on functions (rotations of ) are infinite-dimensional. The spinor representation is a genuinely new structure.
- Double-valued. A rotation by gives : the spinor returns to itself only up to a sign. A rotation is needed to return exactly. No scalar or vector does this; the minus sign is the signature of spin-½.
- Reducible. The representation splits into two pieces, — the two two-component pieces are the Weyl spinors, which transform under rotations identically (the of §1) and under boosts oppositely. They are the left- and right-handed parts of the Dirac spinor.
For rotations alone, — the spin operator of §1 as the generator, confirming from the transformation side that the extra components carry angular momentum .
6. General Solution
The Dirac equation is linear, so the general free solution superposes the four independent solutions per momentum — two spins, two energy signs. In the standard normalization,
with . The coefficients and are complex numbers here — the amplitudes of the particle and antiparticle branches, fixed by the initial conditions. One note, left for later: in the quantized theory these coefficients are promoted to creation and annihilation operators — annihilates an electron, creates a positron — and this expansion becomes the electron field operator. That promotion is the subject of QFT; on this page the coefficients remain numbers. The spinors derived in §3, for a spin direction (, ),
reproduce the structure of Relativistic QM §4: for (positive energy) the upper pair is large at low momentum; for (negative energy) the roles reverse.
The general solution is where the whole page comes together. The -part carries the particles, the -part the antiparticles of §3 and §4; both are dressed with the spin structure of §1, and both transform as the spinors of §5. Why the second coefficient is written conjugated, : under that promotion a number's complex conjugate becomes an operator's adjoint, , the positron creation operator — the notation above is already the shape of the quantized field. That is the bridge to the next stage, QFT.