If classical physics is a “science of calculus,” then quantum physics is, at its core, a “geometry of linear algebra.”
In the models of classical mechanics, the state of a particle can be described by a point in phase space. Quantum theory adopts a different framework: in the finite-dimensional setting that this chapter mainly discusses, a pure state is represented by a unit vector in a complex inner product space, but vectors that differ by a global phase represent the same physical state; observables are represented by Hermitian matrices. The work of Heisenberg, Schrödinger, Dirac, and others developed this theory. For infinite-dimensional operators such as position and momentum, one must also deal with domains and self-adjointness, and not every finite-dimensional conclusion can be carried over directly.
The algebraic theorems of the previous two chapters provide a precise mathematical language for quantum theory. The physical correspondences below rest on the postulates of quantum theory for states, measurement, and evolution; they do not follow from theorems of linear algebra alone:
The eigenvalues of a Hermitian matrix are all real: the possible outcomes of an ideal projective measurement are represented by these real eigenvalues.
Eigenvectors belonging to different eigenvalues are orthogonal: the corresponding orthogonal pure states can be perfectly distinguished by a suitable ideal measurement; arbitrary distinct pure states need not be.
The projection form of the spectral theorem, : the Born rule gives the probabilities ; in an ideal projective measurement, when , the conditional state is . This is the measurement model we adopt, and a measurement need not always change the original state.
Unitary operators (unitary matrices), which preserve inner products and lengths: these correspond to the smooth time evolution of a closed quantum system (the Schrödinger equation) and guarantee that the total probability of the system is always strictly conserved at 1.
When two Hermitian operators do not commute, they have no common complete orthonormal eigenbasis. This does not rule out individual common eigenstates, however, and one cannot conclude from it that no state can have definite values of both. The lower bound in the Robertson relation depends on the chosen state and may be zero; it describes the statistical spread of the state, not instrument error.
For readers interested only in pure mathematics and algorithms, this chapter can be treated as optional; skipping it does not affect the path to the singular value decomposition (SVD) in Chapter 11. But if you are willing to pause here for a while, you will see with your own eyes how eigenvalues, orthogonal projections, and unitary transformations—which looked abstract, perhaps even a little dry, on the blackboard—shed their purely symbolic garb and become a magnificent poem describing the deepest structure of nature.
Chapter Structure and Learning Objectives¶
The introduction mentioned the uncertainty principle and Hermitian matrices, sketching the fascinating central paradox of quantum mechanics: measurement itself takes part in the physical process instead of being a mere passive observer. The task of this chapter is precisely to turn this paradox into exact mathematical statements with the linear algebra tools we sharpened in Chapters 8 and 9.
Before entering the main text, it is worth getting a feel for the line of questions that runs through the chapter. §10.1 starts from the most familiar point—the expectation value and variance of discrete probability—to prepare the language for the later theory of quantum measurement; elementary as they look, these notions will play a key role in the quantum world. §10.2 then asks: what exactly is the “state” of a quantum system? The answer is a normalized vector in a Hilbert space, and what is called “superposition” is simply the expansion of a vector in different bases—a realization that makes many mysterious quantum phenomena seem far more natural. When our knowledge of how a system was prepared is only statistical, however, a single state vector is no longer sufficient, so §10.3 introduces the density matrix : the most elegant physical application of the theory of Hermitian matrices, it describes in a unified way both pure states known with certainty and mixed states that are statistical mixtures, while its off-diagonal entries record whether quantum coherence is present.
With this preparation, §10.4 can give a complete answer to the question “what does a measurement actually do?”: the language of projection operators describes rigorously what happens before and after state collapse, and the Pauli matrices provide a concrete model that turns the abstract theory into examples you can compute by hand. The noncommutativity of the Pauli matrices, , foreshadows a deep fact, which §10.5 proves rigorously with the Robertson-Schrödinger inequality: uncertainty is not a technical problem but a necessary consequence of operator algebra. The most famous conjugate pair, position and momentum, also brings a surprising theorem—the canonical commutation relation cannot hold in any finite-dimensional space; a trace argument gives a one-line proof by contradiction and draws a clear boundary between finite-dimensional and infinite-dimensional Hilbert spaces.
After reading this chapter, your understanding of the “uncertainty principle” will no longer stop at the impression that “the quantum world is strange”; you will be able to see exactly which line of mathematics it comes from—and how that line continues beyond §10.5 to become the starting point of the theory of quantum entanglement in Chapter 11.
10.1 A Review of Basic Probability¶
The finite-dimensional part of this chapter assumes nondegenerate observables, or rank-one ideal projective measurements that can distinguish every basis-vector label; apart from the infinite-dimensional exception of position and momentum, which is discussed separately, everything is treated within this scope.
Before delving into quantum mechanics, we need to review some basic notions of probability. The outcomes of quantum measurements are discrete (quantized), so we focus mainly on discrete probability.
Discrete Random Variables and Probability Distributions¶
A discrete random variable can take finitely many or countably infinitely many values; in this section we first restrict to finitely many values, which ensures that the expectation and variance exist. For example:
The number shown by a die:
The outcome of a quantum measurement: the set of eigenvalues of an observable
The probability mass function (PMF) describes the probability of each possible value of a discrete random variable and satisfies:
Normalization condition.
This ensures that the probabilities of all possible outcomes sum to 1.
Expectation Value and Variance¶
The expectation value, also called the mean, is the “average value” of a random variable, written or :
The expectation value gives the “center” of the values taken by the random variable.
The variance measures how spread out a random variable is, written or :
where is the expectation value. The larger the variance, the more spread out the values of the random variable.
The standard deviation measures the spread in the same units as the original variable.
Independence¶
When we analyze several random variables, we need to consider how they relate to one another.
The joint probability distribution describes the probability that two discrete variables simultaneously take particular values.
Two discrete random variables and are independent if and only if:
This means that knowing the value of does not change the probability distribution of , and vice versa.
Connections to Quantum Mechanics¶
These discrete probability notions appear concretely in quantum mechanics (§10.2-10.5) as follows:
1. Discreteness of measurement outcomes (§10.4):
When an observable is measured, the outcome can only be one of its eigenvalues
One never obtains an arbitrary value between eigenvalues
This is precisely where the word “quantum” (quantum = discretized) comes from
2. The Born rule:
When the pure state is measured with
the probability of obtaining the eigenvalue is
This is the fundamental source of probability in quantum mechanics
3. Expectation value of an observable:
This is exactly the discrete expectation value defined in this section
4. The uncertainty principle (§10.5):
Uncertainty is defined as the standard deviation:
The uncertainty relation:
This uses the variance and expectation value of this section directly
10.2 Fundamentals of Quantum States¶
The mathematical foundation of quantum mechanics rests on the theory of vector spaces, and quantum states are precisely the elements of such a vector space. This section introduces the basic notions of quantum states and lays the groundwork for the later discussion of density matrices and quantum measurement.
Pure States¶
In quantum mechanics, a pure state is the most basic way of representing the state of a quantum system. A pure state is represented by a state vector or a wave function, which completely characterizes the quantum state of the system.
In Dirac notation, a pure state is written (read “ket psi”). Its conjugate transpose is (read “bra psi”). This notation corresponds exactly to the column vectors, and the conjugate transposes of column vectors, that we used in earlier chapters.
A pure state can be written as a linear combination of basis vectors:
where are complex amplitudes and is an orthonormal basis (note that here is exactly the same as the familiar standard basis vector ; only the notation differs).
Superposition States and Eigenstates¶
In quantum mechanics we often need to distinguish superposition states from eigenstates, but this distinction is made relative to a measurement basis.
Dimension of a Quantum System¶
The dimension of a quantum system is the dimension of its state space (Hilbert space); it reflects the complexity and the information capacity of the system.
Physical systems of common dimensions:
| Dimension | Physical system | Typical examples | Remarks |
|---|---|---|---|
| 2 | Qubit | Electron spin, photon polarization | Basic unit of quantum information |
| 3 | Qutrit | Three-level atom, orbital angular momentum of photons | Particles with nuclear spin |
| 4 | Two qubits | Two entangled qubits | |
| 8 | Three qubits | Basic unit of quantum error-correcting codes | |
| qubits | Quantum computer | Dimension grows exponentially | |
| Infinite-dimensional system | Harmonic oscillator, position states of a free particle | Requires tools of functional analysis |
The Probabilistic Nature of Quantum States¶
One of the central features of quantum mechanics is its intrinsically probabilistic nature. Even though a pure state completely describes the system, measurement outcomes are still probabilistic.
The Relationship Between Measurement and Bases¶
The discussion above hints at a deep fact: measurement is closely tied to the choice of basis. Different measurements correspond to different orthogonal bases.
The full theory of measurement, including projective measurement, the mechanism of state collapse, and noncommutativity, is developed systematically in §10.4.
10.3 Mixed States and the Density Matrix¶
In §10.2 we discussed pure states—quantum states that can be completely described by a single state vector. In many practical situations, however, the systems we face cannot be described by a pure state, and we must introduce the notion of a mixed state.
Why Do We Need Mixed States?¶
In real physical systems we often encounter the following situations, in which a description by a pure state no longer applies:
In these situations a single state vector cannot completely describe the state of the system, and we need to introduce the density matrix to handle such a statistical mixture.
Definition of a Mixed State¶
Since pure states cannot describe statistical mixtures, we need a new mathematical tool:
The Density Matrix¶
The density matrix is the general mathematical tool for describing quantum systems (both pure states and mixed states).
Properties of the Density Matrix¶
As a Hermitian operator, the density matrix has all the important properties of Hermitian matrices discussed in Chapter 9 and earlier in this chapter:
Examples of Mixed States¶
Let us understand mixed states and density matrices through concrete examples.
import numpy as np
import matplotlib.pyplot as plt
C_BG = "#F8F8F8"
C_GRID = "#D6D6D6"
C_AXIS = "#000000"
C_V1 = "#57068C"
C_V2 = "#006385"
C_T1 = "#2AD2C9"
# Fonts: the English edition needs no CJK font, so nothing is downloaded when _lang == 'en';
# the Chinese editions use this same block to fetch Noto Sans TC/SC where no CJK font is installed
import os, urllib.request
import matplotlib.font_manager as fm
_lang = 'en'
_cjk = ['Microsoft JhengHei', 'PingFang TC', 'Noto Sans CJK TC', 'Noto Sans TC']
_have = {f.name for f in fm.fontManager.ttflist}
if _lang != 'en' and not _have & set(_cjk):
_font = os.path.join(os.path.expanduser('~'), '.cache', 'fonts', 'NotoSansTC.ttf')
try:
if not os.path.exists(_font):
os.makedirs(os.path.dirname(_font), exist_ok=True)
urllib.request.urlretrieve('https://github.com/google/fonts/raw/main/ofl/notosanstc/NotoSansTC%5Bwght%5D.ttf', _font + '.part')
os.replace(_font + '.part', _font)
fm.fontManager.addfont(_font)
_have.add('Noto Sans TC')
except OSError as err:
print('Could not download the CJK font; Chinese text in figures may not display:', err)
plt.rcParams['font.family'] = [f for f in _cjk if f in _have] + ['DejaVu Sans']
plt.rcParams['axes.unicode_minus'] = False
np.set_printoptions(precision=4, suppress=True, linewidth=100)
def print_header(title):
print("=" * 60)
print(f" {title}")
print("=" * 60)
def print_step(step, desc):
print(f"\n▶ Step {step}: {desc}")
print("-" * 40)
# ============================================================
print_header("Double-Slit Experiment | Probability Density Distribution")
# ============================================================
# --- Step 0: physical parameters ---
print_step(0, "Physical parameters (Fraunhofer far-field approximation)")
lam = 500e-9 # wavelength λ = 500 nm
d = 0.1e-3 # slit separation d = 0.1 mm
L = 1.0 # slit-to-screen distance L = 1 m
a = 0.02e-3 # slit width a = 0.02 mm; (d+a)^2/(λL) << 1
k = 2 * np.pi / lam
# fringe spacing Δx = λL/d, sinc half-width = λL/a
fringe_spacing = lam * L / d
sinc_halfwidth = lam * L / a
print(f" wavelength λ = {lam*1e9:.0f} nm")
print(f" slit separation d = {d*1e3:.1f} mm → fringe spacing Δx = {fringe_spacing*1e3:.3f} mm")
print(f" slit width a = {a*1e3:.2f} mm → sinc half-width = {sinc_halfwidth*1e3:.1f} mm")
print(f" slit-to-screen distance L = {L:.1f} m")
print(f" far-field parameter (d+a)^2/(λL) = {(d+a)**2/(lam*L):.3f}")
# --- Step 1: set up the screen coordinate ---
print_step(1, "Set up the screen coordinate x")
x = np.linspace(-0.05, 0.05, 6000) # ±50 mm, covering the main diffraction envelope
# --- Step 2: consistent common Fraunhofer envelope and path phases ---
print_step(2, "Common sinc envelope and the phase difference between the slits")
theta = np.arctan(x / L)
envelope = np.sinc(a * np.sin(theta) / lam)
phase = k * d * np.sin(theta) / 2
phi0 = envelope * np.exp(-1j * phase)
phi1 = envelope * np.exp(1j * phase)
# --- Step 3: probability densities in the three cases ---
print_step(3, "Compute the probability density P(x) in the three cases")
# density matrix entries (α = β = 1/√2)
rho00 = 0.5
rho11 = 0.5
rho01 = 0.5 # off-diagonal entry (real)
# Case A: coherent superposition (pure state, ρ01 ≠ 0)
# P(x) = ρ00 P0 + ρ11 P1 + 2 Re[ρ01 φ0 φ1*]
P_A = (rho00 * np.abs(phi0)**2
+ rho11 * np.abs(phi1)**2
+ 2 * np.real(rho01 * phi0 * np.conj(phi1)))
# Case B: double slit with observation (mixed state, ρ01 = 0)
# P(x) = ρ00 P0 + ρ11 P1 → incoherent sum of the diffraction envelopes of the two slits
P_B = rho00 * np.abs(phi0)**2 + rho11 * np.abs(phi1)**2
# Case C: only slit 0 open (single-slit diffraction)
P_C = np.abs(phi0)**2
# normalize (area = 1)
P_A /= np.trapezoid(P_A, x)
P_B /= np.trapezoid(P_B, x)
P_C /= np.trapezoid(P_C, x)
print(f" ρ00 = {rho00:.2f}, ρ11 = {rho11:.2f}, ρ01 = {rho01:.2f}")
print(f" Case A (coherent superposition): visibility V = 2|ρ01| = {2*abs(rho01):.2f}")
print(f" Case B (with observation) : ρ01 = 0, V = 0, only the common diffraction envelope remains")
print(f" Case C (single slit) : sinc² envelope, peak at x = 0")
# --- Step 4: plotting ---
print_step(4, "Plot the three probability density distributions")
x_mm = x * 1000 # convert to mm
fig, axes = plt.subplots(3, 1, figsize=(8, 9), sharex=True)
fig.patch.set_facecolor(C_BG)
configs = [
(P_A, C_V1,
r"Case A: coherent superposition — interference fringes appear ($\mathcal{V} = 2|\rho_{01}| = 1$)"),
(P_B, C_V2,
r"Case B: double slit with observation — common envelope, fringes vanish ($\mathcal{V} = 0$)"),
(P_C, C_T1,
r"Case C: single-slit diffraction — $\mathrm{sinc}^2$ envelope (slit 0)"),
]
for ax, (P, color, title) in zip(axes, configs):
ax.set_facecolor(C_BG)
ax.fill_between(x_mm, P, color=color, alpha=0.22)
ax.plot(x_mm, P, color=color, linewidth=1.4)
ax.grid(True, color=C_GRID, linewidth=0.8, alpha=0.7)
ax.set_ylabel("Probability density (normalized)", color=C_AXIS, fontsize=10)
ax.set_title(title, color=color, fontsize=11, fontweight='bold', pad=6)
ax.tick_params(colors=C_AXIS)
ax.set_xlim(x_mm[0], x_mm[-1])
axes[2].set_xlabel(r"Position on the screen $x$ (mm)", color=C_AXIS, fontsize=11)
fig.suptitle(
"Double-slit experiment: probability density and the off-diagonal entry $\\rho_{01}$\n"
r"$P(x)=\rho_{00}P_0(x)+\rho_{11}P_1(x)+2\,\mathrm{Re}[\rho_{01}\,\phi_0(x)\phi_1^*(x)]$",
fontsize=11, color=C_AXIS, y=1.01
)
plt.tight_layout(rect=[0, 0, 1, 0.98])
plt.show()Consider a non-diagonal mixed state:
This density matrix has nonzero off-diagonal entries, which means the system has some degree of coherence. Let us find its spectral decomposition.
Step 1. Compute the eigenvalues
The characteristic polynomial:
Solving gives:
Check: ✓, and ✓
Step 2. Find the eigenvectors
For :
From row 0: →
Normalize: →
Solving gives:
Similarly, for :
Check orthogonality: ✓
Step 3. The spectral decomposition
Physical interpretation:
This mixed state is equivalent to: the system is in the state with probability 78.3% and in the state with probability 21.7%
These two states and are orthogonal and form the eigenbasis of the system
Purity check:
This confirms that is a mixed state.
For a given density matrix :
The spectral decomposition is unique (assuming the eigenvalues are all distinct)
But a general decomposition is not unique
In the nondegenerate case the rank-one eigenprojections are unique, although the eigenvectors can still be multiplied by a phase; in the degenerate case, uniqueness holds only after the projections belonging to the same eigenvalue are combined. This is also why the spectral decomposition is especially important in quantum information theory.
Two Levels of Probability in a Mixed State¶
A subtle point about mixed states is that they contain two levels of probability:
In the mixed state :
First level: classical probability
The probability that the system is in a particular pure state
Comes from incomplete knowledge of how the system was prepared or of its history
This is epistemic uncertainty
Second level: quantum probability
In a particular pure state , the probability that a measurement yields the basis state
Comes from the intrinsic randomness of quantum superposition (the Born rule)
This is ontological uncertainty
In the mixed state , when we measure in the basis , the probability of obtaining is:
Important conclusion: the diagonal entries of the density matrix directly give the probabilities of obtaining when measuring in that basis.
For the detailed theory of measurement, see §10.4.
An ensemble of spin-1/2 systems is in the superposition state with probability 0.6 and in with probability 0.4.
Write the density matrix in matrix form.
Verify that satisfies the three basic properties of a density matrix (Hermiticity, trace one, positive semidefiniteness).
Compute the purity and decide whether this is a pure state or a mixed state.
Find the spectral decomposition of (give the eigenvalues exactly).
Solution to Exercise 1
From and :
Hermiticity: the matrix is real symmetric, so ✓. Trace one: ✓. Positive semidefiniteness: follows from the eigenvalues found in step 4 ✓.
, , so this is a mixed state.
The characteristic polynomial is ():
Check ✓ and ✓. The corresponding orthonormal eigenvectors (numerically) are and , which gives the spectral decomposition.
This section answered the question: when our knowledge of how a quantum system was prepared is incomplete, how do we describe its state and its statistical properties? The answer is the density matrix —a Hermitian, positive semidefinite matrix with trace one that handles pure states and mixed states in a unified way. Its standard form is guaranteed directly by the spectral theorem of §9.3: , where the eigenvalue is the probability weight of each eigenstate, and the purity distinguishes exactly between pure states (equal to 1) and mixed states (strictly less than 1).
Pure states and mixed states are distinguished by purity: for a pure state and for a mixed state; coherence, on the other hand, depends on the chosen basis. In the path basis, the double-slit interference term is , which also depends on the propagation amplitudes and phases. A path detector or certain couplings to the environment can suppress the off-diagonal entries of the reduced state, but this does not mean that the overall correlations necessarily disappear for good.
With the density matrix as a tool, §10.4 can give a complete description of how the state changes before and after a measurement: the act of measurement makes the off-diagonal entries vanish (decoherence), and after the outcome is read out, the system collapses to a definite eigenstate (the state becomes pure).
10.4 Quantum Measurement¶
In §10.3 we developed the complete theory of the density matrix. We now turn to another central topic of quantum mechanics: quantum measurement. Measurement is the bridge between the abstract mathematical formalism and experimental observation, and it is also the most mysterious and most controversial part of quantum mechanics.
The Nature of Quantum Measurement¶
Quantum measurement differs fundamentally from classical measurement:
Classical measurement:
Passive observation that does not disturb the system being measured
The state of the system is the same before and after the measurement
The measurement outcome is completely determined by the state of the system (no randomness)
All physical quantities can be measured simultaneously with unlimited precision
Quantum measurement:
A physical action that may in general disturb the system; eigenstates can remain unchanged
The state is updated conditionally on the outcome (state collapse); when the outcome is ignored, states compatible with the block structure of the measurement projections can remain unchanged
Measurement outcomes have intrinsic randomness (the Born rule)
Incompatible observables have no common complete eigenbasis; the uncertainty within a state is constrained by the corresponding inequality (the uncertainty principle, §10.5)
Projective Measurement¶
The mathematical description of quantum measurement is built on the notion of an observable:
Let the observable be a nondegenerate Hermitian matrix (or let the measurement distinguish every basis-vector label), with spectral decomposition:
where are the eigenvalues and are the corresponding orthogonal eigenstates.
A projective measurement of on the pure state :
Measurement outcome: the result can only be some eigenvalue
Measurement probability: the probability of obtaining is
State collapse: after the measurement, the system collapses to the corresponding eigenstate
Expectation value: the expectation value of the observable is
The outcome of a quantum measurement must be an eigenvalue of the observable; this is a central feature of quantum mechanics:
A measurement cannot yield an arbitrary value between eigenvalues
For example, if the eigenvalues of are , the measurement outcome can only be 0 or 1, never 0.5
This discreteness is where the word “quantum” comes from
Distinguishing the Act of Measurement from Reading Out the Outcome¶
The Schrödinger’s cat thought experiment reminds us to distinguish among the overall state, the reduced state, and the conditional state. We first use two apparatus branches to represent idealized macroscopic records; before the apparatus is coupled, one cannot simply call the superposition of the atom a cat that has already split into two branches.
Once the coupling of the apparatus has established correlations, the overall pure state of the system and the environment can be written as
Ignoring the environment and taking the partial trace over it gives the reduced state of the apparatus $$\boldsymbol{\rho}_{\mathrm{red}}=
When the environment records are nearly orthogonal, the off-diagonal entries are suppressed and the reduced state is approximately a diagonal mixed state. This is the decoherence approximation; the overall state can still be pure, and the information in the correlations has not disappeared permanently in the mathematical sense.
After an outcome of positive probability is read out, conditioning according to this chapter’s postulate of ideal projective measurement gives the conditional state of the corresponding branch. Decoherence explains why locally visible interference is suppressed, but it does not by itself derive, from unitary evolution alone, a unique measurement outcome in each run of the experiment.
| Level of description | State and information |
|---|---|
| System and environment as a whole | Can remain a pure state containing the correlations between branches |
| Reduced state, ignoring the environment | Approximately a diagonal mixed state when the environment records are nearly orthogonal |
| Conditional state after reading out the outcome | Updated according to the measurement postulate given the known outcome; differs from the averaged state when the outcome is not read out |
To give each row of the table above a precise mathematical counterpart, we need to establish three things in order: the trace formula provides a unified tool for computing expectation values, projective measurement of mixed states gives a precise expression for the probabilities, and only then come the formal definitions of the conditional state and the averaged state.
The Trace Formula for Mixed States¶
The trace formula shows that is the classical weighted average of the expectation values in the individual pure states; this is the central computational power of the density matrix as a statistical tool. In particular, taking gives a unified formula for measurement probabilities.
Projective Measurement of Mixed States¶
A note on notation: is a scalar probability; is a matrix operator. This book uses (a bold Greek letter) exclusively for projection operators and (written as a function with parentheses) exclusively for probabilities.
With the probability formula in hand, we can define precisely the two kinds of post-measurement states.
The Conditional State and the Averaged State¶
Perform a projective measurement of the observable on the state , where the projection operators are :
Conditional state (the measurement outcome is known to be , and ):
This is a pure state; it describes the state of the system “given that the outcome was obtained.”
Averaged state (the outcome is not read out, or we average over all possible outcomes):
where is the probability of obtaining the outcome . This is usually a mixed state (unless the initial state is already an eigenstate), but it is necessarily diagonal (its off-diagonal entries in the measurement basis are 0).
Summary of the key differences:
| Conditional state | Averaged state | |
|---|---|---|
| Premise | The measurement outcome is known | The outcome is unknown or ignored |
| Type of state | Pure state | Usually a mixed state |
| Purity | ||
| Off-diagonal entries | 0 (projection onto a pure state) | 0 (decoherence) |
| Diagonal entries | A single nonzero entry (= 1) | Several nonzero entries (a probability distribution) |
| Physical meaning | A single run with a known outcome | Statistics over many runs, or a single run with an unknown outcome |
| Mathematical form |
Consider a measurement on the pure state .
Before the measurement:
Purity: (a pure state)
Off-diagonal entries are present (quantum coherence)
After the measurement (conditional states):
If +1 is obtained (probability 50%):
Purity: ✓ (a pure state)
If -1 is obtained (probability 50%):
Purity: ✓ (a pure state)
Both conditional states are pure states.
After the measurement (averaged state):
Purity: (a mixed state)
No off-diagonal entries (decoherence)
This is the maximally mixed state
Comparison of the evolutions:
The act of measurement removes the off-diagonal entries (decoherence), but if the outcome is not read out, the system is still in a mixed state.
Conditional state: describes the state of the system in a single run after the measurement outcome is known
Example: “After spin up is obtained, the system is in ”
Experimental situation: the display of the measuring instrument has been read
Averaged state: describes the statistical ensemble over many runs, or a single run whose outcome is not read out
Example: “Measuring on a large number of identical systems, 50% are in and 50% are in ”
Experimental situation: statistics over many repeated runs, or discarding the instrument reading after the measurement
Convention of this book: unless stated otherwise, “the post-measurement state” means the conditional state (the outcome has been read out). When we discuss the averaged state, we will label it explicitly as the “averaged state” or say that the “outcome is not read out.”
Pauli Matrices: The Mathematical Framework of Spin Measurement¶
With the complete conceptual framework of measurement in place, we now introduce the most important concrete example in quantum mechanics: the spin-1/2 system.
Physical background:
Spin is the intrinsic angular momentum of a quantum particle, distinct from orbital angular momentum
For a spin-1/2 particle, a measurement of spin along any direction has only two possible outcomes: (“up”) or (“down”)
The spin state is completely described by the two-dimensional complex vector space
Commutation relations and a preview of the uncertainty principle:
Noncommutativity of the Pauli matrices (see §10.5).
The Pauli matrices do not commute with one another; they satisfy:
Check (taking the first as an example):
✓
This noncommutativity is the mathematical root of the uncertainty principle: spin components along different directions have no common complete eigenbasis (for spin-1/2 they do not even have a common eigenstate), so they cannot have definite values simultaneously in the same state. In §10.5 we will see that the lower bound on the product of the standard deviations is determined by the expectation value of the commutator and a covariance term in that state, and it may be zero.
The commutation relations of the Pauli matrices can be written compactly as:
where is the Levi-Civita symbol (+1 for cyclic permutations of , -1 for cyclic permutations of , and 0 when an index is repeated). This relation has the same form as the Lie algebra structure of quantum orbital angular momentum, revealing the mathematical nature of spin as an intrinsic angular momentum.
Spin Measurement and State Collapse¶
With an explicit definition of the Pauli matrices and a clear conceptual framework for measurement, we can now analyze the process of spin measurement systematically.
Let the pure state be (expanded in the eigenbasis of ).
Case 1: is an eigenstate of
If for some
The measurement outcome is certain:
The state is unchanged after the measurement (conditional state):
Case 2: is a superposition state (several )
The measurement outcome is random, with probabilities
The state changes after the measurement (conditional state):
The phase information of the superposition is permanently lost
Consider the pure spin state (the “up” state along ):
Measuring spin along ():
The spin operator:
Eigenvalues and eigenstates:
, eigenstate
, eigenstate
Since is itself an eigenstate of :
Measurement outcome: with probability 100%
State after the measurement (conditional state): (unchanged)
Measuring spin along ():
First expand in the basis. Since:
solving back gives:
Check:
✓
Measurement outcomes:
With probability 50%, is obtained → the state collapses to the conditional state
With probability 50%, is obtained → the state collapses to the conditional state
Key observation: the state has changed fundamentally!
Averaged state (if the outcome is not read out):
This is a mixed state (), but the conditional states are all pure states.
Computing expectation values:
Measuring in the basis:
Measuring in the basis:
This agrees with the probabilistic prediction: 50% × + 50% × = 0 ✓
Basis Dependence of Measurement and Noncommutativity¶
The same pure state, measured in different bases, collapses to different conditional states:
Measuring in the basis (the eigenstates of the observable ) → collapses to some
Measuring in the basis (the eigenstates of the observable ) → collapses to some
If two Hermitian observables have no common orthonormal eigenbasis (equivalently, they do not commute), order effects can arise in general. Below we distinguish the conditional states with read-out outcomes, the joint distribution of outcomes, and the averaged state without read-out:
(in general)
The order of measurements matters: measuring first and then , versus measuring first and then , may give different conditional states or joint distributions of outcomes; in particular examples, the averaged states may also turn out the same
For the Pauli matrices, since , spin measurements along and along are incompatible. This is precisely the root of the uncertainty principle (§10.5).
Choose a pure state that is neither an eigenstate of nor an eigenstate of :
Normalization check: ✓
Properties of the initial state:
Expectation values: ,
This state “leans down” along and “leans up” along
Order A: measure first, then (reading out both outcomes)
Step 1. Measure and read out the outcome
Probabilities:
→ conditional state
→ conditional state
Averaged state after step 1 (taking all possible outcomes into account):
Key observations:
Purity: (a mixed state)
(the information along is preserved)
(the information along is washed out)
Step 2. Measure and read out the outcome
Expand: ,
(a) If step 1 gave (probability 20%):
+1 with probability 50% → conditional state
-1 with probability 50% → conditional state
(b) If step 1 gave (probability 80%):
+1 with probability 50% → conditional state
-1 with probability 50% → conditional state
Total probabilities:
Final averaged state:
This is the maximally mixed state.
Order B: measure first, then (reading out both outcomes)
Step 1. Measure and read out the outcome
Expand in the basis:
Probabilities:
→ conditional state
→ conditional state
Averaged state after step 1:
Key observations:
Purity: (a mixed state)
(the information along is preserved)
(the information along is washed out)
Step 2. Measure and read out the outcome
A similar computation shows that the final averaged state is again:
The key manifestation of noncommutativity:
| Initial state | After step 1 (averaged state) | Final (averaged state) | |
|---|---|---|---|
| Order A | , | , | |
| Order B | , | , |
Core lessons:
Measurement preserves information selectively: the first measurement preserves the information along its own direction and washes out the other directions
The second measurement completely washes out the remaining information: in this example, for mutually perpendicular spin directions, averaging over the outcomes of both measurements gives the maximally mixed state
Different intermediate states, the same final state:
vs.
Diagonal vs. non-diagonal
Physical meaning of noncommutativity: ⇒ measuring one operator destroys the information about the other
For a spin-1/2 system, if ideal projective measurements are performed along two mutually perpendicular spin directions and we average over the outcomes of both measurements, the final state is . This conclusion is not guaranteed for arbitrary noncommuting directions; and the conditional states obtained by reading out each outcome also differ from this averaged state.
Projective Measurement (Mixed States)¶
Recall Definition 6 and Definition 7: when a projective measurement of the observable (with projection operators ) is performed on a mixed state , the probability of obtaining is ; if the outcome is read out, the conditional state is necessarily a pure state. We now apply this general theory to concrete mixed states described with the Pauli matrices.
Whatever mixed state the system starts in, if this chapter’s rank-one projective measurement is performed and an outcome of positive probability is read out, the conditional state is the pure state .
Physical explanation: reading out the outcome removes the classical uncertainty
Before the measurement: we do not know which pure state the system is in (probability )
After reading out the outcome: we know for certain that the system is in the pure state
The quantum uncertainty is converted into the selection of a measurement outcome (which ), but the conditional state after the outcome is read out is itself a definite pure state.
Contrast with the averaged state: if the outcome is not read out, the system is usually in a mixed state (exceptions such as eigenstates can remain pure) (with diagonal entries distributed according to the probabilities), but the off-diagonal entries have been removed by the act of measurement.
Consider a more general mixed spin state (with off-diagonal entries):
Step 1. Verify that this is a valid density matrix
(a) Hermiticity: ✓
(b) Trace one: ✓
(c) Positive semidefiniteness: the eigenvalues are and (both ) ✓
(d) Verify that this is a mixed state:
✓
Step 2. Physical interpretation
The nonzero off-diagonal entries indicate that the system has quantum coherence.
Measuring spin along () and reading out the outcome:
The measurement basis is .
Probability of obtaining “up” (eigenvalue +1, spin ):
Conditional state (outcome read out as “up”):
Check the purity: ✓ (a pure state)
Probability of obtaining “down” (eigenvalue -1, spin ):
Conditional state (outcome read out as “down”):
Check the purity: ✓ (a pure state)
Averaged state (outcome not read out):
Check the purity: (a mixed state, but without coherence)
Comparison:
Before the measurement: (coherence present)
Averaged state: (no coherence, still mixed)
Conditional state: or (a pure state)
Key observations:
The act of measurement: removes the off-diagonal entries (decoherence)
Reading out the outcome: further purifies the state to a single eigenstate (the conditional state)
Expectation value:
Check: ✓
Measuring spin along () and reading out the outcome:
The measurement basis is .
A similar computation (omitted) gives:
,
The conditional state is or (a pure state)
The averaged state is (a mixed state, diagonal in the basis)
A coincidental symmetry: this particular gives the same probability distributions and expectation values along both and .
Consider the spin-1/2 mixed state , and perform a projective measurement of in the basis .
Verify that is a valid density matrix and compute its purity.
Find the probabilities and of obtaining “up” and “down,” and the conditional state after each outcome is read out.
Write the averaged state when the outcome is not read out and compute its purity; explain what the act of measurement does to the off-diagonal entries.
Compute the expectation value .
Solution to Exercise 2
Hermitian ✓; trace ✓; the eigenvalues are , so it is positive semidefinite ✓. The purity is , a mixed state.
, . After the outcome is read out, the conditional state is necessarily pure: if “up” is obtained → ; if “down” is obtained → .
. The purity is , so it is still a mixed state, but the off-diagonal entries go from —the act of measurement has removed the coherence (decoherence).
.
Summary of the Effects of Measurement on States¶
| Before the measurement | Measured, outcome read out | Measured, outcome not read out | Change in purity (conditional state / averaged state) |
|---|---|---|---|
| Mixed state | Pure state (conditional state) | Mixed state (averaged state, no coherence) | rises to 1 / may stay the same |
| Pure state (eigenstate) | The same pure state | The same pure state | Unchanged (= 1) / unchanged (= 1) |
| Pure state (superposition) | A different pure state (conditional state) | Mixed state (averaged state) | Unchanged (= 1) / becomes |
Unifying principles:
Reading out the outcome: the conditional state is always pure, which removes the classical uncertainty
Not reading out the outcome: the averaged state may be mixed, but the off-diagonal entries are removed
The coherence of the system is erased: the off-diagonal entries of both the conditional state and the averaged state in the measurement basis are 0; if the apparatus and the environment are included in an overall unitary description, the information in the correlations can be retained in the overall state (see the note below) and does not disappear permanently
The measurement basis determines the direction of collapse: noncommuting operators give different collapse outcomes
In a given basis, a suitable Hamiltonian and initial state can change the coherence; not every unitary evolution produces a superposition. Certain system–environment couplings establish distinguishable environment records, and after the partial trace over the environment is taken, the coherence of the reduced state of the system is suppressed as a result. The coupling itself does not guarantee that the environment states are orthogonal, nor does it guarantee a monotone or permanent decay.
The Lindblad equation and exponential decay are approximate models of open systems that adopt suitable assumptions, such as the Markov assumption. Overall unitary evolution can retain the information in the correlations; local decoherence and the disappearance of information from the whole should be understood separately.
The central question of this section is: what exactly does a quantum measurement do to a system, and how can the change of state before and after it be described precisely? Projective measurement gives the probability of each outcome through the Born rule , and the act of measurement itself has two levels: when the measurement is made but the outcome is not read out, the state evolves into the averaged state, whose off-diagonal entries vanish (decoherence); after the outcome is read out, the state collapses to the conditional pure state of the corresponding eigenstate (the state becomes pure). This distinction between “act” and “read-out” is a key concept of quantum information theory, and it marks the fundamental difference between quantum uncertainty and classical uncertainty.
The Pauli matrices , as the three observables of the spin-1/2 system, provide a concrete model that can be computed completely, and their noncommutation relations are not merely algebraic identities but the direct reason why spin measurements along different directions interfere with one another. From the point of view of linear algebra, the spectral decomposition is used to construct the measurement projections; a measurement channel that ignores the outcome erases the coherence between branches, which is a change of state, unlike the unitary changes of coordinates in Chapters 8 and 9, which keep the same operator. The consequences of this noncommutativity are raised in §10.5 to the rigorous inequalities of the uncertainty principle, revealing the deepest mathematical root of quantum mechanics.
10.5 The Uncertainty Principle¶
The uncertainty principle is one of the fundamental principles of quantum mechanics. It describes the limits on the intrinsic standard deviations of the distributions of measurement outcomes on identically prepared states; it is not a relation about instrument error or measurement disturbance. This is not a problem of measurement technique; it stems from the noncommutativity of operators—a central feature of quantum mechanics.
Basic Definitions: The Commutator and the Anticommutator¶
The commutator and anticommutator are important notions for describing the noncommutativity between two operators and :
Commutator:
When , we say that and commute
Anticommutator:
When , we say that and anticommute
⇔ the two observables can be measured simultaneously with arbitrary precision (they have a common complete eigenbasis)
⇔ the two observables are incompatible, and an uncertainty relation holds
The size of the commutator quantifies the “degree of incompatibility” of the two operators
The rigorous mathematical basis for “can be measured simultaneously with arbitrary precision” is given by Theorem 1 below: two Hermitian matrices commute if and only if they can be diagonalized simultaneously—that is, there is a common orthonormal eigenbasis on whose quantum states both physical quantities have definite values at the same time. Conversely, if they do not commute, no such complete basis exists.
Let both be Hermitian matrices. Then if and only if there is a unitary matrix such that and are both diagonal matrices—that is, and have a common orthonormal eigenbasis.
(Sufficiency) If and are both diagonal matrices, then since diagonal matrices commute with one another, .
(Necessity) Suppose . By the spectral theorem (§9.3), the eigenspaces corresponding to the distinct eigenvalues of are pairwise orthogonal, and their dimensions add up to .
leaves each eigenspace invariant: if , then , so .
The restriction of to each is still Hermitian (for any , still holds), so by the spectral theorem again, each has an orthonormal basis consisting of eigenvectors of .
These basis vectors are at the same time eigenvectors of (with eigenvalue ). Taking the union of the bases of the subspaces gives an orthonormal basis of in which every vector is an eigenvector of both and ; the unitary matrix whose columns are these vectors is the required matrix.
Algebraic Properties of the Commutator¶
The expectation value of the commutator, , has several important properties in quantum mechanics, and these properties are essential for understanding the uncertainty principle and quantum dynamics:
Anti-Hermiticity: for Hermitian operators and , the commutator is anti-Hermitian, that is:
Hence its expectation value is necessarily purely imaginary:
Decomposition of operator products: the expectation value of the product of any two operators can be split into a commutator part and an anticommutator part:
Since the anticommutator is Hermitian, its expectation value is real, while the expectation value of the commutator is purely imaginary. This decomposition splits the operator product into a real part and an imaginary part.
Behavior in eigenstates: if is an eigenstate of , that is, ( real), then:
This means that in an eigenstate of , “effectively commutes” with every operator (from the point of view of expectation values).
The Robertson-Schrödinger Inequality¶
The following theorem gives a lower bound for the product of the uncertainties of two observables; it is the mathematical statement of the uncertainty principle:
For any two Hermitian operators and and any quantum state ,
where:
is the standard deviation
is the deviation operator
The second term is , which makes this inequality stronger than the classical Robertson inequality (which contains only the first term)
Define the deviation operators
Define the deviation of each operator from its expectation value:
Clearly and .
Construct auxiliary state vectors
Define the state vectors:
Apply the Cauchy-Schwarz inequality
For any two vectors in an inner product space, the Cauchy-Schwarz inequality gives:
Expanding the left-hand side:
(Note that is Hermitian, so .)
Expanding the right-hand side:
Therefore:
Decompose the operator product
Note that can be split into a commutator part and an anticommutator part:
Key observation: the commutator is unaffected by constant shifts:
Use Hermiticity
For the expectation values:
The commutator is anti-Hermitian ⇒ is purely imaginary
The anticommutator is Hermitian ⇒ is real
These two terms are orthogonal to each other (one lies on the real axis and the other on the imaginary axis), so:
Combining this with the result of step 3 gives:
The theorem above holds for pure states . For a mixed state , the uncertainty relation can be proved by convexity:
If every pure-state component satisfies , then for the mixed state we have at least:
In fact, Jensen’s inequality (together with the Cauchy-Schwarz inequality) gives a stronger result: mixing increases uncertainty, (and likewise for ), so ; that is, the uncertainty product of a mixed state is no smaller than the weighted average of the lower bounds of its pure-state components, and this weighted average is itself no smaller than .
The Spin Uncertainty Relation (a Concrete Example)¶
Before discussing the abstract position-momentum uncertainty relation, we first look at a concrete, computable example: the uncertainty relation for a spin-1/2 system. This example is based on the Pauli matrices introduced in §10.4 and shows the core mechanism of the uncertainty principle.
Consider the uncertainty relation between and . Let the quantum state be with .
Step 1. Compute the expectation value of the commutator
From §10.4 we know that , so:
Step 2. Apply the Robertson-Schrödinger inequality
Substituting the expectation value of the commutator into Theorem 2:
This gives the uncertainty relation for spin measurements:
Similarly, cyclic symmetry gives:
Step 3. Concrete examples
(a) An eigenstate of :
This is the spin-up state along (§10.4):
Compute the expectation values:
Hence (completely certain).
Compute :
Hence (completely uncertain).
Compute :
Check the inequality:
✓
Physical interpretation: when we are completely certain that the particle’s spin is up along (), we know nothing about its spin along (). The inequality holds with equality because the right-hand side is zero.
(b) An eigenstate of :
This is the spin-up state along :
Compute :
Compute :
Compute :
Check the inequality:
✓
Equality holds again.
Physical interpretation: when the particle’s spin along is certain, the spins along and both become completely uncertain, and the product of their uncertainties attains the lower bound .
(c) A tilted state (a nontrivial general case):
This state is tilted by on the Bloch sphere (, so the coefficients use ); it lies neither on the axis nor in the plane, and it shows perfectly how the uncertainty principle constrains a general case.
Why choose this state?
The earlier eigenstate examples showed the boundary cases (0 or 1), but they were a bit like computing “” or “”
This tilted state makes and the lower bound take the intermediate value , while still attains its maximum value 1
It shows how uncertainty “slides” and is “distributed” in a general state
Compute along (which determines the lower bound):
Using the trigonometric identity :
This means that the lower bound on the right-hand side of the inequality is no longer 0 or 1 but an intermediate value .
Compute along :
Compute along :
(Because the coefficients are all real, the expectation value along is 0.)
Check the inequality:
Left-hand side (product of the measurement fluctuations):
Right-hand side (theoretical lower bound):
Conclusion:
✓
A closer look: why is this example more representative?
A nonzero lower bound: because the particle leans toward the axis (), the measurement outcomes for and are forced to maintain a certain degree of “mutual exclusivity.” In this state the component is partly uncertain, while the component is still maximally uncertain, with equal probabilities.
The mechanism by which uncertainty is distributed:
Because the state lies in the plane, we have partial knowledge of (), unlike the complete ignorance in the eigenstate ()
As the price, since is completely orthogonal to this plane, our observation of remains maximally uncertain ()
This shows the “distribution of uncertainty” in quantum mechanics: gaining partial information in one direction must be paid for with uncertainty in other directions
Minimum-uncertainty states (intelligent states): this particular tilted state still makes the inequality hold with equality, which shows that it is the “optimal solution” under this lower-bound constraint—for the given , this state attains the theoretical minimum of .
Comparing the three examples:
| Example | Lower bound | Feature | ||||
|---|---|---|---|---|---|---|
| (a) eigenstate | 0 | 0 | 1 | 0 | 0 | Boundary case: completely certain vs. completely uncertain |
| (b) eigenstate | 1 | 1 | 1 | 1 | 1 | Boundary case: maximal uncertainty attained |
| (c) Tilted state | 1 | General case: shows the dynamic constraint |
Example (c) fills the gap between 0 and 1 and shows how the uncertainty principle constrains uncertainty in non-extreme states.
Key conclusions:
The commutator determines the lower bound: ⇒ the lower bound on the uncertainty product is
The lower bound depends on the state: different have different
The three examples give the complete picture:
Example (a): a eigenstate → the lower bound is 0 (one direction completely certain)
Example (c): a tilted state → the lower bound is (the dynamic constraint of the general case)
Example (b): a eigenstate → the lower bound is 1 (maximal uncertainty attained)
Going from , these three examples show completely how uncertainty “slides” as the state changes
The special place of the tilted state:
It is neither “all” nor “nothing,” and it is closer to realistic physical situations
It shows the mechanism by which uncertainty is distributed: gaining partial information in one direction must be paid for with uncertainty in other directions
It is a minimum-uncertainty state (equality holds), optimal under this constraint
Computability in finite dimensions: the Pauli matrices are matrices, all computations are elementary matrix operations, and students can verify every step directly
Physical intuition: spin components along different directions cannot be definite at the same time; here the eigenstates of spin correspond to the standard axes of the Bloch sphere, and clearly no point can be close to two of these axes at once. The tilted state lets us see how this “incompatibility” shows up in non-extreme situations.
This example shows the core mechanism of the uncertainty principle: through the Robertson-Schrödinger inequality, the noncommutativity of operators leads directly to lower bounds on the uncertainties of physical quantities.
Next we generalize this idea to the infinite-dimensional position-momentum system, where we will see a similar structure, except that the lower bound becomes the universal constant .
The Heisenberg Uncertainty Relation¶
The most famous example of the uncertainty principle is the position-momentum uncertainty relation. Having shown the finite-dimensional example of the Pauli matrices, we now turn to continuous systems, which require infinite-dimensional spaces.
In quantum mechanics, the uncertainty relation between position and momentum can be described in two equivalent mathematical languages: wave mechanics (1926) and matrix mechanics (1925).
The Matrix Mechanics Formulation¶
In an abstract Hilbert space , a quantum state is described by a state vector .
The position operator : a self-adjoint operator (Hermitian operator) acting on , satisfying .
The momentum operator : a self-adjoint operator acting on , satisfying .
The canonical commutation relation: these two operators satisfy the fundamental commutation relation
where J·s is the reduced Planck constant.
The infinite-dimensional representation of matrix mechanics, restated in modern language:
In the language of modern quantum mechanics, the harmonic oscillator basis is used to represent position and momentum as infinite-dimensional matrices. In a certain orthonormal basis, and can be represented as:
where is the mass of the particle and is the angular frequency of the harmonic oscillator. The matrix entries are:
These two infinite-dimensional matrices satisfy the canonical commutation relation .
Key observations:
The matrices are tridiagonal (nonzero entries appear only on the diagonal and next to it)
The matrices are unbounded (the entries grow with : )
The matrices must be infinite-dimensional (we prove below that the canonical commutation relation cannot be realized in finite dimensions)
In matrix form the commutation relation reads:
where is the infinite-dimensional identity matrix.
The Wave Mechanics Formulation¶
In the position representation, the Hilbert space is realized as , and a quantum state is described by the wave function .
The position operator (in the position representation):
That is, the position operator multiplies the wave function by the position coordinate .
The momentum operator (in the position representation):
The momentum operator corresponds to differentiation.
Verifying the commutation relation (position representation):
Below we take a normalized Schwartz function , so that differentiation, multiplication, Fourier integrals, and second moments all make sense; the CCR holds on this common dense domain, and the substitution into the uncertainty relation that follows stays within this scope. Compute the commutator:
This confirms that .
The Connection Between the Two Formulations¶
The position representation and the momentum representation are related by the Fourier transform. Let be the wave function in the momentum representation; then:
In the momentum representation:
The momentum operator:
The position operator:
The two representations give the same commutation relation. The infinite-dimensional matrix representation written above in modern language is the representation in the basis of energy eigenstates of the harmonic oscillator, and it is related to the position and momentum representations by a change of basis.
Deriving the Uncertainty Relation¶
Substituting into the Robertson-Schrödinger inequality (Theorem 2), note the expectation value of the commutator:
Therefore:
This is the famous Heisenberg uncertainty principle:
Comparison with the Pauli matrix example:
Pauli matrices: → lower bound (depends on the state)
Position-momentum: → lower bound (a universal constant, independent of the state)
The key difference: a commutator equal to an operator vs. a commutator equal to a constant times the identity operator!
Physical Meaning¶
A fundamental limit, not a technical problem: this is not a limit on the precision of measuring instruments but a fundamental feature of quantum mechanics. means that there is no quantum state that is simultaneously an eigenstate of and of (the two operators have no common complete system of eigenvectors).
The complementarity principle: the more concentrated the position distribution of the prepared state ( smaller), the more uncertain its momentum ( larger), and vice versa. Position and momentum are complementary observables.
Interpretation in matrix mechanics: the noncommutativity means that the infinite-dimensional matrices and cannot be diagonalized simultaneously; this is the core of the uncertainty principle in matrix mechanics.
Interpretation in wave mechanics: in the position representation, a particle appears as a “wave packet.” The narrower the wave packet (definite position), the broader its Fourier spectrum (uncertain momentum); this corresponds to the time-frequency uncertainty principle of signal processing.
The importance of the quantum scale: because is extremely small, at macroscopic scales ( m) the bound kg·m/s is negligible. At atomic scales ( m), however, the momentum uncertainty becomes significant; for example, an electron bound in an atom necessarily has momentum fluctuations of about kg·m/s.
Historical Note¶
Matrix mechanics in 1925 and wave mechanics in 1926 provided two formulations of quantum theory. The tridiagonal matrices of the harmonic oscillator here use modern notation; they are a concrete link between the matrix and wave-function representations.
Why Can the Position-Momentum Uncertainty Relation Not Be Proved in Finite Dimensions?¶
We have seen that the Pauli matrices (§10.4 and the spin example above) display the mathematical structure of the uncertainty principle perfectly. However, the position-momentum uncertainty relation cannot be realized in a finite-dimensional Hilbert space. This is the essential difference between continuous and discrete degrees of freedom in quantum mechanics.
Let and be linear operators acting on a finite-dimensional Hilbert space . Then they cannot satisfy the canonical commutation relation:
Proof (trace argument):
Suppose there are matrices and satisfying the relation above. Take the trace of both sides:
By the cyclic property of the trace (see Exercise Exercise 12 in the Chapter 3 exercises; it can also be verified directly by interchanging the order of summation in ), we have:
Hence the left-hand side is:
But the right-hand side is:
This is a contradiction! Therefore the canonical commutation relation can hold only in an infinite-dimensional Hilbert space.
Key comparison:
| Property | Pauli matrices (spin) | Position and momentum operators |
|---|---|---|
| Dimension of the Hilbert space | Finite () | Infinite () |
| Nature of the physical quantities | Discrete spectrum () | Continuous spectrum () |
| Form of the commutator | ||
| Type of the commutator | A bounded operator (diagonalizable) | On a common domain it equals , which extends to a bounded operator; themselves are unbounded |
| Trace of the commutator | Not trace class; the usual trace does not apply | |
| Applicable systems | Spin- (e.g., the electron) | Particles in continuous space |
| Matrix representation | Explicit matrices | Infinite-dimensional tridiagonal matrices |
| Lower bound of uncertainty | $ | \langle\sigma_z\rangle |
Deeper physical reasons:
Pauli matrices: spin is a discrete intrinsic degree of freedom. Whatever the direction of measurement, a spin- particle has only two possible values (). This discreteness makes the state space naturally finite-dimensional (), and the commutator is itself a bounded matrix with trace zero.
Position and momentum operators: the position and momentum of a particle are both continuous variables taking values in . Describing such a system requires the infinite-dimensional function space . The right-hand side of the commutator is the identity operator, which in infinite dimensions is bounded but not trace class, so the usual finite trace cannot be used.
Mathematical essence: a commutator equal to a constant times the identity operator (, ) is a property reserved for infinite-dimensional Hilbert spaces; every finite-dimensional representation fails because of the trace contradiction. This is precisely a corollary of the cyclic property of the trace: the trace is invariant under cyclic permutations (Exercise 12), so in finite dimensions (or under suitable trace-class conditions) the trace of a commutator, , must be zero.
Conclusion:
The Pauli matrices give us a finite-dimensional analogue of the uncertainty principle and show clearly how the noncommutativity of operators leads to the uncertainty of physical quantities. This is a concrete example that students can master completely and compute by hand.
Because position and momentum are continuous variables, however, the true position-momentum uncertainty relation
must be proved in the infinite-dimensional Hilbert space (or with the infinite-dimensional matrix representation of matrix mechanics) and cannot be reduced to a finite-dimensional matrix model. This reflects the essential difference between continuous and discrete degrees of freedom in quantum mechanics, and it is also why we need functional analysis and the theory of infinite-dimensional spaces to describe quantum mechanics completely.
Besides position-momentum and spin components, quantum mechanics has many other uncertainty relations:
Energy-time uncertainty:
For example, for an observable with no explicit time dependence, define the Mandelstam–Tamm characteristic time (with nonzero denominator); the Robertson inequality then gives this bound. This is not the lifetime of every quantum state
Deriving this relation requires more care, because time is not an operator in quantum mechanics
Angular momentum components:
The three components of angular momentum do not commute with one another:
Similar to the Pauli matrices, but corresponding to spin-1 or higher-spin systems
Can be realized in finite-dimensional spaces (dimension = )
Number-phase uncertainty: in quantum optics, photon number and phase are conjugate variables
A laser (definite phase) has an uncertain photon number
A Fock state (definite photon number) has a completely uncertain phase
All these uncertainty relations stem from the noncommutativity of the corresponding operators and embody a universal feature of quantum mechanics.
Part 1: consider the state .
Compute , , and .
Compute and , and verify the uncertainty relation for the pair .
Part 2:
Use the cyclic property of the trace (Exercise 12) to prove that for any two matrices we always have , and explain how this rules out the canonical commutation relation in finite-dimensional spaces.
Solution to Exercise 3
Since (the coefficients contain the imaginary unit ):
, so and .
✓ (equality holds; a minimum-uncertainty state).
From :
But . If held in finite dimensions, taking the trace of both sides would give , a contradiction. Hence the canonical commutation relation can hold only in an infinite-dimensional Hilbert space.
This section answered one of the deepest questions of quantum mechanics: why can certain pairs of physical quantities not be measured simultaneously with arbitrary precision—is this a technical limitation or a fundamental property of nature? With the Cauchy-Schwarz inequality at its core, the Robertson-Schrödinger inequality proves rigorously that for any Hermitian operators , the lower bound on the uncertainty product is determined by the expectation value of the commutator and the covariance term in that state, and it may be zero; noncommutativity rules out a common complete eigenbasis. The Pauli matrices give a complete finite-dimensional demonstration— can be verified directly by hand. The position-momentum relation reveals a deeper boundary: the trace argument shows rigorously that the canonical commutation relation cannot hold in any finite-dimensional space, which forces us to use infinite-dimensional Hilbert spaces.
This boundary between finite and infinite dimensions marks the essential difference in mathematical structure between discrete degrees of freedom (such as spin) and continuous ones (such as motion in space). The uncertainty principle is not a regret that measuring instruments are not precise enough; it is a deep imprint that the noncommutative algebra of operators leaves on the laws of nature. The main thread of Chapter 10—from the language of probability, through quantum states and the theory of measurement, to the mathematical roots of uncertainty—is extended further in the next chapter by the Schmidt decomposition, which quantifies the even stranger phenomenon of quantum entanglement.
import numpy as np
np.set_printoptions(precision=4, suppress=True, linewidth=100)
def print_header(title):
print("=" * 60)
print(f" {title}")
print("=" * 60)
# Pauli matrices
sx = np.array([[0, 1], [1, 0]], dtype=complex)
sy = np.array([[0, -1j], [1j, 0]], dtype=complex)
sz = np.array([[1, 0], [0, -1]], dtype=complex)
def expect(A, psi):
"""Expectation value ⟨ψ|A|ψ⟩ (necessarily real for a Hermitian operator)"""
return np.real(psi.conj() @ A @ psi)
def delta(A, psi):
"""Standard deviation ΔA = sqrt(⟨A²⟩ − ⟨A⟩²)"""
return np.sqrt(expect(A @ A, psi) - expect(A, psi) ** 2)
# ============================================================
print_header("Verifying the Robertson-Schrödinger Inequality: Δσx · Δσy ≥ |⟨σz⟩|")
# ============================================================
states = {
"(a) σx eigenstate |↑x⟩": np.array([1, 1], dtype=complex) / np.sqrt(2),
"(b) σz eigenstate |↑z⟩": np.array([1, 0], dtype=complex),
"(c) tilted cos(π/8)|0⟩+sin(π/8)|1⟩":
np.array([np.cos(np.pi/8), np.sin(np.pi/8)], dtype=complex),
}
print(f"\n{'State':36s}{'Δσx':>8s}{'Δσy':>8s}{'LHS':>9s}{'Bound|⟨σz⟩|':>12s}{'':>6s}")
print("-" * 79)
for name, psi in states.items():
dx, dy = delta(sx, psi), delta(sy, psi)
lhs = dx * dy
lower = abs(expect(sz, psi))
ok = "✔" if lhs >= lower - 1e-12 else "✘"
print(f"{name:36s}{dx:8.4f}{dy:8.4f}{lhs:9.4f}{lower:12.4f}{ok:>4s}")
print("\nAll three examples satisfy the inequality, each with equality (minimum-uncertainty states).")
10.6 Chapter Summary¶
Review of the Theoretical Thread¶
This chapter took the mathematical language of discrete probability as its starting point (§10.1) and established the basic tools needed for the theory of quantum measurement: the normalization condition, the expectation value, and the variance. These seemingly elementary notions acquire deep physical counterparts in the quantum world—normalization corresponds to the conservation of probability, the expectation-value formula carries over directly to the quantum average of an observable , and the standard deviation becomes the precise measure of uncertainty. The role of this section is to remind the reader that the strangeness of quantum mechanics lies not in any violation of probability theory but in the entirely new physical interpretation it gives to probability theory.
On this foundation, §10.2 gave the most central answer of quantum mechanics: a pure state is a normalized vector in a Hilbert space, a “superposition state” is the natural result of expanding a vector in different bases, and the Born rule is the bridge between the geometric inner product and measurement probabilities. The description by pure states faces a fundamental limitation, however: many real physical systems—particles in thermal equilibrium, subsystems of entangled systems—cannot be characterized by a single state vector. §10.3 therefore introduced the density matrix , which places pure states (rank-one projections) in a unified framework of Hermitian, positive semidefinite matrices with trace one. The off-diagonal entries of the density matrix reveal quantum coherence, their disappearance is the mathematical embodiment of decoherence, and the purity provides an exact criterion distinguishing pure states from mixed states.
With the density matrix as a tool, §10.4 could describe completely the two levels of quantum measurement: the act of measurement (the off-diagonal entries vanish and coherence is lost) and the read-out of the outcome (the state collapses to a conditional pure state). The Pauli matrices , as the three observables of the spin-1/2 system, provide a model that can be computed completely, and their noncommutation relations are not merely algebraic identities but the direct cause of the mutual interference of measurements. This preparation bears fruit in §10.5: with the Cauchy-Schwarz inequality at its core, the Robertson-Schrödinger inequality turns the noncommutativity of operators rigorously into a lower bound on the uncertainty product. The spin system demonstrates this structure perfectly in a finite-dimensional space, while the position-momentum uncertainty relation , through the trace argument, rigorously rules out any finite-dimensional realization, revealing the essential difference in mathematical structure between continuous and discrete degrees of freedom.
Connections to Other Chapters¶
This chapter is deeply rooted in the theory of inner product spaces of Chapter 9. Hermitian matrices (observables), unitary matrices (quantum evolution), positive semidefinite matrices (density matrices), and the spectral theorem (the mathematical pillar of projective measurement)—these are exactly the tools that Chapter 9 built systematically, and in this chapter they receive their most vivid physical interpretation. In particular, the spectral decomposition is a mathematical theorem in quantum mechanics, but connecting its eigenvalues and projections to measurement statistics also requires the Born rule and the measurement postulates; it shows that all measurement information about an observable is completely encoded in its eigenvalues and eigenprojections. The theorem of Chapter 9 that the eigenvalues of a Hermitian matrix are all real and its eigenvectors orthogonal reads physically as “measurement outcomes are necessarily real, and the post-measurement state is necessarily one of the orthogonal components”—a correspondence that depends on the postulate of finite-dimensional, rank-one ideal projective measurement adopted in this chapter.
The end point of this chapter leads naturally to the singular value decomposition (SVD) of Chapter 11. §10.3 already previewed the problem of the reduced density matrix of an entangled subsystem: when the composite system is in an entangled state, the state of subsystem A must be described by a mixed state, and the standard tool for characterizing this mixed state is precisely the Schmidt decomposition —a special form of the SVD for composite quantum systems. The distribution of the Schmidt coefficients quantifies the degree of entanglement, and the Schmidt number (the number of nonzero coefficients) determines whether entanglement is present. From this point of view, this chapter is the physical motivation for Chapter 11, and Chapter 11 is the algebraic answer to the questions this chapter leaves open.
The Role of This Chapter in the Book¶
This chapter settles a central question that runs through the whole book: can the abstract algebraic machinery built in Chapters 8 and 9—eigenvalues, spectral decomposition, Hermitian matrices—describe the strangest physical phenomena of the real world? The answer is yes, with astonishing precision. Every basic rule of quantum mechanics—the discreteness of measurement outcomes, the superposition of states, the uncertainty principle, measurement collapse—is expressed in linear algebra, but the physical postulates and their range of validity must still be specified: observables are Hermitian matrices, measurements are projections, uncertainty is a necessary consequence of the noncommutativity of operators, and evolution is a unitary transformation. The quantum postulates are expressed in linear algebra; from these postulates together with linear algebra one can derive conclusions such as the measurement statistics and the uncertainty relations, but the physical postulates themselves are not theorems of linear algebra.
This chapter also opens a question, however: how should we describe nonclassical correlations among several quantum systems? An entangled state cannot be decomposed into a tensor product of states of its subsystems; this property of “the whole being greater than the sum of its parts” is the source of the power of quantum computing and quantum communication, and it is also the mathematical basis of the violation of Bell inequalities. Readers who enter Chapter 11 with this chapter’s understanding of density matrices, measurement theory, and uncertainty will find that the singular value decomposition is not merely a technical tool of matrix analysis but the natural language for characterizing quantum entanglement. There, the abstract power of linear algebra meets the physical strangeness of the quantum world once again, and the reader is already equipped with every tool that is needed.