In 1807 Joseph Fourier submitted a memoir on heat conduction to the French Academy of Sciences, in which he represented functions by trigonometric series; Lagrange and Laplace then raised objections to the practice of expanding functions in trigonometric series. The Academy made heat conduction the subject of its prize question for 1811, and Fourier won the prize with an expanded version of his memoir; the systematic account appeared in 1822 as Théorie analytique de la chaleur. The method forced mathematicians to rethink a basic question: in what sense can smooth sine and cosine waves approximate a function with corners or jumps? What it means to “have an expansion” cannot be separated from the class of functions and the mode of convergence.
In modern language: in the trigonometric system forms a complete orthogonal system, and the Fourier series converges to the original function in the mean-square sense; pointwise convergence requires a separate discussion.
Looking back on this history, later generations asked: what allowed Fourier to see past intuition? The key lay not in the summation tricks of calculus but in a deep “geometric perpendicularity” hidden among the trigonometric functions:
This integral is exactly zero, which geometrically means that sine waves of different frequencies are independent of one another and do not interfere. And the Fourier coefficients that give the amplitude of each component,
have exactly the same algebraic form as the formula for projecting a geometric arrow onto orthogonal coordinate axes in a finite-dimensional space—the only difference is that the “dot product,” a sum of products of discrete coordinates there, is elevated here into an integral of continuous functions over an interval.
This observation turns “functions are vectors too” from a mere metaphor into a fact: once we choose a function space, an inner product, and the appropriate completeness conditions, we can study length, orthogonality, and projection. Hilbert spaces extend this geometry to infinite dimensions and have become an important language for quantum theory, partial differential equations, and signal processing.
Here, however, Fourier’s story serves only as a mirror for ideas. Properties that are taken for granted in finite dimensions often require additional conditions in infinite-dimensional spaces—convergence, completeness, the domain of an operator, and so on.
The scope of this book is the rigorous and well-controlled setting of finite-dimensional linear spaces. This takes nothing away from the value of the theory in this chapter: in a bare vector space we have only “addition” and “scalar multiplication”; it is a limp world with no length, no angle, and no way to define a perpendicular. It is precisely the introduction of the inner product that gives the space a rigorous measure of length (the norm), a notion of angle, and optimal projection—turning a purely algebraic structure into a solid, concrete geometric stage.
The orthogonality, Gram-Schmidt orthogonalization, orthogonal projection operators, and spectral decomposition of symmetric matrices that we forge in and lose none of their core ideas in infinite-dimensional function spaces. In other words, what this chapter develops is by no means an algebraic special case confined to finite dimensions; it is a geometric foundation that spans the discrete and the continuous and, once the analytic conditions are supplied, extends to infinite dimensions.
Chapter Structure and Learning Objectives¶
The introduction pointed out that “orthogonality of functions” and “orthogonality of vectors” share exactly the same structure. This analogy rests on a precise axiomatic definition—and that is where this chapter begins.
§9.1 lays the foundation: it defines the inner product axiomatically and explains why the complex field forces us to introduce conjugation. Positive definiteness guarantees that “length” is never negative, the central premise that lets the whole geometric framework stand. From the inner product follow naturally the norm, orthogonality, and the Cauchy-Schwarz inequality, the last of which also foreshadows, in an elegant form, the geometric root of the uncertainty principle in Chapter 10. The section closes with Gram-Schmidt orthogonalization—an algorithm that gives a general method for constructing an orthogonal basis in any inner product space and is the geometric essence of the QR decomposition.
With the inner product in hand, §9.2 naturally asks: which linear transformations fully preserve this geometric structure? Orthogonal matrices (over the reals) and unitary matrices (over the complex numbers) are the answer. Their inverse equals their transpose or conjugate transpose, their 2-norm condition number in numerical computation is always 1, and geometrically they are compositions of rotations and reflections—they only rotate the orientation of the coordinate system and do not change distances or angles between vectors.
§9.3 is the heart of the chapter. Self-adjointness () means that a matrix “respects the inner product,” and from this it follows precisely that the eigenvalues must be real and that eigenvectors belonging to different eigenvalues must be orthogonal. Here the spectral decomposition theorem brings the diagonalization theory of Chapter 8 to its most perfect form: symmetric/Hermitian matrices are not only diagonalizable but can be diagonalized by an orthogonal (unitary) matrix. Positive definiteness further characterizes the “energy character” of such matrices and is an indispensable language in optimization and statistics.
§9.4 lands the theory in geometry through quadratic forms. The principal axis theorem shows that every quadratic form can be brought to a diagonal canonical form by rotating the coordinate system—the eigenvalues determine the shape of the ellipsoid or hyperboloid, and the eigenvectors give the directions of the principal axes. Hermitian forms then extend this picture to the complex field.
Once you have read the whole chapter, your understanding of “why symmetric matrices are special” will deepen from Chapter 8’s “they happen to be diagonalizable” to “self-adjointness gives them an orthogonal structure by nature.” This viewpoint will give you an entirely different depth of field when you meet the introduction to quantum mechanics in Chapter 10 and the singular value decomposition in Chapter 11.
The two symbols in the section titles of this chapter mark their reading priority: ◇ marks an optional extension (for example, §9.2.3 on unitary matrices and quantum computing), which can be skipped without affecting the main line; ◆ marks an advanced proof (for example, the equivalent conditions for positive definiteness), where on a first reading you may grasp the conclusion first and return to its derivation later.
9.1 Inner Product Spaces¶
Why Do We Need Inner Product Spaces?¶
In Chapter 2 we saw how the real inner product connects algebraic computation with geometric measurement—it encodes both the length of a vector () and the angle between vectors ().
We now face a broader question: how can similar geometric notions be defined in the complex vector space ? Carrying the definition of Chapter 2 over to the complex field unchanged runs into a fatal problem:
can be negative or even purely imaginary over the complex numbers! For example, take ; then . This would mean that the “length” of a vector could be negative, which is geometrically meaningless.
The solution is to introduce conjugation:
This guarantees that “length” is never negative, and so keeps geometric intuition valid. This key insight leads us into the systematic theory of inner product spaces.
An inner product space is a vector space equipped with an additional inner product operation, which brings in geometric notions. When defining an inner product on a complex vector space, we must pay particular attention to how conjugation is handled.
9.1.1 Definition and Properties of the Inner Product¶
The most common inner product is the standard inner product on Euclidean space:
import numpy as np
# Inner product on a real vector space
x = np.array([1, 2, 3])
y = np.array([4, 5, 6])
# Standard inner product (two equivalent ways)
inner_product1 = np.dot(x, y)
inner_product2 = x @ y # using the @ operator
print(f"Inner product of real vectors: {inner_product1}") # output: 32
# Inner product on a complex vector space
z1 = np.array([1+2j, 3-1j, 0+1j])
z2 = np.array([2-1j, 1+1j, 3-2j])
# ⚠️ A common mistake
wrong_inner = np.dot(z1, z2) # (10+8j) - wrong!
print(f"Wrong method np.dot(): {wrong_inner}")
# ✅ The correct method
correct_inner = np.vdot(z1, z2) # -4j - correct!
print(f"Correct method np.vdot(): {correct_inner}")
# Verify positive definiteness: ⟨z1|z1⟩ must be a nonnegative real number
self_inner = np.vdot(z1, z1)
print(f"⟨z1|z1⟩ = {self_inner.real} ≥ 0") # 16.0
# Check by hand
manual_result = np.conj(z1) @ z2
print(f"Manual computation check: {manual_result}") # output: -4jThe inner product brings in the notion of a norm (length):
The inner product also brings in the notions of angle and orthogonality:
Using the Cauchy-Schwarz inequality, we can define the angle between vectors:
9.1.2 Orthonormal Bases¶
A powerful feature of inner product spaces is that we can define orthogonal bases and orthonormal bases, which are extremely useful both in computation and in theoretical analysis.
In an orthonormal basis, representing and manipulating vectors becomes remarkably simple:
9.1.3 The Gram-Schmidt Orthogonalization Method¶
Given a set of linearly independent (complex) vectors, how can we turn it into an orthogonal basis or an orthonormal basis? The answer is the Gram-Schmidt orthogonalization method (the Gram-Schmidt process).
The Gram-Schmidt process can be expressed uniformly in matrix form.
§2.5 already demonstrated in code a first version of orthogonalizing two vectors; the definition in §9.1.3 extends the same idea to vectors and expresses it uniformly in the language of matrices: in , the columns of are the normalized results of removing projections step by step, and records the coefficients of each projection.
The QR decomposition applies column operations to , so and have the same image. This property is why the QR decomposition is widely used in numerical methods such as solving least-squares problems and computing eigenvalues.
The theory of inner product spaces provides a unified framework that systematically extends geometric notions—orthogonality, length, and angle—to abstract vector spaces. Within this framework, Gram-Schmidt orthogonalization gives a concrete algorithm for constructing orthogonal and orthonormal bases, which is of fundamental importance both for building orthogonal decompositions in theoretical analysis and for solving linear systems stably in numerical computation.
import numpy as np
import scipy.linalg as sclinalg
def gram_schmidt(vectors, rtol=None, atol=0.0):
"""MGS; input has shape (n, d), and each row is one vector.
The tolerance is set from the largest vector norm in the input. Raises an explicit error on numerical degeneracy;
this is a test of linear independence up to a tolerance, not the exact algebraic rank.
"""
vectors = np.asarray(vectors)
if vectors.ndim != 2:
raise ValueError("vectors must be a two-dimensional array")
n, d = vectors.shape
dtype = np.result_type(vectors.dtype, np.float64)
if n == 0:
return np.empty((0, d), dtype=dtype)
vectors = vectors.astype(dtype)
if rtol is None:
rtol = max(n, d) * np.finfo(float).eps
if rtol < 0 or atol < 0:
raise ValueError("tolerances must not be negative")
threshold = atol + rtol * np.max(np.linalg.norm(vectors, axis=1))
basis = []
for v in vectors:
w = v.copy()
for q in basis:
w -= np.vdot(q, w) * q
norm = np.linalg.norm(w)
if norm <= threshold:
raise ValueError("input is numerically degenerate at the given tolerance; adjust the tolerance or use a rank-revealing decomposition")
basis.append(w / norm)
return np.asarray(basis).reshape(-1, d)
# Test: complex vectors
vectors = np.array([
[1, 1j],
[1+1j, 2]
])
orthonormal_basis = gram_schmidt(vectors)
print("Orthonormal basis:")
print(orthonormal_basis)
# Orthogonality check (using the conjugate transpose)
gram_matrix = orthonormal_basis @ orthonormal_basis.conj().T
print("\nOrthogonality check (Gram matrix):")
print(np.round(gram_matrix, 10))
# Compare with SciPy QR
Q, R = sclinalg.qr(vectors.T)
print("\nResult of the SciPy QR function:")
print(Q.T)
# Compare with NumPy's built-in QR
Q, R = np.linalg.qr(vectors.T)
print("\nResult of the NumPy QR function:")
print(Q.T)
# Compare with SciPy orth, which uses the singular value decomposition. The basis it returns is ordered by the distribution of "variance" or "energy" and usually does not line up with the directions of the original vectors at all.
scipy_orth = sclinalg.orth(vectors.T)
print("\nResult of the SciPy orth function, with an orthogonality check:")
print(scipy_orth.T)
gram_matrix = scipy_orth.conj().T @ scipy_orth
print(np.round(gram_matrix, 10))
9.2 Orthogonal Matrices and Unitary Matrices¶
Orthogonal matrices and unitary matrices are special matrices that preserve the lengths and inner products of vectors, and they play an important role in many applications. Orthogonal matrices correspond to distance-preserving transformations of real vector spaces, and unitary matrices are their natural generalization to complex vector spaces. Both classes of matrices play central roles in numerical computation, quantum mechanics, signal processing, and other fields.
9.2.1 Definitions and Basic Properties of Orthogonal and Unitary Matrices¶
Definitions¶
An orthogonal matrix is the special case of a unitary matrix over the real numbers—when all the entries of a matrix are real, the conjugate transpose reduces to the transpose, and the definition of a unitary matrix reduces naturally to that of an orthogonal matrix.
Basic Properties¶
Orthogonality of Columns and Rows¶
The columns and the rows of an orthogonal or unitary matrix each form an orthonormal basis:
Examples of Orthogonal Matrices¶
Every orthogonal matrix can be written as a product of several basic orthogonal transformations (such as rotations and reflections).
9.2.2 The Geometric Meaning of Orthogonal and Unitary Transformations¶
Orthogonal and unitary transformations preserve the lengths of vectors and the angles between them, which gives them a clear geometric interpretation.
Definition of Orthogonal Transformations and Their Equivalence with Orthogonal Matrices¶
A linear transformation is called an orthogonal transformation (on a real vector space) or a unitary transformation (on a complex vector space) if it preserves the inner product:
for all vectors .
In a finite-dimensional Euclidean space, orthogonal transformations correspond one-to-one to orthogonal matrices:
Let be a linear transformation. Then is an orthogonal transformation if and only if there exists an orthogonal matrix such that for all .
Similarly, on the complex vector space , unitary transformations correspond to unitary matrices.
We give a detailed proof:
() An orthogonal matrix gives an orthogonal transformation.
Clearly an orthogonal matrix defines an orthogonal transformation . By Lemma 1 (the matrix inner product duality lemma), for all vectors :
() An orthogonal transformation corresponds to an orthogonal matrix. Let be a linear transformation that preserves the inner product; we must prove that the matrix of with respect to the standard basis is an orthogonal matrix.
Let be the standard basis of , and let be the matrix of with respect to the standard basis, that is, .
Column of is the coordinate vector of with respect to the standard basis. Since preserves the inner product,
This means that the column vectors of form an orthonormal basis.
Now compute :
Therefore , that is, is an orthogonal matrix.
The proof for unitary transformations on complex vector spaces is entirely analogous; simply replace the transpose by the conjugate transpose.
The cosine formula in -dimensional space is equivalent to the cosine formula in two-dimensional space.
We use the fact that orthogonal transformations preserve the lengths of vectors and the angles between them.
1. Constructing an orthogonal transformation
Let and be two nonzero vectors in . When they are linearly independent, they span a plane; when they are collinear, the question of their angle is trivial.
Using the Gram-Schmidt orthogonalization method, we can construct an orthogonal basis of whose first two basis vectors and span exactly the plane .
2. An orthogonal matrix rotates the plane onto the -plane
Construct an orthogonal matrix (that is, one satisfying ) such that
Since maps each vector of the whole orthogonal basis to a new orthogonal basis, after the transformation by the plane originally spanned by and is mapped onto the -plane.
3. Invariance of length and angle
Orthogonal matrices preserve the lengths and inner products of vectors:
Hence the angle between and (which satisfies ) is unchanged by the transformation .
4. Reducing the problem to two dimensions
Since and both lie in the -plane, they can be written as
In the two-dimensional plane their inner product is , and their lengths are and .
Since orthogonal transformations preserve inner products and lengths, the cosine formula in the original -dimensional space,
is equivalent to the formula in the two-dimensional plane:
5. Conclusion
By an orthogonal transformation, two vectors in -dimensional space can be rotated into a two-dimensional plane. Since orthogonal matrices preserve lengths and angles, the inner product and the lengths of the vectors are the same before and after the transformation, and so the angle between the two vectors stays the same. Therefore .
Geometric Decomposition of Orthogonal Transformations¶
The geometric meaning of an orthogonal transformation depends on its eigenvalues and eigenvectors. On a real vector space, an orthogonal transformation can be decomposed into:
Rotations: corresponding to complex pairs of eigenvalues , acting as rotations in the corresponding two-dimensional subspaces
Reflections: corresponding to the eigenvalue -1, acting as reflections in the corresponding directions
The identity: corresponding to the eigenvalue +1, leaving the corresponding directions unchanged
Every orthogonal matrix can be decomposed as a product of rotations and a reflection:
where each is a two-dimensional rotation (acting as the identity on the other dimensions), and is a reflection matrix (or the identity matrix when ).
Using two-dimensional rotations (Givens rotations), we can rotate the column vectors onto the standard basis one at a time, and then use a reflection matrix to fix the signs as needed, arriving at the identity matrix. Since every Givens rotation is an orthogonal matrix and so is its inverse, .
◇9.2.3 Special Properties of Unitary Matrices and Quantum Computing¶
Unitary matrices are the generalization of orthogonal matrices to the complex field, and they play a central role in quantum mechanics and quantum computing.
The Importance of Unitary Matrices in Quantum Mechanics¶
This section considers pure states of closed systems and ideal quantum gates. Within this scope, a pure state of a physical system is described by a unit vector (a quantum state) in a complex vector space, and the evolution of the system is described by a unitary transformation. This is because:
Conservation of probability: unitary transformations preserve the norm of vectors, and therefore preserve the normalization condition of quantum states
Reversibility: physical evolution (in the absence of measurement) is reversible, which corresponds to the invertibility of unitary matrices
Preservation of inner products: the inner product between normalized quantum states gives a probability amplitude, whose squared modulus gives the probability of the corresponding projective measurement; unitary transformations preserve inner products and therefore also preserve these probabilities
Common Unitary Matrices: Quantum Gates¶
The most basic single-qubit gates in quantum computing are all unitary matrices:
Pauli matrices (all unitary; are real orthogonal matrices):
The Hadamard gate (a real unitary matrix):
To verify that is unitary: since is a real matrix, :
Phase gates (complex unitary matrices):
General rotation gates:
Note that the Hadamard gate is a real unitary matrix (and hence also an orthogonal matrix), while the phase gates and for a general angle are non-real unitary matrices; for special angles may also be a real orthogonal matrix.
The Quantum Fourier Transform¶
The Hadamard gate can be viewed as the special case of the quantum Fourier transform (QFT) for a single qubit.
For qubits, the QFT matrix is a complex unitary matrix defined by
where is a -th root of unity.
The QFT plays a key role in many quantum algorithms, especially in solving periodicity problems, as in Shor’s algorithm (integer factorization) and the quantum phase estimation algorithm.
From the point of view of matrix theory, the QFT can be regarded as a special Vandermonde matrix and inherits many of the properties of Vandermonde matrices.
The quantum Fourier transform and the classical fast Fourier transform (FFT) are closely related mathematically, but differ fundamentally in computational complexity:
The classical FFT needs operations to process data points
The quantum QFT needs only quantum gates to process a -dimensional quantum state
This exponential speedup is one of the sources of the advantage of quantum computing.
9.2.4 Numerical Stability and Applications¶
Orthogonal matrices have important stability advantages in matrix computations.
Why Orthogonal Matrices Are Numerically Stable¶
The condition number of an orthogonal matrix, , is the smallest possible value among all matrices. This means:
Errors are not amplified: when we compute , any rounding error in the input is not amplified
Numerical accuracy is maintained: in iterative algorithms, orthogonal transformations do not amplify the 2-norm of existing errors, although new rounding errors produced by repeated operations may still accumulate
Ill-conditioned problems are avoided: using orthogonal matrices avoids many ill-conditioned situations in numerical linear algebra
In practical numerical computation, when a general matrix needs to be factored, decompositions involving orthogonal matrices (such as the QR decomposition and the singular value decomposition) should be preferred over decompositions that may involve ill-conditioned matrices (such as the eigendecomposition). This is one of the basic principles of numerical linear algebra.
The Frobenius norm of a matrix is defined as the square root of the sum of the squared moduli of all its entries,
It regards an matrix as a vector in and takes its length. Orthogonal (unitary) similarity transformations leave the Frobenius norm unchanged: for an orthogonal matrix , (which follows immediately from the cyclic property of and ). This is precisely how “orthogonal transformations do not amplify errors” shows up at the level of matrices. The full theory of matrix norms (including operator-induced norms) will be developed systematically in §11.2 of Chapter 11.
Main Applications¶
Orthogonal and unitary matrices play central roles in the following fields:
Numerical linear algebra:
QR decomposition: factoring a matrix into the product of an orthogonal matrix and an upper triangular matrix
Singular value decomposition (SVD): the subject of Chapter 11
Iterative methods: such as Arnoldi iteration and the Lanczos method
Computer graphics:
Rotation, scaling, and translation of 3D objects
Viewing transformations and projection
Interpolation in animation (SLERP: spherical linear interpolation)
Signal processing:
The discrete Fourier transform (DFT) and its fast algorithm (FFT)
The discrete cosine transform (DCT): the basis of JPEG compression
Wavelet transforms: multiscale signal analysis
Quantum computing:
Quantum gate operations (such as the various unitary matrices introduced in this section)
Quantum algorithm design (such as Shor’s algorithm and Grover’s algorithm)
Simulation of quantum state evolution
Data analysis and machine learning:
Principal component analysis (PCA): discussed in detail in Chapter 11
Independent component analysis (ICA)
The whitening transformation
Orthogonal and unitary matrices are a natural extension of the theory of inner product spaces. They preserve inner products, lengths, and angles, and geometrically they correspond to rigid motions (rotations and reflections). An orthogonal transformation does not change the geometric structure of space, only the angle from which we view it. Understanding the properties of orthogonal transformations is essential for understanding many advanced concepts and applications in linear algebra; in particular, it lays the foundation for the spectral theorem for symmetric matrices in the next section, the introduction to quantum mechanics in Chapter 10, and the singular value decomposition in Chapter 11.
The question this section answered is: in an inner product space, which matrices give linear transformations that leave the geometric structure completely unchanged? The answer is orthogonal matrices () and their complex generalization—unitary matrices (). Their inverse equals their transpose or conjugate transpose, their 2-norm condition number is always 1, and they are the most stable class of matrices in numerical computation.
The geometric core of this section is that an orthogonal transformation is a composition of rotations and reflections, and a unitary transformation is the analogue of a “rotation” in complex space. In both cases the eigenvalues have modulus 1; from the “measurement viewpoint,” they change only the coordinate system of observation, not the object itself. The columns and the rows of an orthogonal matrix are both orthonormal bases, a property that follows directly from the definition , and from it follows that its determinant must have absolute value 1.
It is worth special attention that an orthogonal matrix itself need not be orthogonally diagonalizable—the complex eigenvalues of a rotation matrix reveal this detail. The special objects that truly can be “diagonalized by an orthogonal basis” are the protagonists of the next section—symmetric matrices and Hermitian matrices.
9.3 Symmetric Matrices and Hermitian Matrices¶
Symmetric matrices and Hermitian matrices are two special and important classes of matrices in linear algebra. Symmetric matrices play a central role in real vector spaces, and Hermitian matrices are the natural generalization of symmetric matrices to complex vector spaces. Both classes hold a fundamental place in physics, engineering, and data analysis. Their eigenvalues are all real, and an orthonormal basis of eigenvectors can be chosen for them, which gives them unique advantages in representing physical systems, analyzing the structure of data, and designing algorithms.
9.3.1 Eigenvalues and Eigenvectors of Symmetric and Hermitian Matrices¶
Definitions of Symmetric and Hermitian Matrices¶
Hermitian matrices are the natural generalization of symmetric matrices; for real matrices the two notions coincide (because the conjugate of a real number is the number itself).
Properties of the Eigenvalues and Eigenvectors¶
The eigenvalues and eigenvectors of symmetric and Hermitian matrices have special properties:
For a symmetric matrix or a Hermitian matrix :
All eigenvalues are real
Eigenvectors belonging to different eigenvalues are orthogonal
We prove that the eigenvalues of a Hermitian matrix are all real (the symmetric case can be regarded as a special case).
Let be an eigenvalue of and a corresponding eigenvector, that is, .
Compute
By the eigenvalue equation , we have
Use the Hermitian property and the duality lemma
Since is a Hermitian matrix, it satisfies . By the matrix inner product duality lemma Lemma 1:
Since ,
Expand the inner product
Using , we can write the right-hand side as
Conclude that the eigenvalue is real
Combining the results of steps 1 and 3:
Since is an eigenvector, , that is, . Dividing both sides by gives
This shows that is real.
Eigenvectors belonging to different eigenvalues are orthogonal
Let and with . Then
Since , .
The key points of this proof are:
Use the Hermitian property
Apply the matrix inner product duality lemma to “move” the matrix within the inner product expression
Compare two ways of computing to obtain
For a real symmetric matrix, and all entries are real, so the same proof applies (in this case the conjugate transpose reduces to the transpose).
The Spectral Decomposition Theorem¶
Let be a symmetric matrix. Then there exist an orthogonal matrix and a diagonal matrix such that
For a Hermitian matrix , there exist a unitary matrix and a real diagonal matrix such that
This decomposition is called the spectral decomposition, where and the are the eigenvalues of .
The spectral decomposition can be rewritten in outer-product form:
where is the -th normalized eigenvector of . This shows that a symmetric/Hermitian matrix can be expressed as a linear combination of rank-one matrices.
Unlike a general matrix, a symmetric matrix never has missing eigenvectors.
Prerequisites
The eigenvalues of a Hermitian matrix are all real.
Eigenvectors belonging to different eigenvalues are orthogonal.
A unitary matrix satisfies , where is the identity matrix.
Proof In the inductive proof below, and denote the unitary matrix and the real diagonal matrix being constructed, that is, the and of the Hermitian case in the statement of the theorem (, ).
We proceed by mathematical induction.
Base case When , with (because the diagonal entries of a Hermitian matrix must be real). Then and satisfy the theorem, that is, .
Induction hypothesis Assume that the theorem holds for every Hermitian matrix; that is, every Hermitian matrix can be written as , where is a unitary matrix and is a diagonal matrix.
Inductive step Now consider a Hermitian matrix .
First, since is a Hermitian matrix, by the fundamental theorem of algebra in §8.1, has at least one eigenvalue with a corresponding eigenvector . We may normalize to be a unit vector, that is, . We can then construct a unitary matrix , where consists of unit vectors orthogonal to that together with it form an orthonormal basis of . (This can be constructed using the QR decomposition.)
Consider the matrix :
Compute each block of :
a. The upper-left block: , because is a unit vector.
b. The upper-right block: . Note that for every we have
c. Similarly, the lower-left block: , because is orthogonal to .
d. The lower-right block: , where is a Hermitian matrix.
Therefore has the form
By the induction hypothesis, the Hermitian matrix can be decomposed as , where is a unitary matrix and is a diagonal matrix.
Define the matrix . Then is a unitary matrix, and
Therefore we obtain
where is a unitary matrix and is a diagonal matrix.
This completes the inductive step and proves that the theorem holds for every Hermitian matrix.
Conclusion By mathematical induction, we have proved that for every Hermitian matrix there exist a unitary matrix and a diagonal matrix such that .
From the spectral decomposition , let ; then
Consider the Hermitian matrix
Find its eigenvalues and eigenvectors:
Characteristic equation:
Solving gives the eigenvalues , (both real, as the properties of Hermitian matrices require)
For , the unit eigenvector is
For , the unit eigenvector is
Construct the unitary matrix and the diagonal matrix:
Verify the spectral decomposition:
import numpy as np
np.set_printoptions(precision=4, suppress=True, linewidth=100)
def print_header(title):
print("=" * 60)
print(f" {title}")
print("=" * 60)
def print_step(step, desc):
print(f"\n▶ Step {step}: {desc}")
print("-" * 40)
def compare_print(label, actual, expected_desc):
print(f"\n[{label}]")
print(f" Computed value:\n{actual}")
print(f" Expected value: {expected_desc}")
# ============================================================
print_header("Example 9.3.1a | Spectral Decomposition of a Real Symmetric Matrix A = PΛP^T")
# ============================================================
# --- Step 1 ---
print_step(1, "Define the symmetric matrix A")
A = np.array([[4, 1],
[1, 4]], dtype=float)
print(A)
# --- Step 2 ---
print_step(2, "Compute the eigenvalues and eigenvectors (eigh guarantees real output)")
eigenvalues, P = np.linalg.eigh(A)
idx = np.argsort(eigenvalues)[::-1] # descending order: λ₀ ≥ λ₁
eigenvalues = eigenvalues[idx]
P = P[:, idx]
print(f" Eigenvalues Λ = diag({eigenvalues})")
print(f" Expected from the text: λ₀ = 5, λ₁ = 3 ✓")
print(f"\n Orthogonal matrix P (columns = normalized eigenvectors):")
print(P)
print(f" Expected from the text: q₀ = (1, 1)/√2, q₁ = (-1, 1)/√2")
# --- Step 3 ---
print_step(3, "Verify that P is orthogonal: P^T P = I")
compare_print("P^T P", np.round(P.T @ P, 10), "the identity matrix I")
# --- Step 4 ---
print_step(4, "Reconstruct from the spectral decomposition: A_rec = P Λ P^T")
A_rec = P @ np.diag(eigenvalues) @ P.T
compare_print("A_rec = PΛP^T", A_rec,
"the original matrix A =\n [[4. 1.]\n [1. 4.]]")
print(f"\n Maximum reconstruction error: {np.max(np.abs(A - A_rec)):.2e}")
# --- Step 5 ---
print_step(5, "Outer-product expansion: A = λ₀ q₀q₀^T + λ₁ q₁q₁^T")
term0 = eigenvalues[0] * np.outer(P[:, 0], P[:, 0])
term1 = eigenvalues[1] * np.outer(P[:, 1], P[:, 1])
print(f" λ₀·q₀q₀^T =\n{term0}")
print(f" λ₁·q₁q₁^T =\n{term1}")
print(f" Sum =\n{np.round(term0 + term1, 10)}")
# ============================================================
print_header("Example 9.3.1b | Spectral Decomposition of a Hermitian Matrix A = UΛU*")
# ============================================================
# --- Step 1 ---
print_step(1, "Define the Hermitian matrix A and verify A = A*")
A_h = np.array([[ 2, 1j],
[-1j, 2]], dtype=complex)
print(A_h)
print(f" A = A*: {np.allclose(A_h, A_h.conj().T)}")
# --- Step 2 ---
print_step(2, "Compute the eigenvalues and eigenvectors (eigh guarantees real eigenvalues)")
eigenvalues_h, U = np.linalg.eigh(A_h)
idx = np.argsort(eigenvalues_h)[::-1] # descending order: λ₀ ≥ λ₁
eigenvalues_h = eigenvalues_h[idx]
U = U[:, idx]
print(f" Eigenvalues (all real): {eigenvalues_h}")
print(f" Expected from the text: λ₀ = 3, λ₁ = 1 ✓")
print(f"\n Unitary matrix U (columns = normalized eigenvectors):")
print(U)
print(f" Expected from the text, ordered by λ=3,1: u₀ = (1, -i)/√2, u₁ = (1, i)/√2; eigh returns ascending order and may give different phases")
# --- Step 3 ---
print_step(3, "Verify that U is unitary: U*U = I")
compare_print("U*U", np.round(U.conj().T @ U, 10), "the identity matrix I")
# --- Step 4 ---
print_step(4, "Reconstruct from the spectral decomposition: A_rec = U Λ U*")
A_rec_h = U @ np.diag(eigenvalues_h) @ U.conj().T
compare_print("A_rec = UΛU*", A_rec_h, "the original matrix A_h")
print(f"\n Maximum reconstruction error: {np.max(np.abs(A_h - A_rec_h)):.2e}")
# --- Step 5 ---
print_step(5, "Verify that eigenvectors for different eigenvalues are orthogonal: ⟨u₀|u₁⟩ = u₀*·u₁")
inner = np.vdot(U[:, 0], U[:, 1])
print(f" ⟨u₀|u₁⟩ = {inner:.2e} (should be 0, by property 2 of the theorem)")9.3.2 Positive Definite and Positive Semidefinite Matrices¶
Definition of Positive Definiteness¶
A symmetric matrix is called:
positive definite if for every nonzero vector
positive semidefinite (or nonnegative definite) if for every vector
Negative definite and negative semidefinite matrices are defined similarly.
A positive definite matrix is usually written and a positive semidefinite matrix . Likewise, a negative definite matrix is written and a negative semidefinite matrix .
Equivalent Conditions for Positive Definiteness¶
Positive definiteness can be characterized in several ways:
For a symmetric matrix , the following conditions are equivalent:
is positive definite
All eigenvalues of are positive
All leading principal minors of are positive (the -th leading principal minor is the determinant of the upper-left submatrix, ; Sylvester’s criterion)
can be factored as , where is lower triangular with positive diagonal entries (the Cholesky decomposition)
There exists a nonsingular square matrix such that
, where is a matrix with linearly independent columns
The proof uses a cyclic chain of implications:
That is, we first prove , then proceed in order along , finally closing the cycle of equivalences.
Step 1: (1) (2)
By the spectral theorem for real symmetric matrices, there exists an orthogonal matrix such that
For any , let (since is orthogonal and hence invertible); then
: if some , take ; then , contradicting positive definiteness.
: if all , then for every , so is positive definite.
Step 2: (2) (3)
For any , let be the upper-left leading principal submatrix of . For any , let
Then , and since is positive definite (by (1) (2)),
Hence is positive definite and its eigenvalues are all positive, so
Step 3: (3) (4) (existence of the Cholesky decomposition)
We use induction on . In the induction below, the subscripts in and denote the size , not the 0-based numbering of the leading principal minors in the previous step.
Base case : condition (3) gives ; take .
Induction hypothesis: the Cholesky decomposition exists for every symmetric matrix satisfying condition (3).
Inductive step: partition as
All the leading principal minors of are inherited from , so by the induction hypothesis .
Let and consider the Schur complement (see Chapter 5)
The block determinant formula gives , and since and , we get . Take
A direct computation verifies , and the diagonal entries of are all positive. This completes the induction.
Step 4: (4) (5)
Let . This is an upper triangular square matrix whose diagonal entries are all positive, so is nonsingular. A direct computation gives
Step 5: (5) (6)
Let . Since is a nonsingular square matrix, , so its columns are linearly independent. Thus
and the columns of are linearly independent, so condition (6) holds.
Step 6: (6) (1)
The columns of being linearly independent is equivalent to . For any we have , and hence
So is positive definite, which brings us back to condition (1).
Positive semidefinite matrices have similar equivalent conditions.
For a symmetric matrix , the following conditions are equivalent:
is positive semidefinite, that is, for all
All eigenvalues of are nonnegative
All principal minors of are nonnegative (for every index subset , the determinant of the submatrix formed by the corresponding rows and columns is )
can be factored as , where is lower triangular with nonnegative diagonal entries (the generalized Cholesky decomposition)
There exists a (possibly singular) square matrix such that
, where is any matrix (its columns need not be linearly independent)
The comparison with the positive definite version is as follows:
| Condition | Positive definite | Positive semidefinite |
|---|---|---|
| (2) Eigenvalues | ||
| (3) Principal minors | Leading principal minors | All principal minors |
| (4) Cholesky | Diagonal entries , decomposition unique | Diagonal entries , decomposition not necessarily unique |
| (5) Square-matrix factorization | nonsingular | may be singular |
| (6) General factorization | Columns of linearly independent | No restriction on |
Consider the symmetric matrix
Test whether it is positive definite:
Compute the eigenvalues:
The characteristic equation is
Solving gives the eigenvalues , ,
Since all the eigenvalues are greater than 0, is positive definite
Test the leading principal minors:
The first leading principal minor:
The second leading principal minor:
The third leading principal minor:
All the leading principal minors are positive, so is positive definite
Properties and Applications of Positive Definite Matrices¶
Positive definite matrices have many important properties:
A positive definite matrix is invertible, and its inverse is also positive definite
The main diagonal entries of a positive definite matrix are all positive
The determinant of a positive definite matrix is positive
The sum of two positive definite matrices is positive definite
The Schur complement of a positive definite matrix is also positive definite
A positive definite matrix can define an inner product on a vector space:
Positive definite matrices are very important in applications:
In optimization theory, a positive definite quadratic objective function (with real symmetric positive definite) has a unique global minimum on
In the method of least squares, the normal-equations matrix is positive semidefinite
Covariance matrices in statistics are positive semidefinite
In control theory, the Lyapunov equation is closely tied to positive definite matrices
In finite element analysis, the stiffness matrix is usually positive definite
9.3.3 Positive Definite Hermitian Matrices and Their Applications¶
Definition of Positive Definite Hermitian Matrices¶
For complex vector spaces, the notion of positive definiteness extends naturally:
A Hermitian matrix is called positive definite if for every nonzero complex vector ,
Similarly, if for every nonzero complex vector , then is called a positive semidefinite Hermitian matrix.
It should be emphasized that although the entries of may be complex, the value of the expression is always real (because is a Hermitian matrix).
Equivalent Conditions for Positive Definite Hermitian Matrices¶
For a Hermitian matrix , the following conditions are equivalent:
is positive definite
All eigenvalues of are positive real numbers
All leading principal minors of are positive real numbers
, where is a complex matrix with linearly independent columns
can be factored as , where is lower triangular with positive diagonal entries (the complex version of the Cholesky decomposition)
Consider the Hermitian matrix
Verify that it is positive definite:
Compute the eigenvalues:
Characteristic equation:
Solving gives the eigenvalues ,
Since all the eigenvalues are positive real numbers, is a positive definite Hermitian matrix.
Compute its Cholesky decomposition :
One can verify that .
import numpy as np
from scipy.linalg import cholesky
# Define the Hermitian matrix A
A = np.array([[3, 2-1j],
[2+1j, 4]], dtype=complex)
print("Original matrix A:")
print(A)
print()
# Verify the Hermitian property
print("Hermitian property check (A = A^*):", np.allclose(A, A.conj().T))
print()
# Compute the eigenvalues
eigenvalues = np.linalg.eigvalsh(A)
print("Eigenvalues:")
print(f"λ₀ ≈ {eigenvalues[0]:.2f} (theoretical value: (7-√21)/2 ≈ 1.21)")
print(f"λ₁ ≈ {eigenvalues[1]:.2f} (theoretical value: (7+√21)/2 ≈ 5.79)")
print(f"All eigenvalues are positive: {np.all(eigenvalues > 0)} → the matrix is positive definite")
print()
# Cholesky decomposition
L = cholesky(A, lower=True)
print("Cholesky factor L:")
print(L)
print()
# Verify LL^* = A
print("Verify LL^* = A:")
A_reconstructed = L @ L.conj().T
print("Reconstructed matrix:")
print(A_reconstructed)
print(f"Maximum error: {np.max(np.abs(A - A_reconstructed)):.2e}")A positive definite Hermitian matrix can define an inner product on a complex vector space:
Given a positive definite Hermitian matrix , we can define an inner product on :
The corresponding norm is
Applications in Signal Processing¶
In signal processing and communications, positive definite Hermitian matrices are also widely used:
Covariance matrices: the autocorrelation matrix of a signal is a positive semidefinite Hermitian matrix; it is positive definite only when there is no direction with zero second moment
Multiple-input multiple-output (MIMO) systems: the Gram matrix of the channel matrix is a positive semidefinite Hermitian matrix; it is positive definite only when has full column rank. Calling it a covariance matrix requires a separately specified stochastic model
Beamforming and spatial filtering: optimization based on positive definite Hermitian matrices
Consider the Hermitian matrix :
Verify that is a Hermitian matrix
Test whether is positive definite
Compute the Cholesky decomposition of
Compute the square root of (that is, the positive definite Hermitian matrix satisfying )
Solution to Exercise 2
Hermitian property. . Take the conjugate of entry by entry and then transpose: lies in row 0, column 1; its conjugate is , which after transposition lands in row 1, column 0, matching the entry of the original in that position; the diagonal entries 5 are real. Hence , and is a Hermitian matrix.
Positive definiteness (by Theorem 8, check that all eigenvalues are positive). The characteristic equation is
so , that is, and . Both eigenvalues are positive, so is positive definite.
Cholesky decomposition . Let (). Comparing entries:
:
:
:
Therefore . (One can verify that .)
The square root is constructed from the spectral decomposition: if , then , where . The orthonormal eigenvectors are
Substituting into the outer-product form gives
where and . This matrix is positive definite Hermitian, and squaring it recovers .
This section answered a central question left open by Chapter 8: what condition guarantees that a matrix has all real eigenvalues and that an orthonormal basis can be chosen from its eigenvectors? The answer is precisely symmetry—a real symmetric matrix satisfies , a complex Hermitian matrix satisfies , and both are self-adjoint operators that “respect the inner product structure.”
The spectral decomposition theorem is the summit of this section: every symmetric/Hermitian matrix can be written as (or ), where (or ) is an orthogonal (unitary) matrix and is a real diagonal matrix. This decomposition embodies both the theory of orthogonal matrices from §9.2 and the eigenvalue theory of §8, and it is the most complete matrix structure theorem in the book so far.
Positive definiteness further refines the hierarchy of symmetric matrices: positive definite ↔ all eigenvalues positive ↔ all leading principal minors positive ↔ a Cholesky decomposition exists. A positive definite matrix can be regarded as the generator of a generalized inner product on a vector space, and such matrices appear widely as Hessian matrices in optimization, covariance matrices in statistics, and stiffness matrices in engineering.
Once we have mastered the spectral decomposition of symmetric matrices, we can understand the geometric essence of quadratic forms—the next section establishes this connection systematically.
9.4 Quadratic Forms and the Principal Axis Theorem¶
A quadratic form is a linear combination of squares and products of variables, widely used in optimization, statistics, quantum mechanics, computer graphics, and other fields. The principal axis theorem shows that a quadratic form can be simplified to a standard form by a suitable change of coordinates, a process that geometrically corresponds to finding the principal directions of a quadric surface. This section discusses the theory of quadratic forms and its applications in real and complex spaces in full, revealing the essential geometric meaning of symmetric and Hermitian matrices.
9.4.1 Definition and Basic Properties of Quadratic Forms¶
Definition of a Quadratic Form¶
Geometrically, quadratic forms correspond to conics and quadric surfaces, such as ellipses, hyperbolas, and ellipsoids. In physics, quadratic forms can represent key quantities such as energy functions and moments of inertia.
A general quadratic polynomial can be written as , consisting of a quadratic form, a linear term, and a constant term. In this chapter we focus mainly on the pure quadratic form .
Quadratic Forms and Symmetric Matrices¶
Every quadratic form can be represented by a symmetric matrix:
For any matrix , the quadratic form can also be represented by the symmetric matrix :
Therefore, when studying quadratic forms we usually assume that the matrix is symmetric.
Classifying Quadratic Forms by Definiteness¶
Quadratic forms can be classified according to the signs of the values they take:
A quadratic form is called:
positive definite if for every nonzero vector
positive semidefinite if for every vector
negative definite if for every nonzero vector
negative semidefinite if for every vector
indefinite if takes both positive and negative values
The definiteness of a quadratic form agrees with the definiteness of the corresponding symmetric matrix.
The definiteness of a quadratic form can be determined from the eigenvalues of the corresponding symmetric matrix:
Let be a real symmetric matrix with corresponding quadratic form :
is positive definite if and only if all eigenvalues of are positive
is positive semidefinite if and only if all eigenvalues of are nonnegative
is negative definite if and only if all eigenvalues of are negative
is negative semidefinite if and only if all eigenvalues of are nonpositive
is indefinite if and only if has both positive eigenvalues and negative eigenvalues
import numpy as np
np.set_printoptions(precision=4, suppress=True, linewidth=100)
def print_header(title):
print("=" * 60)
print(f" {title}")
print("=" * 60)
def print_step(step, desc):
print(f"\n▶ Step {step}: {desc}")
print("-" * 40)
# ============================================================
print_header("Example 9.4.1 | Determining the Definiteness of a Quadratic Form")
# ============================================================
# --- Step 1 ---
print_step(1, "Read off the symmetric matrix A from the coefficients of the quadratic form")
# Q(x,y,z) = 2x² + 3y² + 4z² - 2xy + 4xz - 6yz
# Off-diagonal terms are halved to symmetrize: a_ij = a_ji = coefficient/2
A = np.array([[ 2, -1, 2],
[-1, 3, -3],
[ 2, -3, 4]], dtype=float)
print(A)
print(f" Symmetry check A = A^T: {np.allclose(A, A.T)}")
# --- Step 2 ---
print_step(2, "Check the value of the quadratic form at a specific vector")
x = np.array([1, 1, 1], dtype=float)
Q_val = x @ A @ x
print(f" Q(1,1,1) = x^T Ax = {Q_val}")
print(f" Hand check: 2+3+4−2+4−6 = {2+3+4-2+4-6} ✓")
# --- Step 3 ---
print_step(3, "Compute the eigenvalues (eigh guarantees real output, in ascending order)")
eigenvalues = np.linalg.eigvalsh(A)
print(f" λ₀ ≈ {eigenvalues[0]:.4f}")
print(f" λ₁ ≈ {eigenvalues[1]:.4f}")
print(f" λ₂ ≈ {eigenvalues[2]:.4f}")
# --- Step 4 ---
print_step(4, "Determine the definiteness")
is_pd = bool(np.all(eigenvalues > 0))
print(f" All eigenvalues > 0: {is_pd}")
print(f" → the quadratic form is positive definite ✓")9.4.2 Diagonalization of Quadratic Forms in Real Space and the Principal Axis Theorem¶
Algebraic Derivation of the Diagonalization¶
By a suitable change of coordinates, every quadratic form can be simplified to a form containing only square terms:
For every quadratic form with a symmetric matrix, there exists an orthogonal change of variables such that in the new coordinates the quadratic form becomes
where are the eigenvalues of the matrix and the columns of are the corresponding orthonormal eigenvectors. The coordinate axes of the new coordinate system are called the principal axes of the quadric surface.
This theorem is in fact an application of the spectral theorem for real symmetric matrices. It shows that every real symmetric matrix can be diagonalized by an orthogonal matrix:
where is the diagonal matrix of eigenvalues.
This theorem has an important geometric meaning:
The eigenvalues determine how much the quadric surface “stretches” along each principal axis
The eigenvectors determine the principal directions (principal axes) of the quadric surface
The orthogonal transformation can be regarded as a rotation of the coordinate system that aligns the new coordinate system with the principal axes of the quadric surface
Consider the quadratic form
Its matrix form is
Computing the eigenvalues and eigenvectors:
Eigenvalues: ,
Orthonormal eigenvectors: ,
Construct the orthogonal matrix:
The canonical form in the new coordinate system is
This represents an ellipse whose principal axes lie along the directions of the eigenvectors and ; the length of the semi-major axis is proportional to and the length of the semi-minor axis is proportional to .
The Law of Inertia for Quadratic Forms and the Classification of Quadric Surfaces¶
Although a quadratic form can be given different representations by different changes of coordinates, some properties are invariant:
For a real quadratic form, no matter how a change of coordinates is chosen to turn it into a sum of square terms, the number of positive terms, the number of negative terms, and the number of zero terms are fixed. These numbers are called the positive index of inertia, the negative index of inertia, and the zero index of inertia of the quadratic form, respectively.
The indices of inertia are directly related to the eigenvalues of the symmetric matrix: the positive index of inertia equals the number of positive eigenvalues, the negative index of inertia equals the number of negative eigenvalues, and the zero index of inertia equals the number of zero eigenvalues.
Below we classify only the positive level surfaces of nondegenerate quadratic forms in three-dimensional real space, with no linear terms.
After an orthogonal change of basis the equation becomes , with all three eigenvalues nonzero.
Three positive: an ellipsoid.
Two positive and one negative: a hyperboloid of one sheet.
One positive and two negative: a hyperboloid of two sheets.
Three negative: the empty set.
Changing the level value, zero eigenvalues, or linear terms all change the classification; general quadric surfaces and paraboloids are not developed in this chapter.
9.4.3 Hermitian Forms in Complex Space¶
From Quadratic Forms to Hermitian Forms¶
In a complex vector space, the notion of a quadratic form must be extended to that of a Hermitian form in order to keep its positive definiteness and geometric meaning:
On the complex vector space , a Hermitian form is a mapping defined by
where is a Hermitian matrix (that is, ) and denotes the conjugate transpose of .
The values of a Hermitian form are always real, which is guaranteed by the properties of Hermitian matrices. In fact, for every complex vector , the Hermitian form satisfies
which guarantees that is real.
Comparing Hermitian Matrices and Symmetric Matrices¶
Hermitian matrices (defined in Definition 10) are the natural extension of symmetric matrices to complex vector spaces:
Hermitian matrices share many properties with real symmetric matrices:
All eigenvalues are real
Eigenvectors belonging to different eigenvalues are mutually orthogonal
They can be diagonalized by a unitary matrix: , where is a unitary matrix and is a real diagonal matrix
Positive Definiteness in Complex Vector Spaces¶
The definiteness of a Hermitian form is analogous to that of a real quadratic form:
A Hermitian form is called:
positive definite if for every nonzero vector
positive semidefinite if for every vector
negative definite and negative semidefinite are defined similarly
indefinite if takes both positive and negative values
As in real space, the definiteness of a Hermitian form can be determined from the eigenvalues of the Hermitian matrix:
Let be a Hermitian matrix. The Hermitian form is positive definite if and only if all eigenvalues of are positive. Similarly, positive semidefiniteness, negative definiteness, negative semidefiniteness, and indefiniteness can also be determined from the eigenvalues.
9.4.4 The Principal Axis Theorem in Complex Space¶
Statement of the Principal Axis Theorem for Complex Vector Spaces¶
The principal axis theorem extends naturally to complex vector spaces:
For a Hermitian form on a complex vector space, where is a Hermitian matrix, there exists a unitary change of variables such that
where are the eigenvalues of (all real) and is a unitary matrix whose columns are orthonormal eigenvectors of .
Note that in complex space the canonical form involves the squared moduli rather than simple squares. This reflects the essential difference between Hermitian forms and quadratic forms.
Comparing Unitary and Orthogonal Transformations¶
Unitary transformations are the counterparts of orthogonal transformations in complex space:
Orthogonal transformation: , where ; it preserves lengths and inner products
Unitary transformation: , where ; it preserves lengths and inner products
Geometrically, orthogonal transformations represent rotations and reflections, while unitary transformations can be regarded as “rotations” in complex space. Neither changes the lengths of vectors or the “angles” between them.
Unitary transformations have the following properties:
They preserve the norm of vectors:
They preserve inner products of vectors:
The determinant has modulus 1:
Applying the Spectral Decomposition to the Principal Axis Theorem¶
The proof of the principal axis theorem for complex vector spaces relies on the spectral decomposition of Hermitian matrices:
The Geometric Meaning of the Canonical Form in Complex Space¶
The positive level set of the canonical form of a Hermitian form corresponds geometrically to a hypersurface in complex space. Although it is hard to visualize directly, it can be understood as follows:
A positive definite Hermitian form (all ) corresponds to a “hyperellipsoid” in complex space
An indefinite Hermitian form (with both positive and negative eigenvalues) corresponds to a “hyperboloid” in complex space
A singular positive semidefinite Hermitian form (all , with at least one equal to 0) corresponds to a degenerate hypersurface
In quantum mechanics, Hermitian forms have an important physical meaning. Hermitian operators (matrices) represent observables, and the principal axis theorem corresponds to representing an observable in terms of its eigenstates. The canonical form is directly related to the expansion of a quantum state in the eigenstate basis.
Consider the Hermitian form on defined by the Hermitian matrix :
Computing the eigenvalues and eigenvectors:
Eigenvalues: ,
Normalized eigenvectors (obtained by computation): , ; they correspond in order to the eigenvalues above and are assembled as column vectors into
By the unitary change of variables , where , the Hermitian form is brought to canonical form:
This represents a “hyperellipsoid” in complex space whose “principal axes” lie along the directions of and .
Real vectors and imaginary vectors are handled differently. For example, consider a purely imaginary vector , where is a real vector, and substitute it into the Hermitian form:
This shows that a Hermitian form behaves in the same way on purely real vectors and purely imaginary vectors, which is a manifestation of unitary invariance.
Consider the Hermitian matrix on :
Verify that is indeed a Hermitian matrix
Compute the eigenvalues and eigenvectors of
Write down the canonical form of the corresponding Hermitian form
Explain the geometric meaning of the canonical form
Solution to Exercise 3
Hermitian property. The off-diagonal entry is , and its conjugate is exactly the entry; the diagonal entries are real. Hence .
Eigenvalues. , which gives and (both real, as the properties of Hermitian matrices require).
Eigenvectors.
: solve , that is, , which gives ; normalizing, .
: solve , which gives ; normalizing, .
Check orthogonality: ✓
Canonical form. Let be the unitary matrix and make the change of variables (that is, ). By the spectral decomposition (by Theorem 13):
This is the canonical form of the Hermitian form—a real-valued quadratic expression with real coefficients and no cross terms.
Geometric meaning. In the unitary coordinates , becomes independent scaling along two mutually orthogonal principal axes (the directions of and ), with the eigenvalues 1 and 4 as scaling factors. Since both eigenvalues are positive, for all , so is positive definite; in complex coordinates () describes an “ellipsoid”-type level set whose shape is determined by the axis ratio . This is completely parallel in structure to the principal axis theorem for real quadratic forms (Theorem 10), reflecting the consistency of the real and complex theories.
Quadratic forms and the principal axis theorem are a model of how linear algebra binds algebra and geometry tightly together. The geometric interpretation of the eigenvalues and eigenvectors of symmetric matrices makes abstract algebraic structures visible and intuitive. The extension from real vector spaces to complex vector spaces shows the consistency and universality of mathematical concepts, and provides a solid mathematical foundation for modern physical theories such as quantum mechanics.
# ==================== Visualizing the quadratic form of a Hermitian matrix ====================
import numpy as np
import matplotlib.pyplot as plt
# Create a 2×2 Hermitian matrix
A_hermitian = np.array([
[2,1+1j],[1-1j,3]
], dtype=complex)
print("\n" + "=" * 50)
print("Hermitian matrix example (complex case)")
print("=" * 50)
# Verify the Hermitian property: A = A^†
is_hermitian = np.allclose(A_hermitian, A_hermitian.conj().T)
print(f"Is the matrix Hermitian: {is_hermitian}")
print("\nOriginal matrix A:")
print(A_hermitian)
# Compute the eigenvalues and eigenvectors
# For symmetric matrices, `eigh` is more efficient and more accurate than `eig`, and returns the eigenvalues in ascending order.
eigenvalues_herm, eigenvectors_herm = np.linalg.eigh(A_hermitian)
print("\nEigenvalues (note: the eigenvalues of a Hermitian matrix must be real):")
print(eigenvalues_herm)
print(f"Imaginary parts of the eigenvalues (should be close to zero): {np.max(np.abs(eigenvalues_herm.imag)):.2e}")
print("\nEigenvector matrix (complex vectors):")
print(eigenvectors_herm)
# Verify that the eigenvectors are unitary (using the conjugate transpose for complex vectors)
unitarity_check = eigenvectors_herm.conj().T @ eigenvectors_herm
print("\nInner-product matrix of the eigenvectors Q^† Q (should be the identity matrix):")
print(np.round(unitarity_check, 10))
# Verify the eigendecomposition: A = Q Λ Q^†
A_herm_reconstructed = eigenvectors_herm @ np.diag(eigenvalues_herm) @ eigenvectors_herm.conj().T
reconstruction_error_herm = np.linalg.norm(A_hermitian - A_herm_reconstructed)
print(f"\nReconstruction error ||A - QΛQ^†||: {reconstruction_error_herm:.2e}")
# Check positive definiteness
is_positive_definite_herm = np.all(eigenvalues_herm > 0)
print(f"\nIs the matrix positive definite: {is_positive_definite_herm}")
if is_positive_definite_herm:
# Compute the square root of the Hermitian matrix
A_herm_sqrt = eigenvectors_herm @ np.diag(np.sqrt(eigenvalues_herm)) @ eigenvectors_herm.conj().T
print("\nSquare root of the Hermitian matrix A^(1/2):")
print(np.round(A_herm_sqrt, 4))
# Verify
verification_herm = A_herm_sqrt @ A_herm_sqrt
print("\nCheck (A^(1/2))² ≈ A:")
print(np.round(verification_herm, 10))
print(f"Check error: {np.linalg.norm(verification_herm - A_hermitian):.2e}")
# Range of the eigenvalues of the Hermitian matrix
print(f"Eigenvalue range of the Hermitian matrix: [{eigenvalues_herm[0]:.4f}, {eigenvalues_herm[-1]:.4f}]")
print(f"Condition number of the Hermitian matrix: {eigenvalues_herm[-1]/eigenvalues_herm[0]:.4f}")
# Compute the eigenvalues and eigenvectors
eigenvalues_2d_herm, eigenvectors_2d_herm = np.linalg.eigh(A_hermitian )
print("\n" + "=" * 50)
print("Visualizing a 2D Hermitian matrix")
print("=" * 50)
print("Eigenvalues (real):", eigenvalues_2d_herm)
print("Eigenvector matrix (complex):")
print(eigenvectors_2d_herm)
# 3. Create a real coordinate grid (x, y are the two components of the vector)
x = np.linspace(-2, 2, 200)
y = np.linspace(-2, 2, 200)
X, Y = np.meshgrid(x, y)
# 4. Compute the quadratic form Q(v) = v^H A v, where v = [x, y]^T is a real vector
# Note: when v is a real vector, v^H A v = v^T Re(A) v
# because the imaginary parts at symmetric positions cancel (1j * xy - 1j * yx = 0)
Z = A_hermitian[0,0].real * X**2 + 2 * A_hermitian[0,1].real * X * Y + A_hermitian[1,1].real * Y**2
# 5. Plot
plt.figure(figsize=(8, 8))
# Draw the contour lines
contour = plt.contour(X, Y, Z, levels=15, cmap='plasma')
plt.clabel(contour, inline=True, fontsize=8)
# Draw the unit-energy ellipse (Q(v) = 1)
ellipse = plt.contour(X, Y, Z, levels=[1.0], colors='red', linewidths=3)
# On the real section z∈R², the principal axes are determined by Re(A)
section_values, section_vectors = np.linalg.eigh(A_hermitian.real)
for i in range(len(section_values)):
v = section_vectors[:, i]
# Scale the eigenvectors to fit the plot
scale = 1 / np.sqrt(section_values[i])
plt.quiver(0, 0, v[0]*scale, v[1]*scale, angles='xy', scale_units='xy',
scale=1, color=['blue', 'green'][i], label=f'EV {i+1} axis')
plt.axhline(0, color='black', lw=1)
plt.axvline(0, color='black', lw=1)
plt.gca().set_aspect('equal')
plt.title('Real Section: z in R^2\n$Q(x,y) = x^T Re(A) x = 1$')
plt.xlabel('$x_1$')
plt.ylabel('$x_2$')
plt.legend()
plt.grid(True, alpha=0.3)
plt.show()
print("\nKey observations:")
print("• A Hermitian matrix guarantees that z^† A z is always real")
print("• The eigenvalues are all real, even though the eigenvectors are complex vectors")
print("• The contour lines show the section restricted to R²; the full C² has four real dimensions")
print(f"• Eigenvalue range of the matrix A: [{eigenvalues_2d_herm[0]:.4f}, {eigenvalues_2d_herm[-1]:.4f}]")This section answered a classical geometric question: why can the level surfaces of a quadratic form be brought to principal-axis form by an orthogonal change of basis? The answer is given by the principal axis theorem: every quadratic form ( a real symmetric matrix) can be brought by an orthogonal change of variables to , where the are the eigenvalues of and the columns of are the corresponding orthonormal eigenvectors—the “principal axis directions.”
The geometric core of this section is that the eigenvalues determine the shape of the quadric surface (ellipsoid or hyperboloid) and the eigenvectors determine the directions of the principal axes. For a nondegenerate in three dimensions, three positive eigenvalues give an ellipsoid, two positive and one negative a hyperboloid of one sheet, one positive and two negative a hyperboloid of two sheets, and three negative the empty set; one cannot look only at the sign of the determinant or ignore the level value. This classification theory unifies results of analytic geometry that had long been scattered.
Over the complex field, the Hermitian form becomes after a unitary change of variables; what appears is the squared modulus , not the square of a real part. This detail reflects the essential difference between Hermitian forms and real quadratic forms, and corresponds directly to the probabilistic interpretation of expanding an observable in its eigenstate basis in quantum mechanics.
As the theoretical end point of Chapter 9, this section fuses the inner product, the spectral decomposition, and geometric classification into one, providing a complete mathematical language for the operator theory of quantum mechanics in Chapter 10.
9.5 Summary¶
Review of the Theoretical Thread¶
This chapter started from a question: what are the costs and the rewards of extending the geometric language of real vector spaces to the complex field? §9.1 laid the foundation of the answer. With conjugation introduced, the complex inner product regains positive definiteness, thereby giving fully meaningful notions of length and angle. Within this framework, Gram-Schmidt orthogonalization provides a systematic algorithm for constructing orthonormal bases, which is precisely the geometric interpretation of the QR decomposition.
§9.2 then asked: which linear transformations preserve this geometric structure exactly as it is? Orthogonal matrices and unitary matrices are the answer—their 2-norm condition number is always 1, making them the most stable class of matrices in numerical computation. Geometrically they correspond to rotations and reflections, which change only the viewing angle, not the shape of the object. An orthogonal matrix itself, however, need not be orthogonally diagonalizable, and this detail leads to another important class defined by a different condition: real symmetric matrices.
§9.3 is the heart of the chapter. Symmetric matrices () and Hermitian matrices (), as self-adjoint operators, are naturally compatible with the inner product structure, so their eigenvalues must be real and eigenvectors belonging to different eigenvalues must be orthogonal. The spectral decomposition theorem (or ) fuses the theory of orthogonal matrices from §9.2 with the eigenvalue theory of §8, and is the most complete matrix structure theorem so far. Positive definiteness adds a further scale to this structure: Sylvester’s criterion, the Cholesky decomposition, and the interpretation of covariance matrices form a bridge between theory and application.
§9.4 closes with quadratic forms. The principal axis theorem says that every quadratic form can be brought to a diagonal canonical form by an orthogonal change of variables, with the eigenvalues determining the shape of the surface and the eigenvectors the directions of the principal axes. Geometrically, this result unifies the classification theory of classical curves and surfaces such as ellipses, hyperbolas, and ellipsoids; over the complex field, the canonical form of a Hermitian form echoes directly the eigenstate expansion of observables in quantum mechanics.
Connections to Other Chapters¶
This chapter deepens and completes the eigenvalue theory of Chapter 8. §8 built the eigenvalue framework for general matrices, but could give only conditional answers to questions such as “when is a matrix diagonalizable” and “when are the eigenvalues real.” This chapter takes symmetry (self-adjointness) as a sufficient condition and obtains a perfect special case of the general theory of Chapter 8: under this condition, diagonalization is not only possible but can be achieved by an orthogonal (unitary) matrix, the eigenvalues are automatically all real, eigenvectors belonging to different eigenvalues are automatically orthogonal, and within the eigenspace of a repeated root the eigenvectors can be orthogonalized. The spectral decomposition theorem of §9.3 can be regarded as the strongest form of Chapter 8’s “diagonalization by similarity” for the class of symmetric matrices.
This chapter also builds a double bridge to the introduction to quantum mechanics in Chapter 10 and the singular value decomposition in Chapter 11. In quantum mechanics, observables correspond to Hermitian operators, the evolution of quantum states corresponds to unitary transformations, and measurement outcomes (which must be real) are precisely the physical realization of the spectral theorem of this chapter. The singular value decomposition (SVD), for its part, can be seen as asking, for a general matrix (non-square, non-symmetric), “what is the geometric decomposition closest to that of a symmetric matrix?”; it generalizes both the orthogonal matrices and the spectral decomposition of this chapter and forms the final unification of the theory of linear algebra.
The Role of This Chapter in the Book¶
This chapter resolves a central question: once a vector space has been given a geometric structure (an inner product), which class of matrices is the most natural and the most regular? The answer given by symmetric and Hermitian matrices is: all real eigenvalues, orthogonal eigenvectors, and a perfect spectral decomposition. This is not an isolated mathematical coincidence but the full expression, at the level of linear algebra, of a deep symmetry—self-adjointness. In nature, a self-adjoint Hamiltonian corresponds to real energies; in a closed system whose Hamiltonian does not depend explicitly on time, the expectation value of the energy is conserved; physical observables correspond to Hermitian operators, and the covariance structure of data corresponds to positive semidefinite matrices—symmetry is everywhere, and the theory of this chapter therefore has a universal significance that goes beyond linear algebra itself.
As you read on with the perspective of this chapter, you will find that the singular value decomposition of Chapter 11 is in fact the “non-square version” of the spectral decomposition of this chapter: the left and right singular vectors are the orthogonal eigenvectors of and , respectively. Once you understand this chapter, you already hold the most essential geometric intuition of the SVD—decomposing any linear transformation into the three steps “rotate—stretch—rotate,” each of which is rooted in the inner-product geometry of this chapter.