Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Open In Colab Binder

In 1807 Joseph Fourier submitted a memoir on heat conduction to the French Academy of Sciences, in which he represented functions by trigonometric series; Lagrange and Laplace then raised objections to the practice of expanding functions in trigonometric series. The Academy made heat conduction the subject of its prize question for 1811, and Fourier won the prize with an expanded version of his memoir; the systematic account appeared in 1822 as Théorie analytique de la chaleur. The method forced mathematicians to rethink a basic question: in what sense can smooth sine and cosine waves approximate a function with corners or jumps? What it means to “have an expansion” cannot be separated from the class of functions and the mode of convergence.

In modern language: in L2[−π,π]L^2[-\pi,\pi] the trigonometric system forms a complete orthogonal system, and the Fourier series converges to the original function in the mean-square sense; pointwise convergence requires a separate discussion.

Looking back on this history, later generations asked: what allowed Fourier to see past intuition? The key lay not in the summation tricks of calculus but in a deep “geometric perpendicularity” hidden among the trigonometric functions:

∫−ππsin⁡(mx)sin⁡(nx) dx=0(m,n∈Z>0, m≠n)\int_{-\pi}^{\pi} \sin(mx)\sin(nx)\,dx = 0 \quad (m,n\in\mathbb{Z}_{>0},\ m\neq n)

This integral is exactly zero, which geometrically means that sine waves of different frequencies are independent of one another and do not interfere. And the Fourier coefficients that give the amplitude of each component,

bn=1π∫−ππf(x)sin⁡(nx) dxb_n = \frac{1}{\pi}\int_{-\pi}^{\pi} f(x)\sin(nx)\,dx

have exactly the same algebraic form as the formula for projecting a geometric arrow onto orthogonal coordinate axes in a finite-dimensional space—the only difference is that the “dot product,” a sum of products of discrete coordinates there, is elevated here into an integral of continuous functions over an interval.

This observation turns “functions are vectors too” from a mere metaphor into a fact: once we choose a function space, an inner product, and the appropriate completeness conditions, we can study length, orthogonality, and projection. Hilbert spaces extend this geometry to infinite dimensions and have become an important language for quantum theory, partial differential equations, and signal processing.

Here, however, Fourier’s story serves only as a mirror for ideas. Properties that are taken for granted in finite dimensions often require additional conditions in infinite-dimensional spaces—convergence, completeness, the domain of an operator, and so on.

The scope of this book is the rigorous and well-controlled setting of finite-dimensional linear spaces. This takes nothing away from the value of the theory in this chapter: in a bare vector space we have only “addition” and “scalar multiplication”; it is a limp world with no length, no angle, and no way to define a perpendicular. It is precisely the introduction of the inner product that gives the space a rigorous measure of length (the norm), a notion of angle, and optimal projection—turning a purely algebraic structure into a solid, concrete geometric stage.

The orthogonality, Gram-Schmidt orthogonalization, orthogonal projection operators, and spectral decomposition of symmetric matrices that we forge in Rn\mathbb{R}^n and Cn\mathbb{C}^n lose none of their core ideas in infinite-dimensional function spaces. In other words, what this chapter develops is by no means an algebraic special case confined to finite dimensions; it is a geometric foundation that spans the discrete and the continuous and, once the analytic conditions are supplied, extends to infinite dimensions.

Chapter Structure and Learning Objectives

The introduction pointed out that “orthogonality of functions” and “orthogonality of vectors” share exactly the same structure. This analogy rests on a precise axiomatic definition—and that is where this chapter begins.

§9.1 lays the foundation: it defines the inner product axiomatically and explains why the complex field forces us to introduce conjugation. Positive definiteness guarantees that “length” is never negative, the central premise that lets the whole geometric framework stand. From the inner product follow naturally the norm, orthogonality, and the Cauchy-Schwarz inequality, the last of which also foreshadows, in an elegant form, the geometric root of the uncertainty principle in Chapter 10. The section closes with Gram-Schmidt orthogonalization—an algorithm that gives a general method for constructing an orthogonal basis in any inner product space and is the geometric essence of the QR decomposition.

With the inner product in hand, §9.2 naturally asks: which linear transformations fully preserve this geometric structure? Orthogonal matrices (over the reals) and unitary matrices (over the complex numbers) are the answer. Their inverse equals their transpose or conjugate transpose, their 2-norm condition number in numerical computation is always 1, and geometrically they are compositions of rotations and reflections—they only rotate the orientation of the coordinate system and do not change distances or angles between vectors.

§9.3 is the heart of the chapter. Self-adjointness (A=A∗\mathbf{A} = \mathbf{A}^*) means that a matrix “respects the inner product,” and from this it follows precisely that the eigenvalues must be real and that eigenvectors belonging to different eigenvalues must be orthogonal. Here the spectral decomposition theorem brings the diagonalization theory of Chapter 8 to its most perfect form: symmetric/Hermitian matrices are not only diagonalizable but can be diagonalized by an orthogonal (unitary) matrix. Positive definiteness further characterizes the “energy character” of such matrices and is an indispensable language in optimization and statistics.

§9.4 lands the theory in geometry through quadratic forms. The principal axis theorem shows that every quadratic form can be brought to a diagonal canonical form by rotating the coordinate system—the eigenvalues determine the shape of the ellipsoid or hyperboloid, and the eigenvectors give the directions of the principal axes. Hermitian forms then extend this picture to the complex field.

Once you have read the whole chapter, your understanding of “why symmetric matrices are special” will deepen from Chapter 8’s “they happen to be diagonalizable” to “self-adjointness gives them an orthogonal structure by nature.” This viewpoint will give you an entirely different depth of field when you meet the introduction to quantum mechanics in Chapter 10 and the singular value decomposition in Chapter 11.

The two symbols in the section titles of this chapter mark their reading priority: ◇ marks an optional extension (for example, §9.2.3 on unitary matrices and quantum computing), which can be skipped without affecting the main line; ◆ marks an advanced proof (for example, the equivalent conditions for positive definiteness), where on a first reading you may grasp the conclusion first and return to its derivation later.


9.1 Inner Product Spaces

Why Do We Need Inner Product Spaces?

In Chapter 2 we saw how the real inner product u⋅v\mathbf{u} \cdot \mathbf{v} connects algebraic computation with geometric measurement—it encodes both the length of a vector (∥v∥=v⋅v\|\mathbf{v}\| = \sqrt{\mathbf{v} \cdot \mathbf{v}}) and the angle between vectors (cos⁡θ=u⋅v∥u∥∥v∥\cos\theta = \frac{\mathbf{u} \cdot \mathbf{v}}{\|\mathbf{u}\|\|\mathbf{v}\|}).

We now face a broader question: how can similar geometric notions be defined in the complex vector space Cn\mathbb{C}^n? Carrying the definition u⋅v=∑iuivi\mathbf{u} \cdot \mathbf{v} = \sum_i u_i v_i of Chapter 2 over to the complex field unchanged runs into a fatal problem:

v⋅v=∑ivi2\mathbf{v} \cdot \mathbf{v} = \sum_i v_i^2

can be negative or even purely imaginary over the complex numbers! For example, take v=[i0]\mathbf{v} = \begin{bmatrix} i \\ 0 \end{bmatrix}; then v⋅v=i2=−1<0\mathbf{v} \cdot \mathbf{v} = i^2 = -1 < 0. This would mean that the “length” of a vector could be negative, which is geometrically meaningless.

The solution is to introduce conjugation:

⟨v∣v⟩=∑ivi‾vi=∑i∣vi∣2≥0⟨ \mathbf{v} | \mathbf{v} ⟩ = \sum_i \overline{v_i} v_i = \sum_i |v_i|^2 \geq 0

This guarantees that “length” is never negative, and so keeps geometric intuition valid. This key insight leads us into the systematic theory of inner product spaces.

An inner product space is a vector space equipped with an additional inner product operation, which brings in geometric notions. When defining an inner product on a complex vector space, we must pay particular attention to how conjugation is handled.

9.1.1 Definition and Properties of the Inner Product

The most common inner product is the standard inner product on Euclidean space:

The inner product brings in the notion of a norm (length):

The inner product also brings in the notions of angle and orthogonality:

Using the Cauchy-Schwarz inequality, we can define the angle between vectors:

9.1.2 Orthonormal Bases

A powerful feature of inner product spaces is that we can define orthogonal bases and orthonormal bases, which are extremely useful both in computation and in theoretical analysis.

In an orthonormal basis, representing and manipulating vectors becomes remarkably simple:

9.1.3 The Gram-Schmidt Orthogonalization Method

Given a set of linearly independent (complex) vectors, how can we turn it into an orthogonal basis or an orthonormal basis? The answer is the Gram-Schmidt orthogonalization method (the Gram-Schmidt process).


The Gram-Schmidt process can be expressed uniformly in matrix form.

§2.5 already demonstrated in code a first version of orthogonalizing two vectors; the definition in §9.1.3 extends the same idea to kk vectors and expresses it uniformly in the language of matrices: in V=QR\mathbf{V} = \mathbf{QR}, the columns of Q\mathbf{Q} are the normalized results of removing projections step by step, and R\mathbf{R} records the coefficients of each projection.

The QR decomposition applies column operations to V\mathbf{V}, so Q\mathbf{Q} and V\mathbf{V} have the same image. This property is why the QR decomposition is widely used in numerical methods such as solving least-squares problems and computing eigenvalues.

The theory of inner product spaces provides a unified framework that systematically extends geometric notions—orthogonality, length, and angle—to abstract vector spaces. Within this framework, Gram-Schmidt orthogonalization gives a concrete algorithm for constructing orthogonal and orthonormal bases, which is of fundamental importance both for building orthogonal decompositions in theoretical analysis and for solving linear systems stably in numerical computation.

9.2 Orthogonal Matrices and Unitary Matrices

Orthogonal matrices and unitary matrices are special matrices that preserve the lengths and inner products of vectors, and they play an important role in many applications. Orthogonal matrices correspond to distance-preserving transformations of real vector spaces, and unitary matrices are their natural generalization to complex vector spaces. Both classes of matrices play central roles in numerical computation, quantum mechanics, signal processing, and other fields.

9.2.1 Definitions and Basic Properties of Orthogonal and Unitary Matrices

Definitions

An orthogonal matrix is the special case of a unitary matrix over the real numbers—when all the entries of a matrix are real, the conjugate transpose reduces to the transpose, and the definition of a unitary matrix reduces naturally to that of an orthogonal matrix.

Basic Properties

Orthogonality of Columns and Rows

The columns and the rows of an orthogonal or unitary matrix each form an orthonormal basis:

Examples of Orthogonal Matrices

Every orthogonal matrix can be written as a product of several basic orthogonal transformations (such as rotations and reflections).

9.2.2 The Geometric Meaning of Orthogonal and Unitary Transformations

Orthogonal and unitary transformations preserve the lengths of vectors and the angles between them, which gives them a clear geometric interpretation.

Definition of Orthogonal Transformations and Their Equivalence with Orthogonal Matrices

In a finite-dimensional Euclidean space, orthogonal transformations correspond one-to-one to orthogonal matrices:

Geometric Decomposition of Orthogonal Transformations

The geometric meaning of an orthogonal transformation depends on its eigenvalues and eigenvectors. On a real vector space, an orthogonal transformation can be decomposed into:

  1. Rotations: corresponding to complex pairs of eigenvalues e±iθe^{\pm i\theta}, acting as rotations in the corresponding two-dimensional subspaces

  2. Reflections: corresponding to the eigenvalue -1, acting as reflections in the corresponding directions

  3. The identity: corresponding to the eigenvalue +1, leaving the corresponding directions unchanged

◇9.2.3 Special Properties of Unitary Matrices and Quantum Computing

Unitary matrices are the generalization of orthogonal matrices to the complex field, and they play a central role in quantum mechanics and quantum computing.

The Importance of Unitary Matrices in Quantum Mechanics

This section considers pure states of closed systems and ideal quantum gates. Within this scope, a pure state of a physical system is described by a unit vector (a quantum state) in a complex vector space, and the evolution of the system is described by a unitary transformation. This is because:

  1. Conservation of probability: unitary transformations preserve the norm of vectors, and therefore preserve the normalization condition ∥ψ∥=1\|\psi\| = 1 of quantum states

  2. Reversibility: physical evolution (in the absence of measurement) is reversible, which corresponds to the invertibility of unitary matrices

  3. Preservation of inner products: the inner product between normalized quantum states gives a probability amplitude, whose squared modulus gives the probability of the corresponding projective measurement; unitary transformations preserve inner products and therefore also preserve these probabilities

Common Unitary Matrices: Quantum Gates

Note that the Hadamard gate is a real unitary matrix (and hence also an orthogonal matrix), while the phase gates S,TS,T and Rz(θ)R_z(\theta) for a general angle are non-real unitary matrices; for special angles RzR_z may also be a real orthogonal matrix.

The Quantum Fourier Transform

9.2.4 Numerical Stability and Applications

Orthogonal matrices have important stability advantages in matrix computations.

Why Orthogonal Matrices Are Numerically Stable

The condition number of an orthogonal matrix, κ2(Q)=∥Q∥2∥Q−1∥2=1\kappa_2(\mathbf{Q}) = \|\mathbf{Q}\|_2\|\mathbf{Q}^{-1}\|_2 = 1, is the smallest possible value among all matrices. This means:

  1. Errors are not amplified: when we compute Qx\mathbf{Qx}, any rounding error in the input x\mathbf{x} is not amplified

  2. Numerical accuracy is maintained: in iterative algorithms, orthogonal transformations do not amplify the 2-norm of existing errors, although new rounding errors produced by repeated operations may still accumulate

  3. Ill-conditioned problems are avoided: using orthogonal matrices avoids many ill-conditioned situations in numerical linear algebra

Main Applications

Orthogonal and unitary matrices play central roles in the following fields:

  1. Numerical linear algebra:

    • QR decomposition: factoring a matrix into the product of an orthogonal matrix and an upper triangular matrix

    • Singular value decomposition (SVD): the subject of Chapter 11

    • Iterative methods: such as Arnoldi iteration and the Lanczos method

  2. Computer graphics:

    • Rotation, scaling, and translation of 3D objects

    • Viewing transformations and projection

    • Interpolation in animation (SLERP: spherical linear interpolation)

  3. Signal processing:

    • The discrete Fourier transform (DFT) and its fast algorithm (FFT)

    • The discrete cosine transform (DCT): the basis of JPEG compression

    • Wavelet transforms: multiscale signal analysis

  4. Quantum computing:

    • Quantum gate operations (such as the various unitary matrices introduced in this section)

    • Quantum algorithm design (such as Shor’s algorithm and Grover’s algorithm)

    • Simulation of quantum state evolution

  5. Data analysis and machine learning:

    • Principal component analysis (PCA): discussed in detail in Chapter 11

    • Independent component analysis (ICA)

    • The whitening transformation

Orthogonal and unitary matrices are a natural extension of the theory of inner product spaces. They preserve inner products, lengths, and angles, and geometrically they correspond to rigid motions (rotations and reflections). An orthogonal transformation does not change the geometric structure of space, only the angle from which we view it. Understanding the properties of orthogonal transformations is essential for understanding many advanced concepts and applications in linear algebra; in particular, it lays the foundation for the spectral theorem for symmetric matrices in the next section, the introduction to quantum mechanics in Chapter 10, and the singular value decomposition in Chapter 11.

9.3 Symmetric Matrices and Hermitian Matrices

Symmetric matrices and Hermitian matrices are two special and important classes of matrices in linear algebra. Symmetric matrices play a central role in real vector spaces, and Hermitian matrices are the natural generalization of symmetric matrices to complex vector spaces. Both classes hold a fundamental place in physics, engineering, and data analysis. Their eigenvalues are all real, and an orthonormal basis of eigenvectors can be chosen for them, which gives them unique advantages in representing physical systems, analyzing the structure of data, and designing algorithms.

9.3.1 Eigenvalues and Eigenvectors of Symmetric and Hermitian Matrices

Definitions of Symmetric and Hermitian Matrices

Hermitian matrices are the natural generalization of symmetric matrices; for real matrices the two notions coincide (because the conjugate of a real number is the number itself).

Properties of the Eigenvalues and Eigenvectors

The eigenvalues and eigenvectors of symmetric and Hermitian matrices have special properties:

The Spectral Decomposition Theorem

Unlike a general matrix, a symmetric matrix never has missing eigenvectors.

9.3.2 Positive Definite and Positive Semidefinite Matrices

Definition of Positive Definiteness

Equivalent Conditions for Positive Definiteness

Positive definiteness can be characterized in several ways:

Positive semidefinite matrices have similar equivalent conditions.

The comparison with the positive definite version is as follows:

ConditionPositive definitePositive semidefinite
(2) Eigenvalues>0> 0≥0\geq 0
(3) Principal minorsLeading principal minors >0> 0All principal minors ≥0\geq 0
(4) CholeskyDiagonal entries >0> 0, decomposition uniqueDiagonal entries ≥0\geq 0, decomposition not necessarily unique
(5) Square-matrix factorizationC\mathbf{C} nonsingularC\mathbf{C} may be singular
(6) General factorizationColumns of B\mathbf{B} linearly independentNo restriction on B\mathbf{B}

Properties and Applications of Positive Definite Matrices

Positive definite matrices have many important properties:

  1. A positive definite matrix is invertible, and its inverse is also positive definite

  2. The main diagonal entries of a positive definite matrix are all positive

  3. The determinant of a positive definite matrix is positive

  4. The sum of two positive definite matrices is positive definite

  5. The Schur complement of a positive definite matrix is also positive definite

A positive definite matrix can define an inner product on a vector space:

9.3.3 Positive Definite Hermitian Matrices and Their Applications

Definition of Positive Definite Hermitian Matrices

For complex vector spaces, the notion of positive definiteness extends naturally:

It should be emphasized that although the entries of A\mathbf{A} may be complex, the value of the expression z∗Az\mathbf{z}^* \mathbf{A}\mathbf{z} is always real (because A\mathbf{A} is a Hermitian matrix).

Equivalent Conditions for Positive Definite Hermitian Matrices

A positive definite Hermitian matrix can define an inner product on a complex vector space:

Applications in Signal Processing

In signal processing and communications, positive definite Hermitian matrices are also widely used:

  1. Covariance matrices: the autocorrelation matrix of a signal is a positive semidefinite Hermitian matrix; it is positive definite only when there is no direction with zero second moment

  2. Multiple-input multiple-output (MIMO) systems: the Gram matrix H∗H\mathbf{H}^*\mathbf{H} of the channel matrix H\mathbf{H} is a positive semidefinite Hermitian matrix; it is positive definite only when H\mathbf{H} has full column rank. Calling it a covariance matrix requires a separately specified stochastic model

  3. Beamforming and spatial filtering: optimization based on positive definite Hermitian matrices

9.4 Quadratic Forms and the Principal Axis Theorem

A quadratic form is a linear combination of squares and products of variables, widely used in optimization, statistics, quantum mechanics, computer graphics, and other fields. The principal axis theorem shows that a quadratic form can be simplified to a standard form by a suitable change of coordinates, a process that geometrically corresponds to finding the principal directions of a quadric surface. This section discusses the theory of quadratic forms and its applications in real and complex spaces in full, revealing the essential geometric meaning of symmetric and Hermitian matrices.

9.4.1 Definition and Basic Properties of Quadratic Forms

Definition of a Quadratic Form

Geometrically, quadratic forms correspond to conics and quadric surfaces, such as ellipses, hyperbolas, and ellipsoids. In physics, quadratic forms can represent key quantities such as energy functions and moments of inertia.

Quadratic Forms and Symmetric Matrices

Every quadratic form can be represented by a symmetric matrix:

Classifying Quadratic Forms by Definiteness

Quadratic forms can be classified according to the signs of the values they take:

The definiteness of a quadratic form can be determined from the eigenvalues of the corresponding symmetric matrix:

9.4.2 Diagonalization of Quadratic Forms in Real Space and the Principal Axis Theorem

Algebraic Derivation of the Diagonalization

By a suitable change of coordinates, every quadratic form can be simplified to a form containing only square terms:

This theorem is in fact an application of the spectral theorem for real symmetric matrices. It shows that every real symmetric matrix can be diagonalized by an orthogonal matrix:

A=PΛP⊤\mathbf{A} = \mathbf{P}\mathbf{\Lambda}\mathbf{P}^\top

where Λ=diag(λ0,λ1,…,λn−1)\mathbf{\Lambda} = \text{diag}(\lambda_0, \lambda_1, \ldots, \lambda_{n-1}) is the diagonal matrix of eigenvalues.

This theorem has an important geometric meaning:

  1. The eigenvalues determine how much the quadric surface “stretches” along each principal axis

  2. The eigenvectors determine the principal directions (principal axes) of the quadric surface

  3. The orthogonal transformation can be regarded as a rotation of the coordinate system that aligns the new coordinate system with the principal axes of the quadric surface

The Law of Inertia for Quadratic Forms and the Classification of Quadric Surfaces

Although a quadratic form can be given different representations by different changes of coordinates, some properties are invariant:

The indices of inertia are directly related to the eigenvalues of the symmetric matrix: the positive index of inertia equals the number of positive eigenvalues, the negative index of inertia equals the number of negative eigenvalues, and the zero index of inertia equals the number of zero eigenvalues.

Below we classify only the positive level surfaces Q(x)=1Q(\mathbf{x})=1 of nondegenerate quadratic forms in three-dimensional real space, with no linear terms.

9.4.3 Hermitian Forms in Complex Space

From Quadratic Forms to Hermitian Forms

In a complex vector space, the notion of a quadratic form must be extended to that of a Hermitian form in order to keep its positive definiteness and geometric meaning:

The values of a Hermitian form are always real, which is guaranteed by the properties of Hermitian matrices. In fact, for every complex vector z\mathbf{z}, the Hermitian form H(z)=z∗AzH(\mathbf{z}) = \mathbf{z}^* \mathbf{A}\mathbf{z} satisfies

H(z)=H(z)‾H(\mathbf{z}) = \overline{H(\mathbf{z})}

which guarantees that H(z)H(\mathbf{z}) is real.

Comparing Hermitian Matrices and Symmetric Matrices

Hermitian matrices (defined in Definition 10) are the natural extension of symmetric matrices to complex vector spaces:

Hermitian matrices share many properties with real symmetric matrices:

  1. All eigenvalues are real

  2. Eigenvectors belonging to different eigenvalues are mutually orthogonal

  3. They can be diagonalized by a unitary matrix: A=UΛU∗\mathbf{A} = \mathbf{U}\mathbf{\Lambda}\mathbf{U}^*, where U\mathbf{U} is a unitary matrix and Λ\mathbf{\Lambda} is a real diagonal matrix

Positive Definiteness in Complex Vector Spaces

The definiteness of a Hermitian form is analogous to that of a real quadratic form:

As in real space, the definiteness of a Hermitian form can be determined from the eigenvalues of the Hermitian matrix:

9.4.4 The Principal Axis Theorem in Complex Space

Statement of the Principal Axis Theorem for Complex Vector Spaces

The principal axis theorem extends naturally to complex vector spaces:

Note that in complex space the canonical form involves the squared moduli ∣wi∣2=wi‾wi|w_i|^2 = \overline{w_i} w_i rather than simple squares. This reflects the essential difference between Hermitian forms and quadratic forms.

Comparing Unitary and Orthogonal Transformations

Unitary transformations are the counterparts of orthogonal transformations in complex space:

  1. Orthogonal transformation: x=Py\mathbf{x} = \mathbf{P}\mathbf{y}, where P⊤P=I\mathbf{P}^\top \mathbf{P} = \mathbf{I}; it preserves lengths and inner products

  2. Unitary transformation: z=Uw\mathbf{z} = \mathbf{U}\mathbf{w}, where U∗U=I\mathbf{U}^* \mathbf{U} = \mathbf{I}; it preserves lengths and inner products

Geometrically, orthogonal transformations represent rotations and reflections, while unitary transformations can be regarded as “rotations” in complex space. Neither changes the lengths of vectors or the “angles” between them.

Applying the Spectral Decomposition to the Principal Axis Theorem

The proof of the principal axis theorem for complex vector spaces relies on the spectral decomposition of Hermitian matrices:

The Geometric Meaning of the Canonical Form in Complex Space

The positive level set H(z)=1H(\mathbf{z})=1 of the canonical form H(z)=∑i=0n−1λi∣wi∣2H(\mathbf{z}) = \sum_{i=0}^{n-1} \lambda_i |w_i|^2 of a Hermitian form corresponds geometrically to a hypersurface in complex space. Although it is hard to visualize directly, it can be understood as follows:

  1. A positive definite Hermitian form (all λi>0\lambda_i > 0) corresponds to a “hyperellipsoid” in complex space

  2. An indefinite Hermitian form (with both positive and negative eigenvalues) corresponds to a “hyperboloid” in complex space

  3. A singular positive semidefinite Hermitian form (all λi≥0\lambda_i \geq 0, with at least one equal to 0) corresponds to a degenerate hypersurface

Real vectors and imaginary vectors are handled differently. For example, consider a purely imaginary vector z=ix\mathbf{z} = i\mathbf{x}, where x\mathbf{x} is a real vector, and substitute it into the Hermitian form:

H(ix)=(ix)∗A(ix)=x⊤(−i)(A)(ix)=x⊤Ax=H(x)H(i\mathbf{x}) = (i\mathbf{x})^* \mathbf{A}(i\mathbf{x}) = \mathbf{x}^\top (-i)(\mathbf{A})(i\mathbf{x}) = \mathbf{x}^\top \mathbf{A}\mathbf{x} = H(\mathbf{x})

This shows that a Hermitian form behaves in the same way on purely real vectors and purely imaginary vectors, which is a manifestation of unitary invariance.

Quadratic forms and the principal axis theorem are a model of how linear algebra binds algebra and geometry tightly together. The geometric interpretation of the eigenvalues and eigenvectors of symmetric matrices makes abstract algebraic structures visible and intuitive. The extension from real vector spaces to complex vector spaces shows the consistency and universality of mathematical concepts, and provides a solid mathematical foundation for modern physical theories such as quantum mechanics.


9.5 Summary

Review of the Theoretical Thread

This chapter started from a question: what are the costs and the rewards of extending the geometric language of real vector spaces to the complex field? §9.1 laid the foundation of the answer. With conjugation introduced, the complex inner product ⟨u∣v⟩=u∗v⟨\mathbf{u}|\mathbf{v}⟩=\mathbf{u}^*\mathbf{v} regains positive definiteness, thereby giving Cn\mathbb{C}^n fully meaningful notions of length and angle. Within this framework, Gram-Schmidt orthogonalization provides a systematic algorithm for constructing orthonormal bases, which is precisely the geometric interpretation of the QR decomposition.

§9.2 then asked: which linear transformations preserve this geometric structure exactly as it is? Orthogonal matrices and unitary matrices are the answer—their 2-norm condition number is always 1, making them the most stable class of matrices in numerical computation. Geometrically they correspond to rotations and reflections, which change only the viewing angle, not the shape of the object. An orthogonal matrix itself, however, need not be orthogonally diagonalizable, and this detail leads to another important class defined by a different condition: real symmetric matrices.

§9.3 is the heart of the chapter. Symmetric matrices (A=A⊤\mathbf{A}=\mathbf{A}^\top) and Hermitian matrices (A=A∗\mathbf{A}=\mathbf{A}^*), as self-adjoint operators, are naturally compatible with the inner product structure, so their eigenvalues must be real and eigenvectors belonging to different eigenvalues must be orthogonal. The spectral decomposition theorem A=PΛP⊤\mathbf{A}=\mathbf{P}\mathbf{\Lambda}\mathbf{P}^\top (or UΛU∗\mathbf{U}\mathbf{\Lambda}\mathbf{U}^*) fuses the theory of orthogonal matrices from §9.2 with the eigenvalue theory of §8, and is the most complete matrix structure theorem so far. Positive definiteness adds a further scale to this structure: Sylvester’s criterion, the Cholesky decomposition, and the interpretation of covariance matrices form a bridge between theory and application.

§9.4 closes with quadratic forms. The principal axis theorem says that every quadratic form can be brought to a diagonal canonical form by an orthogonal change of variables, with the eigenvalues determining the shape of the surface and the eigenvectors the directions of the principal axes. Geometrically, this result unifies the classification theory of classical curves and surfaces such as ellipses, hyperbolas, and ellipsoids; over the complex field, the canonical form ∑λi∣wi∣2\sum\lambda_i|w_i|^2 of a Hermitian form echoes directly the eigenstate expansion of observables in quantum mechanics.

Connections to Other Chapters

This chapter deepens and completes the eigenvalue theory of Chapter 8. §8 built the eigenvalue framework for general matrices, but could give only conditional answers to questions such as “when is a matrix diagonalizable” and “when are the eigenvalues real.” This chapter takes symmetry (self-adjointness) as a sufficient condition and obtains a perfect special case of the general theory of Chapter 8: under this condition, diagonalization is not only possible but can be achieved by an orthogonal (unitary) matrix, the eigenvalues are automatically all real, eigenvectors belonging to different eigenvalues are automatically orthogonal, and within the eigenspace of a repeated root the eigenvectors can be orthogonalized. The spectral decomposition theorem of §9.3 can be regarded as the strongest form of Chapter 8’s “diagonalization by similarity” for the class of symmetric matrices.

This chapter also builds a double bridge to the introduction to quantum mechanics in Chapter 10 and the singular value decomposition in Chapter 11. In quantum mechanics, observables correspond to Hermitian operators, the evolution of quantum states corresponds to unitary transformations, and measurement outcomes (which must be real) are precisely the physical realization of the spectral theorem of this chapter. The singular value decomposition (SVD), for its part, can be seen as asking, for a general matrix (non-square, non-symmetric), “what is the geometric decomposition closest to that of a symmetric matrix?”; it generalizes both the orthogonal matrices and the spectral decomposition of this chapter and forms the final unification of the theory of linear algebra.

The Role of This Chapter in the Book

This chapter resolves a central question: once a vector space has been given a geometric structure (an inner product), which class of matrices is the most natural and the most regular? The answer given by symmetric and Hermitian matrices is: all real eigenvalues, orthogonal eigenvectors, and a perfect spectral decomposition. This is not an isolated mathematical coincidence but the full expression, at the level of linear algebra, of a deep symmetry—self-adjointness. In nature, a self-adjoint Hamiltonian corresponds to real energies; in a closed system whose Hamiltonian does not depend explicitly on time, the expectation value of the energy is conserved; physical observables correspond to Hermitian operators, and the covariance structure of data corresponds to positive semidefinite matrices—symmetry is everywhere, and the theory of this chapter therefore has a universal significance that goes beyond linear algebra itself.

As you read on with the perspective of this chapter, you will find that the singular value decomposition of Chapter 11 is in fact the “non-square version” of the spectral decomposition of this chapter: the left and right singular vectors are the orthogonal eigenvectors of AA∗\mathbf{A}\mathbf{A}^* and A∗A\mathbf{A}^*\mathbf{A}, respectively. Once you understand this chapter, you already hold the most essential geometric intuition of the SVD—decomposing any linear transformation into the three steps “rotate—stretch—rotate,” each of which is rooted in the inner-product geometry of this chapter.

Concept Map