Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Open In Colab Binder

If you are seeing matrix multiplication for the first time, it is almost impossible to find it “natural.” When two matrices are multiplied, why must the rows of the first matrix be multiplied by the columns of the second and the products then summed term by term? Why not simply multiply the entries in corresponding positions, as we do for addition? The definition looks like an arbitrary rule laid down on a whim by some mathematician, and one suspects that some secret is being concealed behind it.

The suspicion is entirely reasonable—but the secret runs deeper than you expect: this seemingly odd definition of matrix multiplication took more than a hundred years, and the thought and clashing ideas of several first-rate mathematicians, before it finally settled into place. In form it is a convention of operation; in substance it is the most basic, and the most profound, algebraic cornerstone of linear algebra.

The story has to begin at the end of the eighteenth century. Leibniz, Gauss, and Binet, while solving systems of linear equations and studying determinants, all arranged coefficients in rectangular arrays of numbers, and in the course of their computations all of them implicitly carried out elimination steps resembling matrix multiplication. But nobody at the time stopped to ask: “What sort of thing is this array itself?” In their eyes the array was merely scaffolding to assist a computation; once the house was built and the equations solved, the scaffolding came down.

The breakthrough arrived unexpectedly in 1843. The Irish mathematician William Rowan Hamilton had been trying to extend the complex numbers to three-dimensional space, and had pondered the problem for ten years without success. On October 16 of that year, while walking across Brougham Bridge in Dublin, he saw in a flash that rotations of space required four dimensions, and that multiplication would have to give up commutativity altogether. On the spot he carved the fundamental formula i2=j2=k2=ijk=−1i^2 = j^2 = k^2 = ijk = -1 into the stone of the bridge. This was the first time in the history of mathematics that a scholar seriously declared: the failure of multiplication to commute is not a flaw but an intrinsic feature of rotations in real space.

At almost the same time, in 1844, the twenty-one-year-old German prodigy Gotthold Eisenstein, studying linear transformations, had already begun to use a single letter to stand for an entire system of linear substitutions; he wrote down rules resembling matrix multiplication and pointed out explicitly that this operation is not commutative. Gauss held him in the highest esteem, even ranking him alongside Niels Henrik Abel—sadly, both were geniuses who lit up the mathematical sky and died young: Abel of tuberculosis at twenty-six, and Eisenstein of illness in Berlin at twenty-nine (1852). His profound ideas lay scattered here and there, never integrated into a systematic algebra.

It was Arthur Cayley who took the decisive step. Inspired by the symbolic treatment of substitutions in the work of Eisenstein and others, he published his epoch-making A Memoir on the Theory of Matrices in 1858. Cayley did something none of his predecessors had done: he stopped treating the array of numbers as a tool to assist computation and treated it instead as an algebraic entity with operations of its own. He declared explicitly: “I wish to represent a transformation by a single symbol.” The seemingly awkward “row times column” definition of matrix multiplication was by no means set arbitrarily—it is uniquely determined by the structure of composing substitutions, namely the requirement that the combined effect of first applying the linear mapping A\mathbf{A} and then applying the linear mapping B\mathbf{B} must equal exactly the matrix product BA\mathbf{BA}.

Sixty-seven years later, in 1925, the young Heisenberg, working on quantum mechanics on the island of Helgoland and guided by physical intuition in computing the intensities of radiation in atomic spectra, ran independently into the same wall. He found that to fit the spectral regularities he had to adopt a strange algebra in which the order of multiplication changes the result (XY≠YXXY \neq YX). In a letter to Wolfgang Pauli he wrote uneasily that this strange kind of multiplication filled him with dread. Heisenberg did not even know what this algebra was—until his mentor Max Born saw the paper and recognized, to his surprise, that it was exactly the matrix algebra Cayley had built up completely in pure mathematics seventy years before! Only then did physicists, who had long ignored progress in pure algebra, turn back in a scramble and open the yellowed pages of the mathematicians.

More intriguing still is this: even at that point, the geometric face of the matrix as a “linear transformation of space” remained absent from classical physics for a long time—and the absence runs deeper than one would imagine. Newton’s Mathematical Principles of Natural Philosophy of 1687 is from beginning to end a combination of Euclidean geometric proportion and a newborn calculus; Lagrange made mechanics thoroughly algebraic in his Analytical Mechanics of 1788, and one may search his manuscript from end to end without finding a single matrix enclosed in square brackets; Euler, deriving the equations for the rotation of a rigid body, treated the nine components of the inertia tensor as independent coefficients and struggled through intricate trigonometric substitutions rather than using the concise multiplication of matrices. The whole magnificent edifice of classical mechanics—the Newtonian, Lagrangian, and Hamiltonian systems—was finished and roofed over before the notion of a matrix appeared.

The irony is that when students in engineering or physics today solve problems in classical mechanics and vibration, matrices are simply everywhere. That is because only after quantum mechanics broke the deadlock in 1925 did physicists turn back and rewrite the long-finished classical theory in the exceptionally tidy modern language of matrices. First the avant-garde microscopic world of quanta put matrices to work; then matrices returned the favor by beautifying and unifying macroscopic classical mechanics.

The path of historical development was a winding one: from the early arithmetic of elimination in systems of equations, to the purely symbolic matrix algebra of the mid-nineteenth century, to the adoption under duress by quantum mechanics in 1925, which in turn set off a thorough recasting of classical mechanics in matrix form; and only in the 1930s, through the axiomatic foundations laid by von Neumann and others in operator theory and infinite-dimensional Hilbert space, did the “linear mapping” that matrix multiplication represents finally attain its full elevation into modern geometry.

A convention of operation that looks unremarkable took a hundred and seventy years to travel from Leibniz’s tables of coefficients to Cayley’s matrix algebra, and only with the axiomatized spaces of modern mathematics did it finally show its geometric hand. If matrix multiplication feels awkward the first time you meet it, that is not because you lack mathematical intuition; it is because you are touching, for the first time, a mathematical reality that puzzled the finest minds in history for a long time.


Chapter Structure and Learning Objectives

The story told in the introduction points in the end to a single question: why is matrix multiplication defined in the unnatural way of “multiplying rows by columns and summing”? The five sections of this chapter answer that question together, and build up the whole geometric vocabulary of matrices along the way.

§3.1 begins with rotation in two dimensions. The angle addition formulas supply clues to the matrix and to the rule of multiplication; it is by extending the requirement to all compatible linear mappings, so that (AB)x=A(Bx)(\mathbf{AB})\mathbf{x}=\mathbf{A}(\mathbf{B}\mathbf{x}) holds for every input, that standard matrix multiplication is uniquely determined. Rotation is a concrete starting point for understanding this general requirement.

§3.2 generalizes the matrix from a special case about rotations into a general algebraic object, and builds the complete system of operations: addition, multiplication, transposition, and the inverse. The four interpretations of matrix multiplication (the inner product, combinations of columns, combinations of rows, and a sum of rank-one matrices) are equivalent to one another, and you will find yourself using different ones on different occasions. The central result of this section is the representation theorem for linear mappings: the column vectors of a matrix are exactly where the standard basis vectors land under the mapping. This reduces understanding “the effect of a mapping on an entire space” to reading off a few column vectors, and it is the insight in this chapter most worth chewing over again and again.

§3.3 uses this tool to take systematic stock of the linear mappings commonly met in two and three dimensions: the zero mapping, scaling, projection, reflection, and rotation, each with its own definite matrix form and its own determinant signature. Rodrigues’ rotation formula in three dimensions and the gimbal lock of the Euler angles are the high point of the section, and the noncommutativity of composed rotations is exhibited here in concrete form.

§3.4 introduces a quantitative tool alongside these geometric descriptions: the determinant measures how much a mapping scales volume, ∣det⁡A∣|\det \mathbf{A}| is the volume scale factor, and the sign records whether orientation has been reversed. A determinant of zero means that space has been compressed into a lower dimension and that the mapping has lost its invertibility.

§3.5 takes the determinant criterion as a key and interprets the solution of the linear system Ax=b\mathbf{Ax} = \mathbf{b} as “finding a preimage of the mapping.” When det⁡(A)≠0\det(\mathbf{A}) \neq 0, the inverse mapping exists and the solution is given by the formula for the inverse matrix; when the determinant is zero, the system may have no solution at all or infinitely many, and that suspense is left for Chapter 6 to treat systematically.

By the end of this chapter, matrix multiplication will no longer look to you like an algorithm to be memorized, but like a faithful algebraic translation of the geometric operation of composition; and every column vector of a matrix will be a new direction into which space is sent by that mapping.

3.1 From Rotation Mappings to the Notion of a Matrix

This section introduces the notion of a matrix by way of linear mappings in two dimensions; we shall come to understand how a matrix arises naturally by working through a concrete rotation mapping.

3.1.1 Introducing the Matrix through Rotation Mappings

We first consider a vector x=[xy]\mathbf{x} = \begin{bmatrix} x \\ y \end{bmatrix} in the two-dimensional plane R2\mathbb{R}^2, which may be expressed in polar form as:

x=r[cos⁡γsin⁡γ]\mathbf{x} = r \begin{bmatrix} \cos\gamma \\ \sin\gamma \end{bmatrix}

Here r=∥x∥r = \|\mathbf{x}\| is the length of the vector and γ\gamma is the directed polar angle of the nonzero vector, measured counterclockwise from the positive xx half-axis and understood modulo 2π2\pi; for the zero vector the polar angle may be chosen arbitrarily.

Now we consider rotating this vector counterclockwise through α\alpha degrees. By the angle addition formulas for the trigonometric functions, the rotated vector is:

xnew=r[cos⁡(α+γ)sin⁡(α+γ)]=r[cos⁡αcos⁡γ−sin⁡αsin⁡γsin⁡αcos⁡γ+cos⁡αsin⁡γ]\mathbf{x}_{\text{new}} = r \begin{bmatrix} \cos(\alpha+\gamma) \\ \sin(\alpha+\gamma) \end{bmatrix} = r \begin{bmatrix} \cos \alpha \cos \gamma - \sin \alpha \sin \gamma \\ \sin \alpha \cos \gamma + \cos \alpha \sin \gamma \end{bmatrix}

Substituting rcos⁡γ=xr\cos\gamma = x and rsin⁡γ=yr\sin\gamma = y into the expression above, we obtain:

xnew=[cos⁡α⋅x−sin⁡α⋅ysin⁡α⋅x+cos⁡α⋅y]\mathbf{x}_{\text{new}} = \begin{bmatrix} \cos\alpha \cdot x - \sin \alpha \cdot y \\ \sin \alpha \cdot x + \cos \alpha \cdot y \end{bmatrix}

We note several important features:

  1. Since α\alpha is a fixed constant, the final expression is a linear function of the vector [xy]\begin{bmatrix} x \\ y \end{bmatrix}

  2. Since [xy]\begin{bmatrix} x \\ y \end{bmatrix} stands for an arbitrary point of the plane, this mapping rotates every point of the plane through α\alpha degrees

We may separate the constants from the variables to obtain a more compact matrix form:

[xnewynew]=[cos⁡α−sin⁡αsin⁡αcos⁡α][xy]\begin{bmatrix} x_{\text{new}} \\ y_{\text{new}} \end{bmatrix} = \begin{bmatrix} \cos\alpha & - \sin \alpha \\ \sin\alpha & \cos \alpha \end{bmatrix} \begin{bmatrix} x \\ y \end{bmatrix}

We extract the 2×22 \times 2 array of numbers in the middle, call it a matrix (or a square matrix), and give it the symbol A\mathbf{A}:

A=[cos⁡α−sin⁡αsin⁡αcos⁡α]\mathbf{A} = \begin{bmatrix} \cos\alpha & - \sin \alpha \\ \sin\alpha & \cos \alpha \end{bmatrix}

The original mapping may now be written compactly as:

xnew=Ax=TA(x)\mathbf{x}_{\text{new}} = \mathbf{A}\mathbf{x} = T_{\mathbf{A}}(\mathbf{x})

Here TAT_{\mathbf{A}} denotes the mapping by which the matrix A\mathbf{A} acts on a vector.

The question now is: how should the operation between a matrix and a vector be defined so that it yields the result we want?

Observing the structure of the rotation mapping, we find that if the entries of the matrix and the entries of the vector are paired and multiplied in the following way, the original mapping is recovered:

[→→][↓]\begin{bmatrix} \rightarrow \\ \rightarrow \end{bmatrix} \begin{bmatrix} ↓ \end{bmatrix}

Specifically, the inner product of row 0 of the matrix with the vector gives component 0 of the result, and the inner product of row 1 of the matrix with the vector gives component 1. This is the central idea of matrix-vector multiplication.

3.1.2 A Natural Route to Matrix Multiplication

Before giving a rigorous definition of a matrix, let us look more closely at the rotation mapping. Clearly, besides rotating all at once, we may also rotate in two stages. If, for instance, we want to rotate through α+β\alpha + \beta degrees in the end, we may first rotate through α\alpha degrees and then through β\beta degrees.

From the point of view of mappings this is entirely reasonable: the composite of the two mappings ought to agree with the result of carrying out the rotation in one step. By the angle addition formulas, we want:

[cos⁡(α+β)−sin⁡(α+β)sin⁡(α+β)cos⁡(α+β)]\begin{bmatrix} \cos(\alpha+\beta) & - \sin (\alpha+\beta) \\ \sin(\alpha+\beta) & \cos (\alpha+\beta) \end{bmatrix}

Expanding by the angle addition formulas gives:

[cos⁡αcos⁡β−sin⁡αsin⁡β−sin⁡αcos⁡β−cos⁡αsin⁡βsin⁡αcos⁡β+cos⁡αsin⁡βcos⁡αcos⁡β−sin⁡αsin⁡β]\begin{bmatrix} \cos\alpha\cos\beta - \sin\alpha\sin\beta & -\sin\alpha\cos\beta - \cos\alpha\sin\beta \\ \sin\alpha\cos\beta + \cos\alpha\sin\beta & \cos\alpha\cos\beta - \sin\alpha\sin\beta \end{bmatrix}

and this ought to equal the result of first rotating through α\alpha degrees and then through β\beta degrees:

[cos⁡β−sin⁡βsin⁡βcos⁡β][cos⁡α−sin⁡αsin⁡αcos⁡α]\begin{bmatrix} \cos\beta & - \sin \beta \\ \sin\beta & \cos \beta \end{bmatrix} \begin{bmatrix} \cos\alpha & - \sin \alpha \\ \sin\alpha & \cos \alpha \end{bmatrix}

There is an important observation to be made here: note that the two column vectors of the rotation matrix are essentially the same, differing only by a rotation through ninety degrees:

[−sin⁡αcos⁡α]=[cos⁡(α+π/2)sin⁡(α+π/2)]\begin{bmatrix} - \sin \alpha \\ \cos \alpha \end{bmatrix}=\begin{bmatrix} \cos(\alpha+\pi/2) \\ \sin(\alpha+\pi/2) \end{bmatrix}

This means that we may regard matrix multiplication as the matrix on the left acting separately on each of the column vectors of the matrix on the right:

[→→][↓↓]\begin{bmatrix} \rightarrow \\ \rightarrow \end{bmatrix} \begin{bmatrix} ↓ & ↓ \end{bmatrix}

This gives us an important clue for defining matrix multiplication, a notion we shall develop in detail in the sections that follow.

3.1.3 Verifying Rotation Matrix Multiplication in Python

3.2 Introduction to Matrices

The matrix is one of the most basic notions in linear algebra, rich in theoretical significance and in practical application. This section introduces the basic definition of a matrix and the basic operation of multiplication.

3.2.1 The Definition of a Matrix and Its Basic Operations

The rows and the columns of a matrix may be singled out and regarded as vectors in their own right.

We are now in a position to introduce matrix multiplication. Matrix multiplication is one of the most basic and most important operations in linear algebra, and it is closely bound up with linear mappings.

3.2.2 The Two-Way Correspondence between Linear Mappings and Matrices

Much as in the previous section, we may use a matrix to define a linear mapping in the Cartesian coordinate system.

We now face a natural question in the reverse direction: if we start from a linear mapping, can we always find a matrix that represents it?

An important way of understanding a linear mapping is to observe how it acts on the standard basis vectors. As described in 2.1.3 on the standard basis, linear combinations of the standard basis of Rn\mathbb{R}^n generate any vector whatsoever in Rn\mathbb{R}^n. From the properties of a linear mapping it follows that a linear mapping is completely determined by its action on the basis vectors. If we know T(e0),T(e1),…,T(en−1)T(\mathbf{e}_0), T(\mathbf{e}_1), \ldots, T(\mathbf{e}_{n-1}), then we can determine the action of TT on any vector at all. This shows that although a linear mapping is on the face of it a mapping from an infinite set to an infinite set, most of the time it behaves more like a mapping from a finite set to a finite set.

3.2.3 Basic Matrix Operations in Python

3.2.4 ◆Higher-Order Matrix Operations with einsum in Python

Before demonstrating einsum, we first supply two definitions that will be needed below and that recur in later chapters.

3.3 Typical Linear Mappings in Two and Three Dimensions

3.3.1 The Zero Mapping

3.3.2 The Identity Mapping

3.3.3 The Scaling Mapping

3.3.4 The Projection Matrix

Projection is an important notion in linear algebra: it “casts” a vector of a higher-dimensional space onto a lower-dimensional subspace.

Analysis of the Action of a Projection Matrix on the Standard Basis

Let us come to understand the properties of a projection matrix by observing how it acts on the standard basis vectors:

3.3.5 The Reflection Matrix

A reflection mapping flips a figure across some line or plane, just like the image one sees in a mirror.

Analysis of the Action of a Reflection Matrix on the Standard Basis

3.3.6 The Rotation Mapping

We introduced the two-dimensional rotation matrix in detail in Section 3.1. Let us now review it and extend it to the three-dimensional case.

Three-Dimensional Rotation Matrices

Rotation in three-dimensional space is a good deal more complicated than in two dimensions, because we must specify an axis of rotation. The most basic three-dimensional rotations are those about the coordinate axes:

For linear mappings in three-dimensional space the method of analysis is entirely analogous:

Any three-dimensional rotation whatsoever may be expressed as a combination of these three elementary rotations, and this leads to the notion of the Euler angles.

For a rotation about an arbitrary axis, we may use Rodrigues’ rotation formula:

The complete step-by-step computation (including the correspondence between K\mathbf{K} and the cross product of vectors) may be found in the special topic on skew-symmetric matrices in the exercise notebook, Exercise 61.

By means of this method of analysis based on the standard basis vectors, we can systematically understand and classify the various linear mappings, and this lays a solid foundation for the study of more complicated notions of linear algebra to come.

3.3.7 Verifying the Gimbal Lock of the Euler Angles in Python

3.4 Determinants and Mappings of Space

In what we have studied so far we have come to understand the geometric effect of various linear mappings, but one important quantitative tool is still missing—how are we to measure precisely the effect of a mapping on the “size” of space? The determinant is exactly the notion that resolves this question.

3.4.1 The Column Vectors of a Square Matrix Form an Area

Let us begin with a concrete geometric question: if we have a unit square, how does its area change after a linear mapping?

3.4.2 The Column Vectors of a Square Matrix Form a Volume

In the three-dimensional case, what the determinant describes is the volume formed by the column vectors of the square matrix.

Recall from Chapter 2: the scalar triple product a⋅(b×c)\mathbf{a}\cdot(\mathbf{b}\times\mathbf{c}) of three vectors a,b,c\mathbf{a}, \mathbf{b}, \mathbf{c} is numerically equal to the signed volume of the parallelepiped they span (the sign recording the orientation given by the right-hand rule). This conclusion was established by the Chapter 2 exercises Exercise 30 and Exercise 31. Substituting the three column vectors of a 3×33\times 3 matrix into the scalar triple product therefore yields exactly the volume of the parallelepiped they span—the natural generalization to three dimensions of the two-dimensional “signed area.”

3.4.3 Determinants of Some Special Matrices

Let us compute the determinants of some mapping matrices with which we are already familiar:

3.4.4 Checking the Formula for the Determinant in Python

3.4.5 Computing Determinants and Verifying Their Geometric Meaning in Python

3.5 Linear Systems and Linear Mappings

We now examine the relationship between systems of linear equations and some typical linear mappings from two dimensions to two dimensions and from three dimensions to three dimensions. In fact, solving a system of linear equations is nothing other than searching for a preimage of a linear mapping.

3.5.1 Understanding Linear Systems Geometrically

The Need for an Inverse Mapping, and the Notion of an Inverse Matrix

Starting from the system Ax=b\mathbf{Ax} = \mathbf{b}, we naturally ask: how are we to “undo” the mapping A\mathbf{A} in order to find the original vector x\mathbf{x}?

This is like reasoning backward in everyday life:

  • if we know the result of subjecting an object to some transformation, we want to know what it looked like originally

  • if we know the ciphertext, we want to recover the plaintext by an inverse operation

  • if we know the output of a function, we want to find the corresponding input

In linear algebra, this operation of “undoing” is exactly the role played by the inverse matrix.

When the inverse matrix of A\mathbf{A} exists, solving the linear system Ax=b\mathbf{Ax} = \mathbf{b} becomes entirely straightforward:

x=A−1b\mathbf{x} = \mathbf{A}^{-1}\mathbf{b}

In the remainder of this section we concentrate first on the ideal case—that of an invertible square matrix—and give concrete methods for computing the inverse matrix in two and three dimensions. These basic skills will lay important groundwork for understanding the more complicated situations.

3.5.2 The Geometric Meaning of a Two-Dimensional Linear System

3.5.3 The Inverse Matrix and the Solution of a System

When the matrix of the mapping is invertible, we can “undo” the mapping in order to find the original vector:

3.5.4 Generalization to Three Dimensions

A three-dimensional linear system may likewise be written in the matrix form Ax=b\mathbf{Ax} = \mathbf{b}, and so long as det⁡(A)≠0\det(\mathbf{A}) \neq 0 there is a unique solution x=A−1b\mathbf{x} = \mathbf{A}^{-1}\mathbf{b}.

3.5.5 Solving Linear Systems in Python

3.6 Chapter Summary

Review of the Theoretical Thread

This chapter began from rotation in two dimensions, using the angle addition formulas to introduce tables of coefficients and matrix–vector multiplication. The composition of rotations supplies clues to the rule of matrix multiplication; requiring that all compatible linear mappings satisfy (AB)x=A(Bx)(\mathbf{AB})\mathbf{x}=\mathbf{A}(\mathbf{B}\mathbf{x}) is what uniquely determines standard matrix multiplication. Taking x\mathbf{x} to be each standard basis vector in turn determines the product matrix column by column.

§3.2 builds a complete algebraic system on this foundation. The four equivalent interpretations of matrix multiplication (inner products of rows with columns, combinations of columns, combinations of rows, and a sum of rank-one matrices) reveal the rich geometric structure hidden behind the rule of multiplication. The central result is the representation theorem for linear mappings: every linear mapping Rn→Rm\mathbb{R}^n \to \mathbb{R}^m corresponds to exactly one m×nm \times n matrix, and the column vectors of the matrix are precisely the images of the standard basis vectors. This reduces “understanding the behavior of a mapping of infinite dimension” to “reading off where finitely many column vectors go,” and is the epistemological foundation of Chapter 3.

§3.3 takes this as its tool and organizes systematically the matrix forms of the linear mappings commonly met in two and three dimensions. The zero mapping, scaling, projection (P2=P\mathbf{P}^2 = \mathbf{P}), reflection (det⁡=−1\det = -1, R2=I\mathbf{R}^2 = \mathbf{I}), and rotation (det⁡=1\det = 1, preserving distances and angles) each have their own distinct determinant signature. Rodrigues’ formula for three-dimensional rotations and the gimbal lock of the Euler angles reveal further the noncommutativity of composed rotations and the nontrivial topology of the rotation group—foreshadowings of the deeper discussions to come.

§3.4 introduces the determinant as a precise characterization of “the factor by which a mapping scales the size of space”: ∣det⁡A∣|\det \mathbf{A}| measures the scaling of volume, the sign records a reversal of orientation, and det⁡(A)=0\det(\mathbf{A}) = 0 is equivalent to space being compressed into a lower dimension, and hence equivalent to the mapping not being invertible. §3.5 takes this as its criterion and closes the circuit: solving the system Ax=b\mathbf{Ax} = \mathbf{b} is finding a preimage of the mapping TAT_{\mathbf{A}}, and det⁡(A)≠0\det(\mathbf{A}) \neq 0 guarantees that the preimage exists and is unique, given by the formula for the inverse matrix.

Connections to Other Chapters

This chapter takes up the geometric understanding of vectors in Rn\mathbb{R}^n from Chapter 2—length, angle, linear dependence—and converts those geometric notions into the language of matrices. In particular, the linear combination of vectors becomes in this chapter the notion of “the space spanned by the column vectors of a matrix,” laying the linguistic groundwork for the range and the kernel to come.

This chapter prepares directly for the abstract linear spaces of Chapter 4. The principle established in §3.2 that “a linear mapping is completely determined by the images of the basis vectors” is precisely the central tool of Chapter 4—once a space is allowed to have an arbitrary basis, what a change of basis means, and how a similarity transformation may represent one and the same mapping, will both unfold naturally from the matrix representation theorem of this chapter. The geometric intuition for the determinant is only introduced in this chapter; Chapter 7 will place it on the rigorous axioms of multilinearity and antisymmetry, and will develop the theory of its computation systematically. The eigenvalue problem of Chapter 8 then asks: which basis makes the matrix representation of one and the same mapping simplest? The roots of that question lie deep in §3.2 and §3.4 of this chapter.

The Role of This Chapter in the Book

The identification of linear mappings with matrices is one of the most important conceptual bridges in linear algebra. It turns geometric problems into algebraic computations, and it gives algebraic computations their geometric meaning back. The central problem this chapter solves is: what is a matrix, and where does it come from?—and the answer is not “it is an array of numbers,” but “it is the coordinate representation of a linear mapping with respect to a basis.” This shift of viewpoint matters more than any computational technique.

Seen from the standpoint of nature and of engineering, almost every linear approximation (the linearization of a mechanical system, the linear filtering of a signal, the fully connected layer of a neural network) takes the matrix as its central language. Only by understanding “why matrix multiplication is defined the way it is” can one understand why the stacked matrix multiplications of deep learning are, geometrically, the same thing as the composition of a series of transformations of space. By the end of this chapter, the matrix you see is no longer merely an array of numbers waiting to be computed with, but a geometric operation on space—one that distorts, rotates, projects, or compresses it—and this point of view will be the most important background for reading every chapter that follows.

Concept Map