In 1926 quantum physics ran into a famous split in its formulation: Heisenberg’s “matrix mechanics” dealt with discrete arrays of numbers and transitions, while Schrödinger’s “wave mechanics” solved continuous partial differential equations. The two formulations looked worlds apart, yet physicists discovered to their astonishment that they gave exactly the same predictions. Before long the truth came into view: the language of vector spaces and linear operators supplied a common skeleton linking the two formulations—the discrete infinite-dimensional vectors and the continuous wave functions were, at bottom, only projections of one and the same abstract space onto different “bases.”
This story reveals the real driving force behind abstraction in modern mathematics: abstraction has never been a pursuit of formal vanity or elegance; it arises because objects that look completely different turn out to obey one and the same set of rules of operation—and that coincidence calls out for an explanation of what lies beneath it.
In fact, the germ of that explanation appeared nearly ninety years before quantum mechanics.
In 1844 the Prussian schoolteacher Hermann Grassmann published Die lineare Ausdehnungslehre. The central idea of the book was far ahead of its time: he held that the algebraic structure of “things that can be added and scaled” deserved study as a discipline in its own right, whatever objects it happened to carry—geometric line segments, polynomials, or mechanical quantities. He did not care what a vector “is in itself,” only what a vector “can do under the operations”—and this is precisely the essence of modern axiomatic thinking. The price of being ahead of one’s time, however, was enormous: Gauss privately acknowledged the insight but admitted that the book was “too laborious to read,” and the academic mainstream and university review committees shut it out because of its highly philosophical style. Grassmann waited nearly twenty years and in 1862 published, at his own expense, a heavily rewritten second edition, which still found almost no readers. Disheartened, he eventually turned his energies to the study of Sanskrit and achieved world-class results in comparative linguistics; he died in 1877 with his regret unresolved, never having seen his mathematical vision recognized by the academic world in his lifetime.
It was not until 1888 that the Italian mathematician Giuseppe Peano, deeply inspired by Grassmann, gave in his Calcolo geometrico the first rigorous axiom system for abstract vector spaces in human history—the very eight rules we know today: any set that is closed under addition and scalar multiplication and satisfies these eight rules of computation is, as an algebraic structure, a “vector space,” no matter whether what it holds is arrays of real numbers, geometric arrows, polynomials, matrices, or functions. Even so, Peano’s definition still seemed too far ahead of its time, and his contemporaries were widely puzzled: “We have concrete matrices and differential equations to compute with—why go to the trouble of this airy, insubstantial set of axioms?” The axiom system slept on the bookshelf for more than thirty years more, until the mathematical crisis set off by the explosion of quantum mechanics jolted the whole scientific community awake: mathematicians had long since built for them an ultimate framework capable of holding all of this at once.
From Grassmann to von Neumann, the tool was born nearly a century before the physical world urgently needed it.
The scope of this book is finite-dimensional linear spaces; we will not go deeply into the infinite-dimensional topological analysis of Hilbert spaces and Banach spaces—between finite and infinite dimension there are many fundamental differences of an analytic kind, and that is the home ground of functional analysis. But in the finite-dimensional world Peano’s eight axioms are just as complete, pure, and powerful. In the first three chapters we grew used to computing with vectors, matrices, and geometric transformations in Rn, and we may have sensed dimly that polynomials, functions, and even spaces of matrices all follow “exactly the same rules of the game.”
The task of this chapter is to turn that vague intuition into rigorous mathematical language: What is a vector space? What is a linear mapping? What is dimension? Once these concepts are firmly established at the level of axioms, the theory acquires a universal penetrating power—it applies automatically to signal processing, quantum computing, dynamical systems in economics, and even the embedding spaces of learned representations in contemporary machine learning, without our having to reinvent the wheel in every new field.
Peano’s eight axioms are the starting point of this chapter, and the threshold of the book’s turn toward abstraction. §4.1 defines vector spaces at the level of axioms and pins down the precise meaning of subspace, span, and linear dependence. The concept of dimension receives its most general definition here—no longer “the n in Rn,” but “the size of a maximal linearly independent set,” a definition that applies equally well to polynomial spaces and matrix spaces.
§4.2 extends the linear transformations of Chapter 3 from Rn to mappings between arbitrary vector spaces. The kernel and the image of a linear mapping reveal how a transformation “compresses” space; the rank-nullity theorem is the central result of this section, and it is also the theoretical foundation for the structure of the solutions of systems of linear equations in Chapter 6.
§4.3 surveys systematically the typical examples of finite-dimensional vector spaces: the polynomial spaces Pn, the matrix spaces Mm×n, finite-dimensional subspaces of function spaces, and the construction of direct sums. Together these examples show that vector spaces are far more varied than Rn, and that the axiomatic framework handles all of these cases in a unified way.
Once you have read this chapter, what the word “vector” means to you will have changed completely—it is no longer “a list of numbers,” but any mathematical object that satisfies the eight rules. To the learner’s question “why must it be so abstract?” you will have an answer of your own.
The Steinitz exchange lemma is the key tool for proving that the dimension of a finite-dimensional vector space is unique.
import numpy as np
# ──────────────────────────────────────────────────────────
# Helper functions
# ──────────────────────────────────────────────────────────
def print_header(title: str) -> None:
sep = "─" * 52
print(f"\n{sep}")
print(f" {title}")
print(sep)
def print_step(label: str, value=None) -> None:
if value is None:
print(f"\n[{label}]")
else:
print(f" {label}: {value}")
np.set_printoptions(precision=4, suppress=True, linewidth=100)
# ──────────────────────────────────────────────────────────
# Step 0: define the vectors
# ──────────────────────────────────────────────────────────
print_header("Steinitz Exchange Lemma — Verification in Four-Dimensional Space")
# 𝒲: the standard basis (n = 4)
e = np.eye(4, dtype=float) # e[i] is e_i (0-based)
# 𝒱: three linearly independent vectors (m = 3)
v0 = np.array([1, 1, 0, 0], dtype=float)
v1 = np.array([0, 1, 1, 0], dtype=float)
v2 = np.array([0, 0, 1, 1], dtype=float)
V = [v0, v1, v2]
print_step("The set 𝒱 (linearly independent, m = 3)")
for i, vi in enumerate(V):
print(f" v{i} = {vi}")
print_step("The set 𝒲 (standard basis, n = 4)")
for i in range(4):
print(f" e{i} = {e[i]}")
# ──────────────────────────────────────────────────────────
# Step 1: verify that 𝒱 is linearly independent
# ──────────────────────────────────────────────────────────
print_header("Step 1: Verify That 𝒱 Is Linearly Independent")
rank_V = np.linalg.matrix_rank(np.column_stack(V))
print_step("rank(𝒱)", rank_V)
assert rank_V == 3, "𝒱 is not linearly independent!"
print(" ✓ rank = m = 3, so 𝒱 is linearly independent")
# ──────────────────────────────────────────────────────────
# Step 2: exchange one vector at a time and check that rank = 4 at every step
# ──────────────────────────────────────────────────────────
print_header("Step 2: Replacing One Vector at a Time — the Exchange Process")
# Initial set: all four standard basis vectors
current = list(e) # current[i] is vector i of the current set
replacements = [] # records the index of each e_i that is replaced
for k, vk in enumerate(V):
print_step(f"Replacement round {k}: bring in v{k} = {vk}")
# Express vk as a linear combination of current (solve for the coefficients)
M = np.column_stack(current) # (4, 4) matrix
coeffs, _, _, _ = np.linalg.lstsq(M, vk, rcond=None)
print(f" vk = Σ α_i · current_i, coefficients α = {np.round(coeffs, 4)}")
# Find the first position that is not a member of 𝒱 and has a nonzero coefficient
replaced_idx = None
for idx in range(len(current)):
if idx not in range(k): # must not replace a v already put in
if abs(coeffs[idx]) > 1e-10:
replaced_idx = idx
break
assert replaced_idx is not None, "No vector is available for replacement!"
old_vec = current[replaced_idx].copy()
current[replaced_idx] = vk
rank_now = np.linalg.matrix_rank(np.column_stack(current))
print(f" replace current[{replaced_idx}] = {np.round(old_vec,4)} → v{k}")
print(f" rank of the set after the replacement = {rank_now}")
assert rank_now == 4, f"rank ≠ 4 after the replacement!"
print(f" ✓ rank = 4, so the span is unchanged")
replacements.append(replaced_idx)
# ──────────────────────────────────────────────────────────
# Step 3: output the final basis
# ──────────────────────────────────────────────────────────
print_header("Step 3: The Final Basis ℬ")
for i, b in enumerate(current):
label = f"v{i}" if i < len(V) else f"e{replacements[i] if i < len(replacements) else '?'}"
print(f" ℬ[{i}] = {b} ({label})")
rank_final = np.linalg.matrix_rank(np.column_stack(current))
print_step("rank of the final basis", rank_final)
assert rank_final == 4
print(" ✓ Construction complete: the 3 vectors v_i and 1 retained standard basis vector span ℝ⁴")
print(f"\n Conclusion of the lemma verified: m = 3 ≤ n = 4 ✓")
Using the Steinitz exchange lemma, we can prove that the dimension of a finite-dimensional vector space is unique.
A mapping goes from a set to a set; a linear mapping goes from a vector space to a vector space, and it preserves the operations of vector addition and scalar multiplication.
For mappings in general we consider only properties such as injectivity and surjectivity; for linear mappings we can look more closely, and the kernel and the image are two important concepts in their study.
In a vector space, a vector has only one representation with respect to a given basis, and matrices let us express its coordinates with respect to different bases.
When we have two different bases, the coordinates of the same vector with respect to the two bases are different, but they are related linearly.
4.2.4 The Matrix Representation of a Change of Coordinates¶
4.2.5 Linear Mappings between Different Finite-Dimensional Linear Spaces¶
We now go on to discuss linear mappings between different linear spaces.
This result generalizes the two-way correspondence between linear mappings and matrices that we discussed in [Section 3.2.2].
Recalling the classification of linear mappings discussed in Section 4.2.2, the rank-nullity theorem makes it far easier to determine how a linear mapping is classified:
This section introduces some classic finite-dimensional vector spaces in plain, accessible language. We apply a unified analytical framework to each vector space and discuss in detail its basic information, its typical applications, and its signature linear mappings.
Concrete vector spaces are the cornerstone of the theory of linear algebra and the starting point for learning it. They consist of concrete numerical coordinates and provide the most intuitive realization of the concept of a vector. Spaces of this kind are not only concrete models of the theory of abstract vector spaces but also the carriers of almost all practical computation and applications.
Geometric interpretation: complex multiplication performs a rotation and a scaling at the same time, giving a richer geometric structure than real spaces have
Connections with familiar spaces: each complex number carries two pieces of real information (its real part and its imaginary part)
Core Linear Mappings (the Field Must Be Specified)¶
Complex conjugation (real-linear): z↦z; in general it is not complex-linear.
Extracting the real/imaginary part (real-linear): Re,Im:Cn→Rn, regarding the domain as a real vector space.
Unitary transformations (complex-linear): z↦Uz, where U∗U=I. Unitary transformations preserve the complex inner product and can be used to describe unitary evolution in quantum mechanics.
Matrix spaces are among the most important concrete vector spaces in linear algebra: they not only provide concrete representations of linear mappings but also play a central role in a wide range of mathematical and engineering applications.
The trace: tr(A)=∑i=0min(m,n)−1aii (for square matrices)
Left/right multiplication: TB(A)=BA or TC(A)=AC
Geometric meaning of the mappings: transposition corresponds to the dual of a linear transformation; the trace extracts information from the main diagonal
Geometric interpretation: a real symmetric matrix A defines the symmetric bilinear form b(x,y)=x⊤Ay; only when A is positive definite is this bilinear form an inner product, which then defines the length x⊤Ax.
Connections with familiar spaces: each symmetric matrix corresponds to a quadratic form
Taking eigenvalues (nonlinear): maps a symmetric matrix to the vector of its eigenvalues arranged in a fixed order; in general it does not preserve addition.
Geometric interpretation: corresponds to skew-adjoint transformations, and in three-dimensional space is closely related to the cross product of vectors
Connections with familiar spaces: Skew3(R) is isomorphic to R3
Function spaces extend the concept of a vector from finite-dimensional tuples of numbers to infinite-dimensional objects that are functions, a natural extension of the theory of linear algebra toward analysis. Although the full theory of function spaces involves infinite-dimensional analysis, many important function spaces can be approximated and understood through finite-dimensional subspaces.
Definition and notation: Pn(R)={a0+a1x+⋯+anxn∣ai∈R}
Typical basis: the monomial basis {1,x,x2,…,xn}
Dimension: n+1
Important subspaces: the even polynomials {p∈Pn(R):p(−x)=p(x)} and the odd polynomials {p∈Pn(R):p(−x)=−p(x)}; the zero polynomial belongs to both.
A nonexample: the set of monic polynomials is not a subspace. It does not contain the zero polynomial; for example, 1 is a monic polynomial but 1+1=2 is not, so the set is not closed under addition.
Composite spaces embody the constructive and unifying character of the theory of vector spaces: they show how, starting from known vector spaces, new vector spaces can be constructed by operations such as the direct sum and the tensor product. These constructions not only enrich the variety of vector spaces; more importantly, they provide a way of attacking complex problems by systematic decomposition.
Entanglement: (∣00⟩+∣11⟩)/2 cannot be decomposed into a combination of single-qubit states
Why the tensor product is needed: quantum entanglement embodies nonlocal correlations between particles
2. The elastic deformation of a rubber band
Physical phenomenon: the deformation of a rubber band when it is stretched
Tensor-product description: stress state = internal direction of the material ⊗ direction of the external force
Practical significance: a rubber band shows different stiffness when stretched in different directions
Why the tensor product is needed: the molecular chains of the material have an intrinsic orientation, which interacts with the direction of the applied force
3. How polarized sunglasses work
Physical phenomenon: sunglasses can remove the glare from water surfaces and snow
Tensor-product description: photon state = direction of propagation ⊗ direction of polarization
Practical effect: light reflected from water is mainly horizontally polarized, and a vertical polarizer can block this glare
Everyday experience: rotating polarized sunglasses lets you watch the glare disappear and reappear
Through this systematic survey of the classic finite-dimensional vector spaces, we can establish a unified analytical framework. The table below summarizes the core features of each class of vector space:
Class of vector space
Specific space
Dimension
Typical basis
Signature linear transformations
Main applications
Numerical vector spaces
Rn
n
standard basis {ei}
rotation, reflection, scaling
geometric transformations, physical modeling
Cn
n (complex) / 2n (real)
standard basis / basis separating real and imaginary parts
transposition, trace, left and right multiplication
linear systems, data processing
Symn(R)
2n(n+1)
Eii, Eij+Eji
symmetrization
quadratic forms, principal component analysis
Skewn(R)
2n(n−1)
Eij−Eji
skew-symmetrization, Lie bracket with one argument fixed
angular velocity tensor, Lie algebras
Function spaces
Pn(R)
n+1
{1,x,x2,…,xn}
differentiation, translation, evaluation
numerical analysis, approximation theory
finite Fourier space
2n+1
{1}∪{sinkx,coskx:1≤k≤n}
differentiation, transformation to the frequency domain
signal processing, spectral analysis
Composite spaces
V⊕W
dimV+dimW
bases of the components placed side by side
projection, inclusion, action on components
system decomposition, modular design
V⊗W
dimV×dimW
tensor basis {vi⊗wj}
tensor products of mappings, permutation
quantum systems, multilinear analysis
Key points.
Concrete vector spaces provide intuitive geometric pictures and a computational foundation
Matrix spaces provide concrete representations of linear transformations
Function spaces extend the ideas of linear algebra to analysis
Composite spaces display the constructive and unifying character of the theory of vector spaces
Through the unified analytical framework, we have seen the structural features these classic finite-dimensional vector spaces share as well as the distinctive properties of each. This systematic way of understanding helps us to:
Build conceptual connections: understand the analogies between different spaces
Master the method of analysis: use the unified framework to analyze new vector spaces
Choose suitable tools: choose the vector space best suited to the features of a problem
Deepen theoretical understanding: abstract general principles from concrete examples
Practical significance. These classic spaces are not only important components of mathematical theory but also the mathematical foundation of modern science and technology: from quantum computing to machine learning, from signal processing to engineering design, none of them can do without these vector spaces.
Each of the three sections of this chapter answered one question, and together they form a complete logical chain.
§4.1 asked: what exactly is a “vector”? The answer given by Peano’s eight axioms is that any set that is closed under addition and scalar multiplication and obeys the eight laws of operation is called a vector space—whether what it holds is tuples of coordinates, polynomials, or matrices. The basis is the central tool of this framework: it is at once a maximal linearly independent set and a minimal spanning set, and the Steinitz exchange lemma guarantees that, however a basis is chosen, its size never changes. Dimension thus becomes an intrinsic invariant of a vector space, and the starting point for all the quantitative discussion in the chapter.
§4.2 asked: how does a linear mapping “change” a space? The kernel and the image measure, respectively, the dimensions a mapping compresses and the dimensions it retains; the rank-nullity theorem dimV=dimker(T)+dimIm(T) links the two precisely, revealing that a linear mapping is in essence a dimension-conserving redistribution: the dimensions compressed away and the dimensions passed on always add up to the dimension of the input space. The change-of-basis matrix (transition matrix) and the matrix representation of a mapping translate abstract relations between spaces into the computable language of matrices. Deciding whether a mapping is injective, surjective, or an isomorphism is thereby greatly simplified—between spaces of equal dimension the three are equivalent, and it suffices to confirm one of them.
§4.3 asked: how many kinds of spaces satisfy these axioms? From Euclidean spaces and complex spaces to matrix spaces, polynomial spaces, and Fourier spaces, and on to the composite spaces constructed by direct sums and tensor products, these objects, so different on the surface, display an orderly structure within the same framework. The specific values of their dimensions—2n(n+1), dimV×dimW—are not things to memorize but readings of algebraic structure. The fundamental difference between the direct sum and the tensor product (dimensions add vs. dimensions multiply) runs through quantum computing, signal processing, and machine learning, and is the clearest footprint the axiomatic framework has left on modern science.
This chapter builds on Chapter 3 and leads into the chapters that follow, playing the role of the book’s turn toward abstraction. Chapter 3 established the two-way correspondence between matrices and linear mappings on Rn, and all of its conclusions depended on concrete coordinates; this chapter frees that correspondence from Rn and extends it to arbitrary finite-dimensional vector spaces. The idea of Chapter 3 that “a matrix represents a linear mapping” receives a precise meaning within the theorems of this chapter: once bases have been chosen, each linear mapping between finite-dimensional spaces corresponds to exactly one matrix, and vice versa. This correspondence is itself linear, so there is a natural linear isomorphism between the space of linear mappings and the matrix space—the examples of matrix spaces in §4.3 are a direct illustration of this fact.
Looking ahead, the rank-nullity theorem of §4.2 is the theoretical cornerstone for the structure of the solutions of systems of linear equations in Chapter 6: Ax=b has a solution if and only if b∈Im(TA), and the number of degrees of freedom in the solutions is exactly dimker(TA). The block matrices of Chapter 5 correspond, at the algebraic level, to the direct-sum decompositions of §4.3—decomposing a complex space into relatively independent subspaces is precisely the geometric motivation for block operations. In the eigenvalue problem of Chapter 8, the concept of an invariant subspace deepens the theory of subspaces of this chapter; the inner product spaces of Chapter 9 add the structure of geometric distance on top of the axiomatic framework of §4.1, and Gram-Schmidt orthogonalization is a concrete application of the theory of bases in spaces equipped with a metric.
Grassmann published his Die lineare Ausdehnungslehre in 1844 and brought out a new edition in 1862; his work long went without wide recognition, although Hankel had already acknowledged its contribution in 1867. Peano proposed an axiomatic description of real linear spaces in 1888. This history reminds us that the value of an abstract structure may take time to become apparent; the common language established in this chapter allows matrices, polynomials, and functions to be studied within a single framework.
What this chapter has done is precisely to understand the structure of this common language. The axioms of a vector space, the kernel-image decomposition of a linear mapping, the rank-nullity theorem—once these concepts are established at the level of axioms, they apply automatically to basis-function expansions in signal processing, the state spaces of quantum physics, the linear approximations of economic models, and the embedding spaces of machine learning, with no need to prove them again for each field. This is the true power of abstraction: not to make mathematics harder, but to let many problems that would otherwise have to be solved separately be answered all at once.
What you take away from this chapter is not just a set of definitions and theorems but a way of looking at structure—faced with a new mathematical object, the first question is not “what do its elements look like?” but “which laws of operation does it satisfy, what is its dimension, and how do its linear mappings decompose the space?” This habit of questioning is the ticket of admission to all of the more advanced linear algebra that follows.