Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Open In Colab Binder

Background and Motivation

In data science and machine learning we often face multidimensional data: each sample x(k)\mathbf{x}^{(k)} is an nn-dimensional vector recording the observed values of nn features (for example, height, weight, age, …).

The core question: how can a single matrix compactly describe “the structure of linear correlation among the features”?

The answer is the sample covariance matrix S\mathbf{S}. Starting from geometric intuition, this experiment reveals step by step:

  • why data are “centered,” and the algebraic structure of the centering matrix Hm\mathbf{H}_m

  • three equivalent forms of the covariance matrix and the computational advantages of each

  • how the trace tr(S)\text{tr}(\mathbf{S}) measures the “total spread” of the data

  • the precise relationship between the correlation coefficient rr and the shape of the covariance matrix (interactive exploration)

Notational Conventions

We arrange the data matrix with samples as rows, the prevailing convention in machine learning. Vector indices and matrix indices all start at 0.

SymbolMeaningDimension
mmNumber of samplesPositive integer
nnNumber of featuresPositive integer
x(k)\mathbf{x}^{(k)}Sample kk (a column vector; its transpose is a row vector)Rn\mathbb{R}^n, k=0,…,m−1k=0,\ldots,m{-}1
xj(k)x^{(k)}_jFeature jj of sample kkj=0,1,…,n−1j = 0, 1, \ldots, n-1
X\mathbf{X}Data matrix, whose row kk is (x(k))⊤(\mathbf{x}^{(k)})^{\top}Rm×n\mathbb{R}^{m\times n}
xˉ\bar{\mathbf{x}}Sample mean vector 1m∑kx(k)\tfrac{1}{m}\sum_k \mathbf{x}^{(k)}Rn\mathbb{R}^n
x~(k)\tilde{\mathbf{x}}^{(k)}Centered sample x(k)−xˉ\mathbf{x}^{(k)}-\bar{\mathbf{x}}Rn\mathbb{R}^n
X~\tilde{\mathbf{X}}Centered data matrixRm×n\mathbb{R}^{m\times n}
Hm\mathbf{H}_mCentering matrix Im−1m1m1m⊤\mathbf{I}_m - \tfrac{1}{m}\mathbf{1}_m\mathbf{1}_m^{\top}Rm×m\mathbb{R}^{m\times m}
S\mathbf{S}Sample covariance matrixRn×n\mathbb{R}^{n\times n}, symmetric positive semidefinite

Structure of the Experiment

BlockTopicCore ideasRemarks
ALooking at the dataData matrix, mean vector, geometric meaning of centeringChapter 5 suffices
BThe centering matrix Hm\mathbf{H}_mIdempotence, ◆ orthogonal projection, rank and traceThe geometric interpretation needs §9.2
CThe covariance matrixThree equivalent forms; ◆ generating correlated data with CholeskyCholesky needs §9.3.2
DTrace and Frobenius normTwo views of the total variance; ◆ preview of PCA via eigenvaluesD4 needs §8 + §11

How to use: run the cells in order from top to bottom. When you reach a cell marked # 🖊 Student fill-in, fill in the code before continuing.

Environment Setup

Block A: Looking at the Data

Structure of the Data Matrix

Suppose we have mm samples, each with nn features; together they form the data matrix X∈Rm×n\mathbf{X} \in \mathbb{R}^{m\times n}:

X=[—(x(0))⊤——(x(1))⊤—⋮—(x(m−1))⊤—]=[x0(0)⋯xn−1(0)x0(1)⋯xn−1(1)⋮⋱⋮x0(m−1)⋯xn−1(m−1)]\mathbf{X} = \begin{bmatrix} \text{—} & (\mathbf{x}^{(0)})^{\top} & \text{—} \\ \text{—} & (\mathbf{x}^{(1)})^{\top} & \text{—} \\ & \vdots & \\ \text{—} & (\mathbf{x}^{(m-1)})^{\top} & \text{—} \end{bmatrix} = \begin{bmatrix} x^{(0)}_0 & \cdots & x^{(0)}_{n-1} \\ x^{(1)}_0 & \cdots & x^{(1)}_{n-1} \\ \vdots & \ddots & \vdots \\ x^{(m-1)}_0 & \cdots & x^{(m-1)}_{n-1} \end{bmatrix}

The entry (X)kj=xj(k)(\mathbf{X})_{kj} = x^{(k)}_j is the value of feature jj for sample kk.

Why center the data? Covariance measures “how the fluctuations of the features about their means vary together,” not their absolute magnitudes. Centering (subtracting the mean) makes the origin coincide with the centroid of the cloud of samples and removes the distortion caused by an offset mean.

A1 | Loading the Example Data (Height × Weight)

Throughout the experiment we use the following concrete example: m=5m=5 adult samples with n=2n=2 features (height in cm, weight in kg):

X=[1706517570165601807516055]∈R5×2\mathbf{X} = \begin{bmatrix} 170 & 65 \\ 175 & 70 \\ 165 & 60 \\ 180 & 75 \\ 160 & 55 \end{bmatrix} \in \mathbb{R}^{5 \times 2}

A2 | Computing the Mean Vector and Centering

xˉ=1m∑k=0m−1x(k),X~=X−1mxˉ⊤\bar{\mathbf{x}} = \frac{1}{m}\sum_{k=0}^{m-1} \mathbf{x}^{(k)}, \qquad \tilde{\mathbf{X}} = \mathbf{X} - \mathbf{1}_m \bar{\mathbf{x}}^{\top}

A3 | Interactive Scatter Plot: Original Data vs. Centered Data

Run the following cell and use the interactive figure to see the difference before and after centering.

Things to observe: compare the left and right panels and think about the following questions:

  1. How does the position of the × point (the mean) differ between the two panels?

  2. Do the shape and relative positions of the point cloud change?

  3. Why does the × in the right panel land exactly at the origin?

Block B: The Centering Matrix Hm\mathbf{H}_m

From Centering by Hand to a Matrix Operator

In A2 we centered the data by X~=X−1mxˉ⊤\tilde{\mathbf{X}} = \mathbf{X} - \mathbf{1}_m\bar{\mathbf{x}}^{\top}. Observing that xˉ=1mX⊤1m\bar{\mathbf{x}} = \frac{1}{m}\mathbf{X}^{\top}\mathbf{1}_m and substituting, we get

X~=X−1m1m1m⊤X=(Im−1m1m1m⊤)⏟HmX\tilde{\mathbf{X}} = \mathbf{X} - \frac{1}{m}\mathbf{1}_m\mathbf{1}_m^{\top}\mathbf{X} = \underbrace{\left(\mathbf{I}_m - \frac{1}{m}\mathbf{1}_m\mathbf{1}_m^{\top}\right)}_{\displaystyle\mathbf{H}_m}\mathbf{X}

This defines the centering matrix:

Hm=Im−1m1m1m⊤∈Rm×m\mathbf{H}_m = \mathbf{I}_m - \frac{1}{m}\mathbf{1}_m\mathbf{1}_m^{\top} \in \mathbb{R}^{m\times m}

◆ Geometric Meaning: Orthogonal Projection (Needs §9.2)

Hm\mathbf{H}_m is the orthogonal projection matrix onto “the subspace orthogonal to 1m\mathbf{1}_m.” After Hm\mathbf{H}_m acts on any vector v∈Rm\mathbf{v} \in \mathbb{R}^m, its components sum to zero, which corresponds exactly to the operation of “removing the mean.”

Idempotence, Hm2=Hm\mathbf{H}_m^2 = \mathbf{H}_m, reflects the geometric fact that “projecting an already projected vector once more changes nothing.” You can verify this property numerically in B1 (without understanding the geometric reason); the geometric proof is left to §9.2.

B1 | Constructing Hm\mathbf{H}_m and Verifying Its Properties

B2 | Experimenting with the Centering Matrix

B3 | Interactive Heat Map: The Matrix Structure of Hm\mathbf{H}_m

Block C: The Covariance Matrix

Definition: The Sample Covariance Matrix

The denominator is m−1m-1 (rather than mm) in order to obtain an unbiased estimator.

Three Equivalent Forms

Using block matrix multiplication, one can verify that the definition is equivalent to the following two forms:

S=1m−1X~⊤X~⏟definition form (matrix version)=1m−1X⊤HmX⏟operator form=1m−1 ⁣(X⊤X−mxˉxˉ⊤)⏟expanded form\underbrace{\mathbf{S} = \frac{1}{m-1}\tilde{\mathbf{X}}^{\top}\tilde{\mathbf{X}}}_{\text{definition form (matrix version)}} = \underbrace{\frac{1}{m-1}\mathbf{X}^{\top}\mathbf{H}_m\mathbf{X}}_{\text{operator form}} = \underbrace{\frac{1}{m-1}\!\left(\mathbf{X}^{\top}\mathbf{X} - m\bar{\mathbf{x}}\bar{\mathbf{x}}^{\top}\right)}_{\text{expanded form}}
FormAdvantages
Definition formIntuitive, accumulates sample by sample; requires computing X~\tilde{\mathbf{X}} first
Operator formReveals the projection structure; the first choice for theoretical derivations
Expanded formNeeds only one pass over the data; efficient to implement numerically

C1 | Computing the Three Forms Side by Side (Student Fill-In)

C2 | The Correlation Coefficient rr and the Shape of the Covariance Matrix

Theoretical Background

For two-dimensional data (x0,x1)(x_0, x_1), the Pearson correlation coefficient is defined as

r=(S)01(S)00⋅(S)11=Cov(x0,x1)σ0 σ1,r∈[−1,1]r = \frac{(\mathbf{S})_{01}}{\sqrt{(\mathbf{S})_{00}\cdot(\mathbf{S})_{11}}} = \frac{\text{Cov}(x_0, x_1)}{\sigma_0\,\sigma_1}, \qquad r \in [-1, 1]

If we let σ0=σ1=σ\sigma_0 = \sigma_1 = \sigma (equal variances), the covariance matrix is

S=σ2[1rr1]\mathbf{S} = \sigma^2\begin{bmatrix}1 & r \\ r & 1\end{bmatrix}

To generate exactly samples with this covariance structure from a given rr, we use the Cholesky decomposition:

S=LL⊤,L=σ[10r1−r2]\mathbf{S} = \mathbf{L}\mathbf{L}^{\top}, \quad \mathbf{L} = \sigma\begin{bmatrix}1 & 0 \\ r & \sqrt{1-r^2}\end{bmatrix}

For a standard normal sample Z∈Rm×2\mathbf{Z} \in \mathbb{R}^{m\times 2} (with independent columns), let X=ZL⊤\mathbf{X} = \mathbf{Z}\mathbf{L}^{\top}; then the sample covariance matrix of X\mathbf{X} tends exactly to S\mathbf{S} as m→∞m\to\infty.

This section has two parts:

  1. Three fixed typical scenarios (strong positive correlation / no correlation / strong negative correlation), generated exactly and visualized

  2. Interactive exploration: use the sliders to adjust rr, mm, and σ\sigma in real time and watch how the covariance matrix changes

Things to observe interactively (think before you move a slider, then check):

  1. As r→1r \to 1, how does the shape of the point cloud change? What does the ratio of S01\mathbf{S}_{01} to S00⋅S11\sqrt{\mathbf{S}_{00}\cdot\mathbf{S}_{11}} tend to?

  2. Fix r=0.8r=0.8 and increase mm: how does the gap between ractualr_\text{actual} and rtargetr_\text{target} change? Why?

  3. Change σ\sigma but keep rr fixed: how does tr(S)\text{tr}(\mathbf{S}) change, and does ractualr_\text{actual} change? Why?

Block D: Trace and Frobenius Norm

D1 | Two Views of the Total Variance

D2 (Exercise) | Estimating the Spread Radius of the Data from tr(S)\text{tr}(\mathbf{S})

D3 | Interactive Visualization: The Geometric Meaning of ∥X~∥F2\|\tilde{\mathbf{X}}\|_F^2

◆ D4 | Heat Map and Eigenvalues of the Covariance Matrix (Preview of Chapters 8 and 11)

Block E: Capstone Task | A Complete Analysis of Three-Dimensional Data

So far you have practiced each step of the procedure separately. This block is a complete task: you are given a set of three-dimensional data you have never seen before, and you carry out the analysis from start to finish.

The data: 8 students in a class, with three physiological measures recorded:

SymbolMeasureDescription
x0x_0Reaction time (ms)Time from a visual stimulus to a key press
x1x_1Grip strength (kg)Maximal grip strength of the dominant hand
x2x_2Heart rate (bpm)Resting heart rate

Your task: complete the following five steps on your own, without relying on any precomputed results.

Summary of the Experiment

After completing this experiment, you should be able to:

  1. Explain the algebraic properties of Hm\mathbf{H}_m: Hm1m=0\mathbf{H}_m \mathbf{1}_m = \mathbf{0} (it removes the mean direction), Hm2=Hm\mathbf{H}_m^2 = \mathbf{H}_m (idempotence), and tr(Hm)=m−1\text{tr}(\mathbf{H}_m) = m-1. ◆ The geometric interpretation of Hm\mathbf{H}_m as an orthogonal projection needs §9.2; you can come back and complete it then.

  2. Construct Hm\mathbf{H}_m yourself for any new data, and verify that after centering the mean of each column is zero.

  3. Predict and verify that a constant feature vanishes after centering and contributes nothing to the covariance matrix.

  4. Compute and compare the three equivalent forms of the covariance matrix (definition form, operator form, expanded form).

  5. Interpret the diagonal entries (variances) and off-diagonal entries (covariances) of S\mathbf{S}, compute the Pearson correlation coefficient rr, and judge the strength and direction of the correlation between features.

  6. Use the identity tr(S)=∥X~∥F2/(m−1)\text{tr}(\mathbf{S}) = \|\tilde{\mathbf{X}}\|_F^2 / (m-1) to estimate the scale of the spread from reported figures, and verify it with the full data.

  7. Carry out on your own the complete analysis of a three-dimensional data set: mean → centering → covariance matrix → correlation coefficients → visualization.


ConceptKey formula
Centering operatorX~=HmX\tilde{\mathbf{X}} = \mathbf{H}_m \mathbf{X}, Hm=Im−1m1m1m⊤\mathbf{H}_m = \mathbf{I}_m - \tfrac{1}{m}\mathbf{1}_m\mathbf{1}_m^{\top}
Properties of Hm\mathbf{H}_mSymmetric, idempotent, rank(Hm)=tr(Hm)=m−1\text{rank}(\mathbf{H}_m) = \text{tr}(\mathbf{H}_m) = m-1
Covariance matrix (three forms)S=1m−1X~⊤X~=1m−1X⊤HmX=1m−1(X⊤X−mxˉxˉ⊤)\mathbf{S} = \tfrac{1}{m-1}\tilde{\mathbf{X}}^{\top}\tilde{\mathbf{X}} = \tfrac{1}{m-1}\mathbf{X}^{\top}\mathbf{H}_m\mathbf{X} = \tfrac{1}{m-1}(\mathbf{X}^{\top}\mathbf{X} - m\bar{\mathbf{x}}\bar{\mathbf{x}}^{\top})
Correlation coefficientrij=(S)ij(S)ii(S)jjr_{ij} = \tfrac{(\mathbf{S})_{ij}}{\sqrt{(\mathbf{S})_{ii}(\mathbf{S})_{jj}}}
Total variancetr(S)=1m−1∣X~∣F2\text{tr}(\mathbf{S}) = \tfrac{1}{m-1}|\tilde{\mathbf{X}}|_F^2; cyclic property of the trace: tr(X~⊤X~)=tr(X~X~⊤)\text{tr}(\tilde{\mathbf{X}}^{\top}\tilde{\mathbf{X}}) = \text{tr}(\tilde{\mathbf{X}}\tilde{\mathbf{X}}^{\top})
Preview of Chapter 11S=QΛQ⊤\mathbf{S} = \mathbf{Q}\mathbf{\Lambda}\mathbf{Q}^{\top} (eigendecomposition) → principal component analysis (PCA)