Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Open In Colab Binder

Background and Motivation

The Iris dataset (Fisher, 1936) contains 150 iris flowers. For each flower, 4 features are measured—sepal length, sepal width, petal length, and petal width (all in centimeters)—and the species is recorded (setosa, versicolor, virginica, 50 flowers each).

The core question: the 4 features cannot all be drawn in a plane at once. At most we can pick two features for a scatter plot, but then we lose the information in the other two features. Can we find the “most telling” two-dimensional plane—not a coordinate plane spanned by any two of the original features, but a combination of directions in 4-dimensional space that retains the most variance—and pack as much of the information of all 4 features into it as possible?

This is exactly the question that PCA answers. In order, this experiment will:

  • observe the correlation structure among the original features, and understand why plotting an arbitrary pair of features is not ideal (Block A);

  • apply the SVD to the centered data and read off the principal directions and their loadings (Block B);

  • use a scree plot to see how the variance is distributed among the principal components, and verify the conservation of total variance (Block C);

  • project onto the best plane, spanned by the first 2 principal components, watch the three species separate into clusters, and use a biplot to interpret the biological meaning of the principal components (Block D);

  • close with the reconstruction error rank by rank, bringing Eckart–Young’s “error = sum of the squares of the tail singular values” down to real data (Block E).

Notational Conventions

We arrange the data matrix with samples as rows (the mainstream convention in machine learning). Vectors, matrices, and summation indices all start at 0.

SymbolMeaningDimension
mmNumber of samples (flowers)m=150m = 150
nnNumber of featuresn=4n = 4
X\mathbf{X}Data matrix; row kk contains the features of flower kkRm×n\mathbb{R}^{m\times n}
xˉ\bar{\mathbf{x}}Sample mean vectorRn\mathbb{R}^{n}
X~\tilde{\mathbf{X}}Centered data matrix X−1mxˉ⊤\mathbf{X} - \mathbf{1}_m\bar{\mathbf{x}}^{\top}Rm×n\mathbb{R}^{m\times n}
S\mathbf{S}Sample covariance matrix 1m−1X~⊤X~\frac{1}{m-1}\tilde{\mathbf{X}}^{\top}\tilde{\mathbf{X}}Rn×n\mathbb{R}^{n\times n}, symmetric positive semidefinite
σi\sigma_iThe ii-th singular value of X~\tilde{\mathbf{X}}σ0≥σ1≥⋯\sigma_0 \ge \sigma_1 \ge \cdots
vi\mathbf{v}_iThe ii-th right singular vector (the ii-th principal direction)Rn\mathbb{R}^{n}
z(k)\mathbf{z}^{(k)}Principal component scores of flower kkRn\mathbb{R}^{n}
ρi\rho_iVariance share of the ii-th principal component, σi2/∑jσj2\sigma_i^2/\sum_j\sigma_j^2[0,1][0,1]

Structure of the Experiment

BlockTopicCore concepts
ALooking at the dataThe 4 original features, species labels, pairwise scatter plots and correlations
BCentering and the SVDThe SVD of X~\tilde{\mathbf{X}}, principal directions and loadings, correspondence with the eigenvalues of S\mathbf{S}
CVariance decompositionScree plot, cumulative explained variance, verifying the conservation of total variance
DBest projection and clusteringProjection onto PC0–PC1, biplot, the duality “things cluster by kind, people form groups”
EReconstruction errorTruncated SVD reconstruction rank by rank, error = sum of the squares of the tail singular values

How to use: run the cells in order from top to bottom. Passages marked ◆ are advanced extensions and may be skipped on a first reading.

Environment Setup

Block A: Looking at the Data

A1 | Loading the Iris Dataset

Each row of the Iris data matrix X∈R150×4\mathbf{X}\in\mathbb{R}^{150\times 4} is a flower, and its 4 columns are, in order, sepal length, sepal width, petal length, and petal width. The label vector y∈{0,1,2}150\mathbf{y}\in\{0,1,2\}^{150} records the species (0=setosa, 1=versicolor, 2=virginica).

PCA is an unsupervised method (the derivation of Theorem 11 does not use the labels at all), but we keep the labels anyway—not for training, but to color the scatter plots afterward, so that we can check whether the low-dimensional structure that PCA finds on its own happens to correspond to the true species.

A2 | Pairwise Scatter Plots: Why Plotting an Arbitrary Pair of Features Is Not Ideal

The figure below plots the 4 features against one another in pairs. We see that the petal length/width pair (bottom right) already separates setosa cleanly, but versicolor and virginica still overlap on every single pair of features; on the sepal length/width pair (top left) the three classes are even more entangled.

The root of the problem is that each subplot uses only 2 of the 4 features and throws away the other 2. The goal of PCA is not to pick the best of these subplots but to recombine all 4 features and build the two-dimensional plane in four-dimensional space that retains the most variance.

Block B: Centering and the SVD

B1 | Center, Then Apply the SVD to X~\tilde{\mathbf{X}}

Following the standard procedure of §11.3.1: first center to obtain X~=X−1mxˉ⊤\tilde{\mathbf{X}} = \mathbf{X} - \mathbf{1}_m\bar{\mathbf{x}}^{\top}, then apply the singular value decomposition directly to X~\tilde{\mathbf{X}}:

X~=UΣV⊤=∑i=0r−1σi uivi⊤=∑i=0r−1σi ∣ui⟩⟨vi∣\tilde{\mathbf{X}} = \mathbf{U}\boldsymbol{\Sigma}\mathbf{V}^{\top} = \sum_{i=0}^{r-1}\sigma_i\,\mathbf{u}_i\mathbf{v}_i^{\top} = \sum_{i=0}^{r-1}\sigma_i\,|\mathbf{u}_i\rangle\langle\mathbf{v}_i|

The right singular vectors v0,v1,…\mathbf{v}_0,\mathbf{v}_1,\ldots are the principal directions 0, 1, …; the corresponding σi2/(m−1)\sigma_i^2/(m-1) is the variance explained by each principal component.

B2 | Agreement of the Two Paths + Interpreting the Principal Component Loadings

The first part verifies that the SVD and the eigendecomposition of the covariance matrix give the same variances and directions: the eigenvalues of S\mathbf{S} should equal σi2/(m−1)\sigma_i^2/(m-1), and the eigenvectors should be parallel to vi\mathbf{v}_i (the signs may be opposite; this is an inherent freedom of the SVD and does not affect the subspaces spanned by the principal components).

The second part reads V\mathbf{V} as a loading table: the jj-th component of vi\mathbf{v}_i is “the weight of the jj-th original feature in the ii-th principal component.” Once you can read the loadings, you can read the biological meaning of each principal component.

Block C: Variance Decomposition and the Scree Plot

C1 | How Much Variance Does Each Principal Component Explain?

We bring the two formulas of §11.3.1 down to Iris. For Πk=VkVk⊤\boldsymbol{\Pi}_k = \mathbf{V}_k\mathbf{V}_k^{\top} (the projection onto the first kk principal components),

Vapprox(Πk)=1m−1∑i=0k−1σi2,Verr(Πk)=1m−1∑i=kr−1σi2V_{\text{approx}}(\boldsymbol{\Pi}_k) = \frac{1}{m-1}\sum_{i=0}^{k-1}\sigma_i^2,\qquad V_{\text{err}}(\boldsymbol{\Pi}_k) = \frac{1}{m-1}\sum_{i=k}^{r-1}\sigma_i^2

and Theorem 10 guarantees that whatever the value of kk, the two always add up to the total variance VtotalV_{\text{total}}. We define the variance share of the ii-th principal component as

ρi=σi2∑jσj2\rho_i = \frac{\sigma_i^2}{\sum_{j}\sigma_j^2}

and draw it as a scree plot: the bars are the ρi\rho_i of the principal components, and the line is the cumulative share ∑j≤iρj\sum_{j\le i}\rho_j. The scree plot of Iris shows an extremely steep “elbow”—the first bar alone takes more than 90% of the variance, and this is the fundamental reason why PCA can compress 4 dimensions into 2 with almost no distortion.

Block D: Projecting onto the Best Plane and Watching the Clusters Appear

D1 | Principal Component Scores: Projecting the 4-Dimensional Flowers onto PC0–PC1

The principal component scores of a flower are simply its coordinates in the new coordinate system (the principal directions). Computing them for the whole batch takes a single matrix multiplication:

Z=X~V∈Rm×n,z(k)=V⊤x~(k)\mathbf{Z} = \tilde{\mathbf{X}}\mathbf{V}\in\mathbb{R}^{m\times n},\qquad \mathbf{z}^{(k)} = \mathbf{V}^{\top}\tilde{\mathbf{x}}^{(k)}

Column 0 of Z\mathbf{Z} holds the PC0 scores of all the flowers, and column 1 the PC1 scores. Plotting the first 2 columns as a scatter plot gives the “best two-dimensional projection” of §11.3—among all two-dimensional planes it retains the most variance (Theorem 11).

The key point: PCA does not use the species labels at all; it blindly looks for the plane of maximal variance. Yet as soon as the figure below is colored by the true species, three groups emerge clearly—setosa is flung far away, while versicolor and virginica line up as two segments along PC0 whose boundaries barely touch. This shows that the species differences in Iris are essentially a one-dimensional “flower size” gradient (PC0), which is exactly the direction of maximal variance that PCA picks up automatically.

D2 | Biplot: Overlaying Sample Scores and Feature Loadings in One Figure

A biplot draws two kinds of things at once in the same PC0–PC1 plane, corresponding to the row/column duality of Remark 4:

  • Points = sample scores (the first two columns of Z\mathbf{Z}), that is, “people form groups”—flowers of the same species gather together.

  • Arrows = the original feature axes projected onto the PC plane, that is, “things cluster by kind”—letting us see which features point in the same direction.

Where do the arrows come from? The jj-th original feature axis is the standard basis vector ej∈R4\mathbf{e}_j\in\mathbb{R}^4; rewriting it in the new coordinates given by the principal components, its coordinates are (v0⊤ej, v1⊤ej)=(Vj0, Vj1)(\mathbf{v}_0^{\top}\mathbf{e}_j,\ \mathbf{v}_1^{\top}\mathbf{e}_j) = (V_{j0},\ V_{j1}), that is, the first two components of row jj of V\mathbf{V}. This is precisely the orthogonal projection of ej\mathbf{e}_j onto the PC0–PC1 plane. To make them visible against the cloud of points, all the arrows are then multiplied by one and the same magnification factor (which changes only their lengths, not their directions). How to read the plot:

  • Arrow length (before projection ∥ej∥=1\|\mathbf{e}_j\|=1): the closer to 1, the more nearly the whole feature lies in the plane and is faithfully represented; the shorter, the more of it “sticks out of the plane,” hidden in the discarded PC2/PC3.

  • Angle with the PC0 axis: whether the feature leans toward PC0 or PC1.

  • Angle between two arrows: a small angle ⇒\Rightarrow the two features are positively correlated in this plane.

D3 | Interactive: Choose Any Two Principal Components as Coordinate Axes

When PCA compresses 4 dimensions into 2, it takes PC0–PC1 by default (retaining the most variance). Sometimes, however, a minor principal component hides another kind of structure (the Swiss roll of §11.3.2 is exactly such a counterexample). The two drop-down menus below let you choose freely which two principal components to use for the xx- and yy-axes.

Here we use plotly’s native drop-downs (updatemenus). They are controlled purely on the front end and do not depend on a running kernel, so they update instantly both in JupyterLab and in the statically built HTML book. Try comparing PC0–PC1 (the default, with the best separation) with PC2–PC3 (residual directions, where the three classes overlap almost completely), and see that on Iris “retaining the most variance” and “discriminating best” happen to coincide—but this coincidence is not a matter of course; it happens because the species differences happen to line up with the direction of maximal variance.

Block E: Reconstruction Rank by Rank and the Eckart–Young Error Formula

E1 | Truncated SVD Reconstruction: The Error Is Exactly the Sum of the Squares of the Tail Singular Values

The last piece of the puzzle of the Eckart–Young theorem (Theorem 8): reconstruct the data from the first kk principal components,

X~k=∑i=0k−1σi uivi⊤=X~Πk\tilde{\mathbf{X}}_k = \sum_{i=0}^{k-1}\sigma_i\,\mathbf{u}_i\mathbf{v}_i^{\top} = \tilde{\mathbf{X}}\boldsymbol{\Pi}_k

Its reconstruction error has a closed form,

∥X~−X~k∥F2=∑i=kr−1σi2\|\tilde{\mathbf{X}} - \tilde{\mathbf{X}}_k\|_F^2 = \sum_{i=k}^{r-1}\sigma_i^2

that is, “the sum of the squares of the singular values of the discarded principal components.” In this block we reconstruct for k=0,1,2,3,4k=0,1,2,3,4 one at a time and set the measured Frobenius error side by side with the tail sum of squares predicted by the formula—the most direct numerical verification of the whole theory of §11.3 on real data. We also plot the relative reconstruction error ∥X~−X~k∥F/∥X~∥F\|\tilde{\mathbf{X}}-\tilde{\mathbf{X}}_k\|_F / \|\tilde{\mathbf{X}}\|_F as a curve decreasing in kk, confirming by eye that at k=2k=2 the error is already negligibly small.

E2 (Exercise) | How Many Principal Components Are Needed to Reach 99% Cumulative Variance?

Summary of the Experiment

This experiment runs the PCA theory of §11.3 in full on the Iris dataset. When you have finished it, you should be able to:

  1. Carry out the standard PCA procedure: centering → SVD of X~\tilde{\mathbf{X}} → take the right singular vectors as the principal components, and explain why we do not form the covariance matrix S\mathbf{S} first (the condition number is squared, Remark 5).

  2. Verify that the SVD and the eigendecomposition of S\mathbf{S} give the same variances (λi=σi2/(m−1)\lambda_i = \sigma_i^2/(m-1)) and parallel directions.

  3. Interpret the loading table of the principal components: state that PC0 of Iris is the “flower size axis” and PC1 the “sepal size axis.”

  4. Draw and read a scree plot, and point out that the elbow for Iris is extremely steep (PC0 accounts for 92.5%), which is the basis for a drastic reduction in dimension.

  5. Verify numerically the conservation of total variance: for every kk, Vapprox(k)+Verr(k)=VtotalV_{\text{approx}}(k)+V_{\text{err}}(k)=V_{\text{total}} (Theorem 10).

  6. Project and watch PCA separate the three species in the PC0–PC1 plane without labels (Theorem 11), and use a biplot to overlay “things cluster by kind” (feature loadings) and “people form groups” (sample scores) in a single figure (Remark 4).

  7. Verify the Eckart–Young error formula: the rank-by-rank reconstruction error ∥X~−X~k∥F2=∑i≥kσi2\|\tilde{\mathbf{X}}-\tilde{\mathbf{X}}_k\|_F^2 = \sum_{i\ge k}\sigma_i^2 (Theorem 8).


ConceptKey formulaValues for Iris
Centering + SVDX~=UΣV⊤\tilde{\mathbf{X}} = \mathbf{U}\boldsymbol{\Sigma}\mathbf{V}^{\top}σ=(25.10, 6.01, 3.41, 1.89)\sigma = (25.10,\ 6.01,\ 3.41,\ 1.89)
Variance of the principal componentsλi=σi2/(m−1)\lambda_i = \sigma_i^2/(m-1)—
Variance shareρi=σi2/∑jσj2\rho_i = \sigma_i^2/\sum_j\sigma_j^2(92.5%, 5.3%, 1.7%, 0.5%)(92.5\%,\ 5.3\%,\ 1.7\%,\ 0.5\%)
Conservation of total varianceVapprox(k)+Verr(k)=VtotalV_{\text{approx}}(k)+V_{\text{err}}(k)=V_{\text{total}}Independent of kk
Principal component scoresZ=X~V\mathbf{Z} = \tilde{\mathbf{X}}\mathbf{V}The first 2 columns are plotted
Reconstruction error∣X~−X~k∣F2=∑i≥kσi2|\tilde{\mathbf{X}}-\tilde{\mathbf{X}}_k|_F^2 = \sum_{i\ge k}\sigma_i^2Relative error ≈15%\approx 15\% at k=2k=2 (retaining 97.8%97.8\% of the variance)

Concept Map

This concept map ties the whole experiment into a single thread: centering → SVD is the engine, and it produces two outputs, the singular values (which rank the variances by priority) and the right singular vectors (the best coordinate axes). The former support the scree plot, the conservation law, and the Eckart–Young error formula; the latter, through the scores Z=X~V\mathbf{Z}=\tilde{\mathbf{X}}\mathbf{V}, project the samples onto the best plane and meet the feature loadings in the biplot.

Looking back: this experiment brings the PCA theory of §11.3 (conservation + best projection) down to real data, reusing the centering and total-variance tools of Experiment 5. Looking ahead: put a different matrix into the same SVD engine and the remaining applications of this chapter grow out of it—replace the matrix with term-document counts and you get latent semantic analysis (LSA); use the singular values as a threshold for regularization and you get the Moore–Penrose pseudoinverse and ridge regression of §11.4. Here Iris is only the gentlest introductory example; the real tension (high variance is not necessarily meaningful) was already foreshadowed by the Swiss roll of §11.3.2.