Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Open In Colab Binder

Guide to the Experiment

The experiment has two parts:

Part 1 (a toy corpus): using a small 30×1530 \times 15 matrix, we take LSA apart step by step so that the geometric meaning of the SVD is completely visible; an interactive experiment lets you feel the effect of the truncation dimension kk directly.

Part 2 (a real corpus): using the 20 Newsgroups dataset from sklearn (2151 news posts from 4 of its topics, with 4000 terms), we see how the behavior of the SVD changes when the size of the matrix jumps from 30×1530 \times 15 to 4000×21514000 \times 2151.

Part 1: A Toy Corpus (30 Terms × 15 Documents)

§1 Building the Term-Document Matrix

The corpus consists of 5 documents on each of three topics (15 documents in all), and the vocabulary contains 30 terms (10 per topic):

TopicVocabulary (10 terms)
Linear algebramatrix, vector, eigenvalue, singular value, decomposition, determinant, linear equation, subspace, basis, orthogonal
Quantum mechanicsquantum state, superposition, entanglement, measurement, Hamiltonian, wave function, operator, eigenenergy, density matrix, unitary
Machine learningneural network, gradient, loss function, training, feature, classification, regression, dim reduction, data, optimization

Term-document matrix A∈R30×15\mathbf{A} \in \mathbb{R}^{30 \times 15}: entry (i,j)(i,j) = the number of times term ii occurs in document jj. Each document contains only terms from its own topic, and terms from different topics never co-occur—this is the key assumption, and we will see later how it affects the result of the SVD.

§2 TF-IDF Weighting

Raw term counts have a problem: the term “matrix” occurs in every linear algebra document, and this very “generality” lowers its discriminating power. TF-IDF weighting damps frequent general-purpose terms and amplifies rare but discriminating ones:

aij←cij∑i′ci′j⏟TF (normalized term frequency)×(log⁡n+1dfi+1+1)⏟IDF (inverse document frequency)a_{ij} \leftarrow \underbrace{\frac{c_{ij}}{\sum_{i'} c_{i'j}}}_{\text{TF (normalized term frequency)}} \times \underbrace{\left(\log\frac{n+1}{df_i+1} + 1\right)}_{\text{IDF (inverse document frequency)}}

Here cijc_{ij} is the raw count, dfidf_i is the number of documents in which term ii occurs, and n=15n=15 is the total number of documents.

§3 The SVD of the Term-Document Matrix

A=U Σ V⊤U∈R30×15,  Σ∈R15×15,  V∈R15×15\mathbf{A} = \mathbf{U}\,\boldsymbol{\Sigma}\,\mathbf{V}^{\top} \qquad \mathbf{U}\in\mathbb{R}^{30\times 15},\; \boldsymbol{\Sigma}\in\mathbb{R}^{15\times 15},\; \mathbf{V}\in\mathbb{R}^{15\times 15}
  • Column ii of U\mathbf{U}, ui\mathbf{u}_i: the ii-th left singular vector; row tt of U\mathbf{U} gives the coordinates of term tt in the SVD semantic space

  • Column ii of V\mathbf{V}, vi\mathbf{v}_i: the ii-th right singular vector; row jj of V\mathbf{V} gives the coordinates of document jj in the SVD semantic space
    (or, equivalently, (ΣV⊤):,j(\boldsymbol{\Sigma}\mathbf{V}^{\top})_{:,j} is the weighted low-dimensional coordinate vector of document jj)

  • σi\sigma_i: the importance of the ii-th “semantic topic” (the larger, the more important)

After truncation to kk dimensions, only the first kk singular values are kept, and the best rank-kk approximation of the matrix A\mathbf{A} is

Ak=∑i=0k−1σi uivi⊤\mathbf{A}_k = \sum_{i=0}^{k-1}\sigma_i\,\mathbf{u}_i\mathbf{v}_i^{\top}

§4 Interactive Experiment: The Truncation Dimension kk and Semantic Queries

What the two panels show:

Left panel (singular-value energy plot): the bar chart shows the energy σi2\sigma_i^2 of each singular value; the first kk singular values, which are kept, are shown in dark violet, and the discarded part is shown in light violet. Adjusting kk shows you directly “how much information is kept.”

Right panel (semantic similarity): choose a term from the drop-down menu, and the plot shows the 5 terms with the highest cosine similarity to it in the current kk-dimensional LSA space:

sim(i,j)=Uk[i,:]⊤ Uk[j,:]∥Uk[i,:]∥ ∥Uk[j,:]∥\text{sim}(i, j) = \frac{\mathbf{U}_k[i,:]^{\top}\,\mathbf{U}_k[j,:]}{\|\mathbf{U}_k[i,:]\|\,\|\mathbf{U}_k[j,:]\|}

What to look for:

  • k=1k=1: there is only one semantic direction, all the terms are “squeezed together,” and the similarities are almost all 1

  • k=3k=3: the three topics separate exactly; same-topic terms have similarity ≈ 1, and cross-topic terms have similarity ≈ 0

  • k>5k>5: noise singular vectors come in, and cross-topic “contamination” begins to appear

Suggested queries: try “matrix” (linear algebra) and “entanglement” (quantum mechanics). How do their most similar terms differ between k=3k=3 and k=10k=10?


Part 2: A Real Corpus—20 Newsgroups

The matrix of the toy corpus is 30×1530 \times 15, small enough to see with the naked eye. What does a real-world text matrix look like?

We use the 20 Newsgroups dataset: a classic standard NLP dataset drawn from newsgroups of the 1990s, with 20 topics in all (we select 4). We first look at the size of the matrix, and then compare it with the toy corpus.

The core differences between toy and real:

Toy corpus (30×15, disjoint topics)20 Newsgroups (4 overlapping topics)
Global common mode σ0\sigma_0None (no terms shared across topics; the matrix is block-diagonal)Present (σ0\sigma_0 is clearly the largest, σ0/σ1≈1.54\sigma_0/\sigma_1 \approx 1.54; u0\mathbf{u}_0 nearly all positive = the average document direction)
Where the topic structure liesFrom σ0\sigma_0 on (the leading singular vectors = the topics themselves)From σ1\sigma_1 on (the common mode must be discarded first)
Decay of the topic structureCliff-like (drops after the first 3 dimensions; k=3k=3 suffices)Gradual (no break after the common mode is removed; a larger kk is needed)
Cross-topic similarityExactly 0Nonzero (reflecting the fuzziness of language)
The “magic” of the SVDMathematically inevitable (block-diagonal)Statistically effective (but not exact)

Summary of the Experiment

The core differences between toy and real:

Toy corpus (30×15)20 Newsgroups (4000×2151)
Topic boundariesNo overlap at allFuzzy (terms appear across topics)
Decay of the singular valuesCliff-like (k=3k=3 suffices)Gradual (a larger kk is needed)
Cross-topic similarityExactly 0Nonzero (reflecting the fuzziness of language)
The “magic” of the SVDMathematically inevitableStatistically effective (but not exact)

Further reflection: modern word embeddings (Word2Vec, GloVe) can be understood as neural-network versions of LSA—they likewise map semantically similar words to nearby positions in a low-dimensional space, except that the training objective changes from “reconstructing the matrix” to “predicting the context words.” The SVD provides the cleanest linear algebra foundation for this idea.