Starting from a Question¶
Suppose you want to record “the ratings that 10 users give to 5 movies.” There are two quite different ways to do it:
Approach A: record only the average rating of each user and the average rating of each movie. That is numbers in all.
Approach B: record the rating that each user gives to each movie. That is numbers in all.
Approach A uses the direct-sum structure ; Approach B uses the tensor-product structure .
The two store different amounts of data, and they can answer different questions:
| Question | Direct sum | Tensor product |
|---|---|---|
| What is Xiaoming’s average rating? | ✓ | ✓ |
| What is the average rating of Titanic? | ✓ | ✓ |
| What rating did Xiaoming give Titanic? | ✗ | ✓ |
| Do action fans and romance fans rate the same movie equally high? | ✗ | ✓ |
The price of the direct-sum structure is that it implicitly assumes that every user’s preference distribution has the same shape, differing only in its overall level. In other words, “who is watching” does not affect the ranking of “which movies are popular.” In reality this assumption often fails.
Mathematical Structure¶
Let the user space have dimension and the movie space have dimension . The dimensions of the two ways of combining them are
Through two concrete scenarios (movie recommendation and a casino roulette wheel), this experiment uses actual data to show which information each structure keeps and which it loses.
A note on the language of probability: The second half of the experiment uses “marginal distributions” and “conditional distributions” to describe precisely how the two structures differ. If you are not yet familiar with these notions, just remember: the direct sum corresponds to “computing each mean separately,” the tensor product corresponds to “keeping the actual value of every combination,” and the language of probability is only a precise way of saying the same thing.
Through two concrete scenarios, this experiment shows which information the “direct sum” and the “tensor product” keep, and which they lose, when they serve as data compression structures.
| Structure | Mathematical form | Information kept | Language of probability |
|---|---|---|---|
| Direct sum | Sum of marginals | and (each separately) | Marginal distributions |
| Tensor product | Complete interaction matrix | (joint) | Conditional distributions |
The essential difference between the two:
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import matplotlib.colors as mcolors
from matplotlib.gridspec import GridSpec
np.set_printoptions(precision=3, suppress=True, linewidth=100)
np.random.seed(42)
# --- Color scheme: book palette ---
C_BG = "#F8F8F8"
C_GRID = "#D6D6D6"
C_AXIS = "#000000"
C_V1 = "#57068C" # violet
C_V2 = "#006385" # deep blue
C_T1 = "#2AD2C9" # teal
C_T2 = "#8900E1" # Ultra Violet
C_WARN = "#FF5D47" # orange
C_AUX = "#AB82C5" # light violet
C_AREA = (87/255, 6/255, 140/255, 0.25) # violet, transparent (for matplotlib)
# Fonts: the English edition needs no CJK font, so nothing is downloaded when _lang == 'en';
# the Chinese editions use this same block to fetch Noto Sans TC/SC where no CJK font is installed
import os, urllib.request
import matplotlib.font_manager as fm
_lang = 'en'
_cjk = ['Microsoft JhengHei', 'PingFang TC', 'Noto Sans CJK TC', 'Noto Sans TC']
_have = {f.name for f in fm.fontManager.ttflist}
if _lang != 'en' and not _have & set(_cjk):
_font = os.path.join(os.path.expanduser('~'), '.cache', 'fonts', 'NotoSansTC.ttf')
try:
if not os.path.exists(_font):
os.makedirs(os.path.dirname(_font), exist_ok=True)
urllib.request.urlretrieve('https://github.com/google/fonts/raw/main/ofl/notosanstc/NotoSansTC%5Bwght%5D.ttf', _font + '.part')
os.replace(_font + '.part', _font)
fm.fontManager.addfont(_font)
_have.add('Noto Sans TC')
except OSError as err:
print('Could not download the CJK font; Chinese text in figures may not display:', err)
plt.rcParams['font.family'] = [f for f in _cjk if f in _have] + ['DejaVu Sans']
plt.rcParams['axes.unicode_minus'] = False
C_CMAP = mcolors.LinearSegmentedColormap.from_list(
"nyu", ["#F4EEF9", C_AUX, C_V1, C_T1, C_V2]
)
DIVERGE_CMAP = mcolors.LinearSegmentedColormap.from_list(
"nyu_div", [C_V2, "#F8F8F8", C_WARN]
)
def ax_style(ax, title="", xlabel="", ylabel=""):
"""Apply the textbook style"""
ax.set_facecolor(C_BG)
ax.grid(True, color=C_GRID, linewidth=0.8, alpha=0.7)
ax.set_title(title, color=C_V1, fontsize=11, pad=8)
if xlabel: ax.set_xlabel(xlabel, color=C_AXIS, fontsize=9)
if ylabel: ax.set_ylabel(ylabel, color=C_AXIS, fontsize=9)
ax.tick_params(colors=C_AXIS, labelsize=8)
for spine in ax.spines.values():
spine.set_edgecolor(C_GRID)
print("✓ Environment initialized")Scenario 1: A Movie Recommender System¶
How the Data Are Generated¶
We set up 10 users (5 action fans and 5 romance fans) and 5 movies (2 action, 2 romance, and 1 mixed).
Each user is represented by a two-dimensional preference vector ( = preference for action, = preference for romance), and each movie by , its genre weights.
The true rating is determined by the inner product, with normal noise added:
This means that action fans naturally give action movies high ratings and romance movies low ratings, and vice versa. The data really do contain an interaction effect, so we can expect a part that the direct-sum model cannot capture.
Each user–movie combination produces 2 records, 100 records in all.
# ============================================================
# Scenario 1: data generation
# ============================================================
# --- Step 1: user preference vectors ---
# action fans [~0.9, ~0.1]; romance fans [~0.1, ~0.9]; a little noise keeps the pattern from being perfectly regular
users_feat = np.vstack([
np.clip(np.random.normal([0.85, 0.15], 0.10, (5, 2)), 0, 1),
np.clip(np.random.normal([0.15, 0.85], 0.10, (5, 2)), 0, 1),
])
user_names = [f"U{i+1:02d}" for i in range(10)]
user_types = ["Action fan"]*5 + ["Romance fan"]*5
# --- Step 2: movie genre vectors ---
movie_names = ["Fast & Furious", "John Wick", "Titanic", "Before Sunset", "Mr. & Mrs. Smith"]
movie_types = ["Action", "Action", "Romance", "Romance", "Mixed"]
movies_feat = np.array([
[1.0, 0.0], [0.9, 0.1],
[0.1, 0.9], [0.0, 1.0],
[0.5, 0.5],
])
# --- Step 3: true rating matrix (driven by inner products) ---
true_affinity = users_feat @ movies_feat.T # shape (10, 5), range ≈ [0, 1]
# --- Step 4: generate a raw log of 100 records (2 per combination, with noise) ---
records = []
for u in range(10):
for m in range(5):
for _ in range(2):
raw = np.random.normal(true_affinity[u, m] * 5, 0.4)
rating = round(np.clip(raw, 0.5, 5.0) * 2) / 2 # round to the nearest 0.5
records.append({
"user_id": u, "user": user_names[u], "user_type": user_types[u],
"movie_id": m, "movie": movie_names[m], "movie_type": movie_types[m],
"rating": rating
})
df = pd.DataFrame(records)
print(f"A total of {len(df)} records, ratings range {df.rating.min()}–{df.rating.max()}")
print(f"Global mean {df.rating.mean():.3f}, standard deviation {df.rating.std():.3f}")
print()
print("True inner-product preference matrix (rows = users, columns = movies):")
print(" ", " ".join(f"{n:6s}" for n in movie_names))
for i in range(10):
row = " ".join(f"{true_affinity[i,j]:6.3f}" for j in range(5))
print(f" {user_names[i]}({user_types[i][0]}) {row}")A. Shuffled vs. Sorted: The Same Data from Two Points of View¶
The figure below shows two visualizations of the same raw log of records: the left panel lists the users in a randomly shuffled order (simulating the order of the raw records, with action fans and romance fans mixed together), and the right panel re-sorts them by user type (action fans on top, romance fans at the bottom).
The data do not change at all; only the order of the rows differs. Yet how visible the pattern is differs greatly.
# ============================================================
# A. Shuffled vs. sorted: comparing heat maps of the rating matrix
# ============================================================
# Build the rating matrix (10 x 5; each cell is the mean of its two records)
R = df.pivot_table(values="rating", index="user_id", columns="movie_id", aggfunc="mean").values
# Shuffled order: randomly permute the rows (simulating the order of the original records)
shuffle_idx = np.random.permutation(10)
R_shuffled = R[shuffle_idx]
names_shuffled = [user_names[i] for i in shuffle_idx]
# Sorted order: action fans (0-4) on top, romance fans (5-9) at the bottom
sorted_idx = list(range(10)) # the action fans already come first
R_sorted = R[sorted_idx]
names_sorted = user_names
type_labels = [f"{n}\n({user_types[i][0]})" for i, n in enumerate(user_names)]
fig, axes = plt.subplots(1, 2, figsize=(12, 4.5))
fig.patch.set_facecolor("#FFFFFF")
for ax, R_show, ylabels, title in [
(axes[0], R_shuffled, [f"{n}({user_types[shuffle_idx[i]][0]})" for i,n in enumerate(names_shuffled)],
"① Shuffled order (original records)"),
(axes[1], R_sorted, type_labels,
"② Sorted by preference type (action fans↑ romance fans↓)"),
]:
im = ax.imshow(R_show, cmap=C_CMAP, vmin=0.5, vmax=5.0, aspect="auto")
ax.set_xticks(range(5))
ax.set_xticklabels(movie_names, fontsize=8, rotation=15, ha="right")
ax.set_yticks(range(10))
ax.set_yticklabels(ylabels, fontsize=8)
ax_style(ax, title=title)
ax.grid(False) # no grid on heat maps
# annotate the values
for i in range(10):
for j in range(5):
ax.text(j, i, f"{R_show[i,j]:.1f}", ha="center", va="center",
fontsize=7, color="white" if R_show[i,j] > 3.2 else C_AXIS)
plt.colorbar(im, ax=ax, fraction=0.03, pad=0.02, label="Rating")
# Mark the dividing line in the sorted panel
axes[1].axhline(4.5, color=C_WARN, linewidth=1.5, linestyle="--")
axes[1].text(4.6, 2.0, "Action fans", color=C_WARN, fontsize=8, va="center")
axes[1].text(4.6, 7.0, "Romance fans", color=C_WARN, fontsize=8, va="center")
fig.suptitle("The same data, two orderings — the tensor-product matrix $R$", fontsize=13, color=C_V1, y=1.01)
plt.tight_layout()
plt.show()
print("Observation: in the right panel, the high-rating block at the top left (action fans × action movies) and the one at the bottom right (romance fans × romance movies) are clearly visible;")
print(" the same information is present in the left panel, but once shuffled it is hard to notice. The data themselves have not changed.")B. The Direct-Sum Structure: Summing Marginals¶
The direct-sum model keeps only two sets of marginal statistics:
The prediction matrix is built from outer products, and its rank is at most 2 (in practice close to rank 1):
# ============================================================
# B. Direct-sum structure: marginal statistics + prediction matrix
# ============================================================
global_avg = df.rating.mean()
user_avg = df.groupby("user_id").rating.mean().values # (10,)
movie_avg = df.groupby("movie_id").rating.mean().values # (5,)
R_direct = np.clip(
np.outer(user_avg, np.ones(5)) + np.outer(np.ones(10), movie_avg) - global_avg,
0.5, 5.0
)
fig, axes = plt.subplots(1, 3, figsize=(13, 4))
fig.patch.set_facecolor("#FFFFFF")
# Left: user marginals
ax = axes[0]
colors_u = [C_V1 if t == "Action fan" else C_WARN for t in user_types]
ax.barh(user_names[::-1], user_avg[::-1], color=colors_u[::-1], alpha=0.85)
ax.axvline(global_avg, color=C_AUX, linewidth=1.2, linestyle="--", label=f"Global mean {global_avg:.2f}")
ax.legend(fontsize=8)
ax_style(ax, title="User marginals $\\bar{r}_{i\\cdot}$ (averaged over $j$)", xlabel="Average rating")
ax.set_xlim(0, 5.5)
# Annotate the user types
for i, (name, utype) in enumerate(zip(user_names[::-1], user_types[::-1])):
ax.text(0.15, i, utype[0], va="center", fontsize=7, color="white", fontweight="bold")
# Middle: movie marginals
ax = axes[1]
mcolors_bar = [C_V1, C_V1, C_WARN, C_WARN, C_T1]
ax.bar(movie_names, movie_avg, color=mcolors_bar, alpha=0.85)
ax.axhline(global_avg, color=C_AUX, linewidth=1.2, linestyle="--")
ax_style(ax, title="Movie marginals $\\bar{r}_{\\cdot j}$ (averaged over $i$)", ylabel="Average rating")
ax.set_xticklabels(movie_names, rotation=15, ha="right", fontsize=8)
ax.set_ylim(0, 5.5)
for j, (name, mtype) in enumerate(zip(movie_names, movie_types)):
ax.text(j, 0.2, mtype[0], ha="center", fontsize=8, color="white", fontweight="bold")
# Right: direct-sum prediction matrix (sorted by type)
ax = axes[2]
im = ax.imshow(R_direct[sorted_idx], cmap=C_CMAP, vmin=0.5, vmax=5.0, aspect="auto")
ax.set_xticks(range(5)); ax.set_xticklabels(movie_names, rotation=15, ha="right", fontsize=8)
ax.set_yticks(range(10)); ax.set_yticklabels(type_labels, fontsize=8)
ax.axhline(4.5, color=C_WARN, linewidth=1.2, linestyle="--")
for i in range(10):
for j in range(5):
v = R_direct[sorted_idx[i], j]
ax.text(j, i, f"{v:.1f}", ha="center", va="center",
fontsize=7, color="white" if v > 3.2 else C_AXIS)
ax_style(ax, title="Direct-sum prediction matrix $\\hat{R}$ (rank ≤ 2)")
ax.grid(False)
plt.colorbar(im, ax=ax, fraction=0.03, pad=0.02, label="Predicted rating")
fig.suptitle("B. Direct sum $U \\oplus M$: summing marginals — averaging over users and over movies separately compresses away the joint information",
fontsize=12, color=C_V1, y=1.02)
plt.tight_layout()
plt.show()
print(f"Rank of the direct-sum prediction matrix = {np.linalg.matrix_rank(R_direct)}")
print("Note: each row of the direct-sum matrix (along the movie dimension) only shifts up or down as a whole;")
print(" the predicted distributions of action fans and romance fans have exactly the same shape.")C. The Tensor-Product Structure: Keeping the Joint Information¶
The tensor-product matrix records directly (the observed mean for each user–movie pair). From it we can read off the conditional distribution , and also whether, once we fix the user, the shape of the rating curve over the movies differs from person to person. This is precisely what the direct-sum structure cannot express.
# ============================================================
# C. Tensor product vs. direct sum: comparing conditional distributions + interaction-effect matrix
# ============================================================
# Two representative users: U01 (action fan) and U06 (romance fan)
u_act, u_rom = 0, 5
cond_direct_act = user_avg[u_act] + movie_avg - global_avg # conditional prediction of the direct sum
cond_direct_rom = user_avg[u_rom] + movie_avg - global_avg
cond_tensor_act = R[u_act] # observed values in the tensor product
cond_tensor_rom = R[u_rom]
E = R - R_direct # interaction-effect matrix
fig = plt.figure(figsize=(14, 4.5))
fig.patch.set_facecolor("#FFFFFF")
gs = GridSpec(1, 3, figure=fig, wspace=0.35)
# Left: conditional distributions of the direct sum (shifted)
ax0 = fig.add_subplot(gs[0])
x = range(5)
ax0.plot(x, cond_direct_act, "o--", color=C_V2, linewidth=2, markersize=7,
label=f"{user_names[u_act]} (action fan)")
ax0.plot(x, cond_direct_rom, "s--", color=C_WARN, linewidth=2, markersize=7,
label=f"{user_names[u_rom]} (romance fan)")
ax0.fill_between(x, cond_direct_act, cond_direct_rom, alpha=0.08, color=C_AUX)
ax0.set_xticks(x); ax0.set_xticklabels(movie_names, rotation=15, ha="right", fontsize=8)
ax0.set_ylim(0, 5.5)
ax0.legend(fontsize=8)
ax_style(ax0, title="Direct sum: conditional curves have the same shape (shifted only)", ylabel="Predicted rating")
# Middle: conditional distributions of the tensor product (different shapes)
ax1 = fig.add_subplot(gs[1])
ax1.plot(x, cond_tensor_act, "o-", color=C_V1, linewidth=2.5, markersize=8,
label=f"{user_names[u_act]} (action fan)")
ax1.plot(x, cond_tensor_rom, "s-", color=C_T1, linewidth=2.5, markersize=8,
label=f"{user_names[u_rom]} (romance fan)")
ax1.fill_between(x, cond_tensor_act, cond_tensor_rom, alpha=0.10, color=C_V1)
ax1.set_xticks(x); ax1.set_xticklabels(movie_names, rotation=15, ha="right", fontsize=8)
ax1.set_ylim(0, 5.5)
ax1.legend(fontsize=8)
ax_style(ax1, title="Tensor product: conditional curves have very different shapes", ylabel="Observed rating")
# Right: interaction-effect matrix
ax2 = fig.add_subplot(gs[2])
vmax = max(abs(E.max()), abs(E.min()))
im = ax2.imshow(E[sorted_idx], cmap=DIVERGE_CMAP, vmin=-vmax, vmax=vmax, aspect="auto")
ax2.set_xticks(range(5)); ax2.set_xticklabels(movie_names, rotation=15, ha="right", fontsize=8)
ax2.set_yticks(range(10)); ax2.set_yticklabels(type_labels, fontsize=8)
ax2.axhline(4.5, color=C_AXIS, linewidth=1, linestyle=":")
for i in range(10):
for j in range(5):
v = E[sorted_idx[i], j]
ax2.text(j, i, f"{v:+.1f}", ha="center", va="center",
fontsize=7, color="white" if abs(v) > vmax*0.5 else C_AXIS)
ax2.grid(False)
plt.colorbar(im, ax=ax2, fraction=0.03, pad=0.02, label="Interaction effect")
ax_style(ax2, title="Interaction effect $E = R - \\hat{R}$ (orange = above prediction, blue = below)")
fig.suptitle("C. Tensor product $U \\otimes M$: keeping the joint information captures the interaction effect",
fontsize=12, color=C_V1, y=1.02)
plt.tight_layout()
plt.show()
print(f"Standard deviation of the interaction effect = {E.std():.3f} (the size of the systematic error of the direct sum)")
print(f"Largest positive effect {E.max():+.2f} (an action fan's extra liking for an action movie)")
print(f"Largest negative effect {E.min():+.2f} (an action fan's extra dislike of a romance movie)")Scenario 2: Detecting Collusion at a Casino Roulette Wheel¶
How the Data Are Generated¶
A simplified roulette wheel has 6 numbers (0–5). In each round the player bets on one number, the wheel produces one number, and a hit is paid at 5 to 1. This is exactly the fair payout with zero house edge: if everyone plays honestly, the house neither gains nor loses.
The tensor product in this scenario has two axes:
the number the player bets on (the player’s betting choice)
the number the wheel lands on (the casino’s actual outcome)
On a fair wheel these two should be statistically independent: which number you bet on does not affect where the wheel stops. But this casino has an insider who has rigged the game. In about 26% of the rounds (the collusion rounds), a colluding player bets on one of two “signal numbers” , and at the same time the wheel is steered to the same number, so the bet and the outcome are locked together.
The clever part is this: to avoid being caught by a “marginal audit,” in the remaining honest rounds the insider deliberately suppresses the signal numbers: players bet less on , and the wheel lands less often on . Once the surge in the collusion rounds and the dip in the honest rounds offset each other, the marginal distributions of both the bets and the outcomes stay close to uniform, and the table looks no different from any ordinary one.
The data-generation logic:
With probability , play a collusion round: bet outcome one of the signal numbers;
With probability , play an honest round: as cover, the player and the wheel sample independently from a distribution that “suppresses the signal numbers”;
A hit is paid at 5 to 1; on a miss the house collects the stake.
The question: if risk control watches only the two marginal distributions (“which numbers players like to bet on” and “which numbers the wheel often lands on”), can it catch the collusion? Or is it impossible without looking at the joint distribution (the tensor product)?
An intuition trap: Thanks to the cover, both marginal distributions look close to uniform, and any audit that checks only “how often each single number appears” will let the table pass. The hard evidence of collusion lies in no single marginal; it lies in the joint phenomenon that “the bet and the outcome are locked to the same number,” which exists only in the tensor-product structure.
# ============================================================
# Scenario 2: generating the casino collusion data (bet × outcome)
# ============================================================
np.random.seed(7)
N_WHEEL = 6 # simplified wheel: numbers 0..5
SECRET = [1, 4] # the insider's signal numbers
PAYOUT = 5 # 5:1 payout (fair payout = 6 - 1)
EPS = 0.26 # fraction of collusion rounds
# Collusion piles the mass of the signal numbers {1,4} onto the diagonal, which also raises the marginals of these two numbers.
# To avoid detection by a "marginal audit", the insider deliberately suppresses the signal numbers in the honest rounds as cover,
# so that the surge in the collusion rounds and the dip in the honest rounds cancel out — in the end the bet/outcome marginals are both close to uniform.
b_diff = EPS / (2 * (1 - EPS)) # offset to be compensated
b_secret = (1 - 4 * b_diff) / 6 # (suppressed) probability of a signal number in the honest rounds
b_other = b_secret + b_diff # probability of each other number
p_honest = np.array([b_secret if k in SECRET else b_other
for k in range(N_WHEEL)]) # bet/outcome distribution in the honest rounds
roulette_records = []
for _ in range(800):
if np.random.random() < EPS: # collusion round: bet = outcome = signal number
k = np.random.choice(SECRET)
bet, out, colluded = k, k, True
else: # honest round: bet and outcome independent (both suppress the signal numbers)
bet = np.random.choice(N_WHEEL, p=p_honest)
out = np.random.choice(N_WHEEL, p=p_honest)
colluded = False
won = (bet == out)
profit = PAYOUT if won else -1 # player's gain or loss per 1-unit bet
roulette_records.append({
"bet_num": bet, "out_num": out,
"won": won, "player_profit": profit, "colluded": colluded
})
rdf = pd.DataFrame(roulette_records)
print(f"A total of {len(rdf)} rounds")
print(f"Distribution of bets: {rdf.bet_num.value_counts().sort_index().to_dict()}")
print(f"Distribution of outcomes: {rdf.out_num.value_counts().sort_index().to_dict()}")
print(f"Overall hit rate {rdf.won.mean():.3f} (fair baseline {1/N_WHEEL:.3f})")
print(f"Player's expected gain per bet {rdf.player_profit.mean():+.3f} (>0 means the house is losing money)")# ============================================================
# Scenario 2: direct sum (independence assumption) vs. tensor product (joint distribution)
# ============================================================
n = N_WHEEL
nums = [str(k) for k in range(n)]
# Marginal distributions
P_bet = rdf.bet_num.value_counts().reindex(range(n), fill_value=0).values.astype(float)
P_out = rdf.out_num.value_counts().reindex(range(n), fill_value=0).values.astype(float)
P_bet /= P_bet.sum()
P_out /= P_out.sum()
# Joint distribution (tensor product)
joint = rdf.pivot_table(index="bet_num", columns="out_num",
values="won", aggfunc="size", fill_value=0
).reindex(index=range(n), columns=range(n), fill_value=0).values.astype(float)
joint /= joint.sum()
# Direct-sum prediction (independence assumption) and anomaly matrix
direct = np.outer(P_bet, P_out)
anomaly = joint - direct # anomaly = actual joint - independence assumption
hit_fair = 1 / n
hit_direct = np.trace(direct) # hit rate under the independence assumption
hit_actual = np.trace(joint) # actual hit rate (sum of the diagonal)
fig, axes = plt.subplots(2, 3, figsize=(14, 9))
fig.patch.set_facecolor("#FFFFFF")
def heat(ax, M, title, cmap, vmin, vmax, cbar_label, diverge=False):
im = ax.imshow(M, cmap=cmap, vmin=vmin, vmax=vmax, aspect="auto")
ax.set_xticks(range(n)); ax.set_xticklabels(nums, fontsize=9)
ax.set_yticks(range(n)); ax.set_yticklabels(nums, fontsize=9)
ax.set_xlabel("Outcome number", color=C_AXIS, fontsize=9)
ax.set_ylabel("Bet number", color=C_AXIS, fontsize=9)
thr = vmax * 0.5
for i in range(n):
for j in range(n):
v = M[i, j]
txt = f"{v:+.2f}" if diverge else f"{v:.2f}"
ax.text(j, i, txt, ha="center", va="center", fontsize=7,
color="white" if abs(v) > thr else C_AXIS)
for k in SECRET: # outline the diagonal cells of the signal numbers
ax.add_patch(plt.Rectangle((k-0.5, k-0.5), 1, 1, fill=False,
edgecolor=C_WARN, linewidth=2.2))
ax.grid(False)
plt.colorbar(im, ax=ax, fraction=0.046, label=cbar_label)
ax_style(ax, title=title)
# Top left: marginal of the bets (deliberately in a single color — no number looks suspicious)
ax = axes[0, 0]
ax.bar(nums, P_bet, color=C_V2, alpha=0.85)
ax.axhline(1/n, color=C_WARN, linewidth=1.4, linestyle="--", label=f"Uniform {1/n:.3f}")
ax.legend(fontsize=8)
ax_style(ax, title="Marginal of the bets $P(\\text{bet})$ (looks normal)", xlabel="Number", ylabel="Frequency")
ax.set_ylim(0, 0.26)
# Top middle: marginal of the outcomes
ax = axes[0, 1]
ax.bar(nums, P_out, color=C_V2, alpha=0.85)
ax.axhline(1/n, color=C_WARN, linewidth=1.4, linestyle="--", label=f"Uniform {1/n:.3f}")
ax.legend(fontsize=8)
ax_style(ax, title="Marginal of the outcomes $P(\\text{outcome})$ (looks normal)", xlabel="Number", ylabel="Frequency")
ax.set_ylim(0, 0.26)
# Top right: joint predicted by the direct sum (independence assumption) — the marginals are nearly uniform, so the prediction is flat
heat(axes[0, 2], direct, "Direct sum: $P(\\text{bet})\\times P(\\text{outcome})$ (independence assumption)",
C_CMAP, 0, 0.12, "Probability")
# Bottom left: actual joint of the tensor product — the diagonal cells of the signal numbers bulge abnormally
heat(axes[1, 0], joint, "Tensor product: $P(\\text{bet},\\,\\text{outcome})$ (actual joint)",
C_CMAP, 0, 0.12, "Probability")
# Bottom middle: anomaly matrix (actual - independence assumption) — only the two diagonal cells of the signal numbers light up
vmax_a = np.abs(anomaly).max()
heat(axes[1, 1], anomaly, "Anomaly $=$ actual joint $-$ independence assumption",
DIVERGE_CMAP, -vmax_a, vmax_a, "Deviation", diverge=True)
# Bottom right: comparing hit rates
ax = axes[1, 2]
labels = ["Fair baseline", "Direct-sum prediction", "Actual observation"]
vals = [hit_fair, hit_direct, hit_actual]
bars = ax.bar(labels, vals, color=[C_AUX, C_V2, C_WARN], alpha=0.85)
for b, v in zip(bars, vals):
ax.text(b.get_x() + b.get_width()/2, v + 0.008, f"{v:.3f}",
ha="center", fontsize=9, color=C_AXIS)
ax_style(ax, title="Hit rate: the direct sum agrees exactly with the fair baseline, yet is badly wrong", ylabel="Hit rate")
ax.set_ylim(0, max(vals) * 1.25)
fig.suptitle("Scenario 2 (casino): each marginal looks innocent, but the joint distribution reveals the collusion",
fontsize=13, color=C_V1, y=1.00)
plt.tight_layout()
plt.show()
print("Direct sum (independence assumption) vs. tensor product (joint distribution) — diagonal cells of the signal numbers:")
for k in SECRET:
print(f" bet {k}, outcome {k}: actual {joint[k,k]:.3f} direct-sum prediction {direct[k,k]:.3f}"
f" (actual = {joint[k,k]/direct[k,k]:.1f} × direct sum)")
print(f"\nHit rate: fair {hit_fair:.3f} → direct-sum prediction {hit_direct:.3f} → actual {hit_actual:.3f}")
print(f"If the house looks only at the marginals/direct sum, it will think this is a fair table (player's expectation per bet {hit_direct*PAYOUT-(1-hit_direct):+.3f}, close to 0);")
print(f"in reality it is {hit_actual*PAYOUT-(1-hit_actual):+.3f} (a heavy loss). The whole gap lies in two diagonal cells — hard evidence of collusion.")Summary¶
| Direct sum | Tensor product | |
|---|---|---|
| Information kept | Marginal distributions , | Joint distribution |
| Implicit assumption | User and movie preferences are statistically independent | No independence assumption |
| Rank of the matrix | ||
| When they agree | (that is, no interaction effect) | — |
| Casino example | Knows only “which numbers players like to bet on” and “which numbers the wheel often lands on” | Knows “whether the bet and the outcome are locked together” (the evidence of collusion) |
The rows of the direct-sum prediction matrix (each with a fixed user) differ from one another only by a constant offset, and the curves have exactly the same shape. The tensor product keeps each user’s personalized distribution, whose shape differs from person to person.
Preview of Chapter 11. Tensor-product matrices are often redundant: “the rows of all the action fans are similar” means that the effective dimension of the matrix is far smaller than . The SVD (singular value decomposition) is precisely the tool for measuring this redundancy: it decomposes the tensor-product matrix into a sum of rank-one matrices (each a multiplicative pure tensor , not the additive direct-sum structure of Part B), ordered by importance, which makes low-rank approximation and data compression possible.
Exercises¶
# ============================================================
# Exercise — Question 2: testing the independence of bet and outcome
# ============================================================
# Independence assumption: if the bet and the outcome are independent (fair wheel),
# then joint should equal direct = np.outer(P_bet, P_out), and anomaly should be close to 0 everywhere.
print("=== Marginal distributions of bets/outcomes (look normal) ===\n")
print(f"P_bet: {dict(zip(nums, P_bet.round(3)))}")
print(f"P_out: {dict(zip(nums, P_out.round(3)))}")
print()
print("=== Anomaly matrix anomaly = joint - direct (rows = bets, columns = outcomes) ===")
print(np.round(anomaly, 3))
print(f"\nSignal numbers SECRET = {SECRET}\n")
# TODO 1: find the cell of anomaly with the largest absolute value and check whether it lies at a signal-number position on the diagonal (bet = outcome)
max_cell = None # hint: np.unravel_index(np.argmax(np.abs(anomaly)), anomaly.shape)
print(f"Cell with the largest deviation (bet, outcome) = {max_cell}")
print("# Does it lie at a signal-number position {1,4} on the diagonal?")
print()
# TODO 2: compare the actual hit rate with the hit rate predicted by the direct sum
gap = None # hint: np.trace(joint) - np.trace(direct)
print(f"Actual hit rate - direct-sum hit rate = {gap}")
print(f"(direct-sum hit rate trace(direct) = {np.trace(direct):.3f}, which comes out close to the fair baseline {1/N_WHEEL:.3f})")
print("# Why does the direct-sum hit rate come out equal to the fair baseline? Why does the \"marginal audit\" fail completely here?")
print("# Answer: _______")
print()
# TODO 3: look at the values of the signal numbers {1,4} in P_bet and P_out above
print("# Do the signal numbers deviate from the other numbers clearly enough to convict on the marginals alone?")
print("# Conclusion: the evidence of collusion exists only in the joint distribution (joint / anomaly), not in any single marginal.")