Orthogonal Projections, Complements & the Best-Behaved Matrices in the Course
How to drop a perpendicular onto a subspace, split any vector into "in U" and "perpendicular to U," and discover why symmetric matrices are always, effortlessly diagonalizable
§1Motivation — Getting as Close as Possible
Every lecture since Week 5 has been building toward one very human question: given a target you cannot reach exactly, what is the best you can do? You cannot always solve $A\mathbf{x}=\mathbf{b}$ exactly — but you can always find the closest point in a subspace to a point outside it. That single idea, "closest point," is the entire engine behind GPS positioning, noise-cancelling headphones, recommendation engines, and — as you will use constantly in later courses — least-squares regression, where you fit a line or curve that cannot pass through every data point but gets as close as mathematically possible to all of them at once.
§2Part A — Shortest Distance From a Point to a Plane
Here is the motivating problem for the entire lecture, stated as simply as possible: determine the shortest distance between a point $P$ and a subspace $U$ through the origin (in $\mathbb{R}^3$, picture $U$ as a plane through the origin, and $P$ as a point floating somewhere off that plane).
The dashed segment from $P$ meets the plane $U$ at a right angle. Its foot is the orthogonal projection $\text{Proj}_U(P)$ — and that foot is the closest point in $U$ to $P$.
Intuitively — and this is exactly what the picture shows — the shortest path from $P$ down to the plane is a straight perpendicular drop. Where that perpendicular lands is a special point in $U$, called the orthogonal projection of $P$ onto $U$, written $\text{Proj}_U(P)$.
$$\text{shortest distance from } P \text{ to } U \;=\; \text{distance from } P \text{ to } \text{Proj}_U(P) \;=\; \|P - \text{Proj}_U(P)\|.$$
§3Example 8.1.4 — Finding the Closest Point in a Plane
Find the point in the plane $U = \{(x,y,z) : 2x+y-z=0\}$ closest to the point $P=(2,-1,-3)$, and find that shortest distance.
§4Part C — Orthogonal Complements
If $U$ is a subspace of $\mathbb{R}^n$, its orthogonal complement is $$U^{\perp} = \{\,\mathbf{v} \text{ in } \mathbb{R}^n : \mathbf{v}\cdot\mathbf{u}=0 \text{ for every } \mathbf{u} \text{ in } U\,\} —$$ the set of every vector that is orthogonal to all of $U$ at once, not just to one vector in it.
$V = U \oplus U^{\perp}$ — every vector in $V$ splits uniquely into a piece that lies in $U$ and a piece that lies in $U^{\perp}$, and the two pieces are orthogonal to each other.
Two facts make $U^{\perp}$ worth naming: it is always itself a subspace (a short closure check, exactly like every other subspace proof you've done), and together $U$ and $U^{\perp}$ split the whole space with no overlap and no gaps — every vector decomposes uniquely into a piece from each, which is exactly what Part D will demonstrate by hand.
§5Finding a Basis of U⊥ — the Easy Way, for a Plane
Look again at how $U$ was defined: $U = \{(x,y,z) : 2x+y-z=0\}$. The left-hand side $2x+y-z$ is literally the dot product of $(x,y,z)$ with the fixed vector $\mathbf{n}=(2,1,-1)$. So the defining equation of $U$ is a dot-product-equals-zero condition:
$$U = \{\mathbf{v} : \mathbf{n}\cdot\mathbf{v}=0\}.$$
That says every vector in $U$ is already orthogonal to $\mathbf{n}$ — so $\mathbf{n}$ itself is orthogonal to all of $U$, meaning $\mathbf{n}\in U^{\perp}$. Since $U$ is $2$-dimensional inside $\mathbb{R}^3$, its complement must be $3-2=1$-dimensional, and one nonzero vector in $U^{\perp}$ is enough to span it:
$$U^{\perp} = \operatorname{span}\{(2,1,-1)\}.$$
This matches the leftover vector we found by direct computation two sections ago — no coincidence at all.
§6Part D — Decomposing a Vector into U and U⊥
Here is the general skill Part A's projection formula was secretly teaching: splitting any vector into its "inside $U$" piece and its "perpendicular to $U$" piece.
Let $U = \operatorname{span}\{(1,-2,3),\,(-1,1,1)\}$. Write $(1,5,7)$ as the sum of a vector in $U$ and a vector in $U^{\perp}$.
§7Part E — Symmetric Matrices
We now turn from subspaces to a single, extraordinary class of matrices — the reward for everything built so far.
A square matrix $A$ is symmetric if $A^T = A$ — that is, $A$ equals its own transpose, so $a_{ij}=a_{ji}$ for every $i,j$: the matrix is a mirror image of itself across its main diagonal.
§8Proof — Eigenvectors of a Symmetric Matrix Are Orthogonal
Let $A$ be symmetric, and suppose $A\mathbf{x}=\lambda\mathbf{x}$ and $A\mathbf{y}=\mu\mathbf{y}$ for eigenvalues $\lambda \neq \mu$. We compute the single number $\mathbf{x}^T A\mathbf{y}$ two different ways and compare.
$\textbf{First way — substitute } A\mathbf{y}=\mu\mathbf{y}$ directly:
$$\mathbf{x}^TA\mathbf{y} = \mathbf{x}^T(\mu\mathbf{y}) = \mu\,\mathbf{x}^T\mathbf{y}.$$
$\textbf{Second way — use symmetry, } A^T=A$, to move $A$ onto $\mathbf{x}$ instead:
$$\mathbf{x}^TA\mathbf{y} = (A^T\mathbf{x})^T\mathbf{y} = (A\mathbf{x})^T\mathbf{y} = (\lambda\mathbf{x})^T\mathbf{y} = \lambda\,\mathbf{x}^T\mathbf{y}.$$
$\textbf{Both computed the same quantity, so they must agree:}$
$$\lambda\,\mathbf{x}^T\mathbf{y} = \mu\,\mathbf{x}^T\mathbf{y} \quad\Longrightarrow\quad (\lambda-\mu)\,\mathbf{x}^T\mathbf{y}=0.$$
Since $\lambda\neq\mu$, the factor $(\lambda-\mu)$ is nonzero, which forces the other factor to vanish:
$$\mathbf{x}^T\mathbf{y}=0, \quad \text{i.e.} \quad \mathbf{x}\cdot\mathbf{y}=0.$$
$\blacksquare$ Eigenvectors of a symmetric matrix, corresponding to distinct eigenvalues, are automatically orthogonal. No Gram–Schmidt is even needed across different eigenspaces — the symmetry of $A$ does that work for you.
§9Theorem — Orthogonal Diagonalization
Every symmetric matrix $A$ can be written as
$$A = PDP^T$$
where $D$ is the diagonal matrix of eigenvalues of $A$, and $P$ is an orthogonal matrix — its columns are the corresponding eigenvectors of $A$, each first normalized to unit length before being placed in $P$.
$$P^{-1} = P^T.$$
Transposing a matrix costs nothing — you just relabel rows as columns. No inversion, ever. This is the entire practical reason orthogonal diagonalization is worth learning as its own theorem rather than a special case of ordinary diagonalization: for symmetric matrices, which are everywhere in applications (covariance matrices in statistics, stress tensors in engineering, adjacency matrices of undirected graphs, Hessians in optimization), diagonalizing is dramatically cheaper.§10A Fully Worked Orthogonal Diagonalization
Find an orthogonal matrix $P$ and diagonal matrix $D$ such that $A=PDP^T$, for $A=\begin{bmatrix}2&1\\1&2\end{bmatrix}$.
§11Exercises
Three problems, one for each idea in this lecture. Hints only — work through the method yourself before checking.
Find the point in the plane $x-y+2z=0$ closest to $(3,0,1)$, and the shortest distance.
Let $U = \operatorname{span}\{(2,0,-1),\,(1,3,2)\}$. Write $(4,1,5)$ as the sum of a vector in $U$ and a vector in $U^{\perp}$.
Find an orthogonal matrix $P$ and diagonal matrix $D$ with $A=PDP^T$, for $A=\begin{bmatrix}3&1\\1&3\end{bmatrix}$.
Drop a perpendicular, split a vector in two, and watch a symmetric matrix diagonalize itself for free.
The orthogonal projection formula, the direct-sum split $V=U\oplus U^{\perp}$, and the spectral theorem for symmetric matrices are not three separate topics — they are one idea, applied three times. Once you have an orthogonal basis, everything in linear algebra gets cheaper: coordinates become dot products, projections become sums, and matrix inversion becomes a transpose.
The Final Exam is worth 80 marks. There appears to be a partial midterm-replacement policy for students who scored 40 or above on the midterm, involving a ratio structure (roughly 20:80 and 1:4) that governs how final-exam performance can offset the midterm grade. The exact mechanics are unclear from the lecture notes this was transcribed from — please confirm the precise wording, thresholds, and eligibility directly with the instructor rather than relying on this summary.