Home›Courses›Linear Algebra›Week 6 · Lecture 19
← Lecture 18
On this pageMotivation & History·Shortest Distance·Example 8.1.4·Orthogonal Complements·Basis of U⊥·Decomposing a Vector·Symmetric Matrices·Eigenvectors Are Orthogonal·Orthogonal Diagonalization·Worked Diagonalization·Exercises
Lecture 20 →
MATH-120 · Linear Algebra · Lecture 1913 July 2026

Orthogonal Projections, Complements & the Best-Behaved Matrices in the Course

How to drop a perpendicular onto a subspace, split any vector into "in U" and "perpendicular to U," and discover why symmetric matrices are always, effortlessly diagonalizable

Week 6 begins — Monday, 13 July 2026

§1Motivation — Getting as Close as Possible

Every lecture since Week 5 has been building toward one very human question: given a target you cannot reach exactly, what is the best you can do? You cannot always solve $A\mathbf{x}=\mathbf{b}$ exactly — but you can always find the closest point in a subspace to a point outside it. That single idea, "closest point," is the entire engine behind GPS positioning, noise-cancelling headphones, recommendation engines, and — as you will use constantly in later courses — least-squares regression, where you fit a line or curve that cannot pass through every data point but gets as close as mathematically possible to all of them at once.

📜
A bit of history — principal axes, 1829
Long before anyone wrote the words "eigenvector" or "orthogonal complement," physicists needed to solve a very physical problem: a spinning rigid body (a planet, a wheel, a tumbling asteroid) has a moment of inertia described by a symmetric matrix. The body spins most naturally about certain special, mutually perpendicular axes — its principal axes — and finding them was a real engineering necessity, not an abstract exercise. In 1829, Augustin-Louis Cauchy proved what we now call the spectral theorem for symmetric matrices: that these principal axes always exist, are always perpendicular to each other, and the matrix always simplifies completely in that rotated frame. What you'll prove today, in a few clean lines, is the same theorem Cauchy needed heavy 19th-century machinery for.
🧭
Where the Gram–Schmidt process comes in
The names Jørgen Pedersen Gram (a Danish actuary, 1883) and Erhard Schmidt (a German mathematician, 1907) are attached to the orthogonalization process you already built in Lecture 18. Today it stops being a standalone trick and becomes the tool that makes everything else in this lecture computable — every orthogonal projection below is built from an orthogonal basis, and Gram–Schmidt is how you get one. If you need a refresher on the mechanics, Lecture 18, §"Gram–Schmidt: How" has the full derivation — we will not repeat it here, only use it.

§2Part A — Shortest Distance From a Point to a Plane

Here is the motivating problem for the entire lecture, stated as simply as possible: determine the shortest distance between a point $P$ and a subspace $U$ through the origin (in $\mathbb{R}^3$, picture $U$ as a plane through the origin, and $P$ as a point floating somewhere off that plane).

UOProjU(P)P

The dashed segment from $P$ meets the plane $U$ at a right angle. Its foot is the orthogonal projection $\text{Proj}_U(P)$ — and that foot is the closest point in $U$ to $P$.

Intuitively — and this is exactly what the picture shows — the shortest path from $P$ down to the plane is a straight perpendicular drop. Where that perpendicular lands is a special point in $U$, called the orthogonal projection of $P$ onto $U$, written $\text{Proj}_U(P)$.

The key fact

$$\text{shortest distance from } P \text{ to } U \;=\; \text{distance from } P \text{ to } \text{Proj}_U(P) \;=\; \|P - \text{Proj}_U(P)\|.$$

⚠️
The one thing you must not skip
To actually compute $\text{Proj}_U(P)$, you need an orthogonal basis of $U$ — not just any basis. The projection formula below only works term-by-term because the basis vectors don't interfere with each other; feed it a non-orthogonal basis and the formula silently gives the wrong answer. This is precisely why Gram–Schmidt exists: it is the machine that turns any basis of $U$ into an orthogonal one, and it is a required step, not an optional cleanup, every single time you project onto a subspace of dimension $2$ or higher.

§3Example 8.1.4 — Finding the Closest Point in a Plane

★ 1Example 8.1.4 — closest point to (2,-1,-3) in the plane 2x+y-z=0

Find the point in the plane $U = \{(x,y,z) : 2x+y-z=0\}$ closest to the point $P=(2,-1,-3)$, and find that shortest distance.

👀
Notice something remarkable
The leftover vector $P - \text{Proj}_U(P) = (2,1,-1)$ is exactly the normal vector of the plane $2x+y-z=0$. That is not a coincidence — it is a perfect first example of the idea we build formally next: the leftover piece of $P$, after removing everything that lies in $U$, lands in a special subspace perpendicular to all of $U$. That subspace has a name — the orthogonal complement of $U$ — and we meet it right now.

§4Part C — Orthogonal Complements

Orthogonal complement

If $U$ is a subspace of $\mathbb{R}^n$, its orthogonal complement is $$U^{\perp} = \{\,\mathbf{v} \text{ in } \mathbb{R}^n : \mathbf{v}\cdot\mathbf{u}=0 \text{ for every } \mathbf{u} \text{ in } U\,\} —$$ the set of every vector that is orthogonal to all of $U$ at once, not just to one vector in it.

UU⊥(U-perp)V

$V = U \oplus U^{\perp}$ — every vector in $V$ splits uniquely into a piece that lies in $U$ and a piece that lies in $U^{\perp}$, and the two pieces are orthogonal to each other.

Two facts make $U^{\perp}$ worth naming: it is always itself a subspace (a short closure check, exactly like every other subspace proof you've done), and together $U$ and $U^{\perp}$ split the whole space with no overlap and no gaps — every vector decomposes uniquely into a piece from each, which is exactly what Part D will demonstrate by hand.

§5Finding a Basis of U⊥ — the Easy Way, for a Plane

Example 2A basis of U⊥ for U = {(x,y,z) : 2x+y-z=0}

Look again at how $U$ was defined: $U = \{(x,y,z) : 2x+y-z=0\}$. The left-hand side $2x+y-z$ is literally the dot product of $(x,y,z)$ with the fixed vector $\mathbf{n}=(2,1,-1)$. So the defining equation of $U$ is a dot-product-equals-zero condition:

$$U = \{\mathbf{v} : \mathbf{n}\cdot\mathbf{v}=0\}.$$

That says every vector in $U$ is already orthogonal to $\mathbf{n}$ — so $\mathbf{n}$ itself is orthogonal to all of $U$, meaning $\mathbf{n}\in U^{\perp}$. Since $U$ is $2$-dimensional inside $\mathbb{R}^3$, its complement must be $3-2=1$-dimensional, and one nonzero vector in $U^{\perp}$ is enough to span it:

$$U^{\perp} = \operatorname{span}\{(2,1,-1)\}.$$

This matches the leftover vector we found by direct computation two sections ago — no coincidence at all.

🧠
Why this shortcut works — and when it stops working
A plane through the origin in $\mathbb{R}^3$ is always defined by a single dot-product-equals-zero condition, and its normal vector always spans the (one-dimensional) orthogonal complement — no Gram–Schmidt needed, because there's only one direction to worry about. This is a special case of a general fact: $\dim U + \dim U^{\perp} = \dim V$ always holds. But for a higher dimensional $U$ (say a $3$-dimensional subspace of $\mathbb{R}^5$), reading off $U^{\perp}$ isn't a one-line trick — you would typically need to solve a homogeneous system (find every vector orthogonal to a spanning set of $U$) and, if you want an orthogonal basis of the result, run Gram–Schmidt on whatever basis that system produces.

§6Part D — Decomposing a Vector into U and U⊥

Here is the general skill Part A's projection formula was secretly teaching: splitting any vector into its "inside $U$" piece and its "perpendicular to $U$" piece.

★ 3Write (1,5,7) as a sum of a vector in U and a vector in U⊥

Let $U = \operatorname{span}\{(1,-2,3),\,(-1,1,1)\}$. Write $(1,5,7)$ as the sum of a vector in $U$ and a vector in $U^{\perp}$.

§7Part E — Symmetric Matrices

We now turn from subspaces to a single, extraordinary class of matrices — the reward for everything built so far.

Symmetric matrix

A square matrix $A$ is symmetric if $A^T = A$ — that is, $A$ equals its own transpose, so $a_{ij}=a_{ji}$ for every $i,j$: the matrix is a mirror image of itself across its main diagonal.

✨
The headline fact
A symmetric matrix is always diagonalizable — no exceptions, no "if the eigenvalues happen to work out." Every symmetric matrix, no matter how it's built, has a full set of eigenvectors, and — as we're about to prove — those eigenvectors arrange themselves into orthogonal directions almost for free. One instructor's handwritten notes call this section "Purely Magical." It's a fair description.

§8Proof — Eigenvectors of a Symmetric Matrix Are Orthogonal

Let $A$ be symmetric, and suppose $A\mathbf{x}=\lambda\mathbf{x}$ and $A\mathbf{y}=\mu\mathbf{y}$ for eigenvalues $\lambda \neq \mu$. We compute the single number $\mathbf{x}^T A\mathbf{y}$ two different ways and compare.

$\textbf{First way — substitute } A\mathbf{y}=\mu\mathbf{y}$ directly:

$$\mathbf{x}^TA\mathbf{y} = \mathbf{x}^T(\mu\mathbf{y}) = \mu\,\mathbf{x}^T\mathbf{y}.$$

$\textbf{Second way — use symmetry, } A^T=A$, to move $A$ onto $\mathbf{x}$ instead:

$$\mathbf{x}^TA\mathbf{y} = (A^T\mathbf{x})^T\mathbf{y} = (A\mathbf{x})^T\mathbf{y} = (\lambda\mathbf{x})^T\mathbf{y} = \lambda\,\mathbf{x}^T\mathbf{y}.$$

$\textbf{Both computed the same quantity, so they must agree:}$

$$\lambda\,\mathbf{x}^T\mathbf{y} = \mu\,\mathbf{x}^T\mathbf{y} \quad\Longrightarrow\quad (\lambda-\mu)\,\mathbf{x}^T\mathbf{y}=0.$$

Since $\lambda\neq\mu$, the factor $(\lambda-\mu)$ is nonzero, which forces the other factor to vanish:

$$\mathbf{x}^T\mathbf{y}=0, \quad \text{i.e.} \quad \mathbf{x}\cdot\mathbf{y}=0.$$

Conclusion

$\blacksquare$ Eigenvectors of a symmetric matrix, corresponding to distinct eigenvalues, are automatically orthogonal. No Gram–Schmidt is even needed across different eigenspaces — the symmetry of $A$ does that work for you.

§9Theorem — Orthogonal Diagonalization

Orthogonal Diagonalization Theorem

Every symmetric matrix $A$ can be written as

$$A = PDP^T$$

where $D$ is the diagonal matrix of eigenvalues of $A$, and $P$ is an orthogonal matrix — its columns are the corresponding eigenvectors of $A$, each first normalized to unit length before being placed in $P$.

💡
Why this matters — the practical payoff
Compare with ordinary diagonalization, $A=PDP^{-1}$, which in general requires you to actually compute a matrix inverse — real work, and numerically delicate for large matrices. Because the columns of $P$ here are orthonormal (mutually orthogonal and unit length), $P^TP=I$, which means

$$P^{-1} = P^T.$$

Transposing a matrix costs nothing — you just relabel rows as columns. No inversion, ever. This is the entire practical reason orthogonal diagonalization is worth learning as its own theorem rather than a special case of ordinary diagonalization: for symmetric matrices, which are everywhere in applications (covariance matrices in statistics, stress tensors in engineering, adjacency matrices of undirected graphs, Hessians in optimization), diagonalizing is dramatically cheaper.

§10A Fully Worked Orthogonal Diagonalization

★ 4Orthogonally diagonalize A = [[2,1],[1,2]]

Find an orthogonal matrix $P$ and diagonal matrix $D$ such that $A=PDP^T$, for $A=\begin{bmatrix}2&1\\1&2\end{bmatrix}$.

§11Exercises

Three problems, one for each idea in this lecture. Hints only — work through the method yourself before checking.

Exercise AClosest point in a plane

Find the point in the plane $x-y+2z=0$ closest to $(3,0,1)$, and the shortest distance.

🧭
Hint
Same four moves as Example 8.1.4: parametrize the plane to get a (non-orthogonal) spanning set, run Gram–Schmidt to get an orthogonal basis $\{v_1,v_2\}$, project the point onto each $v_i$ and add, then subtract from the original point to get the distance vector. Cross-check your final distance against $|n\cdot P_0|/\|n\|$ with $n=(1,-1,2)$ — they must agree.
Exercise BDecompose into U and U⊥

Let $U = \operatorname{span}\{(2,0,-1),\,(1,3,2)\}$. Write $(4,1,5)$ as the sum of a vector in $U$ and a vector in $U^{\perp}$.

🧭
Hint
Check the two spanning vectors' dot product first — if it's already $0$ (as in Example 3), you can skip Gram–Schmidt entirely and project directly. Follow the same five steps as Example 3: project onto each basis vector, add for $\text{Proj}_U(p)$, subtract from $p$ for the $U^{\perp}$ piece, and verify by dotting that piece against both basis vectors of $U$.
Exercise COrthogonally diagonalize another symmetric matrix

Find an orthogonal matrix $P$ and diagonal matrix $D$ with $A=PDP^T$, for $A=\begin{bmatrix}3&1\\1&3\end{bmatrix}$.

🧭
Hint
Follow Example 4 exactly: confirm $A$ is symmetric, find the two eigenvalues from $(3-\lambda)^2-1=0$, find one eigenvector per eigenvalue, confirm they're orthogonal (guaranteed, since the eigenvalues are distinct), normalize both to unit length, and assemble $P$ and $D$ with matching column/diagonal order.
🔭
Where this leads next
Everything today — projections, complements, orthogonal diagonalization — has been about decomposing space and simplifying matrices that already act "nicely." Lecture 20 turns the question around: instead of studying special matrices, we study linear transformations themselves — the general maps that matrices represent — and ask what survives when you change coordinate systems entirely.
Looking back

Drop a perpendicular, split a vector in two, and watch a symmetric matrix diagonalize itself for free.

The orthogonal projection formula, the direct-sum split $V=U\oplus U^{\perp}$, and the spectral theorem for symmetric matrices are not three separate topics — they are one idea, applied three times. Once you have an orthogonal basis, everything in linear algebra gets cheaper: coordinates become dot products, projections become sums, and matrix inversion becomes a transpose.

📌 Course Administration — not examinable, informational only

The Final Exam is worth 80 marks. There appears to be a partial midterm-replacement policy for students who scored 40 or above on the midterm, involving a ratio structure (roughly 20:80 and 1:4) that governs how final-exam performance can offset the midterm grade. The exact mechanics are unclear from the lecture notes this was transcribed from — please confirm the precise wording, thresholds, and eligibility directly with the instructor rather than relying on this summary.

Lecture 19 — complete
MATH-120 · Shoaib Khan · LUMS · July 2026
← Lecture 18Lecture 20 →