Geometry Arrives: Dot Product, Distance & Gram–Schmidt
Everything so far has been algebra — spans, bases, dimension. Today we add angles, lengths, and the idea of "perpendicular," and build a machine that manufactures perpendicular bases on demand
§1Motivation — Why Geometry, Why Now?
Every lecture since Week 4 has been purely algebraic: a vector was just an object you could add and scale, obeying eight axioms. That abstraction was powerful — it let matrices and polynomials become "vectors" too — but it threw away something $\mathbb{R}^n$ has that a general vector space doesn't: geometry. In $\mathbb{R}^2$ and $\mathbb{R}^3$, vectors have length, and pairs of vectors have an angle between them. Today we bring that geometry back, formalize it for any $\mathbb{R}^n$, and then use it to build the single most useful construction in applied linear algebra: a way to manufacture a perpendicular basis out of any basis at all.
The plan: first, revisit the dot product, length, and distance from Nicholson §5.3 — tools you have used informally since high school, now placed on rigorous footing in $\mathbb{R}^n$. Then, in §8.1, we ask a sharper question: given any basis of a subspace, can we always replace it with an equally good basis whose vectors are mutually perpendicular? The answer is yes, and the algorithm that does it — the Gram–Schmidt process — is today's main event.
§2The Dot Product
If $\mathbf{x}=(x_1,x_2,\ldots,x_n)$ and $\mathbf{y}=(y_1,y_2,\ldots,y_n)$ are vectors in $\mathbb{R}^n$, their dot product is the number
$$\mathbf{x}\cdot\mathbf{y} = x_1y_1 + x_2y_2 + \cdots + x_ny_n.$$
Notice immediately that $\mathbf{x}\cdot\mathbf{y}$ is a single number (a scalar), not a vector — this is the first genuinely new kind of operation we've met: it takes two vectors in and produces a number out.
If $\mathbf{x}=(1,-2,3,0)$ and $\mathbf{y}=(4,1,-1,5)$, then
$$\mathbf{x}\cdot\mathbf{y} = (1)(4)+(-2)(1)+(3)(-1)+(0)(5) = 4-2-3+0 = -1.$$
§3An Observation: Dot Product as Matrix Product
This is not a new definition — it is the same number, seen through the matrix-multiplication rule from Week 2. Recall that multiplying a $1\times n$ row by an $n\times1$ column produces a $1\times1$ matrix, whose single entry is exactly "multiply matching entries and add them up." That is precisely the dot product formula.
With $\mathbf{x}=\begin{bmatrix}1\\-2\\3\end{bmatrix}$ and $\mathbf{y}=\begin{bmatrix}2\\0\\1\end{bmatrix}$ as columns,
$$\mathbf{x}^T\mathbf{y} = \begin{bmatrix}1&-2&3\end{bmatrix}\begin{bmatrix}2\\0\\1\end{bmatrix} = (1)(2)+(-2)(0)+(3)(1) = 5,$$
which matches $\mathbf{x}\cdot\mathbf{y} = (1)(2)+(-2)(0)+(3)(1)=5$ computed directly. Same number — the matrix-product viewpoint is just useful bookkeeping, and it is why some textbooks write $\mathbf{x}^T\mathbf{y}$ instead of $\mathbf{x}\cdot\mathbf{y}$.
§4The Length of a Vector
In $\mathbb{R}^2$, the Pythagorean theorem gives the length of $(x_1,x_2)$ as $\sqrt{x_1^2+x_2^2}$. The dot product lets us extend this to any $\mathbb{R}^n$ at once, since $\mathbf{x}\cdot\mathbf{x} = x_1^2+x_2^2+\cdots+x_n^2$ is exactly the sum of squares under that root.
As in $\mathbb{R}^3$, the length $\|\mathbf{x}\|$ of the vector $\mathbf{x}$ is defined by
$$\|\mathbf{x}\| = \sqrt{\mathbf{x}\cdot\mathbf{x}} = \sqrt{x_1^2+x_2^2+\cdots+x_n^2},$$
where $\sqrt{\phantom{x}}$ indicates the positive square root. A vector $\mathbf{x}$ of length $1$ is called a unit vector.
For $\mathbf{x}=(1,2,-2,4)$ in $\mathbb{R}^4$: $\;\|\mathbf{x}\| = \sqrt{1^2+2^2+(-2)^2+4^2} = \sqrt{1+4+4+16} = \sqrt{25} = 5$.
Turn $\mathbf{x}=(3,0,4)$ into a unit vector pointing in the same direction. First find its length: $\|\mathbf{x}\|=\sqrt{9+0+16}=\sqrt{25}=5$.
Then scale $\mathbf{x}$ by $\frac{1}{\|\mathbf{x}\|}$: $\;\hat{\mathbf{x}} = \frac{1}{5}(3,0,4) = \left(\frac{3}{5},0,\frac{4}{5}\right)$. Check: $\left(\frac{3}{5}\right)^2+0^2+\left(\frac{4}{5}\right)^2 = \frac{9}{25}+\frac{16}{25}=\frac{25}{25}=1$ ✓. This "divide by the length" move — called normalizing — is one of the most-used moves in this entire lecture.
§5Theorem 5.3.1 — The Algebra of Dot Products
Let $\mathbf{x}$, $\mathbf{y}$, and $\mathbf{z}$ denote vectors in $\mathbb{R}^n$. Then:
- •1. $\mathbf{x}\cdot\mathbf{y} = \mathbf{y}\cdot\mathbf{x}$.
- •2. $\mathbf{x}\cdot(\mathbf{y}+\mathbf{z}) = \mathbf{x}\cdot\mathbf{y}+\mathbf{x}\cdot\mathbf{z}$.
- •3. $(a\mathbf{x})\cdot\mathbf{y} = a(\mathbf{x}\cdot\mathbf{y}) = \mathbf{x}\cdot(a\mathbf{y})$ for all scalars $a$.
- •4. $\|\mathbf{x}\|^2 = \mathbf{x}\cdot\mathbf{x}$.
- •5. $\|\mathbf{x}\|\geq0$, and $\|\mathbf{x}\|=0$ if and only if $\mathbf{x}=\mathbf{0}$.
- •6. $\|a\mathbf{x}\| = |a|\,\|\mathbf{x}\|$ for all scalars $a$.
§6Expanding ‖x + y‖²
Show that $\|\mathbf{x}+\mathbf{y}\|^2 = \|\mathbf{x}\|^2 + 2(\mathbf{x}\cdot\mathbf{y}) + \|\mathbf{y}\|^2$ for any $\mathbf{x}$ and $\mathbf{y}$ in $\mathbb{R}^n$.
§7Orthogonal to a Spanning Set Forces x = 0
Suppose that $\mathbb{R}^n = \operatorname{span}\{\mathbf{f}_1,\mathbf{f}_2,\ldots,\mathbf{f}_k\}$ for some vectors $\mathbf{f}_i$. If $\mathbf{x}\cdot\mathbf{f}_i=0$ for each $i$, where $\mathbf{x}$ is in $\mathbb{R}^n$, show that $\mathbf{x}=\mathbf{0}$.
§8Measuring the Gap Between Two Vectors
Length tells us the size of a single vector. A closely related question: how "far apart" are two vectors $\mathbf{x}$ and $\mathbf{y}$? The natural answer, exactly as in $\mathbb{R}^2$ and $\mathbb{R}^3$, is to measure the length of the vector connecting them.
If $\mathbf{x}$ and $\mathbf{y}$ are two vectors in $\mathbb{R}^n$, we define the distance $d(\mathbf{x},\mathbf{y})$ between $\mathbf{x}$ and $\mathbf{y}$ by
$$d(\mathbf{x},\mathbf{y}) = \|\mathbf{x}-\mathbf{y}\|.$$
Left: the distance between $\mathbf{x}$ and $\mathbf{y}$ is the length of the connecting vector $\mathbf{x}-\mathbf{y}$. Right: going straight from $\mathbf{x}$ to $\mathbf{z}$ is never longer than detouring through $\mathbf{y}$.
For $\mathbf{x}=(1,3,-2)$ and $\mathbf{y}=(4,-1,0)$: $\;\mathbf{x}-\mathbf{y}=(-3,4,-2)$, so $d(\mathbf{x},\mathbf{y}) = \sqrt{(-3)^2+4^2+(-2)^2} = \sqrt{9+16+4}=\sqrt{29}$.
§9Theorem 5.3.3 — Properties of Distance
If $\mathbf{x}$, $\mathbf{y}$, and $\mathbf{z}$ are three vectors in $\mathbb{R}^n$ we have:
- •1. $d(\mathbf{x},\mathbf{y}) \geq 0$ for all $\mathbf{x}$ and $\mathbf{y}$.
- •2. $d(\mathbf{x},\mathbf{y}) = 0$ if and only if $\mathbf{x}=\mathbf{y}$.
- •3. $d(\mathbf{x},\mathbf{y}) = d(\mathbf{y},\mathbf{x})$.
- •4. $d(\mathbf{x},\mathbf{z}) \leq d(\mathbf{x},\mathbf{y}) + d(\mathbf{y},\mathbf{z})$. (Triangle inequality.)
Property 1. Distance is a length ($d(\mathbf{x},\mathbf{y})=\|\mathbf{x}-\mathbf{y}\|$), and Theorem 5.3.1 rule 5 already told us every length is $\geq0$. Nothing new — distance simply inherits non-negativity from length.
Property 2. Also inherited from rule 5: $\|\mathbf{x}-\mathbf{y}\|=0$ exactly when $\mathbf{x}-\mathbf{y}=\mathbf{0}$, i.e. exactly when $\mathbf{x}=\mathbf{y}$. Geometrically: the only way two points have zero distance between them is if they're the same point.
Property 3. Distance shouldn't care which point you call "first." Algebraically, $\mathbf{x}-\mathbf{y}$ and $\mathbf{y}-\mathbf{x}$ are negatives of each other, and Theorem 5.3.1 rule 6 (with $a=-1$) says negating a vector doesn't change its length: $\|\mathbf{x}-\mathbf{y}\|=\|-(\mathbf{y}-\mathbf{x})\|=\|\mathbf{y}-\mathbf{x}\|$.
Property 4 — the triangle inequality. This is the algebraic version of "the direct route is never longer than a detour." Picture three points $\mathbf{x}$, $\mathbf{y}$, $\mathbf{z}$: walking straight from $\mathbf{x}$ to $\mathbf{z}$ covers distance $d(\mathbf{x},\mathbf{z})$, while walking from $\mathbf{x}$ to $\mathbf{y}$ and then $\mathbf{y}$ to $\mathbf{z}$ covers $d(\mathbf{x},\mathbf{y})+d(\mathbf{y},\mathbf{z})$. Common sense says the detour can only be as short as the direct path, never shorter — and this holds true in every $\mathbb{R}^n$, not just the $\mathbb{R}^2$ and $\mathbb{R}^3$ where you can literally draw the triangle.
§10Orthogonality
We now name the single most useful relationship two vectors can have with each other.
- •Two vectors $\mathbf{x}$ and $\mathbf{y}$ in $\mathbb{R}^n$ are orthogonal if $\mathbf{x}\cdot\mathbf{y}=0$.
- •A set of vectors $\{\mathbf{f}_1,\ldots,\mathbf{f}_m\}$ (all nonzero) is an orthogonal set if every pair is orthogonal: $\mathbf{f}_i\cdot\mathbf{f}_j=0$ whenever $i\neq j$.
- •An orthonormal set is an orthogonal set in which every vector is also a unit vector: $\|\mathbf{f}_i\|=1$ for every $i$, in addition to $\mathbf{f}_i\cdot\mathbf{f}_j=0$ for $i\neq j$.
Drag the slider or tap a preset angle. The dot product is positive for acute angles, negative for obtuse ones, and lands on exactly $0$ only at $\theta=90°$ — that instant is orthogonality. Toggle "normalize" to see the same picture with unit-length vectors — an orthonormal pair, once $\theta=90°$.
$\mathbf{x}=(1,1)$ and $\mathbf{y}=(1,-1)$ satisfy $\mathbf{x}\cdot\mathbf{y}=(1)(1)+(1)(-1)=0$, so they are orthogonal — and indeed, plotting them, $\mathbf{x}$ points along the line $y=x$ and $\mathbf{y}$ along $y=-x$, which meet at a right angle. Note also $\|\mathbf{x}\|=\|\mathbf{y}\|=\sqrt{2}$, so $\{\mathbf{x},\mathbf{y}\}$ is orthogonal but not yet orthonormal — dividing each by $\sqrt{2}$ would fix that.
§11Motivation — Why Do We Want Orthonormal Sets?
$$\mathbf{x} = \left(\frac{\mathbf{x}\cdot\mathbf{f}_1}{\|\mathbf{f}_1\|^2}\right)\mathbf{f}_1 + \left(\frac{\mathbf{x}\cdot\mathbf{f}_2}{\|\mathbf{f}_2\|^2}\right)\mathbf{f}_2 + \cdots + \left(\frac{\mathbf{x}\cdot\mathbf{f}_m}{\|\mathbf{f}_m\|^2}\right)\mathbf{f}_m.$$
Compare this with a general (non-orthogonal) basis, where finding the coordinates of $\mathbf{x}$ requires solving a full linear system by Gaussian elimination. With an orthogonal basis, each coordinate is found by a single dot product — no elimination, no matrix inversion, just arithmetic. This is the entire reason orthogonal and orthonormal bases are worth the extra effort to build: they turn "solve a system" into "compute a few dot products." This single fact underlies projections, least-squares fitting, Fourier series, and the QR-decomposition algorithms used inside virtually every numerical linear algebra library.§12The Standard Basis Is Orthonormal
The standard basis $\{\mathbf{e}_1,\mathbf{e}_2,\ldots,\mathbf{e}_n\}$ is an orthonormal set in $\mathbb{R}^n$: for $i\neq j$, $\mathbf{e}_i\cdot\mathbf{e}_j=0$ since the two vectors have their single $1$-entries in different positions (every term in the dot product sum involves at least one zero factor), and $\|\mathbf{e}_i\| = \sqrt{0^2+\cdots+1^2+\cdots+0^2}=\sqrt{1}=1$ for every $i$. This is the simplest possible orthonormal set — and the benchmark every other orthonormal set is compared to.
§13Normalizing an Orthogonal Set
If $\{\mathbf{x}_1,\mathbf{x}_2,\ldots,\mathbf{x}_k\}$ is an orthogonal set, then $\left\{\frac{1}{\|\mathbf{x}_1\|}\mathbf{x}_1,\; \frac{1}{\|\mathbf{x}_2\|}\mathbf{x}_2,\; \ldots,\; \frac{1}{\|\mathbf{x}_k\|}\mathbf{x}_k\right\}$ is an orthonormal set, and we say that it is the result of normalizing the orthogonal set $\{\mathbf{x}_1,\mathbf{x}_2,\ldots,\mathbf{x}_k\}$.
Rescaling each vector to length $1$ cannot disturb orthogonality — by Theorem 5.3.1 rule 3, $(a\mathbf{x}_i)\cdot(b\mathbf{x}_j) = ab(\mathbf{x}_i\cdot\mathbf{x}_j)$, and if $\mathbf{x}_i\cdot\mathbf{x}_j=0$, this stays $0$ no matter what $a,b$ are. So "make orthogonal" and "make unit length" are independent jobs — normalizing only ever handles the second one.
§14A Worked Orthonormal Set in R⁴
If $\mathbf{f}_1=\begin{bmatrix}1\\1\\1\\-1\end{bmatrix}$, $\mathbf{f}_2=\begin{bmatrix}1\\0\\1\\2\end{bmatrix}$, $\mathbf{f}_3=\begin{bmatrix}-1\\0\\1\\0\end{bmatrix}$, and $\mathbf{f}_4=\begin{bmatrix}-1\\3\\-1\\1\end{bmatrix}$, then $\{\mathbf{f}_1,\mathbf{f}_2,\mathbf{f}_3,\mathbf{f}_4\}$ is an orthogonal set in $\mathbb{R}^4$, as is easily verified. Find the corresponding orthonormal set.
§15Theorem 5.3.5 — Orthogonal Sets Are Independent
Every orthogonal set in $\mathbb{R}^n$ is linearly independent.
§16§8.1 — Orthogonal Complements and Projections
We now switch chapters, but not topics — everything above was preparation for this. Recall from Lecture 17 the tool that let us grow an independent set into a basis:
§17Lemma 8.1.1 — The Orthogonal Lemma
Let $\{\mathbf{f}_1,\mathbf{f}_2,\ldots,\mathbf{f}_m\}$ be an orthogonal set in $\mathbb{R}^n$. Given $\mathbf{x}$ in $\mathbb{R}^n$, write
$$\mathbf{f}_{m+1} = \mathbf{x} - \frac{\mathbf{x}\cdot\mathbf{f}_1}{\|\mathbf{f}_1\|^2}\mathbf{f}_1 - \frac{\mathbf{x}\cdot\mathbf{f}_2}{\|\mathbf{f}_2\|^2}\mathbf{f}_2 - \cdots - \frac{\mathbf{x}\cdot\mathbf{f}_m}{\|\mathbf{f}_m\|^2}\mathbf{f}_m.$$
Then:
- •1. $\mathbf{f}_{m+1}\cdot\mathbf{f}_k = 0$ for $k=1,2,\ldots,m$.
- •2. If $\mathbf{x}$ is not in $\operatorname{span}\{\mathbf{f}_1,\ldots,\mathbf{f}_m\}$, then $\mathbf{f}_{m+1}\neq\mathbf{0}$ and $\{\mathbf{f}_1,\ldots,\mathbf{f}_m,\mathbf{f}_{m+1}\}$ is an orthogonal set.
§18Theorem 8.1.1
Let $U$ be a subspace of $\mathbb{R}^n$.
- •1. Every orthogonal subset $\{\mathbf{f}_1,\ldots,\mathbf{f}_m\}$ in $U$ is a subset of an orthogonal basis of $U$.
- •2. $U$ has an orthogonal basis.
§19The Gram–Schmidt Process — Why It's Worth Learning
Theorem 8.1.1 is an existence statement: an orthogonal basis exists. It doesn't tell you how to build one from a basis you already have in hand. The Gram–Schmidt process closes that gap completely: give it any basis at all, however skewed and non-perpendicular, and it hands back an orthogonal basis for the exact same subspace — with a fully explicit, step-by-step recipe.
- •It makes the Expansion Theorem usable. §11 showed that an orthogonal basis turns "solve a linear system" into "compute a dot product." Gram–Schmidt is what lets you actually get an orthogonal basis for any subspace you're handed, not just the lucky ones.
- •It is the engine behind the QR-decomposition, the standard tool numerical software uses to solve least-squares problems (curve fitting, regression, GPS positioning) stably and accurately.
- •It is completely constructive. Unlike many existence theorems in this course, Gram–Schmidt is an algorithm you can run by hand on a small example (as we're about to) or code up in five lines for a computer.
§20The Gram–Schmidt Process — How It Works
The idea is simply to apply the Orthogonal Lemma over and over, once for each vector of a starting basis $\{\mathbf{x}_1,\mathbf{x}_2,\ldots,\mathbf{x}_k\}$ of a subspace $U$, in order.
- 1Keep the first vector as-is: $\mathbf{f}_1 = \mathbf{x}_1$. (A single nonzero vector is trivially an orthogonal set — there's nothing to compare it to yet.)
- 2Strip $\mathbf{x}_2$'s overlap with $\mathbf{f}_1$: $\;\mathbf{f}_2 = \mathbf{x}_2 - \dfrac{\mathbf{x}_2\cdot\mathbf{f}_1}{\|\mathbf{f}_1\|^2}\mathbf{f}_1$. By the Orthogonal Lemma, $\mathbf{f}_2\cdot\mathbf{f}_1=0$ automatically.
- 3Strip $\mathbf{x}_3$'s overlap with both $\mathbf{f}_1$ and $\mathbf{f}_2$: $\;\mathbf{f}_3 = \mathbf{x}_3 - \dfrac{\mathbf{x}_3\cdot\mathbf{f}_1}{\|\mathbf{f}_1\|^2}\mathbf{f}_1 - \dfrac{\mathbf{x}_3\cdot\mathbf{f}_2}{\|\mathbf{f}_2\|^2}\mathbf{f}_2$.
- 4Continue the same pattern. At step $i$, subtract off the projection of $\mathbf{x}_i$ onto every $\mathbf{f}_j$ built so far ($j<i$):
$$\mathbf{f}_i = \mathbf{x}_i - \sum_{j=1}^{i-1} \frac{\mathbf{x}_i\cdot\mathbf{f}_j}{\|\mathbf{f}_j\|^2}\mathbf{f}_j.$$
- 5Stop after $k$ steps. Since $\{\mathbf{x}_1,\ldots,\mathbf{x}_k\}$ was independent, each $\mathbf{x}_i$ is genuinely outside $\operatorname{span}\{\mathbf{f}_1,\ldots,\mathbf{f}_{i-1}\}$ (it's part of an independent set!), so by the Orthogonal Lemma every $\mathbf{f}_i$ comes out nonzero, and $\{\mathbf{f}_1,\ldots,\mathbf{f}_k\}$ ends up an orthogonal basis of $U$.
- 6Optional final step: normalize, dividing each $\mathbf{f}_i$ by $\|\mathbf{f}_i\|$, to get a fully orthonormal basis (Definition 5.9).
§21Full Worked Example
Apply the Gram–Schmidt process to the basis $\mathbf{x}_1=(1,1,0)$, $\mathbf{x}_2=(1,0,1)$, $\mathbf{x}_3=(0,1,1)$ of $\mathbb{R}^3$.
§22Two More — Quicker — Examples
Orthogonalize $\mathbf{x}_1=(1,0,1)$, $\mathbf{x}_2=(1,1,1)$.
Orthogonalize $\mathbf{x}_1=(1,1,1)$, $\mathbf{x}_2=(0,1,1)$, $\mathbf{x}_3=(0,0,1)$.
§23Looking Ahead
§24Exercises
Four problems to practice on your own, in the same spirit as today's examples. Hints only — try each one properly before reading further.
Find $\|\mathbf{x}\|$ for $\mathbf{x}=(2,-1,2,4)$, and then find the unit vector pointing in the same direction as $\mathbf{x}$.
Determine whether $\{(1,2,-1),\,(2,-1,0),\,(1,2,5)\}$ is an orthogonal set in $\mathbb{R}^3$.
Let $\mathbf{x}=(0,0)$, $\mathbf{y}=(3,0)$, $\mathbf{z}=(3,4)$ in $\mathbb{R}^2$. Compute $d(\mathbf{x},\mathbf{y})$, $d(\mathbf{y},\mathbf{z})$, and $d(\mathbf{x},\mathbf{z})$, and confirm the triangle inequality holds.
Apply the Gram–Schmidt process to $\mathbf{x}_1=(1,1,0,0)$, $\mathbf{x}_2=(1,0,1,0)$.
The dot product measures agreement, orthogonality means zero agreement, and Gram–Schmidt is the machine that manufactures perpendicular directions out of any basis at all.
Every computation today reduced to the same handful of moves: a dot product, a length, and a subtraction. What made them powerful was the pattern they revealed — that "closest point," "no overlap," and "coordinates without solving a system" are all the same idea, seen from different angles. That idea is about to become the backbone of orthogonal projections in the next lecture.