Home›Courses›Linear Algebra›Week 5 · Lecture 18
← Lecture 17
On this pageThe Dot Product·Length of a Vector·Distance·Orthogonality·Normalizing·Orthogonal Lemma·Gram–Schmidt: How·Gram–Schmidt: Worked·Exercises
Lecture 19 →
MATH-120 · Linear Algebra · Lecture 189 July 2026

Geometry Arrives: Dot Product, Distance & Gram–Schmidt

Everything so far has been algebra — spans, bases, dimension. Today we add angles, lengths, and the idea of "perpendicular," and build a machine that manufactures perpendicular bases on demand

§1Motivation — Why Geometry, Why Now?

Every lecture since Week 4 has been purely algebraic: a vector was just an object you could add and scale, obeying eight axioms. That abstraction was powerful — it let matrices and polynomials become "vectors" too — but it threw away something $\mathbb{R}^n$ has that a general vector space doesn't: geometry. In $\mathbb{R}^2$ and $\mathbb{R}^3$, vectors have length, and pairs of vectors have an angle between them. Today we bring that geometry back, formalize it for any $\mathbb{R}^n$, and then use it to build the single most useful construction in applied linear algebra: a way to manufacture a perpendicular basis out of any basis at all.

🌍
A real fact worth sitting with
Google's original PageRank algorithm, GPS positioning, JPEG image compression, noise-cancelling headphones, and every least-squares regression line you've ever fit to data all lean on the same two ideas from this lecture: the dot product (which measures how much two directions "agree") and orthogonal projection (which finds the closest point in a subspace to a given point). GPS, for instance, has to solve an over-determined system — more satellite-distance equations than unknowns — and it does so by projecting onto a subspace exactly the way we will today.

The plan: first, revisit the dot product, length, and distance from Nicholson §5.3 — tools you have used informally since high school, now placed on rigorous footing in $\mathbb{R}^n$. Then, in §8.1, we ask a sharper question: given any basis of a subspace, can we always replace it with an equally good basis whose vectors are mutually perpendicular? The answer is yes, and the algorithm that does it — the Gram–Schmidt process — is today's main event.

§2The Dot Product

Definition — dot product

If $\mathbf{x}=(x_1,x_2,\ldots,x_n)$ and $\mathbf{y}=(y_1,y_2,\ldots,y_n)$ are vectors in $\mathbb{R}^n$, their dot product is the number

$$\mathbf{x}\cdot\mathbf{y} = x_1y_1 + x_2y_2 + \cdots + x_ny_n.$$

Notice immediately that $\mathbf{x}\cdot\mathbf{y}$ is a single number (a scalar), not a vector — this is the first genuinely new kind of operation we've met: it takes two vectors in and produces a number out.

Example 1Computing a dot product

If $\mathbf{x}=(1,-2,3,0)$ and $\mathbf{y}=(4,1,-1,5)$, then

$$\mathbf{x}\cdot\mathbf{y} = (1)(4)+(-2)(1)+(3)(-1)+(0)(5) = 4-2-3+0 = -1.$$

§3An Observation: Dot Product as Matrix Product

🔍
Observation
If $\mathbf{x}$ and $\mathbf{y}$ are written as columns, then $\mathbf{x}\cdot\mathbf{y} = \mathbf{x}^T\mathbf{y}$ is a matrix product (and $\mathbf{x}\cdot\mathbf{y} = \mathbf{x}\mathbf{y}^T$ if they are written as rows). Here $\mathbf{x}\cdot\mathbf{y}$ is a $1\times1$ matrix, which we take to be a number.

This is not a new definition — it is the same number, seen through the matrix-multiplication rule from Week 2. Recall that multiplying a $1\times n$ row by an $n\times1$ column produces a $1\times1$ matrix, whose single entry is exactly "multiply matching entries and add them up." That is precisely the dot product formula.

Example 2The same number, two ways

With $\mathbf{x}=\begin{bmatrix}1\\-2\\3\end{bmatrix}$ and $\mathbf{y}=\begin{bmatrix}2\\0\\1\end{bmatrix}$ as columns,

$$\mathbf{x}^T\mathbf{y} = \begin{bmatrix}1&-2&3\end{bmatrix}\begin{bmatrix}2\\0\\1\end{bmatrix} = (1)(2)+(-2)(0)+(3)(1) = 5,$$

which matches $\mathbf{x}\cdot\mathbf{y} = (1)(2)+(-2)(0)+(3)(1)=5$ computed directly. Same number — the matrix-product viewpoint is just useful bookkeeping, and it is why some textbooks write $\mathbf{x}^T\mathbf{y}$ instead of $\mathbf{x}\cdot\mathbf{y}$.

§4The Length of a Vector

In $\mathbb{R}^2$, the Pythagorean theorem gives the length of $(x_1,x_2)$ as $\sqrt{x_1^2+x_2^2}$. The dot product lets us extend this to any $\mathbb{R}^n$ at once, since $\mathbf{x}\cdot\mathbf{x} = x_1^2+x_2^2+\cdots+x_n^2$ is exactly the sum of squares under that root.

Definition 5.6 — length (norm)

As in $\mathbb{R}^3$, the length $\|\mathbf{x}\|$ of the vector $\mathbf{x}$ is defined by

$$\|\mathbf{x}\| = \sqrt{\mathbf{x}\cdot\mathbf{x}} = \sqrt{x_1^2+x_2^2+\cdots+x_n^2},$$

where $\sqrt{\phantom{x}}$ indicates the positive square root. A vector $\mathbf{x}$ of length $1$ is called a unit vector.

Example 3Computing a length

For $\mathbf{x}=(1,2,-2,4)$ in $\mathbb{R}^4$: $\;\|\mathbf{x}\| = \sqrt{1^2+2^2+(-2)^2+4^2} = \sqrt{1+4+4+16} = \sqrt{25} = 5$.

Example 4Building a unit vector

Turn $\mathbf{x}=(3,0,4)$ into a unit vector pointing in the same direction. First find its length: $\|\mathbf{x}\|=\sqrt{9+0+16}=\sqrt{25}=5$.

Then scale $\mathbf{x}$ by $\frac{1}{\|\mathbf{x}\|}$: $\;\hat{\mathbf{x}} = \frac{1}{5}(3,0,4) = \left(\frac{3}{5},0,\frac{4}{5}\right)$. Check: $\left(\frac{3}{5}\right)^2+0^2+\left(\frac{4}{5}\right)^2 = \frac{9}{25}+\frac{16}{25}=\frac{25}{25}=1$ ✓. This "divide by the length" move — called normalizing — is one of the most-used moves in this entire lecture.

§5Theorem 5.3.1 — The Algebra of Dot Products

Theorem 5.3.1

Let $\mathbf{x}$, $\mathbf{y}$, and $\mathbf{z}$ denote vectors in $\mathbb{R}^n$. Then:

  • •1. $\mathbf{x}\cdot\mathbf{y} = \mathbf{y}\cdot\mathbf{x}$.
  • •2. $\mathbf{x}\cdot(\mathbf{y}+\mathbf{z}) = \mathbf{x}\cdot\mathbf{y}+\mathbf{x}\cdot\mathbf{z}$.
  • •3. $(a\mathbf{x})\cdot\mathbf{y} = a(\mathbf{x}\cdot\mathbf{y}) = \mathbf{x}\cdot(a\mathbf{y})$ for all scalars $a$.
  • •4. $\|\mathbf{x}\|^2 = \mathbf{x}\cdot\mathbf{x}$.
  • •5. $\|\mathbf{x}\|\geq0$, and $\|\mathbf{x}\|=0$ if and only if $\mathbf{x}=\mathbf{0}$.
  • •6. $\|a\mathbf{x}\| = |a|\,\|\mathbf{x}\|$ for all scalars $a$.
💡
What this theorem is really saying
Think of these six rules as the "grammar" that makes the dot product and length behave the way your intuition already expects. Rules 1–3 say the dot product acts like ordinary multiplication: order doesn't matter, it distributes over addition, and scalars can be pulled out. Rule 4 is really just restating Definition 5.6 — length is, by construction, the square root of a dot product. Rule 5 says length is a genuine measure of size: never negative, and zero only for the zero vector — nothing else can have "no size." Rule 6 says stretching a vector by a factor $a$ stretches its length by $|a|$ (the absolute value matters: flipping a vector's direction with $a=-1$ doesn't change how long it is).

§6Expanding ‖x + y‖²

★ 5Example 5.3.2 — the algebraic expansion

Show that $\|\mathbf{x}+\mathbf{y}\|^2 = \|\mathbf{x}\|^2 + 2(\mathbf{x}\cdot\mathbf{y}) + \|\mathbf{y}\|^2$ for any $\mathbf{x}$ and $\mathbf{y}$ in $\mathbb{R}^n$.

§7Orthogonal to a Spanning Set Forces x = 0

★ 6Example 5.3.3 — a vector that dots to zero with a full spanning set

Suppose that $\mathbb{R}^n = \operatorname{span}\{\mathbf{f}_1,\mathbf{f}_2,\ldots,\mathbf{f}_k\}$ for some vectors $\mathbf{f}_i$. If $\mathbf{x}\cdot\mathbf{f}_i=0$ for each $i$, where $\mathbf{x}$ is in $\mathbb{R}^n$, show that $\mathbf{x}=\mathbf{0}$.

§8Measuring the Gap Between Two Vectors

Length tells us the size of a single vector. A closely related question: how "far apart" are two vectors $\mathbf{x}$ and $\mathbf{y}$? The natural answer, exactly as in $\mathbb{R}^2$ and $\mathbb{R}^3$, is to measure the length of the vector connecting them.

Definition 5.7 — distance

If $\mathbf{x}$ and $\mathbf{y}$ are two vectors in $\mathbb{R}^n$, we define the distance $d(\mathbf{x},\mathbf{y})$ between $\mathbf{x}$ and $\mathbf{y}$ by

$$d(\mathbf{x},\mathbf{y}) = \|\mathbf{x}-\mathbf{y}\|.$$

📐
The motivation, again from R³
Picture $\mathbf{x}$ and $\mathbf{y}$ as two arrows from a common origin $O$, with tips $X$ and $Y$. The vector $\mathbf{x}-\mathbf{y}$ is exactly the arrow that connects tip $Y$ to tip $X$ — so its length $\|\mathbf{x}-\mathbf{y}\|$ is precisely the straight-line distance between the two tips, matching what "distance" already means in $\mathbb{R}^3$. The formula $d(\mathbf{x},\mathbf{y})=\|\mathbf{x}-\mathbf{y}\|$ is doing nothing more than converting "distance between two points" into "length of the vector between them" — a move we can now make sense of in any $\mathbb{R}^n$, however large.
d(x,y) = ‖x − y‖xyx − yOd(x,z) ≤ d(x,y) + d(y,z)xyzd(x,z) — direct pathd(x,y)d(y,z)

Left: the distance between $\mathbf{x}$ and $\mathbf{y}$ is the length of the connecting vector $\mathbf{x}-\mathbf{y}$. Right: going straight from $\mathbf{x}$ to $\mathbf{z}$ is never longer than detouring through $\mathbf{y}$.

Example 7Computing a distance

For $\mathbf{x}=(1,3,-2)$ and $\mathbf{y}=(4,-1,0)$: $\;\mathbf{x}-\mathbf{y}=(-3,4,-2)$, so $d(\mathbf{x},\mathbf{y}) = \sqrt{(-3)^2+4^2+(-2)^2} = \sqrt{9+16+4}=\sqrt{29}$.

§9Theorem 5.3.3 — Properties of Distance

Theorem 5.3.3

If $\mathbf{x}$, $\mathbf{y}$, and $\mathbf{z}$ are three vectors in $\mathbb{R}^n$ we have:

  • •1. $d(\mathbf{x},\mathbf{y}) \geq 0$ for all $\mathbf{x}$ and $\mathbf{y}$.
  • •2. $d(\mathbf{x},\mathbf{y}) = 0$ if and only if $\mathbf{x}=\mathbf{y}$.
  • •3. $d(\mathbf{x},\mathbf{y}) = d(\mathbf{y},\mathbf{x})$.
  • •4. $d(\mathbf{x},\mathbf{z}) \leq d(\mathbf{x},\mathbf{y}) + d(\mathbf{y},\mathbf{z})$. (Triangle inequality.)
💡
Why each of these is true — no proofs, just the idea

Property 1. Distance is a length ($d(\mathbf{x},\mathbf{y})=\|\mathbf{x}-\mathbf{y}\|$), and Theorem 5.3.1 rule 5 already told us every length is $\geq0$. Nothing new — distance simply inherits non-negativity from length.

Property 2. Also inherited from rule 5: $\|\mathbf{x}-\mathbf{y}\|=0$ exactly when $\mathbf{x}-\mathbf{y}=\mathbf{0}$, i.e. exactly when $\mathbf{x}=\mathbf{y}$. Geometrically: the only way two points have zero distance between them is if they're the same point.

Property 3. Distance shouldn't care which point you call "first." Algebraically, $\mathbf{x}-\mathbf{y}$ and $\mathbf{y}-\mathbf{x}$ are negatives of each other, and Theorem 5.3.1 rule 6 (with $a=-1$) says negating a vector doesn't change its length: $\|\mathbf{x}-\mathbf{y}\|=\|-(\mathbf{y}-\mathbf{x})\|=\|\mathbf{y}-\mathbf{x}\|$.

Property 4 — the triangle inequality. This is the algebraic version of "the direct route is never longer than a detour." Picture three points $\mathbf{x}$, $\mathbf{y}$, $\mathbf{z}$: walking straight from $\mathbf{x}$ to $\mathbf{z}$ covers distance $d(\mathbf{x},\mathbf{z})$, while walking from $\mathbf{x}$ to $\mathbf{y}$ and then $\mathbf{y}$ to $\mathbf{z}$ covers $d(\mathbf{x},\mathbf{y})+d(\mathbf{y},\mathbf{z})$. Common sense says the detour can only be as short as the direct path, never shorter — and this holds true in every $\mathbb{R}^n$, not just the $\mathbb{R}^2$ and $\mathbb{R}^3$ where you can literally draw the triangle.

§10Orthogonality

We now name the single most useful relationship two vectors can have with each other.

Orthogonal, orthogonal set, orthonormal set
  • •Two vectors $\mathbf{x}$ and $\mathbf{y}$ in $\mathbb{R}^n$ are orthogonal if $\mathbf{x}\cdot\mathbf{y}=0$.
  • •A set of vectors $\{\mathbf{f}_1,\ldots,\mathbf{f}_m\}$ (all nonzero) is an orthogonal set if every pair is orthogonal: $\mathbf{f}_i\cdot\mathbf{f}_j=0$ whenever $i\neq j$.
  • •An orthonormal set is an orthogonal set in which every vector is also a unit vector: $\|\mathbf{f}_i\|=1$ for every $i$, in addition to $\mathbf{f}_i\cdot\mathbf{f}_j=0$ for $i\neq j$.
📐
The geometric picture
In $\mathbb{R}^2$ and $\mathbb{R}^3$, the dot product relates to the angle $\theta$ between two vectors by $\mathbf{x}\cdot\mathbf{y} = \|\mathbf{x}\|\|\mathbf{y}\|\cos\theta$. Orthogonal means $\mathbf{x}\cdot\mathbf{y}=0$, which forces $\cos\theta=0$ — that is, $\theta=90°$. Orthogonal is simply the algebraic word for perpendicular, now generalized to any $\mathbb{R}^n$, where you can no longer draw the angle but can always compute it.
🎛 Drag the angle between x and y
θ = 60°x (‖x‖=3)y (‖y‖=2)Ox · y = ‖x‖ ‖y‖ cos θ = 3 × 2 × cos(60°) = 3.00

Drag the slider or tap a preset angle. The dot product is positive for acute angles, negative for obtuse ones, and lands on exactly $0$ only at $\theta=90°$ — that instant is orthogonality. Toggle "normalize" to see the same picture with unit-length vectors — an orthonormal pair, once $\theta=90°$.

Example 8A quick geometric check in R²

$\mathbf{x}=(1,1)$ and $\mathbf{y}=(1,-1)$ satisfy $\mathbf{x}\cdot\mathbf{y}=(1)(1)+(1)(-1)=0$, so they are orthogonal — and indeed, plotting them, $\mathbf{x}$ points along the line $y=x$ and $\mathbf{y}$ along $y=-x$, which meet at a right angle. Note also $\|\mathbf{x}\|=\|\mathbf{y}\|=\sqrt{2}$, so $\{\mathbf{x},\mathbf{y}\}$ is orthogonal but not yet orthonormal — dividing each by $\sqrt{2}$ would fix that.

§11Motivation — Why Do We Want Orthonormal Sets?

🎯
The payoff
Recall the Expansion Theorem from an earlier lecture: if $\{\mathbf{f}_1,\ldots,\mathbf{f}_m\}$ is an orthogonal basis of a subspace $U$, then any $\mathbf{x}$ in $U$ can be written as

$$\mathbf{x} = \left(\frac{\mathbf{x}\cdot\mathbf{f}_1}{\|\mathbf{f}_1\|^2}\right)\mathbf{f}_1 + \left(\frac{\mathbf{x}\cdot\mathbf{f}_2}{\|\mathbf{f}_2\|^2}\right)\mathbf{f}_2 + \cdots + \left(\frac{\mathbf{x}\cdot\mathbf{f}_m}{\|\mathbf{f}_m\|^2}\right)\mathbf{f}_m.$$

Compare this with a general (non-orthogonal) basis, where finding the coordinates of $\mathbf{x}$ requires solving a full linear system by Gaussian elimination. With an orthogonal basis, each coordinate is found by a single dot product — no elimination, no matrix inversion, just arithmetic. This is the entire reason orthogonal and orthonormal bases are worth the extra effort to build: they turn "solve a system" into "compute a few dot products." This single fact underlies projections, least-squares fitting, Fourier series, and the QR-decomposition algorithms used inside virtually every numerical linear algebra library.

§12The Standard Basis Is Orthonormal

Example 9Example 5.3.4 — the standard basis

The standard basis $\{\mathbf{e}_1,\mathbf{e}_2,\ldots,\mathbf{e}_n\}$ is an orthonormal set in $\mathbb{R}^n$: for $i\neq j$, $\mathbf{e}_i\cdot\mathbf{e}_j=0$ since the two vectors have their single $1$-entries in different positions (every term in the dot product sum involves at least one zero factor), and $\|\mathbf{e}_i\| = \sqrt{0^2+\cdots+1^2+\cdots+0^2}=\sqrt{1}=1$ for every $i$. This is the simplest possible orthonormal set — and the benchmark every other orthonormal set is compared to.

§13Normalizing an Orthogonal Set

Definition 5.9 — normalizing

If $\{\mathbf{x}_1,\mathbf{x}_2,\ldots,\mathbf{x}_k\}$ is an orthogonal set, then $\left\{\frac{1}{\|\mathbf{x}_1\|}\mathbf{x}_1,\; \frac{1}{\|\mathbf{x}_2\|}\mathbf{x}_2,\; \ldots,\; \frac{1}{\|\mathbf{x}_k\|}\mathbf{x}_k\right\}$ is an orthonormal set, and we say that it is the result of normalizing the orthogonal set $\{\mathbf{x}_1,\mathbf{x}_2,\ldots,\mathbf{x}_k\}$.

Rescaling each vector to length $1$ cannot disturb orthogonality — by Theorem 5.3.1 rule 3, $(a\mathbf{x}_i)\cdot(b\mathbf{x}_j) = ab(\mathbf{x}_i\cdot\mathbf{x}_j)$, and if $\mathbf{x}_i\cdot\mathbf{x}_j=0$, this stays $0$ no matter what $a,b$ are. So "make orthogonal" and "make unit length" are independent jobs — normalizing only ever handles the second one.

§14A Worked Orthonormal Set in R⁴

★ 10Example 5.3.6 — verifying orthogonality and normalizing

If $\mathbf{f}_1=\begin{bmatrix}1\\1\\1\\-1\end{bmatrix}$, $\mathbf{f}_2=\begin{bmatrix}1\\0\\1\\2\end{bmatrix}$, $\mathbf{f}_3=\begin{bmatrix}-1\\0\\1\\0\end{bmatrix}$, and $\mathbf{f}_4=\begin{bmatrix}-1\\3\\-1\\1\end{bmatrix}$, then $\{\mathbf{f}_1,\mathbf{f}_2,\mathbf{f}_3,\mathbf{f}_4\}$ is an orthogonal set in $\mathbb{R}^4$, as is easily verified. Find the corresponding orthonormal set.

§15Theorem 5.3.5 — Orthogonal Sets Are Independent

Theorem 5.3.5

Every orthogonal set in $\mathbb{R}^n$ is linearly independent.

💡
Why this is true
Suppose $\{\mathbf{f}_1,\ldots,\mathbf{f}_m\}$ is orthogonal and $c_1\mathbf{f}_1+c_2\mathbf{f}_2+\cdots+c_m\mathbf{f}_m=\mathbf{0}$. Dot both sides with a single fixed $\mathbf{f}_k$. On the right, $\mathbf{0}\cdot\mathbf{f}_k=0$. On the left, distribute the dot product across the sum: every term $c_i(\mathbf{f}_i\cdot\mathbf{f}_k)$ with $i\neq k$ vanishes immediately, because the set is orthogonal — that's the whole trick. All that survives is the single term $c_k(\mathbf{f}_k\cdot\mathbf{f}_k) = c_k\|\mathbf{f}_k\|^2$. So the equation collapses to $c_k\|\mathbf{f}_k\|^2=0$, and since $\mathbf{f}_k\neq\mathbf{0}$ (orthogonal sets consist of nonzero vectors), $\|\mathbf{f}_k\|^2\neq0$, forcing $c_k=0$. This argument works for every $k$ separately, so all coefficients are forced to $0$ — independence, essentially for free, no elimination needed. This is the second half of why orthogonal sets are so valuable: they are automatically independent, so an orthogonal spanning set is automatically a basis.

§16§8.1 — Orthogonal Complements and Projections

We now switch chapters, but not topics — everything above was preparation for this. Recall from Lecture 17 the tool that let us grow an independent set into a basis:

🔁
Where we're headed
If $\{\mathbf{v}_1,\ldots,\mathbf{v}_m\}$ is linearly independent in a general vector space, and if $\mathbf{v}_{m+1}$ is not in $\operatorname{span}\{\mathbf{v}_1,\ldots,\mathbf{v}_m\}$, then $\{\mathbf{v}_1,\ldots,\mathbf{v}_m,\mathbf{v}_{m+1}\}$ is independent — that was the Independent Lemma (Lemma 6.4.1). Here is the analog for orthogonal sets in $\mathbb{R}^n$: not only can we add a new independent vector, we can engineer it so the new vector is automatically orthogonal to all the old ones.

§17Lemma 8.1.1 — The Orthogonal Lemma

Lemma 8.1.1

Let $\{\mathbf{f}_1,\mathbf{f}_2,\ldots,\mathbf{f}_m\}$ be an orthogonal set in $\mathbb{R}^n$. Given $\mathbf{x}$ in $\mathbb{R}^n$, write

$$\mathbf{f}_{m+1} = \mathbf{x} - \frac{\mathbf{x}\cdot\mathbf{f}_1}{\|\mathbf{f}_1\|^2}\mathbf{f}_1 - \frac{\mathbf{x}\cdot\mathbf{f}_2}{\|\mathbf{f}_2\|^2}\mathbf{f}_2 - \cdots - \frac{\mathbf{x}\cdot\mathbf{f}_m}{\|\mathbf{f}_m\|^2}\mathbf{f}_m.$$

Then:

  • •1. $\mathbf{f}_{m+1}\cdot\mathbf{f}_k = 0$ for $k=1,2,\ldots,m$.
  • •2. If $\mathbf{x}$ is not in $\operatorname{span}\{\mathbf{f}_1,\ldots,\mathbf{f}_m\}$, then $\mathbf{f}_{m+1}\neq\mathbf{0}$ and $\{\mathbf{f}_1,\ldots,\mathbf{f}_m,\mathbf{f}_{m+1}\}$ is an orthogonal set.
💡
Explanation
Each fraction $\frac{\mathbf{x}\cdot\mathbf{f}_k}{\|\mathbf{f}_k\|^2}\mathbf{f}_k$ is exactly the "amount of $\mathbf{x}$ pointing along $\mathbf{f}_k$" — its projection onto $\mathbf{f}_k$ (this is the same formula that appears in the Expansion Theorem). Subtracting off the projection onto every $\mathbf{f}_k$ in turn strips away all the parts of $\mathbf{x}$ that overlap with the existing orthogonal set, leaving only whatever is genuinely new — the "leftover" component, $\mathbf{f}_{m+1}$, is by construction the part of $\mathbf{x}$ that has zero overlap with each $\mathbf{f}_k$, which is precisely what "orthogonal to every $\mathbf{f}_k$" means. Part 2 is then intuitive: if $\mathbf{x}$ had no genuinely new direction (i.e. $\mathbf{x}\in\operatorname{span}\{\mathbf{f}_1,\ldots,\mathbf{f}_m\}$ already), then subtracting off every projection would leave exactly $\mathbf{0}$ — there'd be nothing left over. So a nonzero leftover $\mathbf{f}_{m+1}$ signals that $\mathbf{x}$ truly added something new, and — being orthogonal to everything before it — it extends the orthogonal set.

§18Theorem 8.1.1

Theorem 8.1.1

Let $U$ be a subspace of $\mathbb{R}^n$.

  • •1. Every orthogonal subset $\{\mathbf{f}_1,\ldots,\mathbf{f}_m\}$ in $U$ is a subset of an orthogonal basis of $U$.
  • •2. $U$ has an orthogonal basis.
💡
Explanation
This is the orthogonal counterpart of Lemma 6.4.2 from Lecture 17, and the argument runs exactly the same way, now powered by the Orthogonal Lemma instead of the Independent Lemma. Part 1: start with your orthogonal subset. If it doesn't yet span $U$, some vector $\mathbf{x}\in U$ lies outside its span — apply the Orthogonal Lemma with that $\mathbf{x}$ to get a nonzero $\mathbf{f}_{m+1}$, automatically orthogonal to everything so far. Repeat. Because $U$ is finite dimensional (Theorem 6.4.1), this process must terminate, and it can only terminate once the set spans $U$ — at which point it's an orthogonal basis. Part 2 is the special case where you start from the empty set: run the same process from scratch, and you build an orthogonal basis of $U$ out of nothing. This guarantees that every subspace of $\mathbb{R}^n$, no matter how it was originally described, has some orthogonal basis waiting to be found — which is exactly the promise the Gram–Schmidt process below makes good on, explicitly.

§19The Gram–Schmidt Process — Why It's Worth Learning

Theorem 8.1.1 is an existence statement: an orthogonal basis exists. It doesn't tell you how to build one from a basis you already have in hand. The Gram–Schmidt process closes that gap completely: give it any basis at all, however skewed and non-perpendicular, and it hands back an orthogonal basis for the exact same subspace — with a fully explicit, step-by-step recipe.

🏆
Why this is one of the most important algorithms in the course
  • •It makes the Expansion Theorem usable. §11 showed that an orthogonal basis turns "solve a linear system" into "compute a dot product." Gram–Schmidt is what lets you actually get an orthogonal basis for any subspace you're handed, not just the lucky ones.
  • •It is the engine behind the QR-decomposition, the standard tool numerical software uses to solve least-squares problems (curve fitting, regression, GPS positioning) stably and accurately.
  • •It is completely constructive. Unlike many existence theorems in this course, Gram–Schmidt is an algorithm you can run by hand on a small example (as we're about to) or code up in five lines for a computer.

§20The Gram–Schmidt Process — How It Works

The idea is simply to apply the Orthogonal Lemma over and over, once for each vector of a starting basis $\{\mathbf{x}_1,\mathbf{x}_2,\ldots,\mathbf{x}_k\}$ of a subspace $U$, in order.

  1. 1Keep the first vector as-is: $\mathbf{f}_1 = \mathbf{x}_1$. (A single nonzero vector is trivially an orthogonal set — there's nothing to compare it to yet.)
  2. 2Strip $\mathbf{x}_2$'s overlap with $\mathbf{f}_1$: $\;\mathbf{f}_2 = \mathbf{x}_2 - \dfrac{\mathbf{x}_2\cdot\mathbf{f}_1}{\|\mathbf{f}_1\|^2}\mathbf{f}_1$. By the Orthogonal Lemma, $\mathbf{f}_2\cdot\mathbf{f}_1=0$ automatically.
  3. 3Strip $\mathbf{x}_3$'s overlap with both $\mathbf{f}_1$ and $\mathbf{f}_2$: $\;\mathbf{f}_3 = \mathbf{x}_3 - \dfrac{\mathbf{x}_3\cdot\mathbf{f}_1}{\|\mathbf{f}_1\|^2}\mathbf{f}_1 - \dfrac{\mathbf{x}_3\cdot\mathbf{f}_2}{\|\mathbf{f}_2\|^2}\mathbf{f}_2$.
  4. 4Continue the same pattern. At step $i$, subtract off the projection of $\mathbf{x}_i$ onto every $\mathbf{f}_j$ built so far ($j<i$):

    $$\mathbf{f}_i = \mathbf{x}_i - \sum_{j=1}^{i-1} \frac{\mathbf{x}_i\cdot\mathbf{f}_j}{\|\mathbf{f}_j\|^2}\mathbf{f}_j.$$

  5. 5Stop after $k$ steps. Since $\{\mathbf{x}_1,\ldots,\mathbf{x}_k\}$ was independent, each $\mathbf{x}_i$ is genuinely outside $\operatorname{span}\{\mathbf{f}_1,\ldots,\mathbf{f}_{i-1}\}$ (it's part of an independent set!), so by the Orthogonal Lemma every $\mathbf{f}_i$ comes out nonzero, and $\{\mathbf{f}_1,\ldots,\mathbf{f}_k\}$ ends up an orthogonal basis of $U$.
  6. 6Optional final step: normalize, dividing each $\mathbf{f}_i$ by $\|\mathbf{f}_i\|$, to get a fully orthonormal basis (Definition 5.9).
🧩
The one-line summary
Gram–Schmidt is just "the Orthogonal Lemma, run $k$ times in a row" — build the orthogonal set one vector at a time, and at each step, subtract off everything the next basis vector has in common with what you've already built.

§21Full Worked Example

★ 11Gram–Schmidt on a basis of R³

Apply the Gram–Schmidt process to the basis $\mathbf{x}_1=(1,1,0)$, $\mathbf{x}_2=(1,0,1)$, $\mathbf{x}_3=(0,1,1)$ of $\mathbb{R}^3$.

§22Two More — Quicker — Examples

Example 12Gram–Schmidt on two vectors

Orthogonalize $\mathbf{x}_1=(1,0,1)$, $\mathbf{x}_2=(1,1,1)$.

Example 13Gram–Schmidt on a staircase basis

Orthogonalize $\mathbf{x}_1=(1,1,1)$, $\mathbf{x}_2=(0,1,1)$, $\mathbf{x}_3=(0,0,1)$.

§23Looking Ahead

🧭
Next: orthogonal complements and projections
With Gram–Schmidt in hand, every subspace of $\mathbb{R}^n$ has a ready-made orthogonal basis. The next step is to ask: given a subspace $U$, what does the set of vectors orthogonal to all of $U$ look like — the orthogonal complement $U^{\perp}$? And given any vector $\mathbf{x}$, what is the closest point to it inside $U$ — its orthogonal projection onto $U$? Both questions are answered directly using the machinery built today: an orthogonal basis of $U$ and the same projection formula from the Orthogonal Lemma.

§24Exercises

Four problems to practice on your own, in the same spirit as today's examples. Hints only — try each one properly before reading further.

Exercise ALength and unit vectors

Find $\|\mathbf{x}\|$ for $\mathbf{x}=(2,-1,2,4)$, and then find the unit vector pointing in the same direction as $\mathbf{x}$.

🧭
Hint
Use Definition 5.6 to get $\|\mathbf{x}\|$ first (it comes out to a whole number). Then divide every entry of $\mathbf{x}$ by that length, exactly as in Example 4.
Exercise BChecking orthogonality

Determine whether $\{(1,2,-1),\,(2,-1,0),\,(1,2,5)\}$ is an orthogonal set in $\mathbb{R}^3$.

🧭
Hint
There are $\binom{3}{2}=3$ pairwise dot products to check, exactly as in Example 5.3.6's Step 1. All three must be zero for the set to qualify.
Exercise CThe distance triangle inequality, concretely

Let $\mathbf{x}=(0,0)$, $\mathbf{y}=(3,0)$, $\mathbf{z}=(3,4)$ in $\mathbb{R}^2$. Compute $d(\mathbf{x},\mathbf{y})$, $d(\mathbf{y},\mathbf{z})$, and $d(\mathbf{x},\mathbf{z})$, and confirm the triangle inequality holds.

🧭
Hint
Use Definition 5.7 three times. These three points form a right triangle (a 3-4-5 triangle, in fact) — so you should find the triangle inequality holds strictly (the direct path is shorter, not equal). Equality would only happen if $\mathbf{x}$, $\mathbf{y}$, $\mathbf{z}$ were collinear with $\mathbf{y}$ sitting directly between the other two.
Exercise DGram–Schmidt in R⁴

Apply the Gram–Schmidt process to $\mathbf{x}_1=(1,1,0,0)$, $\mathbf{x}_2=(1,0,1,0)$.

🧭
Hint
Only two vectors, so only Steps 1–2 of the process are needed — set $\mathbf{f}_1=\mathbf{x}_1$, compute $\mathbf{x}_2\cdot\mathbf{f}_1$ and $\|\mathbf{f}_1\|^2$, then subtract the projection exactly as in Example 12.
Looking back

The dot product measures agreement, orthogonality means zero agreement, and Gram–Schmidt is the machine that manufactures perpendicular directions out of any basis at all.

Every computation today reduced to the same handful of moves: a dot product, a length, and a subtraction. What made them powerful was the pattern they revealed — that "closest point," "no overlap," and "coordinates without solving a system" are all the same idea, seen from different angles. That idea is about to become the backbone of orthogonal projections in the next lecture.

Lecture 18 — complete
MATH-120 · Shoaib Khan · LUMS · July 2026
← Lecture 17Lecture 19 →