RREF, Homogeneous Systems & Linear Combinations
Completing the elimination story — then meeting the two most fundamental ideas in linear algebra
§1Quick Recall
Last two lectures we built the full machinery of Gaussian elimination. Before we push further, let us nail down a few facts that are easy to forget but come up in every problem.
§2Rank Bounds & Full Rank
For any $m \times n$ matrix $A$, the rank satisfies $\operatorname{rank}(A) \le \min(m, n)$. The pivots live in distinct rows (at most $m$) AND in distinct columns (at most $n$), so neither limit can be exceeded.
A matrix has full rank when $\operatorname{rank}(A) = \min(m, n)$ — as large as possible. A square $n \times n$ matrix with $\operatorname{rank}(A) = n$ is called full rank (or nonsingular). This turns out to be equivalent to it having an inverse — something we will prove properly in a later lecture.
(a) A $3 \times 5$ matrix: $\operatorname{rank} \le \min(3,5) = 3$. At most 3 pivots, however many columns there are.
(b) A $6 \times 2$ matrix: $\operatorname{rank} \le \min(6,2) = 2$. Tall matrices are always bounded by their column count.
(c) $I_4$ (the $4 \times 4$ identity): rank $= 4 = \min(4,4)$ — full rank. It already is in REF with four pivots.
Here is a concrete demonstration. We apply two different elimination strategies to the same matrix and see two valid but different REFs:
Start with $A = \begin{pmatrix} 2 & 4 & 6 \\ 1 & 2 & 4 \\ 3 & 6 & 10 \end{pmatrix}$.
Route 1 — scale first: $\tfrac{1}{2}R_1$, then $R_2 - R_1$, $R_3 - 3R_1$, then $R_3 - R_2$:
$$A \xrightarrow{\tfrac{1}{2}R_1} \begin{pmatrix}1&2&3\\1&2&4\\3&6&10\end{pmatrix} \xrightarrow{R_2-R_1,\,R_3-3R_1} \begin{pmatrix}1&2&3\\0&0&1\\0&0&1\end{pmatrix} \xrightarrow{R_3-R_2} \underbrace{\begin{pmatrix}1&2&3\\0&0&1\\0&0&0\end{pmatrix}}_{\textbf{REF}_1}$$
Route 2 — swap first: $R_{12}$, then $R_2 - 2R_1$, $R_3 - 3R_1$, then $-\tfrac{1}{2}R_2$:
$$A \xrightarrow{R_{12}} \begin{pmatrix}1&2&4\\2&4&6\\3&6&10\end{pmatrix} \xrightarrow{R_2-2R_1,\,R_3-3R_1} \begin{pmatrix}1&2&4\\0&0&-2\\0&0&-2\end{pmatrix} \xrightarrow{-\tfrac{1}{2}R_2,\,R_3-R_2} \underbrace{\begin{pmatrix}1&2&4\\0&0&1\\0&0&0\end{pmatrix}}_{\textbf{REF}_2}$$
$\textbf{REF}_1 \ne \textbf{REF}_2$ (different entries in column 3, row 1). But both have pivots in columns 1 and 3 — so rank = 2 in both cases. The pivot count and pivot columns are the same. ✓
Practical tricks for faster row reduction
Gaussian elimination always works, but smart choices save significant effort. Here are the most useful ones:
Consider $\begin{pmatrix}1&2&13&11&7\\2&-1&1&1&0\\1&0&0&1&3\end{pmatrix}$. Naïvely, we'd use row 1 as the pivot and compute $R_2-2R_1$, $R_3-R_1$ — but row 1 has five nonzero entries, meaning every subtraction touches five numbers.
Better: apply $R_{13}$ (swap rows 1 and 3) first. Now row 1 is $(1,0,0,1,3)$ — three zeros. The operation $R_2-2R_1$ only modifies the two nonzero columns; $R_2 - R_1$ similarly cheap. Fewer nonzeros in the pivot row = fewer multiplications in every subsequent step.
$$\begin{pmatrix}1&2&13&11&7\\2&-1&1&1&0\\1&0&0&1&3\end{pmatrix}\xrightarrow{R_{13}}\begin{pmatrix}1&0&0&1&3\\2&-1&1&1&0\\1&2&13&11&7\end{pmatrix}$$
Eliminating column 1: $R_2-2R_1$ and $R_3-R_1$ now only change 2 and 4 entries respectively, instead of 5. The savings compound as the matrix grows.
Trick 2 — Avoid fractions as long as possible. If the pivot is $2$, you can scale $R_1 \to \tfrac{1}{2}R_1$ immediately, but this introduces fractions into every other entry. Often it is cleaner to use the operation $R_2 - 2R_1$ to zero out the entry below first, then scale at the end. Only make the pivot a leading 1 when you are ready to move to the next column.
Trick 3 — Look for a row with a 1 entry and put it on top. A pivot of $1$ requires no scaling, saving one entire step per pivot column.
Trick 4 — Use integer multiples before fractions. If you need to zero out entry $3$ in column 1 with pivot $2$: instead of scaling $\tfrac{1}{2}R_1$ and then subtracting, do $2R_3 - 3R_1$ directly (this keeps integers). The operation $R_i \to 2R_i - 3R_j$ is a valid row operation even though it scales $R_i$ — it is a combination of "multiply" and "add a multiple."
Trick 5 — Zero columns are free. A column of all zeros never produces a pivot. Skip it immediately and move to the next column.
§3Reduced Row-Echelon Form — the Overachiever
A matrix is in reduced row-echelon form (RREF) if it satisfies all three REF conditions, plus:
4. Each leading 1 (pivot) is the only nonzero entry in its entire column — zeros above it as well as below.
Every matrix has a unique RREF — unlike REF. This is one reason RREF is theoretically important: it is a canonical representative of the matrix's row space.
$$\begin{pmatrix}1&0&0&2\\0&1&0&-3\\0&0&1&5\end{pmatrix}$$
Each pivot column is a standard basis vector. Read off: solution is immediate.
$$\begin{pmatrix}1&2&-1&3\\0&0&1&-1\\0&0&0&0\end{pmatrix}$$
REF: staircase ✓, leading 1s ✓. But row 2's pivot (col 3) has nonzero entry above it.
Benefit 1 — No back-substitution. In RREF, each pivot variable is immediately expressed in terms of free variables. No substitution chain needed.
Benefit 2 — Uniqueness. Every matrix has exactly one RREF. Comparing two matrices' RREFs tells you whether they have the same row space.
Benefit 3 — Reading off the null space. The free-variable columns of RREF directly give the null space of $A$ — fundamental for solving $Ax = 0$, coming shortly.
Benefit 4 — Inverting matrices. The RREF of $(A \mid I)$ gives $(I \mid A^{-1})$ when $A$ is invertible — a clean algorithm we will use in Chapter 2.
§4The Gauss–Jordan Reduction Method
Gauss–Jordan reduction extends Gaussian elimination: after reaching REF (the forward pass — working top to bottom), continue with a backward pass — working bottom to top, using each pivot to zero out all entries above it (not just below). The result is RREF. The solution then reads off directly: no back-substitution required.
The two passes:
Solve $\begin{cases} x + 2y - z = 1 \\ 2x + y + z = 8 \\ x - y + 2z = 5 \end{cases}$ by Gauss–Jordan.
$$\left(\begin{array}{ccc|c}1&2&-1&1\\2&1&1&8\\1&-1&2&5\end{array}\right)\xrightarrow{R_2-2R_1,\,R_3-R_1}\left(\begin{array}{ccc|c}1&2&-1&1\\0&-3&3&6\\0&-3&3&4\end{array}\right)$$
$$\xrightarrow{-\tfrac{1}{3}R_2}\left(\begin{array}{ccc|c}1&2&-1&1\\0&1&-1&-2\\0&-3&3&4\end{array}\right)\xrightarrow{R_3+3R_2}\left(\begin{array}{ccc|c}1&2&-1&1\\0&1&-1&-2\\0&0&0&-2\end{array}\right)$$
The bottom row reads $0 = -2$. No solution — inconsistent. Gauss–Jordan stops here: there is nothing to RREF. Notice: $\operatorname{rank}(A) = 2 < 3 = \operatorname{rank}(A\mid b)$.
It is tempting, when a row contains $ax + by = c$ with $a$ unknown, to "divide by $a$" to get a leading 1. Do not do this without knowing $a \ne 0$. If $a = 0$, you have divided by zero — the operation is undefined and everything downstream is garbage.
You may only multiply a row by a nonzero constant. If $a$ is a symbol (like in Exercise 1.3.2), you must consider the case $a = 0$ separately, or avoid dividing altogether until you have established $a \ne 0$. This caution will save you from systematic errors in parametric problems.
§5Homogeneous Systems — Welcome to the Scariest Room
Before we define homogeneous systems properly, I want to show you one first.
Look at every equation: the right-hand side is zero in all four. Now try $x_1 = x_2 = x_3 = x_4 = 0$. Every equation becomes $(\text{something}) \cdot 0 + (\text{something}) \cdot 0 + \cdots = 0$. True. Every time. No matter what the coefficients are.
Welcome to the world of homogeneous systems.
A system $AX = b$ is homogeneous when $b = 0$ — that is, when every right-hand side is zero. It is written $AX = 0$ and also called the null system of $A$.
The solution $X = 0$ (all variables zero) is called the trivial solution and always exists. Any other solution is called a non-trivial solution.
When does $AX = 0$ have a non-trivial solution?
For an $m \times n$ homogeneous system $AX = 0$: a non-trivial solution exists if and only if $\operatorname{rank}(A) < n$ (fewer pivots than variables). Equivalently, the REF has at least one free variable. If $\operatorname{rank}(A) = n$, the only solution is $X = 0$.
You might look at a homogeneous system and think "the coefficients are small, probably only the trivial solution." That intuition is unreliable. Always reduce to REF to count pivots. If $\operatorname{rank}(A) = n$, only $X=0$; if $\operatorname{rank}(A) < n$, non-trivial solutions exist. No guessing.
$$\begin{cases} x - 2y + z = 0 \\ x + ay - 3z = 0 \\ -x + 6y - 5z = 0 \end{cases}$$
§6Linear Combinations
We have been solving systems all lecture. But solving $AX = b$ can be read another way — as asking whether $b$ can be built as a weighted sum of the columns of $A$. This reading leads to one of the most important ideas in all of mathematics.
Given vectors $\mathbf{v}_1, \mathbf{v}_2, \dots, \mathbf{v}_k$, a linear combination is any expression $c_1\mathbf{v}_1 + c_2\mathbf{v}_2 + \cdots + c_k\mathbf{v}_k$ where $c_1, c_2, \dots, c_k$ are scalars (real numbers). The set of all linear combinations of $\mathbf{v}_1, \dots, \mathbf{v}_k$ is called their span.
Let $\mathbf{u} = \begin{pmatrix}1\\0\end{pmatrix}$ and $\mathbf{v} = \begin{pmatrix}0\\1\end{pmatrix}$. Then $3\mathbf{u} + 5\mathbf{v} = \begin{pmatrix}3\\5\end{pmatrix}$. In fact $\begin{pmatrix}x\\y\end{pmatrix} = x\mathbf{u} + y\mathbf{v}$ for any real $x, y$. Every vector in $\mathbb{R}^2$ is a linear combination of $\mathbf{u}$ and $\mathbf{v}$. The pair $\{(1,0)^T,(0,1)^T\}$ is the standard basis — the coordinate axes themselves.
Let $\mathbf{a} = \begin{pmatrix}1\\2\end{pmatrix}$ and $\mathbf{b} = \begin{pmatrix}2\\4\end{pmatrix} = 2\mathbf{a}$. Since $\mathbf{b}$ is a multiple of $\mathbf{a}$, every linear combination $c_1\mathbf{a} + c_2\mathbf{b} = (c_1+2c_2)\mathbf{a}$ stays on the line through $\mathbf{a}$. You can never reach $\begin{pmatrix}1\\0\end{pmatrix}$, for example. The span is a line, not all of $\mathbb{R}^2$.
The standard basis of $\mathbb{R}^3$ is $\mathbf{e}_1=\begin{pmatrix}1\\0\\0\end{pmatrix}$, $\mathbf{e}_2=\begin{pmatrix}0\\1\\0\end{pmatrix}$, $\mathbf{e}_3=\begin{pmatrix}0\\0\\1\end{pmatrix}$. Any vector $\begin{pmatrix}a\\b\\c\end{pmatrix} = a\mathbf{e}_1 + b\mathbf{e}_2 + c\mathbf{e}_3$. The components of a vector are the coefficients in the linear combination with the standard basis. This is why we write vectors in column form: the column shows the coefficients.
The story goes that Descartes, notorious for sleeping until noon, was lying in bed one morning in 1637 when he noticed a fly crawling on the ceiling of his bedroom. He began wondering: how can I describe precisely where that fly is at any moment? He realised that if he knew the fly's distance from two walls, he could pin down its position completely. Two numbers, two walls — that was enough. He published the idea in an appendix to his Discourse on the Method and the coordinate plane — now called the Cartesian plane in his honour — was born.
The moral: the desire to locate a point precisely in space is exactly what linear combinations formalise. Every point in the plane is a linear combination of the two coordinate directions. Every point in space is a linear combination of three. Descartes, thanks to one lazy morning and one annoying fly, gave us the language.
Let $X = \begin{pmatrix}2\\1\\-1\end{pmatrix}$, $Y = \begin{pmatrix}1\\0\\1\end{pmatrix}$, $Z = \begin{pmatrix}1\\1\\-2\end{pmatrix}$. Can $V = \begin{pmatrix}0\\1\\3\end{pmatrix}$ be written as $aX + bY + cZ = V$?
Setting up the system: $a\begin{pmatrix}2\\1\\-1\end{pmatrix}+b\begin{pmatrix}1\\0\\1\end{pmatrix}+c\begin{pmatrix}1\\1\\-2\end{pmatrix}=\begin{pmatrix}0\\1\\3\end{pmatrix}$ gives three equations.
§7Geometric Picture — Homogeneous vs Non-Homogeneous
Everything we have done algebraically has a clean geometric picture. Understanding it will make the rest of the course click.
In 2D — two lines
Every line $ax + by = 0$ passes through the origin $(0,0)$. The trivial solution $x=y=0$ is always on the line. Two such lines always share the origin, so they are always consistent. They either coincide (infinitely many solutions — a whole line through the origin) or cross only at the origin (unique solution = trivial only).
The line $ax + by = c$ with $c \ne 0$ is a shift of the line $ax + by = 0$ — same slope, moved away from the origin. Two non-homogeneous lines may be parallel (no solution), cross at a single point (unique solution), or coincide (infinitely many). Inconsistency is possible — a new phenomenon compared to the homogeneous case.
Here is the same idea drawn out. On the left, two homogeneous lines — both forced through the origin. On the right, two non-homogeneous lines that have floated off the origin:
Both lines must pass through the origin — the trivial solution is always shared. They meet there (unique = trivial only) unless they coincide.
Neither line need pass through the origin. They can cross at one point, run parallel (no solution), or coincide — inconsistency becomes possible.
Now make it move. The slider below controls the constant c in $x - y = c$. Watch the amber line slide while its homogeneous twin (dashed teal) stays pinned at the origin. The purple arrow is the shift — your particular solution $\mathbf{x}_p$:
This is the whole story of §7 in one picture: a non-homogeneous solution set is just the homogeneous one, picked up and carried by a particular solution. Set $c = 0$ and the carry distance shrinks to nothing — the two lines merge.
Generalisation to higher dimensions
For any system $AX = b$, the corresponding homogeneous system is $AX = 0$ — same coefficient matrix $A$, right-hand side replaced by zero. Geometrically, this is the "translation to the origin" of $AX = b$: every solution set of $AX = b$ (if it exists) is a translated copy of the solution set of $AX = 0$. One is a flat object through the origin; the other is the same flat object shifted by a particular solution.
Simple examples: a line away from the origin in $\mathbb{R}^2$ corresponds to a line through the origin with the same slope. A plane in $\mathbb{R}^3$ not through the origin corresponds to the parallel plane through the origin. The homogeneous version always has more symmetry — it includes the origin.
If we know a solution to $AX = 0$, can we guess a solution to $AX = b$?
Suppose $\mathbf{x}_0$ is any particular solution to $AX = b$ (we found one by Gaussian elimination), and $\mathbf{x}_h$ is any solution to the homogeneous system $AX = 0$. What is $A(\mathbf{x}_0 + \mathbf{x}_h)$? What does this tell you about the structure of the complete solution set of $AX = b$? Think about it — we will prove the answer rigorously next lecture.