The Language of Matrices
From sequences to systems — building the bookkeeping that runs modern science
In 1801 the astronomer Giuseppe Piazzi spotted a faint moving object in the night sky — a tiny rocky body between Mars and Jupiter. He tracked it for 41 days before it vanished behind the sun. The world feared it was lost forever. Then a 24-year-old mathematician named Carl Friedrich Gauss sat down with a system of equations, invented a systematic method for solving it, and predicted exactly where the object — now known as Ceres — would reappear. Ten months later, it was found almost precisely where Gauss had said. That method, refined over two centuries, is the engine of this course. We call it Gaussian elimination, and by the end of this lecture you will understand its foundation.
§1Sequences & the nth Term
Mathematics rarely introduces a big idea from nowhere. Every concept has a simpler ancestor. The ancestor of a matrix is a sequence — something you have been working with since school without perhaps realising it.
A sequence is an ordered list of numbers. Order matters: $2, 4, 6, 8$ and $8, 6, 4, 2$ are different sequences even though they share the same four numbers. Each number is called a term, and its position in the list is its index.
The even numbers: $2, 4, 6, 8, 10, \dots$ are easy to write out — but suppose I ask for the 1000th term. Writing out 1000 numbers is absurd. Notice each term equals $2 \times \text{its position}$. So the term at position $n$ is simply $2n$. The 1000th term is $2000$.
A compact formula replaces an infinite list. That is the power of algebra.
§2The General Term $a_n$
We make this compact by naming the sequence. Call it $a$. Write the term at position $n$ as $a_n$ — read "a sub n." The tiny $n$ is the index, sitting just below and to the right of the letter.
A sequence is written $(a_n)_{n \ge 1}$, where $a_n$ is a formula giving the value at position $n$. The condition $n \ge 1$ (or $n \ge 0$, depending on context) tells you where the list starts.
(a) $a_n = 2n,\ n \ge 1$ gives $2, 4, 6, 8, \dots$ (even numbers). Here $a_7 = 14$.
(b) $a_n = n^2,\ n \ge 1$ gives $1, 4, 9, 16, 25, \dots$ (perfect squares). Here $a_{10} = 100$.
(c) $a_n = \tfrac{1}{n},\ n \ge 1$ gives $1, \tfrac{1}{2}, \tfrac{1}{3}, \tfrac{1}{4}, \dots$ — terms shrinking toward zero.
(d) The condition matters: $a_n = n - 3,\ n \ge 0$ gives $-3, -2, -1, 0, 1, \dots$ while the same formula with $n \ge 1$ gives $-2, -1, 0, 1, \dots$ — a different sequence starting one step later.
Given the terms, find the rule. This is harder but very useful.
(a) $3, 7, 11, 15, \dots$ — the gap between terms is always $4$, so the sequence is arithmetic: $a_n = 4n - 1$ for $n \ge 1$. (Check: $a_1 = 3$, $a_2 = 7$. ✓)
(b) $2, 6, 12, 20, 30, \dots$ — the gaps are $4, 6, 8, 10$, growing by 2 each time, so the rule is not linear. Notice: $2 = 1\cdot 2$, $6 = 2\cdot 3$, $12 = 3\cdot 4$, $20 = 4\cdot 5$. So $a_n = n(n+1)$. Spotting structure beats blind guessing.
(c) $1, -1, 1, -1, \dots$ — signs alternate. $a_n = (-1)^{n+1}$ works: $a_1 = (-1)^2 = 1$, $a_2 = (-1)^3 = -1$. Sequences can involve more than polynomial rules.
One crucial observation: a single index $n$ moves along one line of positions — first, second, third. Now suppose the data is not a line but a grid. One index is no longer enough.
§3Adding a Second Index — Birth of the Matrix
Introduce a second index. Write $a_{ij}$ for the entry in row $i$, column $j$. The full collection of entries arranged in a grid, with $i$ running from $1$ to $m$ and $j$ from $1$ to $n$, is written:
$$\left(a_{ij}\right)_{\substack{1 \le i \le m \\ 1 \le j \le n}}$$
Every pair $(i,j)$ names exactly one entry. This double-indexed collection is a matrix.
§4The Definition, Order, and Anatomy of a Matrix
A matrix is a rectangular array of numbers arranged in rows and columns. The $(a_{ij})$ written out in full for $1 \le i \le m$, $1 \le j \le n$:
$$A = \begin{pmatrix} a_{11} & a_{12} & \cdots & a_{1n} \\ a_{21} & a_{22} & \cdots & a_{2n} \\ \vdots & \vdots & \ddots & \vdots \\ a_{m1} & a_{m2} & \cdots & a_{mn} \end{pmatrix}$$
The entry $a_{ij}$ sits in row $i$, column $j$. Row first, column second — always. A matrix with $m$ rows and $n$ columns has order $m \times n$ ("$m$ by $n$"). We name matrices with capital letters: $A = (a_{ij})$.
$B = \begin{pmatrix} 1 & 5 & -2 \\ 0 & 3 & 7 \end{pmatrix}$ has $2$ rows and $3$ columns: order $2 \times 3$, holding $6$ entries in total. Here $b_{12} = 5$ (row 1, column 2) and $b_{23} = 7$ (row 2, column 3).
Notice: $b_{12} \ne b_{21}$ in general — the order of the indices is not interchangeable.
Row matrix: One row, order $1 \times n$: $\quad \begin{pmatrix} 4 & 1 & -3 & 9 \end{pmatrix}$. Used to represent a data point with multiple features.
Column matrix (vector): One column, order $m \times 1$: $\quad \begin{pmatrix} 4 \\ 1 \\ 9 \end{pmatrix}$. The language of vectors in machine learning, physics, economics.
Square matrix: $n \times n$. Most of the interesting theory (eigenvalues, determinants, inverses) lives here.
Zero matrix $O$: All entries $0$. The additive identity: $A + O = A$.
Identity matrix $I_n$: Diagonal $1$s, off-diagonal $0$s. The multiplicative identity: $AI_n = I_n A = A$ (when $A$ is $n\times n$). Think of it as the number $1$ for matrices.
Diagonal matrix: Square matrix with zeros everywhere except the main diagonal.
A symmetric matrix satisfies $a_{ij} = a_{ji}$ for every $i, j$ — it equals its own "mirror" across the main diagonal:
$$S = \begin{pmatrix} 4 & 2 & -1 \\ 2 & 7 & 5 \\ -1 & 5 & 3 \end{pmatrix}.$$
Check the $2$s in positions $(1,2)$ and $(2,1)$, the $-1$s in $(1,3)$ and $(3,1)$, the $5$s in $(2,3)$ and $(3,2)$. All correct. In data science, covariance matrices are symmetric.
An upper-triangular matrix has zeros everywhere below the main diagonal. A lower-triangular one has zeros above. These shapes matter because row-echelon form — coming shortly — is triangular.
Suppose $A = (a_{ij})$ is $3 \times 4$ with rule $a_{ij} = 3i - j$. Compute each entry:
$$A = \begin{pmatrix} 2 & 1 & 0 & -1 \\ 5 & 4 & 3 & 2 \\ 8 & 7 & 6 & 5 \end{pmatrix}.$$
Row 1 ($i=1$): $3(1)-1=2$, $3(1)-2=1$, $3(1)-3=0$, $3(1)-4=-1$. Each row increases by 3 going down; each column decreases by 1 going right. This is the "sequence with two indices" made concrete.
Diagonals
In a square matrix $A$, the main diagonal consists of entries $a_{ii}$ where row index equals column index: $a_{11}, a_{22}, a_{33}, \dots$ — running top-left to bottom-right. The secondary diagonal runs top-right to bottom-left. In this course the main diagonal is what matters — it controls the identity matrix, the trace, eigenvalues, and determinants.
In $A = \begin{pmatrix} \mathbf{5} & 1 & \mathbf{3} \\ 2 & \mathbf{0} & 4 \\ \mathbf{7} & 6 & \mathbf{-2} \end{pmatrix}$: main diagonal (bold-top-left to bottom-right) = $5, 0, -2$; secondary diagonal (bold-top-right to bottom-left) = $3, 0, 7$. The zero on the main diagonal is perfectly fine.
§5A Short History — Where Matrices Came From
The word "matrix" arrived in 1850, coined by the English mathematician James Joseph Sylvester. He chose the Latin for "womb" — the idea that the array is the object from which smaller arrays (determinants) are born. Sylvester's life was remarkable: Oxford barred him from a degree because he was Jewish. He practised law for fourteen years while doing mathematics in his spare time. He eventually became a professor at Johns Hopkins, where he founded the first American mathematics research journal, and later returned to Oxford as the Savilian Professor of Geometry.
Sylvester's close friend Arthur Cayley took the decisive step. In his 1858 Memoir on the Theory of Matrices, Cayley defined how to add and multiply matrices as algebraic objects in their own right — not merely as shorthand for systems. This was the leap that made linear algebra a branch of mathematics rather than a computational trick. Cayley too had a double life: he spent fourteen years as a conveyancing barrister before becoming Sadlerian Professor of Pure Mathematics at Cambridge. He produced nearly a thousand papers. His students reported he could write mathematics faster than most people could write longhand, and never had to cross anything out.
But looming over both is Gauss, whose 1801 prediction of Ceres opened this lecture. The elimination method he used is now called Gaussian elimination. Gauss was also the first to apply what we now call least-squares fitting — the foundation of modern statistics and machine learning — to the problem of fitting an orbit to imperfect telescope measurements. He was 24. He did not publish it until he was 32, by which point a French mathematician named Legendre had independently discovered the same method. The resulting priority dispute is one of the bitterest in mathematical history.
§6Solving Linear Equations
Here begins the first main theme of the course. We start with the simplest building block, then scale up until we hit the wall that matrices are designed to break.
A linear equation in variables $x_1, x_2, \dots, x_n$ has the form $a_1 x_1 + a_2 x_2 + \cdots + a_n x_n = b$, where the $a_i$ and $b$ are constants. "Linear" means each variable appears to the first power only — no squares, no products like $xy$, no $\sqrt{x}$.
Geometrically: in two variables, $ax + by = c$ is a straight line in the plane. That is why it is called linear. In three variables, $ax + by + cz = d$ is a plane in space. We will return to this shortly — it is more subtle than it looks.
Linear: $\quad 3x - 2y = 7, \quad x + y + z = 0, \quad 2x_1 - x_2 + 5x_3 - x_4 = 11$.
Not linear: $\quad x^2 + y = 4$ (square), $\quad xy = 1$ (product), $\quad \sin(x) = 0.5$ (transcendental). Even $\sqrt{x} + y = 3$ is not linear.
A system of linear equations is a collection of linear equations that must all be satisfied simultaneously. Solving the system means finding the values of the unknowns that make every equation true at once.
$$\begin{cases} x + y = 3 \\ x - y = 1 \end{cases}$$
Elimination: add both equations to cancel $y$: $(x+y)+(x-y) = 3+1$, so $2x = 4$, giving $x = 2$. Substitute back: $2 + y = 3$, so $y = 1$.
Solution $(x,y) = (2,1)$. Verify in both equations: $2+1=3$ ✓ and $2-1=1$ ✓. Always check.
Geometrically: each equation is a line. The solution is the intersection point of the two lines.
The Three Possible Outcomes
Two lines in a plane can relate in exactly three ways. This gives the fundamental trichotomy — which persists into all higher dimensions.
A system of linear equations has either exactly one solution, no solution, or infinitely many solutions. There is no fourth possibility — you cannot have, say, exactly two solutions. This fact, which seems almost obvious in 2D, remains true in any number of dimensions and will be proved rigorously later in the course.
From 2D to 3D — Something Subtle
In the plane $\mathbb{R}^2$, the equation $2x + 3y = 4$ is a line. Now move to space $\mathbb{R}^3$ and ask: what does the same equation describe there?
Your first instinct might be "still a line" — but it is wrong. The equation says nothing about $z$. So $z$ is completely free — any value of $z$ works. The solution set in $\mathbb{R}^3$ is every point $(x,y,z)$ where $2x+3y=4$ and $z$ is arbitrary. Imagine the line $2x+3y=4$ in the floor of a room; now sweep it straight up and down through every height. The result is a vertical plane.
When a variable does not appear in an equation, it is free — unconstrained, able to take any value. In $\mathbb{R}^3$, an equation missing $z$ sweeps its 2D graph into a plane. In $\mathbb{R}^4$, missing $z$ and $w$ would sweep a 2D solution into a 2D affine subspace. The general principle: each missing variable adds one dimension to the solution set.
Interact with this below. Set $c = 0$ to remove $z$ and watch the tilted plane flatten into a vertical sheet.
$$\begin{cases} x + y + z = 6 \\ x - y + z = 2 \\ 2x + y - z = 1 \end{cases}$$
Step 1. Eliminate $x$ from equations 2 and 3.
Eq 1 $-$ Eq 2: $2y = 4$, so $y = 2$ immediately.
Eq 3 $- 2\times$Eq 1: $-y - 3z = -11$; substituting $y=2$: $z = 3$.
Step 2. Back-substitute into Eq 1: $x + 2 + 3 = 6 \Rightarrow x = 1$.
Solution $(1, 2, 3)$. Check in Eq 3: $2(1)+2-3 = 1$ ✓. With 2 variables we needed one step; with 3, we needed three coordinated stages.
$$\begin{cases} x + y + z = 6 \\ 2x - y + z = 3 \\ x + 2y - z = 3 \end{cases}$$
Eq 2 $- 2\times$Eq 1 gives $-3y - z = -9$; Eq 3 $-$ Eq 1 gives $y - 2z = -3$. From the second, $y = 2z-3$. Substituting: $-3(2z-3)-z = -9 \Rightarrow -7z = -18 \Rightarrow z = \tfrac{18}{7}$. Now $y$ and $x$ are also fractions. Nothing went wrong — this is what generic systems look like. The point: elimination works but the work compounds alarmingly as variables increase.
With two variables, one elimination step was enough. With three, we needed two stages. With four variables we would need three stages; five variables, four stages; and so on. The work grows as the square of the number of variables. For the kind of systems that arise in modern applications — machine learning models regularly solve systems with millions of variables — raw elimination on the original equations is completely impractical. We need a smarter representation.
§7The Matrix Method — Strip the Variables, Keep the Numbers
Here is the key insight: in every step of elimination, the variables $x, y, z$ play no role. We add and subtract equations — but it is only the coefficients and constants that change. The variables are passengers. So we throw them off the bus.
Take a concrete system. Three equations, three unknowns:
$$\begin{cases} 2x + y - z = 8 \\ -3x - y + 2z = -11 \\ -2x + y + 2z = -3 \end{cases}$$
Now strip away the symbols. Line up the coefficients in a grid, stack the unknowns in a column, put the constants on the right. Nothing is lost — every number stays exactly where it was:
$$\underbrace{\begin{pmatrix} 2 & 1 & -1 \\ -3 & -1 & 2 \\ -2 & 1 & 2 \end{pmatrix}}_{\large A} \underbrace{\begin{pmatrix} x \\ y \\ z \end{pmatrix}}_{\large X} = \underbrace{\begin{pmatrix} 8 \\ -11 \\ -3 \end{pmatrix}}_{\large b}$$
Collect the coefficients into one matrix, the unknowns into a column, the constants into another column. In general the system becomes:
$$\underbrace{\begin{pmatrix} a_{11} & a_{12} & \cdots & a_{1n} \\ a_{21} & a_{22} & \cdots & a_{2n} \\ \vdots & \vdots & \ddots & \vdots \\ a_{m1} & a_{m2} & \cdots & a_{mn} \end{pmatrix}}_{\large A\ (\text{coefficient matrix})} \underbrace{\begin{pmatrix} x_1 \\ x_2 \\ \vdots \\ x_n \end{pmatrix}}_{\large X\ (\text{unknowns})} = \underbrace{\begin{pmatrix} b_1 \\ b_2 \\ \vdots \\ b_m \end{pmatrix}}_{\large b\ (\text{constants})}$$
$$A\,X = b$$
Before building $A$, align the variables in the same order in every equation. A missing variable contributes a coefficient of $0$ — write it. If the equations are not aligned, the matrix is meaningless — different columns will correspond to different variables in different rows.
System: $\begin{cases} 2x + 3y = 8 \\ x - y = -1 \end{cases}$. Both equations already list $x$ then $y$, so:
$\quad A = \begin{pmatrix}2&3\\1&-1\end{pmatrix},\quad X = \begin{pmatrix}x\\y\end{pmatrix},\quad b = \begin{pmatrix}8\\-1\end{pmatrix}.$
Now consider $\begin{cases} y + 2x = 8 \\ x - y = -1\end{cases}$. The first equation lists $y$ before $x$. If you copy the coefficients naively you get $A' = \begin{pmatrix}1&2\\1&-1\end{pmatrix}$ — wrong, because now column 1 means $y$ in row 1 but $x$ in row 2. Always re-sort to $x$ first, then $y$, before extracting coefficients.
The Augmented Matrix
For solving, we attach the constants to the coefficient matrix as one extra column separated by a bar. This is the augmented matrix $(A\mid b)$:
$$(A\mid b)=\left(\begin{array}{cccc|c}a_{11}&a_{12}&\cdots&a_{1n}&b_1\\a_{21}&a_{22}&\cdots&a_{2n}&b_2\\\vdots&\vdots&\ddots&\vdots&\vdots\\a_{m1}&a_{m2}&\cdots&a_{mn}&b_m\end{array}\right)$$
The bar marks where the equals sign was. Everything left of it is a coefficient; the single column right of it holds the constants. Operating on the augmented matrix is equivalent to performing elimination on the equations — but without ever writing $x$, $y$, or $z$.
System: $\begin{cases} x + y + z = 6 \\ 2x - y + z = 3 \\ x + 2y - z = 3 \end{cases}$
As $AX = b$: $\quad \begin{pmatrix}1&1&1\\2&-1&1\\1&2&-1\end{pmatrix}\begin{pmatrix}x\\y\\z\end{pmatrix}=\begin{pmatrix}6\\3\\3\end{pmatrix}$
Augmented: $\quad \left(\begin{array}{ccc|c}1&1&1&6\\2&-1&1&3\\1&2&-1&3\end{array}\right)$
System: $\begin{cases} 2x - y = 4 \\ z = 3 \end{cases}$. The first has no $z$; the second has no $x$ or $y$. Fill with zeros:
$$\left(\begin{array}{ccc|c}2&-1&0&4\\0&0&1&3\end{array}\right)$$
Temptation: should we add a third row of zeros to make it square? The row $(0\ 0\ 0\mid 0)$ reads $0=0$ — true always, containing zero information. This redundant (trivial) row can be appended endlessly without changing the system. It is not wrong, but it is clutter. Leave it out.
Suppose elimination on some system produces:
$$\left(\begin{array}{ccc|c}1&2&-1&4\\0&1&3&2\\0&0&0&5\end{array}\right).$$
Read the bottom row: $0x + 0y + 0z = 5$, i.e. $0 = 5$. Impossible. No values of $x, y, z$ can satisfy this. The system has no solution. Whenever elimination produces a row of the form $(0\ 0\ \cdots\ 0 \mid k)$ with $k \neq 0$, you can stop: the system is inconsistent. One glance at the augmented matrix answers the question.
A Matrix That Reads Off Its Own Solution
All the hard work is justified by one payoff. Suppose, after operating on the augmented matrix, we arrive at:
$$\left(\begin{array}{ccc|c}1&2&-1&3\\0&1&4&5\\0&0&1&2\end{array}\right)$$
Bottom row: $z = 2$. Middle row: $y + 4(2) = 5 \Rightarrow y = -3$. Top row: $x + 2(-3) - 2 = 3 \Rightarrow x = 11$. No elimination loops — just read from the bottom up. This is back-substitution, and it is trivial once the matrix is in the right shape. The shape is called row-echelon form.
§8Row-Echelon Form & Pivots
What made that matrix so readable? Three things:
(1) The first nonzero entry of each row is a 1. (2) Everything below that 1 in its column is zero. (3) Each leading 1 sits further right than the one in the row above — a staircase stepping down and to the right.
A matrix is in row-echelon form (REF) if:
1. All zero rows sit at the bottom.
**2.** The first nonzero entry in each nonzero row is a $\mathbf{1}$ — called the leading 1 (or pivot) for that row.
3. Each leading 1 is strictly to the right of the leading 1 in every row above it.
A row-echelon matrix is in reduced row-echelon form if, in addition, each leading 1 is the only nonzero entry in its entire column — zeros above it as well as below. RREF eliminates the need for back-substitution: the solution reads off directly.
After elimination on $\begin{cases}x+2y=5\\3x+5y=12\end{cases}$ we reach REF:
$$\left(\begin{array}{cc|c}1&2&5\\0&1&3\end{array}\right).$$
Back-substitute: $y=3$, then $x+6=5 \Rightarrow x=-1$. Now do one more step — subtract $2\times$row 2 from row 1 — to reach RREF:
$$\left(\begin{array}{cc|c}1&0&-1\\0&1&3\end{array}\right).$$
Read off: $x=-1$, $y=3$. No substitution. RREF does more work up front to eliminate all work at the end.
In $\left(\begin{array}{ccccc}1&3&0&2&-1\\0&0&1&4&5\\0&0&0&1&2\\0&0&0&0&0\end{array}\right)$ the pivots are the three leading 1s, in columns 1, 3, 4 — column 2 is simply skipped (allowed). Zero row safely at the bottom. The pivot columns (1, 3, 4) are fundamental: they correspond to the "determined" variables. The free columns (2, 5) correspond to variables you can set freely. This distinction drives everything in the next two lectures.
§9Practise — Is It in REF?
Before we learn how to produce row-echelon form, you need to be able to recognise it. Work through each matrix below. Apply the three conditions from the definition. Click your answer and read the explanation.