Eigenvalues and Eigenvectors
The special directions a matrix only stretches — the characteristic polynomial, basic eigenvectors, and why they run through all of applied mathematics
§1Quick Recall
Before we meet the new idea, let us gather the three tools from the last two lectures that we will lean on constantly today.
• A square matrix $A$ turns a vector $\mathbf{x}$ into another vector $A\mathbf{x}$. Think of $A$ as a machine that moves arrows around the plane (or space).
• The determinant $\det A$ is a single number, and $A$ is invertible if and only if $\det A \neq 0$ (Lecture 10). We will need this test again in a moment.
• A homogeneous system $M\mathbf{x} = \mathbf{0}$ has a nonzero solution exactly when $\det M = 0$. Gaussian elimination (Lecture 3) finds all of those solutions.
§2Why Eigenvalues? The Big Picture
Eigenvalues are, without exaggeration, one of the two or three most useful ideas in all of applied mathematics. Here is a small taste of where they quietly run the show.
1. Google was built on an eigenvector. The original PageRank algorithm ranks every web page by finding one special eigenvector of an enormous matrix (billions by billions) describing which pages link to which. A page's importance is its entry in that eigenvector. A multi-trillion-dollar company grew out of one eigenvector.
2. Why bridges and buildings fall down. Every structure has natural frequencies of vibration — these are eigenvalues of a stiffness matrix. If an earthquake or a marching crowd hits one of those frequencies, the structure resonates and can tear itself apart. Engineers compute eigenvalues precisely to avoid this.
3. Quantum mechanics is eigenvalue theory. The allowed energy levels of an atom are literally the eigenvalues of an operator called the Hamiltonian. When a neon sign glows, the colours you see are differences of eigenvalues.
4. Data science and PCA. When you compress an image, recommend a movie, or reduce a huge dataset to its most important patterns, you are almost always computing eigenvectors of a covariance matrix (principal component analysis).
The unifying reason all of this works: eigenvectors are the directions in which a complicated matrix acts as simply as a single number. If you understand what $A$ does to its eigenvectors, you understand what $A$ does to everything, because most vectors can be built out of eigenvectors. That is the payoff we are chasing.
§3The Definition
Here is the entire idea in one equation. We are looking for a nonzero vector $\mathbf{x}$ that the matrix $A$ does not turn — it only scales it by some number $\lambda$ (the Greek letter "lambda").
Let $A$ be an $n\times n$ matrix. A number $\lambda$ is called an eigenvalue of $A$ if there is a nonzero vector $\mathbf{x}$ such that
$$A\mathbf{x} = \lambda\mathbf{x}.$$
Every such nonzero vector $\mathbf{x}$ is called an eigenvector of $A$ corresponding to $\lambda$ (a $\lambda$-eigenvector).
$\textbf{1. The vector must be nonzero.}$ Notice $A\mathbf{0} = \lambda\mathbf{0}$ is true for every number $\lambda$, so allowing $\mathbf{x}=\mathbf{0}$ would make every number an "eigenvalue" — useless. Eigenvectors are nonzero by definition.
$\textbf{2. The eigenvalue $\lambda$ can be zero.}$ It is the vector that must be nonzero, not the number. In fact $\lambda = 0$ is an eigenvalue exactly when $A$ is not invertible (you will prove this in Exercise 3.3.3).
§4Seeing It: The Geometric Picture
The definition becomes obvious once you picture it. Draw a vector $\mathbf{x}$ as an arrow. Apply $A$ to get the arrow $A\mathbf{x}$. Ask one question: does the new arrow lie on the same line through the origin as the old one?
• Yes, same line → $\mathbf{x}$ is an eigenvector. The arrow may get longer, shorter, or flip to point backwards, but its line is unchanged. The scale factor is the eigenvalue $\lambda$.
• No, it swung off the line → $\mathbf{x}$ is not an eigenvector. $A$ rotated it into a genuinely new direction.
Special cases worth picturing. If $\lambda > 1$ the eigenvector is stretched; if $0 < \lambda < 1$ it is compressed toward the origin; if $\lambda < 0$ it is flipped to the opposite side (and scaled); if $\lambda = 1$ the vector is left exactly where it was; and if $\lambda = 0$ the vector is crushed onto the origin. A pure rotation matrix (turning everything by, say, $90^\circ$) has no real eigenvectors at all — every arrow is turned — which is why its eigenvalues turn out to be complex (Exercise 3.3.5).
§5Checking a Candidate by the Definition
If someone hands you a matrix, a number, and a vector, checking whether they fit the definition is pure arithmetic: just multiply and compare.
Let $A = \begin{bmatrix} 3 & 5 \\ 1 & -1 \end{bmatrix}$ and $\mathbf{x} = \begin{bmatrix} 5 \\ 1 \end{bmatrix}$. Is $\mathbf{x}$ an eigenvector?
Multiply:
$$A\mathbf{x} = \begin{bmatrix} 3 & 5 \\ 1 & -1 \end{bmatrix}\begin{bmatrix} 5 \\ 1 \end{bmatrix} = \begin{bmatrix} 15+5 \\ 5-1 \end{bmatrix} = \begin{bmatrix} 20 \\ 4 \end{bmatrix} = 4\begin{bmatrix} 5 \\ 1 \end{bmatrix} = 4\mathbf{x}.$$
The output is exactly $4$ times the input, so $A\mathbf{x} = \lambda\mathbf{x}$ holds with $\lambda = 4$. Therefore $\mathbf{x} = (5,1)$ is an eigenvector and $\lambda = 4$ is an eigenvalue of $A$. No characteristic polynomial needed — the definition did all the work.
§6Finding Eigenvalues from Scratch
Start from the defining equation and rearrange it until the determinant test from Lecture 10 can be applied. We want a nonzero $\mathbf{x}$ with
$$A\mathbf{x} = \lambda\mathbf{x}.$$
$\textbf{Step 1 — move everything to one side.}$ Rewrite $\lambda\mathbf{x}$ as $\lambda I\mathbf{x}$ (inserting the identity so both sides are matrix-times-vector), then subtract:
$$\lambda I\mathbf{x} - A\mathbf{x} = \mathbf{0} \quad\Longrightarrow\quad (\lambda I - A)\mathbf{x} = \mathbf{0}.$$
$\textbf{Step 2 — recognise a homogeneous system.}$ This is a homogeneous system with coefficient matrix $\lambda I - A$. We need it to have a nonzero solution $\mathbf{x}$ (remember, eigenvectors cannot be zero).
$\textbf{Step 3 — apply the determinant test.}$ From Lecture 10, a square homogeneous system $M\mathbf{x} = \mathbf{0}$ has a nonzero solution if and only if $\det M = 0$. With $M = \lambda I - A$, the eigenvalues are exactly the numbers $\lambda$ making
$$\det(\lambda I - A) = 0.$$
$\textbf{Step 4 — read it as a polynomial.}$ When you expand $\det(\lambda I - A)$, the result is a polynomial in $\lambda$. Its roots are the eigenvalues. This polynomial deserves a name.
The characteristic polynomial of a square matrix $A$ is
$$c_A(x) = \det(xI - A),$$
a polynomial in the variable $x$. We write $c_A(x)$ to mean "the characteristic polynomial of the matrix $A$." If $A$ is $n\times n$, then $c_A(x)$ has degree exactly $n$.
To find the eigenvalues and eigenvectors of an $n\times n$ matrix $A$:
$\textbf{1.}$ Form the matrix $xI - A$ and compute the characteristic polynomial $c_A(x) = \det(xI - A)$.
$\textbf{2.}$ Find the roots of $c_A(x) = 0$. These roots are the eigenvalues $\lambda_1, \lambda_2, \ldots$
$\textbf{3.}$ For each eigenvalue $\lambda$, solve the homogeneous system $(\lambda I - A)\mathbf{x} = \mathbf{0}$ by Gaussian elimination. The nonzero solutions are the $\lambda$-eigenvectors.
Take $A = \begin{bmatrix} 3 & 5 \\ 1 & -1 \end{bmatrix}$ again, but this time discover its eigenvalues with no vector given.
$\textbf{Step 1 — characteristic polynomial.}$ Form $xI - A = \begin{bmatrix} x-3 & -5 \\ -1 & x+1 \end{bmatrix}$ and take its determinant:
$$c_A(x) = \det(xI - A) = (x-3)(x+1) - (-5)(-1) = x^2 - 2x - 3 - 5 = x^2 - 2x - 8.$$
$\textbf{Step 2 — find the roots.}$ Factor: $x^2 - 2x - 8 = (x-4)(x+2)$. So the roots are $x = 4$ and $x = -2$.
The eigenvalues are $\lambda_1 = 4$ (the one we verified earlier) and $\lambda_2 = -2$ (the hidden second one). The procedure recovered both from nothing but the matrix.
§7The Master Theorem
Everything above is packaged into one clean statement.
Let $A$ be an $n\times n$ matrix.
$\textbf{1.}$ The eigenvalues $\lambda$ of $A$ are the roots of the characteristic polynomial $c_A(x)$ of $A$.
$\textbf{2.}$ The $\lambda$-eigenvectors $\mathbf{x}$ are the nonzero solutions to the homogeneous system $$(\lambda I - A)\mathbf{x} = \mathbf{0}$$ of linear equations with $\lambda I - A$ as coefficient matrix.
In practice, solving the equations in part 2 is a routine application of Gaussian elimination. But finding the eigenvalues — the roots in part 1 — can be genuinely hard, often requiring computers. Our examples and exercises are built so that the roots come out as easy integers, but do not be misled: for the matrices in real applications, eigenvalues are usually not so obliging. There are entire numerical methods (see Section 8.5 of Nicholson) devoted just to approximating them.
§8How Many Eigenvectors? Basic Eigenvectors
A single eigenvalue does not come with just one eigenvector — it comes with a whole family. Every nonzero solution $\mathbf{x}$ of $(\lambda I - A)\mathbf{x} = \mathbf{0}$ is an eigenvector. And if $\mathbf{x}$ is an eigenvector, so is any nonzero multiple $k\mathbf{x}$, because
$$A(k\mathbf{x}) = k(A\mathbf{x}) = k(\lambda\mathbf{x}) = \lambda(k\mathbf{x}).$$
Geometrically this is just the statement that the whole eigen-line consists of eigenvectors. Recall from Lecture 3 (Theorem 1.3.2) that the solutions of a homogeneous system are all linear combinations of certain basic solutions produced by the Gaussian algorithm. We give those a name here.
Any set of nonzero multiples of the basic solutions of $(\lambda I - A)\mathbf{x} = \mathbf{0}$ is called a set of basic eigenvectors corresponding to $\lambda$. In practice we scale each basic solution to clear fractions, giving the tidiest possible integer eigenvector to represent the whole line (or plane) of eigenvectors.
§9Full Worked Examples
Find the characteristic polynomial, the eigenvalues, and basic eigenvectors of $A = \begin{bmatrix} 4 & 2 \\ 1 & 3 \end{bmatrix}$.
Find the characteristic polynomial, eigenvalues, and basic eigenvectors of $A = \begin{bmatrix} 2 & 0 & 0 \\ 1 & 2 & -1 \\ 1 & 3 & -2 \end{bmatrix}$.
§10Interactive: The Eigenvector Playground
Type any $2\times2$ or $3\times3$ matrix below. The tool computes its eigenvalues and eigenvectors, and (for $2\times2$) draws each eigenvector together with $A$ applied to it — watch how $A\mathbf{x}$ always lands back on the same dashed eigen-line. Try a rotation like $\begin{bmatrix} 0 & -1 \\ 1 & 0 \end{bmatrix}$ to see what "no real eigenvector" looks like.
§11A and Aᵀ Share Their Eigenvalues
Show that a square matrix $A$ and its transpose $A^{\mathsf{T}}$ have the same characteristic polynomial, and hence the same eigenvalues.
- $\lambda$ is an eigenvalue of $A$ with eigenvector $\mathbf{x} \neq \mathbf{0}$ means $A\mathbf{x} = \lambda\mathbf{x}$: $A$ only scales $\mathbf{x}$, never turns it.
- Eigenvectors must be nonzero; the eigenvalue itself may be zero.
- Characteristic polynomial: $c_A(x) = \det(xI - A)$, of degree $n$ for an $n\times n$ matrix.
- $\lambda$ is an eigenvalue $\iff c_A(\lambda) = 0$ (Theorem 3.3.2, part 1).
- Eigenvectors for $\lambda$: nonzero solutions of $(\lambda I - A)\mathbf{x} = \mathbf{0}$ (Theorem 3.3.2, part 2).
- Scale basic solutions to integers → basic eigenvectors; any nonzero multiple is still an eigenvector.
- $A$ and $A^{\mathsf{T}}$ have the same characteristic polynomial, hence the same eigenvalues.
§12Solutions to Section 3.3 Exercises
Let $A$ be $n\times n$ and $A_1 = A - \alpha I$ with $\alpha \in \mathbb{R}$. Show $\lambda$ is an eigenvalue of $A$ if and only if $\lambda - \alpha$ is an eigenvalue of $A_1$. How do the eigenvectors compare?
Given $A = \begin{bmatrix} a & b \\ c & d \end{bmatrix}$, show: (a) $c_A(x) = x^2 - (\operatorname{tr}A)x + \det A$, where $\operatorname{tr}A = a + d$; (b) the eigenvalues are $\tfrac12\big[(a+d) \pm \sqrt{(a-d)^2 + 4bc}\big]$.
Let $A$ be $n\times n$ and $r \neq 0$ a real number.
Let $A$ be an invertible $n\times n$ matrix.
Suppose $\lambda$ is an eigenvalue of a square matrix $A$ with eigenvector $\mathbf{x} \neq \mathbf{0}$.
An $n\times n$ matrix $A$ is nilpotent if $A^m = 0$ for some $m \geq 1$.
You can now find the special directions a matrix only stretches. Next we use them to make matrices simple.
When a matrix has enough independent eigenvectors, we can rewrite it in a coordinate system where it becomes purely diagonal — a process called diagonalization. That single trick makes computing $A^{100}$, solving systems of differential equations, and understanding long-term behaviour almost effortless. Eigenvalues are the doorway; diagonalization is the room they open onto.