Home›Courses›Linear Algebra›Week 4 · Lecture 12
← Lecture 11
On this pageRecall·Why Powers·Graph Paths·Diagonal Matrices·Diagonalizing·When Possible?·Examples·Powers Aⁿ·Similar Matrices·Invariants·Cayley–Hamilton·Charpoly Info·λ² Question·Dynamical Systems·Trajectories·Exercises
Lecture 13 →
MATH-120 · Linear Algebra · Lecture 1229 June 2026

Diagonalization and Dynamical Systems

Splitting a matrix into P·D·P⁻¹, similar matrices, the Cayley–Hamilton theorem, and using eigenvalues to predict the long-term future

§1Quick Recall

In Lecture 11 we learned to find the special directions of a matrix. Two ideas from there are the foundation of everything today.

• A number $\lambda$ and a nonzero vector $\mathbf{x}$ form an eigenvalue–eigenvector pair of $A$ when $A\mathbf{x} = \lambda\mathbf{x}$: the matrix only stretches $\mathbf{x}$, it does not turn it.

• We find eigenvalues as the roots of the characteristic polynomial $c_A(x) = \det(xI - A)$, and eigenvectors as the nonzero solutions of $(\lambda I - A)\mathbf{x} = \mathbf{0}$.

🎯
The goal of this lecture
Eigenvectors let us rebuild a matrix in a simpler form. If we line up the eigenvectors as columns of a matrix $P$, then $P^{-1}AP$ becomes a diagonal matrix — the simplest kind there is. This single move makes hard problems (like computing $A^{100}$ or predicting the far future of a population) almost trivial.

§2The Motivation: Why We Care About Powers of a Matrix

Here is a problem that looks impossible by hand. Suppose $A$ is a $2\times2$ matrix and you need $A^{50}$. Multiplying $A$ by itself fifty times is hopeless — the numbers explode and the arithmetic is enormous. Yet powers of matrices appear everywhere:

• Populations over time. If a population this year is $\mathbf{v}_0$ and next year is $A\mathbf{v}_0$, then in $k$ years it is $A^k\mathbf{v}_0$. To predict the far future you need high powers of $A$.

• Counting paths in a network. As we are about to see, $A^k$ literally counts the number of routes of length $k$ between points in a graph.

• Google, Markov chains, random walks. All are questions about $A^k$ for large $k$.

Diagonalization is the trick that makes $A^k$ easy. The reason is a small miracle: powers of a diagonal matrix are trivial. That is where we start.

§3A Concrete Reason: Counting Paths in a Graph

Consider a small network of four locations. We draw a dot (vertex) for each location and a line (edge) between two locations if you can travel directly between them.

1234
A graph on 4 vertices with edges 1–2, 1–3, 2–3, 3–4.

We record this network in an adjacency matrix $A$, where the entry in row $i$, column $j$ is $1$ if there is an edge between vertex $i$ and vertex $j$, and $0$ otherwise. For the graph above:

$$A = \begin{bmatrix} 0 & 1 & 1 & 0 \\ 1 & 0 & 1 & 0 \\ 1 & 1 & 0 & 1 \\ 0 & 0 & 1 & 0 \end{bmatrix}.$$

(Row 1 has $1$s in columns 2 and 3 because vertex 1 connects to vertices 2 and 3, and so on. The matrix is symmetric because the edges have no direction.)

The path-counting theorem

The $(i,j)$-entry of $A^k$ is exactly the number of walks of length $k$ from vertex $i$ to vertex $j$ (a walk of length $k$ is a sequence of $k$ edges, steps allowed to repeat).

$\textbf{Why this is true (the key idea).}$ Look at $A^2$. Its $(i,j)$-entry is $\sum_m A_{im}A_{mj}$. The term $A_{im}A_{mj}$ equals $1$ exactly when there is an edge $i\to m$ and an edge $m\to j$ — that is, one walk $i\to m\to j$ of length $2$. Summing over all middle vertices $m$ counts every length-2 walk. The same bookkeeping repeats for $A^3$, $A^4$, and beyond.

Computing the powers gives:

$$A^2 = \begin{bmatrix} 2 & 1 & 1 & 1 \\ 1 & 2 & 1 & 1 \\ 1 & 1 & 3 & 0 \\ 1 & 1 & 0 & 1 \end{bmatrix}, \qquad A^3 = \begin{bmatrix} 2 & 3 & 4 & 1 \\ 3 & 2 & 4 & 1 \\ 4 & 4 & 2 & 3 \\ 1 & 1 & 3 & 0 \end{bmatrix}.$$

Reading off: the $(1,1)$-entry of $A^2$ is $2$, so there are two walks of length 2 from vertex 1 back to itself ($1\to2\to1$ and $1\to3\to1$). The $(3,3)$-entry is $3$: three length-2 loops at vertex 3 (via $1$, via $2$, via $4$). And $A^3_{13} = 4$ says there are four length-3 walks from vertex 1 to vertex 3.

💡
The punchline
To answer "how many routes of length 100 connect these two cities?" you need $A^{100}$. Diagonalization is how you compute such a power without doing 100 matrix multiplications.

§4Why Diagonal Matrices Are So Easy

Diagonal matrix

A square matrix $D$ is diagonal if every entry off the main diagonal is zero. We write

$$D = \operatorname{diag}(\lambda_1, \lambda_2, \ldots, \lambda_n) = \begin{bmatrix} \lambda_1 & 0 & \cdots & 0 \\ 0 & \lambda_2 & \cdots & 0 \\ \vdots & \vdots & \ddots & \vdots \\ 0 & 0 & \cdots & \lambda_n \end{bmatrix}.$$

Diagonal matrices multiply and add entry-by-entry down the diagonal. If $D = \operatorname{diag}(\lambda_1,\ldots,\lambda_n)$ and $E = \operatorname{diag}(\mu_1,\ldots,\mu_n)$, then

$$DE = \operatorname{diag}(\lambda_1\mu_1, \ldots, \lambda_n\mu_n), \qquad D + E = \operatorname{diag}(\lambda_1+\mu_1, \ldots, \lambda_n+\mu_n).$$

The consequence we care about most: powers are effortless. Multiplying $D$ by itself just raises each diagonal entry to that power:

$$D^k = \operatorname{diag}(\lambda_1^k, \lambda_2^k, \ldots, \lambda_n^k).$$

So $\operatorname{diag}(2,3)^{10} = \operatorname{diag}(2^{10}, 3^{10}) = \operatorname{diag}(1024, 59049)$ — no matrix multiplication at all. If only every matrix were diagonal. Diagonalization is the art of making a matrix diagonal by changing coordinates.

§5What Diagonalization Means

Diagonalizable matrix

An $n\times n$ matrix $A$ is diagonalizable if there is an invertible matrix $P$ such that $$P^{-1}AP = D$$ is diagonal. The matrix $P$ is called a diagonalizing matrix for $A$. Equivalently, $A = PDP^{-1}$.

So what are $P$ and $D$? This is the heart of the lecture, and the answer connects directly to Lecture 11.

Theorem 3.3.4 — the recipe for P and D

Let $A$ be an $n\times n$ matrix.

$\textbf{1.}$ $A$ is diagonalizable if and only if it has $n$ eigenvectors $\mathbf{x}_1, \ldots, \mathbf{x}_n$ such that the matrix $P = [\,\mathbf{x}_1 \; \mathbf{x}_2 \; \cdots \; \mathbf{x}_n\,]$ (eigenvectors as columns) is invertible.

$\textbf{2.}$ In that case $P^{-1}AP = \operatorname{diag}(\lambda_1, \ldots, \lambda_n)$, where $\lambda_i$ is the eigenvalue belonging to the eigenvector $\mathbf{x}_i$.

🧩
In one sentence
$P$ is the matrix whose columns are the eigenvectors, and $D$ is the diagonal matrix whose diagonal entries are the matching eigenvalues, in the same order.

$\textbf{Why the recipe works.}$ Saying $P^{-1}AP = D$ is the same as saying $AP = PD$. Write $P = [\mathbf{x}_1 \cdots \mathbf{x}_n]$ by columns. The left side $AP$ has columns $A\mathbf{x}_1, \ldots, A\mathbf{x}_n$. The right side $PD$, with $D = \operatorname{diag}(\lambda_1,\ldots,\lambda_n)$, has columns $\lambda_1\mathbf{x}_1, \ldots, \lambda_n\mathbf{x}_n$. Matching columns gives

$$A\mathbf{x}_i = \lambda_i\mathbf{x}_i \quad \text{for each } i.$$

That is exactly the eigenvalue equation. So $P^{-1}AP$ is diagonal precisely when the columns of $P$ are eigenvectors and the diagonal of $D$ holds their eigenvalues.

The concrete 3×3 picture. If $A$ is $3\times3$ with eigenvalues $\lambda_1, \lambda_2, \lambda_3$ and corresponding eigenvectors $\mathbf{v}_1, \mathbf{v}_2, \mathbf{v}_3$, then

$$P = \big[\,\mathbf{v}_1 \; \mathbf{v}_2 \; \mathbf{v}_3\,\big], \qquad D = \begin{bmatrix} \lambda_1 & 0 & 0 \\ 0 & \lambda_2 & 0 \\ 0 & 0 & \lambda_3 \end{bmatrix}.$$

🔀
The order is your choice
If you reorder the eigenvector columns of $P$, the eigenvalues on the diagonal of $D$ reorder to match. Putting $\mathbf{v}_2$ first gives $\lambda_2$ in the top-left slot. So you can arrange the eigenvalues along $D$ in any order you like — just keep each eigenvector paired with its own eigenvalue.
⚠️ The repeated-eigenvalue case

When an eigenvalue repeats, the recipe still works — if that eigenvalue supplies enough independent eigenvectors. A twice-repeated eigenvalue needs two independent eigenvectors (a whole plane of them) to fill two columns of $P$. If it only gives one, $P$ has too few columns to be invertible, and $A$ is not diagonalizable. We make this precise next.

§6How to Tell If a Matrix Is Diagonalizable

This is the first key question. There are three checkpoints, from quickest to most careful.

Multiplicity

An eigenvalue $\lambda$ has multiplicity $m$ if it occurs $m$ times as a root of the characteristic polynomial. For instance, $c_A(x) = (x-2)(x+1)^2$ has $\lambda = 2$ with multiplicity 1 and $\lambda = -1$ with multiplicity 2.

The three tests for diagonalizability

$\textbf{Test 1 (definition).}$ $A$ is diagonalizable if there exist an invertible $P$ and diagonal $D$ with $A = PDP^{-1}$.

$\textbf{Test 2 (distinct eigenvalues — the easy win).}$ If an $n\times n$ matrix has $n$ distinct eigenvalues, it is automatically diagonalizable (Theorem 3.3.6). Distinct eigenvalues always give independent eigenvectors.

$\textbf{Test 3 (the general rule — Theorem 3.3.5).}$ $A$ is diagonalizable if and only if every eigenvalue of multiplicity $m$ produces exactly $m$ basic eigenvectors — that is, the solution of $(\lambda I - A)\mathbf{x} = \mathbf{0}$ has exactly $m$ free parameters.

Diagonalization Algorithm

$\textbf{Step 1.}$ Find the distinct eigenvalues of $A$ (roots of $c_A(x)$).

$\textbf{Step 2.}$ For each eigenvalue, find its basic eigenvectors as basic solutions of $(\lambda I - A)\mathbf{x} = \mathbf{0}$.

$\textbf{Step 3.}$ $A$ is diagonalizable if and only if there are $n$ basic eigenvectors in total.

$\textbf{Step 4.}$ If so, put those eigenvectors as columns of $P$; then $P$ is invertible and $P^{-1}AP$ is diagonal with the matching eigenvalues.

📝
A common trap
"$n$ distinct eigenvalues" is sufficient but not necessary. A matrix with repeated eigenvalues can still be diagonalizable — it just has to pass Test 3. Example 2 below shows exactly this: a repeated eigenvalue that supplies its full quota of eigenvectors.

§7Worked Examples: Every Scenario

★ 1Distinct eigenvalues — the clean case

Diagonalize $A = \begin{bmatrix} 2 & 0 & 0 \\ 1 & 2 & -1 \\ 1 & 3 & -2 \end{bmatrix}$ (the matrix from Lecture 11, Example 4).

★ 2Repeated eigenvalue — still diagonalizable

Diagonalize $A = \begin{bmatrix} 0 & 1 & 1 \\ 1 & 0 & 1 \\ 1 & 1 & 0 \end{bmatrix}$.

★ 3Repeated eigenvalue — NOT diagonalizable

Show that $A = \begin{bmatrix} 1 & 1 \\ 0 & 1 \end{bmatrix}$ is not diagonalizable.

Try it yourself

Enter a $2\times2$ matrix below. The tool tells you whether it is diagonalizable, gives $P$ and $D$, and computes $A^n$ for you. Test the three scenarios: $\begin{bmatrix} 4&2\\1&3 \end{bmatrix}$ (distinct), $\begin{bmatrix} 2&0\\0&2 \end{bmatrix}$ (repeated but scalar — diagonalizable), $\begin{bmatrix} 1&1\\0&1 \end{bmatrix}$ (defective).

🎛 Diagonalize & Power Calculator (2×2)
ENTER MATRIX A
[
]
POWER n
n = 3
Diagonalizable ✓
λ₁ = 5, eigenvector (1, 0.500)
λ₂ = 2, eigenvector (1, -1)
D = diag(5, 2)
A^3 = P · D^3 · P⁻¹ =
8678
3947
Numerical values — verify by hand. Try [[4,2],[1,3]] (clean), [[1,1],[0,1]] (defective), [[2,0],[0,2]] (scalar).

§8The Payoff: Computing Aⁿ

Here is where all the setup pays for itself. If $A = PDP^{-1}$, then squaring telescopes beautifully:

$$A^2 = (PDP^{-1})(PDP^{-1}) = PD\underbrace{(P^{-1}P)}_{I}DP^{-1} = PD^2P^{-1}.$$

The inner $P^{-1}P$ collapses to the identity. The same cancellation repeats for every power, giving the master formula:

Power formula

If $A = PDP^{-1}$ with $D$ diagonal, then for every $n \geq 1$, $$A^n = PD^nP^{-1} = P\operatorname{diag}(\lambda_1^n, \ldots, \lambda_k^n)P^{-1}.$$ Computing $A^n$ costs just two matrix multiplications, no matter how large $n$ is — because $D^n$ is free.

Example 4Computing Aⁿ (Exercise 3.3.8a)

Find $A^n$ for $A = \begin{bmatrix} 6 & -5 \\ 2 & -1 \end{bmatrix}$, using the given $P = \begin{bmatrix} 1 & 5 \\ 1 & 2 \end{bmatrix}$.

§9Similar Matrices — The Bigger Idea

Diagonalization is a special case of a broader relationship. The move $P^{-1}AP$ appears so often it earns its own name.

Similar matrices

Two $n\times n$ matrices $A$ and $B$ are similar, written $A \sim B$, if there is an invertible matrix $P$ with $$A = PBP^{-1} \quad(\text{equivalently } B = P^{-1}AP).$$ A matrix is diagonalizable precisely when it is similar to a diagonal matrix.

$\textbf{The right mental model.}$ Similar matrices are the same linear transformation seen from two different coordinate systems. The matrix $P$ is the "dictionary" translating between the two viewpoints. Nothing essential about the transformation changes — only the numbers used to describe it.

🤔
How many matrices are similar to a given one?
Ask: how many matrices are similar to $\begin{bmatrix} 1 & 2 \\ -1 & 0 \end{bmatrix}$? The answer is infinitely many — one for each invertible $P$. Every choice of $P$ gives a matrix $B = P^{-1}AP$ that looks different but represents the same transformation. To produce five of them, just pick five different invertible matrices $P$ and compute $P^{-1}AP$ for each. The interesting question is not how to make them, but what they all share — which is the next section.

§10What Similar Matrices Share (Invariants)

Similar matrices are not identical, but they agree on every quantity that describes the underlying transformation. These shared quantities are called invariants.

Similar matrices share their characteristic polynomial

If $A \sim B$, then $c_A(x) = c_B(x)$. Consequently they have the same eigenvalues, the same determinant, and the same trace.

$\textbf{Proof, step by step.}$ Suppose $B = P^{-1}AP$. We compute $c_B(x) = \det(xI - B)$ and show it equals $c_A(x)$.

$\textbf{Step 1.}$ Substitute $B$: $\;xI - B = xI - P^{-1}AP$.

$\textbf{Step 2.}$ Rewrite $xI$ using $P^{-1}P = I$: since $xI = P^{-1}(xI)P$ (the scalar $x$ passes through), we get $xI - B = P^{-1}(xI)P - P^{-1}AP = P^{-1}(xI - A)P$.

$\textbf{Step 3.}$ Take determinants and use the product rule $\det(XYZ) = \det X\det Y\det Z$ (Lecture 10):

$$c_B(x) = \det(xI - B) = \det(P^{-1})\det(xI - A)\det(P).$$

$\textbf{Step 4.}$ Since $\det(P^{-1})\det(P) = \det(P^{-1}P) = \det(I) = 1$, the two $P$-factors cancel:

$$c_B(x) = \det(xI - A) = c_A(x). \qquad\blacksquare$$

Because eigenvalues are the roots of the characteristic polynomial, determinant is the product of eigenvalues, and trace is their sum, all three are automatically shared once the polynomials match.

More shared invariants

If $A \sim B$ via $B = P^{-1}AP$, then also:

• $\textbf{Rank}$ and $\textbf{invertibility}$ agree (multiplying by invertible $P$, $P^{-1}$ cannot change rank).

• $\textbf{Powers stay similar}$: $B^k = P^{-1}A^kP$, proven by the same telescoping cancellation as the power formula.

• $\textbf{Similarity is an equivalence relation}$: $A \sim A$ (take $P = I$); if $A \sim B$ then $B \sim A$; and if $A \sim B$, $B \sim C$ then $A \sim C$. So "similar" cleanly partitions all matrices into families.

⚠️ A warning: the converse fails

Equal characteristic polynomials do not force similarity. In Exercise 3.3.26 you meet two matrices with the identical polynomial $(x+1)^2(x-2)$ where one is diagonalizable and the other is not — so they cannot be similar. Same eigenvalues is necessary for similarity, but not sufficient.

§11The Cayley–Hamilton Theorem

👥
A note on the mathematicians
$\textbf{Arthur Cayley}$ (1821–1895) was an English mathematician who, alongside a full career as a lawyer, essentially invented matrix algebra as we know it. $\textbf{William Rowan Hamilton}$ (1805–1865) was an Irish prodigy who could read several languages as a child and later discovered the quaternions — reportedly carving the defining equation into a Dublin bridge in a flash of insight. The theorem that carries both names says something almost magical: every matrix satisfies its own characteristic equation.
Cayley–Hamilton Theorem

Every square matrix $A$ satisfies its own characteristic polynomial: if $c_A(x)$ is the characteristic polynomial of $A$, then substituting the matrix $A$ for the variable gives the zero matrix, $$c_A(A) = 0.$$ (Here a constant term $c_0$ becomes $c_0 I$, since you cannot add a plain number to a matrix.)

$\textbf{What "$c_A(A)$" means.}$ If $c_A(x) = x^2 - 5x + 6$, then $c_A(A) = A^2 - 5A + 6I$. Cayley–Hamilton promises this equals the zero matrix. This is the evaluation of a polynomial at a matrix: replace each power of $x$ by the same power of $A$, and replace the constant by that constant times $I$.

$\textbf{Proof for the diagonalizable case}$ (the general case needs Chapter 8). Suppose $A = PDP^{-1}$ with $D = \operatorname{diag}(\lambda_1, \ldots, \lambda_n)$. We use two facts:

$\textbf{Fact 1.}$ For any polynomial $p$, $p(A) = P\,p(D)\,P^{-1}$. (Each power $A^k = PD^kP^{-1}$ by the power formula, and the $P, P^{-1}$ factors pull outside the whole sum.)

$\textbf{Fact 2.}$ For a diagonal matrix, $p(D) = \operatorname{diag}\big(p(\lambda_1), \ldots, p(\lambda_n)\big)$ — you just evaluate $p$ at each diagonal entry.

Now take $p = c_A$, the characteristic polynomial. Every eigenvalue $\lambda_i$ is a root, so $c_A(\lambda_i) = 0$ for each $i$. Therefore

$$c_A(D) = \operatorname{diag}\big(c_A(\lambda_1), \ldots, c_A(\lambda_n)\big) = \operatorname{diag}(0, \ldots, 0) = 0.$$

Feeding this back through Fact 1:

$$c_A(A) = P\,c_A(D)\,P^{-1} = P\cdot 0\cdot P^{-1} = 0. \qquad\blacksquare$$

⚙️
Why this is useful
Cayley–Hamilton lets you express high powers of $A$ in terms of low ones. For a $2\times2$ matrix, $A^2 = (\operatorname{tr}A)A - (\det A)I$, so every power $A^n$ can be rewritten using just $A$ and $I$. It also gives a slick way to compute $A^{-1}$ as a polynomial in $A$. The same trick powers many algorithms in control theory and computer graphics.
Example 5Diagonalizable matrices inherit eigenvalue identities (Example 3.3.11)

If $\lambda^3 = 5\lambda$ for every eigenvalue of a diagonalizable matrix $A$, show that $A^3 = 5A$.

§12What the Characteristic Polynomial Encodes

The final question: how much of a matrix is captured by its characteristic polynomial? Let us dissect the $2\times2$ case fully.

For $A = \begin{bmatrix} a & b \\ c & d \end{bmatrix}$, expand $c_A(x) = \det(xI - A)$:

$$c_A(x) = \begin{vmatrix} x-a & -b \\ -c & x-d \end{vmatrix} = (x-a)(x-d) - bc = x^2 - (a+d)x + (ad - bc).$$

Two familiar quantities jump out of the coefficients:

• The coefficient of $x$ is $-(a+d)$. The number $a + d$ — the sum of the diagonal — is the trace, written $\operatorname{tr}(A)$.

• The constant term is $ad - bc$, which is exactly $\det(A)$.

Trace and determinant live in the characteristic polynomial

For any $2\times2$ matrix, $$c_A(x) = x^2 - \operatorname{tr}(A)\,x + \det(A).$$ Reading off the roots: the two eigenvalues sum to the trace and multiply to the determinant.

$\textbf{The pattern continues for $3\times3$.}$ Expanding $c_A(x)$ for a $3\times3$ matrix gives

$$c_A(x) = x^3 - \operatorname{tr}(A)\,x^2 + (\text{sum of principal } 2\times2 \text{ minors})\,x - \det(A).$$

The leading behaviour and the two ends are always the same story: the $x^{n-1}$ coefficient is $-\operatorname{tr}(A)$, and the constant term is $(-1)^n\det(A)$. So trace and determinant are always encoded in the characteristic polynomial, for any size.

Monic polynomial

A polynomial is monic if its leading coefficient (the coefficient of the highest power) is $1$. For example $x^2 - 7x + 10$ is monic; $2x^2 - 3$ is not.

🔑
Why monic matters here
Every characteristic polynomial $c_A(x) = \det(xI - A)$ is monic of degree $n$: the highest term always comes out as $x^n$ with coefficient exactly $1$ (it is the product of the $n$ diagonal entries $x - a_{ii}$). This is why we can factor $c_A(x) = (x-\lambda_1)(x-\lambda_2)\cdots(x-\lambda_n)$ with leading coefficient $1$, and why the eigenvalues sum to the trace and multiply to $(-1)^n$ times the constant term. Being monic keeps all these bookkeeping relations clean.

$\textbf{Generalizing to $n\times n$.}$ For an $n\times n$ matrix, $c_A(x)$ is a monic degree-$n$ polynomial, its roots are the $n$ eigenvalues (with multiplicity), the sum of the roots is $\operatorname{tr}(A)$, and the product of the roots is $(-1)^n$ times the constant term, which equals $\det(A)$. The characteristic polynomial is a compact fingerprint carrying the eigenvalues, the trace, and the determinant all at once.

§13A Key Question: Eigenvalues of A²

$\textbf{Question.}$ If $\lambda$ is an eigenvalue of $A$, what is the matching eigenvalue of $A^2$?

$\textbf{Answer.}$ It is $\lambda^2$, with the same eigenvector. Here is the one-line reason:

$$A\mathbf{v} = \lambda\mathbf{v} \;\Longrightarrow\; A^2\mathbf{v} = A(A\mathbf{v}) = A(\lambda\mathbf{v}) = \lambda(A\mathbf{v}) = \lambda(\lambda\mathbf{v}) = \lambda^2\mathbf{v}.$$

👁
Seeing why it works
Geometrically: applying $A$ stretches $\mathbf{v}$ by $\lambda$. Applying $A$ a second time stretches the result by $\lambda$ again. Two stretches by $\lambda$ compound to a single stretch by $\lambda \cdot \lambda = \lambda^2$. The direction never changes because $\mathbf{v}$ is an eigenvector at every step. So $A^k\mathbf{v} = \lambda^k\mathbf{v}$ for every power $k$.

$\textbf{The relationship between $c_A(x)$ and $c_{A^2}(x)$.}$ Since the eigenvalues of $A^2$ are the squares of the eigenvalues of $A$, their characteristic polynomials are tightly linked. The precise identity (Exercise 3.3.22) is

$$c_{A^2}(x^2) = (-1)^n\,c_A(x)\,c_A(-x).$$

$\textbf{Why this holds.}$ Factor $c_A(x) = \prod_i (x - \lambda_i)$. Then $c_A(x)c_A(-x) = \prod_i (x-\lambda_i)(-x-\lambda_i) = \prod_i -(x-\lambda_i)(x+\lambda_i) = (-1)^n\prod_i (x^2 - \lambda_i^2)$. Meanwhile $c_{A^2}(y) = \prod_i (y - \lambda_i^2)$, so $c_{A^2}(x^2) = \prod_i (x^2 - \lambda_i^2)$. Comparing the two gives the identity. The squares $\lambda_i^2$ are precisely the eigenvalues of $A^2$, exactly as the geometric argument predicted.

§14Linear Dynamical Systems: Predicting the Future

Now we cash in everything. A linear dynamical system is a sequence of vectors where each one is obtained from the previous by multiplying by a fixed matrix.

Linear dynamical system

Given a starting vector $\mathbf{v}_0$ and a square matrix $A$, the sequence defined by $$\mathbf{v}_{k+1} = A\mathbf{v}_k \quad (k = 0, 1, 2, \ldots)$$ is a linear dynamical system. Unwinding the recurrence, $\mathbf{v}_k = A^k\mathbf{v}_0$.

So the whole future is governed by powers of $A$ — and we now know how to compute those with diagonalization. If $A = PDP^{-1}$ with eigenvalues $\lambda_i$ and eigenvectors $\mathbf{x}_i$, a short calculation gives an exact formula for every step. Writing $\mathbf{b} = P^{-1}\mathbf{v}_0 = (b_1, \ldots, b_n)^{\mathsf{T}}$:

$$\mathbf{v}_k = b_1\lambda_1^k\,\mathbf{x}_1 + b_2\lambda_2^k\,\mathbf{x}_2 + \cdots + b_n\lambda_n^k\,\mathbf{x}_n.$$

Each eigenvector contributes a term that grows or shrinks like its eigenvalue raised to the $k$. This is the exact behaviour of the system, written in "eigen-coordinates."

Example 6A bird population (Example 3.3.12)

Adult and juvenile female birds satisfy $a_{k+1} = \tfrac12 a_k + \tfrac14 j_k$ and $j_{k+1} = 2a_k$, starting from $a_0 = 100$, $j_0 = 40$. Find formulas for $a_k$ and $j_k$, and the long-term behaviour.

Dominant eigenvalue

An eigenvalue $\lambda_1$ is dominant if it has multiplicity 1 and $|\lambda_1| > |\lambda_i|$ for every other eigenvalue. When a dominant eigenvalue exists, factoring $\lambda_1^k$ out of the exact formula shows all other terms shrink relative to the first, so for large $k$, $$\mathbf{v}_k \approx b_1\lambda_1^k\,\mathbf{x}_1.$$ The long-term direction of the system is simply the dominant eigenvector.

Example 7Solving a linear recurrence (Example 3.3.14)

A sequence satisfies $x_0 = 1$, $x_1 = -1$, and $x_{k+2} = 2x_k - x_{k+1}$ for all $k \geq 0$. Find a closed formula for $x_k$.

§15Picturing the Behaviour: Trajectories

Plotting the sequence $\mathbf{v}_0, \mathbf{v}_1, \mathbf{v}_2, \ldots$ as points in the plane gives the trajectory of the system. The eigenvalues completely determine the shape. There are four archetypes, decided by the sizes of $|\lambda|$.

• Attractor — both $|\lambda| < 1$. Every trajectory spirals or slides into the origin. The origin is stable.

• Repellor — both $|\lambda| > 1$. Every trajectory flees away from the origin.

• Saddle — one $|\lambda| > 1$ and one $|\lambda| < 1$. Trajectories approach along one eigen-line and escape along the other. Only the special starting points on the shrinking line reach the origin.

• Spiral — complex eigenvalues. Trajectories rotate around the origin (spiralling in if $|\lambda| < 1$, out if $|\lambda| > 1$). No real eigen-lines exist, so the motion turns.

Switch between the four cases and watch how the paths change:

🎛 Trajectory Plotter — vₖ₊₁ = Avₖ
Each coloured path is a trajectory from a different starting point v₀. White dot = origin.
🌐
Where this leads: Google PageRank
The web is a giant dynamical system. Each page's importance depends on the importance of pages linking to it — a self-referential loop that is exactly $\mathbf{v} = A\mathbf{v}$ for a connectivity matrix $A$. Google's original PageRank is the dominant eigenvector of that matrix (scaled so its entries are positive and sum to 1). Repeatedly applying $A$ to any starting guess converges to that dominant eigenvector — the far-future state of the dynamical system. The single most important algorithm on the internet is a diagonalization problem in disguise.
Summary of key points
  • $A$ is diagonalizable if $P^{-1}AP = D$ for an invertible $P$ and diagonal $D$; then $A = PDP^{-1}$.
  • $P$ = eigenvectors as columns; $D$ = matching eigenvalues on the diagonal (Theorem 3.3.4).
  • $n$ distinct eigenvalues $\Rightarrow$ diagonalizable. In general: each eigenvalue of multiplicity $m$ must give $m$ basic eigenvectors (Theorem 3.3.5).
  • Power formula: $A^n = PD^nP^{-1}$, and $D^n$ just raises each diagonal entry to the $n$.
  • $A \sim B$ (similar) means $B = P^{-1}AP$; similar matrices share $c_A(x)$, eigenvalues, trace, determinant, and rank.
  • Cayley–Hamilton: every matrix satisfies its own characteristic polynomial, $c_A(A) = 0$.
  • $c_A(x)$ is monic of degree $n$; it encodes the eigenvalues, with $\operatorname{tr}(A)$ and $\det(A)$ in its coefficients.
  • If $\lambda$ is an eigenvalue of $A$, then $\lambda^k$ is one of $A^k$ (same eigenvector).
  • Dynamical system $\mathbf{v}_k = A^k\mathbf{v}_0 = \sum_i b_i\lambda_i^k\mathbf{x}_i$; long-term behaviour is set by the dominant eigenvalue.

§16Solutions to Section 3.3 Exercises (Diagonalization)

The eigenvalue/eigenvector-only exercises (3.3.3–3.3.7, 3.3.18–3.3.23) were solved in Lecture 11. Here we cover the diagonalization exercises.

Exercise 3.3.1Find eigenvalues, eigenvectors, and (if possible) P diagonalizing A

For each matrix, we give $c_A(x)$, the eigenvalues, and whether $A$ is diagonalizable (and $P$ when it is). Row reduction details are omitted for brevity.

Exercise 3.3.8Find P⁻¹AP, then compute Aⁿ
Exercise 3.3.9A, B diagonalizable but AB not; and D + A not diagonalizable
Exercise 3.3.10A is diagonalizable ⟺ Aᵀ is diagonalizable
Exercise 3.3.11Diagonalizable ⟹ so are Aⁿ, kA, p(A), U⁻¹AU, kI + A
Exercise 3.3.13If ±1 are the only eigenvalues of diagonalizable A, then A⁻¹ = A
Exercise 3.3.14If 0 and 1 are the only eigenvalues of diagonalizable A, then A² = A
Exercise 3.3.15If every eigenvalue λ ≥ 0, then A = B² for some B
Exercise 3.3.16If P⁻¹AP and P⁻¹BP are both diagonal, then AB = BA
Exercise 3.3.17Find all nilpotent diagonalizable matrices
Exercise 3.3.26Same characteristic polynomial, but one diagonalizable and one not

Let $A = \begin{bmatrix} 2 & 3 & -3 \\ 1 & 0 & -1 \\ 1 & 1 & -2 \end{bmatrix}$ and $B = \begin{bmatrix} 0 & 1 & 0 \\ 3 & 0 & 1 \\ 2 & 0 & 0 \end{bmatrix}$. Show $c_A(x) = c_B(x) = (x+1)^2(x-2)$, but $A$ is diagonalizable and $B$ is not.

Exercise 3.3.27Only diagonalizable matrix with a single eigenvalue is λI; is a given matrix diagonalizable?
Exercise 3.3.28Characterize diagonalizable A with A² − 3A + 2I = 0
Exercise 3.3.29Block-diagonal matrices: diagonalize A = diag(B, C)
Exercise 3.3.30Block-diagonal: c_A = c_B · c_C, and eigenvectors
Looking ahead

Diagonalization turned powers, populations, and even the web into simple eigenvalue arithmetic.

But not every matrix is diagonalizable, and some of the most useful transformations — rotations, oscillations — have complex eigenvalues. Ahead lies the study of matrices that resist diagonalization, orthogonal diagonalization of symmetric matrices (where the eigenvectors become perpendicular axes), and the singular value decomposition, the workhorse behind modern data compression and machine learning. Every one of them grows from the seed planted here: find the right coordinates, and a hard matrix becomes easy.

Lecture 12 — complete
MATH-120 · Shoaib Khan · LUMS · June 2026
← Lecture 11Lecture 13 →