Matrix Theory: Subspaces, Rank, Inner Products and Gram–Schmidt, Special Matrices, Similarity, Definiteness and the Singular Value Decomposition

Section 2 of the GATE Statistics paper is Matrix Theory, and it reads like a first course in linear algebra over both ℝⁿ and ℂⁿ. The shared Linear Algebra and Vector Spaces chapters teach rank and consistency, eigenvalue shortcuts, Cayley–Hamilton, LU, quadratic forms and singular values; this chapter adds what the Statistics list names beyond them: subspaces, span, basis and dimension, row and column spaces and the reduced row echelon form; inner products in ℝⁿ and ℂⁿ and Gram–Schmidt orthonormalisation; the eigenvalues of symmetric, skew-symmetric, Hermitian, skew-Hermitian, orthogonal and unitary matrices; change of basis, equivalence and similarity; the properties of positive definite and semidefinite matrices; and the singular value decomposition. Every covariance matrix in the later sections is a positive semidefinite symmetric matrix, and every regression is a projection — this chapter is where both are prepared.

1. Subspaces, span, basis, dimension, row and column spaces, and rank

A subspace of ℝⁿ or ℂⁿ is a non-empty set closed under addition and scalar multiplication (so it contains 0). The span of vectors is the set of their linear combinations, the smallest subspace containing them. Vectors are linearly independent if the only combination giving 0 is the trivial one; a basis is an independent spanning set, and all bases have the same size, the dimension. For subspaces U, W: dim(U + W) = dim U + dim W − dim(U ∩ W) — so in ℝ⁵ a 3-dimensional and a 4-dimensional subspace whose sum is ℝ⁵ meet in a 2-dimensional subspace.

The row space and column space of an m × n matrix A have the same dimension, the rank; the null space {x : Ax = 0} has dimension n − rank, the nullity. Row operations preserve the row space and the null space (not the column space), and reduce A to its unique reduced row echelon form (RREF): leading 1s, zeros above and below them, each leading 1 to the right of the one above. The pivot columns of A — the original columns, not those of the RREF — form a basis of the column space. The solution set of a consistent Ax = b is a particular solution plus the null space.

  • Trace: tr(A + B) = tr A + tr B and tr(AB) = tr(BA) even when AB ≠ BA; tr A = sum of eigenvalues.
  • Determinant: det(AB) = det A det B, det Aᵀ = det A, det(cA) = cⁿ det A for n × n; det A = product of eigenvalues. det(A + B) ≠ det A + det B in general.
  • Inverse: exists iff det A ≠ 0 iff rank n; A⁻¹ = adj A/det A; (AB)⁻¹ = B⁻¹A⁻¹; (Aᵀ)⁻¹ = (A⁻¹)ᵀ.
  • rank(AB) ≤ min(rank A, rank B), rank(AᵀA) = rank A, rank(A + B) ≤ rank A + rank B.

2. Inner products and Gram–Schmidt orthonormalisation

The standard inner product is ⟨x, y⟩ = Σxᵢyᵢ on ℝⁿ and ⟨x, y⟩ = Σxᵢȳᵢ = y*x on ℂⁿ, with the conjugate so that ⟨x, x⟩ = Σ|xᵢ|² ≥ 0. It is linear in the first argument, conjugate-symmetric (⟨y, x⟩ = conj⟨x, y⟩), and positive definite. Cauchy–Schwarz: |⟨x, y⟩| ≤ ‖x‖‖y‖, with equality iff x and y are dependent — in statistics, the reason a correlation lies in [−1, 1]. The projection of v on a unit vector e is ⟨v, e⟩e.

Gram–Schmidt turns independent v₁, …, v_k into an orthonormal basis of their span: w₁ = v₁, wⱼ = vⱼ − Σi<j ⟨vⱼ, eᵢ⟩eᵢ, eⱼ = wⱼ/‖wⱼ‖. For v₁ = (1, 1, 0), v₂ = (1, 0, 1): e₁ = (1, 1, 0)/√2; ⟨v₂, e₁⟩ = 1/√2, so w₂ = (1, 0, 1) − ½(1, 1, 0) = (½, −½, 1), with ‖w₂‖² = 3/2, and e₂ = (1, −1, 2)/√6. Writing the result as A = QR, with Q orthonormal columns and R upper triangular, is the QR decomposition.

3. Eigenvalues, Cayley–Hamilton and the special matrices

The eigenvalues are the roots of the characteristic polynomial det(λI − A), of degree n; their sum is the trace and their product the determinant. Cayley–Hamilton: A satisfies its own characteristic polynomial. For A = [[2, 1], [1, 2]], λ² − 4λ + 3 = 0, so A² = 4A − 3I and A⁻¹ = (4I − A)/3; since the eigenvalues are 3 and 1, tr A⁵ = 3⁵ + 1 = 244.

Special matrices and their eigenvalues
MatrixDefinitionEigenvalues
Real symmetricAᵀ = Areal; orthogonal eigenvectors
HermitianA* = Areal; orthogonal eigenvectors
Real skew-symmetricAᵀ = −Apurely imaginary or 0
Skew-HermitianA* = −Apurely imaginary or 0
OrthogonalAᵀA = I|λ| = 1; det = ±1
UnitaryA*A = I|λ| = 1
IdempotentA² = A0 or 1; trace = rank

The proofs are one line each: if Ax = λx with x ≠ 0 and A Hermitian, then λ‖x‖² = xAx is real because (xAx)* = xAx = x*Ax. If A is unitary, ‖Ax‖ = ‖x‖ forces |λ| = 1. A real orthogonal matrix of odd order always has a real eigenvalue, which must be ±1, because complex eigenvalues come in conjugate pairs.

4. Change of basis, equivalence, similarity and diagonalisability

If the columns of an invertible P are a new basis, a vector with old coordinates x has new coordinates P⁻¹x, and the linear map with matrix A in the old basis has matrix P⁻¹AP in the new one — the change of basis matrix at work. A and B are similar if B = P⁻¹AP; they then share the characteristic polynomial, eigenvalues (with algebraic and geometric multiplicities), trace, determinant and rank — but not their eigenvectors, which are transformed by P⁻¹. A and B are equivalent if B = PAQ for invertible P, Q; equivalence preserves only the rank, and two matrices of the same size are equivalent iff they have the same rank.

A is diagonalisable — similar to a diagonal matrix — iff it has n independent eigenvectors, iff every eigenvalue has geometric multiplicity equal to its algebraic multiplicity. Distinct eigenvalues suffice; Hermitian (and, more generally, normal: AA* = A*A) matrices are unitarily diagonalisable, A = UDU*. Then Aᵏ = PDᵏP⁻¹. For A = [[4, 1], [2, 3]] with eigenvalues 5 and 2, A³ = [[86, 39], [78, 47]], whose trace 133 = 5³ + 2³ checks the arithmetic.

⚠️ Same invariants, still not similar
[[1, 1], [0, 1]] and I have the same characteristic polynomial (λ − 1)², trace, determinant and rank, but they are not similar: P⁻¹IP = I for every P, so the only matrix similar to I is I itself. Matching invariants is necessary for similarity, never sufficient. The two ARE equivalent, since both have rank 2.

5. Positive definite and semidefinite matrices, quadratic forms and the SVD

A Hermitian (real symmetric) A is positive definite (PD) if x*Ax > 0 for all x ≠ 0 and positive semidefinite (PSD) if x*Ax ≥ 0. Equivalent tests for PD: all eigenvalues positive; all leading principal minors positive (Sylvester); A = BᵀB for an invertible B (Cholesky, B triangular). For PSD: all eigenvalues ≥ 0; all principal minors ≥ 0; A = BᵀB for some B. Consequences: a PD matrix has positive diagonal entries, positive determinant and a PD inverse; sums of PD matrices are PD; |aᵢⱼ|² ≤ aᵢᵢaⱼⱼ. Every covariance matrix is PSD, because Var(aᵀX) = aᵀΣa ≥ 0.

⚠️ Leading minors are enough only for definiteness
diag(0, −1) has leading principal minors 0 and 0, both non-negative, yet it is not PSD (e₂ᵀAe₂ = −1). Semidefiniteness needs every principal minor ≥ 0 — here the 1 × 1 minor −1 fails. Non-negative leading minors do not imply PSD.

A quadratic form q(x) = xᵀAx is written with symmetric A: 2x² + 6xy + 5y² has A = [[2, 3], [3, 5]], det 1 > 0 and a₁₁ > 0, so it is positive definite. The singular value decomposition A = UΣV* holds for every m × n matrix, with U, V unitary and Σ diagonal with σ₁ ≥ σ₂ ≥ … ≥ 0, the singular values — the square roots of the eigenvalues of A*A. The number of non-zero σᵢ is the rank, σ₁ = max ‖Ax‖/‖x‖, and Σσᵢ² = Σ|aᵢⱼ|². For A = [[3, 0], [4, 5]], AᵀA = [[25, 20], [20, 25]] has eigenvalues 45 and 5, so σ = √45 ≈ 6.71 and √5 ≈ 2.24; their product 15 = |det A|.

Key takeaways

  • Row rank = column rank; nullity = n − rank; dim(U + W) = dim U + dim W − dim(U ∩ W). Pivot columns of the original matrix are a basis of its column space.
  • On ℂⁿ, ⟨x, y⟩ = y*x; Gram–Schmidt subtracts projections on the earlier orthonormal vectors and normalises.
  • Hermitian/symmetric → real eigenvalues; skew-Hermitian → imaginary or 0; unitary/orthogonal → |λ| = 1; idempotent → 0 or 1.
  • Similar matrices share eigenvalues, trace, determinant and rank but not eigenvectors; equivalent matrices share only the rank. [[1, 1], [0, 1]] is not similar to I.
  • PD ⇔ eigenvalues > 0 ⇔ leading minors > 0; PSD needs all principal minors ≥ 0. Covariance matrices are PSD.
  • Singular values are √eig(A*A); their count is the rank, the largest is the 2-norm, and their product is |det A| for square A.

Practice questions (14)

Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.

  1. The dimension of the null space of A = [[1, 2, 1, 0], [2, 4, 3, 1], [3, 6, 4, 1]] is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 2

    R₂ − 2R₁ = (0, 0, 1, 1) and R₃ = R₁ + R₂, so the rank is 2. A has 4 columns, so the nullity is 4 − 2 = 2. Using the 3 rows instead of the 4 columns gives 1 — rank–nullity counts columns.
  2. U and W are subspaces of ℝ⁵ with dim U = 3, dim W = 4 and U + W = ℝ⁵. The dimension of U ∩ W is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 2

    dim(U ∩ W) = dim U + dim W − dim(U + W) = 3 + 4 − 5 = 2. The intersection cannot be {0} here: two subspaces whose dimensions add to more than 5 must share a non-zero vector in ℝ⁵.
  3. Gram–Schmidt is applied to v₁ = (1, 1, 0), v₂ = (1, 0, 1). The squared length ‖w₂‖² of w₂ = v₂ − projv₁v₂ (before normalising), correct to one decimal place, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 1.5

    projv₁v₂ = (⟨v₂, v₁⟩/⟨v₁, v₁⟩)v₁ = (1/2)(1, 1, 0). So w₂ = (1/2, −1/2, 1) and ‖w₂‖² = 1/4 + 1/4 + 1 = 1.5. Check: ⟨w₂, v₁⟩ = 1/2 − 1/2 = 0. Subtracting ⟨v₂, v₁⟩v₁ without dividing by ‖v₁‖² = 2 gives (0, −1, 1), which is not orthogonal to v₁.
  4. Which statements about eigenvalues are true?

    1. Every eigenvalue of a Hermitian matrix is real
    2. Every eigenvalue of a unitary matrix has modulus 1
    3. Every eigenvalue of a skew-Hermitian matrix is real
    4. Every 3 × 3 real orthogonal matrix has 1 or −1 as an eigenvalue
    Show answer

    Answer: A — Every eigenvalue of a Hermitian matrix is real; B — Every eigenvalue of a unitary matrix has modulus 1; D — Every 3 × 3 real orthogonal matrix has 1 or −1 as an eigenvalue

    (A) λ‖x‖² = x*Ax is real. (B) ‖Ax‖ = ‖x‖ gives |λ| = 1. (C) False: iA is Hermitian when A is skew-Hermitian, so A’s eigenvalues are purely imaginary or 0 — e.g. [[i]] has eigenvalue i. (D) The characteristic polynomial is a real cubic, so it has a real root, and a real eigenvalue of an orthogonal matrix has modulus 1.
  5. For A = [[2, 1], [1, 2]], the trace of A⁵ is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 244

    Trace 4 and determinant 3 give eigenvalues 3 and 1. A⁵ has eigenvalues 3⁵ = 243 and 1, so tr A⁵ = 244. Raising the trace itself, 4⁵ = 1024, treats the trace as an eigenvalue.
  6. If B = P⁻¹AP for an invertible P, which of the following need NOT be the same for A and B?

    1. The eigenvectors
    2. The characteristic polynomial
    3. The trace
    4. The rank
    Show answer

    Answer: A — The eigenvectors

    det(λI − P⁻¹AP) = det(P⁻¹(λI − A)P) = det(λI − A), so the characteristic polynomial, and with it trace and determinant, agree; rank is unchanged by invertible factors. But if Ax = λx then B(P⁻¹x) = λ(P⁻¹x): the eigenvectors of B are P⁻¹ times those of A.
  7. Let A = [[1, 1], [0, 1]]. Which statements are true?

    1. The characteristic polynomial of A is (λ − 1)²
    2. A is diagonalisable
    3. A is similar to the identity matrix
    4. A is equivalent to the identity matrix (B = PAQ for invertible P, Q)
    Show answer

    Answer: A — The characteristic polynomial of A is (λ − 1)²; D — A is equivalent to the identity matrix (B = PAQ for invertible P, Q)

    (A) A is triangular with diagonal 1, 1. (B) A − I = [[0, 1], [0, 0]] has a one-dimensional null space: geometric multiplicity 1 < 2, not diagonalisable. (C) P⁻¹IP = I for every P, so only I is similar to I. (D) Equivalence needs only equal rank, and both have rank 2 (take P = A⁻¹, Q = I).
  8. For A = [[4, 1], [2, 3]], the (1, 2) entry of A³ is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 39

    A² = [[18, 7], [14, 11]] and A³ = A²A = [[72 + 14, 18 + 21], [56 + 22, 14 + 33]] = [[86, 39], [78, 47]]. Check: tr A³ = 133 = 5³ + 2³ (eigenvalues 5 and 2), and det A³ = 86 × 47 − 39 × 78 = 1000 = 10³. Cubing the entry 1 gives 1, which ignores that matrix powers mix entries.
  9. Let M = [[1, 2], [2, 4]]. Which statements are true?

    1. M is positive semidefinite
    2. M is positive definite
    3. M = vvᵀ for some vector v
    4. xᵀMx = (x₁ + 2x₂)² for every x
    Show answer

    Answer: A — M is positive semidefinite; C — M = vvᵀ for some vector v; D — xᵀMx = (x₁ + 2x₂)² for every x

    Trace 5 and determinant 0 give eigenvalues 5 and 0, so M is PSD but not PD (x = (2, −1) gives xᵀMx = 0). With v = (1, 2), vvᵀ = M, and xᵀMx = (vᵀx)² = (x₁ + 2x₂)². A zero determinant is exactly what separates semidefinite from definite here.
  10. All leading principal minors of a real symmetric matrix are non-negative. Then the matrix:

    1. is necessarily positive semidefinite
    2. is necessarily positive definite
    3. need not be positive semidefinite
    4. is necessarily singular
    Show answer

    Answer: C — need not be positive semidefinite

    diag(0, −1) has leading minors 0 and 0 but the quadratic form takes the value −1 at (0, 1), so it is not PSD. Sylvester’s criterion with strict inequalities characterises PD; for PSD every principal minor, not just the leading ones, must be ≥ 0. And diag(1, 1) shows the matrix need not be singular.
  11. The largest singular value of A = [[3, 0], [4, 5]], correct to two decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 6.71

    AᵀA = [[9 + 16, 20], [20, 25]] = [[25, 20], [20, 25]], with eigenvalues 25 ± 20 = 45 and 5. σ₁ = √45 = 6.708 → 6.71. The eigenvalues of A itself are 3 and 5 (it is triangular), and giving 5 confuses eigenvalues with singular values; check: σ₁σ₂ = √225 = 15 = |det A|.
  12. The determinant of the symmetric matrix of the quadratic form q(x, y) = 2x² + 6xy + 5y² is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 1

    The off-diagonal entries are half the cross coefficient: A = [[2, 3], [3, 5]], det = 10 − 9 = 1. With a₁₁ = 2 > 0 the form is positive definite. Putting 6 off the diagonal gives 10 − 36 = −26 and the wrong verdict "indefinite".
  13. For n × n matrices A and B, which identity always holds?

    1. tr(AB) = tr(BA)
    2. AB = BA
    3. det(A + B) = det A + det B (the determinant is additive)
    4. rank(A + B) = rank A + rank B (rank is additive)
    Show answer

    Answer: A — tr(AB) = tr(BA)

    tr(AB) = Σᵢ Σⱼ aᵢⱼbⱼᵢ = tr(BA) for all square A, B, even when AB ≠ BA. The determinant is multiplicative, not additive (A = B = I in 2 × 2 gives 4 ≠ 2), and rank is only subadditive: rank(A + B) ≤ rank A + rank B.
  14. Let J be the 5 × 5 matrix with every entry 1. The determinant of J + 2I is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 112

    J = 11ᵀ has rank 1, so its eigenvalues are 5 (eigenvector 1) and 0 four times. J + 2I has eigenvalues 7 and 2, 2, 2, 2, so det = 7 × 2⁴ = 112. Adding 2 only to the non-zero eigenvalue gives 7, forgetting that the shift applies to all five.