跳到主要内容

Preface-5

Five Factorizations of a Matrix

Here are the organizing principles of linear algebra. When our matrix has a special property, these factorizations will show it. Chapter after chapter, they express the key idea in a direct and useful way.

The usefulness increases as you go down the list. Orthogonal matrices are the winners in the end, because their columns are perpendicular unit vectors. That is perfection.

2 by 2 Orthogonal Matrix =[cos⁡θ−sin⁡θsin⁡θcos⁡θ]\begin{bmatrix} \cos\theta & -\sin\theta \\ \sin\theta & \cos\theta \end{bmatrix} = Rotation by Angleθ\theta

Here are the five factorizations from Chapters 1, 2, 4, 6, 7:

1 A=CRA = CR =RR combines independent columns in CC to give all columns of AA

2 A=LUA = LU = Lower triangularLL times Upper triangular UU

4 A=QRA = QR = Orthogonal matrixQQ times Upper triangular RR

6 S=QΛQTS = Q\Lambda Q^T = (OrthogonalQQ) (Eigenvalues inΛ\Lambda) (OrthogonalQTQ^T)

7 A=UΣVTA = U\Sigma V^T = (OrthogonalUU) (Singular values inΣ\Sigma) (OrthogonalVTV^T)

May I call your attention to the last one? It is the Singular Value Decomposition (SVD). It applies to every matrixAA. Those factors UU and VV have perpendicular columns—all of length one. Multiplying any vector by UU or VV leaves a vector of the same length—so computations don't blow up or down. And Σ\Sigma is a positive diagonal matrix of "singular values". If you learn about eigenvalues and eigenvectors in Chapter 6, please continue a few pages to singular values in Section 7.1.

Deep Learning

For a true picture of linear algebra, applications have to be included. Completeness is totally impossible. At this moment, the dominating direction of applied mathematics has one special requirement: It cannot be entirely linear!

One name for that direction is "deep learning". It is an extremely successful approach to a fundamental scientific problem: Learning from data. In many cases the data comes in a matrix. Our goal is to look inside the matrix for the connections between variables. Instead of solving matrix equations or differential equations that express known input-output rules, we have to find those rules. The success of deep learning is to build a function F(x,v)F(x,v) with inputs xx and vv of two kinds:

The vectors vv describes the features of the training data.

The matrices xx assign weights to those features.

The functionF(x,v)F(x,v) is close to the correct output for that training data vv.

When vv changes to unseen test data,F(x,v)F(x,v) stays close to correct.

This success comes partly from the form of the learning function FF, which allows it to include vast amounts of data. In the end, a linear functionFF would be totally inadequate. The favorite choice forFF is piecewise linear. This combines simplicity with generality.