Preface-5
Five Factorizations of a Matrix
Here are the organizing principles of linear algebra. When our matrix has a special property, these factorizations will show it. Chapter after chapter, they express the key idea in a direct and useful way.
The usefulness increases as you go down the list. Orthogonal matrices are the winners in the end, because their columns are perpendicular unit vectors. That is perfection.
2 by 2 Orthogonal Matrix = = Rotation by Angle
Here are the five factorizations from Chapters 1, 2, 4, 6, 7:
1 = combines independent columns in to give all columns of
2 = Lower triangular times Upper triangular
4 = Orthogonal matrix times Upper triangular
6 = (Orthogonal) (Eigenvalues in) (Orthogonal)
7 = (Orthogonal) (Singular values in) (Orthogonal)
May I call your attention to the last one? It is the Singular Value Decomposition (SVD). It applies to every matrix. Those factors and have perpendicular columns—all of length one. Multiplying any vector by or leaves a vector of the same length—so computations don't blow up or down. And is a positive diagonal matrix of "singular values". If you learn about eigenvalues and eigenvectors in Chapter 6, please continue a few pages to singular values in Section 7.1.
Deep Learning
For a true picture of linear algebra, applications have to be included. Completeness is totally impossible. At this moment, the dominating direction of applied mathematics has one special requirement: It cannot be entirely linear!
One name for that direction is "deep learning". It is an extremely successful approach to a fundamental scientific problem: Learning from data. In many cases the data comes in a matrix. Our goal is to look inside the matrix for the connections between variables. Instead of solving matrix equations or differential equations that express known input-output rules, we have to find those rules. The success of deep learning is to build a function with inputs and of two kinds:
The vectors describes the features of the training data.
The matrices assign weights to those features.
The function is close to the correct output for that training data .
When changes to unseen test data, stays close to correct.
This success comes partly from the form of the learning function , which allows it to include vast amounts of data. In the end, a linear function would be totally inadequate. The favorite choice for is piecewise linear. This combines simplicity with generality.