It’s pretty fundamental in a variety of multivariate statistical methods. If the rows of X are multi variate observations, then XX’ is the Gram matrix (of dot products). This can be used in clustering and regression. If…
Statistician here. I agree that some of this stuff seems counterintuitive on the surface. Once you make the connection with high-dimensional Gaussians, it can become more "obvious": if Z is standard n-dimensional…
It’s differential everywhere except at x=0. At x=0 it actually has a subdifferential—think of it as the set of slopes of lines that are tangent at that point.
It’s pretty fundamental in a variety of multivariate statistical methods. If the rows of X are multi variate observations, then XX’ is the Gram matrix (of dot products). This can be used in clustering and regression. If…
Statistician here. I agree that some of this stuff seems counterintuitive on the surface. Once you make the connection with high-dimensional Gaussians, it can become more "obvious": if Z is standard n-dimensional…
It’s differential everywhere except at x=0. At x=0 it actually has a subdifferential—think of it as the set of slopes of lines that are tangent at that point.