Information Geometry: When Probability Distributions Become Points
A first walk through statistical manifolds, Fisher information and the geometry hidden inside probability.
We usually imagine a point as something like . Information geometry asks us to entertain a stranger possibility:
What if each point is an entire probability distribution?
Consider the family of normal distributions
Every pair identifies one distribution, so the family itself forms a two-dimensional parameter space.
But what should distance mean?
Ordinary Euclidean distance between parameter vectors depends on how we choose the parameters. Information geometry instead derives a local metric from the statistical model itself: the Fisher information.
For parameters ,
This acts like a position-dependent inner product on tangent directions.
KL divergence is nearby—but not a distance
The Kullback–Leibler divergence
is asymmetric, so it is not a metric. Yet its local second-order behavior is intimately related to Fisher information.
That is the doorway into a beautiful subject where probability, differential geometry, statistics and machine learning meet.