Lognormal Distribution

Overview

Lognormal Distribution

The lognormal distribution is like a lop-sided normal distribution. It is unique in that it takes only positive values, as opposed to the normal distribution which ranges over the entire real line.

Definition

A variable is lognormally distributed if the logarithm of the variable is distributed normally. i.e. we have the following:
{% y = ln(x) %}
where
{% x = e^y %}
and y is distributed normally.
{% y \sim\ N(\mu_y, \sigma_y ^2) %}
The probability density of x is given by
{% f(x) = \sqrt{1/ 2 \pi x^2 \sigma_y ^2} e ^{-0.5[(y-\mu_y)/\sigma_y]^2} %}

Moment Formulas

The following relations hold for the moments of the lognormal distribution:
{% \mu_x = exp[\mu_y + \sigma_y^2/2] %}
{% \sigma_x = exp[2\mu_y + \sigma_y^2][exp(\sigma_y^2) - 1] %}
{% \mu_y = ln[\mu_x^2/ \sqrt{\mu_x^2 + \sigma_x^2}] %}
{% \sigma_y^2 = ln[1 + \sigma_x^2/\mu_x^2] %}

Lognormal vs Normal

Comparison of the distribution of a normal variable and a lognormal variable, both with mean =1 and variance = 1.

Simulating a Lognormal Variable

import numpy as np from scipy.stats import lognorm # 1. Define the parameters of the underlying normal distribution mu = 0.0 # Mean of the log-transformed data sigma = 0.5 # Standard deviation of the log-transformed data # 2. Map parameters to SciPy's notation shape_parameter = sigma # 's' in scipy scale_parameter = np.exp(mu) # 'scale' in scipy loc_parameter = 0 # Keep at 0 unless shifting the distribution # 3. Simulate random variables using .rvs() sample_size = 10000 simulated_data = lognorm.rvs( s=shape_parameter, loc=loc_parameter, scale=scale_parameter, size=sample_size, random_state=42 # Seed for reproducibility ) # Verification: Print the empirical mean vs theoretical expectation print(f"Empirical Mean: {np.mean(simulated_data):.4f}") print(f"Theoretical Mean: {np.exp(mu + (sigma**2)/2):.4f}")