Overview
Lognormal Distribution
The lognormal distribution is like a lop-sided normal distribution. It is unique in that it takes only positive values, as opposed to the normal distribution which ranges over the entire real line.Definition
A variable is lognormally distributed if the logarithm of the variable is distributed normally. i.e. we have the following:
{% y = ln(x) %}
where
{% x = e^y %}
and y is distributed normally.
{% y \sim\ N(\mu_y, \sigma_y ^2) %}
The probability density of x is given by
{% f(x) = \sqrt{1/ 2 \pi x^2 \sigma_y ^2} e ^{-0.5[(y-\mu_y)/\sigma_y]^2} %}
Moment Formulas
The following relations hold for the moments of the lognormal distribution:
{% \mu_x = exp[\mu_y + \sigma_y^2/2] %}
{% \sigma_x = exp[2\mu_y + \sigma_y^2][exp(\sigma_y^2) - 1] %}
{% \mu_y = ln[\mu_x^2/ \sqrt{\mu_x^2 + \sigma_x^2}] %}
{% \sigma_y^2 = ln[1 + \sigma_x^2/\mu_x^2] %}
Lognormal vs Normal
Comparison of the distribution of a normal variable and a lognormal variable, both with mean =1 and variance = 1.Simulating a Lognormal Variable
import numpy as np
from scipy.stats import lognorm
# 1. Define the parameters of the underlying normal distribution
mu = 0.0 # Mean of the log-transformed data
sigma = 0.5 # Standard deviation of the log-transformed data
# 2. Map parameters to SciPy's notation
shape_parameter = sigma # 's' in scipy
scale_parameter = np.exp(mu) # 'scale' in scipy
loc_parameter = 0 # Keep at 0 unless shifting the distribution
# 3. Simulate random variables using .rvs()
sample_size = 10000
simulated_data = lognorm.rvs(
s=shape_parameter,
loc=loc_parameter,
scale=scale_parameter,
size=sample_size,
random_state=42 # Seed for reproducibility
)
# Verification: Print the empirical mean vs theoretical expectation
print(f"Empirical Mean: {np.mean(simulated_data):.4f}")
print(f"Theoretical Mean: {np.exp(mu + (sigma**2)/2):.4f}")