Single Probablity of Default
Overview
The simplest way to estimate the probability of default from a set of historical loan
records is to mark the loans that have defaulted with a 1 and a 0 on loans that have not defaulted,
and then to take the average of this column.
Note, once a loan has defaulted, the records of the defaulted
loan after the default time should be removed. That is, there should be only one record per account
that has 1 in the default column. The code below includes a filter function which filters out all
records
The code given here is given in the script named "estimate" in the desktop.
Script
The estimate function takes an array of dictionaries with the following properties (columns)
- account - is an id that identifies the account
- date
- default
- is 1 if the loan has defaulted at the given date, and 0 otherwise
'''
Expects a list of dictionaries with a column named "default" which is set to 1 if the loan has defaulted
and a 0 otherwise
'''
def estimate(data, periods=1):
fdata = filter(data)
# Calculate the average of the 'default' column
average = sum(item["default"] for item in fdata) / len(fdata)
prob_no_default =(1-average)**periods
return 1 - prob_no_default
'''
filter the dataset to remove records of any account after it has already defaulted
'''
def filter(data):
results = []
sorted_data = sorted(data, key=lambda x: (x['account'], x['date']))
last = None
for item in sorted_data:
if last != None:
if last['account'] == item['account'] and last['default'] == 1:
pass
else: results.append(item)
else:
results.append(item)
last = item
return results