4 of 5Real lending and bank data from four open research datasets
What a readable model costs
problems where a regression a lender can read came within 0.01 of the most accurate model
- What was observed
- Five prediction problems on real data: default on personal loans, take up of an offer, an account's receipts next quarter, default on card statements, and default from card transactions.
- What a generic approach says
- Assume the most complex model is the most accurate, or that a readable one costs nothing.
- What the engine read
- For each problem, the most accurate model was built beside the closest regression a lender can read. The readable model came within 0.002 for loan default (0.701 against 0.703), 0.005 for offer take up (0.923 against 0.928), 0.005 for receipts (0.780 against 0.785, with bends for the curves) and 0.008 for card statement default (0.7898 against 0.7978 on that competition's measure). From card transactions, boosting kept a real lead: 0.750 against 0.777.
- What it meant for the decision
- The cost of a readable model is measured before the model is chosen, not assumed. Often it is close to nothing; where it is not, the lender sees the price and decides.