On one worked case, a merchant categoriser overstates serviceability threefoldSee the worked case
Method study

The whole pipeline, on public data, with the failures left in.

No client, no confidential book and nothing to take on trust. A public credit dataset, run end to end: baseline, engineered signals, validation, measured lift, the segments where it breaks and the specifications that did not work. Every number here is reproducible by anyone who wants to check it.

Discuss running it on your book →
Why this exists

A vendor claim you can audit.

Most vendor evidence cannot be checked: the data is private, the baseline is unstated and the failures are absent. This is the opposite of that, on purpose.

01

What it proves

That the method is real and that it is applied with discipline: a stated baseline, out of time validation, results reported per segment, and a record of what was tried and discarded.

02

What it does not prove

That you will see this lift. Public data is not your book. The value on your portfolio depends on your baseline, your population and your data quality, and nobody can tell you the number in advance.

03

How to check it

The dataset is public and the approach is described in enough detail to repeat. If your modelling team disagrees with a choice made here, that is a conversation worth having before any engagement.

Set up

The question, the data and the baseline.

Stated before any modelling, and not changed afterwards.

01

The question

[STATE THE DECISION: e.g. which applicants fall 90 days behind within 12 months, and what a lender would do differently knowing it]

02

The data

[DATASET NAME AND SOURCE], [N] cases covering [PERIOD]. Chosen because [WHY THIS ONE: transaction level, outcomes matured, representative of SME or consumer lending].

03

The population

[INCLUSIONS AND EXCLUSIONS, AND WHY]. Cases removed: [N] for [REASON].

04

The baseline

[WHAT IT IS COMPARED AGAINST: a published scorecard, a bureau style model fitted on the same data, or a simple logistic model on standard variables]. Baseline performance: [FIGURE].

05

The split

Out of time holdout: [PERIOD]. Never used in fitting or selection. [N] cases.

06

Success measure

Agreed in advance: [WHAT WOULD COUNT AS WORTH HAVING, IN BOTH STATISTICAL AND COMMERCIAL TERMS].

What was engineered

The signals, and what each one is trying to capture.

Definitions in full, so a reviewer can disagree with a construction rather than with a label.

SignalHow it is builtWhat it is trying to capture
01

[SIGNAL NAME]

[FAMILY: cashflow, conduct, concentration, trajectory]

[field_name]

[DEFINITION: source fields, window, transformation, how missing data is handled]

[THE BEHAVIOUR IT IS MEANT TO REFLECT]

02

[SIGNAL NAME]

[FAMILY]

[field_name]

[DEFINITION]

[THE BEHAVIOUR IT IS MEANT TO REFLECT]

03

[SIGNAL NAME]

[FAMILY]

[field_name]

[DEFINITION]

[THE BEHAVIOUR IT IS MEANT TO REFLECT]

[N] candidates engineered, [N] retained. The rest are listed in the experiment record below.

Result

What it added over the baseline.

Out of time, reported per cohort, with the commercial translation stated separately from the statistical one.

VALIDATION / [TARGET]out of time, [N] cohorts
BASELINE[FIGURE][what it is]
WITH ENGINEERED SIGNALS[FIGURE]same population, same split
STABILITY[PSI OR EQUIVALENT]across cohorts
WHAT IT WOULD MEAN AT A LENDER

[COMMERCIAL TRANSLATION: at the same expected loss, approvals move by X; or at the same approval rate, expected loss moves by Y. State the assumptions used to convert.]

Conclusion[WHAT YOU CONCLUDE]
Not tested[WHAT THIS DOES NOT SHOW]

Public dataset. Results on a lender’s own book will differ, in both directions.

Where it breaks

The segments the model is worst on.

Published because a model that passes overall and fails on one segment is the normal case, and because this is the first thing a validation team looks for.

01

[SEGMENT]

[PERFORMANCE HERE], against [OVERALL]. Likely cause: [WHY].

02

[SEGMENT]

[PERFORMANCE HERE]. Likely cause: [WHY].

03

[SEGMENT]

[PERFORMANCE HERE]. Likely cause: [WHY].

04

What would fix it

[WHAT DATA OR METHOD WOULD BE NEEDED, AND WHETHER IT IS WORTH IT].

Experiment record

Everything that was tried, including what failed.

The specifications that did not work are more informative than the one that did, and they are the part vendors never publish.

EXPERIMENT LOG / [SIGNAL SET][N] runs completed
RUNWHAT WAS TESTEDRESULTVERDICT
R-01baseline, held constant
[FIGURE][FIGURE]Baseline
[R-XX][WHAT CHANGED]
[FIGURE][FIGURE]Selected
[R-XX][WHAT CHANGED]
[FIGURE][FIGURE]Rejected
[R-XX][WHAT CHANGED]
[FIGURE][FIGURE]Not carried

[ONE LINE ON WHY THE SELECTED RUN WAS PREFERRED OVER THE HIGHEST SCORING ONE, IF THEY DIFFER.]

Reading it honestly

What a lender should take from this, and what they should not.

Stated plainly, because the temptation to over read a vendor study runs in both directions.

01

Take this

[WHAT TRANSFERS: the method, the discipline, the specific signal constructions that are not dataset dependent].

02

Do not take this

[WHAT DOES NOT TRANSFER: the size of the lift, the ranking of signals, anything driven by this dataset’s population].

03

What would settle it on your book

[THE SMALLEST PIECE OF WORK THAT WOULD GIVE A REAL ANSWER: data required, elapsed time, what you would learn].

Start with one change

What would you change if you could measure it?

Name one change you have argued about. We will tell you whether your book can measure it live, or from history.

Discuss a focused pilot →