PenFit / Model audit
Model governance

Testing PenFit for structural bias.

PenFit was deliberately stress-tested to see whether its scoring architecture favoured a retirement approach even when the underlying questionnaire evidence was neutral and symmetric.

2,000,000symmetric synthetic individual response profiles. These are not two million real people.

The profiles were generated across all five valid response levels specifically to test the model rather than represent a real population.

57.9all six on the fully neutral A/B profile
6 / 6strategies capable of ranking first
11.6–21.2%range of first-place frequency
82.3–90.3range of highest winning scores
13 / 13leave-one-characteristic-out checks passed

Symmetric stress-test results

The goal is not to force six identical outcomes. The check looks for structural dominance, an impossible-to-win strategy, unequal neutral starting points or excessive dependence on one characteristic.

Retirement approachRanked firstHighest winning alignmentAudit result
Flexible Income17.4%82.3PASS
Secure Income21.2%90.3PASS
Target Income14.2%85.8PASS
Flexible and Secure Income Mix19.8%86.0PASS
Flexible Income then Secure Income Later11.6%84.8PASS
Flexible Income with Secure Income Build-Up15.8%86.0PASS

What changed to control bias

The working calibration was changed only where a specific structural issue had been identified.

1
Balanced questionnaire structure

Each of the 12 retirement values appears twice and every answer uses the same +2 / +1 / 0 / −1 / −2 evidence scale.

2
Common neutral starting point

Strategy-specific starting advantage was removed. An all-A/B response gives every approach 57.9.

3
Normalised score responsiveness

A strategy is not rewarded simply because its natural score range moves more as answers change.

4
No artificial equalisation

No strategy bonus, equal-win target or automatic 100% top score is imposed.

Response-position control: PenFit Individual randomises which statement is presented as A or B and automatically reverses the evidence sign when the order changes. Presentation order therefore does not change the underlying model meaning.

Double-counting and sensitivity checks

The audit also challenged whether correlated characteristics such as retained capital, liquidity and reversibility were giving one economic property too many votes.

Correlated-characteristic test

Capping overlapping security/flexibility characteristic families still allowed all six strategies to rank first. First-place frequencies remained approximately 12.6%–21.8%, with maximum winning scores within about ±4.8 points of their mean.

13 leave-one-out tests

Each retirement characteristic was removed in turn. No single characteristic made any strategy impossible to select, so no single characteristic appears to be driving the result.

Audit conclusion: this is evidence of model neutrality under the conditions tested, not proof that any model is completely free of bias. Real-individual testing and ongoing monitoring remain part of validation and governance.