Credit Risk Scoring Lab

A closed-loop MLOps platform for credit risk, not a notebook of models: eleven from-scratch statistical techniques for the scoring question itself, wired into fifteen more that detect drift, retrain, deploy, canary-test, cut over, monitor realized performance, and audit the whole system — honestly, including the state where retraining fired and nobody closed the loop yet.

This page shows the project's results. The methodology, the design decisions and the limitations are documented in the repository's README.

01 · Polyglot scorecard (R + Python + C)

Scoring engine benchmark
*Throughput over 1,000,000 rows, log scale. The compiled engine scores 270.6M

02 · Bidirectional R↔Python interop

Combined credit and market risk heatmap
*The centrepiece: expected loss per credit-risk band stressed against three

03 · Lifetime PD from survival analysis

Estimated hazard vs. the true generating hazard
*The estimated hazard against the one the simulator actually used. The seasoning
PD term structure by risk band
*Left: cumulative PD by horizon for each risk band, with the 12-month

04 · Bayesian hierarchical scorecard

Shrinkage by segment
*Left: each segment's estimated effect with and without pooling, against how

05 · Monotonic constraints + conformal decisioning

Monotonicity audit
*The unconstrained model violates monotonicity for up to 97.6% of applicants on
Conformal coverage by class
*Why the class-conditional (Mondrian) variant matters. Left: coverage tracks the

06 · Optimal-binning scorecard

WOE by binning method
*The same three variables, cut three ways. The DP produces a clean monotone WOE
PSI and CSI by vintage
*Left: PSI stays under 0.025 for eighteen stable vintages — no false alarms —

07 · Fair lending bias audit

Proxy detection
*Left: each feature's group signal against its risk signal. Everything sits

08 · Differentially private scoring

Privacy versus utility
*Left: what privacy costs — AUC against ε, averaged over ten independent runs
Leak detectability
*The same data asked properly: not "how much leakage" but "can it be told apart

09 · Reject inference & selection bias

Recovery of rho
*Grey is the truth. Without an exclusion restriction (red) the model invents

10 · Through-the-cycle vs. point-in-time PD

PIT vs TTC capital
*Grey: RWA density using the through-the-cycle PD — flat by construction.

11 · Federated credit scoring

Calibration by bank
*Left: each bank's own model, trained only on its own customers, applied to