Fraud Detection Techniques Lab

Two real-dataset approaches to credit card fraud, from the same lab: a supervised model-iteration benchmark, and an unsupervised-vs-supervised comparison that quantifies exactly what confirmed fraud labels are worth.

① Catching Credit Card Fraud (2023 dataset, supervised iteration)  ·  ② Autoencoder vs. Supervised XGBoost (ULB dataset, label-value quantified)

① Catching Credit Card Fraud

Fraud classification on 568,629 real 2023 credit card transactions — iterating from a Logistic Regression + SMOTE baseline to CatBoost/XGBoost to a PyTorch MLP trained with Focal Loss, validated for train/test distributional health, and calibrated by a business cost matrix instead of the default 0.5 threshold.

Python CatBoost/XGBoost PyTorch ONNX MIT

Source on GitHub · part of the Fraud Detection Techniques Lab

Dataset

Credit Card Fraud Detection Dataset 2023 (Kaggle, uploaded by Nelgiriyewithana) — 568,629 real transactions, PCA-anonymized components V1–V28 (same anonymization scheme as the classic ULB dataset) plus Amount. License: see the dataset page on Kaggle. Downloaded for this run on 2026-10-01.

This means there are no raw entity columns (card/device/IP) for dynamic per-entity aggregation, unlike IEEE-CIS's original multi-table schema — disclosed explicitly rather than pretended away. Adversarial validation (train-vs-test classifier) confirms a healthy split: AUC 0.5005, statistically indistinguishable.

Techniques

Results — real run, 20% held-out test set

ModelROC-AUCPR-AUCCost @ 0.5Cost @ best thresholdReduction
LogReg + SMOTE0.99420.9950257,950134,44047.9%
CatBoost1.00000.999993077017.2%
XGBoost1.00001.000047026044.7%
XGBoost, Optuna-tuned (30 trials)—0.999988—180—
MLP + Focal Loss1.00001.00001,4301,15019.6%

Business cost = 100 units per undetected fraud (false negative) + 10 units per legitimate transaction wrongly flagged (false positive).

Honest caveat

These near-perfect scores reflect this specific dataset's characteristics — artificially class-balanced (50/50, real-world card fraud is closer to 0.1–1%) and apparently close to linearly separable in its PCA space — not a claim that production fraud detection achieves ROC-AUC 1.0. This is a well-documented property of this exact Kaggle dataset, disclosed here rather than presented as a realistic production benchmark.

Reproduce it

git clone https://github.com/Rxyxs/fraud-detection-techniques-lab.git
cd fraud-detection-techniques-lab/01-credit-card-fraud-multilang
python -m venv venv && venv\Scripts\activate
pip install -r requirements.txt
kaggle datasets download -d nelgiriyewithana/credit-card-fraud-detection-dataset-2023 -p data/raw --unzip
pytest -q

② Deep Autoencoder vs. Supervised XGBoost

A quantified answer to a question every fraud/AML team eventually asks: "how much are we losing by not having confirmed labels yet?" Three unsupervised PyTorch architectures (Autoencoder, VAE, Deep SVDD) trained on only legitimate transactions — the realistic day-one scenario — benchmarked against a supervised XGBoost trained once fraud confirmations exist, on the real ULB/Worldline dataset (284,807 transactions, 492 confirmed frauds). A cost-sensitive threshold optimizer turns each model's raw score into an actual alerting decision.

PyTorch XGBoost SHAP MIT

Source on GitHub · part of the Fraud Detection Techniques Lab

Dataset

Credit Card Fraud Detection (Kaggle, ULB Machine Learning Group / Pozzolo et al., mirrored via OpenML dataset 1597) — 284,807 European cardholder transactions, September 2013, 492 confirmed frauds (0.172%). PCA components V1–V28; Time/Amount untransformed. Downloaded for this run on 2026-10-01 (exact row count verified: 284,807).

Headline comparison — real held-out test set

ModelROC-AUCPR-AUC
Autoencoder (unsupervised, zero labels used)0.9310.242
XGBoost (supervised, real labels)0.9650.834
Hybrid (XGBoost + autoencoder feature)0.9690.829

ROC-AUC alone suggests the autoencoder is "almost as good" (0.931 vs. 0.965). PR-AUC tells the real story: going from zero labels to real confirmed labels is a 3.4x jump in average precision (0.242 → 0.834).

Model comparison

AE vs. VAE vs. Deep SVDD — same protocol, different objective

ModelROC-AUCPR-AUC
Autoencoder (standard)0.9310.242
VAE0.9470.515
Deep SVDD0.9460.743

Deep SVDD's PR-AUC (0.743) is roughly 3x the standard autoencoder's (0.242) and approaches XGBoost's supervised 0.834 — without ever seeing a fraud label. Dropping reconstruction entirely and concentrating normal transactions into a hypersphere separates this dataset's fraud pattern substantially better than learning to reconstruct it.

Cost-sensitive threshold optimization — real financial impact

ModelAlertsTP/FP/FNTotal costReduction
Autoencoder28647/239/27$5,731.9932.4%
VAE14355/88/19$4,683.2644.8%
Deep SVDD14358/85/16$4,808.8843.3%
XGBoost14363/80/11$4,636.0845.4%

Cost matrix: fixed $5 per false-positive review; false-negative cost = the real dollar amount of that missed fraud. Baseline (no model): $8,483.36 total fraud on 74 confirmed test-set frauds.

Cost vs alert budget

Honest caveat — a negative result, reported as found

Adding the autoencoder's reconstruction error as an extra XGBoost feature moved ROC-AUC up marginally (0.965 → 0.969) but PR-AUC down slightly (0.834 → 0.829) — a wash, not an improvement. XGBoost, given the raw 30 PCA features directly, already extracts whatever signal the autoencoder's single reconstruction-error number summarizes. Reported as a negative result on purpose, rather than quietly dropped because it didn't confirm the hoped-for story.

Reproduce it

git clone https://github.com/Rxyxs/fraud-detection-techniques-lab.git
cd fraud-detection-techniques-lab/04-autoencoder-vs-supervised
python -m venv venv && venv\Scripts\activate
pip install -r requirements.txt
python data/download_dataset.py
pytest -q

Pablo Reyes — github.com/Rxyxs. Code: MIT.