Hunting the Higgs Boson

Binary classification of proton-proton collision events (Higgs-to-tau-tau signal vs. background) using the real ATLAS dataset released by CERN Open Data — the same data behind the 2014 HiggsML Challenge on Kaggle, not a synthetic simulation.

Python LightGBM PyTorch DuckDB FastAPI MIT

Source on GitHub · part of the Scientific Classification Lab (also includes exoplanet transit classification)

Dataset

ATLAS Higgs Challenge dataset, CERN Open Data Portal, record 328 (CC0) — 818,238 real simulated collision events, 30 physics-derived features, official train / public-test / private-test split reproduced via the original KaggleSet column (250,000 / 100,000 / 450,000 events) so results are directly comparable to the historical 2014 Kaggle competition leaderboard. Downloaded for this run on 2026-10-01 (atlas-higgs-challenge-2014-v2.csv.gz, 65,630,848 bytes).

The real-data problem: physically-defined missing values

11 columns use -999.0 as a sentinel — not random missingness. Dijet variables are undefined whenever an event has fewer than 2 reconstructed jets; verified empirically: 100% of -999.0 in those columns occurs exactly at PRI_jet_num ∈ {0,1}. Fixed with median imputation grouped by PRI_jet_num (train-only, no leakage) plus an explicit _missing flag per affected column.

Missing dijet values by jet count

Results — real run, official train/public/private split

Evaluated with AMS (Approximate Median Significance), the actual HiggsML Challenge metric — not accuracy or plain AUC.

ModelAMS (public test)AUCAccuracy
Decision Tree (baseline)2.93120.87560.7722
LightGBM3.55340.91060.7906
PyTorch MLP (Dropout+BatchNorm)3.57780.91020.7900
LightGBM, Optuna-tuned (40 trials)3.6414——

Held-out private test (450,000 events, never touched during model selection): untuned MLP AMS 3.5728 (0.14% from its public-test value); the Optuna-tuned LightGBM reaches AMS 3.6294 on the same held-out set (0.33% from its public-test value) — confirming the tuning gain is real, not overfit to the public split.

AMS comparison across models

Activation function comparison (custom Focal Loss)

ActivationAMS (public test)AUC
ReLU3.56990.9096
GELU3.53280.9094
Swish (SiLU)3.47630.9074
AMS by activation function

Honest caveat

The original 2014 Kaggle challenge's winning solutions (heavily tuned ensembles) reached AMS ≈ 3.8–3.9. This project's iteration (baseline → GBM → NN → Optuna-tuned GBM, no ensembling) reaching 3.63–3.64 on the official held-out set is an honest, unexaggerated result of that iteration process — not a leaderboard-matching claim.

Reproduce it

git clone https://github.com/Rxyxs/scientific-classification-lab.git
cd scientific-classification-lab/01-higgs-boson-particle-classification
python -m venv venv && venv\Scripts\activate
pip install -r requirements.txt
curl -L -o data/raw/atlas-higgs-challenge-2014-v2.csv.gz \
  https://opendata.cern.ch/api/files/1dd5c95f-9224-4f0b-9d98-8ba96601aa4a/atlas-higgs-challenge-2014-v2.csv.gz
gunzip data/raw/atlas-higgs-challenge-2014-v2.csv.gz
mv data/raw/atlas-higgs-challenge-2014-v2.csv data/raw/atlas-higgs.csv
pytest -q

Pablo Reyes — github.com/Rxyxs. Code: MIT.