⛑️ RAG Seguridad Minera Chile

NLP + RAG system for incident classification and regulatory Q&A on Chilean mining safety law (DS 132)

This page shows the project's results. The methodology, the design decisions and the limitations are documented in the repository's README.

7. Validated results

What Cross-Encoder re-ranking changes
Honest note on Context Relevance@4 staying flat: almost every query in the evaluation set has exactly one truly relevant article, and that article was already inside the hybrid retriever's top-4 *before* re-ranking — so precision@4 is capped at 1/4 = 0.250 in both conditions by construction of the dataset, not because re-ranking did nothing. What re-ranking changed is *where* that relevant article lands within the top-4, which is exactly what MRR and NDCG measure, and both improved.
Severity classifier: confusion matrix and per-class scores
Honest note on the classifier: the headline accuracy hides two things worth stating.
Corpus composition: severity distribution and incident types
The severity imbalance in the generated corpus (GRAVE 71, FATAL 40, LEVE 33 — a 2.2x spread between the extremes) is the most likely reason LEVE is the weakest class, even with class_weight="balanced" in the logistic regression.
Citation faithfulness per query
Citation faithfulness is 1.0 on 25 of 27 queries. The metric is purpose-built and narrow on purpose: it checks that every article the answer cites is one that was actually retrieved, which catches a *fabricated* citation but not a wrong reading of a real one. It needs no LLM judge, which is deliberate — the whole pipeline runs without an external LLM, and an evaluation that required one would undercut that.