Legal RAG Error Decomposition: Predicting Retrieval Failure, Reasoning Failure, and Hallucination Risk on Legal RAG Bench

Authors

  • Chloe Robinson Electrical Engineering and Computer Science, University of Queensland, Brisbane, QLD, Australia Author

DOI:

https://doi.org/10.66372/JGER.V4I2.1

Keywords:

Legal RAG, retrieval-augmented generation, hallucination, legal information retrieval, error decomposition, groundedness, benchmark evaluation, risk prediction

Abstract

Retrieval-augmented generation (RAG) systems are attractive for legal assistance because they connect generative language models to authoritative source material, yet their failures are not monolithic. A pipeline may miss the controlling passage, retrieve it but apply it incorrectly, or produce a fluent answer that is unsupported by the supplied context. This paper presents an empirical error-decomposition study on Legal RAG Bench, a reasoning-intensive benchmark containing 4,876 legal passages and 100 expert-written questions. The analysis covers the complete six-cell factorial design of three retrievers and two generators, corresponding to 600 question-by-configuration outcomes. For each configuration, we examine correctness, groundedness, retrieval accuracy, hallucination, retrieval error, reasoning error, and a composite RAG accuracy that counts only correct and grounded answers as complete successes. The released cell totals are also represented as a 600-row binary analysis table for cross-configuration risk modeling. Retrieval is the dominant source of performance variation: Kanon 2 Embedder reaches 86.0% retrieval accuracy and 91.5% average RAG accuracy, compared with 52.0% and 72.5% for Text Embedding 3 Large and 53.0% and 68.0% for Gemini Embedding 001. Generator choice changes groundedness and hallucination risk, but it does not offset weak retrieval. Cross-fitted prediction is strongest for retrieval failure (AUROC = 0.663) and retrieval error (AUROC = 0.661), whereas reasoning failure is rare and difficult to predict (AUROC = 0.439). The results support reporting component-attributed legal RAG errors rather than a single end-to-end accuracy score.

Author Biography

  • Chloe Robinson, Electrical Engineering and Computer Science, University of Queensland, Brisbane, QLD, Australia

     

     

     

Downloads

Published

2026-07-03

How to Cite

Legal RAG Error Decomposition: Predicting Retrieval Failure, Reasoning Failure, and Hallucination Risk on Legal RAG Bench. (2026). Journal of Global Engineering Review, 4(2), 1-15. https://doi.org/10.66372/JGER.V4I2.1