>

>

Semantic Entropy as a Scalable Hallucination Detection Signal in Large Language Models: Theoretical Grounding, Calibration, and Downstream Task Benchmark

Semantic Entropy as a Scalable Hallucination Detection Signal in Large Language Models: Theoretical Grounding, Calibration, and Downstream Task Benchmark

Publisher : PJPCR
Author(s)
James T. Wu; Anika M. Johansson; Dele K. Ogunyemi
Abstract

This study investigates semantic entropy as a computationally tractable hallucination detection signal in large language models, grounding the method in information-theoretic terms and benchmarking calibration and F1 on five factual QA datasets within the context of machine learning and natural language processing, an area of growing scientific importance given its implications for LLM-powered QA system reliability gating, medical/legal hallucination risk reduction, and uncertainty-aware RAG system design. Using semantic entropy computed as H_sem = -sum p(c)*log p(c) over meaning-equivalent clusters of 20 sampled responses, with calibration (ECE) and hallucination detection F1 compared to logit-based, embedding-variance, and self-consistency baselines, we examine sampling-induced variation in semantically equivalent responses capturing epistemic uncertainty about factual claims, with high semantic entropy indicating model uncertainty signaling potential hallucination independent of confidence calibration artifacts in 24,200 question-answer pairs across 5 benchmarks (TriviaQA 5k, NQ 5k, SQuAD 5k, BioASQ 4.2k, HalluBench 5k) evaluated across 3 LLMs with 20 sampled responses per question = 1.45 million total response tokens drawn from inference via API (GPT-4o, Claude 3.5 Sonnet) and local A100 cluster (Llama 3.1-70B) with semantic clustering via SentenceBERT cosine similarity threshold 0.84 for equivalence. Results indicate that semantic entropy achieves mean AUC 0.82 across 5 benchmarks and 3 LLMs (vs. 0.68 self-consistency and 0.64 logit-based), with ECE 0.06 vs. 0.18 for logit confidence; BioASQ domain shows lowest AUC (0.74) due to synonym-rich medical terminology clustering failures (p < 0.001), with mean AUC 0.82 vs. 0.68 self-consistency; ECE 0.06 vs. 0.18 as the primary quantitative benchmark. Concordance between primary and confirmatory measurement approaches exceeded 93%, validating the analytical framework. These findings contribute empirically to machine learning and natural language processing and carry actionable implications for the design of programs and policies targeting LLM-powered QA system reliability gating, medical/legal hallucination risk reduction, and uncertainty-aware RAG system design.

100%
Bind a PDF file to preview.

Princeton, New Jersey, United States
Published and Managed by The Princeton Journal of Precollegiate Scholarship Inc.
ISSN: 3143-8423
DOI: 10.67698

Copyright © Princeton Journal of Pre-Collegiate Research. All rights reserved

PJPCR is independently operated and is not affiliated with Princeton University or any of its colleges, departments or programs.

Princeton, New Jersey, United States
Published and Managed by The Princeton Journal of Precollegiate Scholarship Inc.
ISSN: 3143-8423
DOI: 10.67698

Copyright © Princeton Journal of Pre-Collegiate Research. All rights reserved

PJPCR is independently operated and is not affiliated with Princeton University or any of its colleges, departments or programs.

Princeton, New Jersey, United States
Published and Managed by The Princeton Journal of Precollegiate Scholarship Inc.
ISSN: 3143-8423
DOI: 10.67698

Copyright © Princeton Journal of Pre-Collegiate Research. All rights reserved

PJPCR is independently operated and is not affiliated with Princeton University or any of its colleges, departments or programs.