Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools

Panagiotis Lymperopoulos (Tufts University), Vasanth Sarathy (Tufts University)

Abstract

Modern Large Language Models (LLMs) increasingly rely on external tools-such as classifiers and knowledge retrieval systems-to deliver accurate answers when their pre-trained knowledge falls short. While this integration broadens their utility, it also raises a critical issue: ensuring the trustworthiness of the combined outputs. In high-stakes settings like medical decision-making, it is vital to evaluate uncertainty in both the LLM's response and the external tool's output. In this work we introduce a novel framework that jointly assesses the combined uncertainty of the LLM and its external tools and derive practical and effective approximations to estimate uncertainty. Our approach is validated on two synthetic QA datasets and an experiment with retrieval-augmented generation (RAG) systems, demonstrating enhanced reliability when external information is required for the LLM to produce answers.