Skip to content
AI.info

Research

It's LIT! Reliability-Optimized LLMs with Inspectable Tools

Overview Research area: Machine learning — trustworthy AI, large language model tool use, and interpretable/debuggable LLM reasoning. Technical level: Intermediate. Readers should be comfortable with

It's LIT! Reliability-Optimized LLMs with Inspectable Tools
arXiv
2511.14903
Published
2025-11-18
Authors
Ruixin Zhang, Jon Donnelly, Zhicheng Guo, Ghazal Khalighinejad, Haiyang Huang, Alina Jade Barnett, Cynthia Rudin

AI summary

Overview

Research area: Machine learning — trustworthy AI, large language model tool use, and interpretable/debuggable LLM reasoning.

Technical level: Intermediate. Readers should be comfortable with LLM prompting, tool/function calling, and basic evaluation metrics (F1, R²), but no deep mathematics is required.

Scope: The paper introduces LIT (LLMs with Inspectable Tools), a prompting-and-cost framework that steers LLMs to choose reliable and easy-to-debug external tools, plus a new 1,300-question benchmark and a suite of 8 tools built on two datasets.

What This Paper Is About

LLMs often solve problems through opaque reasoning, so when they answer wrong, nobody can tell why or fix it. Existing tool-calling approaches optimize only for whether the task succeeds, ignoring that some tools are far more reliable and far easier to troubleshoot than others. The paper's goal is to force LLMs to prefer the most reliable and inspectable solution path available, assigning a cost to each tool and picking the lowest-cost sequence that still gets the answer right.

Key Contributions

  1. A prompting framework that accounts for tool reliability and inspectability. LIT instructs an LLM to generate multiple candidate solutions (up to four, to limit duplication and token use), compute the total cost of each based on the tools it calls, and select the lowest-cost solution without sacrificing accuracy. Full prompts appear in the paper's Appendix C.

  2. A suite of 8 tools with varying reliability and inspectability: Calculator, DBLoader, PandasInterpreter, PythonInterpreter, Forecaster, TextualClassifier, LLMInferencer, and Finish. Each is designed to give deterministic results for reproducibility.

  3. A new benchmark of 1,300 questions built from 13 templates (each producing 100 variations) over two external databases: the Harvard USPTO Patent Dataset and a new dataset of NeurIPS 2023 paper metadata.

  4. Customizable reliability cost functions. Each tool's cost is the sum of three components — robust performance across inputs (P), ease of debugging (D), and complexity of arguments (C) — that users can set according to their own preferences.

Main Findings

  • LIT improves reliability/inspectability in most cases: LIT produced a solution with similar or better reliability/inspectability in 61 out of 65 cases, and improved or matched the average cost for easy- and medium-difficulty questions in every case except one (Q2).

  • Accuracy is generally maintained or improved: LIT improved or gave comparable performance in 48 out of 65 settings, meaning the transparency gains do not generally trade off against task performance.

  • Hard questions resist the framework: For hard questions, LLMs produced high costs both with and without LIT, suggesting these questions cannot easily be solved with the available reliable and inspectable tools.

  • Medium questions are the weak point: LIT performed well on easy and hard questions but worse on medium-difficulty ones. Notably, no model did well on Question 7, with R² scores close to 0 — barely better than a naïve mean predictor.

  • A concrete debugging example: On Question 10, the black-box solution used a pre-trained BERT model inside TextualClassifier, while LIT selected a pre-trained logistic regression model with comparable accuracy. Because the logistic regression coefficients can be inspected directly, its cost is 7 versus 20 for the black-box BERT solution.

  • Tools carry very different costs: In the paper's cost table, Calculator costs 2 and DBLoader costs 3, while LLMInferencer costs 30 (P=1, D=15, C=14) and TextualClassifier with BERT costs 20 (P=2, D=10, C=8). Finish costs 0 because it merely concludes the task and checks the return data type.

  • Token cost is the main limitation: The prompting strategy increases both input tokens and generated output tokens, since the model must compare multiple solutions, which may be a challenge under context limits.

Methodology in Plain English

The researchers started from a simple observation: tools are not equally trustworthy. A calculator gives the exact right answer for arithmetic, but an ARIMA-based forecasting model depends on assumptions like stationarity and linearity that often do not hold. Likewise, a BERT classifier is accurate but its many parameters make its decisions hard to explain, whereas logistic regression exposes coefficients that reveal how predictions are made.

The team turned that intuition into numbers. Each tool gets a cost built from three questions: Does it perform robustly across inputs (P)? Is it easy to debug (D)? Are its arguments simple (C)? Costs are summed, so the calculator totals 2 and the LLMInferencer totals 30. The code-based tools (PythonInterpreter, PandasInterpreter) scale with the square root of the number of lines times the number of packages.

At inference time, the LIT prompt supplies these cost formulas along with 5 worked examples, then asks the model to draft several candidate solution chains, each composed of sequential tool calls. The model computes each chain's total cost, picks the cheapest one it can, and executes it step by step until it calls Finish. No model training takes place — LIT is purely a prompting and selection procedure. Half of the 1,300 questions were used for validation and half for testing. The team ran this across GPT-3.5-Turbo, GPT-4-Turbo, Gemini-1.5-Pro-001, Claude-3.5-Sonnet-Latest, and Meta-Llama-3.1-70B-Instruct-Turbo, then compared costs and accuracy against the unmodified baseline using one-sided independent two-sample t-tests.

Why This Matters

Research impact: The paper argues that prior tool-calling work optimized only for task correctness, leaving the choice of tool opaque even when more reliable options were available. It also fills a stated gap — before this work there were no metrics, baseline methods, standard frameworks, or question banks for studying tool selection reliability in LLMs. The benchmark and configurable cost functions give other researchers something concrete to build on.

Real-world applications:

  • Patent analysis: The Harvard USPTO Patent Dataset (4.5 million applications filed between 2004 and 2018) supports questions like estimating average filing-to-issuance time or predicting whether an application will be accepted.
  • Scientific literature and conference analytics: The NeurIPS 2023 dataset supports questions about author publication counts, author-number distributions, and predicting oral acceptance from abstracts.
  • Regulated and high-stakes decision support: The paper frames opaque reasoning as a blocker for domains where end users must trust the solution.
  • Production LLM debugging: Because inspectable tools expose their logic, developers can locate where a multi-step agent went wrong instead of guessing.

Industry relevance: Any organization deploying LLM agents that call tools can adopt LIT without retraining a model, since it is a prompting-plus-selection layer compatible with any tool-calling paradigm. The cost functions are explicitly user-configurable, so a team can encode its own tradeoffs between reliability, debuggability, and argument complexity. The paper also notes that the uninspectable baseline is always available, so a deployment can fall back to it when LIT offers no advantage.

Future Directions

  • Reduce token usage. The authors identify this as the primary limitation and call for improving LIT's token efficiency to ease context-limit constraints.
  • Improve medium-difficulty performance. These questions are solvable by multiple different tools and therefore admit many solutions at different costs, yet no model handled Q7 well (R² near 0). The paper calls this a direct avenue for future work.
  • Extend beyond sequential dependent tool calling. The experiments cover only sequential dependent calls, though the paper states LIT is compatible with single-step and iterative multi-step paradigms with feedback.
  • Grow the benchmark and tool suite. The customizable cost functions and tool collection are designed to be extended by other researchers with additional tools and domain-specific reliability metrics.

Target Audience

Researchers and practitioners working on trustworthy machine learning, LLM agents, and tool-augmented language models will get the most from this paper. It is also relevant to engineers building production systems where LLM outputs need to be audited or debugged, and to benchmark designers interested in evaluating tool selection rather than just task success. Readers focused on interpretability and explainable AI will find the cost-function design and the regression-versus-BERT comparison especially useful.

Authors’ abstract

Large language models (LLMs) have exhibited remarkable capabilities across various domains. The ability to call external tools further expands their capability to handle real-world tasks. However, LLMs often follow an opaque reasoning process, which limits their usefulness in high-stakes domains where solutions need to be trustworthy to end users. LLMs can choose solutions that are unreliable and difficult to troubleshoot, even if better options are available. We address this issue by forcing LLMs to use external -- more reliable -- tools to solve problems when possible. We present a framework built on the tool-calling capabilities of existing LLMs to enable them to select the most reliable and easy-to-troubleshoot solution path, which may involve multiple sequential tool calls. We refer to this framework as LIT (LLMs with Inspectable Tools). In order to support LIT, we introduce a new and challenging benchmark dataset of 1,300 questions and a customizable set of reliability cost functions associated with a collection of specialized tools. These cost functions summarize how reliable each tool is and how easy it is to troubleshoot. For instance, a calculator is reliable across domains, whereas a linear prediction model is not reliable if there is distribution shift, but it is easy to troubleshoot. A tool that constructs a random forest is neither reliable nor easy to troubleshoot. These tools interact with the Harvard USPTO Patent Dataset and a new dataset of NeurIPS 2023 papers to solve mathematical, coding, and modeling problems of varying difficulty levels. We demonstrate that LLMs can achieve more reliable and informed problem-solving while maintaining task performance using our framework.

Read the original paper