Skip to content
AI.info

Research

TextBandit: Evaluating Probabilistic Reasoning in LLMs Through Language-Only Decision Tasks

Overview Research area: Natural Language Processing / evaluation of large language model reasoning, specifically probabilistic reasoning and sequential decision-making under uncertainty. Technical lev

TextBandit: Evaluating Probabilistic Reasoning in LLMs Through Language-Only Decision Tasks
arXiv
2510.13878
Published
2025-10-13
Authors
Jimin Lim, Arjun Damerla, Arthur Jiang, Nam Le

AI summary

Overview

  • Research area: Natural Language Processing / evaluation of large language model reasoning, specifically probabilistic reasoning and sequential decision-making under uncertainty.
  • Technical level: Intermediate. The paper assumes familiarity with multi-armed bandit concepts (exploration vs. exploitation, regret, Thompson Sampling, UCB) but explains its benchmark design in accessible terms.
  • Scope in one sentence: The paper introduces TextBandit, a text-only multi-armed bandit benchmark, and uses it to compare four open-source LLMs (Qwen3-4B, Qwen3-8B, Llama-3.1-8B, Phi-2) against four standard bandit algorithms (Thompson Sampling, UCB, Epsilon-Greedy, Random Choice).

What This Paper Is About

Large language models have shown Bayesian-like behavior on constrained tasks, but it is unclear whether they can make repeated decisions under uncertainty when the only information available is natural language. To test this, the authors build a multi-armed bandit environment where the model never sees probabilities or numbers, only feedback phrased as "you earned a token" or "you did not earn a token," and must figure out which option is better over time. The goal is to measure whether probabilistic reasoning can emerge from linguistic cues alone.

Key Contributions

  1. A new benchmark (TextBandit) that evaluates LLM decision-making under uncertainty using purely textual feedback, with no numerical cues or explicit probabilities. The authors state that to their knowledge, no prior benchmark evaluates LLMs in this manner.
  2. A head-to-head comparison of four open-source LLMs against four established bandit baselines (Thompson Sampling, UCB, Epsilon-Greedy, and Random Choice) on the same task.
  3. An empirical demonstration that one LLM can outperform classical bandit algorithms on best-arm selection: Qwen3-4B reached 89.2%, exceeding the best baseline (Thompson Sampling at 51.1%).
  4. Evidence that model size does not predict better decision-making in this setting, with the larger Qwen3-8B and Llama-3.1-8B underperforming the smaller Qwen3-4B, which the authors attribute in part to overthinking.

Main Findings

  • Qwen3-4B dominated all baselines and all other LLMs on best-arm selection. It achieved an 89.2% best-arm selection rate and a final cumulative reward of 11150. The next best model, Qwen3-8B, reached 37.5% and 4686.
  • Most LLMs underperformed the classical baselines. Qwen3-8B (37.5%), Llama-3.1-8B (31.6%), and Phi-2 (25.4%) all fell below Thompson Sampling (51.1%) and UCB (47.6%); Phi-2 also fell below Random Choice (31.8%).
  • Baseline results as reported: Thompson Sampling had the best baseline best-arm selection rate at 51.1% with a final cumulative reward of 8297; UCB had 47.6% with a cumulative reward of 4696; Epsilon-Greedy had 38.1% with 6029; Random Choice had 31.8% with 5783.
  • Performance degraded as the number of arms increased. Llama-3.1-8B's average accuracy on the optimal arm fell from 31.56% in the two-arm test to 7.37% in the five-arm test. Phi-2 went from 25.45% to 17.78%, Qwen3-4B from 89.22% to 6.53%, and Qwen3-8B from 37.49% to 17.09%.
  • Larger models did not produce better decisions. Average accuracy across all tests was 44.56% for Qwen3-4B (4B parameters), 33.96% for Qwen3-8B (8B), 30.7% for Llama-3.1-8B (8B), and 27.6% for Phi-2 (2.7B). The authors note Phi-2 was too small, while Llama-3.1-8B and Qwen3-8B were "too big and overthinking too much."
  • Regret trends matched reward trends. Lower cumulative regret indicated more efficient decision-making. The four-arm prompt produced the lowest cumulative regret across all models, while the five-arm prompt produced the highest. Qwen3-4B excelled specifically in the two-arm case.
  • Qwen3-8B spent an exceptionally large amount of time reasoning instead of producing concise answers during testing.

Methodology in Plain English

The researchers set up a slot-machine-style decision game. Each environment has between 2 and 5 arms, and each arm has a fixed but hidden success probability. In the two-arm configuration, one arm succeeds 65% of the time and the other 30% of the time; these probabilities are never revealed to the model.

At every round, the model receives a prompt containing three things: a natural-language instruction framing the task as a decision situation, a plain-language history of all previous choices and outcomes in the current episode (for example, "Slot machine 1 won," "Slot machine 2 lost"), and a request to output the number of the arm it wants next. The model's number is extracted, the outcome is simulated from the hidden probability, and the new result is appended to the history for the next round.

Notably, models are given no chain-of-thought scaffolding and no memory between runs. Each step is a single-shot completion, and the same prompt format is used across all models and arm counts. The only variation is the number of arms and the accumulated outcome history.

Evaluation covered 500 independent runs, each consisting of 25 decision-making iterations. Three metrics were tracked: cumulative reward (a token for each success, nothing for a failure), cumulative regret (the opportunity cost of not picking the optimal arm), and best-arm selection rate (how often the model chose the 65% arm). The same task was run for Thompson Sampling, UCB, Epsilon-Greedy, and Random Choice as comparisons.

The four LLMs evaluated were chosen to span different architectures, parameter sizes, and training methodologies: Qwen3-4B (4B, decoder-only transformer), Qwen3-8B (8B, decoder-only transformer), Llama-3.1-8B (8B, decoder-only transformer), and Phi-2 (2.7B, transformer). Code and evaluation scripts are released at https://github.com/ChainedTears/TextBandit, and experiments ran on GPU instances hosted on RunPod using the Hugging Face Transformers library.

Why This Matters

Impact on research. The paper argues that basic probabilistic reasoning can emerge from language-only interaction, without numerical cues, and offers a minimal benchmark for studying that capability. It also challenges the assumption that scaling model size improves decision-making under uncertainty: the best performer was not the largest model tested, and the authors suggest larger models trained for complex reasoning may overthink simple reinforcement environments.

Real-world applications (drawn from the paper's framing of decision-making under uncertainty):

  • Decision-making settings where traditional Bayesian or reinforcement learning approaches are impractical because they require complex math or data that is not readily available.
  • Systems where only linguistic feedback is available and explicit reward numbers are not exposed to the agent.
  • Low-information, fast-feedback environments where lightweight models may be preferable to large ones.
  • Evaluating and comparing model behavior in interactive, non-numeric decision loops before deployment.

Industry relevance. The results suggest that model selection for interactive decision tasks should not default to the largest available model. They also show that classical bandit algorithms remain strong general-purpose baselines (three of four outperformed three of the four LLMs on best-arm selection), and that removing chain-of-thought materially changes measured performance.

Future Directions

  • Testing larger-scale models such as Qwen-32B or GPT-4, which were excluded here due to computation limits, to see whether they converge more stably or adapt better under uncertainty.
  • Moving beyond the static, simple reward structure to non-stationary or dynamic bandit environments with delayed rewards and multi-step planning.
  • Reintroducing chain-of-thought prompting to measure how reasoning scaffolds affect the exploration-exploitation balance, since the single-shot design may have underestimated model capability.
  • Evaluating fine-tuned model variants and scaling trends to determine whether the poor showing of larger models is a training artifact or a genuine property of the setting.

Target Audience

Researchers and practitioners working on LLM evaluation, probabilistic reasoning, and sequential decision-making, particularly those interested in benchmarks that isolate language-based adaptation from numerical computation. It is also relevant to engineers choosing models for interactive or agentic systems, and to readers familiar with multi-armed bandit literature who want to see how LLMs compare against textbook algorithms. A working understanding of exploration vs. exploitation, regret, and bandit baselines will help, though the benchmark design itself is explained without heavy mathematics.

Authors’ abstract

Large language models (LLMs) have shown to be increasingly capable of performing reasoning tasks, but their ability to make sequential decisions under uncertainty only using natural language remains underexplored. We introduce a novel benchmark in which LLMs interact with multi-armed bandit environments using purely textual feedback, "you earned a token", without access to numerical cues or explicit probabilities, resulting in the model to infer latent reward structures purely off linguistic cues and to adapt accordingly. We evaluated the performance of four open-source LLMs and compare their performance to standard decision-making algorithms such as Thompson Sampling, Epsilon Greedy, Upper Confidence Bound (UCB), and random choice. While most of the LLMs underperformed compared to the baselines, Qwen3-4B, achieved the best-arm selection rate of 89.2% , which significantly outperformed both the larger LLMs and traditional methods. Our findings suggest that probabilistic reasoning is able to emerge from language alone, and we present this benchmark as a step towards evaluating decision-making capabilities in naturalistic, non-numeric contexts.

Read the original paper