Skip to content
AI.info

Research

AgentOpt v0.1 Technical Report: Client-Side Optimization for LLM-Based Agent

Overview Research area: Systems and machine learning for LLM-based AI agents, specifically client-side (developer-side) optimization of multi-step agent pipelines. Technical level: Advanced. The paper

AgentOpt v0.1 Technical Report: Client-Side Optimization for LLM-Based Agent
arXiv
2604.06296
Published
2026-04-07
Authors
Wenyue Hua, Sripad Karne, Qian Xie, Armaan Agrawal, Nikos Pagonas, Kostis Kaffes, Tianyi Peng

AI summary

Overview

Research area: Systems and machine learning for LLM-based AI agents, specifically client-side (developer-side) optimization of multi-step agent pipelines.

Technical level: Advanced. The paper combines systems engineering (HTTP-layer interception, caching, concurrency control) with combinatorial black-box optimization and multi-armed bandit theory.

Scope: The paper introduces AgentOpt, a framework-agnostic Python package for client-side agent optimization, focusing on choosing which model to assign to each role in a multi-step agent pipeline under accuracy, cost, and latency constraints.

What This Paper Is About

Existing work on making AI agents efficient has concentrated on the server side, using caching, speculative execution, traffic scheduling, and load balancing to reduce the cost of serving agentic workloads. The authors argue that this leaves out an equally important problem on the client side, where a developer must decide how to spend their own available resources: which models to use for which pipeline roles, how to allocate API budget across steps, and which quality-cost-latency tradeoff to accept. The paper introduces AgentOpt, a Python package that treats model combination selection across pipeline roles as a black-box optimization problem and searches for the Pareto frontier of model assignments given a small labeled evaluation set.

Key Contributions

  1. Formalizing client-side optimization as a distinct systems problem. The authors define client-side optimization as improving an agentic workflow using resources under the developer's control (candidate models, model-to-role assignment, local and remote tools, API budget allocation, batching, caching, and scheduling), and argue it is fundamentally different from server-side optimization because the objective is personalized to the application.

  2. Identifying model combination selection as the first-class optimization lever, and introducing the "combo abstraction." The paper argues that other client-side optimizations (caching, scheduling, speculative execution, tool heuristics) operate only after a model assignment is fixed, so model selection is upstream of the rest. It shows that multi-step agent optimization must be evaluated at the level of full model combinations rather than individual calls, treated as a black-box sequential decision problem naturally formulated as a Markov decision process.

  3. Implementing ten search algorithms for sample-efficient exploration of the combination space. These include UCB-E (Matrix UCB-E), UCB-E with Low-Rank Factorization, Arm Elimination, Epsilon-LUCB, Threshold Successive Elimination, Bayesian Optimization, brute-force search, random search, and Hill Climbing, with Arm Elimination as the default adaptive method in AgentOpt.

  4. Building a framework-agnostic, non-intrusive execution substrate. AgentOpt intercepts LLM calls at the HTTP transport layer, supports multiple selectors through a shared API, and provides caching and bounded parallel evaluation, returning ranked combinations and a Pareto frontier over accuracy, cost, and latency.

Main Findings

  • Model selection dominates other client-side levers by a wide margin. At matched accuracy, the cost gap between the best and worst model combinations ranges from 13× to 32× in the authors' experiments. The paper argues this gap cannot be closed by serving optimizations alone.

  • On BFCL, a cheaper model matches a frontier model at 32× lower cost. Qwen3 Next 80B matches Claude Opus 4.6 in accuracy while reducing cost by 32×.

  • On MathQA, the cost gap between expensive and budget-efficient combinations reaches 24× while maintaining similar accuracy.

  • The strongest standalone model can be the worst in a specific pipeline role. On HotpotQA, Claude Opus 4.6, the strongest model in the benchmark by standalone capability, is the worst planner across all 81 model combinations. As a planner it often answers directly from parametric knowledge and bypasses the solver's search tools, preventing the downstream system from executing the intended reasoning process.

  • Combinations, not individual models, are the correct unit of evaluation. Ministral 3 8B, the cheapest planner in the benchmark, more reliably delegates to the solver; paired with Opus as solver this combination achieves 74.27% accuracy, whereas using Opus for both roles yields only 31.71%.

  • UCB-E recovers near-optimal accuracy with a fraction of the evaluation budget. Across four benchmarks, Matrix UCB-E recovers near-optimal accuracy while reducing evaluation budget by 62–76% relative to brute-force search.

  • Exhaustive evaluation is prohibitive by construction. With N roles and candidate sets M_i, the search space size is the product of |M_i| (that is, |M|^N when roles share a pool), and exhaustive search over K datapoints costs O(K|C|) pipeline evaluations.

  • Framework independence is achieved through transport-layer interception. AgentOpt patches httpx.Client.send() and httpx.AsyncClient.send() and uses Python contextvars to attribute each intercepted call to the correct datapoint and model combination, recording model name, input and output token counts, wall-clock latency, cache status, and identifiers.

  • Caching preserves latency fidelity. The caching layer stores responses keyed by a hash of the request payload, in memory and optionally in SQLite on disk, and cached entries retain their original latency measurements so latency comparisons are not biased by cache hits.

  • Parallelism is bounded at two levels. select_best(parallel=True, max_concurrent=20) controls concurrency, with one semaphore limiting how many combinations are evaluated at once and another limiting datapoints processed concurrently within each combination, keeping a global bound on in-flight API calls.

  • The fourth benchmark and dataset sizes are not named in the available content. The paper reports results across four benchmarks and names BFCL, HotpotQA, and MathQA; the identity of the fourth benchmark and the specific evaluation-set sizes are not reported in the provided text.

Methodology in Plain English

The authors start from the observation that an agent pipeline has several roles (for example planner, solver, critic) and that a developer typically has several candidate models available for each role. Every way of assigning models to roles is a "combination," and the number of combinations grows multiplicatively with the number of roles, so testing all of them on a full dataset quickly becomes too expensive.

They frame the choice as a black-box optimization problem: you cannot tell how good a combination is without running the whole pipeline end-to-end on labeled examples, and the score of one role's model depends on which models sit downstream of it. So instead of picking the best model per role independently, they search over whole combinations.

To do this practically, they built a Python package with three layers. The execution substrate intercepts HTTP requests so that any agent framework calling an LLM through httpx is tracked automatically, without changing agent code. The tracker records token usage, latency, and cache hits, and attributes each call to the current datapoint and combination using context variables. The selector decides which combinations to try next, and all selectors share the same interface so a developer can swap brute-force search for a bandit method or Bayesian optimization without rewriting the agent.

The bandit-style selectors treat each combination as an arm, treat one evaluation of that combination on one datapoint as a pull, and allocate more evaluations to combinations that are still uncertain. Matrix UCB-E scores combinations with a mean plus an exploration term; Matrix UCB-E-LRF additionally fits low-rank factorizations to the partially observed score matrix, using an ensemble with random dropout to estimate uncertainty after a random warmup; Arm Elimination removes combinations whose upper confidence bounds fall below the lower confidence bounds of the leaders; Epsilon-LUCB allocates samples to the most ambiguous leader-challenger pair until the leader is separated by at least epsilon; and Threshold Successive Elimination targets all combinations above a quality threshold rather than a single winner. Hill Climbing instead assumes a neighborhood structure where one role assignment changes at a time, with multiple random restarts, and Bayesian Optimization fits a surrogate model over the discrete combination space using BoTorch with an acquisition function such as (Log) Expected Improvement.

The output is not a single winner but a structured record of the explored design space: aggregate accuracy, latency, token usage, and estimated price for each evaluated combination, plus a Pareto-frontier diagram, ranked tables, and export of the selected configuration to CSV or YAML.

Why This Matters

Impact on research. The paper reframes agent efficiency as two complementary problems rather than one. Server-side systems such as Autellix, ThunderAgent, Continuum, AIOS, vLLM, and SGLang optimize shared infrastructure across many users; client-side optimization optimizes a specific workflow under a specific utility function. The paper argues this second layer cannot be handled by providers because quality-latency-cost preferences are application-specific and unavailable to the model provider. It also draws a sharp mathematical distinction from conventional LLM routing: in single-call routing the unit is one call, while in agent pipelines routing decisions are coupled across stages and must be optimized over full combinations or pipeline-level policies.

Real-world applications:

  • Coding assistants, where a developer may accept a modest accuracy drop for a large cost or latency reduction in an interactive setting.
  • Clinical or legal decision-support systems, where the deployment may prioritize reliability over cost, illustrating why the tradeoff cannot be set centrally.
  • Cost-constrained agent products that must allocate a fixed per-query API budget across planner, solver, critic, and retriever stages.
  • Agent builders composing local tools, remote APIs, and multiple candidate models, who need to choose model assignments before any serving-level optimization applies.

Industry relevance. The 13× to 32× cost gaps at matched accuracy, and the specific BFCL result where Qwen3 Next 80B matches Claude Opus 4.6 at 32× lower cost, mean that an incorrect model assignment can dominate all downstream infrastructure improvements a provider might offer. Because AgentOpt is framework-agnostic, non-intrusive, and requires no proxy servers or per-SDK wrappers, it targets existing production agent stacks (the paper cites Langgraph, AutoGen, OpenClaw, Manus, and ClaudeCode) rather than requiring a new programming framework.

Future Directions

  • Extending beyond model selection to other client-side levers. The paper describes client-side optimization as covering model assignment, tool invocation, API budget allocation across steps, and operational choices such as batching, caching, and scheduling within the application, but states that model combination selection is the first functionality AgentOpt supports, leaving the others as open territory.
  • Generalizing the objective beyond accuracy, cost, and latency. The formulation allows user-defined utility functions and constraints, and different applications may use exact-match accuracy, pass rate, or LLM-as-judge metrics, but how well these alternatives behave as optimization objectives is not established in the available content.
  • Handling larger combination spaces and harder problems. Matrix UCB-E-LRF is described as designed to improve sample efficiency on harder problems where performance differences between combinations are subtle, implying that ranking near-tied combinations remains a hard case worth further study.
  • Understanding when each search algorithm is preferable. The package ships ten algorithms with different assumptions about the structure of the objective (independent arms, neighborhood structure, or surrogate models), and the paper notes that full-dataset methods and bandit-style methods benefit differently from the two-level concurrency scheme, leaving the algorithm-selection question open.

Target Audience

This paper is most useful to agent developers and applied ML engineers who assemble multi-step pipelines from multiple models and need to control cost, latency, and quality under their own budget; to systems researchers studying agent-serving efficiency who want to understand the complementary client-side layer; and to readers interested in applying multi-armed bandit and black-box optimization techniques to combinatorial LLM configuration problems. Familiarity with LLM pipelines and basic bandit concepts helps, but the package design section is accessible to practitioners who mainly want to use the tool.

Authors’ abstract

AI agents are increasingly deployed in real-world applications, including systems such as Manus, OpenClaw, and coding agents. Existing research has primarily focused on \emph{server-side} efficiency, proposing methods such as caching, speculative execution, traffic scheduling, and load balancing to reduce the cost of serving agentic workloads. However, as users increasingly construct agents by composing local tools, remote APIs, and diverse models, an equally important optimization problem arises on the client side. Client-side optimization asks how developers should allocate the resources available to them, including model choice, local tools, and API budget across pipeline stages, subject to application-specific quality, cost, and latency constraints. Because these objectives depend on the task and deployment setting, they cannot be determined by server-side systems alone. We introduce AgentOpt, the first framework-agnostic Python package for client-side agent optimization. We first study model selection, a high-impact optimization lever in multi-step agent pipelines. Given a pipeline and a small evaluation set, the goal is to find the most cost-effective assignment of models to pipeline roles. This problem is consequential in practice: at matched accuracy, the cost gap between the best and worst model combinations can reach 13--32$\times$ in our experiments. To efficiently explore the exponentially growing combination space, AgentOpt implements eight search algorithms, including Arm Elimination, Epsilon-LUCB, Threshold Successive Elimination, and Bayesian Optimization. Across four benchmarks, Arm Elimination recovers near-optimal accuracy while reducing evaluation budget by 24--67\% relative to brute-force search on three of four tasks. Code and benchmark results available at https://agentoptimizer.github.io/agentopt/.

Read the original paper