Skip to content
AI.info

Research

Small Language Models for Efficient Agentic Tool Calling: Outperforming Large Models with Targeted Fine-tuning

Overview Research area: Efficient generative AI, specifically small language models (SLMs) for agentic tool calling and tool-augmented reasoning, with an emphasis on enterprise cost optimization. Tech

Small Language Models for Efficient Agentic Tool Calling: Outperforming Large Models with Targeted Fine-tuning
arXiv
2512.15943
Published
2025-12-17
Authors
Polaris Jhandi, Owais Kazi, Shreyas Subramanian, Neel Sendas

AI summary

Overview

Research area: Efficient generative AI, specifically small language models (SLMs) for agentic tool calling and tool-augmented reasoning, with an emphasis on enterprise cost optimization.

Technical level: Intermediate. Readers should understand supervised fine-tuning (SFT), model parameters, and benchmark evaluation, but the paper is written at an applied engineering level rather than a theory level.

Scope: The paper reports a single fine-tuning experiment on facebook/opt-350m trained on the ToolBench dataset and evaluated against much larger baseline models on six ToolBench test categories.

What This Paper Is About

Running large language models in production is expensive, requires substantial infrastructure, and often depends on closed APIs that raise privacy, latency, and robustness concerns. The authors ask whether a small model, carefully fine-tuned for one narrow job, can beat far larger general-purpose models at agentic tool calling — the task of deciding which API to call, with which arguments, in the correct structured format. Their goal is to test whether targeted training can replace brute-force scaling for this specialized workflow.

Key Contributions

  1. A targeted fine-tuning pipeline for a 350M-parameter model. The authors fine-tuned facebook/opt-350m for a single epoch using the Hugging Face TRL Supervised Fine-Tuning (SFT) trainer on the ToolBench dataset, with training conducted on Amazon SageMaker (instance type ml.g5.8xlarge).

  2. A demonstrated parameter-efficiency result. The fine-tuned SLM achieved a 77.55% pass rate on ToolBench, compared with ToolLLaMA-DFS at 30.18%, ChatGPT-CoT at 26.00%, ToolLLaMA-CoT at 16.27%, and Claude-CoT at 2.73%.

  3. Category-level evidence across six ToolBench splits. The paper reports per-category results for G1-instruction, G1-category, G1-tool, G2-category, G2-instruction, and G3-instruction, showing consistent performance rather than gains concentrated in one scenario type.

  4. A resource-and-hyperparameter recipe. The paper documents the exact configuration — learning rate 5×10⁻⁵, 100 warmup steps, effective batch size 32 via gradient accumulation over 4 steps, gradient clipping with max_norm=0.3, FP16 mixed precision, gradient checkpointing, and the AdamW optimizer with 0.01 weight decay — totaling 5,860 training steps.

Main Findings

  • Overall pass rate: The fine-tuned OPT-350M reached a 77.55% pass rate, while ToolLLaMA-DFS (7B) reached 30.18%, ChatGPT-CoT (175B) reached 26.00%, ToolLLaMA-CoT (7B) reached 16.27%, and Claude-CoT (52B) reached 2.73%. The reported gaps relative to the SLM are -47.37%, -51.55%, -61.28%, and -74.82% respectively.

  • Scale of the advantage: The paper describes the improvement over ChatGPT-CoT as a 2.98x gain, and states the model used 20-500x fewer parameters than the baselines.

  • Per-category consistency: Results ranged from 74% to 80.5% across categories — G1_instr 78.5, G1_cat 74.0, G1_tool 79.0, G2_cat 80.5, G2_instr 74.5, G3_instr 80.0 — which the authors describe as a 6.5% range. The corresponding baseline averages were 30.2 (ToolLLaMA-DFS), 26.0 (ChatGPT-CoT), 16.3 (ToolLLaMA-CoT), and 2.7 (Claude-CoT).

  • Strongest absolute margin on the hardest split: On G3_instr, the SLM scored 80.0 against 22.0 for ToolLLaMA-DFS, 5.0 for ChatGPT-CoT, 6.0 for ToolLLaMA-CoT, and 4.0 for Claude-CoT.

  • Proposed explanation: The authors attribute the result to three factors — parameter efficiency (baselines suffer "parameter dilution" from general-purpose training), behavioral focus (learning structured Thought-Action-Action Input patterns instead of verbose or creative responses), and evaluation alignment between the training data and ToolBench's criteria.

  • Training data scale: The ToolBench source data covers over 16,000 real-world APIs from RapidAPI Hub; after transformation into structured training sequences, the dataset comprised 187,542 examples.

  • Win rate not reported: ToolEval defines a Win Rate metric comparing solution quality between models, but the paper does not report win-rate numbers — only pass rates.

  • Acknowledged weakness: The authors state that the model was optimized for ToolBench evaluation criteria and may not generalize to other tool-calling frameworks or real-world API ecosystems.

Methodology in Plain English

The researchers took an off-the-shelf 350-million-parameter model, OPT-350M (released by Meta AI in 2022 as part of the OPT family), and taught it one skill: producing ToolBench-formatted tool calls.

First, they reshaped the ToolBench multi-turn instruction data into flat training sequences. System prompts, user queries, and assistant responses were concatenated with delimiters so the model sees clean instruction-following examples. The transformation scripts were generated using Amazon Q, the generative AI assistant from AWS. This produced 187,542 training examples.

Then they fine-tuned for exactly one epoch on Amazon SageMaker using Hugging Face TRL's SFT trainer. The hyperparameters were chosen for a "high-learning, high-stability" balance: a conservative learning rate of 5×10⁻⁵, 100 warmup steps, an effective batch size of 32 built from gradient accumulation over 4 steps, and aggressive gradient clipping at max_norm=0.3. FP16 mixed precision and gradient checkpointing kept memory use manageable for long tool-chain sequences. With 187,542 records divided by an effective batch size of 32, training ran for 5,860 steps.

Evaluation used ToolEval, an automated framework that employs ChatGPT as the judge. It scores Pass Rate (did the solution complete the instruction within a limited API call budget?) and Win Rate (how does solution quality compare between models?). The judge does not require live API execution. The test set contains 1,100 queries across six categories: G1-instruction (200), G1-category (200), G1-tool (200), G2-instruction (200), G2-category (200), and G3-instruction (100). All models ran under identical inference settings — maximum sequence length 8192 tokens, batch size 8 per device, temperature 0.1, and a maximum of 10 reasoning iterations per query. Each query received at least 4 assessments, aggregated by majority voting, with the environment running Python 3.9 and PyTorch, evaluation steps every 1000 iterations, and 4 dataloader workers.

Why This Matters

The paper argues that the standard assumption — bigger models are necessary for complex reasoning — does not hold for narrow, well-specified tasks like tool calling. If a 350M-parameter model can reach 77.55% pass rate where a 175B-parameter model reaches 26.00%, then the constraint on deploying agentic AI is training data and design quality rather than compute budget. The authors frame this as democratizing access to tool-calling agents for organizations without large-scale ML infrastructure.

Real-world areas the findings bear on (the paper discusses enterprise deployment broadly rather than naming specific domains):

  • Enterprise workflow automation, where routine tasks such as document summarization, query answering, and structured data interpretation are handled repeatedly and the cost of an LLM call per task dominates the budget.
  • API-driven agent deployments, where a model must select among external APIs and produce correctly formatted calls, a setting the paper ties directly to RapidAPI Hub tooling.
  • Resource-constrained or privacy-sensitive environments, where avoiding closed frontier-model APIs reduces data-privacy, latency, and robustness risk — concerns the paper raises in its introduction.
  • High-volume production serving, where the paper claims the 350M parameter count translates into cost savings in both training and inference while maintaining what it calls state-of-the-art results.

Industry relevance: All four authors are affiliated with Amazon Web Services (Seattle, WA, USA), and the work uses AWS tooling — Amazon SageMaker for training and Amazon Q for data transformation scripts. The paper is an applied engineering case study aimed at practitioners deciding whether to fine-tune a small model instead of paying for a large one.

Future Directions

  • Test the generalization boundary. The authors state it is unknown whether the model transfers to other tool-calling frameworks or real-world API ecosystems with different interaction patterns, and call for research into those boundaries.

  • Find the optimal parameter count per domain. The paper suggests the "sweet spot" likely varies by task complexity and that the relationship between task complexity and model capacity warrants systematic investigation.

  • Build hybrid systems. The authors propose combining the efficiency of targeted small models with the adaptability of larger systems, so that a small model handles structured tool calls while something larger handles ambiguous or context-heavy requests.

  • Address maintenance and data dependency. The limitations section notes that as APIs evolve, the specialized model may need frequent retraining, and that performance is bounded by the quality, coverage, and biases of the ToolBench training data — with the risk of brittle behavior on novel API designs. The paper also flags that fine-tuning still requires significant compute and high-quality data, which may limit accessibility.

  • Examine the theoretical basis. The conclusion calls for investigating why targeted fine-tuning is so effective for small models in the first place.

Target Audience

Machine learning engineers and applied scientists who are choosing between fine-tuning a small model and calling a large one; enterprise architects evaluating the cost, latency, and privacy trade-offs of deploying agentic AI in production; and researchers studying tool-augmented language models who want a concrete, single-epoch data point on the parameter-versus-performance question. Readers looking for a controlled ablation study or win-rate comparisons will not find them here — the paper reports one training configuration and pass-rate results only.

Authors’ abstract

As organizations scale adoption of generative AI, model cost optimization and operational efficiency have emerged as critical factors determining sustainability and accessibility. While Large Language Models (LLMs) demonstrate impressive capabilities across diverse tasks, their extensive computational requirements make them cost-prohibitive for routine enterprise use. This limitation motivates the exploration of Small Language Models (SLMs), which can deliver comparable performance in targeted applications while drastically reducing infrastructure overhead (Irugalbandara et al., 2023). In this work, we investigate the feasibility of replacing LLM-driven workflows with optimized SLMs. We trained a domain-adapted SLM to execute representative tasks traditionally handled by LLMs, such as document summarization, query answering, and structured data interpretation. As part of the experiment, we investigated the fine-tuning of facebook/opt-350m model (single epoch only) using the Hugging Face TRL (Transformer Reinforcement Learning), specifically the Supervised Fine-Tuning (SFT) trainer. The OPT-350M model was released by Meta AI in 2022 as part of the OPT (Open Pretrained Transformer) family of models. Similar studies demonstrate that even models at the 350M parameter scale can meaningfully contribute to instruction-tuning pipelines (Mekala et al., 2024). Experimental results demonstrated that our fine-tuned SLM achieves exceptional performance with a 77.55\% pass rate on ToolBench evaluation, significantly outperforming all baseline models including ChatGPT-CoT (26.00\%), ToolLLaMA-DFS (30.18\%), and ToolLLaMA-CoT (16.27\%). These findings emphasize that thoughtful design and targeted training of SLMs can significantly lower barriers to adoption, enabling cost-effective, large-scale integration of generative AI into production systems.

Read the original paper