Research
Evaluating Zero-Shot and One-Shot Adaptation of Small Language Models in Leader-Follower Interaction
Overview Research area: Human-Robot Interaction (HCI / assistive robotics), specifically on-device role classification with small language models. Technical level: Intermediate — the paper combines NL

- arXiv
- 2602.23312
- Published
- 2026-02-26
- Authors
- Rafael R. Baptista, André de Lima Salgado, Ricardo V. Godoy, Marcelo Becker, Thiago Boaventura, Gustavo J. G. Lahr
AI summary
Overview
Research area: Human-Robot Interaction (HCI / assistive robotics), specifically on-device role classification with small language models.
Technical level: Intermediate — the paper combines NLP benchmarking (prompt engineering, fine-tuning, classification metrics) with edge-deployment concerns (latency, throughput, parameter budget), but the framing is applied and accessible to readers with basic machine-learning literacy.
Scope: A benchmark study that tests whether a 0.5-billion-parameter language model can decide, from natural-language input, whether a robot should lead or follow a human partner — under zero-shot and one-shot interaction modes, using either prompt engineering or fine-tuning.
What This Paper Is About
Assistive robots need to know when to take initiative (lead) and when to defer to the user (follow), but large language models are too slow and too resource-hungry to run onboard. This paper builds a leader–follower dialogue dataset and systematically tests whether a small language model, Qwen2.5-0.5B, can handle that role classification cheaply and accurately. The central question is whether prompt engineering or fine-tuning works better, and whether letting the model ask one clarifying question before deciding helps or hurts.
Key Contributions
- A new public dataset for leader–follower HRI communication. Built from DailyDialog samples and augmented with synthetic paraphrases from three LLMs (DeepSeek, Gemini, and GPT-4), yielding 5,400 leader–follower-oriented questions (1,800 per model). The authors state this is the first publicly available dataset tailored specifically to leader–follower HRI communication, addressing a gap they also note was highlighted in a systematic review by Corradini et al.
- Two interaction-mode dataset configurations. One zero-shot configuration (question plus role label) and one one-shot configuration that adds a clarifying question and an LLM-simulated, intentionally ambiguous user reply, using what the authors call a "scarecrow" validation methodology that stands in for a human interlocutor during early-stage validation.
- A comparative benchmark of adaptation strategies. Baseline (unadapted Qwen2.5-0.5B), prompt engineering, and fine-tuning, evaluated across zero-shot and one-shot modes with Monte Carlo Cross-Validation over 30 independent iterations.
- An efficiency analysis and a sentence-length analysis. The paper reports throughput and latency alongside accuracy, and analyzes classification accuracy as a function of input character length for the fine-tuned models in both interaction modes.
Main Findings
- Zero-shot fine-tuning wins decisively on accuracy. Fine-tuning reached 86.66% ± 6.77% accuracy, versus 53.87% ± 10.38% for prompt engineering and 55.00% ± 10.98% for the baseline (z-test, p < 0.001). Fine-tuning also recorded the best precision (84.81% ± 10.07%), recall (90.05% ± 7.98%), and F1-score (86.72% ± 6.53%).
- Unadapted and prompt-engineered models are conservative. In the zero-shot setting, the baseline's precision was 75.17% ± 21.61% but its recall was only 16.39% ± 6.53%, and prompt engineering showed the same skew (56.80% ± 19.20% precision versus 31.03% ± 12.58% recall), meaning the models missed most leader–follower transitions.
- One-shot performance collapses toward chance. Fine-tuning dropped to 51.65% ± 13.40% accuracy, statistically comparable to the baseline (48.00% ± 9.09%) and prompt engineering (45.07% ± 8.17%) (z-test, p > 0.001). One-shot fine-tuned recall fell to 8.90% ± 13.29% and precision to 43.69% ± 45.02%.
- The zero-shot to one-shot degradation is significant. Fine-tuned accuracy fell from 86.66% to 51.65% (p < 0.001); the baseline fell from 55.00% to 48.00%; prompt engineering fell from 53.87% to 45.07%.
- Fine-tuning is also the most efficient strategy. It achieved the lowest latency (around 22.2 ms per sample in both modes) and the highest throughput (432.1 ± 148.2 tokens/s zero-shot; 1851.69 ± 380.09 tokens/s one-shot). Prompt engineering incurred the highest latency, reaching 213.8 ± 309.1 ms in one-shot mode — nearly a twofold increase over its zero-shot latency of 111.9 ± 40.0 ms.
- Sentence length hurts one-shot classification but not zero-shot. Using binning with a step size of 10 characters, zero-shot accuracy stayed above 80% for shorter sentences with a localized dip at the 39-character bin, while one-shot accuracy declined from approximately 65% in the shortest bin to nearly 25% for the longest sentences.
- Likely cause is architectural capacity, not the adaptation method. The authors attribute the one-shot breakdown to the increased context length — the original question, a clarifying question, and a synthetic response — exceeding what a 0.5B-parameter model can process while preserving semantic fidelity.
Methodology in Plain English
The researchers first had to create data, because none existed for this task. They pulled 415 questions from the DailyDialog corpus that matched leader–follower dynamics, labeling a question LEADER when a participant asked for guidance or directions and FOLLOWER when a participant started a task alone or suggested being accompanied. Of these, 315 questions (148 FOLLOWER, 167 LEADER) were used as seeds for augmentation, and 100 original human-sourced questions (50 of each class) were held back as the test set. Three LLMs each generated roughly six paraphrases per seed question, producing 5,400 synthetic questions, 1,800 per model. To check that the paraphrases kept their meaning, the team embedded real and synthetic sentences with Sentence-BERT and measured maximum cosine similarity against real samples of the same class.
From that pool they built two test configurations. The zero-shot version just asks the model to classify a question. The one-shot version simulates a more realistic exchange: the model sees a clarifying question and a deliberately ambiguous reply generated by an LLM acting as a stand-in human.
They then tried three ways of getting Qwen2.5-0.5B to make the call. The baseline used the model as-is with only a task instruction. Prompt engineering used hand-designed prompts — one system prompt for zero-shot, and two for one-shot (one to elicit a clarifying question, a second to make the final classification). Fine-tuning used the Autotrain framework for text classification, appending a lightweight linear head to the pretrained model, trained for 10 epochs with a 50-step learning-rate warmup, weight decay of 0.01, batch size 16, FP16 mixed precision, and a stratified 20% validation split; the best model was chosen by accuracy. Fine-tuning was run on an NVIDIA server with L40S 48GB GPUs. In the one-shot condition, two separate fine-tuned models were trained — one for the clarifying question, one for final classification. Every configuration was run through Monte Carlo Cross-Validation with 30 independent iterations so the reported numbers would be statistically robust.
Why This Matters
Impact on research. The paper supplies the first publicly available dataset for leader–follower HRI communication, which makes this class of experiments reproducible rather than one-off. It also reframes a common assumption — that richer, multi-turn dialogue is always better — by showing that for sub-1B models on the edge, extra context can actively degrade a safety-relevant decision.
Real-world applications:
- Assistive robots in hospitals and clinics that must decide whether to guide a patient to a location or accompany them.
- Remote homecare, a scenario the authors specifically name for elderly coffee farmers in rural Brazil, where internet-independent inference matters.
- Edge-deployed dyadic systems where the classified role acts as a high-level trigger selecting either a proactive guidance policy or a reactive, impedance-based following policy.
- Mixed-initiative interfaces generally, where a system must balance user agency against proactive support.
Industry relevance. The results speak directly to the practical constraints of embedded robotics: fine-tuning beat prompt engineering on both accuracy and latency, and its latency stayed flat at about 22 ms even when the one-shot interaction generated more tokens. For teams deciding how to spend a small on-device parameter budget, that combination of better decisions and lower latency is the central engineering takeaway.
Future Directions
- Extending the dataset. The authors acknowledge their dataset is synthetic and limited in scale, and call for larger, multimodal, and human-annotated samples to strengthen ecological validity.
- Making one-shot interaction viable. Because one-shot dialogue mimics real dyadic coordination, the authors argue future work must build specialized edge evaluation frameworks — including specialized context-pruning pipelines — so multi-turn flexibility does not cost initiative-detection reliability.
- Evaluating the generated content, not just the labels. The paper calls for assessing the semantic quality of the model's clarifying questions and the contextual veracity of the scarecrow simulation responses, via either rigorous mathematical evaluation of semantic fidelity or human-in-the-loop subjective judgment.
- Broadening the technical search space. Exploring alternative SLM architectures, compression strategies, and multilingual settings, and validating the models in real-world robotic experiments embedded in dynamic tasks and environments.
Target Audience
Robotics and HRI researchers working on assistive or collaborative systems; applied NLP practitioners interested in how far sub-1B models can be pushed on a narrow classification task; embedded and edge-ML engineers balancing accuracy against latency, memory, and power; and designers of mixed-initiative or leader–follower interfaces who need to know whether multi-turn clarification is worth the cost on constrained hardware.
Authors’ abstract
Leader-follower interaction is an important paradigm in human-robot interaction (HRI). Yet, assigning roles in real time remains challenging for resource-constrained mobile and assistive robots. While large language models (LLMs) have shown promise for natural communication, their size and latency limit on-device deployment. Small language models (SLMs) offer a potential alternative, but their effectiveness for role classification in HRI has not been systematically evaluated. In this paper, we present a benchmark of SLMs for leader-follower communication, introducing a novel dataset derived from a published database and augmented with synthetic samples to capture interaction-specific dynamics. We investigate two adaptation strategies: prompt engineering and fine-tuning, studied under zero-shot and one-shot interaction modes, compared with an untrained baseline. Experiments with Qwen2.5-0.5B reveal that zero-shot fine-tuning achieves robust classification performance (86.66% accuracy) while maintaining low latency (22.2 ms per sample), significantly outperforming baseline and prompt-engineered approaches. However, results also indicate a performance degradation in one-shot modes, where increased context length challenges the model's architectural capacity. These findings demonstrate that fine-tuned SLMs provide an effective solution for direct role assignment, while highlighting critical trade-offs between dialogue complexity and classification reliability on the edge.