Skip to content
AI.info

Research

XInsight: Integrative Stage-Consistent Psychological Counseling Support Agents for Digital Well-Being

XInsight: Integrative Stage-Consistent Psychological Counseling Support Agents for Digital Well-Being Overview Research area: Human-Computer Interaction (cs.HC), specifically LLM-based multi-agent sys

arXiv
2603.06583
Published
2026-01-21
Authors
Fei Wang, Jiangnan Yang, Junjie Chen, Yuxin Liu, Kun Li, Yanyan Wei, Dan Guo, Meng Wang

AI summary

XInsight: Integrative Stage-Consistent Psychological Counseling Support Agents for Digital Well-Being

Overview

  • Research area: Human-Computer Interaction (cs.HC), specifically LLM-based multi-agent systems for web-delivered psychological counseling and digital mental well-being. The paper carries ACM CCS tags for multi-agent planning, cognitive science, and psychology, and was published at the ACM Web Conference 2026 (WWW '26), April 13–17, 2026, Dubai, United Arab Emirates (DOI 10.1145/3774904.3793062, CC BY-NC-ND 4.0).
  • Technical level: Advanced. It assumes familiarity with LLM agent orchestration (five specialized agents, staged workflows, tool-taking memory), psychotherapy paradigms (SFBT, CBT, MBCT), and clinical fidelity instruments (FIT, CTS-R, MBCT-AS).
  • Scope in one sentence: The paper proposes a five-agent, three-stage counseling framework, a matching benchmark, and a scale-guided LLM evaluation protocol, then tests them against prior CBT-oriented multi-agent counseling systems.

What This Paper Is About

Most LLM counseling chatbots are opaque, handle only a single stage of therapy, and are not grounded in established therapeutic practice, which limits their usefulness on web platforms that aim to support mental well-being. The paper's goal is to rebuild counseling support as a stage-consistent multi-agent workflow that mirrors the classical Exploration–Insight–Action paradigm, routes each client to an appropriate therapeutic school, and records the conversation as standardized psychological documents. To judge whether that actually works, the authors also build a benchmark and an evaluation protocol anchored in clinical scales rather than ad-hoc dialogue heuristics.

Key Contributions

  1. Paradigm-aligned architecture. Counseling is recast as a three-stage, multi-agent workflow governed by a unified Reason-Intervene-Reflect cycle, preserving stage-consistent control while keeping dialogue fluid in web settings.
  2. Adaptive Therapeutic Routing (ATR). A routing procedure selects among SFBT, CBT, and MBCT, paired with a unified Therapeutic agent that executes school-consistent submodules in a coherent sequence.
  3. Structured psychological grounding. A Recording agent converts multi-turn conversations into standardized clinical instruments — the Case Conceptualization Form (F_case), the Therapeutic Record (O_ther), and the Relapse Prevention Plan (P_rel) — providing interpretable memory and enabling longitudinal reasoning.
  4. Benchmark and evaluation. The authors collect XInsight-Bench and define a Scale-Guided LLM Evaluation (SGLE) protocol that combines therapy-specific clinical scales (FIT, CTS-R, MBCT-AS) with general counseling criteria (HPEC).

Main Findings

  • Best backbone reaches the top scores. With Qwen3 (14B) as default backbone on XInsight-Bench, XInsight scores 74.92 on FIT (SFBT), 57.96 on CTS-R (CBT), and 22.17 on MBCT-AS (MBCT). The comparison backbones reported in the paper's prose are Qwen2.5 (14B) at 69.94 / 54.53 / 19.91 and Gemma 3 (12B) at 65.14 / 55.05 / 19.99.
  • The MBCT margin is the widest. Qwen3 exceeds Qwen2.5 and Gemma 3 on MBCT-AS by 2.26 and 2.18 points respectively, which the authors attribute to structured paradigm alignment.
  • Smaller models generalize less well across schools. InternLM2.5 (7B), Llama 3.1 (8B), and Mistral (7B) perform reasonably on CTS-R but drop markedly on FIT and MBCT-AS, which the authors read as evidence that paradigm-driven orchestration, not model size alone, drives counseling quality. (Note: for several backbones, the numbers quoted in the running text differ from those listed in Table 2.)
  • Multi-stage beats single-stage baselines. XInsight outperforms CBT-LLM and CACTUS on XInsight-Bench@CBT under CTS-R, with the paper reporting consistent per-case gains.
  • Longer interactions. Table 1 reports an average of 53.7 turns for XInsight (Qwen3-14B), versus 1.0 turn for the GPT-3.5-Turbo baseline, 16.6 turns for CACTUS (LLaMA-3-8B), and 1.0 turn for the LLaMA-3.1-70B baseline. XInsight is also the only framework in that table marked as combining multi-agent, multi-stage, and multi-therapy capabilities.
  • Human-perspective criteria improve. On HPEC (Prof, Com(i), Com(ii), Flu(i), Flu(ii), Sim, Safe), XInsight (CBT) scores 8.79 / 8.95 / 8.76 / 8.68 / 8.37 / 8.42 / 0.00, compared with CBT-LLM at 6.16 / 7.16 / 6.37 / 7.68 / 8.47 / 5.95 / 0.00 and CACTUS at 7.11 / 8.05 / 7.79 / 8.36 / 8.94 / 7.11 / 0.00. XInsight also reports strong marks when extended to SFBT (7.84 / 8.47 / 8.05 / 7.37 / 5.58 / 7.37 / 0.00) and MBCT (8.32 / 8.79 / 8.47 / 7.89 / 7.11 / 8.05 / 0.00).
  • Routing is accurate. ATR achieves F1 = 0.98 on MBCT (accuracy 0.95, precision 1.00, recall 1.00), 0.90 on CBT (accuracy 1.00, precision 0.81, recall 0.90), and 0.89 on SFBT (accuracy 0.83, precision 0.95, recall 0.95). Most confusion occurs between CBT and SFBT, which the authors attribute to semantic overlap.
  • Stages are load-bearing. Replacing a stage with a one-sentence prompt lowers CTS-R to 55.95 (Stage 1), 54.60 (Stage 2), and 55.84 (Stage 3), versus 57.96 for the full model (drops of 2.01, 3.36, and 2.12). Deleting stages entirely yields 55.79 (w/o Stage 1), 51.38 (w/o Stage 2), and 55.42 (w/o Stage 3) — gaps of 2.17, 6.58, and 2.54. Stage 2 (Insight) is the most critical.
  • The Recording agent matters. Removing it drops CTS-R to 50.11, which is 7.85 points below the full model, while a MemGPT-style memory design scores 55.55, 2.41 points below XInsight.
  • Case study. The worked example traces a 44-year-old skilled worker presenting with anxiety and depression through all three stages, including CBT restructuring of a global pessimistic belief that "everything will fall apart," and ends with a Relapse Prevention Plan covering high-risk situations, early warning signs, action plans, and self-encouragement scripts.
  • Not reported in the available content: the size of XInsight-Bench (number of cases, turns, or participants), any explicit limitations section, and the tail of the conclusion (the text is truncated mid-sentence).

Methodology in Plain English

The authors treat a counseling session as a pipeline of specialized roles rather than one chatbot.

  • Five agents. An Exploration agent opens the conversation, builds rapport, monitors mood, gives psychoeducation, sets goals, coaches on negative thoughts, and encourages activity. A Routing agent reads the structured case formulation and picks the therapeutic school that best fits the client. A Therapeutic agent then behaves as three interchangeable specialists — solution-focused (SFBT), cognitive-behavioral (CBT), or mindfulness-based (MBCT) — each with its own sequence of submodules (CBT, for example, moves from automatic thoughts up to core beliefs). A Consolidation agent runs review, skill integration, and relapse prevention. A Recording agent sits across all of them.
  • Reason, Intervene, Reflect. At every stage each agent first computes a context-aware goal, then executes a therapeutic submodule, then updates the structured artifacts. Those artifacts became the memory that later agents read from.
  • Note-taking as memory. Inspired by the Zettelkasten method, the Recording agent compresses each exchange into an "atomic memory unit" tagged with structured information, topic, keywords, memory ID, and an embedding. Units are filed into three standardized tools (case formulation, therapeutic record, relapse prevention plan), and later agents retrieve relevant ones using topic-filtered, cosine-similarity search.
  • Benchmark construction. XInsight-Bench cases were generated with GPT-4o using curated exemplars, then filtered and reviewed by licensed counselors, producing balanced distributions across demographics and domains aligned to CBT, MBCT, and SFBT.
  • Evaluation protocol. SGLE uses GPT-4o as a psychometric rater, scoring with FIT for SFBT, CTS-R for CBT, and MBCT-AS for MBCT, plus a general Human-Perspective Evaluation Criteria set covering things like professionalism and safety from the client's viewpoint.
  • Experimental setup. Everything is built on MetaGPT with temperature fixed at 0. To prevent loops, exploration is capped at 15 turns and each insight/action submodule at 6 consecutive selections. Candidates tested as agent backbones include InternLM2.5 (7B), GLM-4 (9B), Llama 3.1 (8B), Mistral (7B), Falcon-H1 (7B), Gemma 3 (12B), Qwen2.5 (14B), and Qwen3 (14B).

Why This Matters

Impact on research. The paper argues that the bottleneck in counseling agents is not raw model scale but architectural alignment with how therapy actually proceeds. By contributing a benchmark and a clinical-scale evaluation protocol, it gives the field a reproducible way to compare multi-therapy counseling systems instead of relying on dialogic heuristics or single-school evaluations.

Real-world applications:

  • Web-based mental health triage and support. The staged design fits browser-based or app-based platforms where users need flexible, on-demand help, and where the session must be legible to a clinician afterward.
  • Clinician handoff and documentation. The Case Conceptualization Form, Therapeutic Record, and Relapse Prevention Plan convert free-form chat into artifacts a human professional could review, supporting continuity of care.
  • Relapse prevention and follow-up. The Action stage's high-risk-situation and early-warning-sign planning maps directly onto maintenance programs outside active therapy.
  • Cross-therapy matching. ATR could help route users toward solution-focused, cognitive-behavioral, or mindfulness-based content depending on their presenting profile — the reported F1 of 0.98 on MBCT and 0.90 on CBT speaks to how reliably this can be automated.

Industry relevance. The work is explicitly positioned as complementing, not replacing, professional mental health care. For technology companies building digital well-being products, it offers a blueprint for multi-agent counseling features with auditability, and the benchmark plus SGLE protocol give product teams a way to test such features against clinical rubrics rather than ad-hoc ratings. The finding that smaller, cheaper models handle CBT reasonably but falter on FIT and MBCT-AS is directly relevant to cost-sensitive deployment decisions.

Future Directions

  • Improve cross-school generalization in smaller models. InternLM2.5 (7B), Llama 3.1 (8B), and Mistral (7B) dropped markedly on FIT and MBCT-AS, so a key open question is how to make paradigm-driven orchestration effective at lower model scale.
  • Resolve the CBT/SFBT routing confusion. Errors cluster between these two schools, and both sit below the near-perfect MBCT routing (F1 = 0.98), so sharper routing criteria for semantically overlapping schools are needed.
  • Extend beyond the three-stage, three-school scope. XInsight-Bench is aligned to CBT, MBCT, and SFBT only; adding other therapeutic schools would test whether the RIR cycle is genuinely paradigm-general.
  • Validate the artifacts with real clinical outcomes. The evaluation uses GPT-4o as the rater and HPEC criteria; whether the generated Case Conceptualization Forms, Therapeutic Records, and Relapse Prevention Plans translate into meaningful client outcomes beyond rubric scores remains open, as does human professional review at scale.

Target Audience

This paper suits researchers and practitioners working at the intersection of conversational AI and mental health — particularly those building or evaluating multi-agent LLM systems for web-based support. It will also be useful to HCI researchers studying human-AI interaction in sensitive domains, clinical researchers interested in how standardized therapy instruments can be operationalized computationally, and technically experienced product or engineering leads in digital mental health who need a structured reference architecture and a benchmark to test against. Readers without a background in LLM agents or psychotherapy modalities will find the methodology dense, though the abstract and Figure 2 overview carry the main idea.

Authors’ abstract

Web-based platforms are becoming a primary channel for psychological support, yet most LLM-driven chatbots remain opaque, single-stage, and weakly grounded in established therapeutic practice, limiting their usefulness for web applications that promote digital well-being. To address this gap, we present \textbf{XInsight}, a counseling-inspired multi-agent framework that models psychological support as a stage-consistent workflow aligned with the classical \textit{Exploration-Insight-Action} paradigm. Building on structured client representations, XInsight orchestrates specialized agents under a unified \textit{Reason-Intervene-Reflect} cycle: an Exploration agent organizes background and concerns into a structured Case Conceptualization Form, a Routing agent performs Adaptive Therapeutic Routing (ATR) across SFBT, CBT, and MBCT, a unified Therapeutic agent executes school-consistent submodules, and a Consolidation agent guides review, skill integration, and relapse-prevention planning. A Recording agent continuously transforms open-ended web dialogues into standardized psychological artifacts, including case formulations, therapeutic records, and relapse-prevention plans, enhancing interpretability, continuity, and accountability. To support rigorous and transparent assessment, we introduce \textbf{XInsight-Bench} with a Scale-Guided LLM Evaluation (SGLE) protocol that combines therapy-specific clinical scales with general counseling criteria. Experiments show improved paradigm alignment, multi-therapy integration, interaction depth, and interpretability over existing multi-agent counseling systems, indicating that XInsight provides a practical blueprint for integrating counseling-inspired support agents into web applications for digital well-being.

Read the original paper