Skip to content
AI.info

Research

Towards AI as Colleagues: Multi-Agent System Improves Structured Ideation Processes

Towards AI as Colleagues: Multi-Agent System Improves Structured Ideation Processes Overview Research area: Human-Computer Interaction (HCI) — specifically human–AI collaborative creativity, multi-age

Towards AI as Colleagues: Multi-Agent System Improves Structured Ideation Processes
arXiv
2510.23904
Published
2025-10-27
Authors
Kexin Quan, Dina Albassam, Mengke Wu, Zijian Ding, Jessie Chin

AI summary

Towards AI as Colleagues: Multi-Agent System Improves Structured Ideation Processes

Overview

Research area: Human-Computer Interaction (HCI) — specifically human–AI collaborative creativity, multi-agent LLM systems, and interaction design for ideation support. The paper is published at CHI '26 (Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, Barcelona, April 13–17, 2026; DOI 10.1145/3772318.3790375) and appears as arXiv:2510.23904v2 [cs.HC].

Technical level: Intermediate. The system architecture (React frontend, Flask API, GPT-4o, persona ranking and history compression) is described at a design level rather than with implementation code, but the paper assumes familiarity with HCI study conventions such as within-subjects designs, Likert-scale instruments, and non-parametric testing.

Scope in one sentence: The paper introduces MultiColleagues, a multi-agent conversational ideation system with role-differentiated AI "colleagues," an Explore/Focus mode switch, and an AI facilitator, and evaluates it against ChatGPT in a 20-participant within-subjects study.

What This Paper Is About

Most AI systems are built to execute predefined steps, which makes them useful process coordinators but weak at joint problem-solving or contributing genuinely new ideas. The authors argue that LLMs could instead be experienced as colleagues rather than tools or copilots, and they built MultiColleagues to test that idea in the context of brainstorming. The goal is to determine whether a socially orchestrated multi-agent design changes how people collaborate, explore ideas, and judge the quality and novelty of what they produce, compared with a standard single-agent chat workflow.

Key Contributions

  1. A multi-agent "AI as colleague" system. MultiColleagues deploys a roster of distinct AI personas with defined professional backgrounds, communication styles, and domain expertise, who converse with each other and with the user rather than simply answering prompts. The system integrates dynamic turn-taking, explicit divergent–convergent phase shifts, and user-facing facilitation in one unified system — a combination the authors identify as absent from prior work.

  2. A design framework with three stated goals. (DG1) Support adaptive human–AI co-ideation dynamics; (DG2) enable rich, multi-perspective co-ideation; (DG3) facilitate purposeful and transparent human–AI collaborative control. Each design goal is mapped to concrete implementation mechanisms (the Explore/Focus dual-mode framework, persona orchestration, and pause points plus the AI facilitator).

  3. A controlled within-subjects evaluation against ChatGPT. Twenty participants used both MultiColleagues and a ChatGPT baseline (both powered by GPT-4o) in counterbalanced order, so the comparison isolates interaction paradigm rather than underlying model.

  4. A multi-layered analysis pipeline. Survey ratings analyzed with the Wilcoxon signed-rank test, thematic analysis of interview transcripts, plus an LLM-assisted semantic segmentation of user contributions into main topics and sub-topics (yielding a "branching ratio"), with noun extraction used as a proxy for concept introduction.

Main Findings

  • Stronger perceived social presence: The abstract reports that MultiColleagues "fostered stronger perceived social presence" than the single-agent baseline.

  • Higher perceived outcome quality and novelty: Participants rated their ideation outcomes as higher in quality and novelty when using MultiColleagues.

  • More elaboration during ideation: The authors report that participants showed more elaboration during ideation with the multi-agent system.

  • Personas and orchestration were designed to build on each other: Unlike workflow-based systems that structure generation through predefined stages, MultiColleagues has AI personas build on each other's reasoning, surface differences, and adapt through dynamic turn-taking and facilitator oversight, with a small randomization factor (20%) in speaker ranking to avoid rigid patterns.

  • Human oversight was preserved deliberately: Strategic interaction points after each AI response let users choose to "Continue" autonomous AI discussion or "Call Facilitator," based on the authors' finding that uninterrupted AI generation creates information overload and diminishes human creative contribution and strategic oversight capacity.

  • Facilitation worked as a metacognitive regulator: The AI Facilitator monitors conversation dynamics, intervenes when discussions deviate from productive ideation patterns, and prompts users on whether to keep exploring or move to convergent evaluation.

  • A comparison table positions the work against five prior systems: LLM Discussion (COLM'24), SWTW (CHI'24), Weaver (CHI EA'25), Supermind (CI'24), and CoExploreDS (CHI'25). Per the paper's Table 1, MultiColleagues is the only one marked as having dynamic turn selection, divergent–convergent phases, human-facing orchestration, explicit agent identities, an interactive (in-situ) study, user agent-picking, and a controlled comparison against ChatGPT.

  • Statistic-level results are not reported in the available content. The provided text is truncated before the results tables and inferential statistics, so specific medians, p-values, and effect sizes cannot be stated here.

Methodology in Plain English

The authors first reviewed prior multi-agent and AI-assisted ideation systems, then identified a gap: existing tools implement pieces of good brainstorming (divergent/convergent structuring, role play, facilitation) but rarely combine them. They built MultiColleagues to fill that gap.

The system has two layers: a React frontend and a Flask API, with GPT-4o generating natural language. Users pick a team of AI personas, state a problem, and then brainstorm in one of two thinking modes — Explore (broad idea generation) or Focus (evaluating and synthesizing) — switching manually as they see fit. Personas take turns: each produces an initial response, an AI-driven evaluation picks the opening speaker, and subsequent turns are chosen by a ranking mechanism weighing contextual relevance, conversation history, and unexpressed perspectives. When conversations get long, an automated pipeline summarizes older persona contributions while keeping recent turns in full.

For evaluation, 20 participants (9 male, 11 female; aged 20–39, M = 26.7, SD = 3.8; 15 students and 5 early-career professionals; baseline creativity M = 5.39, SD = 1.16 across 11 items on a 7-point scale) completed approximately 70-minute remote Zoom sessions. Each participant used both MultiColleagues and ChatGPT with the same self-chosen problem, in counterbalanced order, with roughly 10-minute ideation sessions per system. Participants thought aloud and recorded ideas in a side-by-side Google Doc. After each system they filled out a 12-item survey on 7-point Likert scales, drawn from an initial pool of 30 items and refined with a pilot of 5 participants, covering Experience (3 items), Outcomes (3 items), and System Design & User Control (4 items) as listed. A semi-structured comparative interview followed.

Analysis combined quantitative and qualitative strands. Paired Likert ratings were compared with the non-parametric Wilcoxon signed-rank test because the data are ordinal and the design is within-subjects. Two researchers performed thematic analysis of interview transcripts independently before reconciling. For user contributions, GPT-5 with role-based prompts segmented each message into main topics and sub-topics; each conversation was processed three times and averaged, with inter-rater reliability of Cohen's κ = 0.86. A "branching ratio" (sub-topics divided by main topics) characterized whether a participant developed ideas linearly or explored multiple parallel directions, and noun extraction was used to estimate conceptual density.

Why This Matters

The paper argues for a shift in how AI is framed in collaborative work: from process partner or tool to colleague that shares intent and strengthens group dynamics. It also makes a design argument that mechanisms previously studied in isolation — turn-taking, phase structuring, role differentiation, facilitation — should be integrated, because real brainstorming blends them.

Real-world applications suggested by the work:

  • Team brainstorming and workshops, where a human facilitator could delegate role-differentiated perspectives and phase management to AI colleagues while retaining strategic control.
  • Early-stage design and product ideation, illustrated in the paper's own scenarios: designing collaboration tools for remote teams, and a "trustworthy mood-aware karaoke" experience in autonomous vehicles with privacy, simplicity, and local data processing constraints.
  • Interdisciplinary exploration by individuals working alone, where personas function as lightweight boundary objects that expose assumptions and tensions a single assistant would leave hidden.
  • Creativity-support tools generally, where structured Explore/Focus switching and explicit pause points could be adopted to counter the tendency of single-agent autoregressive models toward convergent, homogeneous output.

Industry relevance: the paper targets the dominant single-agent chat workflow for everyday ideation and benchmarks directly against ChatGPT-style prompting. For teams building AI assistants, it suggests that role differentiation and explicit cognitive-phase controls are design levers for perceived quality, novelty, and social presence — not just model capability.

Future Directions

  • Long-term team formation. The authors explicitly state their goal was to examine immediate coordination and convergence processes rather than long-term team formation, leaving open how durable multi-agent "colleagueship" develops over repeated sessions.
  • Ablating individual mechanisms. The authors note that an ablated multi-agent baseline could isolate specific mechanisms; determining which of persona diversity, dynamic turn-taking, Explore/Focus switching, or facilitation drives the reported benefits remains an open question.
  • Scaling and generalizing the evidence. The findings rest on 20 participants in short (approximately 10-minute per system) ideation sprints with self-chosen problems, so broader populations, longer sessions, and diverse task types are unexamined.
  • Rigorous analysis of idea structure. The paper sets up branching-ratio and conceptual-density measures of user contributions; the reported outcomes of those analyses are not included in the available content, and validating them against human expert judgment would be a natural next step.

Target Audience

This paper is most useful to HCI and CSCW researchers studying human–AI collaboration and creativity support; interaction designers and product teams building multi-agent or ideation tools; and LLM practitioners interested in how orchestration choices (turn-taking, persona ranking, history compression, human intervention points) translate into user-perceived outcomes. Readers looking for benchmark-style multi-agent performance results or detailed implementation code will find less here — the contribution is a designed, evaluated interaction paradigm rather than an engineering artifact.

Authors’ abstract

Most AI systems today are designed to manage tasks and execute predefined steps. This makes them effective for process coordination but limited in their ability to engage in joint problem-solving with humans or contribute new ideas. We introduce MultiColleagues, a multi-agent conversational system that shows how AI agents can act as colleagues by conversing with each other, sharing new ideas, and actively involving users in collaborative ideation processes. In a within-subjects study with 20 participants, we compared MultiColleagues to a single-agent baseline. Results show that MultiColleagues fostered stronger perceived social presence, and participants rated their outcomes as higher in quality and novelty, with more elaboration during ideation. These findings demonstrate the potential of AI agents to move beyond process partners toward colleagues that share intent, strengthen group dynamics, and collaborate with humans to advance ideas.

Read the original paper