Skip to content
AI.info

Research

ShareChat: A Dataset of Chatbot Conversations in the Wild

Overview Research area: Natural Language Processing, specifically large-scale dataset construction and evaluation of conversational LLM chatbots across commercial platforms. Technical level: Intermedi

ShareChat: A Dataset of Chatbot Conversations in the Wild
arXiv
2512.17843
Published
2025-12-19
Authors
Yueru Yan, Tuc Nguyen, Bo Su, Melissa Lieffers, Thai Le

AI summary

Overview

  • Research area: Natural Language Processing, specifically large-scale dataset construction and evaluation of conversational LLM chatbots across commercial platforms.
  • Technical level: Intermediate. The dataset statistics and case studies are accessible, but some analyses (toxicity correlation tests, latency correlation statistics, intention-classification pipelines) assume familiarity with standard NLP evaluation practice.
  • Scope in one sentence: The paper introduces ShareChat, a 142,808-conversation corpus drawn from publicly shared URLs on five commercial chatbot platforms that preserves platform-native features such as citations, thinking traces, and code artifacts, and demonstrates its value through three case studies.

What This Paper Is About

Existing academic benchmarks evaluate LLMs through uniform, text-only interfaces, which erases the fact that real users interact with specific products that differ in design, features, and safety policy. The authors address this by collecting a large multi-platform corpus of conversations that users themselves chose to share publicly, retaining the metadata and affordances each platform exposes. The goal is to enable evaluation and research questions that single-platform or stripped-down corpora cannot support.

Key Contributions

  1. A share-URL-based corpus across five platforms simultaneously. ShareChat contains 142,808 conversations and 660,293 turns from ChatGPT, Perplexity, Grok, Gemini, and Claude, collected from URLs users shared publicly. The authors describe this post-hoc public-sharing access model as complementary to consent-based gateways such as WildChat and opt-in plugins such as ShareLM.

  2. Longer, denser, and more linguistically diverse conversations than prior corpora. ShareChat averages 4.62 turns per conversation and 1,115.30 chatbot tokens per turn, spanning 95 languages. For comparison, the paper reports WildChat averaging 1.61 turns with 519.73 chatbot tokens and covering 76 languages, and LMSYS-Chat-1M averaging 2.02 turns and covering 65 languages.

  3. Preservation of native platform affordances. Each conversation record retains platform-specific metadata where available, including source citations, thinking traces, code artifacts, analysis blocks, turn timestamps, model version identifiers, and view/share counts. Per the paper's feature table, citations appear for Perplexity and Grok, thinking blocks for Grok and Claude, code and analysis blocks for Claude, turn timestamps for ChatGPT and Grok, model version for ChatGPT, Grok, and Gemini, and view/share counts for Perplexity.

  4. Three case studies demonstrating evaluative utility. Conversation completeness analysis, source grounding analysis, and timestamp/latency analysis show platform-dependent patterns the authors argue are inaccessible to single-platform or affordance-stripped corpora.

Main Findings

  • Longer interactions. ShareChat conversations average 4.62 turns, and the median turn count is 2.0, compared with 1.0 for Alpaca, Dolly, LMSYS, and WildChat in the paper's Figure 1 comparison.
  • Denser model responses. Mean chatbot output is 1,115.30 tokens across the multi-platform corpus, versus 519.73 for WildChat. Mean user turn length is 135.04 tokens.
  • Platform imbalance. ChatGPT contributes the largest share at 102,740 conversations and 542,148 turns (5.28 average turns), followed by Perplexity (17,305 conversations, 24,378 turns), Grok (14,415 conversations, 53,094 turns), Gemini (7,402 conversations, 36,422 turns), and Claude (946 conversations, 4,251 turns). The Limitations section states ChatGPT accounts for over 70% of conversations and Claude less than 1%.
  • Language distribution. English accounts for 61.8% of conversations, Japanese 18.0%, and every remaining language less than 3%. The corpus covers 95 languages; per-platform counts are 84 (ChatGPT), 47 (Perplexity), 54 (Grok), 41 (Gemini), and 18 (Claude).
  • Lower observed toxicity than prior corpora. Overall user toxicity is 2.9% via OpenAI Moderation, compared with 6.05% for WildChat on its own corpus and 3.08% for LMSYS-Chat-1M re-evaluated with the same pipeline. Overall LLM toxicity via OpenAI is 3.2%, below WildChat's 5.18% and comparable to LMSYS-Chat-1M's 4.12%.
  • User and model toxicity correlate within platforms. At the conversation level, Spearman correlations between mean user-turn toxicity and mean LLM-turn toxicity fall in the range 0.45–0.66 under Detoxify and 0.45–0.63 under OpenAI Moderation, with Bonferroni-corrected p < 10⁻⁶⁵. The paper notes Claude's rates rest on a smaller subsample with wider uncertainty.
  • Topic specialization by platform. Seeking Information is the most prevalent intent overall at 39.6%, followed by Other/Unknown at 19.0% and Technical Help at 12.4%. Perplexity concentrates 63.3% of requests on Seeking Information (with Self-Expression at 1.2% and Multimedia at 0.7%); Claude has the highest Technical Help share at 17.0%; Seeking Information rises across ChatGPT (34.2%), Gemini (38.3%), and Grok (42.8%); Grok shows the highest Multimedia usage at 4.4%; Writing stays near 10–11% across ChatGPT, Gemini, and Grok.
  • Conversation completeness varies by platform. Claude achieves the highest full-completion rate at 88%, followed by ChatGPT at 83% and Gemini at 77%, while Perplexity shows the lowest at 67%. ChatGPT and Claude show a median of 2 extracted intentions per conversation; Gemini, Grok, and Perplexity show a median of 1.
  • Divergent citation strategies. Of Grok's 14,415 conversations, 8,242 (57.18%) include sources; of Perplexity's 17,305 conversations, 8,545 (49.38%) include sources. Grok's top cited domain is X with 40,624 references, nearly triple Wikipedia at 12,507. Perplexity's leading source is English Wikipedia at 2,919 citations, followed by Reddit at 1,642 and NIH at 1,339. Grok conversations typically cite fewer than 5 sources; Perplexity routinely supports deeper research sessions.
  • Divergent latency dynamics. Timestamps cover 99.97% of ChatGPT turns and 100% of Grok turns. ChatGPT shows longer mean user response times than Grok (1,580 vs. 931 seconds) but shorter mean LLM response times (18.4 vs. 24.6 seconds). Within conversations, ChatGPT shows a negative correlation between turn index and LLM response time (Pearson r = −0.238; Spearman ρ = −0.427), while Grok shows a positive correlation (Pearson r = 0.315; Spearman ρ = 0.254). The authors caution that ChatGPT's window (May 2023–Aug 2025) conflates platform behavior with four model upgrades, whereas Grok's window (Dec 2024–Oct 2025) is more stable.
  • Response length weakly predicts user processing time. Binned mean user interval rises monotonically with LLM response length, plateauing beyond roughly 4–6k characters, but raw Pearson correlations are near zero (ChatGPT r = 0.034; Grok r = 0.021).
  • Toxicity detection methods disagree by role. Aggregated across platforms, Detoxify produces higher rates on user turns at four of five platforms, while OpenAI Moderation produces higher rates on LLM turns at four of five. Perplexity shows the lowest toxic-conversation prevalence at roughly 3% to 4%, while Grok shows the highest at 11% to 13%.

Methodology in Plain English

The authors discovered shared conversation URLs by querying the Internet Archive (Wayback Machine) for URL patterns matching each platform, such as chatgpt.com/share/* for ChatGPT and perplexity.ai/search/* for Perplexity. They note that shared URLs appear on social media sites such as X, Reddit, and Discord, but monitoring those ecosystems exhaustively does not scale. For each shared page, they used automated browser control with Selenium to render the page, click through elements that hide content behind interaction (for example, Claude's thinking bar), and parse the result into structured JSON containing ordered turns, prompts, responses, and platform metadata.

They stripped personally identifiable information using Microsoft Presidio across names, phone numbers, emails, credit cards, driver licenses, and URLs in multiple languages, then audited the removal with an LLM-based evaluator (GPT-OSS-120B, per the acknowledgments). Collection was conducted under IRB approval.

For the three case studies, they applied off-the-shelf and open models: Llama-3.1-8B-Instruct for topic classification into 24 categories under few-shot prompting with 4-bit quantization and batch size 32, and Qwen3-8B for a three-stage completeness pipeline that extracts user intentions, classifies each as complete, partial, or incomplete, and computes a weighted conversation-level score. Human validation used 100 English user messages for the topic classifier (Cohen's κ = 0.701 between two annotators; classifier accuracy 82.0%, macro-F1 76.9% against adjudicated gold labels) and 133 intentions from 64 English conversations for completeness (96.24% of extracted intentions judged correct; LLM fulfillment accuracy 87.40%, F1 89.90%).

Why This Matters

Impact on research. The paper argues that forcing every model into one neutral interface distorts how users actually behave, and that plain-text logs discard the structural signals (citations, thinking traces, timestamps) that shape prompting and evaluation. ShareChat offers stratified, platform-aware evaluation and longer multi-turn contexts for studying phenomena like reliability degradation over extended conversations, which the authors note is hard to study in corpora that are predominantly single-turn.

Real-world applications.

  • Building and auditing retrieval-augmented generation systems, using the preserved citation metadata to assess citation accuracy, source diversity, and grounding quality from authentic system outputs rather than synthetic setups.
  • Latency-aware product evaluation, using turn-level timestamps to model interaction pacing and user-facing performance over extended sessions.
  • Studying conversational breakdown detection and when unresolved user intentions persist across turns, supported by the completeness labels.
  • Analyzing how platform positioning shapes user intent, for example the contrast between Perplexity's search-oriented usage and Claude's technical-help usage.

Industry relevance. The corpus captures naturalistic task distributions that each commercial product actually faces, which the authors present as a basis for benchmarking models against real user intents instead of uniform synthetic prompts. The findings that users treat platforms as specialized tools, and that latency trends diverge in opposite directions between ChatGPT and Grok as conversations lengthen, give concrete test cases for how architectural choices affect user-facing behavior.

Future Directions

  • Expanding under-represented platforms. The authors report that collection is ongoing and that future releases will expand minority platforms; the current skew toward ChatGPT and away from Claude is partly a sampling artifact they cannot disentangle from real-world sharing behavior.
  • Collecting missing metadata. Claude creation timestamps were not captured in this iteration and are planned for future versions, and per-platform collection windows differ, meaning cross-platform comparisons partly conflate platform with collection era.
  • Extending validation beyond English. Both human validations were conducted on English conversations, leaving classifier and completeness performance on lower-resource languages unquantified.
  • Exploring additional research directions. The authors name conversational breakdown detection, cross-platform transfer learning, and platform-aware evaluation frameworks as directions the dataset supports but that this paper does not exhaustively analyze, along with longitudinal study as platform features evolve.

Target Audience

Researchers and practitioners who build, train, or evaluate conversational LLM systems will benefit most, particularly those working on multi-turn dialogue, RAG and citation grounding, platform-aware benchmarking, and reinforcement learning from conversation-level feedback. Dataset builders interested in collection ethics and PII handling will also find the access-model discussion relevant, as will product and evaluation teams that need realistic task distributions rather than synthetic prompts. Readers seeking transferable modeling techniques will find less here, since the paper's focus is corpus construction and descriptive analysis rather than new model architectures.

Authors’ abstract

By evaluating Large Language Models (LLMs) through uniform, text-only interfaces, current academic benchmarks obscure how the unique designs and affordances of distinct commercial platforms shape real-world user behavior and system performance. To bridge this gap, we present ShareChat, the first large-scale corpus of 142,808 conversations (660,293 turns) collected from publicly shared URLs on ChatGPT, Perplexity, Grok, Gemini, and Claude. ShareChat preserves native platform affordances, including citations, thinking traces, and code artifacts, across 95 languages and the period from April 2023 to October 2025, complementing existing corpora that homogenize these interactions. To demonstrate the dataset's evaluative utility, we present three case studies: a conversation completeness analysis assessing cross-platform differences in intent satisfaction, a source grounding analysis comparing citation strategies between search-augmented systems, and a temporal analysis revealing divergent response latency dynamics. Together, these analyses demonstrate research questions that are inaccessible to single-platform or stripped-affordance corpora. The dataset is publicly available.

Read the original paper