Research
Expectation, Backlash, Recovery, and Excitement: How Model Releases Shape Reddit Perceptions of Conversational AI Systems
Overview Research area: Computational social science / Natural Language Processing — large-scale analysis of social media discourse about conversational AI systems (CAISes), specifically how model rel
- arXiv
- 2608.24654
- Published
- 2026-08-25
- Authors
- Vahid Rahimzadeh, Yury Zhauniarovich, Savvas Zannettou
AI summary
Overview
Research area: Computational social science / Natural Language Processing — large-scale analysis of social media discourse about conversational AI systems (CAISes), specifically how model releases change user sentiment and discussion themes on Reddit.
Technical level: Intermediate. The paper combines LLM-based text pipelines (mention extraction, concept induction, sentiment classification) with time-series sentiment analysis; readers need some familiarity with sentiment classification metrics (Krippendorff's alpha, weighted F1, FDR correction) but the framing is largely conceptual.
Scope: A long-term study of 668,063 Reddit posts from 20 subreddits (October 2022 to December 2025) measuring post-release sentiment and thematic shifts around model releases from nine providers.
What This Paper Is About
Most prior work on how people perceive conversational AI systems captures a single static snapshot through surveys or aggregate social media analysis, even though these systems change constantly through new model releases, feature updates, safety interventions, and access-policy shifts. This paper treats each model release as a recurring, provider-specific "intervention" and asks whether user sentiment and discussion themes shift measurably around those events. The goal is to determine whether reactions like the "#Keep4o" backlash are exceptional cases or reflect a broader pattern of intervention-sensitive user perception across providers.
Key Contributions
-
A large-scale, release-centered Reddit dataset. The authors assemble 668,063 posts from 20 subreddits covering nine providers (October 2022 to December 2025), drawn from 849,603 submissions after excluding 181K removed or deleted posts, with an accompanying hierarchical model taxonomy and validated mention-extraction pipeline.
-
A four-level model taxonomy and validated LLM mention extractor. The taxonomy uses a
<provider, family, generation, tier>4-tuple built from the top 250 entries of the LLM Arena text-to-text leaderboard (as of November 28, 2025), yielding 189 entries spanning 9 providers, 15 model families, 50 generations, and 39 tier variants. The extractor was validated on 397 manually annotated posts containing 601 model mentions, with F1 scores of 1.00 for provider, 0.956 for family, 0.967 for generation, and 0.976 for tier. -
An adapted LLooM-based concept pipeline producing an auditable "Canonical Concept Bank" (CCB). The 363K posts with model mentions were split into 365 non-overlapping batches, producing 5–30 raw concepts per batch (mean 13.6) and 4,979 raw concepts, which LLM-assisted deduplication reduced to 236 canonical concepts with explicit inclusion criteria.
-
Provider- and release-level findings on sentiment and themes. The paper links aggregate sentiment deltas to specific concepts (e.g., Expectation Gap, Feature Removal Frustration, Integration & Automation, Bias & Censorship Concerns) to explain why each release shifted perception, not just that it did.
Main Findings
-
Overall Reddit discussion skews critical: negative accounts for 38% of posts, compared with 20% positive and 42% neutral, meaning critical discussion is nearly twice as common as favorable discussion.
-
OpenAI and Google shift from neutral toward negative over time. For OpenAI, neutral sentiment declines from roughly 45% in late 2022 to 33% by late 2025 (−0.117 pp/week) while negative rises from 27% to 46% (+0.116 pp/week), with little long-term change in positive sentiment (−0.008 pp/week). Google follows a similar but less steep trajectory: negative rises +0.101 pp/week (vs. +0.116 for OpenAI) and neutral declines −0.078 pp/week (vs. −0.117). Google also shows short-lived positive convergence around the Gemini 2.5 and Gemini 3.0 releases.
-
Anthropic is the main exception. From late 2023 to late 2025, positive sentiment rises from roughly 20% to 35% while negative sentiment falls from 49% to 27% — the only provider with a long-term shift from negative toward positive.
-
Anthropic has the clearest favorable release profile. On average, Anthropic releases are correlated with a +3.0/−4.4 percentage-point (pp) change in positive/negative sentiment, while OpenAI shows a smaller average change of +0.9/−0.5 pp, and Google, Meta, and xAI remain closer to the origin. DeepSeek shows the least favorable average, largely driven by DeepSeek R1.
-
DeepSeek R1 is the most negative release outlier, with a −8.1/+16.0 pp shift in positive/negative sentiment, and single-mention posts growing from 119 in the pre-window to 2,253 in the post-window (roughly 19-fold). The mechanism is capacity failure rather than model disappointment: rising concepts include Feature Expectation Gap, Trust & Reliability, and Reliability Availability Frustration, but Model Comparison & Evaluation carries the strongest positive sentiment share among rising concepts (technical capability and price-performance), alongside Self-Hosting & Privacy and Bias & Censorship Concerns.
-
GPT-4o is the clearest expectation-gap release. Expectation Gap rises +19.4 pp after release — the largest concept-level surge in the entire data — driven by the distance between the "omni" demo and the capabilities actually available at launch (real-time voice, image output, video/camera access, interruption-based conversation). Access & Availability and Pricing & Limits also rise, as does Emotional Anthropomorphic Engagement (partly around voice interaction and the Sky voice controversy), while older GPT-4-era themes (Productivity Assistant, Exploratory Experimentation, Trust & Reliability) decline. The paper links this to roughly 14 months without a major OpenAI flagship release.
-
GPT-5 is the strongest adverse OpenAI release, with a −3.5/+10.3 pp shift. Rising concepts are all negatively discussed, spanning relational attachment (Feature Removal Frustration, Forced Upgrade Frustration) and instrumental dependency (Performance Decline Perception, Version Downgrade Expectation, Inconsistent Behavior Frustration).
-
GPT-5.1 is a partial recovery, not an enthusiastic one. Narrower complaints still rise, but their magnitudes are smaller than the large declines in GPT-5-era backlash concepts (Feature Removal Frustration, Trust & Reliability, Perceived Limitations & Censorship, Emotional Anthropomorphic Engagement), shifting discussion away from platform-level revolt.
-
Claude 4 is the most favorable release-window shift, correlating with +5.3 pp positive and −10.5 pp negative sentiment. The driver is Claude Code: coding concepts (Integration & Automation, Coding Productivity, Productivity & Coding Use) dominate rising concepts, with the first two emerging from zero pre-window coverage. Claude Code mentions rise from 13/100 in pre-window falling concepts to 64/100 in post-window rising concepts. Pricing & Limits still rises, so access grievances persist.
-
Claude 4.1 is the only Anthropic release with an unfavorable shift (−0.5/+6.4 pp).
-
Grok 3 shows divided reception — the only release where both positive and negative sentiment increase meaningfully (+4.2/+8.0 pp), with single-mention posts growing from 331 to 1,839. Rising concepts include negative ones (Feature Expectation Gap, Reliability Skepticism, Reliability Availability Frustration, Inconsistent Behavior Frustration) alongside Personalization and Anthropomorphism. The authors attribute the distinctive pattern to politicization: Elon Musk appears across rising concepts as the source of the "truth-seeking AI" expectation, an object of system-prompt and censorship debates, and a political figure.
-
Concept shifts are robust. Of the 60 concept changes discussed, 58 remain significant after FDR correction; the two exceptions are DeepSeek R1 concepts affected by its small pre-release window. Recomputing with windows down to 0.6× the provider-specific size yields Pearson r ≥ .73 with reported values.
Methodology in Plain English
The researchers started by collecting public Reddit dumps covering all submissions from October 2022 to December 2025. To decide which communities to study, they reviewed the LLM Arena text-to-text leaderboard (as of November 28, 2025), manually kept providers and products with public user-facing interfaces, and searched Reddit manually for corresponding subreddits — arriving at 20 subreddits. After filtering to those subreddits and removing deleted or removed posts, the corpus came to 668,063 posts.
To find model mentions, they built a four-level taxonomy of model names from the top 250 leaderboard entries and used an LLM to map each post to taxonomy levels, deliberately forbidding the model from inferring models that are not explicitly mentioned. They validated this on hand-annotated posts. This produced 505,250 final model mentions across 363,546 posts.
For perception concepts, they adapted LLooM, a qualitative-research-inspired LLM framework. It extracts excerpts, summarizes them, clusters them into interpretable concepts with natural-language inclusion criteria, and scores each post on a five-point Likert scale. Because the corpus was split into 365 batches for context-window limits, the authors added an LLM-assisted deduplication step that merged exact-name duplicates and near-duplicates into a 236-concept Canonical Concept Bank. Posts were rescored against the batch's concepts, with a concept counted as present only when the verdict was "Strongly Agree" or "Agree."
Sentiment was classified per single-mention post as positive, neutral, or negative by an LLM annotator validated against 150 human-annotated posts labeled by three annotators (three-way nominal Krippendorff's alpha of 0.71; 0.82 weighted F1 against the majority label).
The analysis treats each model release as an intervention point, comparing symmetric windows before and after the release date, using provider-specific window sizes because providers differ in release frequency. Two deltas are computed per release: the change in the share of positive/neutral/negative posts, and the change in prevalence of each concept, focusing on concepts with at least 10 posts in the pre- or post-release window. Only six providers with more than 10K model mentions enter the release-window analysis, and only posts mentioning a single model with a text body are used.
Why This Matters
Impact on research: The paper argues that static snapshots of user perception are insufficient for continuously evolving systems. It proposes treating model releases as recurring socio-technical events and provides a replicable, auditable pipeline (taxonomy-grounded mention extraction, LLooM-derived concepts, symmetric-window comparison) that scales a qualitative-style inquiry to hundreds of thousands of posts. It also extends beyond the single-event #Keep4o case study, which was based on 1.5K Twitter/X posts, to a systematic multi-provider comparison.
Real-world applications:
- Release communication and launch planning: Aligning capability announcements with what users can actually access at launch (the GPT-4o Expectation Gap) and pairing releases with concrete workflows that make practical value immediately visible (the Claude 4 / Claude Code case).
- Model deprecation and transition policy: Keeping established models available during transitions where possible, given that GPT-5 backlash centered on Forced Upgrade Frustration and Feature Removal Frustration.
- Capacity and infrastructure provisioning: Preparing for release-driven demand spikes, as seen when DeepSeek R1 single-mention posts grew roughly 19-fold and the dominant concepts were reliability and availability frustrations.
- Ongoing perception monitoring: Tracking user reactions after release and responding to emerging concerns rather than treating release reception as a one-time event, since concepts continue to shift in the post-release window.
Industry relevance: The findings map directly onto product-management decisions — subscription and usage-limit design (Pricing & Limits rises even around the well-received Claude 4), safety and refusal policies (Bias & Censorship Concerns around DeepSeek R1), and the reputational entanglement of a provider's leadership with model perception (Grok 3). The paper also notes that providers must account for Anthropic's distinctive positive trend versus OpenAI's and Google's neutral-to-negative erosion.
Future Directions
-
Multimodal and non-text signals. The authors explicitly limit analysis to text-based Reddit posts, excluding image-only or video-only posts that may carry relevant perception information. A natural extension is analyzing multimedia content alongside text.
-
Beyond Reddit and beyond early adopters. The authors caution that Reddit communities likely skew toward technically engaged and early-adopting users and that some posts may be AI-generated or AI-assisted, which they cannot reliably identify. Future work could compare platforms and validate against general-population samples.
-
Finer temporal resolution around releases. The symmetric-window design may smooth short-lived reactions — the authors note Gemini 2.5 shifts persisted only several weeks while the average Google analysis window spans 104 days. Shorter or event-triggered windows could capture immediate reactions that longer windows absorb.
-
Reducing concept-pipeline redundancy and improving automation. Because LLooM ran on 365 independently analyzed batches, the resulting Canonical Concept Bank may contain more concepts than a single-joint-batch analysis would produce. Joint or hierarchical concept induction is an open methodological question, as is quantifying residual error from LLM-based mention extraction, sentiment labeling, and concept assignment.
Target Audience
This paper is most useful to researchers in computational social science and NLP who study public perception of AI systems, to industry teams responsible for model release strategy, deprecation policy, pricing and limits, and developer-product bundling, and to analysts tracking provider reputation over time. It will also interest readers who followed the "#Keep4o" case and want to know whether such backlash is exceptional or part of a general pattern of intervention-sensitive user perception. The methodology sections will be most accessible to readers already comfortable with LLM-based annotation pipelines and social media sentiment analysis.
Authors’ abstract
Conversational AI systems (CAISes) continuously change through model releases, feature updates, safety interventions, and access-policy shifts, yet user perceptions are often studied as static snapshots. We conduct a long-term, large-scale analysis of Reddit discussions to examine how users perceive CAIS model release interventions across providers. By combining sentiment classification and thematic concept analysis, we show that CAIS perceptions are dynamic and intervention-sensitive. Anthropic exhibits the clearest positive release profile through Claude Code and product-model fit, OpenAI shows backlash-and-recovery dynamics around GPT-5 and GPT-5.1, Grok-3 is shaped by provider identity and political discourse, and DeepSeek-R1 combines engineering praise with concerns about censorship, access, and reliability. These findings show that model releases are not merely technical updates, but user-facing interventions that reshape sentiment, expectations, and public discussion.