Research
Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs
Overview Research area: Natural Language Processing, specifically unsupervised text clustering and the use of large language models as semantic validators rather than embedding generators. Technical l

- arXiv
- 2604.07562
- Published
- 2026-04-08
- Authors
- Tunazzina Islam
AI summary
Overview
Research area: Natural Language Processing, specifically unsupervised text clustering and the use of large language models as semantic validators rather than embedding generators.
Technical level: Intermediate. The framework itself is conceptually simple and modular, but the paper assumes familiarity with clustering pipelines (TF-IDF, UMAP, HDBSCAN), topic models (LDA), sentence embeddings (SBERT, BERTopic), and standard cluster-quality metrics.
Scope in one sentence: The paper proposes and evaluates a three-stage LLM reasoning layer that verifies, merges, and labels the output of any unsupervised clustering algorithm, tested on vegan discourse collected from X and Bluesky.
What This Paper Is About
Unsupervised clustering is widely used to discover themes in large text collections, but its outputs are often incoherent, redundant, or hard to interpret, and there is no labeled data available to check them. This paper treats clusters produced by an existing algorithm as hypotheses and uses a large language model as a semantic judge that decides whether each cluster is internally coherent, whether two clusters overlap, and what human-readable label each cluster should carry. The goal is a post-hoc refinement layer that improves the meaning and interpretability of unsupervised structure without needing any gold-standard annotations.
Key Contributions
- A reasoning-based framework that validates and refines unsupervised semantic structure by using LLMs as semantic judges, with three explicit reasoning stages: coherence verification, redundancy adjudication, and label grounding.
- A systematic evaluation, including human validation with two expert annotators, comparing reasoning-based refinement against embedding-only approaches (HDBSCAN without refinement and SBERT-based refinement) and against LDA, BERTopic, SBERT assignment, TopicGPT, Llama 3.2, Mistral Large 2, and GPT-4o assignment baselines.
- A two-stage label grounding process that first generates candidate labels per cluster and then consolidates semantically similar labels into a smaller set of final labels, providing fully unsupervised interpretability.
- Release of cross-platform datasets and evaluation resources for future work on interpretable and reliable unsupervised text analysis (public repository linked in the paper).
Main Findings
- Cluster quality on X: LLM-based refinement reaches a Silhouette score of 0.674 versus 0.122 for raw HDBSCAN and 0.156 for SBERT-based refinement. Cluster count drops from 359 (HDBSCAN) to 250 (SBERT-rf) and 232 (LLM-rf). On the Davies–Bouldin Index (lower is better), SBERT-rf achieves 0.569 versus 0.635 for LLM-rf and 2.322 for HDBSCAN.
- Cluster quality on Bluesky: LLM-rf reaches a Silhouette score of 0.979 versus 0.052 for SBERT-rf and -0.017 for HDBSCAN, and the best Davies–Bouldin Index at 0.227 versus 0.282 for SBERT-rf and 2.739 for HDBSCAN. Cluster counts are 37 (HDBSCAN), 34 (SBERT-rf), and 36 (LLM-rf).
- Intra-cluster coherence: On X, HDBSCAN and LLM-refinement both show medians around 0.60 cosine similarity between text pairs and produce many clusters with cohesion in the 0.9–1.0 range, which SBERT-refinement rarely achieves. On Bluesky, HDBSCAN and LLM-refinement show medians around 0.38–0.40 versus roughly 0.35 for SBERT-refinement, with most clusters in the 0.3–0.5 range and HDBSCAN/LLM-refinement reaching around 0.87–0.89.
- Statistical testing: For X, a Kruskal–Wallis test shows a significant difference among methods (H = 16.187, p < 0.001). Mann–Whitney U post-hoc comparisons show SBERT-refinement significantly outperformed both HDBSCAN (p < 0.001) and LLM-refinement (p < 0.01), while the difference between HDBSCAN and LLM-refinement was not significant (p = 0.4808). For Bluesky, neither the Kruskal–Wallis test nor any pairwise comparison reached significance.
- Human-aligned assignment: GPT-4o achieved 78.4% assignment accuracy on X and 89.8% on Bluesky, the highest of the compared systems. Other results: Mistral Large 2 (71.6% X, 71.8% Bluesky), TopicGPT (72.8% X, 68.4% Bluesky), Llama 3.2 (66.6% X, 60.0% Bluesky), SBERT (56.2% X, 53.6% Bluesky), BERTopic (38.7% X, 42.4% Bluesky), and LDA (30.4% X, 36.2% Bluesky).
- Human evaluation reliability: Two expert annotators judged 500 randomly selected tweets from X and 500 posts from Bluesky (random seed 42), reaching an inter-annotator agreement of 0.82 by Cohen's Kappa.
- Final label sets: After the framework's steps, 14 labels were generated for X and 22 for Bluesky.
- Temporal and volume robustness: Restricting X to its densest 28-day window (2020.01.14–2020.02.10) and down-sampling Bluesky to match post counts improved cluster quality for both platforms and sharpened the contrast: X showed higher cohesion (Silhouette 0.60 vs. 0.08) and tighter separation (Davies–Bouldin 0.61 vs. 1.14). A chi-square test on the theme × platform table confirmed a non-random association (χ² = 80.0, df = 23, p ≈ 3 × 10⁻⁸).
- Platform discourse differences: X leans toward informational and aspirational themes such as advocacy and daily motivation, while Bluesky emphasizes conversational and satirical content including humor and sociopolitical critique. In the UMAP projection of 1,000 texts (500 per platform), Bluesky posts dominate the "Vegan & Sustainable Products" region, while both platforms contribute to "Food & Recipes," "Ethical Lifestyle," and "Animal Rights."
- Error patterns: GPT-4o misclassified short personal posts when advocacy, lifestyle, and ethics themes overlapped; food mentions triggered dining-experience labels for ads or generic content; on Bluesky, abstract themes such as social and ethical commentary were confused by implicit moral cues, sarcasm, or vague language; and keyword over-reliance caused mentions of skincare or donations to be labeled as ethical consumption or advocacy.
Methodology in Plain English
The researchers collected vegan-related posts from two platforms. From X they gathered 330,464 tweets from 204,670 users via the Twitter streaming API between October 2019 and February 2020, and after noticing 63,751 suspended users, they sampled at the user level to obtain a final set of 20,000 tweets from 275 users. From Bluesky they used a firehose pipeline in June 2025 to collect 13,032 English posts, of which 1,752 are unique. Both collections used a keyword filter (vegan, veganism, plantbased, meatfree, and others listed in an appendix).
Texts were clustered with a standard pipeline: TF-IDF vectorization, MaxAbsScaler normalization, Truncated SVD for dimensionality reduction, UMAP for further reduction, and HDBSCAN for clustering, with DBCV used to select parameters. The best settings were min_cluster_size 15 and min_samples 3 for X (DBCV 0.35) and min_cluster_size 10 and min_samples 2 for Bluesky (DBCV 0.53).
Then the LLM takes over. For each cluster, GPT-4o generates a short summary from the top 5 documents closest to the cluster centroid (robustness was checked at k = 3, 5, and 7). The LLM then judges whether that summary is actually supported by those representative texts; if not, the cluster is discarded as incoherent. Next, cluster summaries are embedded with Sentence-BERT, cosine similarity is computed between them, and clusters above a similarity threshold are merged. The merge threshold was grid-searched over {0.75, 0.80, 0.85, 0.90} using Silhouette score, Davies–Bouldin Index, and cluster count, and 0.85 was selected. Labels are then generated per cluster, and labels with SBERT similarity above 0.85 are grouped and consolidated by the LLM into a single final label. Finally, the LLM reassigns each individual document to its best-fitting consolidated label. All LLM steps used GPT-4o with default parameters.
Why This Matters
Impact on research. The paper reframes clustering as a proposal step and moves the validation burden from embedding geometry to explicit natural-language reasoning. It shows that improvements come from semantic consolidation, not from tuning representations, and it provides cross-platform datasets and human-validated evaluation resources. The paper is explicit that it does not propose a new topic model and that the statistically comparable behavior on X (p = 0.48) indicates gains are not driven by trivial pruning.
Real-world applications.
- Social media monitoring: turning noisy, rapidly shifting short-text streams into coherent, labeled theme sets without manual annotation.
- Advocacy and campaign strategy: identifying which platforms foreground which facets of a movement, for example prioritizing a community-driven environment for animal rights mobilization versus scrutinizing high-volume product promotions elsewhere.
- Consumer protection and greenwashing detection: surfacing where promotional or product-marketing themes dominate discourse so regulators can target scrutiny.
- Computational social science: enabling timely, interpretable analysis of public sentiment and emerging narratives at scale where annotation is costly or infeasible.
Industry relevance. The framework is algorithm-agnostic and can be inserted as a post-hoc refinement layer on top of existing unsupervised systems, which lowers adoption cost. The paper reports latency and API cost considerations: limiting prompt size to top-5 documents reduces latency and API cost, and cost details are provided in an appendix.
Future Directions
- Extend beyond English-language social media and the veganism domain, adapting prompts and evaluation criteria for new settings, since the framework itself is claimed to be domain-agnostic.
- Collect a 2025 X sample to remove the residual historical confound, because the X data predates Bluesky by about 5 years and the 28-day subsample narrows but does not eliminate this gap.
- Replace TF-IDF with dense encoders such as E5 or GTR-T5 to test whether dense initializations further improve cluster coherence in cross-platform discourse analysis.
- Incorporate document-level coherence metrics adapted for clustering outputs, alongside the current intra-cluster embedding similarity, Silhouette score, Davies–Bouldin Index, and human validation.
- Investigate and mitigate the human biases that LLMs may embed from their training data, an issue the paper states is not addressed, and consider whether fine-tuning (excluded here due to resource constraints) would help.
Target Audience
Researchers and practitioners in NLP and computational social science who work with unsupervised text analysis and want more interpretable, less redundant cluster structure without labeled data. It is also relevant to social media analysts, discourse researchers, and applied machine learning engineers who need to validate the output of existing clustering pipelines, and to readers interested in LLMs serving as semantic judges rather than as generators or embedders. Some familiarity with clustering metrics and embedding models is helpful, so the paper is most accessible at an intermediate level.
Authors’ abstract
Unsupervised methods are widely used to induce latent semantic structure from large text collections, yet their outputs often contain incoherent, redundant, or poorly grounded clusters that are difficult to validate without labeled data. We propose a reasoning-based refinement framework that leverages large language models (LLMs) not as embedding generators, but as semantic judges that validate and restructure the outputs of arbitrary unsupervised clustering algorithms.Our framework introduces three reasoning stages: (i) coherence verification, where LLMs assess whether cluster summaries are supported by their member texts; (ii) redundancy adjudication, where candidate clusters are merged or rejected based on semantic overlap; and (iii) label grounding, where clusters are assigned interpretable labels in a fully unsupervised manner. This design decouples representation learning from structural validation and mitigates common failure modes of embedding-only approaches. We evaluate the framework on real-world social media corpora from two platforms with distinct interaction models, demonstrating consistent improvements in cluster coherence and human-aligned labeling quality over classical topic models and recent representation-based baselines. Human evaluation shows strong agreement with LLM-generated labels, despite the absence of gold-standard annotations. We further conduct robustness analyses under matched temporal and volume conditions to assess cross-platform stability. Beyond empirical gains, our results suggest that LLM-based reasoning can serve as a general mechanism for validating and refining unsupervised semantic structure, enabling more reliable and interpretable analyses of large text collections without supervision.