Research
A Multimodal Framework for Depression Detection during Covid-19 via Harvesting Social Media: A Novel Dataset and Method
Overview Research area: Mental health informatics / social media mining — multimodal machine learning applied to depression detection, with a specific focus on the Covid-19 period. Technical level: In

- arXiv
- 2511.00424
- Published
- 2025-11-01
- Authors
- Ashutosh Anshul, Gumpili Sai Pranav, Mohammad Zia Ur Rehman, Nagendra Kumar
AI summary
Overview
- Research area: Mental health informatics / social media mining — multimodal machine learning applied to depression detection, with a specific focus on the Covid-19 period.
- Technical level: Intermediate. The paper combines established NLP techniques (LDA topic modeling, lexicon dictionaries, OCR) with CNN transfer learning and an ensemble classifier; readers should be comfortable with basic machine learning and neural network terminology.
- Scope: The paper presents a multimodal feature-extraction and ensemble-classification framework for detecting whether a social media user is depressed, validated on an existing benchmark (the Tsinghua Dataset) and on a newly constructed Covid-19 Twitter dataset. Note: the provided content is truncated at the dataset-description stage of the experimental section, so results beyond headline accuracies, ablations, comparison tables, and parameter-sensitivity analyses are not available in the text supplied.
What This Paper Is About
Existing depression-detection systems mostly read only the text of short social media posts, which often lack enough context and suffer from data sparsity. This paper argues that a social media user can be better described by combining several signals — the URLs they share, the text embedded inside the images they post, the raw visual character of those images, their topic and emotion profile, and their account-level activity — and proposes a framework plus a new Covid-19-era dataset to test that idea.
Key Contributions
- A novel multimodal framework that represents each user through external knowledge extracted from URLs, user-profile information, and historical posts, rather than tweet text alone.
- Two specific context-enrichment techniques: an extrinsic feature built by fetching and titling the web pages behind URLs in tweets, and optical character recognition (OCR) applied to text-bearing images posted in tweets — plus five distinct sets of modality-specific features describing a user.
- A new deep learning model, the Visual Neural Network (VNN), built on ResNet50 via transfer learning, which produces 128-dimensional image embeddings used to form a visual feature vector for prediction.
- A curated, well-labeled Covid-19 dataset of depressed and non-depressed Twitter users, collected during the outbreak and used both for evaluation and qualitative analysis, alongside evaluation on an additional benchmark dataset.
Main Findings
- Benchmark performance: On the publicly available Tsinghua Dataset, the model achieved an accuracy of 93.1%, which the authors state surpasses existing state-of-the-art methods. The abstract summarizes the overall gain over state-of-the-art methods on the benchmark dataset as 2%–8%.
- Covid-19 dataset performance: On the authors' novel Covid-19 dataset, the model achieved an accuracy of 91.7%, which the abstract describes as "promising results."
- Ensembling helps: Logistic Regression, XGBoost, and a Neural Network were each trained individually, but the ensemble of all three performed better than any individual classifier. The ensemble uses stacking with a max-voting rule (a user is predicted depressed if at least two of the three base models say so).
- Topic modeling detail: Latent Dirichlet Allocation was run with the number of topics T set to 15, and the model identified topics better when trained only on tweets posted by depressed users.
- Modality combination: Five feature families were extracted per user — visual, topic-based, emotional, depression-specific, and user-specific — producing a 274-dimensional feature vector per user.
- Emotional feature composition: Ten emotion intensities (Joy, Fear, Anger, Anticipation, Disgust, Trust, Surprise, Positive, Negative, Sadness) were extracted with NRClex, which uses an affect dictionary of approximately 27,000 words; emoji sentiment was encoded as a three-dimensional representation; Empath lexicon categories gave a 194-dimensional set reduced to 90 dimensions with PCA.
- Dataset statistics: The Tsinghua Dataset, after preprocessing, contains 2,522 depressed users and 4,533 non-depressed users, with a total of 2,059,326 tweets — 312,729 posted by depressed users and 1,746,597 by non-depressed users.
- Dataset reliability measure: Krippendorff's alpha was used to assess annotator agreement on the new dataset; the paper notes the coefficient ranges from 0 to 1 with higher values indicating stronger agreement, but the actual coefficient value is not reported in the available content.
- Not reported in the available content: The specific comparison methods, evaluation metric definitions, per-modality ablation numbers, parameter sensitivity results, and the qualitative case study of a sample user are referenced in the paper's structure but their results fall outside the truncated text.
Methodology in Plain English
The authors start by treating each Twitter user as a bundle of evidence rather than a bag of tweets.
First, they pull in outside context. Tweets are short, so when a user shares a link, the framework fetches the linked web page and appends the page title to the tweet text — the assumption being that a depressed user's shared links may point to depression- or anxiety-related material, and the page title sets the context of what the tweet is about.
Second, they mine the images. Many posted images contain text, so each image goes through OCR, which finds character outlines, groups them into blobs, assembles text lines, breaks them into words, and recognizes them with a classifier. That recovered text is appended to the associated tweet. For images with little or no text, they resize each image to 112×112 and feed it into the Visual Neural Network, a ResNet50 backbone adapted through transfer learning. Inside the VNN, the ResNet output (a 4×4×2048 matrix) is flattened to 32,768 values and passed through three dense blocks (512, 256, 128 units, each a dense layer plus dropout plus batch normalization); the 128-dimensional block output becomes the user's image embedding. Because users post a variable number of images, the embeddings are averaged per user, and a zero vector is used for users who posted no images.
Third, they profile the user's language and behavior. LDA assigns topic distributions to tweets. A dictionary-based approach scores ten emotion intensities plus a three-way emoji sentiment. The Empath lexicon library measures word-category usage and adds a dedicated depression-and-anxiety lexicon category. User-specific signals include tweet count over one month, follower count, friend count, favorite count, status count, and emotion extracted from the user's profile description.
All of these features are concatenated into a single 274-dimensional vector per user. Three classifiers — Logistic Regression, XGBoost, and a feed-forward Neural Network with dense blocks of 256, 128, and 64 and a binary cross-entropy loss — are trained on the same data, and their predictions are combined by max voting to give the final label.
For data, the authors reused the Tsinghua Dataset and built their own. The new dataset was crawled from the Twitter API for tweets posted between May 2021 and January 2022, spanning the second wave through the onset of the third wave. Depressed users were identified by a strict self-report rule: a tweet containing phrases such as "I'm diagnosed with depression," "I was diagnosed with depression," "I am diagnosed with depression," or "I've been diagnosed with depression" served as an anchor tweet, and all of that user's tweets within a one-month window around the anchor were collected. Non-depressed users were randomly selected, their one-month tweet histories between May 2021 and January 2022 gathered, and manually screened to confirm no depressive symptoms such as hopelessness or negative talk.
Why This Matters
- Research impact: The work pushes depression detection away from text-only pipelines toward multimodal, context-augmented representations, and it contributes a purpose-built Covid-19-era dataset that other researchers can build on. It also targets a recognized weakness — data sparsity in short posts — by pulling signal from URLs, OCR text, and images.
- Real-world applications:
- Early-warning or triage systems that flag social media users who may need mental health resources, at a scale no questionnaire-based survey could reach.
- Public-health monitoring during pandemics or other large-scale crises, since the paper documents a 27.6% rise in major depressive disorders and a 25.6% rise in anxiety disorders worldwide during 2020.
- Support for platforms that want to surface wellbeing resources to at-risk users, based on multimodal behavior rather than keyword matching alone.
- Clinical and epidemiological research that needs population-scale, low-cost mental health signal, given the paper's argument that surveys and interviews (for example, Google Forms using DASS-21) are time-consuming and expensive.
- Industry relevance: Social media platforms, digital mental health companies, and moderation or trust-and-safety teams all handle multimodal content (text, images, links, emojis) at scale; the paper's feature-engineering recipe and its ensemble of off-the-shelf classifiers are relatively inexpensive to reproduce compared with end-to-end large models.
Future Directions
- Close the evaluation gap: Report the full comparison against state-of-the-art methods, per-modality ablation results, parameter sensitivity, and the qualitative case study of a sample user, since only headline accuracies are visible in the available content.
- Privacy and ethics of URL and image mining: Fetching linked web pages and OCR-ing user images raises consent, storage, and surveillance questions that the framework does not yet address.
- Generalization beyond Twitter and beyond Covid-19: The model was built on Twitter data, and the new dataset is anchored to the May 2021–January 2022 period; whether the features transfer to other platforms and to non-pandemic periods is untested.
- Label quality at scale: The self-report anchor-phrase method and manual screening of randomly chosen non-depressed users do not scale easily, so automated or semi-automated labeling with reported inter-annotator agreement values would strengthen future datasets.
Target Audience
This paper is most useful to researchers and graduate students working on computational mental health, multimodal social media analysis, or NLP for wellbeing; to data scientists building content-moderation or user-support systems at social platforms; and to public-health researchers who need population-scale, low-cost signals of depression. Readers looking for clinical validation or deployment-ready tools should note that this is a methodological and dataset contribution evaluated on Twitter data, not a clinical study.
Authors’ abstract
The recent coronavirus disease (Covid-19) has become a pandemic and has affected the entire globe. During the pandemic, we have observed a spike in cases related to mental health, such as anxiety, stress, and depression. Depression significantly influences most diseases worldwide, making it difficult to detect mental health conditions in people due to unawareness and unwillingness to consult a doctor. However, nowadays, people extensively use online social media platforms to express their emotions and thoughts. Hence, social media platforms are now becoming a large data source that can be utilized for detecting depression and mental illness. However, existing approaches often overlook data sparsity in tweets and the multimodal aspects of social media. In this paper, we propose a novel multimodal framework that combines textual, user-specific, and image analysis to detect depression among social media users. To provide enough context about the user's emotional state, we propose (i) an extrinsic feature by harnessing the URLs present in tweets and (ii) extracting textual content present in images posted in tweets. We also extract five sets of features belonging to different modalities to describe a user. Additionally, we introduce a Deep Learning model, the Visual Neural Network (VNN), to generate embeddings of user-posted images, which are used to create the visual feature vector for prediction. We contribute a curated Covid-19 dataset of depressed and non-depressed users for research purposes and demonstrate the effectiveness of our model in detecting depression during the Covid-19 outbreak. Our model outperforms existing state-of-the-art methods over a benchmark dataset by 2%-8% and produces promising results on the Covid-19 dataset. Our analysis highlights the impact of each modality and provides valuable insights into users' mental and emotional states.