Skip to content
AI.info

Research

Prompting-in-a-Series: Psychology-Informed Contents and Embeddings for Personality Recognition With Decoder-Only Models

Overview Research area: Natural Language Processing, specifically personality recognition from text using large language models (LLMs), combining computational methods with psychological personality t

Prompting-in-a-Series: Psychology-Informed Contents and Embeddings for Personality Recognition With Decoder-Only Models
arXiv
2512.06991
Published
2025-12-07
Authors
Jing Jie Tan, Ban-Hoe Kwan, Danny Wee-Kiat Ng, Yan-Chai Hum, Anissa Mokraoui, Shih-Yu Lo

AI summary

Overview

Research area: Natural Language Processing, specifically personality recognition from text using large language models (LLMs), combining computational methods with psychological personality theory.

Technical level: Intermediate. The paper assumes familiarity with prompting strategies (zero-shot, few-shot, chain-of-thought), parameter-efficient fine-tuning methods (LoRA, QLoRA), and standard classification metrics.

Scope: The paper proposes and evaluates the PICEPR ("Prompting-in-a-Series") algorithm, a modular, dual-pipeline framework that uses decoder-only LLMs as personality feature extractors and generators for personality recognition, tested on the Big-5 Essays dataset and the MBTI Kaggle dataset.

What This Paper Is About

The core problem is that personality recognition from written text is a hard, under-explored task for LLMs: most existing approaches either require fine-tuning an encoder-only model or rely on heavy pre-inferencing to build reference libraries, which is slow and resource-intensive. The authors ask whether a decoder-only LLM can be split into a series of small, specialised modules that chain reasoning steps together, so that the model both classifies personality more accurately and produces psychologically meaningful explanations. Their goal is to build a modular "Prompting-in-a-Series" algorithm that either classifies personality directly through prompting, or generates augmented data used to fine-tune a fast encoder-based classifier.

Key Contributions

  1. A modular "Prompting-in-a-Series" framework (PICEPR) that reorganises personality classification into five specialised LLM components — Summary LLM (S), Mimic LLM (M), Psycho LLM (P), Classify LLM (C), and Vector LLM (V) — arranged across two pipelines: a Contents pipeline (S, P, C) that classifies via prompting alone, and an Embeddings pipeline (S, P, M, V) that transfers decoder-only LLM knowledge into a fine-tuned encoder model.

  2. A broad empirical study of open- and closed-source LLM backbones on personality tasks, comparing gpt-4o-2024-08-06 (gpt4o), gemini-1.5-flash (gemini), gpt-3.5-turbo-0125 (gpt3.5), Meta-Llama-3.1-8B-Instruct (llama), and Mistral-7B-Instruct-v0.3 (mistral) as text generators, plus embedding models text-embedding-ada-002 (gptada), text-embedding-004 (gemb), mistral-embed-23.12 (mistemd), paraphrase-MiniLM-L6-v2 (minilm), all-mpnet-base-v2 (mpnet), and sentence-bert-base-nli-mean-tokens (sbert).

  3. Analysis of whether decoder-only LLMs can eliminate the need for task-specific fine-tuning or referencing libraries, in comparison to training encoder-only models for personality tasks.

  4. A reporting of model reliability through an error-rate metric and statistical hypothesis testing (McNemar's test), alongside standard accuracy metrics, showing where LLMs fail to produce parseable structured output.

Main Findings

  • PICEPR sets a new state of the art with a 5–15% improvement over the compared baselines for personality recognition, according to the abstract.

  • The Contents pipeline's PICEPR variant (C^PR) wins for every backbone model on the Essays dataset. For example, with gpt4o on Openness, C^PR reached BA 0.6667, F1 0.6870, and RA 0.6680, compared with BA 0.5137, F1 0.6786, RA 0.5283 for the plain chain-of-thought baseline (C^R). The same pattern held for Conscientiousness (BA 0.7081, F1 0.7198, RA 0.7085), Extraversion (BA 0.6535, F1 0.6640, RA 0.6538), Agreeableness (BA 0.6793, F1 0.6876, RA 0.6781), and Neuroticism (BA 0.6721, F1 0.6786, RA 0.6721).

  • The two-shot PICEPR variant (C^PR2) performs worse than one-shot C^PR on the Essays dataset. For gpt4o, Openness BA dropped to 0.6352 and Conscientiousness BA to 0.7124, while Extraversion BA rose to 0.6654 relative to C^PR.

  • Fine-tuning the decoder-only Classify LLM (C^RT) degrades performance on the Essays dataset compared with prompting-based PICEPR. For gpt4o, Openness BA was 0.5765, Conscientiousness BA 0.5660, Extraversion BA 0.5305, Agreeableness BA 0.5518, and Neuroticism BA 0.5364.

  • On the Kaggle (MBTI) dataset, PICEPR also dominates. For gpt4o, the Sensing/Intuition dimension reached BA 0.8858, F1 0.9569, RA 0.9268 under C^PR, versus BA 0.7280, F1 0.8852, RA 0.8115 under C^R. Judging/Perceiving reached BA 0.8158, F1 0.7779, RA 0.8190; Extra/Introversion reached BA 0.8579, F1 0.7596, RA 0.8807; and the remaining dimension reached BA 0.9039, F1 0.9091, RA 0.9032.

  • Larger closed-source models outperform smaller open-source ones. gpt4o and gemini generally led, while llama and mistral showed the lowest scores (for instance, mistral on Kaggle Sensing/Intuition under C^R: BA 0.5827, F1 0.7650, RA 0.6455).

  • Error rates were near zero for most models but severe for gemini under fine-tuning. In Table III, gemini's error rate under C^RT was 0.609 on Essays and 0.750 on Kaggle; gpt3.5 showed small nonzero error rates (e.g., 0.002 for S, 0.008 for P, 0.017 for C^R, 0.020 for C^RT, 0.006 for C^PR, 0.027 for C^PR2 on Kaggle) while gpt4o, llama, and mistral recorded 0. The paper states gemini's C^RT results are not referable because the fraction of valid JSON outputs that could be parsed was less than one-third of the dataset.

  • The PICEPR approaches were statistically significant. McNemar's test gave p-value < α (α = 0.05) for both the Contents and Embeddings pipelines.

  • The Essays dataset is inherently harder. The paper notes it was collected in a formal setting, producing more formal text, which limits achievable accuracy; the Kaggle dataset has label imbalance exceeding a 6:4 ratio on several dimensions.

  • The Embeddings pipeline result tables (Table VI and Table VII) are referenced in the discussion but their numeric contents are not included in the available paper text, so specific Embeddings pipeline figures are not reported here.

Methodology in Plain English

The authors treat personality recognition as a multi-label classification problem rather than running four or five separate binary classifications, which avoids repeated training. They use two datasets: the Essays dataset (Big-5 model, 2,467 samples split into 1,578 training, 395 validation, and 494 test) and the Kaggle dataset (MBTI model, 8,675 samples split into 5,552 training, 1,388 validation, and 1,735 test). Splits follow Tan's train-validation-evaluation split method. Text preprocessing is deliberately minimal so the LLM sees raw or lightly processed text.

The PICEPR algorithm chains five LLM roles. The Summary LLM writes a short synopsis of the user's personality by combining given labels with textual evidence — labels are provided during training but removed at test time. The Psycho LLM scores 77 personality facets as binary 0 or 1 values in JSON. The Classify LLM then takes the summary plus the facet list and outputs the final personality label as JSON, using chain-of-thought so analysis precedes the answer. In the Embeddings pipeline, the Mimic LLM synthesises augmented social media-style text — both positive examples matching the labels and negative examples that do not — and the Vector LLM converts text into fixed-length embeddings. A multilayer perceptron with an input layer of size r + 77 (where r is the embedding length and 77 is the facet count) performs the final classification.

Training details: fine-tuning of the decoder-only Classify LLM uses QLoRA to save compute, optimising cross-entropy loss. The encoder-side training uses multi-binary cross-entropy with a sigmoid activation, switched to a focal loss variant to handle class imbalance, plus a contrastive loss that pulls same-personality pairs together and pushes different ones apart. Structured JSON output is requested via a structured output engine, with a JSON repair tool used for minor formatting problems; rows that cannot be repaired are excluded and counted as errors, which is how the error rate is computed. Evaluation uses Regular Accuracy, Balanced Accuracy, and F1 score.

Why This Matters

For research, the paper tests a specific idea: that decomposing a decoder-only LLM into a series of narrow, chained prompting modules can beat both direct prompting and fine-tuning, and can also synthesise training data for lightweight encoder models. It provides an empirical baseline landscape across five text-generating models and six embedding models, plus a reliability analysis (error rates) that is often missing from LLM classification papers. It also connects personality psychology directly to model design through the 77-facet Psycho LLM.

Real-world applications the paper and its framing point toward:

  • Recommendation systems that adapt to user personality for more personalised suggestions.
  • Virtual assistants and human-computer interaction that adjust tone and responses based on inferred user traits.
  • Explainability tooling for AI systems, since personality recognition is presented as a way to make AI behaviour more interpretable and user-aligned.
  • Content generation and marketing, where the paper mentions personality informing advertisement content and predicting buying intent.

Industry relevance: the Embeddings pipeline matters commercially because it shows a route to migrating knowledge from expensive decoder-only LLMs into encoder models offering fast, low-computational-cost classification — attractive for deployment at scale. The paper also flags practical engineering issues (JSON schema compliance, repair tooling, and error rates) that any production LLM pipeline must handle.

Future Directions

  • Reducing the pre-inferencing and inference-speed drawbacks of decoder-only approaches relative to encoder-only models, which the literature review explicitly identifies as a target for improvement.
  • Determining when fine-tuning helps versus hurts. The Essays results show decoder fine-tuning (C^RT) underperforming prompting-based PICEPR, and two-shot prompting (C^PR2) underperforming one-shot — both open questions about optimal configuration.
  • Improving structured output reliability, especially given that gemini's fine-tuned runs produced valid JSON for less than one-third of the dataset and had to be treated as non-referable.
  • Extending the analysis to the Embeddings pipeline, whose detailed results (Tables VI and VII) are referenced but not available in the provided content, and clarifying how different encoder/decoder combinations affect performance.

Target Audience

Researchers and graduate students in NLP and computational social science working on personality recognition, text classification, or LLM prompting strategies; machine learning engineers evaluating whether to deploy decoder-only LLM pipelines versus fine-tuned encoder models in production; and psychologists or HCI researchers interested in how the Big-5 and MBTI frameworks, and 77 personality facets, can be operationalised inside language models.

Authors’ abstract

Large Language Models (LLMs) have demonstrated remarkable capabilities across various natural language processing tasks. This research introduces a novel "Prompting-in-a-Series" algorithm, termed PICEPR (Psychology-Informed Contents Embeddings for Personality Recognition), featuring two pipelines: (a) Contents and (b) Embeddings. The approach demonstrates how a modularised decoder-only LLM can summarize or generate content, which can aid in classifying or enhancing personality recognition functions as a personality feature extractor and a generator for personality-rich content. We conducted various experiments to provide evidence to justify the rationale behind the PICEPR algorithm. Meanwhile, we also explored closed-source models such as \textit{gpt4o} from OpenAI and \textit{gemini} from Google, along with open-source models like \textit{mistral} from Mistral AI, to compare the quality of the generated content. The PICEPR algorithm has achieved a new state-of-the-art performance for personality recognition by 5-15\% improvement. The work repository and models' weight can be found at https://research.jingjietan.com/?q=PICEPR.

Read the original paper