Research
Is Passive Expertise-Based Personalization Enough? A Case Study in AI-Assisted Test-Taking
Is Passive Expertise-Based Personalization Enough? A Case Study in AI-Assisted Test-Taking Overview Research area: Human-Computer Interaction, specifically personalization of Large Language Model (LLM

- arXiv
- 2511.23376
- Published
- 2025-11-28
- Authors
- Li Siyan, Jason Zhang, Akash Maharaj, Yuanming Shi, Yunyao Li
AI summary
Is Passive Expertise-Based Personalization Enough? A Case Study in AI-Assisted Test-TakingOverview
Research area: Human-Computer Interaction, specifically personalization of Large Language Model (LLM)-based enterprise conversational assistants for knowledge-intensive, task-oriented work.
Technical level: Intermediate. The paper combines prompt engineering, prompt optimization frameworks (DSPy), lightweight user modeling, and a controlled human-subject study; readers benefit from familiarity with LLM assistants and standard usability instruments such as NASA-TLX.
Scope (one sentence): A small-scale, within-subject user study testing whether an enterprise AI assistant that passively adapts responses to a user's expertise level improves exam performance, task load, and conversational experience in a timed certification-exam task on a customer data management platform (anonymized as "Platform A").
What This Paper Is About
Prior work has argued that tailoring responses to a user's expertise level should help novices and experts alike, but the actual effects in knowledge-intensive tasks have not been systematically studied. The authors build a version of an enterprise AI assistant with passive expertise-based personalization (the system infers and applies expertise; the user has little control) and compare it against a baseline assistant. They then test both systems with participants answering timed product-knowledge exam questions, asking whether personalization actually improves performance and experience, and where it fails.
Key Contributions
- An extensible framework for passive expertise-based personalization, applied to both a product-knowledge agent and a primary agent in an agentic assistant design, using domain-level and overall expertise.
- An empirical examination of that framework through a small-scale, within-subject user study in which participants completed two timed sets of certification exam questions, only one of which was completed with the personalized assistant.
- Identification of task-specific limitations of passive personalization, notably that non-experts do not always benefit from more informative personalized responses under time pressure even though they prefer them in a less pressured setting.
- Qualitative and quantitative characterization of expert and non-expert behavior, including how they use the assistant and how their queries differ.
Main Findings
- Exam scores (questions where the AI assistant was used): Average percentage scores were 58.1 overall for baseline and 62.0 overall for personalized. By level: novices 49.2 to 56.0, intermediates 56.3 to 51.3, experts 68.3 to 74.2. The authors read this as personalization possibly helping novices and experts but not intermediates.
- Assistant usage volume: Participants sent the baseline assistant 88 inquiries and the personalized assistant 89 inquiries.
- How responses were used: Baseline responses led to 34.5% guessed answers, 47.3% directly given answers, and 18.2% extrapolated answers; personalized responses led to 42.3% guessed, 23.1% direct, and 34.6% extrapolated. Non-expert participants guessed slightly less with personalization (28.2% versus 29.5%).
- Task load (NASA-TLX, 13 participants): Personalization reduced physical demand (2.23 to 1.92), temporal demand (4.85 to 4.69), perceived performance shortfall (4.38 to 4.08), and effort (4.53 to 4.38). It increased mental demand (5.00 to 5.31) and frustration (3.69 to 3.84).
- Task load for non-experts (nine participants, Appendix D.2): The same pattern held with a larger increase in mental demand, a gap of 0.44 compared to 0.31 for the full pool. Baseline values were mental 5.22, physical 2.11, temporal 5.00, performance 4.56, effort 4.78, frustration 4.11; personalized values were 5.66, 2.00, 4.78, 3.89, 4.78, and 4.33.
- Conversational experience (5-point ratings): The personalized assistant was rated higher on helpfulness (3.08 versus 2.84) and relevance (3.07 versus 3.00), but lower on understandability (3.46 versus 3.61) and expertise alignment (2.92 versus 3.00).
- Expertise alignment by level (Appendix D.1): Experts rated the personalized assistant better aligned (3.00 versus 2.25 for baseline), while novices preferred baseline (3.60 versus 3.00) and intermediates also preferred baseline (3.00 versus 2.75).
- Novice preference under relaxed conditions: In a follow-up study with four novice participants comparing responses to their own sampled queries, the personalized assistant was preferred 78.6% of the time for helpfulness and relevance, and 85.7% for understandability and expertise alignment (baseline rates were 21.4%, 14.3%, 21.4%, and 14.3% respectively).
- Query characteristics invert for novices: Novice queries in the study were more similar to the exam questions (partial string similarity 74.7 versus 67.5 for experts, significant at p < 0.05), longer (20.1 words versus 10.8, p < 0.001), contained more jargon (1.97 versus 1.22, p < 0.05), and were judged by the LLM to exhibit more expertise (0.670 versus 0.439, p < 0.05). Experts asked fewer questions (10.3 versus 17.6 on average).
- Expertise classification dimensions were validated on internal data: Comparing 781 intermediate and 417 expert queries, experts showed higher LLM-judged expertise (0.528 versus 0.431, p < 0.05), higher words-per-sentence ratio (11.0 versus 9.66, p < 0.001), more jargon (0.194 versus 0.104, p < 0.001), and higher Flesch Reading Ease (86.0 versus 79.3, p < 0.001).
- Qualitative behavior differences: Novices and intermediates used the assistant mainly as an educational resource, while experts used it primarily as a sanity check; experts were more likely to recognize and ignore assistant hallucinations.
Methodology in Plain English
The assistant follows an agentic design. When a user submits a product-knowledge query, a primary agent estimates the user's overall expertise level and forwards a slightly reworded query to a product-knowledge (PK) agent. The PK agent runs a domain classifier, retrieves the user's expertise for that domain, and personalizes its answer while reporting the domain-specific expertise level. The primary agent then applies stylistic constraints based on overall expertise to produce the final response.
Domain classification: The PK agent uses Retrieval-Augmented Generation, and the authors built an LLM-based domain classifier optimized with the SIMBA optimizer from DSPy. They used 300 randomly selected queries as the training set and the rest as validation, achieving 82.2% validation top-1 accuracy.
Expertise signals: Four dimensions were examined: an LLM judgment prompt adapted from prior work (originally five tiers, reduced to Novice (0), Intermediate (1), and Expert (2)); words-per-sentence ratio; jargon usage computed by fuzzy string matching against an official dictionary of Platform A terminology; and readability, using a library that produces nine readability scores (the Flesch Reading Ease score is reported). These were tested on 1,198 internal use-case-specific queries (781 intermediate, 417 expert) with independent t-tests.
Response adaptation: Rather than fine-tuning, the authors used prompt modification. Content control was driven by a directive specifying how to phrase responses for novice, expert, and intermediate users, adapted from an established instructional-strategy framework. Stylistic control (readability, jargon density, message length) was handled by a developer prompt optimized with DSPy's MIPROv2 optimizer and GPT-4o-mini, proposing 15 prompts evaluated on 150 internal queries against a compliance metric. Compliance scores went from 53.1 to 75.8 for readability, 78.1 to 75.8 for jargon, and stayed at 100.0 for length. GPT-4o was the default LLM unless otherwise specified, and the data collection drew on 1,785 internal user queries gathered during 2024–2025.
User study: 16 users of Platform A were recruited through internal channels. Three experts completed the exams unaided and were removed, leaving 13 participants for analysis (five novices, four intermediates, four experts). One person worked in administration, one in product, and the rest in engineering. Participants completed two sets of eight product-knowledge questions drawn from 24 manually curated certification-exam questions; Appendix C describes the filtering process and states that 18 questions remained after removing items with unresolved disagreement, with a resulting Krippendorff's alpha of 0.899 on difficulty ratings. A coin flip determined which assistant version appeared first, and each set was allotted 10 minutes. After each question, participants reported whether they knew the answer, guessed, were given the answer directly, or extrapolated it. After each set they completed a post-survey with NASA-TLX and 5-point Likert items on helpfulness, understandability, relevance, and expertise alignment. Expertise level was assigned via a role-to-expertise mapping created with expert consultation and assumed static throughout the interaction. A follow-up preference study asked four novice participants to compare baseline and personalized responses for three to four of their own queries each.
Why This Matters
The paper shows that personalization based on user expertise is not a straightforward win: it moderates task load and improves perceived helpfulness and relevance, but it can backfire under time pressure and can misread users when expertise is inferred from query text. In this task, novice queries superficially resembled expert queries, meaning a system that dynamically inferred expertise from language alone could serve novices exactly the terse, jargon-heavy responses they cannot use. The authors argue for combining passive and active personalization so users can adjust responses to the situation.
Real-world applications:
- Enterprise AI assistants that answer product documentation and support questions for mixed-skill user bases.
- Onboarding and enablement tools where employees move between domains of varying familiarity.
- Certification and exam preparation tools that need to balance informativeness against time constraints.
- Conversational interfaces that expose explicit style controls, such as response-length toggles, so users can trade detail for speed.
Industry relevance: The results speak directly to how vendors of enterprise copilots should design personalization controls — the study suggests interface-level user agency (letting users change expertise level or response style per task) may matter as much as backend inference, and that response latency and verbosity introduced by personalization carry measurable costs.
Future Directions
- Add active personalization: Give users control over their declared expertise level and stylistic dimensions, such as response length, and measure whether this improves both task performance and experience.
- Run a larger, more role-diverse study: The authors note that their participant pool did not allow strong conclusions per expertise level and that many participants reported the same role (engineering), possibly limiting generalizability.
- Improve expertise inference carefully: The finding that novice queries resemble expert queries in this task raises open questions about dynamic expertise classification and how to avoid misclassifying novices.
- Build a jargon difficulty taxonomy: The authors state that without an official taxonomy of jargon difficulty for Platform A, they currently lack the means to further simplify personalized responses for non-experts, and that verbosity and latency of personalized responses need addressing.
Target Audience
Researchers and practitioners in human-computer interaction and conversational AI who work on personalization, user modeling, or enterprise assistants. It is also useful for product teams designing LLM-based assistants for mixed-expertise audiences, for learning-science researchers interested in expertise-adaptive instruction, and for anyone planning small-scale within-subject studies of AI assistant features.
Authors’ abstract
Novice and expert users have different systematic preferences in task-oriented dialogues. However, whether catering to these preferences actually improves user experience and task performance remains understudied. To investigate the effects of expertise-based personalization, we first built a version of an enterprise AI assistant with passive personalization. We then conducted a user study where participants completed timed exams, aided by the two versions of the AI assistant. Preliminary results indicate that passive personalization helps reduce task load and improve assistant perception, but reveal task-specific limitations that can be addressed through providing more user agency. These findings underscore the importance of combining active and passive personalization to optimize user experience and effectiveness in enterprise task-oriented environments.