Skip to content
AI.info

Research

Rethinking Search: A Study of University Students' Perspectives on Using LLMs and Traditional Search Engines in Academic Problem Solving

Overview Research area: Human-Computer Interaction (cs.HC) — specifically how university students choose between LLM-based tools (e.g., ChatGPT) and traditional search engines (e.g., Google) for acade

Rethinking Search: A Study of University Students' Perspectives on Using LLMs and Traditional Search Engines in Academic Problem Solving
arXiv
2510.17726
Published
2025-10-20
Authors
Md. Faiyaz Abdullah Sayeedi, Md. Sadman Haque, Zobaer Ibn Razzaque, Robiul Awoul Robin, Sabila Nawshin

AI summary

Overview

Research area: Human-Computer Interaction (cs.HC) — specifically how university students choose between LLM-based tools (e.g., ChatGPT) and traditional search engines (e.g., Google) for academic problem solving, plus the design of a prototype hybrid tool.

Technical level: Beginner-Friendly. The paper is a user study written in largely non-technical language; the only statistical machinery is one-way repeated measures ANOVA and a chi-square test, both reported plainly.

Scope (1 sentence): A mixed-methods study of 109 surveyed students and 12 interviewed students that measures perceived usability, efficiency, satisfaction and task performance for LLMs versus search engines, and proposes a conceptual chatbot-in-the-search-interface prototype combining the two.

Published as arXiv:2510.17726v1 [cs.HC] on 20 Oct 2025 under a CC BY 4.0 license. Authors are from United International University, Bangladesh, and Indiana University Bloomington, USA.

What This Paper Is About

University students now routinely bounce between Google and large language models to do academic work, but no one has clearly measured how they perceive each tool across the same set of criteria, or how the constant switching affects their actual task performance. The paper's goal is to quantify those perceptions and performance differences, identify the natural usage patterns students fall into, and then design a prototype that fuses GPT's conversational summarization with Google's source credibility. The authors frame the core tension as a trade-off between speed/ease (LLMs) and trustworthiness/verifiability (search engines).

Key Contributions

  1. A parallel, same-rater comparison of LLMs versus search engines on four matched dimensions — usage frequency, satisfaction, efficiency, and ease of use — measured with 5-point Likert scales across 109 students, enabling within-subjects statistical testing rather than comparing separate studies.

  2. Observational task-performance data from 12 in-person interviews in which participants completed six structured academic tasks (summarization, coding, circuit analysis, business chart interpretation, formal email drafting, concept comparison) and were grouped into four tool-usage patterns: GPT-only, Google-only, tool-balancing, and random-choice.

  3. Empirical evidence that hybrid tool use beats single-tool use, with tool-balancing users reaching 90% average accuracy versus 83% for GPT-only, 78% for Google-only, and 82–88% for random-choice users.

  4. A user-informed conceptual prototype — an embedded chatbot positioned in the corner of a search interface — that layers GPT-style summaries and follow-up dialogue on top of standard search results while attaching links to original sources, developed as a Figma design.

Main Findings

  • LLMs outscored search engines on every dimension. Usage frequency mean 2.79 (median 3.0, mode 3) for LLMs versus 2.33 (median 2.0, mode 2) for search engines; satisfaction 2.47 versus 1.99; efficiency 2.65 versus 2.06; ease of use 2.74 versus 2.10.

  • The differences were statistically significant. One-way repeated measures ANOVA gave F(1,108) = 14.82 for usage frequency, 18.95 for satisfaction, 21.37 for efficiency, and 17.04 for ease of use, all with p < 0.001. Standard deviations were around 0.9 for usage and satisfaction, 0.94 and 1.00 for efficiency, and 1.12 for LLM ease of use.

  • Demographics did not predict tool preference. A chi-square test yielded χ²(6, N=109) = 2.012 with p = 0.570, so the null hypothesis was not rejected — preference was not significantly associated with age group, gender, or academic department.

  • Balanced tool use produced the highest accuracy at the highest time cost. Tool-balancing users averaged 90% accuracy over 30 minutes; GPT-only users 83% in 19 minutes; Google-only users 78% in 24 minutes; the eight random-choice users ranged from 82% to 88% accuracy and 22 to 29 minutes. Overall completion times ranged from 13 to 42 minutes.

  • Search engines were slow because of manual synthesis, LLMs were fast but distrusted. Google-only performance was attributed to the effort of navigating, synthesizing and rephrasing content from multiple sources; GPT was fast because its conversational form "reduces the need to browse multiple webpages."

  • Four qualitative themes emerged: task suitability and tool preference; perceptions of reliability and accuracy; workflow efficiency and cognitive load; and usability and interaction experience.

  • Students described the tools as complementary, not competing. GPT was favored for quick answers, summarization, drafting and coding support; Google for fact-checking, citing sources, and exploring multiple viewpoints.

  • There was strong demand for a hybrid tool. Participants wanted concise responses with embedded citations and links, framed as a way to reduce cognitive load and eliminate repetitive tool-switching.

Methodology in Plain English

The authors ran two connected studies.

First, an online survey of 109 university students drawn mostly from technology fields (CSE, EEE, Data Science) but also from Business Administration and Medicine. The questionnaire asked the same four questions twice — once about traditional search engines and once about LLM-based tools — covering usage frequency, satisfaction with accuracy, efficiency, and ease of use, each on a 5-point scale from Never to Always. Answers were converted to numbers 0 through 4, summarised with means, medians, modes and standard deviations, and compared using one-way repeated measures ANOVA (appropriate because each person rated both tools). A chi-square test checked whether overall tool preference was linked to age, gender or department.

Second, in-person interviews with 12 students, mostly from CSE, EEE and BBA at United International University in Dhaka. Each participant completed six academic tasks of varying cognitive demand, some discipline-specific (a Python factorial function for CSE, a resistive circuit for EEE, a quarterly revenue bar chart for BBA) and some general (summarizing a research abstract, drafting a formal email, comparing quantitative versus qualitative research methods). Tasks were graded by two evaluators using a pre-written 0–10 rubric covering things like coverage, correctness, code quality, professional tone and comparative logic. Participants were sorted into four groups by how they used tools, and both accuracy and completion time were recorded. Open-ended survey answers and interview transcripts were then coded line-by-line by two researchers and grouped into four themes.

Why This Matters

Impact on research. The paper pushes past the usual setup of evaluating LLMs and search engines in isolation or on single narrow tasks. By measuring both tools with the same instrument and the same participants, and by adding behavioural performance data alongside self-reports, it gives a cleaner picture of the hybrid behaviour researchers have mostly described anecdotally. The 90% versus 83% versus 78% accuracy split is a concrete data point for the argument that tool integration outperforms any single tool.

Real-world applications:

  • Academic search product design — the prototype sketch gives a concrete interaction pattern (corner chatbot, source-linked summaries, toggle between raw results and AI interpretation) that product teams could build toward.
  • Library and research-skills instruction — the qualitative themes identify exactly where students misjudge tool reliability ("Sometimes GPT gives an answer that sounds right but isn't actually correct"), which points to specific training targets.
  • Discipline-specific study tooling — the finding that GPT was favored for programming help and Google for research papers suggests tools could adapt default behaviour to the subject matter.
  • Assessment design in universities — knowing that a balanced-tool workflow takes 30 minutes versus 19 minutes for GPT-only gives instructors a realistic sense of task time budgets.

Industry relevance. The paper directly addresses the gap the authors say exists in Perplexity AI and Google's Search Generative Experience (SGE), which they describe as largely static and lacking personalization, real-time adaptation, and task-specific reasoning. Search providers, AI assistant developers, and edtech companies all have a stake in the design direction it argues for — AI-generated answers that carry their sources with them.

Future Directions

  1. Expand participant diversity. The interview sample (n = 12) was small and drawn primarily from technology disciplines at a single university, so the authors plan to broaden recruitment to test whether the tool-balancing advantage holds beyond CSE-dominated populations.

  2. Validate the qualitative themes. The thematic analysis lacked inter-coder reliability checks; the authors call for subsequent work to verify the four themes, and note that the survey data were self-reported and potentially subject to recall or social desirability bias.

  3. Build and test a functional prototype. The current design is conceptual and has not been implemented or evaluated in a real academic environment, so its practical impact on usability and learning outcomes is unverified.

  4. Account for unmeasured moderators. Digital literacy, prior experience with AI tools, and task complexity were not controlled for and could plausibly influence both tool preference and performance — an open question raised explicitly in the limitations. The authors also speculate that the assistant could eventually learn user preferences, discipline-specific language, and search habits over time.

Target Audience

HCI and information-retrieval researchers studying AI-assisted search and mixed-initiative interfaces; educators and university administrators deciding how to position AI tools in coursework; UX and product designers building search, study, or knowledge-management tools; and students or instructors who want an evidence-based picture of when ChatGPT-style tools help and when they need to be checked against a search engine. The paper is written at an accessible level and the statistics are limited to ANOVA and chi-square, so readers without a strong quantitative background can follow it.

Authors’ abstract

With the increasing integration of Artificial Intelligence (AI) in academic problem solving, university students frequently alternate between traditional search engines like Google and large language models (LLMs) for information retrieval. This study explores students' perceptions of both tools, emphasizing usability, efficiency, and their integration into academic workflows. Employing a mixed-methods approach, we surveyed 109 students from diverse disciplines and conducted in-depth interviews with 12 participants. Quantitative analyses, including ANOVA and chi-square tests, were used to assess differences in efficiency, satisfaction, and tool preference. Qualitative insights revealed that students commonly switch between GPT and Google: using Google for credible, multi-source information and GPT for summarization, explanation, and drafting. While neither tool proved sufficient on its own, there was a strong demand for a hybrid solution. In response, we developed a prototype, a chatbot embedded within the search interface, that combines GPT's conversational capabilities with Google's reliability to enhance academic research and reduce cognitive load.

Read the original paper