Research
Depth and Autonomy: A Framework for Evaluating LLM Applications in Social Science Research
Overview Research area: Applications of large language models (LLMs) in qualitative social science research, and the methodological evaluation of those applications. Technical level: Beginner-Friendly
- arXiv
- 2510.25432
- Published
- 2025-10-29
- Authors
- Ali Sanaei, Ali Rajabzadeh
AI summary
Overview
Research area: Applications of large language models (LLMs) in qualitative social science research, and the methodological evaluation of those applications.
Technical level: Beginner-Friendly to Intermediate. The abstract describes a conceptual framework and a literature survey rather than new model architectures or experimental benchmarks, so no specialized machine learning background is required.
Scope (one sentence): The paper proposes a two-dimensional framework — interpretive depth and autonomy — for classifying LLM applications in qualitative social science research and for deriving practical design recommendations, supported by a review of published social science papers that use LLMs as a tool.
What This Paper Is About
Qualitative social science researchers are increasingly adopting LLMs, but that adoption runs into persistent problems: interpretive bias, low reliability, and weak auditability. The authors argue that the way to address these problems is not to give models more freedom, but to structure their use carefully — decomposing tasks into manageable segments and limiting how much the model decides on its own. The paper's goal is to supply a simple way of describing where any given LLM application sits on two axes, and to turn that description into concrete design advice.
Key Contributions
- A two-dimensional framework. LLM usage in qualitative research is situated along two dimensions — interpretive depth (how much interpretive judgment the model performs) and autonomy (how much freedom the model is given) — providing a straightforward classification scheme.
- Practical design recommendations. The framework is used to derive guidance for how researchers should configure LLM use, rather than leaving it as a descriptive taxonomy.
- A mapping of the existing literature. The authors present the state of published social science research with respect to these two dimensions, drawing on all social science papers available on Web of Science that use LLMs as a tool and not strictly as the subject of study.
- A delegation analogy as a design principle. The paper frames appropriate LLM delegation in terms of how a researcher would delegate work to a capable undergraduate research assistant — decompose the work, keep the model's latitude low, and add interpretive responsibility only with supervision.
Main Findings
- Two dimensions organize the field. Interpretive depth and autonomy are presented as the axes along which LLM applications in qualitative research can be usefully distinguished and compared.
- The literature is surveyed along these axes. The abstract states that the authors present the state of the literature on both dimensions, covering published Web of Science social science papers that use LLMs as a tool rather than as an object of study. The specific distribution of papers across the two dimensions, and any counts or trends, are not reported in the abstract.
- Low autonomy is the recommended default. The framework counsels against granting models expansive freedom; instead, researchers should decompose tasks into manageable segments.
- Interpretive depth should be raised selectively. Depth is to be increased only where warranted and only under supervision, rather than applied uniformly across a research pipeline.
- The claimed payoff is plausible, not demonstrated. The abstract frames the outcome as one that "one can plausibly reap" — the benefits of LLMs while preserving transparency and reliability. It does not report measured improvements in reliability, bias reduction, or auditability.
- The problems motivating the work are named but not quantified. Interpretive bias, low reliability, and weak auditability are identified as persistent challenges to LLM adoption in the field; the abstract offers no figures on their magnitude.
Methodology in Plain English
The paper takes a conceptual and review-based approach rather than an experimental one. The authors first define two dimensions — how much interpretation a model performs and how much autonomy it is given — and then use those dimensions as a lens for organizing existing work. To characterize the field, they look at social science papers published on Web of Science that employ LLMs as a tool, excluding work in which the model itself is the object of study. From this mapping they draw design recommendations, anchored in an analogy to how a researcher would supervise an undergraduate research assistant: splitting work into pieces, keeping the assistant's discretion limited, and reserving harder interpretive calls for supervised stages. The abstract does not describe the search procedure, the number of papers reviewed, the coding scheme, or any quantitative synthesis.
Why This Matters
Impact on research. Qualitative social science depends on transparency and on being able to trace how interpretations were reached. If LLMs are inserted into that process without structure, bias and unverifiable outputs become hard to detect. A shared vocabulary of depth and autonomy gives researchers, reviewers, and readers a way to state and evaluate exactly how a model was used — which is itself a step toward auditability.
Where the framework would apply (the abstract names the setting, qualitative social science, but does not enumerate specific tasks):
- Thematic analysis and qualitative coding, where a model might assist with segmenting or labeling material.
- Screening and triage of large document sets during literature review.
- Summarization or extraction of passages from interviews and field notes.
- Any multi-stage workflow where an LLM's role could be dialed up or down depending on the interpretive stakes.
Industry relevance. The framework speaks directly to a broader pattern in applied settings, where teams deploy LLMs in analysis pipelines and face the same trade-off between capability and control. The undergraduate research assistant delegation model is a design heuristic that transfers readily: split tasks, constrain discretion, supervise the interpretive steps. The paper's argument that capability should be granted incrementally rather than by default is relevant to anyone building human-in-the-loop LLM tooling.
Future Directions
- Empirical validation. The abstract claims only that benefits can "plausibly" be reaped under low autonomy and selective depth. Testing whether these settings actually reduce interpretive bias, improve reliability, or strengthen auditability is the obvious next step.
- Operationalizing the dimensions. Turning "interpretive depth" and "autonomy" into measurable, reproducible ratings would make the framework usable for comparing applications rather than only describing them.
- Broadening the evidence base. The review is bounded to Web of Science social science papers using LLMs as a tool; extending it to other databases, disciplines, and to non-published deployments would test how general the mapping is.
- Supervision and decomposition protocols. The paper recommends decomposing tasks and supervising raised interpretive depth but does not, per the abstract, specify how to decide where depth is "warranted" — a question of concrete procedure and tooling.
Target Audience
Qualitative social scientists and research methodologists considering or already using LLMs; reviewers, ethics boards, and research-integrity staff who need to assess how a model was used in a study; and developers or product teams building LLM-assisted research and analysis tools who want a principled rationale for limiting model autonomy and staging interpretive responsibility.
Authors’ abstract
Large language models (LLMs) are increasingly utilized by researchers across a wide range of domains, and qualitative social science is no exception; however, this adoption faces persistent challenges, including interpretive bias, low reliability, and weak auditability. We introduce a framework that situates LLM usage along two dimensions, interpretive depth and autonomy, thereby offering a straightforward way to classify LLM applications in qualitative research and to derive practical design recommendations. We present the state of the literature with respect to these two dimensions, based on all published social science papers available on Web of Science that use LLMs as a tool and not strictly as the subject of study. Rather than granting models expansive freedom, our approach encourages researchers to decompose tasks into manageable segments, much as they would when delegating work to capable undergraduate research assistants. By maintaining low levels of autonomy and selectively increasing interpretive depth only where warranted and under supervision, one can plausibly reap the benefits of LLMs while preserving transparency and reliability.