Skip to main content
All Collections

Psychology

LLMs in qualitative psychology

No spamNo marketingOnly data
ES

Evgeny Smirnov

21 papers · 6 Must Read · 2023–2026

Last updated Aug 2, 2026

Sorted by publication date, newest first. New papers are marked so you can spot recent additions.

Introduction

Emerging applications of large language models in qualitative psychological research methodology as well as the critics.

1
Must Read
beginner

Smirnov, Evgeny · 2026 · Center for Open Science

At a GlanceAI

Critical review argues against blanket rejection of large language models in reflexive qualitative research.

SummaryAI

This commentary responds to an open letter in which 419 qualitative researchers rejected generative AI for reflexive qualitative research. It organizes 27 concerns into technical, methodological, ethical-political, epistemological, and ontological domains, arguing that many address outdated tools, misuse, or unsettled philosophical questions rather than inherent limits of LLMs. The paper advocates constructive debate about responsible use instead of categorical rejection, framing LLMs as increasingly unavoidable research tools.

Line-by-line explanation of why many arguments against the use of LLMs in psychology are NOT plausible.

ES

Method:AI
Critical commentary reviewing and evaluating 27 objections to LLM use in reflexive qualitative research.
Background:AI
Qualitative research methods, reflexivity, and basic familiarity with generative AI are needed.

At a GlanceAI

Position paper argues that metaphysical objections should not categorically exclude generative AI from qualitative analysis.

SummaryAI

This response challenges calls to reject generative AI outright in qualitative analysis. It argues that the proposed prohibition rests largely on philosophical assumptions rather than solely methodological concerns, and that treating these assumptions as settled could limit scholarly debate and innovation. The paper advocates keeping discussion of appropriate genAI use in qualitative research open rather than imposing a categorical ban.

An attempt to defend the usage of LLMs. A good attempt but it's too polite and does not focus on the real gaps in Jowsey's arguments.

ES

Method:AI
A philosophical position paper critiques a categorical ban on generative AI in qualitative analysis.
Background:AI
Basic knowledge of qualitative research methods and debates over generative AI in research.
3
Niche
intermediate

Misra, Ranjita, Dahal, Rajashree, Kirk, Brenna et al. · 2026 · International Journal of Qualitative Methods

At a GlanceAI

Open-source LLMs can speed qualitative coding but still produce many repetitive or context-poor codes.

SummaryAI

This study tests whether privacy-friendlier open-source LLMs, Gemma2 and Llama3.1, can assist with thematic analysis of sensitive patient-interview data. Compared with researcher-generated codes, the models showed partial alignment, and deductive prompting produced more and more nuanced codes than inductive prompting. However, only about 45% of LLM codes supplied meaningful context and 22–39% were duplicates, indicating that LLMs are useful assistants rather than replacements for qualitative researchers. Domain-specific models and stronger validation may improve their reliability for healthcare communication research.

Note the fact that the models are outdated. Newer ones can narrow the gap.

ES

Method:AI
The study compares codes from two open-source LLMs with researcher-led inductive thematic analysis of 34 patient interviews.
Background:AI
Basic knowledge of qualitative thematic analysis, coding, and large language models is needed.
4
Must Read
beginner

Jowsey, Tanisha, Braun, Virginia, Clarke, Victoria et al. · 2025 · Qualitative Inquiry

At a GlanceAI

419 qualitative researchers reject GenAI for reflexive analysis as methodologically incompatible and socially unjust.

SummaryAI

This position statement brings together 419 experienced qualitative researchers from 32 countries to oppose generative AI in reflexive qualitative research. The authors argue that approaches such as reflexive thematic analysis depend on a situated human researcher's subjective interpretation and reflexive engagement, making GenAI methodologically incongruent. They also frame rejection of GenAI as an issue of social and environmental justice, extending the debate beyond questions of analytic convenience or accuracy.

Debates, debates, debates... However, how strong are the arguments? (spoiler: nope; why? because the authors do not realize what LLMs really is and how to use it...)

ES

Method:AI
A multi-author position statement argues against using generative AI in reflexive qualitative analysis.
Background:AI
Foundational qualitative research methods, especially reflexivity and reflexive thematic analysis.
5
Skip
beginner

Khalid, Muhammad Talal, Witmer, Ann-Perry · 2025 · Social Science Computer Review

At a GlanceAI

A four-step prompt-engineering framework aims to make LLM-assisted thematic analysis more rigorous and reproducible.

SummaryAI

LLMs may reduce the time and cost of inductive thematic analysis, but ad hoc prompting can make results difficult to trust, explain, or reproduce. This paper reviews how existing studies incorporate LLMs and prompt engineering into thematic-analysis workflows, then proposes a structured four-step prompting process. It also discusses advanced prompting techniques and identifies research gaps, offering practical guidance for more methodologically rigorous LLM-supported qualitative research.

Prompt engineering was important... a few years ago. Today it's better to focus on the methods that can overcome this limitation.

ES

Method:AI
A literature review is used to map LLM-assisted inductive thematic analysis and derive a structured prompt-engineering process.
Background:AI
Basic knowledge of qualitative thematic analysis, large language models, and prompt engineering is needed.
6
Skip
intermediate

Vikan, Magnhild, Aryan, Ramtin, Kannelønning, Mari Serine et al. · 2025 · Qualitative Health Research

At a GlanceAI

Testing finds base LLMs offer limited help for reflexive thematic analysis and cannot replace human interpretation.

SummaryAI

This exploratory study directly tests whether a base LLM can support reflexive thematic analysis (RTA), a qualitative approach in which researchers’ situated interpretation is central. Using an offline model on a health-science interview, the authors developed and refined prompts across RTA phases and identified pitfalls including misleading prompts, self-referential outputs, translation problems, and errors. They conclude that current base models do not improve RTA efficiency, though they may help reveal gaps in researchers’ perspectives. The work argues that careful human familiarization and methodological competence remain necessary to preserve RTA’s epistemological integrity.

Wrong models, wrong method, wrong results...

ES

Method:AI
The authors tested an offline base LLM on a health-science interview across reflexive thematic analysis phases, iteratively refining prompts.
Background:AI
Basic qualitative research methods, especially reflexive thematic analysis and the role of researcher reflexivity.
7
Niche
beginner
★ Essential

Yue, Yongjie, Liu, Dong, Lv, Yilin et al. · 2025 · Journal of Medical Internet Research

At a GlanceAI

Tutorial evaluates ChatGPT 4-Turbo for grounded theory, finding faster, more diverse coding but weaker contextual depth than humans.

SummaryAI

This tutorial provides a practical early assessment of using ChatGPT 4-Turbo for grounded-theory analysis, based on semistructured interviews with blind gamers. It finds that the model can reliably support many coding tasks while increasing coding efficiency and diversity, potentially updating how grounded-theory workflows are conducted. However, its coding was weaker than manual analysis in contextual depth, connections between ideas, and organization, so it should support rather than replace qualitative researchers.

Grounded theory application tutorial, though the methods could be better.

ES

Method:AI
A step-by-step case study compares ChatGPT 4-Turbo coding with software-assisted manual coding in grounded-theory analysis of interviews with blind gamers.
Background:AI
Basic qualitative research methods, especially grounded theory and interview coding, plus familiarity with generative AI.
8
Must Read
intermediate

Dunivin, Zackary Okun · 2025 · EPJ Data Science

At a GlanceAI

A hybrid workflow shows how LLMs can scale qualitative coding while retaining human-led reflexive interpretation.

SummaryAI

The paper addresses how researchers can use LLMs to code large qualitative datasets without reducing coding to purely mechanical classification. It proposes retaining human-led codebook development and refinement, then rewriting code definitions and using structured prompts so models can apply them more reliably. A socio-historical case study suggests frontier language models can interpret paragraph-length text, while the paper stresses ethical safeguards and researchers' continuing interpretive leadership.

If you read one methods paper before coding with an LLM, read this one. It takes both the hermeneutics and the engineering seriously — a rare combination.

ES

Method:AI
A hybrid qualitative-coding workflow combines human codebook development with iterative prompt design and LLM-based text categorization.
Background:AI
Basic qualitative content analysis and familiarity with large language models are needed.

At a GlanceAI

A framework helps qualitative researchers align LLM-assisted analysis with their research values and methods.

SummaryAI

This guide addresses a key gap in AI-assisted qualitative research: an LLM is not a neutral tool, and its use must fit a study’s underlying research values. The authors contrast “Small-q” and “Big-Q” approaches to qualitative research, review seminal LLM-based analysis work, and propose an approach-based model for assessing alignment between values, methods, and AI use. Exploratory reflexive content-analysis examples show how the same AI assistance may be applied differently under each approach. The paper also highlights ethical issues and offers questions and further resources for researchers considering AI-supported qualitative analysis.

Method:AI
A conceptual review and exploratory illustration of how LLM-assisted analysis aligns with contrasting qualitative research approaches.
Background:AI
Basic qualitative research methods and a general understanding of large language models are needed.
10
Must Read
beginner
★ Essential

Smirnov, Evgeny · 2024 · Qualitative Research in Psychology

At a Glance

Shows how LLMs can speed up and triangulate content analysis even for complex areas

SummaryAI

This paper matters because it translates large language models (especially GPT-4) from a vague "helpful tool" into concrete, text-based qualitative research workflows. Using a set of simulations, it illustrates how LLMs can support study planning (e.g., generating plausible mock interviews), perform directed and conventional content analysis, and evaluate narrative properties like causal coherence. The novelty is the methodological framing: LLMs are positioned as a fast, repeatable "second coder/validator" that can bolster trustworthiness (credibility, dependability, confirmability) when researchers document prompts, repeat runs, and report classification metrics. The implication is a pragmatic path for qualitative psychologists to reduce analysis time while adding structured checks—though the paper also flags key limitations (synthetic data, single-human validation) and insists on human expert verification.

LLMs can identify even complex themes, such as existential concerns, very accurately and perform both direct and conventional content analysis.

ES

Method:
LLM-based simulations
Background:AI
Familiarity with qualitative methods in psychology (content/narrative analysis, trustworthiness criteria) and basic LLM concepts (prompting, RAG).
11
Worth Reading
beginner

Rathje, Steve, Mirea, Dan-Mircea, Sucholutsky, Ilia et al. · 2024 · Proceedings of the National Academy of Sciences

At a GlanceAI

GPT enables accurate, low-code multilingual detection of psychological constructs without task-specific training data.

SummaryAI

The study tests GPT-3.5 Turbo, GPT-4, and GPT-4 Turbo on 47,925 manually annotated tweets and news headlines in 12 languages. Across sentiment, emotions, offensiveness, and moral foundations, GPT substantially outperformed English-language dictionary methods and approached or sometimes exceeded specialized fine-tuned models. Newer GPT versions improved performance, especially for lesser-spoken languages, while becoming less expensive. The findings suggest that prompt-based LLMs can make multilingual psychological text analysis more accessible for researchers without extensive coding skills or labeled training data.

One of the first uses of coding in psychology using LLMs

ES

Method:AI
Benchmarking GPT versions against manual labels, dictionary methods, and fine-tuned models across multilingual text datasets.
Background:AI
Basic familiarity with natural language processing, psychological text measures, and model evaluation.
12
Niche
intermediate

Prescott, Maximo R., Yeager, Samantha, Ham, Lillian et al. · 2024 · JMIR AI

At a GlanceAI

ChatGPT and Bard produced similar themes far faster than humans, but weaker coding agreement supports hybrid qualitative analysis.

SummaryAI

This study tests whether generative AI can accelerate qualitative analysis needed to improve digital health interventions. On 40 SMS reminders for HIV medication adherence, ChatGPT and Bard recovered many human-generated inductive themes while completing analysis in about 20 minutes versus roughly 567 minutes for human coders. However, coding agreement with humans was only fair to moderate, and people better identified nuanced, interpretive themes. The findings support using LLMs to reduce workload while retaining human oversight rather than replacing qualitative researchers.

One of the first pros & cons analysis on the usage of LLMs in psychology + comparison

ES

Method:AI
The study compares human coders with ChatGPT and Bard on inductive and deductive thematic analyses of SMS health-intervention messages.
Background:AI
Basic qualitative research methods, especially thematic analysis and intercoder reliability, plus familiarity with generative AI.
13
Niche
intermediate

Bijker, Rimke, Merkouris, Stephanie S., Dowling, Nicki A. et al. · 2024 · Journal of Medical Internet Research

At a GlanceAI

ChatGPT showed fair-to-strong reliability for assisting qualitative content analysis, especially when building inductive coding schemes.

SummaryAI

This study tests whether ChatGPT can reduce the labor involved in qualitative content analysis of online discussions about lowering sugar consumption. Across repeated coding runs, ChatGPT showed stronger agreement for data-driven, inductive coding than for adapting and applying a pre-existing theoretical framework. The results support using an LLM as a potentially useful assistant or second coder, but emphasize that researchers must iteratively evaluate reliability at every stage of the analysis.

Quantifies coding results

ES

Method:AI
The study used prompt-engineered ChatGPT conversations to extract and code behavior-change mechanisms in 537 forum posts, comparing outputs across inductive and theory-guided schemes.
Background:AI
Basic knowledge of qualitative content analysis, coding reliability, and behavior-change frameworks is helpful.
14
Niche
intermediate

Wachinger, Jonas, Bärnighausen, Kate, Schäfer, Louis N. et al. · 2024 · Qualitative Health Research

At a GlanceAI

ChatGPT broadly matched human qualitative themes but needs careful review, especially when generating codes and theory-linked interpretations.

SummaryAI

The study directly tests whether ChatGPT can support qualitative analysis by comparing its transcript-based outputs with those of an experienced human researcher. Across prompts, ChatGPT identified substantially overlapping themes, generated a plausible codebook and supporting quotes, and connected findings to theoretical debates. Its outputs tended toward descriptive themes and require careful human review, but the results suggest LLMs could assist qualitative research and teaching while also challenging established standards for rigor and nuance.

Method:AI
Comparing ChatGPT and an experienced researcher’s qualitative analysis of an interview transcript across multiple prompts.
Background:AI
Basic qualitative research methods, including thematic analysis, coding, and theory-informed interpretation.
15
Niche
intermediate

Gao, Jie, Guo, Yuchen, Lim, Gionnieve et al. · 2024 · Proceedings of the CHI Conference on Human Factors in Computing Systems

At a GlanceAI

CollabCoder proposes an LLM-supported workflow to make rigorous inductive collaborative qualitative analysis more accessible.

SummaryAI

CollabCoder addresses the challenge of conducting collaborative qualitative analysis rigorously while reducing barriers for researchers. It introduces a workflow that incorporates large language models into inductive coding and analysis rather than treating them only as automated labelers. The work is relevant to researchers seeking practical ways to use LLMs in qualitative research while retaining a structured collaborative process.

Method:AI
The authors present CollabCoder, an LLM-supported workflow for inductive collaborative qualitative analysis.
Background:AI
Basic familiarity with qualitative coding, inductive analysis, and large language models is helpful.
16
Niche
intermediate

Roberts, John, Baker, Max, Andrew, Jane · 2024 · Critical Perspectives on Accounting

At a GlanceAI

Critical assessment of how LLM assistance could reshape qualitative research, including both benefits and risks.

SummaryAI

The article addresses the growing use of large language models as assistants in qualitative research. Its central contribution is to frame LLM use as offering potential practical benefits while also raising important methodological and ethical risks. The discussion is especially relevant to researchers considering where AI support may be useful without undermining interpretation, reflexivity, or research quality.

Method:AI
A critical conceptual discussion of how large language models may assist qualitative research.
Background:AI
Familiarity with qualitative research and the basic capabilities and limits of large language models.
17
Niche
intermediate

Davison, Robert M., Chughtai, Hameed, Nielsen, Petter et al. · 2024 · Information Systems Journal

At a GlanceAI

Ethical guidance for using generative AI in qualitative data analysis.

SummaryAI

This article addresses the ethical challenges of applying generative AI to qualitative data analysis, where interpretation, participant privacy, and researcher responsibility are central. It brings ethical considerations for this emerging use of AI into focus within information systems research. The discussion can help researchers assess when and how generative AI may be used without compromising qualitative research integrity.

Bla-bla-bla. I'd suggest to start with not cheating on the data in psychology...

ES

Method:AI
A conceptual ethics analysis of generative AI use in qualitative data analysis.
Background:AI
Basic knowledge of qualitative research methods, research ethics, and generative AI.
18
Must Read
intermediate

Tai, Robert H., Bentley, Lillian R., Xia, Xin et al. · 2024 · International Journal of Qualitative Methods

At a GlanceAI

LLMs can support reliable deductive coding of qualitative text when guided by a predefined codebook.

SummaryAI

The study tests whether an LLM can assist qualitative researchers with a core but labor-intensive task: applying predefined codes to interview text. Using three interview excerpts, a codebook, and 160 repeated LLM analyses, the authors compare model outputs with traditional human coding evaluations. They argue that LLMs can provide systematic code identification and evidence for coding decisions, potentially reducing misalignment in qualitative analysis, while emphasizing current limitations and research-practice implications.

An importance of a codebook used together with LLMs in qualitative analysis

ES

Method:AI
The authors repeatedly prompt an LLM to apply a predefined qualitative codebook to interview excerpts and compare its coding with human evaluations.
Background:AI
Basic qualitative research methods, especially deductive coding and codebooks, plus a general understanding of LLMs.

At a GlanceAI

GPT-3.5 can recover most human-identified themes from interview data, suggesting a viable but bounded role for LLMs in inductive thematic analysis.

SummaryAI

This study tests whether an LLM can support inductive thematic analysis, a qualitative method traditionally grounded in human interpretation of explicit and latent meaning. Using two previously analyzed open-access interview datasets, the author finds that GPT-3.5-Turbo inferred most of the main themes identified by earlier researchers. The paper is notable for examining how Braun and Clarke’s six analysis phases can only partly be reproduced with an LLM, framing both the promise and limits of automated qualitative analysis. It offers practical recommendations for using LLMs in qualitative research rather than treating them as replacements for human analysts.

The limitations of the LLMs in thematic analysis, though it was performed on a very old model (and hence, the conclusions are not current)

ES

Method:AI
The study uses GPT-3.5-Turbo to conduct inductive thematic analysis on two interview datasets and compares its themes with prior human analyses.
Background:AI
Familiarity with qualitative research, semi-structured interviews, and Braun and Clarke’s thematic analysis framework is helpful.
20
Must Read
beginner
★ Essential

Demszky, Dorottya, Yang, Diyi, Yeager, David S. et al. · 2023 · Nature Reviews Psychology

At a GlanceAI

A review framework for using large language models responsibly in psychological research and applications.

SummaryAI

Large language models can analyze and generate language at a scale that could expand psychology’s ability to study behavior, deliver interventions, and communicate research. This review synthesizes potential applications of these models in psychology alongside concerns about bias, validity, privacy, and misuse. It provides a timely foundation for researchers who want to use LLMs while treating their outputs as tools requiring careful evaluation rather than as authoritative psychological evidence.

The paper everyone cites for a reason.

ES

Method:AI
A narrative review examines how large language models can support psychological research and practice while outlining their limitations and risks.
Background:AI
Basic knowledge of psychology research and large language models is helpful.
21
Worth Reading
beginner

Morgan, David L. · 2023 · International Journal of Qualitative Methods

At a GlanceAI

ChatGPT reproduces concrete qualitative themes well but misses subtler interpretive insights from manual analysis.

SummaryAI

This study tests whether ChatGPT can reanalyze qualitative datasets without the labor-intensive process of manual coding. Across two previously analyzed datasets, it recreated concrete, descriptive themes reasonably well but was less able to identify subtle interpretive themes. The results suggest AI can reduce effort in parts of qualitative analysis, while reinforcing that human judgment remains essential within the broader analytic process.

One of the first honest head-to-head tests of ChatGPT against human thematic analysis.

ES

Method:AI
Using ChatGPT to reanalyze two previously coded qualitative datasets and comparing its themes with the original analyses.
Background:AI
Basic familiarity with qualitative research and thematic analysis is helpful.