Skip to main content
v2026.11,772 entries · CC-BY 4.0

Using NVivo for Braun and Clarke Thematic Analysis

A phase-by-phase workflow guide mapping Braun and Clarke’s six thematic analysis phases onto NVivo’s actual features — memos, nodes, node hierarchies, and coding-comparison queries — including where NVivo’s design pulls against reflexive TA specifically.

Written and maintained by CASRAI Editorial Board

Last updated

A workflow guide, not a method primer

This page maps Virginia Braun and Victoria Clarke’s six-phase thematic analysis process onto the specific operations NVivo provides for each phase. It assumes you already know what thematic analysis is; for the method itself — the theoretical background, what distinguishes reflexive TA from coding-reliability and codebook approaches, and common mistakes — see CASRAI’s step-by-step guide to Braun and Clarke’s six phases. For what NVivo does as a general research-data tool beyond thematic analysis, see the NVivo software guide. This page sits between the two: how the six phases actually get performed inside the software, node by node, memo by memo.

One point worth stating up front, because it shapes every section below: NVivo’s terminology does not map cleanly onto Braun and Clarke’s. NVivo calls a coded label a “node.” Braun and Clarke call the coded label a “code,” and reserve “theme” for a higher-order pattern of shared meaning built up through interpretation, not a bigger bucket of codes. Treating an NVivo node hierarchy as if it were automatically a set of themes is the single most common way this software gets misused for reflexive TA — the sections below flag where that risk shows up.

Phase 1: Familiarization — memos, not nodes

Braun and Clarke’s first phase is immersive, repeated reading of the full dataset, with the analyst noting early observations before any formal coding starts. In NVivo, this phase belongs to memos, not nodes. Attach a memo to each source (an interview transcript, a set of field notes, a focus-group recording) and write your initial impressions directly into it: recurring language, surprising responses, your own reactions as the analyst, and provisional ideas about what might matter later. NVivo links a memo back to its source, so this record stays attached to the material it came from rather than floating free in a separate document.

Resist the pull to start coding during familiarization just because the coding pane is one click away. Reflexive TA treats familiarization as genuinely open reading, not a first coding pass in disguise — opening a node structure too early anchors your reading to categories you haven’t earned yet. Use annotations (short comments tied to a specific passage) for line-level noticing, and reserve memos for the broader, source-level reflections that later become raw material for your reflexivity statement. CASRAI’s reflexivity and positionality statement guide covers how to turn exactly this kind of running memo into a defensible methods-section statement.

Phase 2: Generating Initial Codes — nodes

This is the phase NVivo is built for. Work through the dataset systematically and apply a node to every segment relevant to your research question — a sentence, a few lines, an entire response, whatever unit of meaning the passage actually represents. NVivo supports this two ways: manual coding, where you select text and drag it onto an existing or newly created node, and AI-assisted coding suggestions in current NVivo versions, which propose codes for you to review and confirm rather than apply automatically. For reflexive TA specifically, treat AI-assisted suggestions as a prompt to consider, not a shortcut past your own close reading — the analyst’s interpretive judgment is the method’s actual analytic engine, and offloading first-pass coding to a suggestion engine works against that.

Two structural choices matter here and are easy to get wrong under time pressure:

  • Code inductively, from the data, not from a pre-built list. Braun and Clarke’s phase 2 is explicitly bottom-up. Don’t import a codebook from a prior study or a supervisor’s template and start slotting extracts into it — that’s codebook TA, a legitimate but different approach with its own conventions. If your project genuinely calls for a predefined coding frame, say so explicitly in your methods section rather than letting it happen by default because NVivo made it easy to paste in existing nodes.
  • Code generously and code the same extract multiple times where it genuinely fits more than one code. NVivo doesn’t limit how many nodes a stretch of text can carry, and initial coding is meant to be comprehensive and inclusive, not selective — you thin the code list down in later phases, not now.

Use coding stripes (the colored bars NVivo displays alongside a source, one per node applied to that line) to visually verify your own coding as you go: a passage with stripes that don’t match your memory of what you coded is worth a second look before you move on. For a fully worked example of applying codes to raw interview text, see CASRAI’s coding qualitative interview data guide.

Phase 3: Searching for Themes — node hierarchies

Once initial coding is complete, phase 3 asks you to start clustering codes into candidate themes based on shared meaning, not just surface similarity. In NVivo, this is where you build a node hierarchy: create parent nodes representing a candidate theme, and drag related child nodes underneath them. NVivo’s hierarchy chart or the project map view gives you a visual layout of the structure as it develops, which is useful for spotting an accidental single-code “theme” (a parent node with only one child, usually a sign the grouping isn’t a theme yet) or an overloaded one that’s really two ideas wearing one label.

Run a matrix coding query at this stage to cross-tabulate your candidate node clusters against case attributes (participant role, site, condition) — not to test a hypothesis, but to see whether a pattern you’re calling a theme actually recurs across the dataset or is really one or two participants’ idiosyncratic framing. This is exploratory, not confirmatory: a theme in reflexive TA doesn’t need to appear in every case to be real, but a query result that shows it’s entirely confined to a single source is worth noting honestly in your analysis rather than glossing over.

Where this phase pulls against reflexive TA specifically: the node-hierarchy interface rewards early structural tidiness — a clean tree of parents and children looks finished. Braun and Clarke are explicit that phase 3 output is candidate themes, provisional and expected to be reworked or discarded in phase 4. Don’t let a hierarchy that looks complete in NVivo’s outline view substitute for the harder, slower interpretive work of asking whether a cluster of codes actually shares a central organizing concept.

Phase 4: Reviewing Themes — the two-level check

Braun and Clarke’s phase 4 has two distinct checks: do the coded extracts under each candidate theme actually cohere (level 1), and does the set of themes, taken together, accurately represent the dataset as a whole (level 2). NVivo supports level 1 directly: open each parent node and read through every extract coded to it in sequence — NVivo’s node summary or node report pulls all coded content together regardless of which source it came from, which is exactly the coherence check phase 4 asks for. Where an extract doesn’t fit, recode it, split the node into two more precise ones, or fold it into a different candidate theme.

Level 2 is harder to automate and NVivo doesn’t really try — re-read the full dataset again against your revised theme set, the same way you did in phase 1, and check whether the themes collectively tell an accurate story or whether something you’re seeing in the raw material has no theme representing it yet. This is a step performed by re-reading, not by running a query, however tempting it is to treat query output as sufficient because it’s already inside the software.

On NVivo’s coding comparison query specifically: NVivo can compare how two coders applied nodes to the same material and report an inter-rater agreement statistic (kappa or percentage agreement). This is a genuinely useful feature — for coding-reliability TA, where consistent application of a fixed codebook across coders is the whole point. It is not a fit for reflexive TA. Braun and Clarke’s own later work is explicit that reflexive TA rejects inter-rater reliability as a marker of coding quality, because it treats coding as a single-analyst (or collaboratively negotiated, not statistically reconciled) interpretive act, not a measurement exercise with a “correct” answer two coders should independently converge on. If your team is running the coding comparison query and reporting the resulting kappa in a reflexive-TA writeup, that is a methodological mismatch worth resolving before submission, not a bonus reliability figure. See CASRAI’s guide to choosing an inter-rater reliability coefficient for where that statistic genuinely belongs, and consider member checking as a validity strategy that fits reflexive TA’s actual epistemology instead.

Phase 5: Defining and Naming Themes — memos on the finished nodes

For each theme that survives phase 4, write a detailed analysis in a memo attached to that theme’s parent node: what the theme captures, its scope and boundaries, how it relates to the other themes, and a short vivid excerpt or two that illustrates it. NVivo’s node description field (a short free-text field on every node) is a reasonable place for a one-line working title, but it’s too short for the substantive definitional work phase 5 actually requires — do that in the linked memo, not the description field, so the reasoning behind the theme’s final name and scope survives as part of your audit trail rather than getting compressed into a label.

Name each theme with an analytic phrase that captures its content, not a topic-summary label. “Frustration with delayed feedback” says more than “Feedback,” and NVivo’s flat node-name field doesn’t distinguish between the two for you — the discipline has to come from how you write the name, not from the software.

Phase 6: Producing the Report — extracts, visualizations, and the audit trail

For the write-up, NVivo’s node reports export every extract coded to a theme, source by source, which is the fastest way to locate strong illustrative quotations without re-searching the raw transcripts by hand. Matrix coding query results and cluster diagrams can support a findings section visually, but treat them as illustrations of a pattern you’ve already established through the interpretive work in phases 3–5, not as the evidence itself — a matrix output describes co-occurrence, it doesn’t establish that a theme is analytically meaningful.

The memo trail built up across phases 1 through 5 is also your methods-section raw material. A reviewer or examiner asking how you moved from raw data to reported themes is asking a question your familiarization and theme-definition memos already answer, provided you wrote them as you went rather than reconstructing them retrospectively. This is the practical case for treating memo-writing as a required step in every phase above, not an optional nicety layered on afterward.

Frequently asked questions

Are NVivo nodes the same thing as themes?

No. A node is NVivo’s container for a coded label — closer to Braun and Clarke’s “code” than their “theme.” A theme is a higher-order pattern of shared meaning that usually gets built by grouping several related nodes together and then doing the interpretive work of phases 3 and 4 to confirm the grouping actually holds together conceptually. A parent node in your NVivo hierarchy is a candidate theme, not a confirmed one, until it survives that review.

Do you need NVivo to do Braun and Clarke thematic analysis?

No. Thematic analysis was developed and is still routinely conducted with manual methods — highlighting printed transcripts, spreadsheet-based coding, index cards. NVivo and other CAQDAS tools make the process more manageable and auditable at scale (particularly for large datasets or team-coded projects), but the software doesn’t perform the analysis; the analyst’s interpretive judgment does. See CASRAI’s NVivo vs. ATLAS.ti vs. MAXQDA comparison if you’re deciding whether you need a CAQDAS tool at all, and which one.

Can NVivo run inter-rater reliability statistics for thematic analysis?

It can — the coding comparison query calculates kappa or percentage agreement between two coders’ node application on the same material. Whether you should report that figure depends on which school of thematic analysis you’re using. It fits coding-reliability TA, where a fixed codebook is applied consistently across coders. It’s a poor fit for reflexive TA, which explicitly rejects inter-rater reliability as a quality marker in favor of researcher reflexivity and depth of engagement with the data.

What’s the practical difference between coding and theme development in this workflow?

Coding (phase 2) is applying nodes to extracts — a largely mechanical, if judgment-heavy, tagging task. Theme development (phases 3–5) is deciding which codes share an underlying pattern of meaning, testing that grouping against the full dataset, and writing an account of what the pattern actually means. NVivo’s interface can make the two feel like the same activity, because both involve dragging things into a node tree, but Braun and Clarke treat them as distinct analytic stages with different goals, and collapsing them is a common source of thin, code-list-shaped “themes” in published TA writeups.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · included with Regulatory Radar

Ask about Using NVivo for Braun and Clarke Thematic Analysis

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.