Written and maintained by CASRAI Editorial Board
Last updated
Verbatim, denaturalised, and Jefferson transcription are not interchangeable house styles — they encode three different theories of what counts as data. A verbatim transcript treats every utterance, filler, and false start as potential evidence. A denaturalised transcript treats the words as the evidence and cleans away everything that gets in the way of reading them. A Jefferson transcript treats the pauses, overlaps, and stress patterns as the evidence itself. Picking the wrong one does not just cost re-transcription time later — it can mean the recording never captured the level of detail the analysis actually needed.
This guide compares what each convention preserves, gives a decision rule for matching convention to analytic approach, and covers how to format and cite transcript excerpts once you have them. For the interview design and data-collection side of this, see CASRAI’s guides to designing semi-structured interviews and interviews as a research method; for what happens to a transcript after it is produced, see qualitative data: coding and analysis.
The three conventions at a glance
| Convention | What it preserves | What it strips or normalises | Typical analytic use |
|---|---|---|---|
| Verbatim | Every word spoken, including fillers, false starts, and repetitions | Little to nothing — punctuation is added for readability, but no wording is removed | Thematic analysis, content analysis, grounded theory, general interview/focus-group coding |
| Denaturalised | The substance and structure of what was said | Fillers (“um,” “uh”), false starts, stutters, and non-standard grammar — cleaned toward readable prose | Narrative analysis, thematic analysis where speech production is not the object of study, member-checking drafts, most published-quote extracts |
| Jefferson | Timed pauses, overlapping speech, latching, stress, volume, pitch shifts, breathing, elongation, cut-offs, and speech rate | Nothing — this is the maximal-detail convention | Conversation analysis (CA), and any study where the sequential, interactional organisation of talk is itself the phenomenon |
The underlying distinction between “verbatim” and “denaturalised” transcription comes from a paper by Mary Jo MacLean, Manfred Meyer, and Alma Estable, who used the terms naturalized and denaturalized transcription to describe two different positions on how much of the messiness of real speech belongs in a transcript. A naturalized transcript tries to capture speech production as faithfully as possible — closer in spirit to Jefferson notation. A denaturalized transcript treats speech errors, hesitations, and dialect features as noise relative to the ideational content the researcher actually wants to analyse, and removes them.
Verbatim transcription: every utterance, unfiltered
A verbatim transcript records what was said as completely as ordinary prose punctuation allows: every filler word, every repeated word, every false start and self-correction, in the order it happened. Nothing is paraphrased and nothing is smoothed over, but a verbatim transcript stops short of Jefferson-level detail — it typically does not time pauses to the tenth of a second, mark overlapping speech precisely, or notate pitch and stress.
This is the default convention for most qualitative interview and focus-group studies, because it is the minimum level of fidelity that keeps a coder honest to what a participant actually said rather than what the transcriber assumed they meant. Thematic analysis (Braun and Clarke’s approach — see CASRAI’s step-by-step guide to thematic analysis), most content analysis, and most grounded theory coding (see the grounded theory guide) work directly from verbatim transcripts, because the unit of analysis in these approaches is the idea or theme a participant expressed, not the moment-by-moment mechanics of how they said it.
A verbatim transcript is also the right baseline when you are not yet certain which analytic approach you will use, or when multiple team members will code the same data for different purposes — it preserves the option to move to more detailed notation later, whereas a denaturalised transcript has already discarded information a later analysis might need.
Denaturalised transcription: cleaned for readability
A denaturalised transcript keeps the substance of what a participant said while removing false starts, filler words, stutters, and grammatical irregularities that do not carry analytic meaning. “Um, I guess, like, I dunno, I think it was — it was probably around, um, six months?” becomes something closer to “I think it was probably around six months.” The content and the claim are unchanged; the speech-production noise around them is gone.
This convention suits analyses where the researcher’s interest is in what was said rather than how it was produced: narrative analysis (see narrative analysis as a qualitative research method), thematic coding aimed at meaning rather than talk-in-interaction, and — very commonly — the quote extracts that actually appear in a published paper, which are almost always lightly denaturalised even when the working transcript underneath them was verbatim. It is also the more defensible choice for member-checking drafts sent back to participants, since a heavily disfluent verbatim transcript can read as an unflattering or inaccurate representation of what someone meant to say, even though it accurately captures what they said.
The trade-off is real: denaturalising discards information. Once fillers, hesitations, and false starts are gone, you cannot go back and analyse hesitation patterns, self-correction, or how confidently something was said — those are gone from the record, not just hidden from the reader. Denaturalise the transcript you present; keep the verbatim (or Jefferson) source transcript as the analytic record if there is any chance a later question will turn on how something was said.
Jefferson transcription: full paralinguistic detail for conversation analysis
Jefferson notation — developed by Gail Jefferson for conversation analysis — is the convention built specifically to capture the interactional mechanics of talk: exactly when a pause occurred and how long it lasted, precisely where one speaker’s turn overlapped another’s, where a word was stretched, cut off, spoken loudly or quietly, or produced with rising or falling pitch. In conversation analysis, these features are not incidental texture — they are the data. CA’s core claim is that turn-taking, pausing, and overlap are themselves systematically organised, and that organisation is only visible if the transcript preserves it.
The core symbol set, in its commonly taught form:
| Symbol | Meaning |
|---|---|
| (0.5) | Timed pause, in seconds and tenths of a second |
| (.) | A micro-pause, too short to time meaningfully |
| [ … ] | Overlapping speech — brackets aligned across two speakers’ lines mark where the overlap starts and ends |
| = | Latching — no gap at all between one turn ending and the next starting |
| word: | Elongation of the preceding sound; more colons mark a longer stretch |
| word- | A cut-off — the speaker stops abruptly mid-word or mid-utterance |
| word | Stress or emphasis (typically shown with underlining) |
| WORD | Noticeably louder speech than the surrounding talk |
| °word° | Noticeably quieter speech (“degree signs”) |
| ↑ / ↓ | A marked rise or fall in pitch on the following sound |
| >word< | Speech delivered faster than the surrounding talk |
| <word> | Speech delivered slower than the surrounding talk |
| .hhh / hhh | Audible in-breath / out-breath |
| ( ) | Speech the transcriber could not make out |
| ((laughs)) | Transcriber’s description of a non-verbal feature, in double parentheses |
This level of detail is expensive to produce — a single minute of natural conversation can take well over half an hour to transcribe to full Jefferson standard — so it is worth reserving for studies where the interactional organisation of talk really is the research question, rather than applying it by default to interview data that a thematic or content analysis will end up reading past anyway. CASRAI’s guide to discourse analysis and the broader overview of qualitative research methodologies cover where CA and discourse-analytic approaches sit relative to the more common thematic/content approaches.
Which convention fits your analysis method
The decision is really about what your analytic method treats as evidence, not about which convention is “more rigorous” in the abstract — a Jefferson transcript is not automatically better for a study that never needed pause timing in the first place.
| Analysis method | Convention to use | Why |
|---|---|---|
| Thematic analysis (Braun & Clarke) | Verbatim, with denaturalised quotes in the write-up | Unit of analysis is meaning/theme, not speech production; denaturalising the published extract improves readability without changing the code |
| Content analysis | Verbatim | Systematic coding of manifest content depends on complete wording; fillers are usually excluded from coding frames but not from the transcript itself |
| Grounded theory | Verbatim | Constant comparison works from participants’ own words; premature cleaning risks losing in-vivo codes drawn from exact phrasing |
| Narrative analysis | Denaturalised (sometimes lightly naturalized for storytelling features) | Focus is on story structure and meaning-making, not moment-by-moment speech production |
| Discourse analysis | Verbatim to Jefferson, depending on tradition | Critical/Foucauldian discourse analysis usually works from verbatim; discursive-psychology traditions closer to CA use Jefferson-level detail |
| Conversation analysis (CA) | Jefferson | Turn-taking, overlap, and pause timing are the object of study, not background noise |
If you are not sure yet which analysis you will run, transcribe verbatim and keep the audio. You can denaturalise a verbatim transcript for publication in minutes; you cannot add Jefferson-level pause timing and overlap marking to a transcript after the fact without re-listening to the recording from scratch, so decide before transcription — not after — whenever conversation analysis is even a possibility.
Formatting and citing transcript excerpts
A few conventions carry across all three transcription styles when you quote a transcript in a paper:
- Line numbers. Number transcript lines (or turns) so you can cite a specific excerpt precisely — “(Participant 4, lines 112–118)” rather than a vague page reference. This is standard practice for CA excerpts and increasingly expected for verbatim/denaturalised excerpts too.
- Anonymisation. Replace real names with pseudonyms or participant codes (P1, P2…) consistently across the transcript and the write-up, and remove or generalise any other identifying detail (employer names, specific locations) per your ethics approval — see CASRAI’s guide to research ethics in qualitative research for the consent and confidentiality obligations this sits inside.
- Editorial marks. Square brackets mark anything you have added or changed for clarity — “[the training programme] was a waste of time” — and an ellipsis in brackets,
[...], marks an omitted section, distinct from Jefferson’s unbracketed pause/cut-off symbols. - A symbol key. If you use Jefferson (or any non-obvious) notation in a published excerpt, include a short key — either as a footnote on first use or in a methods appendix — so a reader unfamiliar with CA conventions can follow the excerpt.
- In-text citation. Treat a quoted excerpt like any other primary data citation in your reference style (APA, for instance, treats interview/transcript data as personal communication or as an appendix-referenced source depending on whether it is publicly archived) — check your target journal’s author guidelines, since practice varies by field.
Practical transcription workflow
- Transcribe close to the interview. Doing it yourself, or reviewing an AI/vendor transcript carefully, while the interview is still fresh catches misheard words and missing context an automated tool cannot.
- Automated transcription is a first pass, not a final one for any convention beyond basic verbatim — automatic speech-recognition tools reliably miss overlapping speech, mis-time pauses, and normalise disfluencies you may specifically want preserved, so budget a manual correction pass, especially before Jefferson-level notation.
- Decide the convention before transcription starts, and put it in your protocol — retrofitting Jefferson detail onto an already-denaturalised transcript means re-transcribing from the original recording.
- Keep the original audio and a verbatim (or Jefferson) master transcript even if a denaturalised version is what appears in the paper — reviewers or a later re-analysis may need to check the underlying wording.
- If you are coding in CAQDAS software (NVivo, ATLAS.ti, MAXQDA, and similar), import the verbatim or Jefferson transcript, not a denaturalised one, so the coding record stays traceable to what was actually said — see CASRAI’s overview of CAQDAS workflows for what these tools do and don’t do with transcript formatting.
Frequently asked questions
Do I need to transcribe every “um” and “uh”?
For a verbatim transcript, yes — fillers are part of what was actually said, even if you strip them out of the quotes you eventually publish. Whether they need to survive into your coding frame is a separate question from whether they belong in the transcript itself.
Can I switch conventions partway through a study?
Not without cost. If you started denaturalising and later need Jefferson-level detail (for example, a reviewer asks for interactional evidence you did not originally plan to analyse), you will need to re-transcribe from the original recordings, because denaturalising discards the information a Jefferson transcript depends on.
Is Jefferson notation only for conversation analysis?
It is most closely associated with CA, but discursive psychology and some ethnomethodological work use it too, whenever the interactional detail of talk — not just its content — is the object of study.
Do published papers use raw verbatim or Jefferson quotes, or cleaned-up ones?
Most published extracts are denaturalised to some degree, even in studies that transcribed verbatim, because a heavily disfluent quote can be harder to read and can unintentionally make a participant sound less articulate than they intended. CA papers are the main exception — they typically publish full Jefferson-notated excerpts, because the notation itself is the evidence being presented.
Does a transcription convention affect intercoder reliability?
It can. If two coders are working from transcripts produced under different conventions (one denaturalised, one verbatim), agreement statistics can be affected by transcript differences rather than genuine coding disagreement. Keep the transcription convention constant across a study’s data set for this reason — see CASRAI’s guide to qualitative research methodologies for how this fits into the broader design decisions a qualitative study has to make consistently.








