
Thematic Analysis
- 3 installs
- 3.2k repo stars
- Updated August 4, 2026
- brycewang-stanford/awesome-agent-skills-for-empirical-research
thematic-analysis is a Claude skill that conducts a rigorous thematic analysis of qualitative data following Braun and Clarke's six-phase framework.
About
This skill walks a user through conducting a thematic analysis of qualitative data using Braun and Clarke's 2006 six-phase framework. A researcher uses it to code interviews, focus groups, or open-ended survey responses into themes and write up a findings section. It outputs a Word document and an annotated thematic map, and does not cover IPA, grounded theory, or discourse analysis.
- Runs Braun & Clarke's six-phase thematic analysis on qualitative data
- Produces a Word document write-up and an annotated thematic map
- Includes a 15-point quality checklist and five common pitfalls
Thematic Analysis by the numbers
- 3 all-time installs (skills.sh)
- Ranked #1,661 of 2,064 Data Science & ML skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
thematic-analysis capabilities & compatibility
- Capabilities
- qualitative coding · research synthesis
- Use cases
- research · documentation
What thematic-analysis says it does
Conduct rigorous thematic analysis (TA) of qualitative data following Braun and Clarke's (2006) six-phase framework.
Produces a Word document write-up and an annotated thematic map.
Covers all six phases, the four upfront analytic decisions, the 15-point quality checklist, and the five common pitfalls.
npx skills add https://github.com/brycewang-stanford/awesome-agent-skills-for-empirical-research --skill thematic-analysisAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 3 |
|---|---|
| repo stars | ★ 3.2k |
| Last updated | August 4, 2026 |
| Repository | brycewang-stanford/awesome-agent-skills-for-empirical-research ↗ |
What it does
A qualitative researcher codes interview or focus-group transcripts into themes and produces a rigorous findings write-up.
Who is it for?
Researchers coding interview, focus-group, or open-ended survey data into themes.
Skip if: IPA, grounded theory, discourse analysis, conversation analysis, or narrative analysis.
When should I use this skill?
The user mentions thematic analysis, TA, Braun and Clarke, qualitative coding, or identifying themes in transcripts.
What you get
A defensible themed findings section, a Word write-up, and an annotated thematic map.
- Word document write-up
- annotated thematic map
By the numbers
- six-phase framework
- 15-point quality checklist
- five common pitfalls
Files
Thematic Analysis — Braun & Clarke's Six-Phase Framework
This skill walks a user through conducting a rigorous thematic analysis (TA) on qualitative data, following the six-phase framework from Braun and Clarke (2006). It produces a Word document (.docx) write-up of the analysis and an annotated thematic map (PNG).
The skill is grounded in one source:
Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77–101.
Where this skill cites the paper, treat those statements as the method's published position, not Claude's own.
Before you begin
Read these reference files as needed:
references/upfront-decisions.md— The four analytic decisions to settle before coding starts. Consult during Phase 1 (interview).references/coding-guide.md— How to generate codes well. Consult during Phase 3 (generating initial codes).references/theme-development.md— How to move from codes to themes, with worked examples. Consult during Phases 4–6.references/thematic-map.md— How to build and annotate the thematic map. Consult during Phase 6.references/quality-checklist.md— The 15-point checklist for assessing the analysis. Consult before producing the final write-up.references/pitfalls.md— The five common pitfalls. Consult after the first draft of the write-up.
Also read these skills before generating outputs:
- docx skill (
/mnt/skills/public/docx/SKILL.md) — Required for the Word document. - apa-referencing skill (
/mnt/skills/user/apa-referencing/SKILL.md) — If the user wants citations to existing literature in the analysis, format them in APA 7th Edition.
If the user has a writing-style skill, do not apply it to the manuscript body — see "Writing register" under Phase 7. A writing-style skill may still apply to ancillary outputs (a plain-language summary, a blog version of the findings) if the user asks for those separately.
---
Step 0: Establish the research question(s) or objective(s) first
Before any of the six phases begin, elicit the research question(s) or objective(s) explicitly. This is the first action of the skill and is non-negotiable. The research question disciplines what counts as interesting in the data, which codes earn their keep, and which patterns rise to the level of a theme. Coding without a clear question tends to drift into surface description.
Prompt the user along these lines:
Before we begin the analysis, please state the research question(s) or objective(s) for this study. If there is more than one, list them in order of priority. If they are still in draft form, share the draft — we can sharpen them together before coding starts.
If the user is unsure or only has a study aim, help them work a draft into a workable analytic question. A good TA research question is broad enough to allow patterned meaning to surface across the data set, but narrow enough to discipline what is included and excluded.
Record the agreed research question(s) verbatim. They will be referenced explicitly in every subsequent phase:
- Phase 3 (coding): when generating codes, keep the research question in view. Ask of each segment, "does this speak to the research question, directly or obliquely?" Code inclusively, but the question anchors the work.
- Phase 4 (searching for themes): themes must capture something important in relation to the research question, not just frequent content.
- Phase 5 (reviewing themes): the second-level review checks the candidate map against the data set as a whole — also check it against the research question. A theme that is internally coherent but unrelated to the question is a candidate for the discard pile.
- Phase 7 (report): the introduction states the research question(s) verbatim, and the findings are organised to answer them.
If the analytic approach is theoretical/deductive, the research question is also tied to the theoretical framework being applied — make this link explicit at this stage, before any coding begins.
Save the agreed research question(s) to the workspace as step0_research_questions.md. Refer back to this file at the start of each subsequent phase.
---
Phase 1 (skill workflow): Interview
Before any analysis, gather what is needed to plan the TA. Offer the user two paths up front.
Path A: Upload existing materials
Ask the user whether they have any of the following:
- Interview, focus group or open-ended survey transcripts (the data corpus itself)
- A research question or research aim document
- A proposal or protocol that specifies the analytic approach
- An existing coding frame or codebook (for theoretical/deductive work)
- Field notes, memos, or reflexive journal entries
Read uploaded transcripts using the appropriate tool (file-reading skill for .txt/.md, docx skill for .docx, pdf-reading skill for PDFs, xlsx skill for spreadsheet-formatted survey data). Then summarise what is in the corpus and ask the user to confirm.
Path B: Conversational interview
If the user has no materials, gather the essentials conversationally. Adapt to what they offer; do not interrogate.
Essential information to collect
About the project:
- Working title of the study
- Research question(s) — already collected in Step 0; carry these forward, do not re-elicit
- The data corpus: what kind of data, how many items, who from
- The data set for this particular analysis (may be the whole corpus or a subset — see Braun & Clarke, p. 6)
The four upfront analytic decisions (see references/upfront-decisions.md for full guidance):
1. Rich description of the whole data set, or detailed account of one aspect? 2. Inductive (bottom-up) or theoretical/deductive (top-down) coding? 3. Semantic themes (surface meaning) or latent themes (underlying ideas, assumptions, ideologies)? 4. Epistemology: essentialist/realist, contextualist, or constructionist?
These decisions are inter-related. Tendencies cluster: realist + semantic + inductive + rich description; constructionist + latent + theoretical + detailed account. But other combinations are valid — what matters is that the choices are explicit and internally consistent.
Walk the user through each decision. Do not assume realist + semantic + inductive by default just because the paper notes this is the common (often unspoken) default. Ask.
Confirm the plan
Before moving to Phase 2, produce a short plan summary and ask the user to confirm:
- Research question
- Data set (what items, how many)
- The four decisions (with one-sentence rationale for each)
- Whether engagement with prior literature happens before or after coding (inductive work usually delays it; theoretical work requires it upfront)
---
Phase 2 (skill workflow): Familiarising yourself with the data
The first of Braun and Clarke's six phases. This phase is immersion.
Ask the user to confirm that transcription (if needed) has been done. The transcript must be at minimum a rigorous orthographic verbatim record — every word spoken, including non-verbal utterances where they carry meaning (laughter, sighs, "um", "you know"). TA does not require Jefferson-style detail.
In this phase:
- Read every data item at least once before coding starts.
- Read actively — search for meanings, oddities, contradictions, patterns.
- Take notes as you read. Jot down initial ideas, hunches, and possible codes. These are not yet codes; they are a starting list to feed Phase 3.
Output of this phase: a familiarisation note for the user — a paragraph per data item summarising what struck you, plus a running list of initial ideas across the data set. Save this to the workspace as phase2_familiarisation.md.
If the data set is too large for full re-reading in one pass, do it in batches and combine the notes.
---
Phase 3 (skill workflow): Generating initial codes
Before coding starts, re-read the research question(s) saved in step0_research_questions.md. Coding is inclusive but not undisciplined — the question is the compass.
A code identifies a feature of the data — semantic content or latent meaning — that appears interesting to the analyst. A code is the most basic segment of raw data that can be assessed in a meaningful way (Braun & Clarke, 2006, p. 18, citing Boyatzis).
Codes are not themes. Codes are smaller, narrower, more numerous. Themes come later.
For full guidance on what good coding looks like (including data-driven vs theory-driven approaches, manual vs software coding, inclusive coding, and contradictions), read references/coding-guide.md.
In this phase:
- Work systematically through every data item. Give equal attention to each.
- Code for as many potential themes/patterns as possible — you do not yet know what will matter.
- Code extracts inclusively — keep a little surrounding context so meaning is not lost.
- A single extract can be coded under multiple codes, or none.
- Retain accounts that depart from the dominant story; do not smooth them out.
Output of this phase: a coded data table. For each data item, list the extracts and the code(s) applied to each. Save as phase3_codes.md. At the end, produce a consolidated code list with every code and the data extracts that sit under it.
A short worked example showing data → code, modelled on Braun and Clarke's Figure 1:
| Data extract | Codes applied |
|---|---|
| "it's too much like hard work I mean how much paper have you got to sign to change a flippin' name no I I mean no I no we we have thought about it half heartedly and thought no no I jus- I can't be bothered" | (1) Talked about with partner; (2) Too much hassle to change name |
---
Phase 4 (skill workflow): Searching for themes
A theme captures something important about the data in relation to the research question, and represents some level of patterned response or meaning across the data set.
Prevalence matters but is not decisive. A theme can appear in many items briefly, or in a few items at length. Researcher judgement — guided by the research question — decides what is a theme.
In this phase:
- Sort codes into candidate themes. Some codes will become themes, some will be sub-themes, some will be discarded, some will sit in a temporary "miscellaneous" pile.
- Look for relationships between codes, between themes, and between levels of themes (overarching themes vs sub-themes).
- Produce an initial thematic map — a visual sketch (mind-map style) showing candidate themes and how codes feed into them. See
references/thematic-map.md.
Output of this phase: a draft thematic map (saved as phase4_initial_map.png or as a markdown outline if a visual is not yet practical) and a candidate theme list with the codes under each.
End this phase with candidate themes, sub-themes, and all coded extracts grouped under them. Do not discard anything yet — Phase 5 will tell you whether the themes hold.
---
Phase 5 (skill workflow): Reviewing themes
Refining the candidate themes. Some candidate themes will not survive. Some will collapse together. Some will split.
Use Patton's dual criterion (cited in Braun & Clarke, 2006, p. 20):
- Internal homogeneity — data within a theme cohere meaningfully.
- External heterogeneity — themes are clearly distinct from each other.
This phase has two levels of review.
Level 1 — Review at the level of the coded extracts. Read all the collated extracts under each candidate theme. Do they form a coherent pattern? If yes, move on. If no, decide whether the theme is broken or whether some extracts simply belong elsewhere. Rework as needed.
Level 2 — Review against the entire data set. Re-read the full data set. Two questions: (a) Does the candidate thematic map accurately reflect the meanings in the data set as a whole? (b) Has any new relevant data been missed in earlier coding? If so, code it now.
When refinements stop adding anything substantial, stop. Endless re-coding has diminishing returns — Braun and Clarke compare further fiddling to "rearranging the hundreds and thousands on an already nicely decorated cake" (p. 21).
Output of this phase: a refined thematic map (phase5_refined_map.png) and a refined theme list.
---
Phase 6 (skill workflow): Defining and naming themes
Now define what each theme is and what it is not.
For each theme:
- Identify the essence — what aspect of the data this theme captures.
- Write a short definition (a couple of sentences). If you cannot describe the scope of a theme in two or three sentences, the theme needs more work.
- Identify any sub-themes (themes within a theme). Use these to give structure to large themes.
- Give the theme a concise, punchy name that immediately signals to the reader what the theme is about. Working titles from earlier phases are usually too descriptive — sharpen them now.
- Check the theme against the others: does it overlap too much? Does it add something distinctive to the overall story?
Output of this phase: the final theme list with definitions, sub-themes, and final names. Save as phase6_definitions.md.
Also produce the final thematic map (phase6_final_map.png) — this is the version that will appear in the write-up.
---
Phase 7 (skill workflow): Producing the report
The final write-up. This is the last phase of Braun and Clarke's framework and the deliverable of the skill.
Run the quality checklist first
Before drafting the report, run through the 15-point checklist in references/quality-checklist.md. Flag any items the analysis does not yet meet and fix them.
Then read references/pitfalls.md and audit the draft against the five common pitfalls. The most frequent failures: (1) describing extracts instead of analysing them, and (2) using interview questions as themes.
Writing register: strictly academic
The write-up must use a formal academic register suitable for peer-reviewed publication. This is the deliverable standard for the manuscript body and it overrides any personal writing-style skill the user has loaded. Those preferences apply to blogs, op-eds and informal pieces — not to the findings of a thematic analysis.
Concretely, the manuscript body follows these conventions:
- Third person throughout. No "I think", "I feel", "in my view". The reflexivity statement (if the user wants one) is the only place where first person is admissible, and it is used sparingly.
- Hedged, evidenced claims: "Participants tended to...", "The data suggest...", "This pattern is consistent with...", rather than "Participants clearly...", "Obviously...".
- No colloquialisms, no figurative analogies, no rhetorical questions to the reader, no exclamations.
- Active voice when describing what the analyst did ("I identified...", "The analysis generated...", "This study constructed..."). Do not use the passive voice as a vehicle for evasion ("themes emerged" is forbidden — see Important Reminders).
- Tense conventions: methods in past tense; results in present tense (or past when describing what participants said); discussion blends both as appropriate.
- Citations integrated grammatically, not appended as parenthetical afterthoughts. APA 7th Edition throughout — use the apa-referencing skill for any reference list entries and in-text citations.
- Sentences disciplined: one main idea per sentence where possible. Paragraphs follow a claim → evidence (extract plus analytic commentary) → link-back structure, where the link-back ties the point to the research question or the broader theme.
- No metadiscourse padding. Avoid "It is interesting to note that...", "It should be mentioned that...", "It is worth pointing out...". State the point directly.
- Analytic commentary on extracts must go beyond paraphrase. Address what the extract means, what it assumes, what it implies, and why participants might frame the matter in this way rather than another (see Braun & Clarke, 2006, p. 24).
If the user has a writing-style skill loaded, apply it only to ancillary outputs they request separately — for instance, a plain-language summary or a blog adaptation of the findings — not to the manuscript itself.
Generate the Word document
Use the docx skill to produce a manuscript-style .docx with this structure:
Title
Author / affiliation (if provided)
1. Introduction
- Research question(s) and rationale
- Brief note on the analytic approach and the four decisions
(e.g. "An inductive, semantic, realist thematic analysis was conducted,
aiming for a rich description across the full data set.")
2. Method
- Data corpus and data set
- Participants / sources (anonymised)
- Data collection (brief, if relevant)
- Analytic procedure — describe the six phases in your own words,
citing Braun and Clarke (2006). Make the "how" explicit, not implicit.
- Researcher positionality / reflexivity (if the user wants this)
3. Findings
- Overview paragraph that names the themes and sketches the overall story
- One section per theme. For each theme:
* A definition paragraph
* Sub-themes (if any), each with a brief definition
* 2 to 4 illustrative data extracts per theme, each followed by
analytic commentary that goes BEYOND paraphrase
* Where relevant, link to existing literature
- Include the final thematic map as a figure
4. Discussion (optional, depending on what the user wants)
- Overall story across themes
- Theoretical implications
- Practical implications
- Limitations
- Future directions
References (APA 7th Edition, including Braun & Clarke, 2006)For the findings section, do not paraphrase the extracts — paraphrasing is the most common failure of weak TA. The commentary should answer questions like: What does this theme mean? What assumptions underpin it? What are its implications? Why might participants talk about this in this way rather than another? (See Braun & Clarke, 2006, p. 24.)
Data extracts in the report should be:
- Verbatim from the transcript
- Anonymised (use participant pseudonyms or codes, e.g. P03, Kate F07a)
- Long enough to retain meaning, short enough to be readable — typically one to four sentences per extract
- Vivid and illustrative — pick the extracts that capture the point most clearly, not the first one you find
Save the file as <study_title>_thematic_analysis.docx in /mnt/user-data/outputs/.
Include the thematic map as a figure
The thematic map (phase6_final_map.png) goes into the findings section as a figure. See references/thematic-map.md for how to generate it (use matplotlib with networkx-style layout, or a simple node-and-edge diagram).
Caption the figure with theme names, sub-theme names, and a short explanation of relationships if relevant.
Present the files
Use present_files to give the user the .docx and the .png. Lead with the .docx.
---
Iteration
After the first version is produced, expect revisions. Common iteration requests:
- Rename a theme
- Split or merge themes
- Add or remove an extract
- Re-balance commentary vs description
- Shift register (more academic / more accessible)
- Re-draw the thematic map
- Add or trim the discussion section
Treat each iteration as a targeted edit, not a full rewrite, unless the user asks for one.
---
Important reminders
- TA is a method, not just a technique. Make the four upfront decisions explicit in the write-up.
- Themes do not "emerge". The analyst identifies them. Do not use "themes emerged" in the write-up — it is a passive construction that hides the analyst's active role (see Braun & Clarke, 2006, pp. 7, Note 4, and item 15 of the checklist).
- Prevalence does not require a number. If you describe prevalence, be consistent ("most participants", "a number of participants", "in over half the data set") and pick a unit of analysis (data item, participant, occurrence) and stick to it.
- Coding can take much longer than expected. Do not rush it. Phase 1 (familiarisation) is the bedrock of everything that follows.
- If the data are weak (thin, surface-level, common-sense), a good analysis is much harder. Flag this to the user early rather than over-claiming at the write-up stage.
Coding Guide
Detailed guidance for Phase 3 (generating initial codes).
What a code is
A code identifies a feature of the data — either semantic content or latent meaning — that the analyst finds interesting. A code is the most basic segment of raw data that can be assessed in a meaningful way (Boyatzis, 1998, cited in Braun & Clarke, 2006, p. 18).
Codes are not themes. Codes are smaller, narrower, and more numerous. Themes are built from codes in later phases.
Data-driven vs theory-driven coding
The coding approach should match Decision 2 (inductive vs theoretical).
Data-driven coding (inductive)
- Codes come from what is in the data.
- Read the data without trying to fit it into a pre-existing frame.
- Code diversely — be open to anything that catches your attention.
- Codes may have no obvious connection to the questions that were asked.
Theory-driven coding (theoretical)
- Codes come from the researcher's analytic interest or theoretical frame.
- Approach the data with specific questions in mind.
- Code for those specific features. Other features may be left uncoded if they are outside the scope of the question.
Practical coding rules
1. Work systematically through every data item. Give equal attention to each. Do not over-code the early transcripts and under-code the later ones.
2. Code for as many potential patterns as possible, time permitting. You do not yet know what will matter in later phases. It is far cheaper to discard a code in Phase 4 than to go back and re-code in Phase 5.
3. Code extracts inclusively. Keep a little surrounding context so meaning is not lost when the extract is read in isolation. A common criticism of coding is decontextualisation — guard against it.
4. A single extract can have multiple codes, or none. Do not force one-code-per-extract.
5. Retain accounts that depart from the dominant story. Contradictions and outliers are not problems to be smoothed out. They are signals. Code them. They will earn their place in the analysis if they are real.
6. No data set is without contradiction. A theme that supposedly applies to 100% of the data without nuance is suspicious. Expect tensions and inconsistencies, and represent them honestly.
Manual vs software coding
Either is fine. Choose what fits the user's workflow.
Manual options:
- Annotate transcripts with notes in the margin
- Use highlighters or coloured pens for different codes
- Use sticky notes to tag segments
- Copy extracts into a table or document, one row per extract, with codes assigned
Software options:
- NVivo, Atlas.ti, MAXQDA — paid, full-featured
- Taguette, Quirkos — lighter alternatives
- Word documents with comments, or a spreadsheet with one row per extract
- Any of these is acceptable for TA
What matters is that the coding can be reviewed, refined, and audited.
What good codes look like
A good code:
- Names the feature concisely. "Too much hassle to change name" is a code. "Things people said" is not.
- Is descriptive enough that another analyst could apply it consistently.
- Is at the right level of granularity — not so broad it becomes a theme, not so narrow it applies to a single extract.
- Stays close to the data, especially at the inductive/semantic end.
A simple format for tracking codes
A two-column table per data item works well:
| Data extract (verbatim, with speaker ID and approximate line/page) | Codes applied |
|---|---|
| "I just couldn't see myself doing it. Not after everything that happened." (P03, p.4) | (1) Reluctance based on past experience; (2) Self-concept |
| "My mum thinks I should. My sister thinks I shouldn't. I'm stuck." (P03, p.5) | (3) Family pressure; (4) Conflicting advice; (5) Indecision |
At the end of the phase, produce a consolidated code list: one row per code, with every extract assigned to that code aggregated underneath it. This becomes the input to Phase 4.
Common coding mistakes
- Using the interview questions as codes. This is not coding, it is sorting. The coding should cut across questions.
- Coding only what is interesting to the researcher in advance. This narrows the analytic field before the data has had a chance to push back.
- Coding extracts so short they lose meaning. A single word is rarely codeable. A clause is usually the minimum.
- Coding only the dominant voices. Some participants will say more, some less. Equal attention to each data item does not mean equal volume of codes — but the analyst should listen to every voice, including the quiet or contradictory ones.
- Stopping too early. First-pass coding usually under-codes. A second pass after a day or two often catches material the first pass missed.
When to stop coding
Phase 3 ends when:
- Every data item has been worked through systematically.
- The code list feels saturated — new data items are mostly producing codes that already exist, with occasional additions.
- The consolidated code list is ready to be sorted into candidate themes (Phase 4).
If new items keep producing genuinely new codes, the data set may be too diverse for a single analysis, or the coding may be drifting. Pause and review.
Five Common Pitfalls
Braun and Clarke (2006, pp. 25–26) identify five common failures of thematic analysis. Audit the draft against each one before finalising the write-up.
These are not minor stylistic issues. Each one signals a substantive problem with the analysis itself, not just the prose.
---
Pitfall 1: Failure to actually analyse the data
What it looks like: A findings section that is a string of extracts with little or no analytic narrative between them. The commentary, where present, paraphrases the extracts rather than interpreting them.
Why it happens: Confusion about what "analysis" actually involves in TA. The analyst treats their job as presenting evidence to the reader, rather than making sense of it.
How to fix it: For every extract in the write-up, ask: what does this extract mean? What does it tell us about the research question? What is the analyst's claim here? If the commentary only restates the extract in different words, rewrite it to make an interpretative point.
A useful test: cover the extracts and read only the analytic commentary. Does it stand alone as an argument about the topic? If yes, the analysis is doing real work. If no, the commentary is leaning on the extracts and the analyst needs to add interpretation.
---
Pitfall 2: Using interview (or focus group) questions as themes
What it looks like: The themes in the write-up correspond one-to-one with the questions in the interview schedule. Theme 1: "Why participants chose to study online" matches Question 1 of the interview. And so on.
Why it happens: The analyst has sorted responses by question rather than analysed them. No analytic work has crossed the question boundaries.
How to fix it: Re-do Phase 4 (searching for themes) properly. Themes should cut across questions and find patterns that the interview schedule did not anticipate. If after re-doing Phase 4 the themes still mirror the questions, either (a) the questions exhausted the topic and the data set is genuinely thin, or (b) the analyst has not yet broken out of the schedule's framing. The second is more likely.
---
Pitfall 3: A weak or unconvincing analysis
What it looks like:
- Themes that overlap too much, so the reader cannot tell where one ends and another begins
- Themes that lack internal coherence — the extracts inside a theme do not hang together
- Themes supported by only one or two extracts (anecdotalism — reifying a few instances into a pattern)
- A theme that fails to capture the majority of the data it claims to be about
Why it happens: Phase 5 (reviewing themes) was rushed or skipped. Internal homogeneity and external heterogeneity have not been checked.
How to fix it: Go back to Phase 5. Run both levels of review. Be ruthless. Merge overlapping themes. Split incoherent themes. Demote thin themes to sub-themes or discard them.
It is much better to report three solid themes than six weak ones.
---
Pitfall 4: A mismatch between the data and the analytic claims
What it looks like: The commentary makes claims that the extracts do not support — or worse, the extracts appear to contradict the claims. The reader can read the same extract and reach a different conclusion than the analyst.
Why it happens: The analyst became attached to a particular reading and stopped checking it against the data. Alternative readings were not considered.
How to fix it: For every analytic claim, ask:
- Could this extract be read differently?
- Is the claim supported by the extract, or am I projecting?
- What about the extracts that do not support this claim — am I ignoring them?
A pattern in qualitative data is rarely 100% consistent. If the analysis claims it is, the reader will be suspicious. Acknowledge variation. Address contradictions explicitly rather than burying them.
Pick the most compelling extracts for the write-up — but compelling means "clearly illustrates the analytic point", not "happens to fit my preferred reading".
---
Pitfall 5: A mismatch between theory and analytic claims
What it looks like:
- The Method section claims an inductive, semantic, realist approach
- The findings section then makes constructionist claims about how participants' accounts are shaped by patriarchal discourses
Or:
- The Method section claims a constructionist analysis
- The findings section then treats participants' talk as a transparent window onto their inner experiences
Why it happens: The four upfront decisions were made on paper but not actually applied to the analysis. The theoretical framework and the analytic moves do not match.
How to fix it: Either (a) revise the analysis to match the stated theoretical position, or (b) revise the stated theoretical position to match the analysis that was actually done. Do not paper over the gap.
This is the most subtle pitfall — it requires the analyst to read their own work with theoretical awareness. The quality checklist (items 12, 13, 14) helps catch it.
---
A bonus failure that is not numbered but is mentioned in the paper
Failing to spell out the theoretical assumptions or the analytic process at all.
Even a good analysis loses credibility if the method section is silent on:
- Which version of TA was used
- The four upfront decisions
- How the six phases were actually applied
- Who did the coding (sole analyst, multiple coders, agreement procedures, etc.)
Method opacity does not protect the analysis. It only makes the reader (or reviewer) suspicious.
---
How to audit the draft
After the first draft of the write-up:
1. Read the findings section with each pitfall in mind, one at a time. 2. Mark every place where a pitfall is present, even partially. 3. For Pitfalls 1, 2, and 5, the fix is usually in the analysis itself, not the prose. Be prepared to go back to earlier phases. 4. For Pitfalls 3 and 4, the fix is usually in Phase 5 or Phase 6 — revisit theme review and theme definition. 5. Re-draft.
The pitfalls are not independent. Pitfall 1 (description not analysis) and Pitfall 5 (theory–analysis mismatch) often co-occur, because both stem from the analyst not fully committing to an interpretative stance. Fixing one often eases the other.
15-Point Quality Checklist
Adapted from Braun and Clarke (2006, Table 2). Run through this checklist before producing the final write-up. Flag every item the analysis does not meet, then fix it.
The checklist is organised into five process areas: transcription, coding, analysis, overall, and written report.
---
Transcription
1. The data have been transcribed to an appropriate level of detail, and the transcripts have been checked against the recordings for "accuracy".
For TA, a rigorous orthographic verbatim transcript is sufficient. Jefferson-style detail is not required, but transcripts should capture every word, plus non-verbal utterances where they carry meaning (laughter, sighs, "um", "you know").
If transcripts were prepared by someone else, the analyst should still spot-check against the audio.
---
Coding
2. Each data item has been given equal attention in the coding process.
Equal attention does not mean equal volume of codes — some items will yield more than others. It means the analyst worked through every item systematically rather than concentrating effort on the most vocal or interesting participants.
3. Themes have not been generated from a few vivid examples (an anecdotal approach), but instead the coding process has been thorough, inclusive, and comprehensive.
This guards against the "anecdotalism" pitfall — reifying one or two striking examples into a pattern that does not actually exist across the data set.
4. All relevant extracts for each theme have been collated.
Every extract that fits a theme should be filed under it. Do not pick only the most quotable ones at the coding stage — that decision belongs in Phase 6 when extracts are selected for the write-up.
5. Themes have been checked against each other and back to the original data set.
Both levels of Phase 5 review must have happened — extracts against theme (Level 1) and themes against the whole data set (Level 2).
6. Themes are internally coherent, consistent, and distinctive.
Internal homogeneity (the data within a theme hold together meaningfully) and external heterogeneity (themes are clearly distinct from each other). Patton's dual criterion.
---
Analysis
7. Data have been analysed — interpreted, made sense of — rather than just paraphrased or described.
The most common failure of weak TA. The findings section must do more than restate what participants said in slightly different words. Ask of every paragraph of commentary: what does this add beyond paraphrase?
8. Analysis and data match each other — the extracts illustrate the analytic claims.
If a claim cannot be supported by an extract, the claim does not belong. If an extract suggests a different reading from the one being argued, address it — do not ignore it.
9. Analysis tells a convincing and well-organised story about the data and topic.
The themes should add up to more than the sum of their parts. Together they should answer the research question.
10. A good balance between analytic narrative and illustrative extracts is provided.
Too much extract, too little analysis = description, not analysis. Too much analysis, too little extract = claims unsupported by data. Aim for roughly balanced — extracts should anchor the analysis, not dominate it.
---
Overall
11. Enough time has been allocated to complete all phases of the analysis adequately, without rushing a phase or giving it a once-over-lightly.
Phase 1 (familiarisation) is the bedrock. Phase 3 (coding) is labour-intensive. Phases 4 and 5 require iteration. None of these should be rushed.
If the user is on a tight deadline, the analysis will suffer. Flag this honestly rather than producing a fast, weak analysis.
---
Written report
12. The assumptions about, and specific approach to, thematic analysis are clearly explicated.
The four upfront decisions must appear in the Method section. The reader should not have to infer whether the analysis is inductive or theoretical, semantic or latent, realist or constructionist.
13. There is a good fit between what you claim you do and what you show you have done — i.e. described method and reported analysis are consistent.
If the method says "inductive, semantic, realist", the findings should not slip into latent-level interpretation of underlying ideologies. Internal consistency between method and findings is essential.
14. The language and concepts used in the report are consistent with the epistemological position of the analysis.
Realist work uses language of experience, motivation, and meaning. Constructionist work uses language of discourse, social construction, and positioning.
Mixing these inconsistently signals that the epistemological position has not been worked out.
15. The researcher is positioned as active in the research process; themes do not just "emerge".
Avoid passive constructions like:
- "Themes emerged from the data"
- "The data revealed three themes"
- "Patterns were discovered in the analysis"
Use active constructions:
- "Three themes were identified" (acceptable, though still passive)
- "I identified three themes" or "We identified three themes" (better)
- "The analysis produced three themes" (acceptable)
The analyst chooses what counts as a theme. Hiding that choice behind "emergence" misrepresents the analytic process.
---
How to use the checklist
After the first draft of the write-up:
1. Read each item. 2. For each item, state explicitly whether the analysis meets it. Do not skip items. 3. For any item the analysis does not meet, fix the underlying issue rather than glossing over it in the write-up. 4. If an item is genuinely not applicable (rare), note why.
Do not include the checklist itself in the write-up — it is an internal audit, not a public artefact.
Building the Thematic Map
The thematic map is a visual representation of the analysis. It shows the themes, their sub-themes, and the relationships between them.
Braun and Clarke produce three maps over the analysis — initial (Phase 4), developed (Phase 5), and final (Phase 6). The final map goes into the write-up. The earlier maps are working documents; share them with the user during iteration if useful, but they do not have to be polished.
What the final map should show
- Overarching themes (the main themes) — usually one shape per theme, centrally placed.
- Sub-themes — smaller shapes connected to their parent theme.
- Relationships between themes, where they exist — connecting lines, with optional labels.
What it should not show:
- Every code (the map is not a code map — Phase 3 produces that)
- Discarded candidate themes from earlier phases
- Counts, percentages, or other numeric clutter
- Decorative elements that do not carry analytic meaning
Braun and Clarke's Figure 4 is the simplest version: two overarching themes ("Vagina as asset" and "Vagina as liability"), each with three sub-themes radiating outward. No connecting lines between themes. Clean and readable.
How to draw it — matplotlib approach
For a final map suitable for inclusion in the docx, use matplotlib. The output should be a PNG at 300 DPI (or higher) so it renders cleanly in print.
A simple radial layout works for most cases:
import matplotlib.pyplot as plt
import matplotlib.patches as mpatches
import numpy as np
fig, ax = plt.subplots(figsize=(12, 8), dpi=300)
ax.set_xlim(0, 12)
ax.set_ylim(0, 8)
ax.axis('off')
# Style settings — keep monochrome or low-saturation
THEME_FACE = '#E8E8E8'
SUBTHEME_FACE = '#FFFFFF'
EDGE = '#333333'
FONT = {'family': 'serif', 'fontsize': 11}
def draw_theme(ax, x, y, label, width=2.6, height=1.0):
"""Draw an overarching theme as an ellipse."""
ellipse = mpatches.Ellipse((x, y), width, height,
facecolor=THEME_FACE,
edgecolor=EDGE, linewidth=1.5)
ax.add_patch(ellipse)
ax.text(x, y, label, ha='center', va='center',
fontweight='bold', **FONT)
def draw_subtheme(ax, x, y, label, width=2.0, height=0.6):
"""Draw a sub-theme as a rectangle with rounded corners."""
rect = mpatches.FancyBboxPatch((x - width/2, y - height/2),
width, height,
boxstyle="round,pad=0.05",
facecolor=SUBTHEME_FACE,
edgecolor=EDGE, linewidth=1.0)
ax.add_patch(rect)
ax.text(x, y, label, ha='center', va='center', **FONT)
def connect(ax, x1, y1, x2, y2):
ax.plot([x1, x2], [y1, y2], color=EDGE, linewidth=0.8, zorder=0)
# Theme 1 — left side
draw_theme(ax, 3, 6, 'Theme name 1')
draw_subtheme(ax, 1, 4, 'Sub-theme 1a')
draw_subtheme(ax, 3, 3.5, 'Sub-theme 1b')
draw_subtheme(ax, 5, 4, 'Sub-theme 1c')
connect(ax, 3, 5.5, 1, 4.3)
connect(ax, 3, 5.5, 3, 3.8)
connect(ax, 3, 5.5, 5, 4.3)
# Theme 2 — right side
draw_theme(ax, 9, 6, 'Theme name 2')
draw_subtheme(ax, 7, 4, 'Sub-theme 2a')
draw_subtheme(ax, 9, 3.5, 'Sub-theme 2b')
draw_subtheme(ax, 11, 4, 'Sub-theme 2c')
connect(ax, 9, 5.5, 7, 4.3)
connect(ax, 9, 5.5, 9, 3.8)
connect(ax, 9, 5.5, 11, 4.3)
plt.tight_layout()
plt.savefig('phase6_final_map.png', dpi=300, bbox_inches='tight',
facecolor='white')
plt.close()Adapt the layout to the number of themes and sub-themes. For three or more themes, arrange them in a triangle or along the top of the figure rather than left-right.
When themes have relationships
If overarching themes are related (e.g. one feeds into another, or they sit in tension), add a connecting line between them with a short label:
ax.annotate('', xy=(7.7, 6), xytext=(4.3, 6),
arrowprops=dict(arrowstyle='<->', color=EDGE, lw=1.0))
ax.text(6, 6.3, 'in tension with', ha='center', style='italic',
fontsize=10, color=EDGE)Map captions in the write-up
The figure caption should be more than "Figure 1: Thematic map". Give the reader something useful:
Figure 1. Final thematic map showing two overarching themes ("Deciding in the dark" and "Trusting the gut") and the six sub-themes that constitute them. Sub-themes within each overarching theme are presented in no particular order.
If there are relationships between themes, name them in the caption.
When to save which version
| Phase | File name | What it shows |
|---|---|---|
| 4 | phase4_initial_map.png | Candidate themes, rough — may have many themes, some of which will be discarded |
| 5 | phase5_refined_map.png | Refined themes after Phase 5 review |
| 6 | phase6_final_map.png | The final map for the write-up |
Only the Phase 6 map needs to be polished. Phases 4 and 5 maps can be quick sketches — they exist to help the analyst think, not to be read by the user.
Style notes
- Keep the map readable in greyscale. Many readers print papers in black and white.
- Avoid colour-coding by category if the meaning depends on the colour — readers with colour-vision differences may miss the distinction.
- Use one shape for themes and a different shape for sub-themes. This carries meaning even without colour.
- Keep labels short. Long labels make the map hard to read. If the theme name does not fit, the theme name may be too long for the write-up too.
Theme Development
Detailed guidance for Phases 4 to 6 — the move from codes to themes.
What a theme is
A theme captures something important about the data in relation to the research question and represents some level of patterned response or meaning across the data set (Braun & Clarke, 2006, p. 10).
A theme is not the same as a code:
- A code is a feature of a single extract or a small set of extracts.
- A theme is a higher-level pattern built from multiple codes.
Prevalence is not the test
The "keyness" of a theme does not depend on a quantifiable measure. A theme can:
- Appear in many data items briefly
- Appear in a few data items at length
- Be present in only part of the data set, but still capture something central to the research question
There is no hard rule like "must appear in 50% of items". Researcher judgement, guided by the research question, decides what is a theme.
That said, prevalence still matters for reporting. See quality-checklist.md for how to describe prevalence without quantifying it falsely.
Three things to do in Phase 4 (Searching for themes)
1. Sort the codes. Put each code on a card (physical or digital) and group them. Some groupings will be obvious; some will need rearranging. Use mind-maps, tables, or sticky notes — whichever the user prefers.
2. Identify candidate themes and candidate sub-themes. Some codes will become themes. Some will be sub-themes within a larger theme. Some will not fit anywhere — park these in a "miscellaneous" pile for now.
3. Sketch a thematic map. Show the candidate themes and how codes feed into them. This is the initial thematic map — it is rough, and that is fine. See thematic-map.md.
At the end of Phase 4: a draft thematic map and a candidate theme list with all coded extracts grouped under each theme.
Two levels of review in Phase 5
Level 1 — Check the themes against their coded extracts.
For each candidate theme, read all the collated extracts under it. Ask:
- Do these extracts cohere meaningfully?
- Is the theme telling one consistent story, or two?
- Are any extracts misfiled?
Decisions to make:
- If the theme holds, move on.
- If the theme is incoherent, rework it. Either redefine the theme, move the extracts to a different theme, or split the theme into two.
- If a theme has too few supporting extracts, it may not be a theme. Consider folding it into a sub-theme or discarding it.
Level 2 — Check the themes against the whole data set.
Re-read the entire data set. Ask:
- Does the thematic map accurately reflect the data set as a whole?
- Has any relevant material been missed during coding? Code it now.
This re-reading often produces minor adjustments rather than wholesale change. If it produces wholesale change, that may mean Phase 3 was rushed.
When to stop refining
Re-coding can go on indefinitely. Stop when refinements stop adding anything substantive. Braun and Clarke's metaphor (2006, p. 21): further fiddling is like "rearranging the hundreds and thousands on an already nicely decorated cake".
A practical heuristic: if two rounds of review produce only cosmetic changes, the analysis has settled.
Phase 6: Defining and naming themes
For each theme, do four things:
1. Identify the essence
What is this theme really about? Write a single sentence that captures the core idea.
If you cannot describe the scope of the theme in two or three sentences, it is too diffuse. Either tighten the definition or split the theme.
2. Identify sub-themes (if any)
Sub-themes are themes within a theme. They are useful for:
- Giving structure to a large or complex theme
- Showing hierarchy of meaning in the data
- Distinguishing between related but distinct aspects of one overall idea
Example from Braun and Clarke (2006, p. 22), based on Braun and Wilkinson's (2003) study of women's talk about the vagina:
- Overarching theme: Vagina as liability — sub-themes: nastiness and dirtiness; anxieties; vulnerability.
- Overarching theme: Vagina as asset — sub-themes: satisfaction; power; pleasure.
Each sub-theme has its own definition and its own extracts. Together they constitute the overarching theme.
3. Check against other themes
For each theme:
- Does it overlap significantly with another theme? If yes, consider merging or sharpening the distinction.
- Does it add something distinctive to the overall story? If not, it may be redundant.
4. Give it a final name
Working titles are often too long or too descriptive. Final names should be:
- Concise — typically three to seven words
- Punchy — they should grab attention
- Immediately informative — the reader should know what the theme is about from the name alone
Compare:
- Working title: "How participants talked about not having enough information when making the decision"
- Final name: "Deciding in the dark"
The final name should sit comfortably as a sub-heading in the findings section.
The overall story
By the end of Phase 6, the themes together should tell a coherent story about the data — a story that answers the research question.
A useful test: write a one-paragraph synopsis of the overall analysis using only the theme names and definitions. If the paragraph does not make sense as a standalone story, the themes have not yet cohered.
What weak theme development looks like
- Themes that overlap so much they could be merged
- Themes whose extracts could just as easily live under another theme
- Themes that paraphrase the interview question
- Themes whose definition is "things participants said about X" — that is a topic, not a theme
- A theme list that does not add up to a story — a collection of observations, not an analysis
The Four Upfront Decisions
Before any coding starts, the user must settle four analytic decisions. These should be discussed during Phase 1 of the skill workflow (the interview) and stated explicitly in the Method section of the write-up.
The decisions are inter-related but distinct. Walk the user through each one in turn.
---
Decision 1: Rich description of the whole data set, or detailed account of one aspect?
This decision is about scope.
Rich description of the whole data set
- The themes identified must reflect the predominant or important content across the entire data set.
- Best suited to under-researched topics, or when participants' views on the topic are not yet known.
- Some depth and complexity is lost, but the reader gets an overall picture.
- Useful for short outputs (article, dissertation chapter) where space is tight.
Detailed account of one particular aspect
- A more nuanced analysis of one theme, or a small group of related themes.
- May focus on a specific question, or on a particular latent theme running through the data.
- Example: Clarke and Kitzinger's (2004) talk show study focused on how lesbian and gay parents "normalise" their families — one specific aspect, examined in depth.
How to ask
"Are you trying to give a broad, rich picture of everything in the data — or zoom in on one specific thing that runs through it?"
---
Decision 2: Inductive or theoretical (deductive)?
This decision is about what drives the coding.
Inductive (bottom-up, data-driven)
- Themes are strongly linked to the data themselves.
- Codes do not have to map back to interview questions, and they are not driven by the researcher's prior theoretical interests.
- Bears some similarity to grounded theory (without grounded theory's theoretical commitments).
- Researchers cannot fully escape their theoretical and epistemological commitments — but they actively avoid forcing the data into a pre-existing frame.
Theoretical (top-down, deductive, analyst-driven)
- Coding is driven by the researcher's theoretical or analytic interest.
- Tends to provide less rich description overall, more depth on the chosen aspect.
- Maps onto a specific research question rather than letting one evolve through coding.
- Example: A researcher interested in Hollway's discourses of heterosex (male sexual drive, have/hold, permissive) would code for those specifically, rather than coding diversely for any heterosex talk.
How to ask
"Are you coming to the data with a specific theoretical lens or framework you want to apply — or do you want the codes to come from what is actually in the data?"
Practical implication
This decision shapes when the user engages with prior literature.
- Inductive work usually delays literature engagement so it does not narrow the analytic field of vision too early.
- Theoretical work requires literature engagement before coding starts.
---
Decision 3: Semantic or latent themes?
This decision is about the level at which themes are identified.
Semantic themes (explicit, surface)
- Themes are identified within the explicit, surface meanings of the data.
- The analyst does not look beyond what the participant said.
- Typical workflow: description → organisation into patterns → interpretation of significance (often in relation to prior literature).
- Best suited to a realist or experiential epistemology.
Latent themes (interpretative, underlying)
- Themes go beyond surface meaning.
- The analyst identifies underlying ideas, assumptions, conceptualisations, and sometimes ideologies that shape the surface content.
- The development of the theme itself involves interpretative work — analysis is theorised, not just descriptive.
- Best suited to a constructionist epistemology. Overlaps with some forms of thematic discourse analysis.
Braun and Clarke's "jelly" metaphor: a semantic approach describes the surface of the jelly; a latent approach identifies the features that gave the jelly its particular form.
A TA typically focuses primarily on one level, not both.
How to ask
"Do you want to report what participants explicitly said about the topic — or interpret the underlying ideas, assumptions, or ideologies that shape what they said?"
---
Decision 4: Epistemology — essentialist/realist, contextualist, or constructionist?
This decision is about what you assume the data represent.
Essentialist / realist
- Assumes a straightforward, largely unidirectional relationship between language, meaning, and experience.
- Language reflects and enables the articulation of meaning and experience.
- The analyst can theorise motivations, experiences, and meanings directly from the data.
- Most common (and often unspoken) default in TA. The skill should not let this go unstated.
Contextualist (e.g. critical realism)
- Sits between essentialism and constructionism.
- Acknowledges how individuals make meaning of their experience AND how the broader social context shapes those meanings.
- Retains a focus on the material and other limits of "reality".
- Useful when neither pole feels right.
Constructionist
- Meaning and experience are socially produced and reproduced, not residing inside individuals.
- TA in this paradigm does not focus on motivation or individual psychology.
- Instead, it theorises the socio-cultural contexts and structural conditions that enable the individual accounts.
- Latent TA tends to be constructionist (not all latent TA is, but the overlap is strong).
How to ask
"What do you take the data to be evidence of? Direct accounts of participants' experiences and feelings (realist)? Socially shaped accounts that reflect broader cultural discourses (constructionist)? Or somewhere in between (contextualist)?"
---
How the decisions cluster
Tendencies, not rules:
- Cluster A — Realist + semantic + inductive + rich description across the data set. The most common combination. Often the unspoken default.
- Cluster B — Constructionist + latent + theoretical + detailed account of a specific aspect. The more interpretative end. Overlaps with thematic discourse analysis.
Other combinations are valid. What matters is internal consistency. For example, do not code at a latent level using constructionist language and then treat participant talk as a transparent window onto inner experience.
---
What to put in the write-up
After the four decisions are settled, draft one paragraph for the Method section that states all four explicitly. Example:
"The analysis followed Braun and Clarke's (2006) six-phase framework. It was inductive (codes were generated from the data rather than from a pre-existing frame), focused at the semantic level (themes were identified within the explicit content of participants' accounts), and conducted within a realist epistemology (participant talk was taken as a reflection of their experiences and meanings). The aim was a rich description of the data set as a whole rather than a detailed account of one aspect."
Substitute the relevant choices. The point is that none of the four decisions should be left for the reader to infer.
Related skills
FAQ
What framework does this skill follow?
It follows Braun and Clarke's (2006) six-phase thematic analysis framework.
What does it produce?
A Word document write-up of the analysis and an annotated thematic map.
Does it cover grounded theory?
No. It does not cover IPA, grounded theory, discourse analysis, conversation analysis, or narrative analysis.