Skip to content
Transkio
Back to blog
TutorialsJuly 21, 2026 · 7 min read

How to Transcribe a Focus Group

Focus group transcription is hard because everyone talks at once. Here's how to record, transcribe, and label a multi-speaker session so the transcript stays usable.

By Transkio Team


Eight people in a room, one moderator, and a topic they all have opinions about. That's a focus group, and it produces the messiest audio you'll ever try to transcribe. People interrupt. Two participants agree loudly at the same time. Someone at the far end of the table mutters something the recorder barely catches. Focus group transcription is a real skill because the recording is working against you from the first minute.

If you run these sessions for market research, UX studies, or policy work, this guide walks through getting a transcript that's actually usable — from mic placement to labeling who said what when three people are talking over each other.

Why focus groups are the hard case

A one-on-one interview has two clear voices taking turns. A focus group has six to ten voices with no turn-taking discipline, wildly different distances from the mic, and constant overlap. Every automated tool struggles with this, and honestly, so do human transcriptionists.

The upside: you can stack the deck in your favor before you ever hit record. Most of the pain in focus group transcription is created in the room, not at the keyboard. If the source material for these sessions is new to you, the broader field of the focus group as a research method explains why the group dynamic — the thing that makes them valuable — is also what makes them hard to capture.

Set the room up to be recordable

  • Use more than one mic. A single recorder in the center will favor whoever sits closest. A couple of boundary mics on the table, or a mic per two or three people, dramatically improves what you capture.
  • Seat people deliberately. Spread the talkative ones out. Keep quiet participants closer to a mic.
  • Kill room noise. Hard surfaces echo. A carpeted room with soft furnishings sounds far cleaner than a glass-walled conference room.
  • Have the moderator manage overlap. A good moderator gently steers people to speak one at a time and repeats names — "Go ahead, Priya" — which turns out to be gold for later labeling.

The moderator's secret weapon: name-calling

If your moderator says a participant's name before or as they speak, you get an audio cue for who's talking that no algorithm can match. It feels a little stilted in the moment, but "What do you think, Marcus?" followed by Marcus answering means you can label that turn with confidence weeks later. Brief your moderator to do this, especially early on while you're still learning voices.

Recording the session

Once the room is set, the recording itself is simple. Capture to a device you trust — a dedicated recorder or a phone in airplane mode so a call doesn't interrupt it. Do a real sound check with people sitting where they'll actually sit, and listen back to the far end of the table specifically. That's where audio goes to die.

If it's a remote focus group over video, record locally and ask everyone to use headphones. Video call platforms often duck or cut audio when multiple people speak at once, which is exactly the moment you most want captured. A local recording of the raw meeting audio usually beats the platform's processed stream.

Consent, first

Because a focus group involves several people, consent is more involved than a solo interview. Everyone in the room needs to agree to be recorded, and they should know how the recording and transcript will be handled. If an ethics board approved your study, follow that protocol.

Worth stating plainly: Transkio is a general-purpose transcription app, not a compliance-certified service, and it doesn't offer a certified transcript or a human transcription service. If your session covers sensitive topics or your agreement restricts sending data to third-party tools, sort that out before you upload anything.

Transcribing multi-speaker audio

Here's where focus groups differ most from other jobs. You've got two problems to solve: getting the words right, and getting the attribution right.

Getting the words down

Start from an AI draft to save time, then correct heavily. With Transkio you upload the recording or capture it in the browser, and you get a text draft to edit. Expect the accuracy on group audio to be lower than on a clean solo recording — crosstalk and distance are exactly the conditions models handle worst.

AI-generated transcripts may contain errors — please review before relying on them. With focus groups that review is not optional polish; it's most of the work. Play the audio, follow along, and fix the spots where the model dropped an overlapping line or blended two people into one run of text.

Keep an eye on your limits, too. The free plan gives you 60 trial minutes and then 30 minutes a month, with files up to 100 MB — a single 90-minute group can eat that quickly, so a serious research schedule usually needs a paid tier. For the mechanics of getting spoken words into editable text, our audio-to-text tool page covers the basics.

Getting attribution right

This is the part people underestimate. Even when the words are perfect, knowing who said each line is what makes a focus group transcript analyzable — you're often comparing how different participant types respond.

Automated speaker detection (a paid feature, Elite+ in Transkio) can split a recording into distinct voices, and it helps. But be realistic: on eight-person audio with overlap, no tool nails every turn. It'll merge two similar voices or split one person across two labels. Treat auto-labeling as a first pass, then correct against your moderator's name cues and your seating notes.

When you can't tell who spoke

Sometimes you genuinely can't identify a speaker — the audio's too muddy or three people said it at once. Don't guess. Use an honest marker like [unclear speaker] or [multiple] and move on. In analysis, an honestly-flagged unknown is far safer than a confident misattribution that skews your findings toward the wrong participant group.

Labeling conventions that hold up

Pick a scheme before you start and keep it consistent:

  • Use stable codes: Moderator, P1, P2, and so on, rather than real names if you're anonymizing.
  • Keep a secured, separate key linking codes to participants.
  • One speaker turn per paragraph.
  • Bracket overlaps and interruptions if your method cares about group dynamics — [P3 interrupts] can itself be data.
  • Timestamp at turn changes so you can jump back to verify a contested quote.

From transcript to analysis

A focus group transcript usually feeds thematic analysis or coding software, so export in a format that stage can use.

Formats and exports

Plain-text and basic exports (TXT, SRT, VTT) come with the free plan. A formatted Word file — often what coding tools and collaborators prefer — is a paid export (Pro+). If you're heading into NVivo, ATLAS.ti, or MAXQDA, check what they want; clean speaker labels and consistent formatting let them auto-detect turns.

Compared with a solo interview, expect focus group prep to take longer at every step. If you also run individual interviews in the same study, our guide on research interview transcription pairs well with this one, and the general interview transcription overview covers the parts both share.

A realistic time estimate

Because you're constantly rewinding to untangle overlap and confirm speakers, focus group transcription runs longer per audio-minute than almost any other kind. Even starting from an AI draft, plan for several hours of editing per hour of recording — and more if you skipped the mic-per-few-people setup. That planning number is the single most useful thing to take from this article, because underestimating it is how researchers end up with a backlog of un-transcribed sessions.

Quick checklist

  1. Get consent from everyone in the room.
  2. Use multiple mics; seat talkers apart.
  3. Brief the moderator to say names and manage overlap.
  4. Sound-check with people in their real seats.
  5. Record locally; keep the raw audio.
  6. Generate an AI draft, then correct heavily against the audio.
  7. Attribute speakers using name cues and seating notes; flag genuine unknowns honestly.
  8. Anonymize inline with stable codes.
  9. Export in the format your analysis tool wants.

Focus groups will never be the easy transcription job — the overlap is baked into what makes them useful. But if you record with multiple mics, brief your moderator to name participants, and treat the AI draft as a starting point rather than a finished transcript, you'll end up with something you can actually code. The work moves from the keyboard back into the room, which is exactly where it's cheapest to fix.

Turn Your Next Recording Into Text.

Upload a file or record a meeting in your browser — get an accurate, editable transcript in minutes.

Transcribe for free
  • 30 free minutes, no card required
  • Transcripts in minutes, not hours
  • 50+ languages