Transcription Services

Audio Transcription Services

Clear, accurate transcripts of your recordings in 300+ languages, in the style your work needs.

Last updated

Audio transcription turns a recording into a written transcript. Verbatim transcription captures every word, filler and false start; intelligent verbatim removes those and keeps what the speaker meant; edited transcription goes further and tidies the grammar. Prism Linguistics transcribes audio in more than 300 languages, with speaker labels and timecodes, and can translate the transcript into English.

A recording sitting on a drive is hard to use. A transcript is something you can work with: a researcher can code it, a solicitor can cite it, an editor can pull quotes from it. Getting the style right for that purpose is half the job.

Choosing a style

Verbatim, intelligent verbatim and edited transcription

There is no single "correct" transcript. It depends entirely on what you are going to do with it, and picking the wrong style is the most common reason a transcript gets sent back.

What is verbatim transcription?
Verbatim transcription records everything: every "um", every false start, every repeated word, stammers, laughter and pauses, usually with non-verbal sounds marked. Nothing is smoothed. It is used where the exact words are the evidence, so legal recordings, disciplinary interviews and conversation-analysis research.
What is intelligent verbatim transcription?
Intelligent verbatim transcription keeps the speaker's own words and sentence structure but drops the fillers, stumbles and repetitions that add nothing. "So, um, I mean, I think, I think we should probably start again" becomes "I think we should probably start again". It is the default for interviews, focus groups and meetings.
What is edited transcription?
Edited transcription goes one step further and tidies grammar, cuts digressions and produces something closer to written prose. Useful for publication, show notes and speeches. It is the least faithful of the three, so it is the wrong choice for anything evidential.
Which style should I choose?
If the transcript could end up in front of a court, a regulator or a discourse analyst, choose verbatim. If it is going to be read, coded or quoted from, choose intelligent verbatim. If it is going to be published, choose edited. Tell us what it is for and we will recommend one.

Speaker identification and timecoded transcription

Two things turn a wall of text into something usable: knowing who said it, and knowing when. We label speakers throughout, and where you give us names or roles we use those rather than Speaker 1 and Speaker 2, which matters enormously when you come back to the file six months later. Two voices in a quiet room are straightforward. A focus group of eight, several of whom sound alike and talk across each other, is genuinely harder. Thirty seconds at the top of the recording where everyone states their name is the cheapest quality improvement available to anyone.

Timecoded transcription adds a time reference so any line can be traced back to the recording. We can place timecodes at fixed intervals, at every change of speaker, or only at points you nominate. Researchers coding data usually want them frequent. A solicitor citing a specific admission wants them at the moment that matters.

Audio quality: what makes a recording hard to transcribe

We would rather be honest about this than promise a number we cannot hit. Clear audio with one person speaking at a time transcribes beautifully. Some recordings do not, and no transcriber, human or machine, can recover words that were never captured.

  • Overlapping speech. Two people talking at once produces one intelligible line at best. Group discussions where everyone joins in are the hardest material we handle.
  • Background noise. A cafe, a busy ward, traffic through an open window, or air conditioning sitting right on top of the speech frequencies.
  • Microphone distance. A phone left at the far end of a boardroom table will lose whoever is sitting furthest from it.
  • Accents and dialect. Manageable, and usually solved by assigning a transcriber familiar with the variety, but it does take longer.
  • Specialist terminology. Drug names, case citations, product codes. A short glossary or a list of participant names sent with the audio removes most of this risk.

Where a section genuinely cannot be made out, we mark it inaudible with a timecode rather than guess. A guessed word in a research transcript is worse than a gap, because the gap is visible and the guess is not. Poor audio also costs more, simply because it takes longer, so a decent recorder and a quiet room pay for themselves.

Foreign language transcription and transcription with translation

We transcribe audio in more than 300 languages, and foreign language transcription is a different job from translation. Someone has to listen to the recording in the source language and write down what was said, accurately, before anything can be translated at all. That work goes to a native speaker of the language on the recording.

From there you have three options, and clients often want the third: a transcript in the original language only, an English translation only, or both side by side so a reader can check a quotation against the source. Research teams reporting to an ethics committee or a funder almost always want both, because the original is the data and the translation is the interpretation.

Translated transcripts are produced by our translation services team, so the same standards apply: a native speaker of the target language does the work, and a second linguist reads it before delivery. Interview recordings needing Romanian translation for a housing casework file, or asylum interview audio needing Urdu translation services for a solicitor, are both regular work. If the recording is spoken evidence rather than a document, our interpreting services cover the live side.

Who uses audio transcription services

The work is broader than most people expect, and the requirements barely overlap.

Academic research

Research transcription of interviews, focus groups and oral histories, timecoded and laid out so they import cleanly into the qualitative coding software your team already uses.

Legal and investigations

Legal transcription of recorded interviews, hearings and disclosure audio, taken verbatim where the exact wording has to survive scrutiny. Also see legal translation services.

Market research

Depth interviews, IDIs and focus groups turned round quickly, with respondent labels kept consistent across a whole wave of fieldwork.

Media and podcasts

Interview transcription for editing, show notes and accessibility, handled alongside the rest of our media language services.

Healthcare

Clinical dictation, case notes and recorded consultations, transcribed by people used to medical terminology and handled with the care patient data requires.

Meetings and governance

Board meetings, AGMs, panel hearings and conference sessions, turned into a written record that stands up as minutes.

Confidentiality, consent and UK GDPR

Recordings are often more sensitive than documents, because a voice identifies a person and people say things aloud that they would never write down. We treat them accordingly.

Only the transcriber and project manager assigned to your job can reach the files, and their access ends when the job does. Our handling is UK GDPR compliant. We will sign a non-disclosure agreement where your organisation, client or ethics committee requires one, and we are used to working inside a university's data management plan or a firm's own security schedule. If you need transcripts pseudonymised, with participant names replaced by codes, say so with the brief and we will produce them that way.

How to order audio transcription

  1. 1

    Send the audio and tell us the basics

    Upload the file through the quote form or ask us for a secure link for anything large. Tell us the total runtime, the language, roughly how many speakers there are, and what the transcript is for.

  2. 2

    Agree the style and the deadline

    We recommend verbatim, intelligent verbatim or edited, confirm whether you want timecodes and speaker names, and come back with a fixed price and a date. A sample of the audio helps us quote accurately rather than defensively.

  3. 3

    We transcribe, check and deliver

    A human transcriber does the work and it is checked against the audio before delivery, with anything inaudible flagged rather than guessed. You get the transcript in Word or your preferred format, with translation alongside it if you asked for one.

What drives the cost of transcription

We quote per recording rather than off a rate card, because an hour of audio is not a fixed amount of work. Four things move the number.

  • Runtime. The starting point, and the only figure most people think of.
  • Audio quality. The biggest hidden variable. A clean recording might take three or four times its runtime to transcribe; a difficult one takes considerably longer.
  • Number of speakers. One dictated voice is quick. Eight people in a focus group means constant attribution and a lot of rewinding.
  • Style and turnaround. Full verbatim takes longer than intelligent verbatim. Foreign language transcription takes longer again, and adding translation is a separate stage. Urgent work carries a premium.

Our pricing page sets out how we price language work generally. For transcription the honest answer is that a short sample of your audio tells us more than any rate card, and the quote that follows is fixed.

Video files, or subtitles on screen?

This page covers audio-only recordings. If what you have is footage, our video transcription service handles the file directly and can return the transcript timecoded against the picture. If the end goal is text appearing on screen for viewers rather than a document you read, that is subtitling and closed captions, which is a different craft with its own reading-speed rules. Everything else we do is listed on the services page.

Questions people ask

Audio transcription FAQs

What is the difference between verbatim and intelligent transcription?
Verbatim transcription captures everything, including false starts, repetitions and fillers such as "um" and "you know". It is used where exactly what was said matters, for example legal recordings and some research. Intelligent or clean transcription tidies that out and gives a readable transcript of what the speaker meant. We will recommend one based on what the transcript is for.
Can you add timecodes to the transcript?
Yes. We can add timecodes at set intervals, at each change of speaker, or at points you specify. Timecodes make it easy to find a moment in the recording later, which is useful for research coding, editing and evidence work.
Do you transcribe recordings that are not in English?
Yes. We transcribe audio in more than 300 languages. We can give you a transcript in the original language, an English translation, or both side by side. For interview-based research, having the original and the translation together is often the most useful.
How accurate are your transcripts?
Our transcripts are produced by people, not software alone, and checked before delivery. Accuracy depends a lot on the recording: clear audio with one speaker at a time transcribes very well, while a noisy room with people talking over each other is harder. We will flag any sections we genuinely could not make out rather than guess.
How should I send you the audio?
Most common formats are fine, including MP3, WAV, M4A and recordings from phones and dictation apps. For anything large, we will share a secure upload link. Tell us roughly how long the recording is and how many speakers there are, and we will quote and give you a turnaround.
Can you tell who is speaking on the recording?
Yes. We label speakers throughout, and where you give us names or roles we use those instead of Speaker 1 and Speaker 2. Two voices in a quiet room are straightforward. A focus group of eight people, several of whom sound alike and talk across each other, is harder, and a short introduction at the start of the recording where everyone says their name makes a real difference.
How much does audio transcription cost?
Transcription is quoted per recording rather than from a fixed rate, because the same hour of audio can take very different amounts of time. Four things move the price: the runtime, the quality of the audio, the number of speakers, and how quickly you need it back. Verbatim styles and foreign language transcription take longer than clean English audio with one speaker. Send us the runtime and a sample and we will quote. Our pricing page explains how we price language work generally.
Are my recordings kept confidential?
Yes. Recordings and transcripts are shared only with the transcriber and project manager assigned to your job, our handling is UK GDPR compliant, and we will sign a non-disclosure agreement where your organisation or ethics committee requires one. Much of what we transcribe is research interview data, legal material or clinical dictation, so this is routine rather than a special arrangement.

Got a recording to transcribe?

Tell us the length, the language and how many speakers. We will quote and confirm a turnaround.