> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coloop.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Language quality

> What level of transcription and translation quality to expect from CoLoop in each language, and when to budget for researcher review.

Transcription accuracy varies by language. This page tells you what to expect from each one and how much review to budget for. For the full list of languages CoLoop accepts, see [Supported languages](/docs/setting-up-a-project/supported-languages).

## How the pipeline works

CoLoop processes multilingual research in three steps: transcription, then translation, then analysis.

Analysis runs on the original-language transcript, even after you translate a file to English and even when the output you read is in English. Themes and insights come from the source transcript rather than from a translation of it, so nothing drifts in the translation step. Direct participant quotes stay in the language they were spoken in.

<Info>
  Audio quality affects accuracy in every language. Record in a quiet room with a good microphone wherever you can.
</Info>

## What to expect by language

Ratings describe the transcript CoLoop produces. These are the languages most common in international research programs.

| Language                   | Transcription | Notes                                        |
| -------------------------- | ------------- | -------------------------------------------- |
| English                    | 🟢 High       | UK, US, and Australian variants              |
| French                     | 🟢 High       | Includes Quebecois                           |
| German                     | 🟢 High       |                                              |
| Spanish                    | 🟢 High       | Multiple dialects                            |
| Portuguese                 | 🟢 High       | BR and PT variants                           |
| Italian                    | 🟢 High       |                                              |
| Dutch                      | 🟢 High       |                                              |
| Polish                     | 🟢 High       |                                              |
| Russian                    | 🟢 High       |                                              |
| Turkish                    | 🟢 High       |                                              |
| Swedish, Norwegian, Danish | 🟢 High       |                                              |
| Japanese                   | 🟢 High       |                                              |
| Mandarin Chinese           | 🟢 High       | Simplified                                   |
| Traditional Chinese        | 🟢 High       | Taiwan, HK, Macau                            |
| Korean                     | 🟡 Good       |                                              |
| Arabic                     | 🟡 Good       | Modern Standard Arabic; dialects vary        |
| Hindi                      | 🟡 Good       |                                              |
| Indonesian, Malay          | 🟡 Good       |                                              |
| Czech, Slovak, Romanian    | 🟡 Good       |                                              |
| Thai                       | 🟡 Good       |                                              |
| Vietnamese                 | 🟡 Good       |                                              |
| Hebrew                     | 🟡 Good       |                                              |
| Cantonese                  | 🟡 Good       | Distinct from Mandarin; select it explicitly |
| Urdu                       | 🟡 Good       |                                              |
| Swahili                    | 🟠 Moderate   |                                              |
| Tamil                      | 🟠 Moderate   |                                              |
| Marathi                    | 🟠 Moderate   |                                              |
| Bengali                    | 🔴 Lower      |                                              |
| Gujarati                   | 🔴 Lower      |                                              |
| Burmese, Khmer, Lao        | 🔴 Lower      |                                              |

Every language CoLoop transcribes can also be translated to English.

## How to plan your project

* 🟢 High: run these end to end without special handling.
* 🟡 Good: have a native speaker check a sample of transcripts before you run full analysis.
* 🟠 Moderate and 🔴 Lower: correct transcripts before you analyze them, and consider human transcription where the audio is poor or the interview is unstructured.

Enter your key words and phrases at the transcription stage, in any language. Brand names, product names, and domain vocabulary are then corrected rather than guessed at.

## Medical and clinical research

Medical Mode adds a correction pass over the terminology general-purpose models most often get wrong: medication names, procedures, conditions, and dosages. It covers English, Spanish, German, and French, and you can combine it with your own key phrases for terminology specific to your study.

<Note>
  Medical Mode is off by default and enabled per account. For clinical research in other languages, standard transcription applies, so enter your key phrases.
</Note>

## Reading the accuracy bands

The bands above are based on Word Error Rate (WER), the standard measure for transcription accuracy. WER counts the corrections, insertions, and deletions needed to turn an automated transcript into a perfect one, as a percentage of total words.

A high WER does not mean every other word is wrong. Errors cluster around proper nouns, technical terms, strong accents, and fast speech, while the surrounding context stays intact. WER tells you how much researcher review to budget for.

| Band        | WER        | What it means in practice                                                                                                                                                                                                     |
| ----------- | ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🟢 High     | Up to 10%  | The transcript reads naturally. Occasional errors on proper nouns or specialist terms; you can follow the conversation with minimal review. Key phrases you enter at the transcription stage push the error rate lower still. |
| 🟡 Good     | 10% to 25% | Clearly intelligible but imperfect. Some sentences need correction, mostly around names, accents, or domain vocabulary. Budget for a light review pass.                                                                       |
| 🟠 Moderate | 25% to 50% | Errors are more frequent, concentrated in complex words, fast speech, and strong accents. Meaning is usually recoverable, but review the transcript before you analyze it.                                                    |
| 🔴 Lower    | Above 50%  | Errors are frequent enough that you need to correct the transcript before analysis. Usable, but researcher involvement is significantly higher.                                                                               |

## How files are routed

CoLoop picks the transcription and translation service that handles each language best. Routing is automatic. Regional dialects such as Quebecois French, Brazilian Portuguese, and Mexican and Argentine Spanish are recognized without a separate setting.

For the list of providers that handle your data, see the [CoLoop subprocessor list](https://trust.coloop.ai/subprocessors).

<Warning>
  Cantonese is a distinct spoken language from Mandarin. Select Cantonese, not Chinese, when you set up the file, and review a sample before you analyze it.
</Warning>

## What CoLoop does to limit errors

| Control                  | What it does                                                                                                                                                                                              |
| ------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Confidence scoring       | Words the model was unsure of are highlighted for you to review. The specialist Chinese model and the meeting bot return no per-word confidence, so files transcribed through either show no highlighting |
| Native-language analysis | Transcripts are analyzed in their original language, so meaning does not drift through a translation step                                                                                                 |
| Verbatim quotes          | Direct participant quotes stay in the language they were spoken in                                                                                                                                        |
| Key phrases              | Brand names, product names, and domain terms you supply are corrected rather than guessed at                                                                                                              |
| AI correction            | A correction pass fixes technical terms, repeated hallucinations, and acronyms, and logs every change it makes                                                                                            |
| Transcript editing       | You can review and correct transcripts in any language before analysis begins                                                                                                                             |
