Skip to content
Gladia Help Center home
InboxAsk a human

Understanding the Gladia Transcription Process Breakdown

Introduction

When using Gladiaโ€™s transcription API, understanding how your audio is processed helps you interpret response times and optimize performance.
Each transcription request passes through multiple internal stages before the final text output is generated.

Knowing these steps helps you understand where time is spent and how to optimize your requests โ€” for example, by defining languages in advance or managing concurrency according to your plan.

See also: Optimizing transcription performance by specifying probable languages

The Transcription Pipeline: Step by Step

Each transcription request follows a predictable sequence of stages.

1. Queue (Job Scheduling)

Before processing begins, each transcription request enters a queue.
This initial waiting time depends on:

  • The number of active transcriptions currently running on Gladiaโ€™s clusters

  • The overall system load

  • Your planโ€™s concurrency limits

This queuing mechanism ensures fair resource allocation across users.
During high-demand periods or when account limits are reached, short queue delays may appear โ€” especially visible on short audio files where queue time can represent a significant share of the total duration.

For more information, see the official documentation on Concurrency and Rate Limits.

2. Pre-processing

Once a job is dispatched, the audio is prepared for transcription.
This step includes:

  • Audio normalization (format, codec, bitrate, sample rate). The type of audio format can slightly impact processing time, as some formats require longer conversion before transcription begins.
    You can find detailed estimates here: Conversion Time by Audio Format

  • Voice Activity Detection (VAD) to isolate speech from silence

Pre-processing ensures consistent and clean input for accurate transcription.

3. Language Discovery

If no language is specified in your request, the system automatically detects which language(s) are spoken in the audio.
This step ensures accuracy but can increase total processing time.

If you already know the language, you can skip this phase entirely by defining it in your request.
Learn more here: Specifying probable languages

4. Inference

This is the core phase where speech is converted into text.
The inference time scales naturally with audio length and content complexity and is typically the most stable and predictable part of the process.

5. Post-processing and Formatting

After inference, the transcription is refined and formatted.
Depending on your configuration, this may include:

  • Diarization (speaker separation)

  • Sentence structuring, Custom Vocabulary / Custom Spelling

  • Other enabled add-ons

These enhancements improve readability and deliver a clear, well-structured result.

Account-Level Concurrency and Plan Limits

Processing speed and queue time can also vary depending on your account type.
Hereโ€™s an overview of concurrency limits per plan:

  • Enterprise plan โ†’ Unlimited usage, on-demand concurrency

  • Paid plan (Self Serve, Scaling) โ†’ Unlimited usage, up to 25 concurrent pre-recorded

  • Free plan โ†’ โ‚ฌ50 credits (one-time, no renewal), 3 concurrent pre-recorded transcriptions, and 1 live

When concurrency limits are reached, additional requests are queued until processing slots become available.
Full details are available in the Gladia Docs โ€“ Concurrency and Rate Limits.

Multi-language and Code-Switching Support

If your audio contains multiple speakers or languages, Gladia can automatically detect and transcribe them.
You can also define several probable languages manually to improve accuracy and consistency.

Learn more in this article: Handling audio with multi-language speakers

Retry policy: when to resubmit (and when not to)

Pre-recorded transcription is asynchronous. The job is accepted when either of these happens:

Both mean Gladia has queued the job. Processing is not finished when that happens.

Once a job is accepted, Gladia continues trying to process it on our side for up to 2 hours after we receive the request. During incidents or high load, queue and processing times can be longer than usual โ€” a job that is still queued or processing is not a failed submission.

Do not resubmit the same audio if you already got a 200 or a transcription.created webhook. Keep that job ID and wait for completion via polling, success/error webhook, or callback. Re-POSTing does not make the transcript arrive faster: it creates a new job that competes for the same capacity and can worsen the backlog for everyone, including you.

If the HTTP client times out but Gladia still accepted the job, you may still receive transcription.created. Prefer that webhook (or a follow-up GET by ID if you already stored one) before re-POSTing the same audio.

  • POST returned 200, or you received transcription.created โ†’ do not resubmit; poll / wait on that job ID

  • POST failed (timeout, connection error, 5xx, no response) and no transcription.created โ†’ safe to retry the POST with backoff

  • transcription.error via webhook or callback, or job status error โ†’ you may submit again, but check the error first and retry only when it makes sense for that failure

  • Job status queued or processing โ†’ do not resubmit

  • HTTP 429 โ†’ back off; see Concurrency and Rate Limits

Errors and when to retry. A transcription.error event (webhook or callback) means processing failed for that job ID. At that point a new POST is allowed โ€” but not every error is worth retrying:

  • Transient / infrastructure-style failures (timeouts, temporary 5xx, short-lived capacity issues): retry with backoff can help

  • Request or media issues (invalid / unreachable audio URL, unsupported or corrupt file, bad parameters): fix the input first โ€” resubmitting the same payload will fail again and still create a new billable job if accepted

Inspect the error details (webhook or callback payload, and/or GET on the job ID) before deciding.

No deduplication today. Each successful POST creates a new job and is billed separately, even if the audio file and your custom_metadata are identical. Until a server-side deduplication option exists, โ€œone accepted job per piece of contentโ€ is a client-side responsibility.

Conclusion

Every transcription request follows the same pipeline, from queue management to post-processing.
A short queue at the start is normal and depends on both system demand and your planโ€™s concurrency limits.

By specifying probable languages, understanding concurrency behavior, and optimizing your configuration, you can achieve faster, more predictable, and consistent transcription performance.

If a job was accepted (HTTP 200 or transcription.created), prefer waiting on that job ID rather than resubmitting โ€” aggressive retries create duplicate jobs, extra billing, and longer queues without speeding up your result.