Understanding the Gladia Transcription Process Breakdown
Introduction
When using Gladiaโs transcription API, understanding how your audio is processed helps you interpret response times and optimize performance.
Each transcription request passes through multiple internal stages before the final text output is generated.
Knowing these steps helps you understand where time is spent and how to optimize your requests โ for example, by defining languages in advance or managing concurrency according to your plan.
See also: Optimizing transcription performance by specifying probable languages
The Transcription Pipeline: Step by Step
Each transcription request follows a predictable sequence of stages.
1. Queue (Job Scheduling)
Before processing begins, each transcription request enters a queue.
This initial waiting time depends on:
The number of active transcriptions currently running on Gladiaโs clusters
The overall system load
Your planโs concurrency limits
This queuing mechanism ensures fair resource allocation across users.
During high-demand periods or when account limits are reached, short queue delays may appear โ especially visible on short audio files where queue time can represent a significant share of the total duration.
For more information, see the official documentation on Concurrency and Rate Limits.
2. Pre-processing
Once a job is dispatched, the audio is prepared for transcription.
This step includes:
Audio normalization (format, codec, bitrate, sample rate). The type of audio format can slightly impact processing time, as some formats require longer conversion before transcription begins.
You can find detailed estimates here: Conversion Time by Audio FormatVoice Activity Detection (VAD) to isolate speech from silence
Pre-processing ensures consistent and clean input for accurate transcription.
3. Language Discovery
If no language is specified in your request, the system automatically detects which language(s) are spoken in the audio.
This step ensures accuracy but can increase total processing time.
If you already know the language, you can skip this phase entirely by defining it in your request.
Learn more here: Specifying probable languages
4. Inference
This is the core phase where speech is converted into text.
The inference time scales naturally with audio length and content complexity and is typically the most stable and predictable part of the process.
5. Post-processing and Formatting
After inference, the transcription is refined and formatted.
Depending on your configuration, this may include:
Diarization (speaker separation)
Sentence structuring, Custom Vocabulary / Custom Spelling
Other enabled add-ons
These enhancements improve readability and deliver a clear, well-structured result.
Account-Level Concurrency and Plan Limits
Processing speed and queue time can also vary depending on your account type.
Hereโs an overview of concurrency limits per plan:
Enterprise plan โ Unlimited usage, on-demand concurrency
Paid plan (Self Serve, Scaling) โ Unlimited usage, up to 25 concurrent pre-recorded
Free plan โ โฌ50 credits (one-time, no renewal), 3 concurrent pre-recorded transcriptions, and 1 live
When concurrency limits are reached, additional requests are queued until processing slots become available.
Full details are available in the Gladia Docs โ Concurrency and Rate Limits.
Multi-language and Code-Switching Support
If your audio contains multiple speakers or languages, Gladia can automatically detect and transcribe them.
You can also define several probable languages manually to improve accuracy and consistency.
Learn more in this article: Handling audio with multi-language speakers
Retry policy: when to resubmit (and when not to)
Pre-recorded transcription is asynchronous. The job is accepted when either of these happens:
Your
POST /v2/pre-recordedreturns HTTP 200 (with a job ID)You receive the
transcription.createdwebhook
Both mean Gladia has queued the job. Processing is not finished when that happens.
Once a job is accepted, Gladia continues trying to process it on our side for up to 2 hours after we receive the request. During incidents or high load, queue and processing times can be longer than usual โ a job that is still queued or processing is not a failed submission.
Do not resubmit the same audio if you already got a 200 or a transcription.created webhook. Keep that job ID and wait for completion via polling, success/error webhook, or callback. Re-POSTing does not make the transcript arrive faster: it creates a new job that competes for the same capacity and can worsen the backlog for everyone, including you.
If the HTTP client times out but Gladia still accepted the job, you may still receive transcription.created. Prefer that webhook (or a follow-up GET by ID if you already stored one) before re-POSTing the same audio.
POST returned 200, or you received
transcription.createdโ do not resubmit; poll / wait on that job IDPOST failed (timeout, connection error, 5xx, no response) and no
transcription.createdโ safe to retry the POST with backofftranscription.errorvia webhook or callback, or job statuserrorโ you may submit again, but check the error first and retry only when it makes sense for that failureJob status
queuedorprocessingโ do not resubmitHTTP 429 โ back off; see Concurrency and Rate Limits
Errors and when to retry. A transcription.error event (webhook or callback) means processing failed for that job ID. At that point a new POST is allowed โ but not every error is worth retrying:
Transient / infrastructure-style failures (timeouts, temporary 5xx, short-lived capacity issues): retry with backoff can help
Request or media issues (invalid / unreachable audio URL, unsupported or corrupt file, bad parameters): fix the input first โ resubmitting the same payload will fail again and still create a new billable job if accepted
Inspect the error details (webhook or callback payload, and/or GET on the job ID) before deciding.
No deduplication today. Each successful POST creates a new job and is billed separately, even if the audio file and your custom_metadata are identical. Until a server-side deduplication option exists, โone accepted job per piece of contentโ is a client-side responsibility.
Conclusion
Every transcription request follows the same pipeline, from queue management to post-processing.
A short queue at the start is normal and depends on both system demand and your planโs concurrency limits.
By specifying probable languages, understanding concurrency behavior, and optimizing your configuration, you can achieve faster, more predictable, and consistent transcription performance.
If a job was accepted (HTTP 200 or transcription.created), prefer waiting on that job ID rather than resubmitting โ aggressive retries create duplicate jobs, extra billing, and longer queues without speeding up your result.