Transcribe
Send pre-recorded Arabic audio — either a file as multipart/form-data or a public audio_url we fetch for you — and receive a high-quality transcript with total duration and word-level timestamps. Optimized for asynchronous processing of interviews, meetings, media clips and customer calls.
Endpoint
Authenticate with your API key in the x-api-key header. See Authentication.
| Header | Value |
|---|---|
x-api-key | YOUR_MUNSIT_API_KEY |
Request
Send the body as multipart/form-data. Supply the audio in exactly one of two ways: file to upload the bytes yourself, or audio_url to have Munsit fetch it. Sending both, or neither, is rejected — see choosing a source.
| Field | Type | Required | Description |
|---|---|---|---|
file | file | One of | Audio file in a supported format — see what you can send. Mutually exclusive with audio_url. |
audio_url | string | One of | Public https URL of the audio to transcribe. Presigned links from S3, GCS, R2 and similar are supported. Mutually exclusive with file. |
model | string | No | ASR model to use: munsit (default) or munsit-en-ar (mixed Arabic-English with code-switching). |
language | string | No | Only used with munsit-en-ar. Sets the main language of the audio: auto (default) detects it as the audio goes, ar for mostly Arabic speech, en for mostly English speech. Every value still transcribes speech that mixes both languages. ar-en is accepted as another name for auto. Any other value is ignored and auto is used. munsit is Arabic only and ignores the field. See Language. |
hotwords | string | No | Comma-separated custom vocabulary (multi-word phrases allowed). Biases recognition toward names, brands and domain terms. With munsit-en-ar, at most 20 words are used. |
return_confidence | boolean | No | When true, each timestamps entry includes a confidence score (0–1). The response is otherwise unchanged. |
return_timestamps | boolean | No | Defaults to true on munsit; set false to return an empty timestamps array. On munsit-en-ar timestamps are off unless you set this to true. |
return_turns | boolean | No | When true, adds the turns array and tags each turn with smart-turn is_complete and turn_probability. These two fields appear only with this flag. |
return_gender | boolean | No | When true, adds the turns array with a gender object (label, score) per turn, plus a whole-file rollup in analysis. |
return_sentiment | boolean | No | When true, adds the turns array with a sentiment object (label, score) per turn, plus a whole-file rollup in analysis. |
diarize | boolean | No | When true, labels every word with its speaker and adds utterances, speakers and diarization to the response. Works with both models. Use true or false; an unrecognized value is ignored. See Speaker labels. |
Choosing a source
Both sources produce the same transcript, are billed the same way on decoded audio duration, and store the audio identically. They differ only in who moves the bytes.
| Use | When | Why |
|---|---|---|
file | The audio is on the machine making the call | One request, nothing to host. The upload happens over your connection, so a large file on a slow uplink takes as long as that link allows. |
audio_url | The audio already lives in object storage or on a CDN | You never upload it twice. Munsit fetches it server-side, typically far faster than a client upload, and your request body stays a few hundred bytes. |
Rules for audio_url: the scheme must be https, the host must be publicly resolvable, and the URL must serve the audio directly. Redirects are not followed, and private, loopback and link-local addresses are refused. The fetched file is subject to the same size and duration limits as an upload.
Custom vocabulary
Pass rare terms the recognizer is unlikely to know — customer names, product codes, brand words. In our benchmarks, biasing recovered rare terms from 0% to 77% recall with overall accuracy unchanged or slightly better. Short lists of 5–30 genuinely rare terms work best; very long lists dilute the effect.
model=munsit-en-ar, hotwords takes up to 20 words, counted across all entries; any further words are ignored. Separate entries with commas or new lines. If using them would leave the transcript empty, the request is retried automatically without them.hotwords, language and return_confidence apply to POST /audio/transcribe — not to diarization or minutes of meetings. For live audio, streaming takes hotwords and language as query parameters and always returns confidence.Language
munsit-en-ar transcribes Arabic, English, and speech that switches between them, whatever language you set. language tells the model which language most of the speech is in. When you know it, set it: a mostly Arabic call with some English terms is transcribed more accurately with ar than with auto. Use auto, or leave the field out, when the mix is even or you do not know.
| Value | Behavior |
|---|---|
auto | Detects the language as the audio goes. Best for an even mix, or when you do not know the main language. This is the default; ar-en is another name for it. |
ar | The speech is mostly Arabic. English words and phrases in it are still transcribed. |
en | The speech is mostly English. Arabic words and phrases in it are still transcribed. Do not use it for mostly Arabic audio: the transcript can come back partly or wholly in English and lose most of what was said. |
auto, ar, en or ar-en is ignored and auto is used, so a spelling such as EN or en-US does not select English. Values are case-sensitive and written in lowercase. The munsit model is Arabic only: it accepts the field and ignores it.Per-turn analysis
Set return_turns, return_gender and/or return_sentiment to break the transcript into turns and annotate each one. Any of the three adds the turns array — the gender and sentiment flags imply it — and each flag contributes only its own fields. analysis holds the whole-file rollup and appears only with return_gender or return_sentiment; return_turns on its own does not produce it. Short single-speaker recordings typically come back as one turn.
Sentiment and Gender events per turn.GET /history/speech-to-text/:id also returns turns and analysis for transcriptions created with one of these flags set.Speaker labels
Set diarize=true to find out who said what. Every word gets a speaker number, and the transcript is grouped into utterances, each spoken by one speaker. It works with both models, and the transcript text itself does not change. The per-word speaker number sits inside timestamps, which return_timestamps=false empties; utterances and speakers are returned either way.
| Topic | Behavior |
|---|---|
| Speaker numbers | Whole numbers starting at 0, in the order speakers first talk. They tell speakers apart within one recording; they do not say who a person is, and the same person can get a different number in another request. |
| Utterances | A new utterance starts when the speaker changes or after a pause of 1 second or more. start and end are the first word's start and the last word's end. |
| Overlapping speech | Each word goes to one speaker, the one who covers most of it, so two people talking at once still produce a single stream of words. |
| Speaker limit | Up to 8 speakers are told apart. speaker_cap_reached is true when 8 were found; in a recording with more, some speakers share a number. |
| No speech | A silent recording returns empty utterances and speakers. |
| When it cannot run | If speaker detection cannot be completed you still get the full transcript, with diarization.status set to "unavailable". utterances, speakers and the per-word speaker are then left out. |
diarize applies to POST /audio/transcribe — not to live streaming or minutes of meetings. For speaker segments with their own endpoint and sentiment per speaker, see Diarization.Example request
How it works: send the audio, Munsit analyzes the recording and converts the Arabic speech into text, and you get back a transcript with duration and word-level timestamps. The first tab uploads a file; swap the -F "file=@…" line for -F "audio_url=…" to have Munsit fetch it instead.
Response fields
The payload arrives under data, alongside statusCode and message.
| Field | Type | Description |
|---|---|---|
transcriptionId | string (UUID) | Transcription identifier. Pass it as the path parameter to sentiment analysis and keyword extraction. |
transcription | string | Full transcript text. |
duration | number | Audio duration in seconds. |
timestamps | array of objects (word, start, end; plus confidence with return_confidence, speaker with diarize) | Word-level timestamps. |
turns | array of objects (turn_id, start, end, text; plus is_complete / turn_probability with return_turns, gender with return_gender, sentiment with return_sentiment) | Present when any of the three flags is true — return_gender and return_sentiment imply it. Short single-speaker clips typically return one turn. |
analysis | object (turns count; plus gender and sentiment when those flags are set) | Whole-file rollup. Present only with return_gender or return_sentiment — return_turns alone does not produce it. |
utterances | array of objects (speaker, start, end, text) | Who said what. Present only with diarize, when diarization.status is "ok". See Speaker labels. |
speakers | array of objects (speaker, speech_seconds) | Talk time per speaker, in seconds. Present under the same conditions as utterances. |
diarization | object (status, speaker_cap_reached) | Present only with diarize. status is "ok" or "unavailable". |
attributes | object | Internal metadata blob persisted with the transcription; mirrors fields above under different names (and repeats the full timestamps array). Prefer the named fields — treat this as unstable. |
summary | string | Always present. Empty string on plain transcription; carries the generated summary on minutes of meetings. |
stats | object (fileName, fileSize, mimeType, creditsConsumed) | Upload metadata and the credits billed for this request. |
Audio retention
Audio submitted to the hosted API is stored server-side, whether you uploaded it or Munsit fetched it from audio_url. The transcript text is persisted alongside it so it can be retrieved later through the history endpoints. The transcription response does not include a link to the stored audio.
| What | Detail |
|---|---|
| What is stored | The audio and the resulting transcript, keyed to your account. Passing audio_url instead of a file does not avoid this — the fetched audio is stored the same way. |
| Access | Stored audio is private to your account. The transcription response does not return a link to it. |
| Removal | Delete a transcription and its stored audio through the history endpoints, or contact support for bulk removal. |
| Avoiding retention | If your deployment cannot retain audio at all, use self-hosting, where storage is under your control. |
Go further
Do more with your transcript.
Working with an AI assistant? Every page is available as Markdown: add .md to the URL, or send an Accept: text/markdown header. For the whole documentation in one request, point it at llms-full.txt; the page index is llms.txt. Or use Copy Page, top right.