Speech to Text › Transcribe

Transcribe

Send pre-recorded Arabic audio — either a file as multipart/form-data or a public audio_url we fetch for you — and receive a high-quality transcript with total duration and word-level timestamps. Optimized for asynchronous processing of interviews, meetings, media clips and customer calls.

1

Endpoint

POST /api/v1/audio/transcribe

Authenticate with your API key in the x-api-key header. See Authentication.

HeaderValue
x-api-keyYOUR_MUNSIT_API_KEY
2

Request

Send the body as multipart/form-data. Supply the audio in exactly one of two ways: file to upload the bytes yourself, or audio_url to have Munsit fetch it. Sending both, or neither, is rejected — see choosing a source.

FieldTypeRequiredDescription
filefileOne ofAudio file in a supported format — see what you can send. Mutually exclusive with audio_url.
audio_urlstringOne ofPublic https URL of the audio to transcribe. Presigned links from S3, GCS, R2 and similar are supported. Mutually exclusive with file.
modelstringNoASR model to use: munsit (default) or munsit-en-ar (mixed Arabic-English with code-switching).
languagestringNoOnly used with munsit-en-ar. Sets the main language of the audio: auto (default) detects it as the audio goes, ar for mostly Arabic speech, en for mostly English speech. Every value still transcribes speech that mixes both languages. ar-en is accepted as another name for auto. Any other value is ignored and auto is used. munsit is Arabic only and ignores the field. See Language.
hotwordsstringNoComma-separated custom vocabulary (multi-word phrases allowed). Biases recognition toward names, brands and domain terms. With munsit-en-ar, at most 20 words are used.
return_confidencebooleanNoWhen true, each timestamps entry includes a confidence score (0–1). The response is otherwise unchanged.
return_timestampsbooleanNoDefaults to true on munsit; set false to return an empty timestamps array. On munsit-en-ar timestamps are off unless you set this to true.
return_turnsbooleanNoWhen true, adds the turns array and tags each turn with smart-turn is_complete and turn_probability. These two fields appear only with this flag.
return_genderbooleanNoWhen true, adds the turns array with a gender object (label, score) per turn, plus a whole-file rollup in analysis.
return_sentimentbooleanNoWhen true, adds the turns array with a sentiment object (label, score) per turn, plus a whole-file rollup in analysis.
diarizebooleanNoWhen true, labels every word with its speaker and adds utterances, speakers and diarization to the response. Works with both models. Use true or false; an unrecognized value is ignored. See Speaker labels.
3

Choosing a source

Both sources produce the same transcript, are billed the same way on decoded audio duration, and store the audio identically. They differ only in who moves the bytes.

UseWhenWhy
fileThe audio is on the machine making the callOne request, nothing to host. The upload happens over your connection, so a large file on a slow uplink takes as long as that link allows.
audio_urlThe audio already lives in object storage or on a CDNYou never upload it twice. Munsit fetches it server-side, typically far faster than a client upload, and your request body stays a few hundred bytes.

Rules for audio_url: the scheme must be https, the host must be publicly resolvable, and the URL must serve the audio directly. Redirects are not followed, and private, loopback and link-local addresses are refused. The fetched file is subject to the same size and duration limits as an upload.

4

Custom vocabulary

Pass rare terms the recognizer is unlikely to know — customer names, product codes, brand words. In our benchmarks, biasing recovered rare terms from 0% to 77% recall with overall accuracy unchanged or slightly better. Short lists of 5–30 genuinely rare terms work best; very long lists dilute the effect.

bash
curl -X POST "https://api.munsit.com/api/v1/audio/transcribe" \ -H "x-api-key: $MUNSIT_API_KEY" \ -F "file=@call.wav" \ -F "hotwords=عبد القادر,أديب" \ -F "return_confidence=true"
Response — timestamps entry
{ "word": "الأشياء", "start": 0.24, "end": 0.31, "confidence": 0.994 }
With model=munsit-en-ar, hotwords takes up to 20 words, counted across all entries; any further words are ignored. Separate entries with commas or new lines. If using them would leave the transcript empty, the request is retried automatically without them.
Transcribe only. hotwords, language and return_confidence apply to POST /audio/transcribe — not to diarization or minutes of meetings. For live audio, streaming takes hotwords and language as query parameters and always returns confidence.
5

Language

munsit-en-ar transcribes Arabic, English, and speech that switches between them, whatever language you set. language tells the model which language most of the speech is in. When you know it, set it: a mostly Arabic call with some English terms is transcribed more accurately with ar than with auto. Use auto, or leave the field out, when the mix is even or you do not know.

ValueBehavior
autoDetects the language as the audio goes. Best for an even mix, or when you do not know the main language. This is the default; ar-en is another name for it.
arThe speech is mostly Arabic. English words and phrases in it are still transcribed.
enThe speech is mostly English. Arabic words and phrases in it are still transcribed. Do not use it for mostly Arabic audio: the transcript can come back partly or wholly in English and lose most of what was said.
bash
curl -X POST "https://api.munsit.com/api/v1/audio/transcribe" \ -H "x-api-key: $MUNSIT_API_KEY" \ -F "file=@call.wav" \ -F "model=munsit-en-ar" \ -F "language=en"
A value other than auto, ar, en or ar-en is ignored and auto is used, so a spelling such as EN or en-US does not select English. Values are case-sensitive and written in lowercase. The munsit model is Arabic only: it accepts the field and ignores it.
6

Per-turn analysis

Set return_turns, return_gender and/or return_sentiment to break the transcript into turns and annotate each one. Any of the three adds the turns array — the gender and sentiment flags imply it — and each flag contributes only its own fields. analysis holds the whole-file rollup and appears only with return_gender or return_sentiment; return_turns on its own does not produce it. Short single-speaker recordings typically come back as one turn.

bash
curl -X POST "https://api.munsit.com/api/v1/audio/transcribe" \ -H "x-api-key: $MUNSIT_API_KEY" \ -F "file=@call.wav" \ -F "return_turns=true" \ -F "return_gender=true" \ -F "return_sentiment=true"
Response — 200
{ "statusCode": 200, "data": { "transcription": "أهلا كيف حالك... بخير شكرا", "duration": 6.2, "turns": [ { "turn_id": 0, "start": 0.0, "end": 2.8, "text": "أهلا كيف حالك", "is_complete": true, "turn_probability": 0.94, "gender": { "label": "male", "score": 0.88 }, "sentiment": { "label": "neutral", "score": 0.81 } }, { "turn_id": 1, "start": 3.1, "end": 6.2, "text": "بخير شكرا", "is_complete": true, "turn_probability": 0.97, "gender": { "label": "female", "score": 0.91 }, "sentiment": { "label": "positive", "score": 0.86 } } ], "analysis": { "turns": 2, "gender": { "dominant": "female", "by_duration_s": { "male": 2.8, "female": 3.1 } }, "sentiment": { "dominant": "positive", "counts": { "neutral": 1, "positive": 1 } } } }, "message": "Success" }
This sentiment is a quick per-utterance signal returned alongside the transcript. For deeper, LLM-based analysis of an existing transcription — emotions, trends, critical moments — use Sentiment analysis instead. For live audio, streaming emits Sentiment and Gender events per turn.
GET /history/speech-to-text/:id also returns turns and analysis for transcriptions created with one of these flags set.
7

Speaker labels

Set diarize=true to find out who said what. Every word gets a speaker number, and the transcript is grouped into utterances, each spoken by one speaker. It works with both models, and the transcript text itself does not change. The per-word speaker number sits inside timestamps, which return_timestamps=false empties; utterances and speakers are returned either way.

bash
curl -X POST "https://api.munsit.com/api/v1/audio/transcribe" \ -H "x-api-key: $MUNSIT_API_KEY" \ -F "file=@call.wav" \ -F "diarize=true"
Response — 200
{ "statusCode": 200, "data": { "transcription": "أهلا كيف حالك بخير شكرا", "duration": 6.2, "timestamps": [ { "word": "أهلا", "start": 0.12, "end": 0.48, "speaker": 0 } ], "utterances": [ { "speaker": 0, "start": 0.12, "end": 2.8, "text": "أهلا كيف حالك" }, { "speaker": 1, "start": 3.1, "end": 6.2, "text": "بخير شكرا" } ], "speakers": [ { "speaker": 0, "speech_seconds": 2.68 }, { "speaker": 1, "speech_seconds": 3.1 } ], "diarization": { "status": "ok", "speaker_cap_reached": false } }, "message": "Success" }
TopicBehavior
Speaker numbersWhole numbers starting at 0, in the order speakers first talk. They tell speakers apart within one recording; they do not say who a person is, and the same person can get a different number in another request.
UtterancesA new utterance starts when the speaker changes or after a pause of 1 second or more. start and end are the first word's start and the last word's end.
Overlapping speechEach word goes to one speaker, the one who covers most of it, so two people talking at once still produce a single stream of words.
Speaker limitUp to 8 speakers are told apart. speaker_cap_reached is true when 8 were found; in a recording with more, some speakers share a number.
No speechA silent recording returns empty utterances and speakers.
When it cannot runIf speaker detection cannot be completed you still get the full transcript, with diarization.status set to "unavailable". utterances, speakers and the per-word speaker are then left out.
File transcription only. diarize applies to POST /audio/transcribe — not to live streaming or minutes of meetings. For speaker segments with their own endpoint and sentiment per speaker, see Diarization.
8

Example request

How it works: send the audio, Munsit analyzes the recording and converts the Arabic speech into text, and you get back a transcript with duration and word-level timestamps. The first tab uploads a file; swap the -F "file=@…" line for -F "audio_url=…" to have Munsit fetch it instead.

bash
curl -X POST "https://api.munsit.com/api/v1/audio/transcribe" \ -H "x-api-key: $MUNSIT_API_KEY" \ -F "file=@meeting.mp3" \ -F "model=munsit"# or let Munsit fetch it, instead of -F "file=@…"curl -X POST "https://api.munsit.com/api/v1/audio/transcribe" \ -H "x-api-key: $MUNSIT_API_KEY" \ -F "audio_url=https://storage.example.com/meeting.mp3" \ -F "model=munsit"
Response — 200
{ "statusCode": 200, "data": { "transcriptionId": "805059bf-7c3f-4a1e-9d2b-1f0c6ae83b47", "transcription": "لك كلما عمقت الآخرين أصبحت قزما...", "duration": 53.661375, "timestamps": [ { "word": "الأشياء", "start": 0.24, "end": 0.31 } ], "summary": "", "stats": { "fileName": "meeting.mp3", "fileSize": "1.42 MB", "mimeType": "audio/mpeg", "creditsConsumed": 7 } }, "message": "Success" }
Files must be under 60 minutes. For longer recordings, split the audio into shorter segments for best performance — or use Streaming, which has no duration limit.
9

Response fields

The payload arrives under data, alongside statusCode and message.

FieldTypeDescription
transcriptionIdstring (UUID)Transcription identifier. Pass it as the path parameter to sentiment analysis and keyword extraction.
transcriptionstringFull transcript text.
durationnumberAudio duration in seconds.
timestampsarray of objects (word, start, end; plus confidence with return_confidence, speaker with diarize)Word-level timestamps.
turnsarray of objects (turn_id, start, end, text; plus is_complete / turn_probability with return_turns, gender with return_gender, sentiment with return_sentiment)Present when any of the three flags is true — return_gender and return_sentiment imply it. Short single-speaker clips typically return one turn.
analysisobject (turns count; plus gender and sentiment when those flags are set)Whole-file rollup. Present only with return_gender or return_sentiment — return_turns alone does not produce it.
utterancesarray of objects (speaker, start, end, text)Who said what. Present only with diarize, when diarization.status is "ok". See Speaker labels.
speakersarray of objects (speaker, speech_seconds)Talk time per speaker, in seconds. Present under the same conditions as utterances.
diarizationobject (status, speaker_cap_reached)Present only with diarize. status is "ok" or "unavailable".
attributesobjectInternal metadata blob persisted with the transcription; mirrors fields above under different names (and repeats the full timestamps array). Prefer the named fields — treat this as unstable.
summarystringAlways present. Empty string on plain transcription; carries the generated summary on minutes of meetings.
statsobject (fileName, fileSize, mimeType, creditsConsumed)Upload metadata and the credits billed for this request.
10

Audio retention

Audio submitted to the hosted API is stored server-side, whether you uploaded it or Munsit fetched it from audio_url. The transcript text is persisted alongside it so it can be retrieved later through the history endpoints. The transcription response does not include a link to the stored audio.

WhatDetail
What is storedThe audio and the resulting transcript, keyed to your account. Passing audio_url instead of a file does not avoid this — the fetched audio is stored the same way.
AccessStored audio is private to your account. The transcription response does not return a link to it.
RemovalDelete a transcription and its stored audio through the history endpoints, or contact support for bulk removal.
Avoiding retentionIf your deployment cannot retain audio at all, use self-hosting, where storage is under your control.
11

Go further

Do more with your transcript.

Working with an AI assistant? Every page is available as Markdown: add .md to the URL, or send an Accept: text/markdown header. For the whole documentation in one request, point it at llms-full.txt; the page index is llms.txt. Or use Copy Page, top right.