Speech to Text
Munsit converts spoken Arabic into accurate, structured text — high-accuracy recognition across dialects and accents, with word-level timestamps, speaker diarization and meeting intelligence built on top.
Your first transcription
One POST with a file. Munsit analyzes the recording, converts the Arabic speech to text, and returns the transcript with its duration and word-level timestamps.
munsit-en-ar model below.
Choose a model
Two models. Every batch endpoint accepts an optional model parameter; if omitted, munsit is used. Live streaming takes the same model values; its language parameter is ar-only in v1. Custom vocabulary (hotwords) is ignored on munsit-en-ar.
| Model | ID | Use it for | Default |
|---|---|---|---|
| Munsit | munsit | Arabic. Optimized for Arabic speech recognition, with strong performance across dialects and accents. | Yes |
| Munsit En-Ar | munsit-en-ar | Mixed Arabic–English spoken content with code-switching support — for speakers who naturally alternate between the two languages within the same conversation or utterance. | — |
What you can send
Twelve audio formats, covering common output from different recording platforms — no transcoding step.
| Workflow | Duration limit | Notes |
|---|---|---|
| Audio transcription | Under 60 minutes | Pre-recorded files, word-level timing. |
| Minutes of meetings | Under 30 minutes | Structured transcripts optimized for meeting use cases. |
| Live streaming | No limit | Unbounded while the connection stays active. Keep the session alive with KeepAlive during pauses; unbroken speech is force-segmented about every 60 seconds. |
Go further
Three core workflows plus real-time streaming.
Working with an AI assistant? Every page is available as Markdown: add .md to the URL, or send an Accept: text/markdown header. For the whole documentation in one request, point it at llms-full.txt; the page index is llms.txt. Or use Copy Page, top right.
