Changelog

Changelog

New endpoints, improvements, and API changes to Munsit — newest first.

Skip word timestamps on live streaming

  • NewWS /api/v1/listen accepts an optional timestamps query parameter. With timestamps=false, every Results event carries the transcript only: words is an empty array and confidence is null. The default is true, and connections that do not send it get the same events as before.
  • APIWith model=munsit-en-ar, timestamps=false also skips the word-timing step, so final transcripts arrive slightly sooner. UtteranceEnd is still sent; with munsit-en-ar its last_word_end is then approximate.

Speaker labels on transcription

  • NewPOST /api/v1/audio/transcribe accepts an optional diarize field. With diarize=true, the response adds utterances (who said what), speakers (talk time per speaker) and diarization (status, speaker_cap_reached), and every timestamps entry gets a speaker number. It works with both munsit and munsit-en-ar.
  • APISpeaker numbers start at 0 in the order speakers first talk and tell speakers apart within one recording. Up to 8 speakers are told apart; speaker_cap_reached is true when 8 were found. A new utterance starts when the speaker changes or after a pause of 1 second or more.
  • If speaker detection cannot be completed, the transcript is still returned with diarization.status set to unavailable, and utterances, speakers and the per-word speaker are left out.
  • Requests without diarize, or with diarize=false, get exactly the response they got before. A value other than true or false is ignored. diarize is not available on live streaming or minutes of meetings.

Munsit MCP server

  • NewThe Munsit MCP server lets AI apps use your Munsit account: find voices, generate speech, transcribe audio and check credits and usage, by asking in plain language. Add https://mcp.munsit.com/mcp to Claude, Claude Code, ChatGPT or any app that supports MCP with OAuth, and sign in with your Munsit account — no API key needed.
  • Thirteen tools; only speech and transcription spend credits, at the same prices as the app. See MCP server for the permissions, limits and regional URLs.

Language mode for munsit-en-ar

  • NewPOST /api/v1/audio/transcribe accepts an optional language field with model=munsit-en-ar. auto, the default, detects the language as the audio goes; ar or en sets the main language of the speech. Every value still transcribes mixed Arabic-English speech; setting the main language gives better accuracy when you know it. Requests without the field behave as before.
  • WS /api/v1/listen takes the same language query parameter with model=munsit-en-ar. Each Results event reports the language of that result: the detected one with auto, otherwise the one you set.
  • ImprovedCustom vocabulary (hotwords) now works with munsit-en-ar on POST /api/v1/audio/transcribe. Up to 20 words are used, counted across all entries.
  • APIOn POST /api/v1/audio/transcribe, a language value other than auto, ar, en or ar-en is ignored and auto is used. On WS /api/v1/listen, an unsupported value closes the connection with error 4002. The munsit model is Arabic only: it ignores the field on batch requests and accepts only ar on live sessions.
  • Live munsit-en-ar sessions do not apply hotwords. Every entry you send is reported in Metadata.dropped_hotwords.

Word timestamps for Text to Speech

  • NewNew POST /api/v1/text-to-speech/{model_id}/with-timestamps returns speech together with character-level timings for the text you submitted. Use it to highlight words as they are spoken, drive captions, or align a transcript to the audio.
  • The response is NDJSON — one JSON object per line. Audio lines carry base64 PCM16 in audio_base64; the final lines carry alignment (aligned to your original text) and normalized_alignment (aligned to the engine's normalized form). Concatenating characters reproduces your input exactly, so array indices map straight back onto your string.
  • Break tags are supported in the text you send: <break time="3s"/> and <break time="500ms"/> insert a pause of the given length.
  • APITimings are emitted once generation finishes, not incrementally with each audio chunk. To highlight from the first word, buffer the full response before playback or request long text sentence by sentence.
  • Requires a model served by the v1.5 engine. Other models return 400 with Word timestamps are not available for model '…'. Pricing, wallet deduction, and history are identical to a standard synthesis request — timestamps cost nothing extra.

Streaming finalization on the legacy STT socket

  • NewSend {"event":"end_of_stream"} to transcribe any remaining buffered audio, at any duration. The server replies with a transcription carrying isFinal: true, then a finalized event with the complete transcript. Closing the socket does not flush buffered audio — send this first.
  • transcription events now carry an isFinal boolean. Clients reading only data are unaffected.
  • New min_buffer_seconds query parameter controls how much audio accumulates before an interim result is emitted. Default 0.5, range 0.1–5.0.
  • FixedUtterances shorter than roughly 1.7 seconds could complete with no transcript and no error. The audio buffer was measured before each incoming chunk was appended, and the first chunk was never evaluated, so short answers never reached the recognizer.
  • Rapid audio frames arriving during connection setup could each start their own initialization, discarding the WAV metadata from the first frame. Every subsequent frame then failed with Not a valid WAV file (missing RIFF header) for the life of the connection. Clients sending 10–20 ms frames were affected on every session.
  • DeprecatedWS /websocket/speech-to-text remains deprecated and receives no new recognition features. New integrations should use WS /api/v1/listen.
v1

New capabilities

  • NewNew Text-to-Speech capabilities.
  • New Speech-to-Text capabilities.

Working with an AI assistant? Every page is available as Markdown: add .md to the URL, or send an Accept: text/markdown header. For the whole documentation in one request, point it at llms-full.txt; the page index is llms.txt. Or use Copy Page, top right.