Changelog
Changelog
New endpoints, improvements, and API changes to Munsit — newest first.
Skip word timestamps on live streaming
- New
WS /api/v1/listenaccepts an optionaltimestampsquery parameter. Withtimestamps=false, everyResultsevent carries the transcript only:wordsis an empty array andconfidenceisnull. The default istrue, and connections that do not send it get the same events as before. - APIWith
model=munsit-en-ar,timestamps=falsealso skips the word-timing step, so final transcripts arrive slightly sooner.UtteranceEndis still sent; withmunsit-en-aritslast_word_endis then approximate.
Speaker labels on transcription
- New
POST /api/v1/audio/transcribeaccepts an optionaldiarizefield. Withdiarize=true, the response addsutterances(who said what),speakers(talk time per speaker) anddiarization(status,speaker_cap_reached), and everytimestampsentry gets aspeakernumber. It works with bothmunsitandmunsit-en-ar. - APISpeaker numbers start at
0in the order speakers first talk and tell speakers apart within one recording. Up to 8 speakers are told apart;speaker_cap_reachedistruewhen 8 were found. A new utterance starts when the speaker changes or after a pause of 1 second or more. - If speaker detection cannot be completed, the transcript is still returned with
diarization.statusset tounavailable, andutterances,speakersand the per-wordspeakerare left out. - Requests without
diarize, or withdiarize=false, get exactly the response they got before. A value other thantrueorfalseis ignored.diarizeis not available on live streaming or minutes of meetings.
Munsit MCP server
- NewThe Munsit MCP server lets AI apps use your Munsit account: find voices, generate speech, transcribe audio and check credits and usage, by asking in plain language. Add
https://mcp.munsit.com/mcpto Claude, Claude Code, ChatGPT or any app that supports MCP with OAuth, and sign in with your Munsit account — no API key needed. - Thirteen tools; only speech and transcription spend credits, at the same prices as the app. See MCP server for the permissions, limits and regional URLs.
Language mode for munsit-en-ar
- New
POST /api/v1/audio/transcribeaccepts an optionallanguagefield withmodel=munsit-en-ar.auto, the default, detects the language as the audio goes;arorensets the main language of the speech. Every value still transcribes mixed Arabic-English speech; setting the main language gives better accuracy when you know it. Requests without the field behave as before. WS /api/v1/listentakes the samelanguagequery parameter withmodel=munsit-en-ar. EachResultsevent reports thelanguageof that result: the detected one withauto, otherwise the one you set.- ImprovedCustom vocabulary (
hotwords) now works withmunsit-en-aronPOST /api/v1/audio/transcribe. Up to 20 words are used, counted across all entries. - APIOn
POST /api/v1/audio/transcribe, alanguagevalue other thanauto,ar,enorar-enis ignored andautois used. OnWS /api/v1/listen, an unsupported value closes the connection with error4002. Themunsitmodel is Arabic only: it ignores the field on batch requests and accepts onlyaron live sessions. - Live
munsit-en-arsessions do not applyhotwords. Every entry you send is reported inMetadata.dropped_hotwords.
Word timestamps for Text to Speech
- NewNew
POST /api/v1/text-to-speech/{model_id}/with-timestampsreturns speech together with character-level timings for the text you submitted. Use it to highlight words as they are spoken, drive captions, or align a transcript to the audio. - The response is NDJSON — one JSON object per line. Audio lines carry base64 PCM16 in
audio_base64; the final lines carryalignment(aligned to your original text) andnormalized_alignment(aligned to the engine's normalized form). Concatenatingcharactersreproduces your input exactly, so array indices map straight back onto your string. - Break tags are supported in the text you send:
<break time="3s"/>and<break time="500ms"/>insert a pause of the given length. - APITimings are emitted once generation finishes, not incrementally with each audio chunk. To highlight from the first word, buffer the full response before playback or request long text sentence by sentence.
- Requires a model served by the v1.5 engine. Other models return
400withWord timestamps are not available for model '…'. Pricing, wallet deduction, and history are identical to a standard synthesis request — timestamps cost nothing extra.
Streaming finalization on the legacy STT socket
- NewSend
{"event":"end_of_stream"}to transcribe any remaining buffered audio, at any duration. The server replies with atranscriptioncarryingisFinal: true, then afinalizedevent with the complete transcript. Closing the socket does not flush buffered audio — send this first. transcriptionevents now carry anisFinalboolean. Clients reading onlydataare unaffected.- New
min_buffer_secondsquery parameter controls how much audio accumulates before an interim result is emitted. Default0.5, range0.1–5.0. - FixedUtterances shorter than roughly 1.7 seconds could complete with no transcript and no error. The audio buffer was measured before each incoming chunk was appended, and the first chunk was never evaluated, so short answers never reached the recognizer.
- Rapid audio frames arriving during connection setup could each start their own initialization, discarding the WAV metadata from the first frame. Every subsequent frame then failed with
Not a valid WAV file (missing RIFF header)for the life of the connection. Clients sending 10–20 ms frames were affected on every session. - Deprecated
WS /websocket/speech-to-textremains deprecated and receives no new recognition features. New integrations should useWS /api/v1/listen.
New capabilities
- NewNew Text-to-Speech capabilities.
- New Speech-to-Text capabilities.
Working with an AI assistant? Every page is available as Markdown: add .md to the URL, or send an Accept: text/markdown header. For the whole documentation in one request, point it at llms-full.txt; the page index is llms.txt. Or use Copy Page, top right.