Base URL: https://api.stt.multispeaker.ganas.ai. Generate a client key in the admin dashboard. Test uploads in the playground.
Upload up to 60 seconds and 25 MiB. Nemotron 3 identifies up to eight speakers; the existing STT recognizer transcribes each detected turn.
curl --fail-with-body https://api.stt.multispeaker.ganas.ai/v1/audio/diarizations \ -H "Authorization: Bearer $STT_API_KEY" -F 'file=@meeting.wav'
Response: {"speakers":["speaker_1"],"segments":[{"start":0.5,"end":3.2,"speaker":"speaker_1","text":"...","is_low_confidence":false}],"text":"..."}
POST /v1/audio/transcriptions handles plain transcription up to 30 seconds. WSS /v1/audio/transcriptions/ws handles live PCM transcription without speaker labels. WSS /ws/telephony supports mu-law telephony streams. Details and limits are in the complete API guide.