Kurdish TTS & STT API

A simple HTTP API for Kurdish text-to-speech and speech-to-text. Text-to-speech covers Sorani (Central Kurdish), Kurmanji (Northern Kurdish) and Badini (Behdinî — Northern Kurdish in the Arabic script, with six dedicated voices); speech-to-text covers Sorani and Kurmanji. Send text, get natural Kurdish speech from hundreds of voices; or send audio, get an accurate transcript — by file upload or live streaming. There is a free tier, and paid plans start at $12/month, or pay yearly and save 20%.

Machine-readable spec: OpenAPI 3.1 (/openapi.json). Get an API key in Settings → API. See pricing & plans. New here? Start with the step-by-step guides.

Authentication

All endpoints except GET /api/get-speakers require an API key in the x-api-key request header. Base URL: https://www.kurdishtts.com.

TTS and STT use separate keys. A TTS key authenticates the text-to-speech endpoint; an STT key authenticates the speech-to-text endpoints. They are not interchangeable — a TTS key returns 401 against an STT endpoint. Generate both in Settings → API. Keep keys server-side; never ship them in client code.

Text-to-Speech — POST /api/tts-proxy

There is one TTS model: V5. It renders every request. model_version does not pick a model or an engine — it picks which speaker group your speaker_id is looked up in: v3 (198 speakers), v4 (664), or v5 (20 curated, the six badini_ speakers, plus the 13 Cast and Studio speakers on plans bought for the API). The groups share no IDs, so send the one your speaker belongs to. Omitting the field selects v3, which is why every example below states it explicitly. The one exception is a badini_ id: it exists in no other group, so it resolves on v5 whatever you send. Converts Kurdish text to speech and returns audio/wav by default, or JSON with base64 audio and word-level timestamps when include_timestamps is true. The dialect is derived from the speaker_id prefix (sorani_… / kurmanji_… / badini_…).

FieldTypeRequiredDescription
textstringYesText to synthesize. Max 500 chars (free) / 5000 (paid).
speaker_idstringYesBuilt-in id from /api/get-speakers, or an owned clone_ id from the Voice Cloning page. Clone API use requires a paid subscription.
dialect"sorani" | "kurmanji" | "badini"NoReading language. Inferred from the speaker_id prefix when it has one. Send it for the Cast and Studio speakers: they read Sorani or Kurmanji and otherwise default to the Sorani of their reference clip. Badini (Behdinî) is Northern Kurdish in the Perso-Arabic script and has its own six badini_ speakers — a badini_ id resolves on v5 whatever model_version you send, and pairing dialect 'badini' with any other voice is refused with a 400.
model_version"v3" | "v4" | "v5"NoSpeaker group your speaker_id is resolved in — not a model or engine choice; every value renders on V5. v3 = 198 speakers, v4 = 664, v5 = 20 curated plus the six badini_ speakers and the 13 Cast and Studio ones on plans bought for the API. Defaults to v3 when omitted, and the groups share no IDs, so send it explicitly.
include_timestampsbooleanNoDefault false. true → JSON with base64 audio + word timestamps.
format"wav" | "opus" | "mp3"NoDefault wav. opus = Ogg/Opus at 24 kHz, ~10–15× smaller — ideal for mobile data. mp3 is also 24 kHz. Ignored when include_timestamps is true.
speednumberNo0.25–4.0, higher = faster (industry convention; inverted internally).
temperature / stabilitynumberNoOptional generation controls. Mutually exclusive; omit both for the engine default.
seedintegerNoOptional reproducibility control.
pitch, top_p, repetition_penalty…numberNoOptional advanced and post-processing controls; forwarded to V5 for compatibility.

Example — cURL

curl -X POST https://www.kurdishtts.com/api/tts-proxy \
  -H "x-api-key: YOUR_TTS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"text":"سڵاو، چۆنیت؟","speaker_id":"sorani_1","model_version":"v4"}' \
  --output speech.wav

Example — Python

import requests

resp = requests.post(
    "https://www.kurdishtts.com/api/tts-proxy",
    headers={"x-api-key": "YOUR_TTS_API_KEY"},
    json={"text": "سڵاو، چۆنیت؟", "speaker_id": "sorani_1", "model_version": "v4"},
)
with open("speech.wav", "wb") as f:
    f.write(resp.content)

Example — JavaScript

const res = await fetch("https://www.kurdishtts.com/api/tts-proxy", {
  method: "POST",
  headers: {
    "x-api-key": "YOUR_TTS_API_KEY",
    "Content-Type": "application/json",
  },
  body: JSON.stringify({ text: "سڵاو، چۆنیت؟", speaker_id: "sorani_1", model_version: "v4" }),
});
const audio = await res.arrayBuffer(); // audio/wav

Example: owned cloned voice

curl -X POST https://www.kurdishtts.com/api/tts-proxy   -H "x-api-key: YOUR_TTS_API_KEY"   -H "Content-Type: application/json"   -d '{"text":"سڵاو، چۆنیت؟","speaker_id":"clone_YOUR_CLONE_ID","dialect":"sorani"}'   --output cloned-voice.wav

Good to know

  • speed follows the industry convention: higher = faster.
  • format: "opus" returns audio/ogg (Opus, 24 kHz) — the response is ~10–15× smaller than WAV. format only applies to binary responses; the JSON/timestamps response stays base64 PCM.
  • v3 and v4 are permanent speaker-group selectors, not a deprecated path. Existing speaker IDs and plan entitlements do not change, and will not.
  • Any paid API plan may select v5 and render its 20 curated speakers and all six badini_ ones. The 13 Cast and Studio speakers are listed to everyone for discovery but render only on plans bought for the API — Developer, Pro, Business, and the legacy starter, starter-v4 and api-pro. Naming one on another plan returns 403 with upgrade_required.
  • What a plan sees inside a group is not uniform. Free gets four speakers each from v3 and v4, plus badini_story_m and badini_narrator_f on v5 — its only access to that group; the legacy plans are deliberately lopsided (starter has every v3 speaker but four from v4; starter-v4 is the reverse). Settings → API lists exactly what your key can call.
  • Creator, Developer, Pro, and Business may send an owned clone_… ID. Free accounts cannot use cloned voices through the API.
  • Clone allowances are Creator 3, Developer 10, Pro 20, and Business 30. Free includes one clone for the fixed website preview only.

Text-to-Speech (streaming) — POST /api/tts-stream

Streams speech while it is being generated instead of waiting for the full file — for short sentences the first audio typically arrives in about a second. Same key, limits, and billing as /api/tts-proxy (characters are debited when the stream starts). SSE and PCM are progressive V5 streams. WAV remains accepted for backward compatibility but is buffered before the response begins.

FieldTypeRequiredDescription
textstringYesText to synthesize. Max 500 chars (free) / 5000 (paid). Send sentence-by-sentence for the fastest first chunk.
speaker_idstringYesBuilt-in id from /api/get-speakers, or an owned clone_ id on a paid subscription.
stream_format"sse" | "pcm" | "wav"NoDefault sse (JSON events, base64 audio). pcm = progressive headerless 16-bit little-endian mono PCM at 24 kHz. wav = buffered compatibility response.
model_version"v3" | "v4" | "v5"NoSpeaker group your speaker_id is resolved in, exactly as on /api/tts-proxy. Same values, same v3 default, same entitlements. Not a model or engine choice.
speednumberNo0.25–4.0, higher = faster (same convention as /api/tts-proxy).
include_timestampsbooleanNoWith SSE, emits speech.word_timestamps events and includes the cumulative list in speech.audio.done.

Example — cURL (raw PCM)

curl -N -X POST https://www.kurdishtts.com/api/tts-stream \
  -H "x-api-key: YOUR_TTS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"text":"Silav, tu çawa yî?","speaker_id":"kurmanji_236","model_version":"v4","stream_format":"pcm"}' \
  --output speech.pcm
# play it: ffplay -f s16le -ar 24000 -ch_layout mono speech.pcm

Example — JavaScript (raw PCM)

const res = await fetch("https://www.kurdishtts.com/api/tts-stream", {
  method: "POST",
  headers: {
    "x-api-key": "YOUR_TTS_API_KEY",
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    text: "Silav, tu çawa yî?",
    speaker_id: "kurmanji_236",
    model_version: "v4", // speaker group the id lives in — see "List voices"
    stream_format: "pcm", // headerless 16-bit LE mono PCM @ 24000 Hz
  }),
});
const reader = res.body.getReader();
while (true) {
  const { done, value } = await reader.read();
  if (done) break; // normal end of the HTTP response = end of speech
  playPcmChunk(value); // Uint8Array of raw samples — feed your audio buffer
}

Example — JavaScript (SSE)

const res = await fetch("https://www.kurdishtts.com/api/tts-stream", {
  method: "POST",
  headers: {
    "x-api-key": "YOUR_TTS_API_KEY",
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    text: "سڵاو، چۆنیت؟",
    speaker_id: "sorani_1",
    model_version: "v4", // speaker group the id lives in — see "List voices"
    stream_format: "sse",
  }),
});
const reader = res.body.pipeThrough(new TextDecoderStream()).getReader();
let buf = "";
while (true) {
  const { done, value } = await reader.read();
  if (done) break;
  buf += value;
  const events = buf.split("\n\n");
  buf = events.pop() ?? ""; // keep the trailing partial event
  for (const line of events) {
    if (!line.startsWith("data: ")) continue;
    const evt = JSON.parse(line.slice(6));
    if (evt.type === "speech.audio.delta") playBase64Pcm(evt.audio); // 16-bit PCM @ 24000 Hz
    if (evt.type === "speech.word_timestamps") updateCaptions(evt.words);
    if (evt.type === "error") throw new Error(evt.error);
  }
}

Good to know

  • PCM chunks after the first are about 2 seconds of audio (96,000 bytes at the default 24 kHz rate) — do not hard-code the chunk size; read until the response ends. The stream ends with the normal end of the HTTP response — there is no in-band end marker.
  • A connection that drops mid-stream means generation failed — discard the audio and retry. In sse mode you also get an in-band error event.
  • Response headers describe the audio (X-Sample-Rate: 24000, X-Channels: 1, X-Bit-Depth: 16) and echo the selected catalog in X-Model-Version.

Speech-to-Text (file upload) — POST /api/stt-proxy

Transcribes an uploaded audio file (WAV/MP3/FLAC/OGG/M4A) sent as multipart/form-data. One credit is debited per successful transcription. Max file size and transcript length depend on your plan (free: 10 MB / 500 chars; starter: 50 MB / unlimited; pro: 100 MB / unlimited).

FieldTypeRequiredDescription
filefileYesAudio file (WAV/MP3/FLAC/OGG/M4A).
dialect"sorani" | "kurmanji"YesKurdish dialect of the audio.

Example — cURL

curl -X POST https://www.kurdishtts.com/api/stt-proxy \
  -H "x-api-key: YOUR_STT_API_KEY" \
  -F "file=@audio.wav" \
  -F "dialect=sorani"

Example — Python

import requests

resp = requests.post(
    "https://www.kurdishtts.com/api/stt-proxy",
    headers={"x-api-key": "YOUR_STT_API_KEY"},
    files={"file": open("audio.wav", "rb")},
    data={"dialect": "sorani"},
)
print(resp.json()["text"])

Response JSON includes text, detected_dialect, detected_script, and language. On the free plan a long transcript may be clipped — indicated by truncated: true and truncation_limit.

Speech-to-Text (live streaming) — POST /api/stt-stream-connect

Real-time transcription over a WebSocket, in two steps:

  1. POST /api/stt-stream-connect with your STT key and { "dialect": "sorani" } → returns a temporary websocket_url (connect within 5 minutes; it does not carry your key).
  2. Open the WebSocket and stream raw 16-bit PCM, mono, 16 kHz audio as binary frames. Send { "type": "control", "event": "finalize" } to flush. The server streams { "text": "…", "is_final": bool } messages and { "type": "control", "event": "done" } when complete.

One streaming session is debited per connect. Session limits and max duration depend on your plan (free: 20 sessions / 2 min; starter: 100 / 10 min; pro: 500 / 30 min).

Example — JavaScript (browser)

const API_BASE = "https://www.kurdishtts.com";
const API_KEY = "YOUR_STT_API_KEY";
let ws;

async function connect(dialect) {
  const res = await fetch(API_BASE + "/api/stt-stream-connect", {
    method: "POST",
    headers: { "x-api-key": API_KEY, "Content-Type": "application/json" },
    body: JSON.stringify({ dialect }),
  });
  if (!res.ok) throw new Error((await res.json()).detail || "Failed to connect");

  const data = await res.json();
  console.log("Sessions remaining:", data.streaming_sessions_remaining);

  ws = new WebSocket(data.websocket_url); // temporary URL, no key inside
  ws.onopen = () => capture();
  ws.onmessage = (event) => {
    const msg = JSON.parse(event.data);
    if (msg.type === "control" && msg.event === "done") return;
    if (msg.text) console.log(msg.is_final ? "Final:" : "Partial:", msg.text);
  };
}

async function capture() {
  const stream = await navigator.mediaDevices.getUserMedia({ audio: true });
  const ctx = new AudioContext({ sampleRate: 16000 }); // 16 kHz required
  const source = ctx.createMediaStreamSource(stream);
  const processor = ctx.createScriptProcessor(4096, 1, 1);
  processor.onaudioprocess = (e) => {
    const input = e.inputBuffer.getChannelData(0);
    const pcm16 = new Int16Array(input.length); // 16-bit mono PCM
    for (let i = 0; i < input.length; i++) {
      pcm16[i] = Math.max(-32768, Math.min(32767, input[i] * 32768));
    }
    if (ws && ws.readyState === WebSocket.OPEN) ws.send(pcm16.buffer);
  };
  source.connect(processor);
  processor.connect(ctx.destination);
}

// Call when the speaker is done:
function finalize() {
  if (ws) ws.send(JSON.stringify({ type: "control", event: "finalize" }));
}

connect("sorani");

Example — Python

import asyncio, json
import requests, websockets
import numpy as np
import sounddevice as sd

API_BASE = "https://www.kurdishtts.com"
API_KEY = "YOUR_STT_API_KEY"

async def stream_stt(dialect="sorani"):
    resp = requests.post(
        API_BASE + "/api/stt-stream-connect",
        headers={"x-api-key": API_KEY, "Content-Type": "application/json"},
        json={"dialect": dialect},
    )
    resp.raise_for_status()
    data = resp.json()
    print("Sessions remaining:", data["streaming_sessions_remaining"])

    async with websockets.connect(data["websocket_url"]) as ws:
        async def receive():
            async for message in ws:
                msg = json.loads(message)
                if msg.get("type") == "control" and msg.get("event") == "done":
                    return
                if "text" in msg:
                    print("Final:" if msg.get("is_final") else "Partial:", msg["text"])

        receiver = asyncio.create_task(receive())

        def callback(indata, frames, time, status):
            pcm16 = (indata[:, 0] * 32767).astype(np.int16)  # 16-bit mono
            asyncio.run_coroutine_threadsafe(ws.send(pcm16.tobytes()), asyncio.get_event_loop())

        with sd.InputStream(samplerate=16000, channels=1, callback=callback, blocksize=4096):
            await asyncio.sleep(30)  # record for 30s

        await ws.send(json.dumps({"type": "control", "event": "finalize"}))
        await receiver

asyncio.run(stream_stt("sorani"))

List voices — GET /api/get-speakers

A public, unauthenticated listing of the speaker groups. Use each returned id as the speaker_id, and send the matching model_version with it — the groups share no IDs. Pass ?model_version=v3 (198 speakers), v4 (664), or v5(the 20 curated speakers, the six badini_ ones, plus the 13 Cast and Studio ones).

curl "https://www.kurdishtts.com/api/get-speakers?model_version=v4"

Each speaker has id, name, dialect (sorani/kurmanji/badini) and gender. Entries in the v5 group carry two more: origin, which is v5 for a curated speaker, badini for one of the six Badini narrators, and cast or designed for a Cast or Studio one, and moods, the emotions that speaker was recorded with.

This endpoint lists a group in full and cannot see your key, so it is a catalog rather than an entitlement. What your plan may actually render is on the Settings → API page. A Cast or Studio speaker appears twice in the v5 listing, once under Sorani and once under Kurmanji, because it reads both. The six badini_ speakers appear once, under Badini only: they are the only voices auditioned for it, and pairing dialect "badini" with any other speaker returns 400.

MCP server — for AI agents

AI agents can call Kurdish TTS/STT natively over the Model Context Protocol. The MCP endpoint is https://www.kurdishtts.com/api/mcp (Streamable HTTP transport) and exposes six tools: synthesize_speech, transcribe_audio, start_streaming_transcription, list_voices, list_dialects and get_plan.

list_voices, list_dialects and get_plan need no API key, so an agent can connect and browse the voice catalog and the plan ladder before anyone signs up. Call get_plan first whenever a tool returns 401 or 403 — it reports what the current key may use and why the call was refused.

The three billed tools authenticate with Authorization: Bearer <api-key> — a TTS or STT key from Settings → API, or both joined as <ttsKey>:<sttKey>. Every tool call is validated, metered and billed against your plan exactly like the HTTP endpoints above — the MCP server is a thin wrapper, not a separate quota.

MCP payload limits: 4,000 characters per synthesis in the default mp3 container — 600 with format: "wav", and 500 on free plans — and 3MB of decoded audio per transcription, which is about 90 seconds of 16kHz WAV but roughly 25 minutes of 64kbps MP3, so send compressed audio. For longer texts or larger files call the HTTP API directly. Example client config:

{
  "mcpServers": {
    "kurdish-tts": {
      "url": "https://www.kurdishtts.com/api/mcp",
      "headers": {
        "Authorization": "Bearer YOUR_TTS_KEY:YOUR_STT_KEY"
      }
    }
  }
}

Step-by-step setup for Claude Code, Claude Desktop and Cursor: Add Kurdish voice to Claude.

Errors & limits

Successful responses use 200. Common error statuses:

  • 400 — bad request: a missing field, an invalid dialect, an invalid speed, a file too large, an emotion (moods are not on the API yet), a speaker_id that is not in the v5 speaker group, or a voice paired with a dialect it cannot read (only the six badini_ speakers read Badini).
  • 401 — missing or invalid API key (check you are using the right key space).
  • 403 — plan inactive or expired, characters exhausted, a speaker outside your plan's slice of the group you selected, a Cast or Studio speaker on a plan not bought for the API, or a reserved speaker. Body includes error and often upgrade_required and upgrade_url.
  • 422 — the engine rejected a generation parameter (e.g. temperature: 0.0); the detail array carries specifics.
  • 503 — transient. The speaker roster could not be reached, so the request was refused rather than validated against a stale list. Retry.

Two that surprise people

  • An unknown model_version is a 403, not a 422, and the message names your plan. The check fails closed on any value it does not recognise, so a typo reads like an entitlement problem. Send v3, v4 or v5.
  • A speaker that is not in the group you selected returns 400 under v5 but 403 under v3 and v4 — the same mistake, two statuses. Handle both if you validate IDs client-side.

Plans, quotas and prices are on the pricing page. The full machine-readable contract is at /openapi.json.